Cooperative defense method, device and equipment based on endogenous threat intelligence knowledge graph
By constructing an endogenous threat intelligence knowledge graph, extracting knowledge triples using security situation data, determining attack paths and formulating defense strategies, the problem of lack of specificity in existing power Internet of Things defense methods is solved, and more efficient data utilization and precise defense measures are achieved.
Patent Information
- Application Number
- CN202510840840.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-16
AI Technical Summary
Existing power Internet of Things security defense methods rely on third-party threat intelligence and lack specificity, resulting in insufficient defense effectiveness and making it difficult to ensure the safe and stable operation of the power network.
By acquiring security situation data, extracting knowledge triples using preset security models, and constructing an endogenous threat intelligence knowledge graph, we can determine attack paths and formulate defense strategies, thereby improving the queryability and analyzability of data, and enhancing the comprehensiveness of attack path identification and the accuracy of defense.
It ensures the targeted nature of data collection, improves the queryability and analyzability of data, improves the comprehensiveness of attack path identification and the accuracy of defense, and enhances the defense capabilities of the power Internet of Things.
Smart Images

Figure CN120658470A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of risk defense technology, and in particular to a collaborative defense method, device and equipment based on an endogenous threat intelligence knowledge graph. Background Art
[0002] As the digital transformation of new power systems accelerates, power networks are increasingly incorporating distributed, cross-regional IoT devices, significantly enhancing system flexibility through wide-area wireless access technologies. However, this transformation also presents new security challenges, such as reduced physical security controllability and blurred network boundaries. While knowledge graph technology offers new approaches to IoT security, existing approaches still rely heavily on directly procured third-party threat intelligence. This intelligence often lacks specificity and struggles to accurately match the actual needs of power companies, hindering the effectiveness of knowledge graphs in providing effective defense information. Therefore, in the context of the power IoT, improving the defense capabilities of the IoT and ensuring the safe and stable operation of power networks has become a key technical issue that urgently needs to be addressed in the current field of risk defense technology. Summary of the Invention
[0003] The present invention provides a collaborative defense method, device and equipment based on an endogenous threat intelligence knowledge graph. The present invention ensures the targeted data collection, improves the queryability and analyzability of the data, and improves the comprehensiveness of attack path identification and the accuracy of defense.
[0004] One aspect of an embodiment of the present invention provides a collaborative defense method based on an endogenous threat intelligence knowledge graph, including:
[0005] Obtain security situation data from at least one data source;
[0006] Extracting knowledge triples from security situation data based on a preset security model;
[0007] Construct an endogenous threat intelligence knowledge graph based on knowledge triples;
[0008] An attack path is determined based on the endogenous threat intelligence knowledge graph, and a defense strategy for the attack path is determined.
[0009] Another aspect of an embodiment of the present invention provides a collaborative defense device based on an endogenous threat intelligence knowledge graph, including:
[0010] A data acquisition module, configured to obtain security situation data from at least one data source;
[0011] An information extraction module, used to extract knowledge triples from security situation data based on a preset security model;
[0012] The knowledge graph module is used to construct an endogenous threat intelligence knowledge graph based on knowledge triples;
[0013] A defense module is used to determine an attack path based on the endogenous threat intelligence knowledge graph and determine a defense strategy for the attack path.
[0014] Another aspect of an embodiment of the present invention provides a device, including:
[0015] at least one processor;
[0016] and a memory communicatively coupled to the at least one processor;
[0017] In which, the memory stores a computer program that can be executed by at least one processor, and the computer program is executed by at least one processor so that the at least one processor can execute the collaborative defense method based on the endogenous threat intelligence knowledge graph of any embodiment of the present invention.
[0018] In an embodiment of the present invention, a series of security situation data reflecting the network security status is obtained from at least one data source, a preset security model is obtained, knowledge triples are extracted from the security situation data using the obtained preset security model, an endogenous threat intelligence knowledge graph is constructed based on the obtained knowledge triples, an attack path is determined in the endogenous threat intelligence knowledge graph, and a defense strategy for preventing network attacks is obtained based on the attack path. In an embodiment of the present invention, by constructing a knowledge graph based on existing security situation data, the targeted nature of data collection is ensured. By constructing a knowledge graph using knowledge triples extracted from security situation data, complex data can be converted into structured knowledge, thereby improving the queryability and analyzability of the data. By determining the attack path based on the knowledge graph, the comprehensiveness of attack path identification and the accuracy of defense are improved.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 This is a collaborative defense method based on endogenous threat intelligence knowledge graph provided according to embodiment 1 of the present invention;
[0022] Figure 2 This is a flow chart of another collaborative defense method based on endogenous threat intelligence knowledge graph provided in accordance with the second embodiment of the present invention;
[0023] Figure 3 This is a flow chart of another collaborative defense method based on endogenous threat intelligence knowledge graph provided in accordance with the third embodiment of the present invention;
[0024] Figure 4 This is a structural diagram of a bidirectional Chinese BERT conditional random field model provided according to Example 3 of the present invention;
[0025] Figure 5 A comparison chart of intelligence processing rates provided according to Example 3 of the present invention;
[0026] Figure 6 A schematic diagram of the structure of a collaborative defense device based on an endogenous threat intelligence knowledge graph according to a third embodiment of the present invention;
[0027] Figure 7 2. It is a schematic structural diagram of a collaborative defense device based on an endogenous threat intelligence knowledge graph according to a fourth embodiment of the present invention;
[0028] Figure 8 This is a block diagram of a device for executing a collaborative defense method based on an endogenous threat intelligence knowledge graph according to the fifth embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0031] Example 1
[0032] Figure 1 A flowchart of a collaborative defense method based on an endogenous threat intelligence knowledge graph is provided for the first embodiment of the present invention. The embodiment of the present invention is applicable to the situation of defending against potential attack intentions of the Internet of Things. The method can be executed by a collaborative defense device based on an endogenous threat intelligence knowledge graph. The collaborative defense device based on an endogenous threat intelligence knowledge graph can be implemented in the form of hardware and / or software. The collaborative defense device based on an endogenous threat intelligence knowledge graph can be configured in a device. Figure 1 As shown, the method includes:
[0033] S110: Obtain security situation data from at least one data source.
[0034] The data source refers to an information base for storing different types of security situation data. For example, the data source may include an asset information base, a security knowledge base, or a vulnerability intelligence base.
[0035] Security situation data can be understood as a series of information used to reflect the network security status. For example, security situation data can include: asset information data in the asset information library, security knowledge data in the security knowledge library, or vulnerability intelligence data in the vulnerability intelligence library.
[0036] Specifically, a series of security situation data reflecting the network security status is obtained from at least one data source by calling an application programming interface provided by the data source or by writing a structured query statement.
[0037] S120 . Extracting knowledge triples from the security situation data based on a preset security model.
[0038] The preset security model can be understood as a predefined algorithmic framework for constructing knowledge triples. For example, the preset security model may include an algorithmic framework based on machine learning or natural language processing. The preset security model can be used to extract data such as entities or relationships in the security situation data from the security situation data.
[0039] Knowledge triples can be understood as a type of structured data, which can include structured data such as subjects, relationships or attributes.
[0040] Specifically, a preset security model for constructing knowledge triples is obtained, and the security situation data can be used as input to the preset security model. The preset security model can extract information such as subjects, relationships or attributes contained in the security situation data based on the input security situation data, and can construct knowledge triples based on data such as subjects, relationships or attributes.
[0041] S130. Construct an endogenous threat intelligence knowledge graph based on knowledge triples.
[0042] Among them, the endogenous threat intelligence knowledge graph can be understood as a semantic network used to describe the relationship between entities. For example, the endogenous threat intelligence knowledge graph can include: nodes or edges, etc. The nodes in the endogenous threat intelligence knowledge graph can represent entities, and the edges represent the relationship between entities.
[0043] Specifically, a knowledge triple of security situation data is obtained, which includes a subject, a relationship, and an attribute. The subject in the obtained knowledge triple is used as a vertex, the relationship as an edge of the vertex, and the attribute as an attribute of the vertex. According to the vertices, the edges between the vertices, and the attributes of the vertices, an endogenous threat intelligence knowledge graph is constructed to describe the relationship between entities in the security situation data.
[0044] S140. Determine the attack path based on the endogenous threat intelligence knowledge graph and determine the defense strategy for the attack path.
[0045] Among them, the attack path refers to the trajectory from the starting node of the attack operation to the final target node. In the attack path, it may include the situation where the same starting node is connected to multiple final target nodes, or it may include the situation where multiple starting nodes are connected to one final target node. In the embodiment of the present invention, the above attack path is only used as an example.
[0046] A defense strategy refers to a series of measures taken to prevent network attacks. For example, a defense strategy may include: establishing a firewall strategy or a data encryption strategy.
[0047] Specifically, an endogenous threat intelligence knowledge graph is obtained, and the attack path between the starting node and the final target node of the attack operation is determined in the endogenous threat intelligence knowledge graph. Corresponding defense strategies for preventing network attacks can be formulated based on the attack path.
[0048] In an embodiment of the present invention, a series of security situation data reflecting the network security status is obtained from at least one data source, a preset security model is obtained, knowledge triples are extracted from the security situation data using the obtained preset security model, an endogenous threat intelligence knowledge graph is constructed based on the obtained knowledge triples, an attack path is determined in the endogenous threat intelligence knowledge graph, and a defense strategy for preventing network attacks is obtained based on the attack path. In an embodiment of the present invention, by constructing a knowledge graph based on existing security situation data, the targeted nature of data collection is ensured. By constructing a knowledge graph using knowledge triples extracted from security situation data, complex data can be converted into structured knowledge, thereby improving the queryability and analyzability of the data. By determining the attack path based on the knowledge graph, the comprehensiveness of attack path identification and the accuracy of defense are improved.
[0049] Furthermore, in an embodiment of the present invention, obtaining security situation data of at least one data source includes at least one of the following: collecting internal security data as security situation data; collecting first information data from an asset information library as security situation data; collecting second information data from a security knowledge library as security situation data; collecting third information data from a vulnerability intelligence library as security situation data; and collecting fourth information data from a basic resource library as security situation data.
[0050] Among them, internal security data refers to security-related data generated within an organization or enterprise. For example, internal security data may include: alarm information generated within the organization or data such as the name of the device that generates alarm behavior within the organization; internal security data can be used as security situation data to construct an endogenous threat intelligence knowledge graph.
[0051] An asset information repository is a database used to store all assets and asset-related attributes within an organization. For example, assets may include hardware devices or software systems, and asset-related attributes may include asset type, location, configuration, status, version, or owner information. All asset information and asset-related attribute information stored in the asset information repository is referred to as first information data.
[0052] The security knowledge base refers to a database for storing a series of security-related knowledge content. For example, the security-related knowledge content may include laws and regulations, company regulations, training materials, or emergency plans. The series of security-related knowledge content is used as the second information data.
[0053] A vulnerability intelligence database refers to a database containing detailed information about known vulnerabilities. For example, this information may include vulnerability intelligence data or threat intelligence data. Vulnerability intelligence focuses on Common Vulnerabilities and Exposures (CVE) numbers, vulnerability database numbers, middleware affected by the vulnerabilities, and versions affected by the vulnerabilities. Threat intelligence focuses on Internet Protocol (IP) intelligence, IP keywords in domain name intelligence, domain name keywords, threat tags, and advanced persistent threat (APT) groups. Vulnerability intelligence data or threat intelligence data are integrated to form third-party information data.
[0054] The basic resource library refers to a database storing basic underlying resource information. For example, the basic underlying resource information may include: IP address, subnet information, user name or network topology structure, etc. The basic underlying resource information is used as the fourth information data.
[0055] Specifically, internal security data related to security generated within the organization is collected to obtain an asset information library, a security knowledge base, a vulnerability intelligence library and a basic resource library. First information data, second information data, third information data and fourth information data can be extracted from the asset information library, the security knowledge base, the vulnerability intelligence library and the basic resource library respectively. The internal security data, third-party security data, first information data, second information data, third information data and fourth information data obtained above can be used separately as security situation data, or a collection of several types of data among the internal security data, third-party security data, first information data, second information data, third information data and fourth information data can be selected as security situation data.
[0056] Furthermore, in an embodiment of the present invention, the entity model of the preset security model includes alarms, attack sources, assets, business systems and threat intelligence, and the relationship model of the entity model includes attack sources triggering alarms, attack sources attacking assets, alarms occurring on assets, attack sources conforming to threat intelligence and assets belonging to business systems.
[0057] The entity model can be understood as a conceptual framework for describing various entities. For example, the entity model can include entity data such as alarms, attack sources, assets, business systems, and threat intelligence.
[0058] A relational model can be understood as a conceptual framework for describing various relationships between entities. For example, a relational model may include relational data such as the attack source triggering an alarm, the attack source attacking an asset, the alarm occurring on an asset, the attack source conforming to threat intelligence, and the asset belonging to a business system.
[0059] Specifically, the security situation data is input into a preset security model, and the preset security model is used to identify the entities and relationships contained in the security situation data. Entity data such as alarms, attack sources, assets, business systems, and threat intelligence in the entity model are extracted from the security situation data. Relational data such as attack sources triggering alarms, attack sources attacking assets, alarms occurring on assets, attack sources complying with threat intelligence, and assets belonging to business systems in the relational model are extracted from the security situation data.
[0060] Example 2
[0061] Figure 2 A flowchart of another collaborative defense method based on an endogenous threat intelligence knowledge graph is provided for the second embodiment of the present invention. The embodiment of the present invention is a refinement of the above embodiment. Specifically, it refines the specific steps of how to extract knowledge triples and how to determine the attack path.
[0062] like Figure 2 As shown in Figure 2, another collaborative defense method based on endogenous threat intelligence knowledge graph includes:
[0063] S210: Obtain security situation data from at least one data source.
[0064] S220: Calling the first large language model to extract entity data corresponding to the entity model in the preset security model from the security situation data.
[0065] Among them, the first large language model can be understood as an algorithm framework for extracting entity data in security situation data. For example, the first large language model can include: a general information extraction large model or a large language model artificial intelligence (LLaMA) model.
[0066] Entity data can be understood as a data set consisting of objectively existing objects with attribute information. For example, entity data may include entity data such as alarms, attack sources, assets, business systems, and threat intelligence.
[0067] Specifically, an entity model of a preset security model is obtained, in which entity data to be extracted is defined. Based on the obtained entity model, the first large language model is called to extract the entity data defined in the entity model from the security situation data.
[0068] For example, before calling the first large language model to extract entity data defined in the entity model in the preset security model in the security situation data, the following steps may also be included: preprocessing the security situation data, and the preprocessing operations may include: data deduplication operation, data denoising operation, data conversion operation, data enhancement operation or data standardization operation, etc.
[0069] S230: Call the second large language model to extract relationship data between different entity data in the security situation data according to the relationship model in the preset security model.
[0070] Among them, the second large language model can be understood as an algorithm framework for extracting relational data within security situation data. For example, the second large language model can include: a Transformer-based bidirectional encoder representation model or a neural network-based relation extraction model, etc.
[0071] Relational data can be understood as a series of data used to describe the relationship between entities. For example, relational data may include: the attack source triggering the alarm, the attack source attacking the asset, the alarm occurring on the asset, the attack source conforming to the threat intelligence, and the asset belonging to the business system.
[0072] Specifically, a relational model of a preset security model is obtained, in which relational data to be extracted is defined. Based on the obtained relational model, a second large language model is called to extract the relational data defined in the relational model from the security situation data.
[0073] S240 . For each entity data, extract the attribute information and entity identifier of the entity data, and extract the relationship identifier of the relationship data associated with the entity data, and save the entity identifier, relationship identifier, and attribute information as a knowledge triple.
[0074] Among them, attribute information can be understood as information used to describe the relevant attributes of entity data. It can be understood that each entity data can have corresponding attribute information. For example, if the entity data is software, the corresponding attribute information of the software may include software identification (ID), software name, software version number, software type and other information. If the entity data is a business system, the corresponding attribute information of the business system may include: business system ID, system name, system type and other information.
[0075] An entity identifier is a symbol used to uniquely distinguish or identify each entity in security situation data. The entity identifier can be used to accurately reference or associate a specific entity data.
[0076] A relationship identifier can be understood as a symbol used to mark the relationship between two entity data. Through the relationship identifier, the relationship between the entity data can be clearly expressed.
[0077] Specifically, for each entity data obtained, the attribute information corresponding to the entity data for describing the relevant attributes of the entity data and the entity identifier for uniquely identifying each entity data in the security situation data can be extracted from the security situation data. The relationship data corresponding to the entity data in the security situation data and the relationship identifier that can be used to uniquely mark the association between two entity data can be obtained from the security situation data. The extracted attribute information, entity identifier and relationship identifier can be combined in a prescribed format to form a knowledge triple.
[0078] S250: Obtain a pre-configured built-in analysis rule library, and extract key fields of all built-in analysis rules in the built-in analysis rule library.
[0079] The built-in analysis rule base can be understood as a collection of built-in analysis rules. These rules are standardized rules used for threat detection and / or security analysis. Each built-in analysis rule describes a specific threat pattern or security event. These rules can be stored in a structured or semi-structured format. The built-in analysis rule base can be configured based on common attack behaviors or the security requirements of your industry.
[0080] Key fields refer to the fields in each built-in analysis rule that are used to identify the core features of the rule. For example, the key fields can be the fields for various entities or identifying time in each built-in analysis rule. Optionally, the built-in analysis rule can be: if the same attack source attacks ≥3 different assets within 12 hours, and these assets belong to the same business system, the danger level is raised to medium, and a defense strategy is issued according to the medium alarm. At this time, the key fields in the built-in analysis rule can be attack source, 12 hours, assets, and business systems, and the attack source, assets, and business systems are all entities defined in the entity model. In the embodiment of the present invention, the above-mentioned built-in analysis rules are only used as examples to illustrate the process of extracting key fields in the built-in analysis rules.
[0081] Specifically, obtain the built-in analysis rule library configured by the user based on common attack behaviors or the security requirements of the industry, use structured query statements or input and output reading operations to extract all built-in analysis rules from the built-in analysis rule library, and use regular expressions or corresponding data parsing libraries to extract key fields from the built-in analysis rules.
[0082] S260. Traverse the target knowledge triples that match the key fields in the endogenous threat intelligence knowledge graph and use the target knowledge triples as attack paths.
[0083] Target knowledge triples are knowledge triples that match key fields within the endogenous threat intelligence knowledge graph. They represent events or activities related to the threats or security described by built-in analysis rules. Target knowledge triples can serve as potential attack paths.
[0084] Specifically, the system obtains key fields from built-in analysis rules and, based on these key fields, extracts all entities, relationships between entities, and attributes within the endogenous threat intelligence knowledge graph that are identical or similar to the key fields. These extracted entities, relationships between entities, and attributes are then used to construct a target knowledge triple, which is then used as the attack path. Optionally, the target knowledge triple may contain connections between the same entity and multiple different entities.
[0085] S270: Search the built-in analysis rule library for a matching attack path, and use the built-in analysis rule as a defense strategy.
[0086] Among them, defense strategy can be understood as the countermeasures that can be taken against the attack path to prevent potential threats or attacks in the network.
[0087] Specifically, through character matching or vector similarity calculation, the built-in analysis rule that is most similar to the attack path is searched in the built-in analysis rule library, and the built-in analysis rule is used as the defense strategy for the attack path.
[0088] In an embodiment of the present invention, security situation data of at least one data source is obtained, entity data corresponding to an entity model in a preset security model is extracted from the security situation data by calling a first large language model, relationship data between different entity data is extracted from the security situation data according to a relationship model in the preset security model by calling a second large language model, attribute information and entity identifiers of the entity data are extracted from each entity data, and relationship identifiers are extracted from relationship data associated with the entity data, the entity identifiers, relationship identifiers and attribute information are saved as knowledge triples, a pre-configured built-in analysis rule base is obtained, key fields of all built-in analysis rules in the built-in analysis rule base are extracted, knowledge triples matching the key fields in the endogenous threat intelligence knowledge graph are used as target knowledge triples, the target knowledge triples can be used as attack paths, built-in analysis rules matching the attack paths are searched for in the built-in analysis rule base, and the built-in analysis rules are used as defense strategies. The embodiments of the present invention can automatically extract entity data and relationship data by utilizing a large language model, thereby improving the efficiency of data processing, constructing an endogenous threat intelligence knowledge graph through the extracted entity, attribute and relationship information, determining the target knowledge triples by traversing the endogenous threat intelligence knowledge graph, and accurately identifying the attack path by mapping key fields with nodes in the endogenous threat intelligence knowledge graph. Based on the accurately determined attack path, targeted countermeasures can be provided in the built-in analysis rule library, thereby improving the targetedness and accuracy of defense.
[0089] Optionally, an embodiment of the present invention refines the first large language model and the second large language model. Specifically: the first large language model includes a word matching input layer, a feature representation layer, an attention encoding layer, and a conditional random field decoding layer. The word matching input layer performs character matching based on a bidirectional maximum matching method; the second large language model includes a bidirectional encoder representation model based on Transformer.
[0090] Specifically, the first large language model can be called to extract entity data in the security situation data, wherein the first large language model can be composed of components such as a word matching input layer, a feature representation layer, an attention encoding layer, and a conditional random field decoding layer. In the word matching input layer, characters or words can be matched based on the bidirectional maximum matching method, which can improve the accuracy of word segmentation. The Transformer-based bidirectional encoder representation model can be used as the second large language model to extract relationship data between different entity data in the security situation data.
[0091] Optionally, the embodiment of the present invention further refines the attention encoding layer of the first large language model. Specifically, the attention score calculation formula of the attention encoding layer of the first large language model includes: Among them, R t-jIt is a relative position code, i represents the number of characters, d k is the dimension of the key vector, t represents the index of the target token, and j represents the index of the context token.
[0092] Among them, the attention score calculation formula can be understood as a function used to calculate the attention weight.
[0093] Relative position encoding can be understood as a calculation formula used to provide position information to the model, and position encoding can be generated based on the relative distance between markers.
[0094] Specifically, in the attention encoding layer of the first large language model, the existing absolute position encoding can be replaced by the relative position encoding R t-j ,in, i represents the number of characters, d k is the dimension of the key vector, t represents the index of the target tag, j represents the index of the context tag, and the attention encoding layer replaces the existing absolute position encoding with the relative position encoding R t-j , which enables the attention encoding layer to distinguish the direction information and distance information in the text information, and allows the model to pay more attention to words in the text information that may be named entities.
[0095] Optionally, the embodiment of the present invention further refines the built-in analysis rules. Specifically, the built-in analysis rules include at least one of the following: if the same attack source attacks assets greater than or equal to a first preset number threshold, and each of the assets belongs to the same business system, then the danger level of the attack path is determined to be a medium risk level, and the built-in analysis rule matching the medium risk level is searched in the built-in analysis rule library; if the same asset is attacked by different types of attacks from the same attack source, then the danger level of the attack path is determined to be a high risk level, and the built-in analysis rule matching the high risk level is searched in the built-in analysis rule library; if the same asset is attacked by attack sources greater than or equal to a second preset number threshold, and each of the attack sources belongs to the same attack type and the same threat intelligence, then the danger level of the attack path is determined to be a high risk level, and the built-in analysis rule matching the high risk level is searched in the built-in analysis rule library.
[0096] The first preset number threshold refers to the number of assets attacked by a single attack source within a specific time period. For example, the first preset number threshold can be 5 or 3, which means that within a certain time window, the same attack source attacked at least 5 or 3 assets.
[0097] The second preset threshold value refers to the number of times a single asset is attacked by different types of attacks from the same attack source within a certain time interval. For example, the first threshold value can be 4 or 3, indicating that the same asset is attacked by at least four or three types of attacks from the same attack source within a certain time interval. Optionally, the attack types may include multi-port scanning or password guessing attacks.
[0098] Specifically, the built-in analysis rule library includes at least the following built-in analysis rules: if the number of assets attacked by a single attack source is greater than or equal to a first preset number threshold, and the attacked assets belong to the same business system, the danger level of the attack path is determined to be a medium risk level, and the built-in analysis rule library that matches the medium risk level is searched for in the built-in analysis rule library; if the same asset is attacked by different types of attacks from the same attack source, the attack path is determined to be a high risk level, and the built-in analysis rule library that matches the high risk level is searched for in the built-in analysis rule library; if the same asset is attacked by attack sources greater than or equal to a second preset number threshold, and each of the attack sources belongs to the same attack type and the same threat intelligence, the danger level of the attack path is determined to be a high risk level, and the built-in analysis rule library that matches the high risk level is searched for in the built-in analysis rule library. Optionally, in some embodiments, the built-in analysis rules can be set as logical conditions within a preset time interval. For example, if, within the preset time interval, the number of assets attacked by a single attack source is greater than or equal to a first preset number threshold, and the attacked assets belong to the same business system, the danger level of the attack path is determined to be a medium-risk level; or if, within the preset time interval, the same asset is attacked by different types of attacks from the same attack source, the attack path is determined to be a high-risk level, and a built-in analysis rule matching the high-risk level is searched in the built-in analysis rule library.
[0099] Example 3
[0100] Figure 3 This is a flowchart of another collaborative defense method based on endogenous threat intelligence knowledge graph provided according to the third embodiment of the present invention; Figure 4 This is a structural diagram of a bidirectional Chinese BERT conditional random field model provided according to Example 3 of the present invention; Figure 5 A comparison chart of intelligence processing rates provided according to Example 3 of the present invention; Figure 6 A schematic diagram of the structure of a collaborative defense device based on an endogenous threat intelligence knowledge graph is provided according to a third embodiment of the present invention. This embodiment of the present invention refines the above embodiment and specifically details how to obtain security situation data, how to pre-process the security situation data, how to construct an endogenous threat intelligence knowledge graph, and how to issue a defense strategy based on the endogenous threat intelligence knowledge graph and built-in analysis rules.
[0101] like Figure 3As shown, another collaborative defense method based on endogenous threat intelligence knowledge graph may include the following specific steps:
[0102] S310: Collect security situation data, including internal alarm data and data from the four libraries of the full-scenario network security situation awareness platform S6000. The four libraries include: asset information library data, security knowledge library data, vulnerability intelligence library data, or basic resource library data.
[0103] In step S310, the basic requirements and standards for collecting data within the enterprise or company can be followed to collect security situation data, including internal alarm data and the four-library data of the State Grid Corporation's full-scene network security situation awareness platform S6000. The four libraries of the State Grid Corporation S6000 specifically refer to: Asset Information Library: Filter, complete, and integrate the company's asset data, and perform data quality monitoring on the data in the processing process to form the final asset information data; Security Knowledge Base: Centrally extract, govern, integrate, and model the State Grid Corporation's various laws and regulations, company specifications, report documents, training materials, network attacks, expert experience data, security equipment parameters, disposal strategies and other network security data, combined with the large language model technology in the vertical field of power system network security , which brings together the work experience accumulated in the company's network security professional work and extracts a network security knowledge base; Vulnerability Intelligence Library: Vulnerability intelligence and threat intelligence are sorted out separately from the perspective of data needs. Threat intelligence focuses on IP intelligence, IP keywords, domain name keywords, threat tags, and APT organizations in domain name intelligence. Vulnerability intelligence focuses on CVE numbers, security vulnerability library numbers, vulnerability-affected middleware, vulnerability-affected versions and other information, and finally builds a national network-level vulnerability intelligence library; Basic Resource Library: Centrally manages policy sets, sample sets, tool sets, and model sets that are common to the entire network and special to the provincial side, and provides network security practitioners with rich, practical, and timely updated equipment features and rules, malicious samples and traffic samples, red and blue team tools, and artificial intelligence models.
[0104] S320: Preprocess the data. The preprocessing may include data deduplication, data denoising, data conversion, data enhancement, or data standardization.
[0105] In step S320, data deduplication is to screen the collected raw data and remove duplicate data; data denoising is to clean the collected raw data, remove data with empty required fields, remove data with incorrect data formats, remove data with empty important fields, and remove data with inconsistent number of fields; data conversion is to support the conversion of collected raw data of the same type but different formats into a standardized database format; data enhancement is to enhance the collected raw data based on the four-library data of S6000, etc., and the enhanced content includes the relevant attributes of the assets, geographical location, intelligence hit information, etc.; data standardization refers to the data unification with reference to the threat intelligence data format standard.
[0106] S330. Use the Bi-Chinese-BERT-CRF model to extract information from the preprocessed data, including 5 types of entities and corresponding attribute information.
[0107] The construction of a known knowledge graph is represented as G = (E, R, S), where E and R are sets of entities and relationships, respectively. S is a knowledge base consisting of triples <entity, relationship, entity> and <concept, attribute, value>. In step S330, five types of entities and five types of relationships are extracted from the collected data, thereby standardizing the multi-source data into a unified triple format. The five types of entities and five types of relationships extracted from the collected data are shown in Table 1. Entity attribute information is shown in Tables 2, 3, 4, 5, and 6. The relationships between the five types of entities are shown in Table 7.
[0108] Table 1 Entity relationship extraction results
[0109]
[0110] Table 2 Alarm entity attribute information table
[0111]
[0112]
[0113] Table 3 Attack source entity attribute information table
[0114] property name illustrate id Attack source ID Unique identifier ip IP address IPv4 / IPv6 format network Network Such as "Malicious IP Pool-North America" geoLocation Geographical location Contains latitude and longitude coordinates firstActiveTime First active time ISO 8601 format lastActiveTime Latest active time ISO 8601 format threatLevel Threat Level Risk Score
[0115] Table 4 Asset entity attribute information table
[0116] property name illustrate id Asset ID Unique identifier name Asset Name ip Asset IP IPv4 / IPv6 format os operating system For example, "CentOS 7.9" port Open ports For example, ["80 / tcp","443 / tcp"] businessSystem Business system Such as "Payment System V3.2" department Department Such as "Cloud Computing Division" operator Operation and Maintenance Manager Name / Employee Number
[0117] Table 5 Business system entity attribute information table
[0118]
[0119] Table 6 Threat intelligence entity attribute information table
[0120]
[0121]
[0122] Table 7 Entity relationship table
[0123]
[0124] like Figure 4 As shown in the figure, the architecture of the Bi-Chinese-BERT-CRF model is mainly divided into four parts: word matching input layer, feature representation layer, feature encoding layer, and conditional random field (CRF) decoding layer.
[0125] Word matching input layer: Unlike English words that are naturally separated by spaces, Chinese words are usually composed of continuous Chinese characters, which makes word boundary recognition complicated, so it is easy to cause the words extracted from entities to be inaccurate. The embodiment of the present application adopts a two-way maximum matching method to match characters with S6000 four-library words. The two-way maximum matching method combines the advantages of the forward maximum matching method and the reverse maximum matching method, and selects the best word segmentation scheme by comparing the word segmentation results of the two. This method can effectively improve the accuracy and efficiency of word segmentation when processing Chinese word segmentation. At the same time, it combines characters with existing four-library words, especially the security knowledge base in the four-library, which covers a large amount of security field knowledge and can enhance the model's ability to capture contextual semantic information and boundary information. The input of the word matching input layer is: internal alarm data and the four-library data of the full-scene network security situation awareness platform S6000. The output is discrete words. Its function is to introduce four-library data. The four-library data covers a large amount of security field knowledge and can enhance the model's ability to capture contextual semantic information and boundary information.
[0126] The feature representation layer adopts a full-word coverage strategy, meaning that when masking, entire words are masked rather than partial characters. This better preserves the semantic integrity of Chinese characters, resulting in more complete word information. This improvement not only improves the model's understanding of Chinese text but also significantly enhances its performance in downstream tasks. The feature representation layer takes discrete words as input and outputs word vectors, which improves the model's understanding of Chinese text and significantly enhances its performance in downstream tasks.
[0127] The formula for calculating the attention score of the feature encoding layer is as follows:
[0128]
[0129]
[0130] Among them, Q, K, V refer to the query, key, and value vectors respectively, H is the input matrix, i represents the number of characters, and W q and W v is the weight matrix, d k is the dimension of the key vector; t represents the index of the target tag, j represents the index of the context tag, Q t and K j are the query vector and key vector of labels t and j respectively; u and v are both learnable parameters, R t-j It is a relative position encoding; is the attention weight, represents the attention score between two tokens, is the deviation of the t-th marker to some relative distance, is the deviation of the jth marker, It is the deviation term of specific distance and direction; softmax is a smooth function.
[0131] Formula 5 and Formula 6 can be obtained from Formula 2:
[0132]
[0133]
[0134] Since sin(-x) = -sin(x), cos(-x) = cos(x), for any offset t, the forward and backward relative position encodings are the same in cos(t), but opposite in sin(t). Therefore, the absolute position encoding in the existing Transformer is replaced by the relative position encoding R t-j Finally, the attention mechanism can distinguish the direction information and distance information in the text information, making the model pay more attention to the words in the sentence that may be named entities.
[0135] The input of the feature encoding layer is the word vector, and the output is the word vector that carries richer information, providing key context information for subsequent decoding. Its role is to pay more attention to the words in the sentence that may be named entities.
[0136] The feature decoding layer leverages global information and considers dependencies between labels, providing the most likely label for each word and helping to improve the accuracy of named entity recognition. The feature decoding layer takes as input a word vector that carries richer information and outputs an entity. Its purpose is to leverage global information and consider dependencies between labels, providing the most likely label for each word, which helps to improve the accuracy of named entity recognition.
[0137] S340. Connect entities through five types of relationships, construct an endogenous threat intelligence knowledge graph in the form of triples, and store it in the Neo4j database.
[0138] In step S340, the Neo4j database is a graph-structured NoSQL database with high performance, large-scale scalability, and high flexibility. While standardizing the final generated form of the system's intelligence data, it can also meet high-concurrency reading and writing requirements, greatly supporting the expansion of the endogenous threat intelligence knowledge graph, and can use the Cypher query language to run queries to retrieve data.
[0139] S350, based on the existing endogenous threat intelligence knowledge graph, uses built-in analysis rules to generate disposal suggestions, and issues defense strategies based on the disposal suggestions.
[0140] In step S350, the data of the endogenous threat intelligence knowledge graph can help network analysts generate treatment suggestions using built-in analysis rules when an attacker invades. The built-in analysis rules may include the following rules:
[0141] The first built-in analysis rule targets short-term cross-asset penetration by a single attack source. This rule recommends raising the risk level to medium and issuing a defense strategy based on a medium-risk alert. For example, if the attack source 40.77.167.243 launches multiple port scans against assets 10.0.0.1 (web service), 10.0.0.2 (database), and 10.0.0.3 (API gateway) within the same business system within 10 minutes, triggering three low-risk alerts. The recommendation is based on the following: The attacker attempted to attack a specific system, demonstrating organized and targeted attack behavior. Recommended Action: Raise the risk level to medium and issue a defense strategy based on a medium-risk alert. The default time window is 12 hours, and the default number threshold is 3.
[0142] The second built-in analysis rule targets multi-dimensional attacks on a single asset. This refers to attacks on the same asset from the same source within a preset time window, including multiple types of attacks such as multi-port scanning, password guessing, or system command execution. It recommends raising the risk level to high and initiating a defense strategy based on the high-risk alert. For example, asset 10.135.78.32 was first subjected to a multi-port scan by 40.77.167.243, followed by a password guessing attack in an attempt to obtain legitimate login credentials and execute system commands on the asset. The basis for generating the recommendation is: The attack chain is complete, indicating that the asset has been compromised. Recommended Action: It is recommended to raise the risk level to high and initiate a defense strategy based on the high-risk alert.
[0143] The third built-in analysis rule targets coordinated cluster attacks from multiple attack sources. This refers to the situation where the same asset is attacked by the same type of attack from at least the threshold number of different attack sources within a preset time window, and the attack source IP addresses match the same threat intelligence entity. The rule recommends raising the risk level to high and issuing a defense strategy based on a high-risk alert. For example, if attack sources 3.3.3.3, 4.4.4.4, 5.5.5.5, and 0.77.167.243 launch multiple attacks against asset 10.0.0.5 within 12 hours, triggering multiple low-risk alerts, and the attack source IP addresses match the same threat intelligence entity, this indicates possible coordination between these attack sources. Recommendation basis: Coordinated attacks from multiple attack sources of known active threats require priority response. Recommended action: Raise the risk level to high and issue a defense strategy based on a high-risk alert.
[0144] The above method can be deployed in a system that executes the method. The system can be deployed in the company for actual application to verify the defensive effect of the system: a set of comparative experiments between internal data and third-party security data are designed. The experimental environment of the system is shown in Table 8. The internal alarm data can include: boundary protection equipment attack samples collected by the unified probe of the company's large-scale security situation awareness platform, intelligence reported by analysts of provincial units, relevant reporting documents such as defender reports during network security actual combat exercises, etc. The data information is shown in Table 9.
[0145] Table 8 Experimental environment diagram
[0146]
[0147]
[0148] Table 9 Experimental data information table
[0149]
[0150] After data processing and entity extraction by this system, the amount of structured data currently available for knowledge graph construction is approximately 4,000,000, and the amount of new data during non-major security periods is approximately 10,000 per day. Observe the data processing rate comparison results based on internal data and third-party security data in the embodiment of this application, as shown in the following example: Figure 5 As shown, the data processing rate of this system is about 85%, and the third-party security data processing rate is about 44%, which proves the effectiveness of the defense method based on the knowledge graph of internal data. The knowledge graph generated based on internal data is closer to the use needs of the defender. Among them, the data processing rate refers to the data that is matched with actual hazards and processed after it is sent to the company branch, divided by the total number of data sent. It can be understood that the knowledge graph in the embodiment of the present invention is the same as the endogenous threat intelligence knowledge graph mentioned in the above embodiment.
[0151] This embodiment also provides a collaborative defense system based on endogenous threat intelligence knowledge graph, such as Figure 6 Shown, including:
[0152] The data collection module 410 is used to collect security situation data, including internal alarm data and the four database data of the full-scenario network security situation awareness platform S6000;
[0153] The data preprocessing module 420 is used to perform preprocessing operations on the data;
[0154] The information extraction module 430 is used to extract information from the preprocessed data using the bidirectional Chinese BERT conditional random field model, including five types of entities and their corresponding attribute information;
[0155] The knowledge graph module 440 is used to connect entities through five types of relationships, construct an endogenous threat intelligence knowledge graph in the form of triples, and store it in the Neo4j database;
[0156] The defense module 450 is used to generate disposal suggestions based on the existing endogenous threat intelligence knowledge graph using built-in analysis rules, and issue defense strategies based on the disposal suggestions.
[0157] Example 4
[0158] Figure 7 This is a schematic diagram of the structure of a collaborative defense device based on an endogenous threat intelligence knowledge graph provided by the fourth embodiment of the present invention. This embodiment is applicable to situations where it is impossible to accurately match the actual needs of power companies. The device can be implemented in software and / or hardware. The device can be integrated into any device that provides defense functions and can execute any of the collaborative defense methods based on the endogenous threat intelligence knowledge graph of the above embodiments. Figure 7 As shown, the collaborative defense device based on the endogenous threat intelligence knowledge graph specifically includes: a data acquisition module 510, an information extraction module 520, a knowledge graph module 530 and a defense module 540.
[0159] The data acquisition module 510 is used to obtain security situation data from at least one data source;
[0160] An information extraction module 520 is used to extract knowledge triples from security situation data based on a preset security model;
[0161] The knowledge graph module 530 is used to construct an endogenous threat intelligence knowledge graph based on knowledge triples;
[0162] The defense module 540 is used to determine the attack intention and attack path of the attack data based on the endogenous threat intelligence knowledge graph, and determine the defense strategy for the attack path intention.
[0163] Optionally, the data collection module 510 is specifically used to collect internal security data as security situation data; collect first information data from the asset information library as security situation data; collect second information data from the security knowledge library as security situation data; collect third information data from the vulnerability intelligence library as security situation data; and collect fourth information data from the basic resource library as security situation data.
[0164] Optionally, in an embodiment of the present invention, the entity model of the preset security model includes alarms, attack sources, assets, business systems and threat intelligence, and the relationship model of the entity model includes attack sources triggering alarms, attack sources attacking assets, alarms occurring on assets, attack sources complying with threat intelligence and assets belonging to business systems.
[0165] Optionally, the information extraction module 520 is specifically used to call the first large language model to extract entity data corresponding to the entity model in the preset security model within the security situation data; call the second large language model to extract relationship data between different entity data within the security situation data according to the relationship model in the preset security model; for each entity data, extract the attribute information and entity identifier of the entity data, and extract the relationship identifier of the relationship data associated with the entity data, and save the entity identifier, the relationship identifier and the attribute information as the knowledge triple.
[0166] Optionally, in an embodiment of the present invention, the first large language model includes a word matching input layer, a feature representation layer, an attention encoding layer, and a conditional random field decoding layer, and the word matching input layer performs character matching based on a bidirectional maximum matching method; the second large language model includes a bidirectional encoder representation model based on Transformer.
[0167] Optionally, in this embodiment of the present invention, the attention score calculation formula of the attention encoding layer of the first large language model includes: Among them, R t-j It is a relative position code, i represents the number of characters, d k is the dimension of the key vector, t represents the index of the target token, and j represents the index of the context token.
[0168] Optionally, the defense module 540 is specifically used to obtain a pre-configured built-in analysis rule base, extract key fields of all built-in analysis rules in the built-in analysis rule base; traverse the target knowledge triples that match the key fields in the endogenous threat intelligence knowledge graph, and use the target knowledge triples as the attack path; search for built-in analysis rules that match the attack path in the built-in analysis rule base, and use the built-in analysis rules as the defense strategy.
[0169] Optionally, in an embodiment of the present invention, the built-in analysis rules include at least one of the following: within a preset time window, if the same attack source simultaneously attacks assets greater than or equal to a first preset number threshold, and each of the assets belongs to the same business system, then the attack path is determined to be a medium-risk level, and the built-in analysis rules that match the medium-risk level are searched in the built-in analysis rule library; within a preset time interval, if the same asset is attacked by different types of attacks from the same attack source, then the attack path is determined to be a high-risk level, and the built-in analysis rules that match the high-risk level are searched in the built-in analysis rule library; within a preset time interval, if the same asset is attacked by attack sources greater than or equal to a second preset number threshold, and each of the attack source entities belongs to the same attack type and the same threat intelligence, then the attack path is determined to be a high-risk level, and the built-in analysis rules that match the high-risk level are searched in the built-in analysis rule library.
[0170] The collaborative defense device based on the endogenous threat intelligence knowledge graph provided by the embodiment of the present invention can execute the collaborative defense method based on the endogenous threat intelligence knowledge graph provided by any embodiment of the present invention, and has the corresponding beneficial effects of the execution method.
[0171] Example 5
[0172] Embodiment 5 of the present invention provides a device for executing a collaborative defense method based on an endogenous threat intelligence knowledge graph, a computer-readable storage medium, and a computer program product.
[0173] Figure 8 A schematic diagram of the structure of an apparatus that can be used to implement an embodiment of the present invention is shown. The apparatus is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The apparatus may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown in the embodiments of the present invention, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present invention described and / or required herein.
[0174] like Figure 8As shown, the device includes at least one processor 11 and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the ROM 12 or the computer program loaded from the storage unit 18 into the RAM 13. Various programs and data required for the operation of the device 10 can also be stored in the RAM 13. The processor 11, ROM 12 and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0175] Multiple components in the device are connected to the I / O interface 15, including an input unit 16, such as a keyboard, mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, optical disk, etc.; and a communication unit 19, such as a network card, modem, wireless communication transceiver, etc. The communication unit 19 allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0176] Processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing units, graphics processing units, various specialized artificial intelligence computing chips, various processors running machine learning model algorithms, digital signal processors, and any suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as the collaborative defense method based on the endogenous threat intelligence knowledge graph.
[0177] In some embodiments, the collaborative defense method based on the endogenous threat intelligence knowledge graph can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the device via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the collaborative defense method based on the endogenous threat intelligence knowledge graph can be performed. Alternatively, in other embodiments, the processor 11 can be configured as a collaborative defense method based on the endogenous threat intelligence knowledge graph by any other appropriate means (for example, by means of firmware).
[0178] Various implementations of the systems and techniques described above in the embodiments of the present invention may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays, application specific integrated circuits, application specific standard products, system-on-chip systems, load programmable logic devices, computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: being implemented in one or more computer programs that are executable and / or interpreted on a programmable system including at least one programmable processor, which may be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0179] The computer programs for implementing the methods of the embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0180] In the context of an embodiment of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, RAM, ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0181] To provide interaction with a user, the systems and techniques described herein can be implemented on a device having: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0182] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks, wide area networks, blockchain networks, and the Internet.
[0183] A computing system may include clients and servers. The clients and servers are generally remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private server services.
[0184] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0185] The above specific embodiments do not constitute a limitation on the scope of protection of the embodiments of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A collaborative defense method based on endogenous threat intelligence knowledge graph, characterized by: The method comprises: Obtain security situation data from at least one data source; Extracting knowledge triples from the security situation data based on a preset security model; Constructing an endogenous threat intelligence knowledge graph based on the knowledge triples; An attack path is determined based on the endogenous threat intelligence knowledge graph, and a defense strategy for the attack path is determined.
2. The method according to claim 1, characterized in that The obtaining of security situation data of at least one data source includes at least one of the following: Collecting internal security data as the security situation data; Collecting first information data from an asset information database as the security situation data; Collecting second information data from a security knowledge base as the security situation data; Collecting third information data from the vulnerability intelligence database as the security situation data; The fourth information data of the basic resource library is collected as the security situation data.
3. The method according to claim 1, characterized in that The entity model of the preset security model includes alarms, attack sources, assets, business systems and threat intelligence, and the relationship model of the entity model includes attack sources triggering alarms, attack sources attacking assets, alarms occurring on assets, attack sources complying with threat intelligence and assets belonging to business systems.
4. The method according to claim 1 or 3, characterized in that The extracting of knowledge triples from the security situation data based on a preset security model includes: Calling a first large language model to extract entity data corresponding to the entity model in the preset security model from the security situation data; Invoking a second large language model to extract relationship data between different entity data in the security situation data according to the relationship model in the preset security model; For each entity data, the attribute information and entity identifier of the entity data are extracted, and the relationship identifier of the relationship data associated with the entity data is extracted, and the entity identifier, the relationship identifier and the attribute information are saved as the knowledge triple.
5. The method according to claim 4, characterized in that: The first large-scale language model includes a word matching input layer, a feature representation layer, an attention encoding layer, and a conditional random field decoding layer. The word matching input layer performs character matching based on a bidirectional maximum matching method; the second large-scale language model includes a bidirectional encoder representation model based on Transformer.
6. The method according to claim 4 or 5, characterized in that The attention score calculation formula of the attention encoding layer of the first large language model is include: Among them, R t-j It is a relative position code, i represents the number of characters, d k is the dimension of the key vector, t represents the index of the target token, and j represents the index of the context token.
7. The method according to claim 1, characterized in that: Determining an attack path based on the endogenous threat intelligence knowledge graph and determining a defense strategy for the attack path includes: Obtaining a pre-configured built-in analysis rule library, and extracting key fields of all built-in analysis rules in the built-in analysis rule library; Traversing the target knowledge triples matching the key fields in the endogenous threat intelligence knowledge graph, and using the target knowledge triples as the attack paths; The built-in analysis rule library is searched for a built-in analysis rule that matches the attack path, and the built-in analysis rule is used as the defense strategy.
8. The method according to claim 7, characterized in that: The built-in analysis rules include at least one of the following: If the same attack source attacks assets greater than or equal to a first preset number threshold, and the assets belong to the same business system, then the risk level of the attack path is determined to be medium, and a built-in analysis rule matching the medium risk level is searched in the built-in analysis rule library; If the same asset is attacked by different types of attacks from the same attack source, the risk level of the attack path is determined to be high, and a built-in analysis rule matching the high risk level is searched in the built-in analysis rule library; If the same asset is attacked by attack sources greater than or equal to a second preset number threshold, and each of the attack sources belongs to the same attack type and the same threat intelligence, the danger level of the attack path is determined to be a high-risk level, and a built-in analysis rule matching the high-risk level is searched in the built-in analysis rule library.
9. A collaborative defense system based on endogenous threat intelligence knowledge graph, characterized by: The system comprises: A data acquisition module, configured to obtain security situation data from at least one data source; An information extraction module, configured to extract knowledge triples from the security situation data based on a preset security model; A knowledge graph module, configured to construct an endogenous threat intelligence knowledge graph based on the knowledge triples; A defense module is used to determine an attack path based on the endogenous threat intelligence knowledge graph and determine a defense strategy for the attack path.
10. A device, characterized in that The device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the collaborative defense method based on the endogenous threat intelligence knowledge graph according to any one of claims 1 to 8.
Citation Information
Cited By
Threat intelligence knowledge graph construction method and device fusing attack intention
CN121212298A