Knowledge graph construction method and device for network security situation awareness
By constructing and optimizing triplets, the initial knowledge graph is optimized based on the importance of entity categories and relationship categories, and the redundancy and weak correlation problems of multi-source heterogeneous data in network security situation awareness is solved, and the efficient optimization of knowledge graphs and the full utilization of data value is achieved.
Patent Information
- Application Number
- CN202510219624.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-24
AI Technical Summary
In network security situation awareness, the multi-source heterogeneous data is large and redundant, making it difficult for existing knowledge graphs to fully utilize the data value, especially the existence of a large number of weakly correlated data, which affects the effective use of data.
By constructing multiple triplets, including entity features and relational features, and optimizing the initial knowledge graph based on the importance of entity categories and relational categories, unnecessary entity features and relational features are eliminated to form an optimized knowledge graph.
The optimization of the initial knowledge graph is achieved, the redundancy is reduced, weakly associated data is eliminated, and the utilization efficiency of data value is improved. It is suitable for knowledge graph optimization of large-scale multi-source heterogeneous data.
Smart Images

Figure CN120196790A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graphs, and more particularly, to a method and apparatus for constructing a knowledge graph for network security situation awareness. Background Art
[0002] The application of Knowledge Graph (KG) in network attacks has gradually attracted attention in recent years, especially in aspects such as network security defense, threat intelligence analysis, and attack behavior reasoning. A knowledge graph can integrate multi-source heterogeneous threat intelligence data (such as malicious IPs, domain names, vulnerabilities, attack patterns, etc.) to construct a threat intelligence knowledge graph, helping security analysts quickly discover potential threats; it can also be used to model the behavior patterns of attackers, discover hidden attack paths or attack intentions through reasoning techniques; and it can enhance the intelligence level of network security defense systems to assist in automated response and decision-making.
[0003] Currently, when a knowledge graph is applied in the field of network attacks, the main drawbacks are mainly reflected in: the extremely large amount of multi-source heterogeneous data generated in network security situation awareness, the data redundancy, and the existence of a large number of weakly associated data. It is difficult to fully exert the data value based on the knowledge graph built from such source data. Summary of the Invention
[0004] The present invention provides a method and apparatus for constructing a knowledge graph for network security situation awareness, which can overcome certain or some defects of the prior art.
[0005] According to the method for constructing a knowledge graph for network security situation awareness of the present invention, it includes:
[0006] Constructing a plurality of triples; wherein, any triple has an entity feature pair composed of two entity features and a corresponding relationship feature for characterizing the association between the corresponding entity feature pair. Any entity feature includes an entity ontology and a corresponding entity category label, and any relationship feature has a relationship ontology and a corresponding relationship category label;
[0007] Constructing an initial knowledge graph based on the plurality of triples;
[0008] Sorting the entity categories in the initial knowledge graph by importance based on the entity category labels;
[0009] Sorting the relationship categories in the initial knowledge graph by importance based on the entity category labels and the relationship category labels to obtain the importance score RP of each relationship category u ;
[0010] Based on the importance order of the entity categories and the importance score RP of each relationship category u, optimize the triples in the initial knowledge graph;
[0011] Based on the optimized triples, complete the construction of the knowledge graph.
[0012] Preferably, based on the entity category labels, sort the entity categories in the constructed knowledge graph by importance, including
[0013] Based on manual annotation, obtain the initial importance sequence L of entity categories s ; where L s ={L i |i∈N +}, L i is the i-th entity category in the initial importance sequence L s , and the importance of the i-th entity category L s in the initial importance sequence L i is greater than the importance of the (i + 1)-th entity category L i+1 ;
[0014] Construct an importance index for entity categories. Based on the initial importance sequence L s , obtain the importance sequence L of entity categories e ; where L e ={L j |j∈N +}, L j is the j-th entity category in the importance sequence L e , and the importance of the j-th entity category L e in the importance sequence L j is greater than the importance of the (j + 1)-th entity category L j+1 ;
[0015] Preferably, the construction of the importance index for entity categories is based on the initial importance sequence L s , and the importance sequence L of entity categories is obtained e , including
[0016] Construct the normalized initial importance P s of each entity category L i in the initial importance sequence L s i ;
[0017] Construct the normalized inter-class association index Q s of each entity category L i in the initial importance sequence L i ; where the normalized inter-class association index Q i is used to characterize the connection between each entity category L i and the rest of the entity categories;
[0018] Construct the initial importance sequence L s for each entity category L i in the normalized intra-class correlation index R i ; among them, the normalized intra-class correlation index R i is used to characterize the connection between entity ontologies of each entity feature in the same entity category L i ;
[0019] Construct the initial importance sequence L s for each entity category L i in the normalized risk index S i ; among them, the normalized risk index S i is used to characterize the connection between the entity category L i and the network threat event;
[0020] Based on the normalized initial importance the normalized inter-class correlation index Q i , the normalized intra-class correlation index R i and the normalized risk index S i , obtain the reconstructed importance of each entity category L s in the initial importance sequence L i ;
[0021] Based on the reconstructed importance sort each entity category L s in the initial importance sequence L i to obtain the importance sequence L e of the entity category.
[0022] Preferably, the importance of the i-th entity category L s in the initial importance sequence L i is where N is the total number of entity categories.
[0023] Preferably, the normalized inter-class correlation index Q s of each entity category L i in the constructed initial importance sequence L i , includes,
[0024] Obtain the correlation value i of the corresponding entity category L with each remaining entity category where, represents the current entity category L i and the a-th entity category L s in the initial importance sequence L aThe number of entity feature pairs among them, the correlation value when a = i Take 0, the correlation value when a < i Record as a positive value, the correlation value when a > i Record as a negative value;
[0025] Obtain the inter-class correlation index of each entity category L i among them Among them, is the normalized initial importance of the a-th entity category L a ;
[0026] Based on the inter-class correlation index of each entity category L i among them Obtain the normalized inter-class correlation index Q of each entity category L i ; Among them, i ;
[0027] Preferably, the normalized intra-class correlation index R of each entity category L s in the constructed initial importance sequence L i includes, i including,
[0028] Obtain the total number of entity features in each entity category L i among them and the total number of strongly associated entity features that form entity feature pairs with entity features in the remaining entity categories The total number of entity feature pairs formed by strongly associated entity features and entity features in the remaining entity categories and the total number of entity feature pairs formed by strongly associated entity features and entity features in the corresponding entity category L i among them
[0029] Obtain the intra-class correlation index of each entity category L i among them Among them,
[0030]
[0031] Obtain the normalized intra-class correlation index R of each entity category L i ; Among them, i ;
[0032] Preferably, the normalized risk index S of each entity category L s in the constructed initial importance sequence L i includes, i including,
[0033] Construct a historical network threat event set; among them, the network threat event set includes multiple occurred network threat events;
[0034] Obtain the occurrence frequency of each entity category L i in the historical network threat event set as a risk indicator
[0035] For the risk indicator perform normalization processing to obtain the normalized risk indicator S i ; among them,
[0036] Preferably, the acquisition method of the reconstruction importance is specifically as follows
[0037]
[0038] Preferably, based on the importance order of entity categories and the importance score RP u of each relationship category, optimize the triples in the initial knowledge graph, including
[0039] Obtain the importance sequence L e of entity categories and the importance score RP u of each relationship category; among them, L e ={L j |j∈N +}, L j is the j-th entity category in the importance sequence L e , and the importance of the j-th entity category L e in the importance sequence L j is greater than the importance of the (j + 1)-th entity category L j+1 ; among them, RP u is the importance score of the u-th relationship category;
[0040] Based on the importance sequence L e of entity categories, in the order of importance from low to high, divide the entity features in each entity category in turn to obtain the core entity feature sequence L j1 、the key entity feature sequence L j2 and the marginal entity feature sequence L j3 ;
[0041] Based on the key entity feature sequence L j2 and the core entity feature sequence L j1 , obtain the eliminated entity feature set C from the marginal entity feature sequence L j3 ;
[0042] Based on the elimination of the entity feature set C, the elimination relationship category set E is obtained;
[0043] Delete from the initial knowledge graph the triples with entity features in the entity feature set C and the relationship categories in the elimination relationship category set E to obtain the optimized triples.
[0044] In addition, the present invention also provides a device, which includes:
[0045] At least one processor; and
[0046] A memory communicatively connected to the at least one processor; wherein,
[0047] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the above methods.
[0048] The present invention has the following beneficial effects:
[0049] It is possible to form corresponding entity category labels and relationship category labels based on the classification of entity ontology and relationship ontology, and then construct corresponding entity features and relationship features; and it is possible to construct initial triples based on this, and then form an initial knowledge graph. After forming the initial knowledge graph, it is possible to eliminate unnecessary entity features and relationship features in the initial knowledge graph based on the importance of entity categories and relationship categories, so as to optimize the initial knowledge graph. The method provided by the present invention can use the importance index as the optimization index, and the optimization is carried out based on the architecture relationship of the knowledge graph, without involving semantic analysis or discrimination of a single entity ontology or relationship ontology, and can reduce the redundancy in the initial knowledge graph and eliminate weakly associated data based on the connection between data relationships, and is applicable to the optimization of large-scale multi-source heterogeneous knowledge graphs. Description of the Drawings
[0050] Figure 1 It is a schematic flowchart of a method for constructing a knowledge graph for network security situation awareness in Embodiment 1.
[0051] Figure 2 It is a schematic flowchart of a method for sorting the importance of entity categories in a knowledge graph.
[0052] Figure 3 It is a schematic flowchart of a method for sorting the importance of relationship categories in a knowledge graph in Embodiment 1.
[0053] Figure 4 It is a schematic flowchart of a method for optimizing an initial knowledge graph in Embodiment 1. Detailed implementation mode
[0054] To further understand the content of the present invention, the present invention will be described in detail in combination with embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.
[0055] Embodiment 1
[0056] As shown in Figure 1 , this embodiment provides a method for constructing a knowledge graph for network security situation awareness, which includes,
[0057] Constructing a plurality of triples; wherein, any triple has an entity feature pair composed of 2 entity features and a corresponding relationship feature for characterizing the association between the corresponding entity feature pair, any entity feature includes an entity ontology and a corresponding entity category label, and any relationship feature has a relationship ontology and a corresponding relationship category label;
[0058] Constructing an initial knowledge graph based on the plurality of triples;
[0059] Sorting the entity categories in the initial knowledge graph by importance based on the entity category labels;
[0060] Sorting the relationship categories in the initial knowledge graph by importance based on the entity category labels and relationship category labels, and obtaining the importance score RP of each relationship category u ;
[0061] Optimizing the triples in the initial knowledge graph based on the importance order of the entity categories and the importance score RP of each relationship category u , and completing the construction of the knowledge graph based on the optimized triples.
[0062] Based on the optimized triples, the construction of the knowledge graph is completed.
[0063] Based on the above method, it is possible to form corresponding entity category labels and relationship category labels based on the classification of entity ontologies and relationship ontologies, and then construct corresponding entity features and relationship features; and it is possible to construct initial triples based on this, and then form an initial knowledge graph. After forming the initial knowledge graph, it is possible to eliminate unnecessary entity features and relationship features in the initial knowledge graph based on the importance of entity categories and relationship categories, and then realize the optimization of the initial knowledge graph. The method provided by the present invention can use the importance index as the optimization index, and the optimization is carried out based on the architecture relationship of the knowledge graph, without involving semantic analysis or discrimination of individual entity ontologies or relationship ontologies, and can reduce the redundancy in the initial knowledge graph and eliminate weakly associated data based on the connection between data relationships, and is applicable to the optimization of knowledge graphs with large amounts of multi-source heterogeneous data.
[0064] Embodiment 2
[0065] seen in Figure 2 , in order to implement the sorting of entity categories in the constructed knowledge graph according to importance based on entity category labels described in Embodiment 1, this embodiment provides a method for sorting entity categories in a knowledge graph according to importance. It includes,
[0066] Based on manual annotation, obtain the initial importance sequence L of entity categories s ; where L s = {L i |i ∈ N +}, L i is the i-th entity category in the initial importance sequence L s , and the importance of the i-th entity category L s in the initial importance sequence L i is greater than the importance of the (i + 1)-th entity category L i+1 ;
[0067] Construct an importance index for entity categories, and based on the initial importance sequence L s , obtain the importance sequence L of entity categories e ; where L e = {L j |j ∈ N +}, L j is the j-th entity category in the importance sequence L e , and the importance of the j-th entity category L e in the importance sequence L j is greater than the importance of the (j + 1)-th entity category L j+1 ;
[0068] Based on the above steps, it is possible to first determine the initial importance ranking in a manual annotation manner, and then reconstruct the initial importance sequence by constructing an importance index, which can preferably introduce manual experience and mathematical indicators, and realize the importance ranking of entity categories that combines subjective and objective evaluations.
[0069] In this embodiment, the construction of the importance index for entity categories, based on the initial importance sequence L s , to obtain the importance sequence L of entity categories e , includes,
[0070] Construct the normalized initial importance of each entity category L s in the initial importance sequence L i
[0071] Construct the normalized initial importance of each entity category L s in the initial importance sequence Li The normalized inter-class correlation index Q i ; Among them, the normalized inter-class correlation index Q i Used to characterize each entity category L i Relationships with other entity classes;
[0072] Construct the initial importance sequence L s For each entity class L i The normalized intra-class correlation index R i ; Among them, the normalized intra-class correlation index R i Used to represent the same entity category L i The relationship between the entity ontology of each entity feature in;
[0073] Construct the initial importance sequence L s For each entity class L i The normalized risk index S i ; Among them, the normalized risk index S i Used to represent entity category L i Links to cyber threat incidents;
[0074] Based on normalized initial importance Normalized inter-class correlation index Q i , Normalized intra-class correlation index R i and normalized risk index S i , get the initial importance sequence L s For each entity class L i The importance of reconstruction
[0075] Reconstruction importance based For the initial importance sequence L s For each entity class L i Sort and obtain the importance sequence L of entity categories e .
[0076] Based on the above, the importance of entity categories can be reconstructed better based on artificial experience, the associations between each entity category, the associations within each entity category, and the associations between each entity category and network threat events, thereby better realizing the optimization of the knowledge graph based on the association between data.
[0077] In this embodiment, the initial importance sequence L s The i-th entity category L i The importance of N is the total number of entity categories. Therefore, the initial importance of each entity category can be obtained simply and quickly.
[0078] In this embodiment, the initial importance sequence L s for each entity category L i in the normalized inter-class association index Q i , including
[0079] obtain the association value corresponding to the entity category L i and each of the remaining entity categories of the association value where represents the current entity category L i and the initial importance sequence L s in the a-th entity category L a the number of entity feature pairs between, a = i when the association value L i a take 0, a less than i when the association value is recorded as a positive value, a greater than i when the association value is recorded as a negative value;
[0080] obtain the inter-class association index for each entity category L i where where is the normalized initial importance of the a-th entity category L a ;
[0081] based on each entity category L i of the inter-class association index obtain the normalized inter-class association index Q for each entity category L i i ; where
[0082] Based on the above, it is possible to better analyze the association degree between each entity category, and thus better obtain the normalized inter-class association index.
[0083] In this embodiment, the normalized intra-class association index R for each entity category L s in the initial importance sequence L i includes i , including
[0084] obtain the total number of entity features in each entity category L i and the total number of strongly associated entity features that form entity feature pairs with entity features in the remaining entity categories the total number of entity feature pairs formed by the strongly associated entity features and entity features in the remaining entity categories and the total number of entity feature pairs formed by the strongly associated entity features and entity features in the corresponding entity category L and the total number of entity feature pairs formed by the strongly associated entity features and entity features in the corresponding entity category L i
[0085] Obtain the intra-class association index for each entity category L i ; among them, Among them,
[0086]
[0087] Obtain the intra-class association index for each entity category L i of the normalized intra-class association index R i ; among them,
[0088] Based on the above, it is possible to preferably analyze the intra-class association degree of each entity category based on the data volume of each entity category and the data volume associated with other entity categories, and then preferably obtain the normalized intra-class association index.
[0089] In this embodiment, the construction of the initial importance sequence L s each entity category L in i of the normalized risk index S i , including,
[0090] Construct a historical network threat event set; among them, the network threat event set includes multiple occurred network threat events;
[0091] Obtain the occurrence frequency of each entity category L i in the historical network threat event set as the risk index
[0092] Perform normalization processing on the risk index to obtain the normalized risk index S i ; among them,
[0093] Based on the above, it is possible to preferably introduce historical risk data, realize the analysis of the importance of each entity category, and preferably realize the acquisition of the normalized risk index.
[0094] In this embodiment, the acquisition method of the reconstructed importance is specifically as follows
[0095]
[0096] Based on this, it is possible to preferably realize the acquisition of the reconstructed importance.
[0097] Embodiment 3
[0098] See Figure 3, in order to rank the relationship categories in the initial knowledge graph according to the entity category labels and relationship category labels in Embodiment 1, this embodiment provides a method for ranking the importance of relationship categories in a knowledge graph. It includes,
[0099] Based on the entity category labels and relationship category labels, obtain the total number T of each relationship category in each entity category uv , and construct an association matrix T between the relationship categories and entity categories; where T uv represents the total number of times the u-th relationship category appears in the v-th entity category, T = (T uv ) H*N , H is the total number of relationship categories, and N is the total number of entity categories;
[0100] Based on the association matrix T, obtain the importance score RP of each relationship category u ; where RP u is the importance score of the u-th relationship category;
[0101] Based on the importance score RP u , obtain the importance sequence RR of the relationship categories; where RR = {RR u |u ∈ N +}, RR u is the u-th relationship category in the importance sequence R, and the importance of the u-th relationship category RR u in the importance sequence R is greater than the importance of the (u + 1)-th relationship category RR u+1 .
[0102] Based on the above, it is possible to preferably evaluate the relationship categories by entity category. By using the relationship categories of more important entity feature data in the knowledge graph for evaluation, it is possible to preferably achieve the objectivity of the importance evaluation of the relationship categories.
[0103] In this embodiment, the obtaining of the importance score RP of each relationship category based on the association matrix T u includes,
[0104] Obtain the proportion LP of each entity category in each relationship category uv ;
[0105] Obtain the information entropy LH of each entity category v ;
[0106] Obtain the weight LW of each entity category v ;
[0107] Obtain the initial importance score RP of each relationship category u * ;
[0108] Obtain the importance score RP for each relationship category u ;
[0109] Among them,
[0110]
[0111] Based on the above, the concept of information entropy can be introduced to evaluate the importance of relationship categories, making the evaluation results more scientific.
[0112] Embodiment 4
[0113] As shown in Figure 4 , in order to implement the importance order based on entity categories and the importance score RP for each relationship category described in Embodiment 1 u , optimize the triples in the initial knowledge graph. This embodiment provides a method for optimizing the initial knowledge graph. It includes,[[]]
[0114] Obtain the importance sequence L of entity categories e and the importance score RP for each relationship category u ; Among them, L e ={L j |j∈N +}, L j is the j-th entity category in the importance sequence L e . The importance of the j-th entity category L e in the importance sequence L is greater than that of the (j + 1)-th entity category L j ; Among them, RP j+1 is the importance score of the u-th relationship category; u
[0115] Based on the importance sequence L of entity categories e , in the order of importance from low to high, divide the entity features in each entity category in turn to obtain the core entity feature sequence L j1 , the key entity feature sequence L j2 and the marginal entity feature sequence L j3 ;
[0116] Based on the key entity feature sequence L j2 and the core entity feature sequence L j1 , obtain the set C of entity features to be removed from the marginal entity feature sequence L j3 ;
[0117] Based on the set C of entity features to be removed, obtain the set E of relationship categories to be removed;
[0118] Delete triples with entity features in the entity feature set C and relationship categories in the excluded relationship category set E from the initial knowledge graph to obtain optimized triples.
[0119] Based on the above, the optimization of the knowledge graph can be better achieved. In particular, during the optimization process, entity categories that are less important can be processed first, which can greatly avoid the occurrence of over-optimization and effectively prevent the deletion of important data.
[0120] In this embodiment, the core entity feature sequence L j1 is obtained specifically as follows.
[0121] Obtain the entity features that have relationship features with the entity features in the remaining entity categories in each entity category L j and use them as core entity features to construct the core entity feature sequence L j1 ; where is the α-th core entity feature in the core entity feature sequence L j1 .
[0122] Based on the above, the screening of core entity features can be better achieved.
[0123] In this embodiment, the key entity feature sequence L j2 and the marginal entity feature sequence L j3 are obtained specifically as follows.
[0124] Obtain the number of relationship features of each relationship category possessed by each core entity feature ; where where represents the number of relationship features of the u-th relationship category at the core entity feature ;
[0125] Obtain the number of relationship features of each relationship category possessed by each non-core entity feature ; where where is the β-th non-core entity feature, is the non-core entity feature and represents the number of relationship features of the u-th relationship category at the non-core entity feature
[0126] Based on the importance score RP u of each relationship category, obtain the feature score Z α of each core entity feature and the feature score Z β of each non-core entity feature; where H is the total number of relationship categories;
[0127] Using the minimum value MIn(Z of the feature scores Z of all core entity features α as a threshold, the feature scores Z α not lower than the threshold MIn(Z β ) of the non-core entity features are used as key entity features, and the feature scores Z α lower than the threshold MIn(Z β ) of the non-core entity features are used as marginal entity features, completing the construction of the key entity feature sequence L α and the marginal entity feature sequence L j2 . j3
[0128] Based on the above, the acquisition of key entity features and marginal entity features can be better realized.
[0129] In this embodiment, based on the key entity feature sequence L j2 and the core entity feature sequence L j1 , the excluded entity feature set C is obtained from the marginal entity feature sequence L j3 , including
[0130] obtaining the entity feature set D included in the shortest path from each key entity feature in the key entity feature sequence L j2 to each core entity feature in each core entity feature sequence L j1 ;
[0131] Taking the set composed of entity features that belong to the marginal entity feature sequence L j3 and do not belong to the entity feature set D as the excluded entity feature set C.
[0132] Based on the above, the acquisition of excluded entity features can be better realized.
[0133] In this embodiment, based on the excluded entity feature set C, the excluded relationship category set E is obtained, including
[0134] obtaining the occurrence times of each relationship category in the excluded entity feature set C;
[0135] constructing the excluded relationship category set E with relationship categories whose occurrence times are in the interval [μ - 3σ, μ + 3σ]; where μ is the mean of the occurrence times of each relationship category, and σ is the variance of the occurrence times of each relationship category.
[0136] Based on the above, the acquisition of excluded relationship categories can be better realized.
[0137] Embodiment 5
[0138] This embodiment provides a device, which includes:
[0139] at least one processor; and
[0140] a memory communicatively connected to the at least one processor; wherein
[0141] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of Embodiments 1-4.
[0142] It is easily understandable that those skilled in the art can combine, split, recombine, etc. the embodiments of the present application based on one or several embodiments provided in the present application to obtain other embodiments, and these embodiments do not exceed the protection scope of the present application.
[0143] The present invention and its implementation manners are schematically described above. The description is not restrictive. What is shown in the embodiments is only part of the implementation manners of the present invention, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design, without creative efforts, a structural manner and an embodiment similar to the technical solution without departing from the gist of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A knowledge graph construction method for network security situation awareness, comprising: Construct multiple triples; wherein any triple has an entity feature pair consisting of two entity features and a corresponding relationship feature for characterizing the association between the corresponding entity feature pairs, any entity feature includes an entity ontology and a corresponding entity category label, and any relationship feature has a relationship ontology and a corresponding relationship category label; Constructing an initial knowledge graph based on the multiple triples; Based on entity category labels, the entity categories in the initial knowledge graph are sorted by importance; Based on entity category labels and relationship category labels, the relationship categories in the initial knowledge graph are sorted by importance, and the importance score RP of each relationship category is obtained. u ; Based on the importance order of entity categories and the importance score RP of each relationship category u , optimize the triples in the initial knowledge graph; Based on the optimized triples, the knowledge graph is constructed.
2. The method for constructing a knowledge graph for network security situation awareness according to claim 1, characterized in that: The entity categories in the constructed knowledge graph are sorted by importance based on entity category labels, including: Based on manual annotation, obtain the initial importance sequence L of entity categories s Among them, L s ={L i |i∈N + }, L i is the initial importance sequence L s The i-th entity category in the initial importance sequence L s The i-th entity category L i The importance of is greater than the i+1th entity category L i+1 The importance of Construct the importance index of entity category based on the initial importance sequence L s , get the importance sequence L of entity categories e Among them, L e ={L j |j∈N + }, L j is the importance sequence L e The j-th entity category in the importance sequence L e The jth entity category L in j The importance of is greater than the j+1th entity category L j+1 The importance of.
3. The method for constructing a knowledge graph for network security situation awareness according to claim 2, characterized in that: The importance index of the constructed entity category is based on the initial importance sequence L s , get the importance sequence L of entity categories e ,include, Construct the initial importance sequence L s For each entity class L i The normalized initial importance P s i ; Construct the initial importance sequence L s For each entity class L i The normalized inter-class correlation index Q i ; Among them, the normalized inter-class correlation index Q i Used to characterize each entity category L i Relationships with other entity classes; Construct the initial importance sequence L s For each entity class L i The normalized intra-class correlation index R i ; Among them, the normalized intra-class correlation index R i Used to represent the same entity category L i The relationship between the entity ontology of each entity feature in; Construct the initial importance sequence L s For each entity class L i The normalized risk index S i ; Among them, the normalized risk index S i Used to represent entity category L i Links to cyber threat incidents; Based on normalized initial importance Normalized inter-class correlation index Q i , Normalized intra-class correlation index R i and normalized risk index S i , get the initial importance sequence L s For each entity class L i The importance of reconstruction Reconstruction importance based For the initial importance sequence L s For each entity class L i Sort and obtain the importance sequence L of entity categories e .
4. The method for constructing a knowledge graph for network security situation awareness according to claim 3 is characterized in that: Initial importance sequence L s The i-th entity category L i The importance of N is the total number of entity categories.
5. The method for constructing a knowledge graph for network security situation awareness according to claim 4 is characterized in that: The initial importance sequence L is constructed s For each entity class L i The normalized inter-class correlation index Q i ,include, Get the corresponding entity category L i With each other entity class The associated value in, Indicates the current entity category L i With the initial importance sequence L s The a-th entity category L a The number of entity feature pairs between them, when a=i, the association value Take 0, the associated value when a is less than i Recorded as a positive value, when a is greater than i, the associated value Recorded as negative value; Get each entity category L i The inter-class correlation index in, is the a-th entity category L a The normalized initial importance of ; Based on each entity category L i The inter-class correlation index Get each entity category L i The normalized inter-class correlation index Q i ;in, 6. The method for constructing a knowledge graph for network security situation awareness according to claim 5, characterized in that: The initial importance sequence L is constructed s For each entity class L i The normalized intra-class correlation index R i ,include, Get each entity category L i The total number of entity features in The total number of strongly associated entity features that form entity feature pairs with entity features in other entity categories The total number of entity feature pairs formed by strongly associated entity features and entity features in other entity categories And the strong association entity features and the corresponding entity category L i The total number of entity feature pairs formed by the entity features in Get each entity category L i The intra-class correlation index in, Get each entity category L i The normalized intra-class correlation index R i ;in, 7. The method for constructing a knowledge graph for network security situation awareness according to claim 6, characterized in that: The initial importance sequence L is constructed s For each entity class L i The normalized risk index S i ,include, Constructing a historical network threat event set; wherein the network threat event set includes multiple network threat events that have occurred; Get each entity category L i The frequency of occurrence in historical network threat events as a risk indicator Risk Indicators Perform normalization processing to obtain the normalized risk index S i ;in, 8. The method for constructing a knowledge graph for network security situation awareness according to claim 7 is characterized in that: Reconstruction importance The specific method of obtaining is:
9. The method for constructing a knowledge graph for network security situation awareness according to claim 1, characterized in that: The importance order based on entity categories and the importance score RP of each relationship category u , optimize the triples in the initial knowledge graph, including, Get the importance sequence L of entity categories e And the importance score RP of each relationship category u Among them, L e ={L j |j∈N + }, L j is the importance sequence L e The j-th entity category in the importance sequence L e The jth entity category L in j The importance of is greater than the j+1th entity category L j+1 The importance of RP u Score the importance of the u-th relationship category; Importance sequence L based on entity category e , divide the entity features in each entity category in order from low to high importance, and obtain the core entity feature sequence L j1 , key entity feature sequence L j2 And the edge entity feature sequence L j3 ; Based on the key entity feature sequence L j2 and core entity feature sequence L j1 , from the edge entity feature sequence L j3 Get the feature set C of the excluded entity; Based on the eliminated entity feature set C, obtain the eliminated relationship category set E; Delete the triplets with entity features in the entity feature set C and the relationship categories in the excluded relationship category set E from the initial knowledge graph to obtain optimized triplets.
10. An apparatus comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
Citation Information
Cited By
Knowledge graph optimization method and device and electronic equipment
CN120744194A
Security and protection monitoring situation awareness method and device based on big data and medium
CN121561358A
A security monitoring situation awareness method and device based on big data, and a medium
CN121561358B