Knowledge graph optimization method and device suitable for network security situation awareness data

By optimizing the knowledge graph of network security situation-aware data, using the importance scores of entity categories and relationship categories, redundant and weakly associated data are eliminated, and the optimized triplets are formed, which solves the problems of large data volume and high redundancy and improves the utilization of data value.

CN120090833APending Publication Date: 2025-06-03STATE GRID ANHUI ELECTRIC POWER CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510219720.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In network security situation awareness, the multi-source heterogeneous data generated is huge and has high redundancy, making it difficult to fully utilize the data value of the knowledge graph based on this data, especially the existence of a large number of weakly correlated data.

Method used

By constructing an initial knowledge graph, we obtain the importance sequence of entity categories and the importance score of relation categories, divide the entity features in turn, obtain core, key and edge entity features, eliminate unnecessary entity features and relation categories, and perform optimization reconstruction to form an optimized triple.

Benefits of technology

The knowledge graph is effectively optimized, redundancy is reduced, weakly associated data is eliminated, over-optimization is avoided, important data is retained, and data value utilization of the knowledge graph is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090833A_ABST
    Figure CN120090833A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of knowledge graph optimization, in particular to a knowledge graph optimization method and device suitable for network security situation awareness data. The method comprises the following steps: constructing an initial knowledge graph; obtaining an importance degree sequence of the entity category and an importance degree score of each relation category; based on the importance degree sequence of the entity categories, sequentially dividing the entity features in each entity category according to the importance degree from low to high, and obtaining a core entity feature sequence, a key entity feature sequence and an edge entity feature sequence; obtaining a rejected entity feature set; obtaining a rejection relationship category set; deleting triads with entity features in the entity feature set and relation categories in the relation category set from the initial knowledge graph, and obtaining optimized triads; and reconstructing the initial knowledge graph based on the optimized triple. According to the method, the knowledge graph can be optimized well.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph optimization, and more specifically, to a method and device for optimizing a knowledge graph applicable to network security situation awareness data. Background Art

[0002] In recent years, the application of Knowledge Graph (KG) in cyberattacks has gradually attracted attention, especially in aspects such as network security defense, threat intelligence analysis, and attack behavior reasoning. The knowledge graph can integrate multi-source heterogeneous threat intelligence data (such as malicious IPs, domain names, vulnerabilities, attack patterns, etc.) to construct a threat intelligence knowledge graph, helping security analysts quickly discover potential threats; it can also be used to model the behavior patterns of attackers, and discover hidden attack paths or attack intentions through reasoning techniques; it can also enhance the intelligence level of network security defense systems and help with automated response and decision-making.

[0003] Currently, when the knowledge graph is applied to the field of cyberattacks, the main drawbacks are mainly reflected in: the extremely large amount of multi-source heterogeneous data generated in network security situation awareness, the high data redundancy, and the existence of a large number of weakly associated data. It is difficult to fully utilize the data value based on the knowledge graph built from such source data. Summary of the Invention

[0004] The present invention provides a method for optimizing a knowledge graph applicable to network security situation awareness data, which can overcome certain or some defects of the prior art.

[0005] According to the method for optimizing a knowledge graph applicable to network security situation awareness data of the present invention, it includes:

[0006] Constructing an initial knowledge graph; wherein, the initial knowledge graph has multiple triples, and any triple has an entity feature pair composed of 2 entity features and a corresponding relationship feature used to represent the association between the corresponding entity feature pair. Any entity feature includes an entity ontology and a corresponding entity category label, and any relationship feature has a relationship ontology and a corresponding relationship category label;

[0007] Obtaining the importance sequence L e of entity categories and the importance score RP u of each relationship category; wherein, L e ={L j |j ∈ N +}, L j is the j-th entity category in the importance sequence L e , and the importance of the j-th entity category L e in the importance sequence L j is greater than that of the (j + 1)-th entity category Lj+1 importance; where RP u is the importance score of the u-th relationship category;

[0008] Based on the importance sequence L of entity categories e , in the order of increasing importance, the entity features in each entity category are sequentially divided to obtain the core entity feature sequence L j1 , the key entity feature sequence L j2 and the marginal entity feature sequence L j3 ;

[0009] Based on the key entity feature sequence L j2 and the core entity feature sequence L j1 , the excluded entity feature set C is obtained from the marginal entity feature sequence L j3 ;

[0010] Based on the excluded entity feature set C, the excluded relationship category set E is obtained;

[0011] Triples with entity features in the entity feature set C and relationship categories in the excluded relationship category set E are deleted from the initial knowledge graph to obtain optimized triples;

[0012] Based on the optimized triples, the initial knowledge graph is reconstructed.

[0013] Preferably, the core entity feature sequence L j1 is obtained as follows,

[0014] Obtain the entity features in each entity category L j that have relationship features with entity features in other entity categories, and use them as core entity features to construct the core entity feature sequence L j1 ; where, is the α-th core entity feature in the core entity feature sequence L j1 ;

[0015] Preferably, the key entity feature sequence L j2 and the marginal entity feature sequence L j3 are obtained as follows,

[0016] Obtain the number of relationship features of each relationship category possessed by each core entity feature ; where, where, represents the number of relationship features of the u-th relationship category at the core entity feature ;

[0017] Obtain each non-core entity feature The number of relationship features of each relationship category in the location wherein is the β-th non-core entity feature is a non-core entity feature is the number of relationship features of the u-th relationship category at the location;

[0018] Based on the importance score RP of each relationship category u , obtain the feature score Z of each core entity feature α and the feature score Z of each non-core entity feature β ; wherein is the total number of relationship categories;

[0019] Using the minimum value MIn(Z α ) of the feature scores Z of all core entity features as the threshold, the non-core entity features with feature scores Z α not lower than the threshold MIn(Z β ) are used as key entity features, and the non-core entity features with feature scores Z α lower than the threshold MIn(Z β ) are used as marginal entity features, and complete the construction of the key entity feature sequence L j2 and the marginal entity feature sequence L j3 . j2 Preferably, based on the key entity feature sequence L

[0020] and the core entity feature sequence L j1 , obtain the set C of entity features to be removed from the marginal entity feature sequence L j3 , including

[0021] Obtain the set D of entity features included in the shortest path from each key entity feature in the key entity feature sequence L j2 to each core entity feature in each core entity feature sequence L j1 ;

[0022] Taking the set composed of entity features that belong to the marginal entity feature sequence L j3 and do not belong to the set D of entity features as the set C of entity features to be removed. e

[0023] Preferably, based on the set C of entity features to be removed, obtain the set E of relationship categories to be removed, including

[0024] Obtain the number of occurrences of each relationship category in the set C of entity features to be removed;

[0025] ​Construct a set E of excluded relationship categories based on the relationship categories whose occurrence times are in the interval [μ - 3σ, μ + 3σ]; where μ is the mean of the occurrence times of each relationship category, and σ is the variance of the occurrence times of each relationship category.

[0026] Preferably, the importance sequence L of entity categories e The acquisition method includes

[0027] Based on manual annotation, obtain the initial importance sequence L of entity categories s ; where L s ={L i |i ∈ N +}, L i is the i-th entity category in the initial importance sequence L s , and the importance of the i-th entity category L s in the initial importance sequence L i is greater than the importance of the (i + 1)-th entity category L i+1 .

[0028] Construct the importance index of entity categories. Based on the initial importance sequence L s , obtain the importance sequence L of entity categories e ; where L e ={L j |j ∈ N +}, L j is the j-th entity category in the importance sequence L e , and the importance of the j-th entity category L e in the importance sequence L j is greater than the importance of the (j + 1)-th entity category L j+1 .

[0029] Preferably, the acquisition method of the importance score RP of each relationship category u includes

[0030] Based on the entity category label and the relationship category label, obtain the total number T of each relationship category in each entity category uv , and construct the association matrix T between the relationship category and the entity category; where T uv represents the total number of times the u-th relationship category appears in the v-th entity category, and T=(T uv ) H*N , H is the total number of relationship categories, and N is the total number of entity categories;

[0031] Based on the association matrix T, obtain the importance score RP of each relationship category u ; where RP u is the importance score of the u-th relationship category.

[0032] Preferably, based on the association matrix T, the importance score RP of each relationship category is obtained u , including

[0033] Obtain the proportion LP of each entity category in each relationship category uv ;

[0034] Obtain the information entropy LH of each entity category v ;

[0035] Obtain the weight LW of each entity category v ;

[0036] Obtain the initial importance score of each relationship category

[0037] Obtain the importance score RP of each relationship category u ;

[0038] Among them,

[0039]

[0040] In addition, the present invention also provides a device, including:

[0041] At least one processor; and

[0042] A memory communicatively connected to the at least one processor; wherein,

[0043] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the above methods.

[0044] The present invention has the following beneficial effects:

[0045] It can preferably realize the optimization of the knowledge graph. In particular, during the optimization process, it can first process the less important entity categories, which can greatly avoid the occurrence of over-optimization and effectively avoid the elimination of important data. Description of the Drawings

[0046] Figure 1 It is a schematic flow chart of a method for constructing a knowledge graph for network security situation awareness in Embodiment 1.

[0047] Figure 2 It is a schematic flow chart of a method for ranking the importance of entity categories in a knowledge graph.

[0048] Figure 3It is a flowchart showing a method for ranking the importance of relationship categories in a knowledge graph in Embodiment 1.

[0049] Figure 4 It is a flowchart showing a method for optimizing an initial knowledge graph in Embodiment 1. Detailed implementation manners

[0050] To further understand the content of the present invention, the present invention will be described in detail in combination with embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.

[0051] Embodiment 1

[0052] As shown in Figure 1 , this embodiment provides a method for constructing a knowledge graph for network security situation awareness, which includes,

[0053] Constructing a plurality of triples; wherein, any triple has an entity feature pair composed of 2 entity features and a corresponding relationship feature for characterizing the association between the corresponding entity feature pair, any entity feature includes an entity ontology and a corresponding entity category label, and any relationship feature has a relationship ontology and a corresponding relationship category label;

[0054] Constructing an initial knowledge graph based on the plurality of triples;

[0055] Based on the entity category labels, ranking the entity categories in the initial knowledge graph by importance;

[0056] Based on the entity category labels and relationship category labels, ranking the relationship categories in the initial knowledge graph by importance to obtain the importance score RP of each relationship category u ;

[0057] Based on the importance order of entity categories and the importance score RP of each relationship category u , optimizing the triples in the initial knowledge graph;

[0058] Based on the optimized triples, completing the construction of the knowledge graph.

[0059] Based on the above method, it is possible to classify entity ontologies and relationship ontologies, and then form corresponding entity category labels and relationship category labels, and further construct corresponding entity features and relationship features; and it is possible to construct initial triples based on this, and then form an initial knowledge graph. After forming the initial knowledge graph, it is possible to eliminate unnecessary entity features and relationship features in the initial knowledge graph based on the importance of entity categories and relationship categories, thereby realizing the optimization of the initial knowledge graph. The method provided by the present invention can use the importance index as the optimization index, and the optimization is carried out based on the architecture relationship of the knowledge graph, without involving semantic analysis or discrimination of individual entity ontologies or relationship ontologies, and can reduce the redundancy in the initial knowledge graph and eliminate weakly associated data based on the connection between data relationships, and is applicable to the optimization of large-scale multi-source heterogeneous knowledge graphs.

[0060] Example 2

[0061] As shown in Figure 2 , in order to implement the sorting of entity categories in the constructed knowledge graph based on the importance of entity category labels described in Example 1, this example provides a method for sorting the importance of entity categories in a knowledge graph. It includes:

[0062] Based on manual annotation, obtain the initial importance sequence L of entity categories s ; where L s = {L i | i ∈ N +}, L i is the i-th entity category in the initial importance sequence L s , and the importance of the i-th entity category L s in the initial importance sequence L i is greater than the importance of the (i + 1)-th entity category L i+1 ;

[0063] Construct the importance index of entity categories, and based on the initial importance sequence L s , obtain the importance sequence L e of entity categories; where L e = {L j | j ∈ N +}, L j is the j-th entity category in the importance sequence L e , and the importance of the j-th entity category L e in the importance sequence L j is greater than the importance of the (j + 1)-th entity category L j+1 ;

[0064] Based on the above steps, it is possible to first determine the initial importance ranking in the way of manual annotation, and then reconstruct the initial importance sequence by constructing importance indicators, which can preferably introduce manual experience and mathematical indicators, and realize the importance ranking of entity categories that integrates subjective and objective evaluations.

[0065] In this embodiment, the importance indicators of entity categories are constructed based on the initial importance sequence L s , and the importance sequence L of entity categories is obtained e , including

[0066] Construct the normalized initial importance of each entity category L in the initial importance sequence L s i

[0067] Construct the normalized inter-class association index Q of each entity category L in the initial importance sequence L s i ; among them, the normalized inter-class association index Q i is used to characterize the connection between each entity category L i and the rest of the entity categories; i

[0068] Construct the normalized intra-class association index R of each entity category L in the initial importance sequence L s i ; among them, the normalized intra-class association index R i is used to characterize the connection between the entity ontologies of each entity feature in the same entity category L i i ;

[0069] Construct the normalized risk index S of each entity category L in the initial importance sequence L s i ; among them, the normalized risk index S i is used to characterize the connection between the entity category L i and network threat events; i

[0070] Based on the normalized initial importance P s i , the normalized inter-class association index Q i , the normalized intra-class association index R i and the normalized risk index S i , obtain the reconstructed importance of each entity category L in the initial importance sequence L s i

[0071] Based on the reconstructed importance Sort the initial importance sequence L s for each entity category L i in it to obtain the importance sequence L e of the entity categories.

[0072] Based on the above, it is possible to preferably reconstruct the importance of entity categories based on artificial experience, the associations between each entity category, the associations within each entity category, and the associations between each entity category and network threat events, so as to preferably optimize the knowledge graph based on the association degree between data.

[0073] In this embodiment, for the i-th entity category L s in the initial importance sequence L i the importance is where N is the total number of entity categories. Therefore, it is possible to simply and quickly obtain the initial importance of each entity category.

[0074] In this embodiment, the normalized inter-category association index Q s of each entity category L i in the constructed initial importance sequence L i includes:

[0075] Obtain the association value i corresponding to the entity category L and each of the remaining entity categories where represents the number of entity feature pairs between the current entity category L i and the a-th entity category L s in the initial importance sequence L a When a = i, the association value is taken as 0. When a < i, the association value is recorded as a positive value. When a > i, the association value is recorded as a negative value;

[0076] Obtain the inter-category association index i of each entity category L where is the normalized initial importance of the a-th entity category L a ;

[0077] Based on the inter-category association index i of each entity category L obtain the normalized inter-category association index Q i of each entity category L i ; where

[0078] Based on the above, it is possible to better analyze the correlation degree between each entity category, and then better obtain the normalized inter-class correlation index.

[0079] In this embodiment, the construction of the initial importance sequence L s each entity category L in i the normalized intra-class correlation index R i , including,

[0080] Obtain the total number of entity features in each entity category L i and the total number of strongly correlated entity features that form entity feature pairs with entity features in the remaining entity categories the total number of entity feature pairs formed by the strongly correlated entity features and entity features in the remaining entity categories and the total number of entity feature pairs formed by the strongly correlated entity features and entity features in the corresponding entity category L and the total number of entity feature pairs formed by the strongly correlated entity features and entity features in the corresponding entity category L i in

[0081] Obtain the intra-class correlation index of each entity category L i Among them,

[0082]

[0083] Obtain the normalized intra-class correlation index R of each entity category L i ; among them, i ;

[0084] Based on the above, it is possible to better analyze the intra-class correlation degree of each entity category with the data volume of each entity category and the data volume associated with other entity categories, and then better obtain the normalized intra-class correlation index.

[0085] In this embodiment, the normalized risk index S of each entity category L in the constructed initial importance sequence L s i i , including,

[0086] Construct a historical network threat event set; the network threat event set includes a plurality of occurred network threat events;

[0087] Obtain the occurrence frequency of each entity category L i in the historical network threat event set as the risk index

[0088] Perform normalization processing on the risk index to obtain the normalized risk index S​​​i ; Among them,

[0089] Based on the above, historical risk data can be preferably introduced to realize the analysis of the importance of each entity category, and preferably, the acquisition of normalized risk indicators is realized.

[0090] In this embodiment, the reconstruction importance is obtained in the following specific way:

[0091]

[0092] Based on this, the acquisition of the reconstruction importance can be preferably realized.

[0093] Embodiment 3

[0094] As shown in Figure 3 , in order to rank the relationship categories in the initial knowledge graph according to importance based on the entity category labels and relationship category labels in Embodiment 1, this embodiment provides a method for ranking the importance of relationship categories in a knowledge graph. It includes:

[0095] Based on the entity category labels and relationship category labels, obtain the total number T of each relationship category in each entity category uv , and construct an association matrix T of relationship categories and entity categories; where T uv represents the total number of times the u-th relationship category appears in the v-th entity category, and T=(T uv ) H*N , H is the total number of relationship categories, and N is the total number of entity categories;

[0096] Based on the association matrix T, obtain the importance score RP of each relationship category u ; where RP u is the importance score of the u-th relationship category;

[0097] Based on the importance score RP u , obtain the importance sequence RR of relationship categories; where RR={RR u |u∈N +},RR u is the u-th relationship category in the importance sequence R, and the importance of the u-th relationship category RR u in the importance sequence R is greater than the importance of the (u + 1)-th relationship category RR u+1 .

[0098] Based on the above, the relationship categories can be preferably evaluated by entity categories. By using the relationship categories of more important entity feature data in the knowledge graph for evaluation, the objectivity of the importance evaluation of relationship categories can be preferably realized.

[0099] In this embodiment, based on the association matrix T, the importance score RP of each relationship category is obtained u , including

[0100] Obtain the proportion LP of each entity category in each relationship category uv ;

[0101] Obtain the information entropy LH of each entity category v ;

[0102] Obtain the weight LW of each entity category v ;

[0103] Obtain the initial importance score of each relationship category

[0104] Obtain the importance score RP of each relationship category u ;

[0105] Among them

[0106]

[0107] Based on the above, the concept of information entropy can be introduced to realize the evaluation of the importance of relationship categories, making the evaluation result more scientific

[0108] Embodiment 4

[0109] As shown in Figure 4 , in order to implement the importance order based on entity categories and the importance score RP of each relationship category described in Embodiment 1 u , optimize the triples in the initial knowledge graph. This embodiment provides a method for optimizing the initial knowledge graph. It includes

[0110] Obtain the importance sequence L of entity categories e and the importance score RP of each relationship category u ; among them, L e ={L j |j∈N +}, L j is the j-th entity category in the importance sequence L e , in the importance sequence L e , the importance of the j-th entity category L j is greater than the importance of the (j + 1)-th entity category L j+1 ; among them, RP u is the importance score of the u-th relationship category;

[0111] Based on the importance sequence L of entity categories e, divide the entity features in each entity category in ascending order of importance to obtain the core entity feature sequence L j1 , the key entity feature sequence L j2 and the marginal entity feature sequence L j3 ;

[0112] Based on the key entity feature sequence L j2 and the core entity feature sequence L j1 , obtain the set C of excluded entity features from the marginal entity feature sequence L j3 ;

[0113] Based on the set C of excluded entity features, obtain the set E of excluded relationship categories;

[0114] Delete the triples with entity features in the set C of entity features and relationship categories in the set E of excluded relationship categories from the initial knowledge graph to obtain the optimized triples.

[0115] Based on the above, the optimization of the knowledge graph can be better achieved. In particular, during the optimization process, the less important entity categories can be processed first, which can greatly avoid the occurrence of over-optimization and effectively avoid the exclusion of important data.

[0116] In this embodiment, the core entity feature sequence L j1 is obtained specifically as follows,

[0117] Obtain the entity features that have relationship features with the entity features in the other entity categories in each entity category L j , and use them as core entity features to construct the core entity feature sequence L j1 ; where is the α-th core entity feature in the core entity feature sequence L j1 .

[0118] Based on the above, the screening of core entity features can be better achieved.

[0119] In this embodiment, the key entity feature sequence L j2 and the marginal entity feature sequence L j3 are obtained specifically as follows,

[0120] Obtain the number of relationship features of each relationship category possessed by each core entity feature ; where represents the number of relationship features of the u-th relationship category at the core entity feature ;

[0121] ​Obtain each non-core entity feature The number of relationship features of each relationship category possessed by the location Among them, is the β-th non-core entity feature, is the non-core entity feature is the number of relationship features of the u-th relationship category at ;

[0122] Based on the importance score RP of each relationship category u , obtain the feature score Z of each core entity feature α and the feature score Z of each non-core entity feature β ; Among them, is the total number of relationship categories;

[0123] Using the minimum value MIn(Z α ) of the feature scores Z of all core entity features as the threshold, the non-core entity features with feature scores Z α not lower than the threshold MIn(Z β ) are used as key entity features, and the non-core entity features with feature scores Z α lower than the threshold MIn(Z β ) are used as marginal entity features, and complete the construction of the key entity feature sequence L α ) and the marginal entity feature sequence L j2 . j3

[0124] Based on the above, the acquisition of key entity features and marginal entity features can be better achieved.

[0125] In this embodiment, based on the key entity feature sequence L j2 and the core entity feature sequence L j1 , obtain the set C of entity features to be excluded from the marginal entity feature sequence L j3 , including,

[0126] Obtain the set D of entity features included in the shortest path from each key entity feature in the key entity feature sequence L j2 to each core entity feature in each core entity feature sequence L j1 ;

[0127] Taking the set composed of entity features that belong to the marginal entity feature sequence L j3 and do not belong to the set D of entity features as the set C of entity features to be excluded.

[0128] Based on the above, the acquisition of entity features to be excluded can be better achieved.

[0129] ​In this embodiment, obtaining the set E of eliminated relationship categories based on the set C of eliminated entity features includes:

[0130] Obtaining the occurrence times of each relationship category in the set C of eliminated entity features;

[0131] Constructing the set E of eliminated relationship categories with the relationship categories whose occurrence times are in the interval [μ - 3σ, μ + 3σ]; where μ is the mean of the occurrence times of each relationship category, and σ is the variance of the occurrence times of each relationship category.

[0132] Based on the above, the acquisition of the eliminated relationship categories can be preferably achieved.

[0133] Embodiment 5

[0134] This embodiment provides a device, which includes:

[0135] At least one processor; and

[0136] A memory communicatively connected to the at least one processor; where

[0137] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of Embodiments 1 - 4.

[0138] It is easy to understand that those skilled in the art can combine, split, reorganize, etc. the embodiments of the present application based on one or several embodiments provided by the present application to obtain other embodiments, and these embodiments do not exceed the protection scope of the present application.

[0139] The above schematically describes the present invention and its implementation manners. This description is not restrictive. What is shown in the embodiments is only part of the implementation manners of the present invention, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the spirit of the present invention, they should all fall within the protection scope of the present invention.

Claims

1. A knowledge graph optimization method applicable to network security situation awareness data, which includes: Constructing an initial knowledge graph; wherein the initial knowledge graph has multiple triples, each triple has an entity feature pair consisting of two entity features and a corresponding relationship feature for characterizing the association between the corresponding entity feature pairs, each entity feature includes an entity ontology and a corresponding entity category label, and each relationship feature has a relationship ontology and a corresponding relationship category label; Get the importance sequence L of entity categories e And the importance score RP of each relationship category u Among them, L e ={L j |j∈N + }, L j is the importance sequence L e The j-th entity category in the importance sequence L e The jth entity category L in j The importance of is greater than the j+1th entity category L j+1 The importance of RP u Score the importance of the u-th relationship category; Importance sequence L based on entity category e , divide the entity features in each entity category in order from low to high importance, and obtain the core entity feature sequence L j1 , key entity feature sequence L j2 And the edge entity feature sequence L j3 ; Based on the key entity feature sequence L j2 and core entity feature sequence L j1 , from the edge entity feature sequence L j3 Get the feature set C of the excluded entity; Based on the eliminated entity feature set C, obtain the eliminated relationship category set E; Delete the triples with entity features in the entity feature set C and the triples with relationship categories in the excluded relationship category set E from the initial knowledge graph to obtain optimized triples; Based on the optimized triples, the initial knowledge graph is reconstructed.

2. The knowledge graph optimization method applicable to network security situation awareness data according to claim 1 is characterized in that: Core entity feature sequence L j1 The specific acquisition is as follows: Get each entity category L j The entity features that have relationship features with the entity features in other entity categories are used as core entity features to construct the core entity feature sequence L j1 ;in, is the core entity feature sequence L j1 The αth core entity feature in .

3. The knowledge graph optimization method applicable to network security situation awareness data according to claim 2 is characterized in that: Key entity feature sequence L j2 And the edge entity feature sequence L j3 The specific acquisition is as follows: Get each core entity feature The number of relationship features of each relationship category that a location has in, Representing core entity features The number of relation features in the u-th relation category; Get each non-core entity feature The number of relationship features of each relationship category that a location has in, is the βth non-core entity feature, Non-core entity features The number of relation features in the u-th relation category; Based on the importance score of each relationship category RP u , get the feature score Z of each core entity feature α and the feature score Z for each non-core entity feature β ;in, is the total number of relationship categories; The feature score Z of all core entity features α The minimum value MIn(Z α ) as the threshold, and the feature score Z β Not less than the threshold value MIn(Z α ) as the key entity features, and the feature score Z β Below the threshold value MIn(Z α ) as the edge entity features to complete the key entity feature sequence L j2 And the edge entity feature sequence L j3 's construction.

4. The knowledge graph optimization method applicable to network security situation awareness data according to claim 3 is characterized in that: The key entity feature sequence L j2 and core entity feature sequence L j1 , from the edge entity feature sequence L j3 Get the feature set C of the eliminated entity, including: Get the key entity feature sequence L j2 Each key entity feature in each core entity feature sequence L j1 The entity feature set D contained in the shortest path of each core entity feature in; Taking the feature sequence L belonging to the edge entity j3 And the set of entity features that does not belong to the entity feature set D is used as the eliminated entity feature set C.

5. The knowledge graph optimization method applicable to network security situation awareness data according to claim 4 is characterized in that: The elimination relationship category set E is obtained based on the elimination entity feature set C, including: Get the number of occurrences of each relationship category in the eliminated entity feature set C; The relationship categories whose occurrence times are in the interval [μ-3σ, μ+3σ] are used to construct the eliminated relationship category set E; where μ is the mean of the occurrence times of each relationship category, and σ is the variance of the occurrence times of each relationship category.

6. The knowledge graph optimization method applicable to network security situation awareness data according to claim 1 is characterized in that: The importance sequence L of entity categories e The methods for obtaining include: Based on manual annotation, obtain the initial importance sequence L of entity categories s Among them, L s ={L i |i∈N + }, L i is the initial importance sequence L s The i-th entity category in the initial importance sequence L s The i-th entity category L i The importance of is greater than the i+1th entity category L i+1 The importance of Construct the importance index of entity category based on the initial importance sequence L s , get the importance sequence L of entity categories e Among them, L e ={L j |j∈N + }, L j is the importance sequence L e The j-th entity category in the importance sequence L e The jth entity category L in j The importance of is greater than the j+1th entity category L j+1 The importance of.

7. The knowledge graph optimization method applicable to network security situation awareness data according to claim 1 is characterized in that: The importance score RP of each relationship category u The methods for obtaining include: Based on the entity category labels and relationship category labels, get the total number T of each relationship category in each entity category uv , construct the association matrix T between relationship categories and entity categories; where T uv represents the total number of times the uth relationship category appears in the vth entity category, T = (T uv ) H*N , H is the total number of relationship categories, N is the total number of entity categories; Based on the correlation matrix T, obtain the importance score RP of each relationship category u ; Among them, RP u Score the importance of the u-th relationship category.

8. The knowledge graph optimization method applicable to network security situation awareness data according to claim 7 is characterized in that: Based on the correlation matrix T, the importance score RP of each relationship category is obtained. u ,include, Get the weight LP of each entity category in each relationship category uv ; Get the information entropy LH of each entity category v ; Get the weight LW of each entity category v ; Get the initial importance score for each relationship category Get the importance score RP of each relationship category u ; in, 9. An apparatus comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Voice interaction task execution method and device based on large model, equipment and medium

    CN120472906A

  • Intelligent data resource table entering system based on knowledge graph

    CN120596486A

  • Knowledge Graph-Based Intelligent Data Resource Entry System

    CN120596486B