Entity alignment method and device in security and secrecy scene
By constructing entity relationship diagrams in secure and confidential scenarios and using entity alignment methods of attention mechanisms and gating mechanisms, the knowledge redundancy and ambiguity of multi-source heterogeneous data are solved, and more efficient entity alignment and data fusion are achieved, improving the accuracy and security of secure and confidential data.
Patent Information
- Application Number
- CN202510430610.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
In the security and confidential scenarios, the entity alignment method has knowledge redundancy and ambiguity, and the method based on knowledge graph embedding and graph neural network is not effective in multi-source heterogeneous data fusion, making it difficult to effectively solve the structural differences and neighborhood characteristics problems of different knowledge graphs.
The entity alignment method based on attention mechanism and gating mechanism is adopted. By constructing an entity relationship diagram, all samples are sampled on one-hop neighbors, and locally samples are performed on two-hop neighbors or above. A neighborhood aggregation matching network is used to generate matching-oriented entity representations, and a distance function is used to predict entity alignment.
It has achieved improvements in the depth, accuracy and security of security and confidential knowledge in security and confidentiality scenarios, and can more accurately analyze and integrate multi-source heterogeneous data, reduce knowledge redundancy and ambiguity, and improve the accuracy of entity alignment.
Smart Images

Figure CN120354429A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of data semantic matching, and particularly relates to an entity alignment method and device in a security and confidentiality scenario. Background Art
[0002] With the development of information technology, requirements and cognitions are always in constant change. Therefore, the logic and data patterns of security and confidentiality knowledge inevitably change frequently, and different groups may have different interpretations of the same knowledge, resulting in phenomena such as ambiguity, conflict, error, and redundancy. This kind of change requires entity alignment at the entity layer to achieve conflict resolution of security and confidentiality knowledge. Entity alignment at the entity layer refers to the fusion of entity and relationship (including attribute) tuples, aiming to determine whether entities from different semantic knowledge bases point to the same object in the real world.
[0003] Specifically, since the input data used in the security and confidentiality semantic knowledge base comes from different data sources, when performing collaborative fusion of knowledge graphs, the problem of entity alignment needs to be considered, and various types of data that have been converted into triples are fused. For example, in the cyber space, the fusion of asset and topology data is mainly achieved by linking through IP addresses, linking the IP fields of network assets and topology data, so as to fuse asset data and topology data. For network assets and network security data, make full use of information such as firmware version and operating system version in security data to connect with asset data to form the has attribute of network assets. In fact, this fusion method benefits from our classification of various data into different levels according to their fields. First, subgraphs are constructed within each level, and then cross-graph fusion is carried out by using the association features between levels.
[0004] From the perspective of the existing technology, the methods for entity alignment include techniques based on knowledge graph embedding and techniques based on graph neural networks. The entity alignment method based on knowledge graph embedding requires a sufficient number of seed sequences and is usually affected by the incompleteness and heterogeneity between different knowledge graphs, resulting in relatively low alignment accuracy. Moreover, the existing entity alignment methods based on graph neural networks still have the problem that there are structural differences between different knowledge graphs, resulting in different neighborhood structures between corresponding entities, so that the information of entities in knowledge graphs from different sources cannot be fully learned, resulting in poor entity alignment effects. Therefore, the current methods still have corresponding defects and problems in solving the fusion of multi-source heterogeneous data in the security and confidentiality scenario, and it is necessary to consider the heterogeneity and neighborhood features between different knowledge structures to improve the alignment and fusion effect of security and confidentiality data. Summary of the Invention
[0005] The purpose of this application is to disclose an entity alignment method and device in a secure and confidential scenario. By the method of this application, knowledge redundancy and ambiguity are eliminated, strict requirements for the depth, accuracy, and security of secure and confidential knowledge are met, and collaboration based on semantic knowledge in all links of the secure and confidential system is achieved, so as to overcome the problems of the prior art.
[0006] The purpose of this application is achieved through the following technical solutions:
[0007] An entity alignment method in a secure and confidential scenario, the entity alignment method comprising:
[0008] S1: Construct an entity relationship graph based on the secure and confidential knowledge system from the set of entity relationship triples obtained through extraction. Sample all one-hop neighbors, and for two-hop and above neighbors, use the attention mechanism for local sampling;
[0009] S2: Introduce a gating mechanism for aggregation to learn the representation of the graph structure. Construct a neighborhood local subgraph for each entity for neighborhood matching, and at the same time jointly encode the output of the matching stage and the graph structure representation to generate an entity representation for matching;
[0010] S3: Use a similarity metric based on a distance function to achieve entity alignment prediction.
[0011] According to a preferred embodiment, in step S1, one-hop neighbors of an entity are considered the most important neighborhoods, and vanilla GCN is used to aggregate the neighbor information of the entity to learn the semantic knowledge base structure embedding.
[0012] According to a preferred embodiment, in step S1, during the process of sampling all one-hop neighbors, first use the pre-trained word embedding method to initialize vanilla GCN, and at the same time introduce the highway networks method to avoid the propagation of noise between GNN layers.
[0013] According to a preferred embodiment, in step S1, the attention mechanism is used to discover distant neighbor entities and find two-hop and above neighbor entities that have an impact on the features and meanings of the current entity. Specifically, in the calculation process, two matrices are used to perform linear transformation on the central entity and neighbors, and normalization processing is performed, so that entities can be analyzed and compared more accurately.
[0014] According to a preferred embodiment, step S2 includes: sampling the one-hop neighbors, two-hop and above neighbors of the central entity to form a domain local subgraph. Consider the entity with the largest amount of information for the central entity during the sampling process, and then use the gating mechanism to aggregate the neighborhood information to form an embedded representation of the entity.
[0015] According to a preferred embodiment, in step S2, a neighborhood aggregation matching network (MANN) model is used for domain matching. The MANN model extracts a distinguishable neighborhood for each entity and constructs a neighborhood local subgraph for cross-graph neighborhood matching.
[0016] According to a preferred embodiment, in step S2, during the neighborhood matching process, the selection of candidate entities is carried out first. The MANN model calculates the similarity between the entities in semantic knowledge base 1 and all entities in semantic knowledge base 2 in their representation spaces, finds the entity in semantic knowledge base 2 that is closest in the embedding space, and uses it as the candidate entity.
[0017] In the neighborhood matching module of the MANN model, semantic knowledge base 1 and semantic knowledge base 2 are superimposed to form a large input graph, and a matching vector is introduced to calculate the matching degree between the entity neighborhoods in semantic knowledge base 1 and all entities in semantic knowledge base 2.
[0018] Then, the output embedding of the neighbor is combined with the matching vector, and the sampled neighbor representation set is aggregated. Thus, the hidden representation of the calculated entity and its neighbor representation are connected to generate the final entity representation for matching.
[0019] According to a preferred embodiment, step S3 includes: calculating the similarity of entities in different security and confidentiality domains based on the entity representation for matching obtained in step S2 and the vector representation of the entity calculated according to the distance function.
[0020] Among them, similar or aligned entities have a distance less than the threshold in the vector space, while dissimilar or unaligned entities have a distance greater than the threshold in the vector space.
[0021] On the other hand, the present application also discloses:
[0022] An entity alignment device in a security and confidentiality scenario, the entity alignment device includes a data processing unit, and the data processing unit is configured to perform data processing according to the foregoing entity alignment method.
[0023] The main solution of the present application and its various further alternative solutions can be freely combined to form multiple solutions, all of which are solutions that can be adopted and claimed by the present application. Those skilled in the art can understand that there are various combinations according to the prior art and common general knowledge after understanding the solution of the present application, all of which are the technical solutions to be protected by the present application, and will not be enumerated here.
[0024] Advantages of the present application:
[0025] The present application has multiple significant advantages in solving the data fusion problem in a security and confidentiality scenario.
[0026] First, the true correlation in security and confidentiality incidents is usually deeply hidden. The method for discovering distant domain entities based on the attention mechanism proposed in this application can analyze and explore security and confidentiality entities with a relatively long relationship distance, find distant neighbor entities that have a greater impact on the central entity, so as to find relevant entities with deep influence from multi-source heterogeneous security and confidentiality data;
[0027] Secondly, this application constructs a neighborhood local subgraph through the closest one-hop neighbor entity and multiple distant neighbor entities with important influence, realizing the aggregation and utilization of security and confidentiality entity information at different levels and distances. The vector representation of the central entity calculated through the neighborhood local subgraph can more fully represent its true meaning; on this basis, calculating the distance-based vectors of entities to determine whether the entities are aligned has higher accuracy, realizing more accurate and efficient fusion analysis of security and confidentiality data. Description of the Drawings
[0028] Figure 1 is an example of distant neighbor selection in this application. Specific Embodiments
[0029] The following uses specific specific examples to illustrate the implementation manners of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0030] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0031] In the description of this application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of this application is usually placed when in use. It is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to this application. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0032] In addition, terms such as "horizontal", "vertical", "hanging", etc. do not mean that the components are required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined.
[0033] In the description of the present application, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0034] In addition, the present application points out that in the present application, if the specifically involved structure, connection relationship, position relationship, power source relationship, etc. are not specifically written, the structure, connection relationship, position relationship, power source relationship, etc. involved in the present application are all those that can be known by those skilled in the art on the basis of the prior art without creative labor.
[0035] Embodiment 1:
[0036] The present application discloses an entity alignment method in a secure and confidential scenario. As shown in the appendix Figure 1 , for two central entity pairs (a, A), their one-hop neighbors are different, only including peer entity pairs (b, B) and (c, C), while the one-hop neighbor d of a corresponds to the long-distance neighbor D of A, and the one-hop neighbors E and F of A correspond to the long-distance neighbors e and f of a. Therefore, the long-distance neighbors e and f can be included in the neighborhood aggregation of a, and the long-distance neighbor D of A is taken into account. On this basis, the graph neural network will learn more representations about the similarity between a and A. However, not all long-distance neighbors contribute positively to the similar representation. Therefore, the present invention uses an attention mechanism to find the long-distance neighborhood entities that contribute more to the central entity. In the entity alignment stage, the output embedding of the neighborhood entities is jointly encoded with the graph structure representation learned through the gating mechanism to generate entity representations for matching. Finally, a distance function is used to perform entity alignment prediction on the entity representations for matching.
[0037] Specifically, the entity alignment method includes the following steps.
[0038] Step S1: Construct an entity relationship graph from the set of entity relationship triples obtained through extraction according to the secure and confidential knowledge system, sample all one-hop neighbors, and use the attention mechanism for local sampling for two-hop and above neighbors.
[0039] In step S1, the one-hop neighbors of the entity are considered to be the most important neighborhoods, and vanilla GCN is used to aggregate the neighbor information of the entity to learn the semantic knowledge base structure embedding. In the process of sampling all the one-hop neighbors, the pre-trained word embedding method is first used to initialize the vanilla GCN, and the highway networks method is introduced to avoid noise propagation between GNN layers.
[0040] Therefore, in order to find the distant neighbors that contribute positively to the central entity, an attention mechanism is used to calculate the two-hop neighbor information of the entity, and a gating mechanism is used to further aggregate the neighborhood information to mine the hidden representation of the entity. Although the graph attention network uses a shared linear transformation in the entity, it ignores that the central entity and the neighbors may be completely different, and this shared transformation may lead to an inability to correctly distinguish them. Therefore, the present invention uses two matrices to perform linear transformations on the central entity and the neighbors, respectively, and performs normalization processing to make them comparable between different entities.
[0041] That is, in step S1, the attention mechanism is used to discover distant neighbor entities and find neighbor entities of two hops or more that are effective for the characteristics and meaning of the current entity. Specifically, two matrices are used to perform linear transformation on the central entity and neighbors during the calculation process, and normalization is performed to enable more accurate analysis and comparison between entities.
[0042] The real correlation relationship in security and confidentiality events is usually hidden deeply. The long-distance domain entity discovery method based on the attention mechanism proposed in this application can analyze and explore security and confidentiality entities with a long relationship distance, and find the long-distance neighbor entities that have a greater impact on the central entity, thereby finding related entities with deep influence from multi-source heterogeneous security and confidentiality data.
[0043] Step S2: Introduce a gating mechanism for aggregation to learn the representation of the graph structure, build a neighborhood local subgraph for each entity for neighborhood matching, and jointly encode the output of the matching stage with the graph structure representation to generate a matching-oriented entity representation.
[0044] Preferably, step S2 includes: sampling one-hop neighbors, two-hop and above neighbors of the central entity to form a domain local subgraph, considering the entity with the largest amount of information about the central entity during the sampling process, and then using a gating mechanism to aggregate the neighborhood information to form an embedded representation of the entity.
[0045] Preferably, in step S2, a neighborhood aggregation matching network MANN model is used for domain matching, and the neighborhood aggregation matching network MANN model extracts a distinguishable neighborhood for each entity and constructs a neighborhood local subgraph for cross-graph neighborhood matching.
[0046] The key to this application is to reduce the heterogeneity of the neighborhood structures of entities in different knowledge graphs. Therefore, the neighborhood aggregation matching network MANN model is used to encode the graph structure information from the perspective of the entity neighborhood to achieve entity alignment and alleviate the impact of structural heterogeneity. NAMN adopts a hierarchical idea. First, all one-hop neighbors are sampled, and for two-hop and above neighbors, an attention mechanism is used for local sampling. Subsequently, a gating mechanism is introduced to aggregate the k-hop neighbor information of the entity to mine the hidden information of the graph structure. In order to eliminate the impact of the one-hop neighbor structure heterogeneity of the entity, the model extracts a distinguishable neighborhood for each entity and constructs a local subgraph of the neighborhood for cross-graph neighborhood matching.
[0047] In order to determine whether an entity is aligned with other entities, the neighbor entities of the entity are key. However, not all neighbor entities have a positive impact on entity alignment. Therefore, this application introduces a local subgraph to select the neighbor entity with the most information about the central entity, and adopts a down-sampling process to select suitable neighbors. If the entity pair has a relationship, a directed edge is added to its corresponding node in the local subgraph, but only the direction of the relationship is retained. In order to select suitable neighbors, a neighbor sampling strategy is adopted.
[0048] Specifically, in step S2, after determining the neighbors that should be considered for the central entity, a neighborhood local subgraph is generated. In the neighborhood matching process, the first step is to select candidate entities. The neighborhood aggregation matching network MANN model calculates the similarity between the entities in the semantic knowledge base 1 and all entities in the semantic knowledge base 2 in their representation space, finds the closest entity in the embedding space of the semantic knowledge base 2, and uses it as a candidate entity; in the neighborhood matching module of the neighborhood aggregation matching network MANN model, the semantic knowledge base 1 and the semantic knowledge base 2 are superimposed to form a large input graph, and a matching vector is introduced to calculate the matching degree between the entity neighborhood in the semantic knowledge base 1 and all entities in the semantic knowledge base 2; then, the output embedding of the neighbors is combined with the matching vector, and the sampled neighbor representation set is summarized; thus, the calculated hidden representation of the entity and its neighbor representation are connected to generate the final matching-oriented entity representation.
[0049] The goal of this application is to complete the entity alignment task by minimizing the distance between aligned entities, while unaligned entities should be represented as negative samples with larger distances. During the optimization process, the Adam optimizer is used to optimize the target, and Xavier initialization is used to initialize all learnable parameters (including the input feature vector of the entity).
[0050] This application constructs a neighborhood local subgraph through the closest one-hop neighbor entity and multiple influential distant neighborhood entities, realizing the aggregation and utilization of secure and confidential entity information at different levels and distances. The vector representation of the central entity obtained through the neighborhood local subgraph can more fully represent its true meaning. On this basis, calculating the vectors of entities based on distance can determine whether the entities are aligned with higher accuracy, achieving more accurate and efficient fusion analysis of secure and confidential data.
[0051] Step S3: Use similarity measurement based on a distance function to implement entity alignment prediction.
[0052] Preferably, step S3 includes: Based on the entity representation for matching obtained in step S2, calculate the vector representation of the entity according to the distance function, and calculate the similarity of entities in different secure and confidential domains. Entities that are similar or aligned have a distance less than the threshold in the vector space, while entities that are not similar or not aligned have a distance greater than the threshold in the vector space.
[0053] Embodiment 2
[0054] Based on Embodiment 1, this embodiment also discloses: An entity alignment device in a secure and confidential scenario, where the entity alignment device includes a data processing unit configured to process data according to the entity alignment method of Embodiment 1.
[0055] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.
Claims
1. An entity alignment method in a secure and confidential scenario, characterized in that The entity alignment method includes: S1: Construct an entity relationship graph based on the security and confidentiality knowledge system from the set of entity relationship triples obtained through extraction. Sample all one-hop neighbors, and for two-hop and above neighbors, use the attention mechanism for local sampling; S2: Introduce a gating mechanism for aggregation to learn the representation of the graph structure. Construct a neighborhood local subgraph for each entity for neighborhood matching, and at the same time jointly encode the output of the matching stage and the graph structure representation to generate an entity representation oriented to matching; S3: Use a similarity metric based on a distance function to achieve entity alignment prediction.
2. The entity alignment method according to claim 1, wherein In step S1, the one-hop neighbors of an entity are considered the most important neighborhoods, and vanilla GCN is used to aggregate the neighbor information of the entity to learn the semantic knowledge base structure embedding.
3. The entity alignment method according to claim 2, wherein In step S1, during the process of sampling all one-hop neighbors, First, use the pre-trained word embedding method to initialize vanilla GCN, and at the same time introduce the highway networks method to avoid the propagation of noise between GNN layers.
4. The entity alignment method according to claim 3, wherein In step S1, use the attention mechanism to discover distant neighbor entities and find neighbor entities of two-hop and above that have an impact on the features and meanings of the current entity, Specifically, in the calculation process, two matrices are used to perform linear transformation on the central entity and neighbors, and normalization processing is performed to enable more accurate analysis and comparison between entities.
5. The entity alignment method according to claim 1, wherein Step S2 includes: Sampling the one-hop neighbors, two-hop and above neighbors of the central entity to form a domain local subgraph. Consider the entity with the largest amount of information for the central entity during the sampling process, and then use the gating mechanism to aggregate the neighborhood information to form the embedded representation of the entity.
6. The entity alignment method according to claim 5, wherein In step S2, use the neighborhood aggregation matching network MANN model for domain matching. The neighborhood aggregation matching network MANN model extracts a distinguishable neighborhood for each entity and constructs a neighborhood local subgraph for cross-graph neighborhood matching.
7. The entity alignment method according to claim 6, wherein In step S2, during the neighborhood matching process, the selection of candidate entities is carried out first. The neighborhood aggregation matching network MANN model calculates the similarity between the entities in semantic knowledge base 1 and all entities in semantic knowledge base 2 in their representation space, finds the entity closest in the embedding space in semantic knowledge base 2, and uses it as the candidate entity; In the neighborhood matching module of the neighborhood aggregation matching network MANN model, semantic knowledge base 1 and semantic knowledge base 2 are superimposed to form a large input graph, and a matching vector is introduced to calculate the matching degree between the entity neighborhood in semantic knowledge base 1 and all entities in semantic knowledge base 2; Then, combine the output embedding of the neighbor with the matching vector and summarize the set of sampled neighbor representations; thus, connect the calculated hidden representation of the entity and its neighbor representation to generate the final entity representation oriented to matching.
8. The entity alignment method according to claim 1, characterized in that Step S3 includes: Based on the entity representation oriented to matching obtained in step S2, calculate the vector representation of the entity according to the distance function, and calculate the similarity of entities in different security and confidentiality domains. Among them, similar or aligned entities have a distance less than the threshold in the vector space, while dissimilar or unaligned entities have a distance greater than the threshold in the vector space.
9. An entity alignment device in a secure and confidential scenario, characterized in that, The entity alignment device includes a data processing unit, The data processing unit is configured to perform data processing according to the entity alignment method described in any one of claims 1 to 8.