Dynamic knowledge graph construction method for three-dimensional security situation analysis
By analyzing the entity's own and inherited attributes and calculating the intersection degree and structural similarity, the problem of misjudgment in traditional entity alignment algorithms is solved, and the entity alignment accuracy and application effectiveness of dynamic knowledge graphs are improved.
Patent Information
- Application Number
- CN202511156586.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Traditional entity alignment algorithms are prone to misjudgment in knowledge graphs, ignoring the conceptual relationships and attribute inheritance characteristics between entities, resulting in low accuracy of alignment results and affecting the application effectiveness of dynamic knowledge graphs.
By analyzing the entity's own attributes and inherited attributes, calculating the first cross-degree and the second cross-degree, combining the structural approximation, and using the association coefficient to evaluate the entity alignment of the target node, a dynamic knowledge graph is constructed.
The accuracy of entity alignment is improved, the application effectiveness of dynamic knowledge graphs is enhanced, and the degree of association between entities is more accurately reflected by considering the attributes and structural relationships of entities.
Smart Images

Figure CN120706524A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of knowledge graph construction, and specifically to a method for constructing a dynamic knowledge graph for three-dimensional security situation analysis. Background Art
[0002] Knowledge graph technology, broadly defined, encompasses the entire knowledge formation process, from extracting, integrating, evaluating, reasoning, and storing knowledge resources. Ultimately, a visual knowledge graph is constructed. Because the relationships between entities in a knowledge graph change over time, dynamically evolving knowledge graphs can be used to dynamically perceive the security status of knowledge and mitigate the risks of inferring knowledge. Entity alignment in dynamic knowledge graphs is a key research area in the knowledge graph field.
[0003] Traditional entity alignment algorithms determine whether they are the same entity by performing string matching on the entity names. This method is prone to misjudgment and incorrectly aligns entities of different concepts. It also ignores the broad or narrow relationships between the concepts represented by different entities, does not consider the inheritance characteristics of attributes between entities, and the structural characteristics of entities in different knowledge graphs. As a result, when aligning entities in different knowledge graphs, the alignment results are less accurate, affecting the application effectiveness of dynamic knowledge graphs. Summary of the Invention
[0004] In order to solve the above technical problems, a dynamic knowledge graph construction method for three-dimensional security situation analysis is provided to solve the existing problems.
[0005] The solution to the technical problem in this application is to provide a method for constructing a dynamic knowledge graph for three-dimensional security situation analysis, including the following steps: Extract entities, relationships between entities, attributes, and time from texts of multiple data sources to construct multiple knowledge graphs. The attributes corresponding to each node in each knowledge graph are recorded as its own attributes. For all entities corresponding to all nodes in the knowledge graph, analyze the differences between the different attributes corresponding to any two entities, and determine the first intersection degree of the two entities based on the distribution of the attributes of the entities in the corresponding text; Based on the connection relationship between each node in the same knowledge graph and the distribution of the first cross degree between the connected nodes, the node to which each node belongs and the node at the same level are determined; based on the own attributes of the node to which it belongs, the inherited attributes of each node are generated; Select a node from each of the two knowledge graphs and record it as a target node; calculate a correction factor based on the difference between the inherited attribute of each target node and the attribute of the other target node, and determine the second cross-degree of the two target nodes in combination with the first cross-degree; Determine the structural similarity of the two target nodes based on the similarity between the target node and its sibling nodes and the difference in the number of sibling nodes; Based on the second intersection degree and the structural similarity, the correlation coefficient of the two target nodes is determined, the entities corresponding to the target nodes are evaluated, and the entities are aligned to construct a dynamic knowledge graph.
[0006] Preferably, determining the first intersection degree of any two entities includes: Analyze the distance between each entity and its corresponding attributes in the text, as well as the frequency of text containing each entity and its corresponding attributes, and calculate the semantic influence; Select the entity with the smallest number of all its own attributes among the two entities and record it as the target entity; Calculating the minimum distance between each self-attribute corresponding to the target entity and all self-attributes corresponding to the other entity in the arbitrary two entities; The first intersection degree is the sum of the ratios of the semantic influence of all the attributes corresponding to the target entity of the arbitrary two entities to the minimum distance.
[0007] Preferably, the calculating of semantic influence includes: The extracted texts of entities corresponding to all nodes in each knowledge graph are combined into a text set; Selecting texts that contain all entities and their own attributes from the text set and recording them as target texts; Calculate the character distance between each entity in the target text and its corresponding attribute; Counting the frequency of occurrence of target text in the text set; Calculate the sum of the distances between each entity and its corresponding attributes in all target texts, and record it as relative spacing; The semantic influence is a ratio of the frequency to the relative distance.
[0008] Preferably, determining the node to which each node belongs and the node at the same level includes: All other nodes connected to each node in each knowledge graph are recorded as connected nodes; Clustering the first intersection degree between each node and all connected nodes to obtain two clusters; The average value of all the first cross-degrees in each cluster is calculated, and the connected node in the cluster with the largest average value is recorded as the belonging node; otherwise, it is recorded as the peer node.
[0009] Preferably, the inherited attributes of each node are attributes that do not exist in the own attributes of each node and are inherited from the own attributes of all corresponding nodes.
[0010] Preferably, the calculation of the correction factor includes: Record all the own attributes and inherited attributes of each target node as target attributes; The reciprocal of the distance between each inherited attribute corresponding to any one of the two target nodes and each target attribute of the other target node is recorded as the matching degree; Among the two target nodes, selecting the maximum matching degree between each inherited attribute corresponding to any one target node and all target attributes of the other target node; and taking the sum of the maximum matching degrees of all inherited attributes corresponding to any one target node as the association degree of any one target node; The correction factor is the sum of the association degrees between the two target nodes.
[0011] Preferably, the second cross-degree is the product of the first cross-degree between the two target nodes and the correction factor.
[0012] Preferably, the relationship is represented by a quadruple between nodes in the knowledge graph as follows: , s is the head entity, p is the relationship, o is the tail entity, t is The time when it is established, where E represents the entity set, R represents the relationship set, and T represents the time interval set. The further measurement process of the relationship between the target node and its sibling nodes is: Based on the relationship between each node and its sibling nodes in each knowledge graph, the quadruple between each node and its sibling nodes is extracted, and the quadruple is converted into a word vector through a pre-trained model.
[0013] Preferably, determining the structural similarity of the two target nodes includes: Select the target node with the smallest number of peer nodes among the two target nodes and record it as the key node; Select the maximum value of the correlation between each word vector corresponding to the key node of the two target nodes and all word vectors corresponding to the other target node, and record it as the maximum correlation; Calculating the cumulative sum of the maximum relevance of all word vectors corresponding to the key node; Calculate the difference between the number of all peer nodes corresponding to the key node and the other target node in the two target nodes, and record it as the number difference; The structural similarity is the ratio of the cumulative sum to the quantitative difference.
[0014] Preferably, determining the correlation coefficient between the two target nodes and evaluating the entities corresponding to the target nodes includes: Normalizing the product of the structural similarity and the second cross degree as the correlation coefficient between the two target nodes; If the correlation coefficient is greater than or equal to a preset threshold, the entities corresponding to the two target nodes are the same entity; otherwise, the entities corresponding to the two target nodes are not the same entity.
[0015] This application has at least the following beneficial effects: This application calculates the semantic influence of each entity's own attributes through the distribution of entities and their own attributes in the text in each data source. Its beneficial effect is that it takes into account the information representation of the entity by each entity's own attributes. Secondly, combined with the difference between the different own attributes corresponding to any two entities, the first intersection degree of the any two entities is calculated. Its beneficial effect is that it takes into account the correlation of the information contained in the own attributes of the two entities to reflect the association between the two entities, and then explains the possibility that the two entities represent the same thing or concept; through the first intersection between each node in each knowledge graph and the rest of the nodes connected to it, the rest of the connected nodes are screened to determine the belonging nodes and the same-level nodes of each node. Its beneficial effect is that it takes into account the belonging nodes that may belong to the same category of concepts in a broad or narrow sense with each node, reflecting the relationship that each node and its own node may belong to the parent class and child class, so that each node inherits the own attributes of the belonging node and obtains the inherited attributes; further, through the difference in inherited attributes between two target nodes in any two knowledge graphs, the correction factor is calculated to determine Determine the second intersection degree of the two target nodes, further illustrate the association between the two entities, and reflect the possibility that the entities corresponding to the two target nodes represent the same thing or concept, so that the two target nodes not only consider the correlation of the information contained in their own attributes, but also consider the association of inherited attributes, and more accurately reflect the degree of association between the entities corresponding to the two target nodes; through the four-tuple relationship represented by the target node and its node at the same level, analyze the word vectors corresponding to the four-tuple between the two target nodes, and calculate the structural similarity of the two target nodes. The beneficial effect is that it considers the structural similarity in the knowledge graph where the two target nodes are located to reflect the possibility that the two target nodes represent the same entity; determine the association coefficient of the two target nodes, evaluate the entities corresponding to the target nodes, and perform entity alignment to construct a dynamic knowledge graph. The beneficial effect is that it characterizes the possibility that the two target nodes are the same entity through the association coefficient, while considering the inheritance of attributes between nodes and the structural difference characteristics of nodes in different knowledge graphs, thereby improving the accuracy of entity alignment between different knowledge graphs. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The following is a further detailed description of the method for constructing a dynamic knowledge graph for three-dimensional security situation analysis of the present application in conjunction with the accompanying drawings.
[0017] Figure 1 A flowchart of the steps of a method for constructing a dynamic knowledge graph for three-dimensional security situation analysis provided in an embodiment of the present application; Figure 2 A flowchart of the steps of the method for obtaining the correlation coefficient provided in an embodiment of the present application. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the following, in conjunction with the accompanying drawings and implementation examples, further describes in detail the method for constructing a dynamic knowledge graph for three-dimensional security situation analysis proposed in this application. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0019] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0020] See also Figure 1 , which shows a flowchart of a method for constructing a dynamic knowledge graph for three-dimensional security situation analysis provided by an embodiment of the present application, the method comprising the following steps: Step 1: Extract entities, relationships between entities, attributes, and time from texts of multiple data sources to construct multiple knowledge graphs. The attributes corresponding to each node in each knowledge graph are recorded as its own attributes.
[0021] A dynamic knowledge graph is generally constructed by adding time attributes to the triples of a static knowledge graph. Therefore, a dynamic knowledge graph is usually in the form of a quadruple or quintuple.
[0022] In static knowledge graphs, entity alignment is the process of determining whether two entities in the same or different knowledge graphs refer to the same real-world object. Specifically, the same entity often has different representations and conceptual hierarchies in different encyclopedias, such as full names versus abbreviations, jargon versus professional terms, and original names versus nicknames. Therefore, in dynamic knowledge graphs, entity alignment is the process of identifying the same entity with different descriptions or representations at different times.
[0023] Use Scrapy to obtain data sets from different data sources, perform data preprocessing on the data sets from different data sources, and use GPT-3.5-turbo to extract entities, relationships between entities, and attributes and time information of entities from the text in the data sets of different data sources, and then build a knowledge graph. Then, each data set of the data source corresponds to a knowledge graph, and thus multiple knowledge graphs are obtained. The knowledge graph is composed of nodes corresponding to multiple entities. Therefore, each entity corresponds to a node, and each node corresponds to multiple attributes. All attributes corresponding to each node are recorded as the own attributes of each node; In this embodiment, Scrapy is used to obtain data sets from two data sources, Baidu Encyclopedia and Wikipedia. Secondly, the data preprocessing process includes removing escape characters and statements containing invalid characters; writing regular rules to unify the disambiguation name format to standardize the data; extracting entry names and their disambiguation names and synonyms and storing them in MySQL for entity linking; extracting entry attribute lists from the entry information box and storing them in MySQL for importing into the graph database and constructing relationships to extract supervision data. It should be noted that the data preprocessing process and the construction of the knowledge graph are well-known technologies and will not be repeated here.
[0024] It should be noted that for an entity representing a person's name, its own attributes may include "date of birth", "nationality", and "occupation"; for an entity representing a "company", its own attributes may include "date of establishment", "headquarters location", "number of employees", and so on.
[0025] At this point, multiple knowledge graphs are obtained.
[0026] Step 2: For the entities corresponding to all nodes in all knowledge graphs, analyze the differences between the different self-attributes corresponding to any two entities, and determine the first intersection degree of the any two entities based on the distribution of each self-attribute of the entity in the corresponding text.
[0027] Since different knowledge graphs are constructed from data sets from different data sources, entity alignment can merge relevant information from different sources to form a more complete data set, which helps improve the quality and comprehensiveness of the data, and also improves the effectiveness of information retrieval. Secondly, entity alignment, also known as knowledge fusion, is the merging of two knowledge graphs, and is one of the key steps in the construction of dynamic knowledge graphs. How to fuse descriptive information about the same entity or concept from multiple sources is an issue that needs attention. For example, the descriptive information of "Tang Sanzang", "Xuanzang", and "Jin Chanzi" is actually the information of the same entity, but has different expressions in different knowledge graphs. Therefore, entity alignment is needed to match the same entity with different descriptions or representations.
[0028] First, depending on the texts from which the entity is extracted, the semantic influence of each attribute in different texts is often different. We analyze the distance relationship between each entity and its corresponding attributes in the corresponding text, as well as the frequency of occurrence of each attribute in different texts, and calculate the semantic influence. Specifically, The extracted texts of entities represented by all nodes in each knowledge graph are combined into a text set; Selecting texts that contain all entities and their own attributes from the text set and recording them as target texts; Counting the frequency of occurrence of target text in the text set; Calculate the character distance between each entity in the target text and its corresponding attribute; Calculate the sum of the distances between each entity and its own attributes in all target texts, and record it as relative spacing; The ratio of the frequency to the relative distance is used as the semantic influence of each attribute corresponding to each entity; It should be noted that the smaller the relative spacing is, the more likely the attribute itself is to directly describe the entity's information; the greater the semantic influence is, the richer the information contained in the text of the attribute itself of the corresponding entity is, and the more likely it is to represent the characteristic information of the entity; conversely, the less information contained in the text of the attribute itself of the corresponding entity is, and the less likely it is to represent the characteristic information of the entity.
[0029] Secondly, by combining the matching of the attributes of different entities with the semantic influence, the first intersection degree is calculated to evaluate the association of the description information of different entities and reflect the possibility that the entities represent the same thing or concept. Specifically: For the entities corresponding to all nodes in all knowledge graphs, take any two entities as an example, where the two entities are recorded as and , select the entity With entity The entity with the smallest number of attributes is recorded as the target entity; assuming that the entity The number of corresponding attributes is small, so the entity Recorded as target entity; Calculate target entity Each corresponding attribute and entity The minimum distance between all corresponding self attributes; In this embodiment, the distance is calculated by Each corresponding attribute and entity The edit distance between the corresponding attributes, select the target entity Each corresponding attribute and entity The minimum value of the edit distance between all corresponding attributes; wherein, the calculation of the edit distance is a well-known technology and will not be repeated here.
[0030] Calculate target entity The ratio of the semantic influence of each corresponding attribute to the minimum distance is recorded as the relative ratio; The target entity The sum of the relative ratios of all corresponding self-attributes is used as the target entity With entity The first degree of intersection between It should be noted that the smaller the minimum distance is, the more closely the two corresponding attributes match and the more identical the semantic information they contain is. The larger the first intersection is, the greater the correlation between the description information of the two entities is, and the more likely the two entities represent the same thing or concept. Conversely, the smaller the correlation between the description information of the two entities is, the more likely the two entities do not represent the same thing or concept.
[0031] It should be understood that the first cross-degree between any two entities is the first cross-degree calculated between the entities corresponding to any two nodes in the same knowledge graph, and the first cross-degree between the entities corresponding to each node in each knowledge graph and the entities corresponding to each node in another knowledge graph.
[0032] At this point, the first intersection degree of any two entities is obtained.
[0033] Step 3: Based on the connection relationship of each node in the same knowledge graph and the distribution of the first cross-degree between the connected nodes, determine the node to which each node belongs and the node at the same level; based on the own attributes of the node to which it belongs, generate the inherited attributes of each node; select a node from each of any two knowledge graphs and record it as the target node; calculate the correction factor based on the difference between the inherited attributes of each target node in the two target nodes and the attributes of the other target node, and determine the second cross-degree of the two target nodes in combination with the first cross-degree.
[0034] Furthermore, in each knowledge graph, each entity corresponds to a node. There is a relationship between the entities corresponding to two nodes, which can form a quadruple. The quadruple represents the relationship between nodes in the knowledge graph. The quadruple includes "subject", "predicate", "object" and "time". For example, (Xiao Ming, lives in Beijing, 2020-2023), then each node in the knowledge graph can be extracted to the corresponding quadruple. The quadruple in the knowledge graph is represented as , s is the head entity, p is the relationship, o is the tail entity, t is The time when it is established, where E represents the entity set, R represents the relationship set, T represents the time interval set, and the four-tuple represents that entity s has a relationship p with object entity o within the time interval t.
[0035] Secondly, the first intersection degree can effectively reflect whether the attributes of two entities are cross-identical through the attribute information contained in the entity. If the first intersection degree is directly used for entity alignment, it is easy to cause two entities with synonymous features to be misaligned, resulting in subsequent mismatches. Therefore, by analyzing the structural relationship between different entities in each knowledge graph, the belonging nodes and peer nodes are obtained, specifically: In each knowledge graph Take the node as an example, and compare the nodes in each knowledge graph with the All other nodes connected to the node are recorded as connected nodes; For the first Clustering the first intersection degree between a node and all its connected nodes to obtain two clusters; In this embodiment, the k-means clustering algorithm is used for clustering to obtain two cluster clusters. The k-means clustering algorithm is a well-known technology and will not be described here. As other implementation methods, the implementer can adopt other methods of the existing technology, such as the DBSCAN clustering algorithm, etc. This embodiment does not impose any special restrictions on this.
[0036] Calculate the average value of all the first cross-degrees in each cluster, and record the connected nodes in the cluster with the largest average value as the belonging nodes; otherwise, record them as peer nodes; It should be noted that the The greater the first cross degree between a node and its connected nodes, the greater the The more likely a node is a subclass of the connected node or a broad or narrow relationship of the same concept, for example, (mammal, is, animal), where mammal is a subclass of animal, and the subclass of the entity "animal" can be inherited by its subclass "mammal", which is manifested as a larger first intersection degree of the attributes of the two entities.
[0037] Furthermore, according to the inheritance characteristics of the attributes between the node's subclass and parent class, the attributes of the node can be obtained. The node itself does not have the attribute, get the The inherited properties of each node are: The first The node's own attributes do not have attributes corresponding to the own attributes of all its nodes, as the first The inherited properties of each node; It should be noted that, for the sake of ease of understanding, it is assumed that The node to which the node belongs is , No. The node has attribute A, The node itself does not have attribute A, then The attribute A of the node to which it belongs will be If it is inherited by nodes, then attribute A is The inherited properties of the node; The attributes corresponding to each node are divided into two categories: one is the attributes extracted from the data source, and the other is the inherited attributes between nodes.
[0038] Since different knowledge graphs have different data sources, the information contained in the data sources is often different. This can cause the attributes represented by the same node to be missing or omitted, which in turn causes the corresponding attributes of entities in different knowledge graphs to be different, that is, some may contain inherited attributes, while others may not. However, analyzing the first cross-degree between entities corresponding to two nodes in different knowledge graphs ignores this situation. Therefore, based on the changes in the inherited attributes between nodes in the two knowledge graphs, a correction factor is constructed to correct the first cross-degree and calculate the second cross-degree. Specifically, Select a node from each of the two knowledge graphs and record them as target nodes. and target node ; Record all the own attributes and inherited attributes of each target node as target attributes; Calculate target node Each inherited attribute of the target node The reciprocal of the distance between each target attribute is recorded as the matching degree; In this embodiment, the target node is calculated Each inherited attribute of the target node The inverse of the edit distance between each target attribute is recorded as the matching degree, where the calculation of the edit distance is a well-known technology and will not be repeated here. Each target attribute includes its own attributes extracted from the data source and the inherited attributes of the node.
[0039] Select the target node Each inherited attribute of the target node The maximum matching degree between all target attributes of the target node The sum of the maximum values of all inherited attributes is used as the target node degree of relevance; It should be noted that the greater the matching degree, the The inherited properties of the target node The more the target attributes in the target node match, the more the semantic information contained is the same, and the greater the correlation, the more likely the target node is to be The inherited properties of the target node The higher the information relevance of the target attribute in .
[0040] Similarly, calculate the target node Each inherited attribute of the target node The reciprocal of the distance between each target attribute is recorded as the matching degree; In this embodiment, the target node is calculated Each inherited attribute of the target node The inverse of the edit distance between each target attribute is recorded as the matching degree, where the calculation of the edit distance is a well-known technology and will not be repeated here. The target attributes include attributes extracted from the data source and inherited attributes of the node.
[0041] Select the target node Each inherited attribute of the target node The maximum matching degree between all target attributes of the target node The sum of the maximum matching degrees of all inherited attributes is used as the target node degree of relevance; The target node The correlation between the target node The sum of the association degrees of With the target node Correction factors between The target node With the target node The product of the first intersection degree and the correction factor is used as the target node With the target node The second degree of intersection between It should be noted that the larger the correction factor, the more similar the nodes to which the two target nodes belong, the more similar the actual attributes between the target nodes, and the larger the actual first cross-degree should be; conversely, the greater the difference between the nodes to which the two target nodes belong, the greater the difference between the actual attributes between the target nodes, the smaller the actual first cross-degree should be, and the larger the obtained second cross-degree, indicating that the greater the correlation between the descriptive information of the entities corresponding to the two target nodes in different knowledge graphs, the more likely the two entities represent the same thing or concept.
[0042] At this point, the second intersection degree between two target nodes in any two knowledge graphs is obtained.
[0043] Step 4: Determine the structural similarity of the two target nodes through the similarity between the relationship between the target node and its sibling nodes, and the difference in the number of sibling nodes; determine the correlation coefficient of the two target nodes based on the second intersection degree and the structural similarity, evaluate the entities corresponding to the target nodes, perform entity alignment, and construct a dynamic knowledge graph.
[0044] Furthermore, through the structural relationship between each target node and its peer nodes in the knowledge graph, the structural association of nodes in different knowledge graphs can be reflected, and the structural similarity can be calculated, specifically: Based on the relationship representation between each node and its sibling nodes in each knowledge graph, the relationship representation is represented by a quadruple. The quadruple between each node and its sibling nodes is obtained and converted into a word vector through a pre-trained model. The number of all word vectors corresponding to each node is consistent with the number of its corresponding sibling nodes. In this embodiment, the pre-training model uses the Word2Vec model to obtain word vectors, wherein Word2Vec is a well-known technology and will not be described in detail here. As other implementation methods, implementers can adopt other methods of existing technologies, such as the GloVe model, etc. This embodiment does not impose any special restrictions on this.
[0045] Select a node from each of the two knowledge graphs and record them as target nodes. and target node ; Target node and target node For example, the target node and target node The node with the least number of peer nodes is recorded as the key node. Assume that the target node The number of sibling nodes is less than the target node The number of peer nodes; Select the target node Each corresponding word vector and target node The maximum value of the correlation between all corresponding word vectors is recorded as the maximum correlation; In this embodiment, the degree of relevance is determined by the target node Each corresponding word vector and target node The cosine similarity between the corresponding word vectors is measured, wherein the calculation of cosine similarity is a well-known technology and will not be repeated here. As other implementation methods, the implementer may adopt other methods of the existing technology, such as Jaccard similarity, etc. This embodiment does not impose any special restrictions on this.
[0046] It should be noted that the larger the maximum correlation is, the more semantically similar the two word vectors are, and the higher the correlation is.
[0047] Calculate target node The cumulative sum of the maximum relevance of all corresponding word vectors; Calculate target node The number of all corresponding peer nodes and the target node The difference between the numbers of all corresponding sibling nodes is recorded as the number difference; In this embodiment, the target node is calculated The number of all corresponding peer nodes and the target node The absolute value of the difference between the numbers of all corresponding peer nodes is recorded as the number difference.
[0048] The ratio of the cumulative sum to the quantity difference is used as the target node With the target node The structural similarity between them.
[0049] In this embodiment, the target node With the target node The calculation formula of the structural similarity between is: in, Target node With the target node The structural similarity between Target node The maximum relevance of each corresponding word vector, Target node The number of all corresponding peer nodes, Target node The number of all corresponding peer nodes, where Less than , To preset a value greater than 0, avoid the denominator being 0, The value range is , in this embodiment, The value is 1. As other implementation methods, the implementer can set it according to actual conditions.
[0050] It should be noted that the larger the cumulative sum is, the higher the target node With the target node The stronger the correlation, the smaller the difference in the number, indicating that the target node With the target node The closer the number of nodes of the same level connected between them is, the greater the similarity of the obtained structure is, indicating that the target node Structure and nodes in the knowledge graph The higher the similarity of the structure in the knowledge graph, the better the target node With the target node The greater the possibility that they are represented as the same entity.
[0051] Furthermore, based on the structural similarity and the second cross-degree, a correlation coefficient is calculated, specifically: The normalized result of the product of the structural approximation and the second cross degree is used as the target node With the target node The correlation coefficient between In this embodiment, the sigmoid function is used for normalization processing, wherein the sigmoid function is a well-known technology and will not be described in detail here. As other implementation methods, the implementer may adopt other methods of the prior art, such as the tanh function, etc., and this embodiment does not impose any special restrictions on this. The step flow chart of the method for obtaining the correlation coefficient provided in the embodiment of the present application is as follows: Figure 2 shown.
[0052] If the correlation coefficient is greater than or equal to the preset threshold, the target node With the target node The corresponding entities are the same entity, otherwise, the target node With the target node The corresponding entities do not belong to the same entity; In this embodiment, the preset threshold value is 0.8. As for other implementation methods, the implementer can set it according to actual conditions.
[0053] It should be noted that, the larger the correlation coefficient is, the more likely it is that the entities corresponding to the two target nodes represent the same entity; conversely, the smaller the correlation coefficient is, the less likely it is that the entities corresponding to the two target nodes represent the same entity.
[0054] The entities corresponding to the nodes in different knowledge graphs are aligned, the information in different knowledge graphs is integrated, and a dynamic knowledge graph is constructed; this improves the efficiency and accuracy of information sharing and integration between different knowledge graphs, thereby improving the efficiency and accuracy of constructing dynamic knowledge graphs.
[0055] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0056] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0057] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the present application. It should be noted that a person skilled in the art can make various modifications and improvements without departing from the spirit of the present application. Therefore, any simple modifications, equivalent variations, and modifications to the above embodiments made in accordance with the technical essence of the present application without departing from the content of the present application's technical solution fall within the scope of protection of the present application's technical solution.
Claims
1. A method for constructing a dynamic knowledge graph for three-dimensional security situation analysis, characterized in that: The method comprises the following steps: Extract entities, relationships between entities, attributes, and time from texts of multiple data sources to construct multiple knowledge graphs. The attributes corresponding to each node in each knowledge graph are recorded as its own attributes. For all entities corresponding to all nodes in the knowledge graph, analyze the differences between the different attributes corresponding to any two entities, and determine the first intersection degree of the two entities based on the distribution of the attributes of the entities in the corresponding text; Based on the connection relationship between each node in the same knowledge graph and the distribution of the first cross degree between the connected nodes, the node to which each node belongs and the node at the same level are determined; based on the own attributes of the node to which it belongs, the inherited attributes of each node are generated; Select a node from each of the two knowledge graphs and record it as a target node; calculate a correction factor based on the difference between the inherited attribute of each target node and the attribute of the other target node, and determine the second cross-degree of the two target nodes in combination with the first cross-degree; Determine the structural similarity of the two target nodes based on the similarity between the target node and its sibling nodes and the difference in the number of sibling nodes; Based on the second intersection degree and the structural similarity, the correlation coefficient of the two target nodes is determined, the entities corresponding to the target nodes are evaluated, and the entities are aligned to construct a dynamic knowledge graph.
2. A method for constructing a dynamic knowledge graph for three-dimensional security situation analysis according to claim 1, characterized in that: Determining the first intersection degree of the arbitrary two entities includes: Analyze the distance between each entity and its corresponding attributes in the text, as well as the frequency of text containing each entity and its corresponding attributes, and calculate the semantic influence; Select the entity with the smallest number of all its own attributes among the two entities and record it as the target entity; Calculating the minimum distance between each self-attribute corresponding to the target entity and all self-attributes corresponding to the other entity in the arbitrary two entities; The first intersection degree is the sum of the ratios of the semantic influence of all the attributes corresponding to the target entity of the arbitrary two entities to the minimum distance.
3. The method for constructing a dynamic knowledge graph for three-dimensional security situation analysis according to claim 2, characterized in that: The calculating of the semantic influence includes: The extracted texts of entities corresponding to all nodes in each knowledge graph are combined into a text set; Selecting texts that contain all entities and their own attributes from the text set and recording them as target texts; Calculate the character distance between each entity in the target text and its corresponding attribute; Counting the frequency of occurrence of target text in the text set; Calculate the sum of the distances between each entity and its corresponding attributes in all target texts, and record it as relative spacing; The semantic influence is a ratio of the frequency to the relative distance.
4. The method for constructing a dynamic knowledge graph for three-dimensional security situation analysis according to claim 1, characterized in that: Determining the node to which each node belongs and the node at the same level includes: All other nodes connected to each node in each knowledge graph are recorded as connected nodes; Clustering the first intersection degree between each node and all connected nodes to obtain two clusters; The average value of all the first cross-degrees in each cluster is calculated, and the connected node in the cluster with the largest average value is recorded as the belonging node; otherwise, it is recorded as the peer node.
5. The method for constructing a dynamic knowledge graph for three-dimensional security situation analysis according to claim 1, characterized in that: The inherited attributes of each node are attributes that do not exist in the own attributes of each node and are inherited from the own attributes of all corresponding nodes.
6. The method for constructing a dynamic knowledge graph for three-dimensional security situation analysis according to claim 1, characterized in that: The calculation of the correction factor includes: Record all the own attributes and inherited attributes of each target node as target attributes; The reciprocal of the distance between each inherited attribute corresponding to any one of the two target nodes and each target attribute of the other target node is recorded as the matching degree; Among the two target nodes, selecting the maximum matching degree between each inherited attribute corresponding to any one target node and all target attributes of the other target node; and taking the sum of the maximum matching degrees of all inherited attributes corresponding to any one target node as the association degree of any one target node; The correction factor is the sum of the association degrees between the two target nodes.
7. The method for constructing a dynamic knowledge graph for three-dimensional security situation analysis according to claim 1, characterized in that: The second cross-degree is a product of the first cross-degree between the two target nodes and the correction factor.
8. The method for constructing a dynamic knowledge graph for three-dimensional security situation analysis according to claim 1, characterized in that: The relationship is represented by the quadruple between nodes in the knowledge graph as follows: , s is the head entity, p is the relationship, o is the tail entity, t is The time when it is established, where E represents the entity set, R represents the relationship set, and T represents the time interval set. The further measurement process of the relationship between the target node and its sibling nodes is: Based on the relationship between each node and its sibling nodes in each knowledge graph, the quadruple between each node and its sibling nodes is extracted, and the quadruple is converted into a word vector through a pre-trained model.
9. The method for constructing a dynamic knowledge graph for three-dimensional security situation analysis according to claim 8, characterized in that: Determining the structural similarity of the two target nodes includes: Select the target node with the smallest number of peer nodes among the two target nodes and record it as the key node; Select the maximum value of the correlation between each word vector corresponding to the key node of the two target nodes and all word vectors corresponding to the other target node, and record it as the maximum correlation; Calculating the cumulative sum of the maximum relevance of all word vectors corresponding to the key node; Calculate the difference between the number of all peer nodes corresponding to the key node and the other target node in the two target nodes, and record it as the number difference; The structural similarity is the ratio of the cumulative sum to the quantitative difference.
10. The method for constructing a dynamic knowledge graph for three-dimensional security situation analysis according to claim 1, characterized in that: Determining the correlation coefficient of the two target nodes and evaluating the entities corresponding to the target nodes includes: Normalizing the product of the structural similarity and the second cross degree as the correlation coefficient between the two target nodes; If the correlation coefficient is greater than or equal to a preset threshold, the entities corresponding to the two target nodes are the same entity; otherwise, the entities corresponding to the two target nodes are not the same entity.
Citation Information
Patent Citations
Internet of Things capability and knowledge mapping and construction method thereof
CN108021718A
Data processing method and device and computer readable storage medium
CN114328799A
RPA-oriented multi-modal interactive entity alignment method
CN116128056A
Multi-source data-based power grid knowledge graph construction method
CN119886298A
Knowledge graph construction method and apparatus, and storage medium and electronic device
WO2025123841A1