Method and device for locating system model data in knowledge graph
By using technical means such as relational graph convolution networks and fuzzy matching query in the global knowledge base, the problem of low data positioning accuracy in system models in the existing technology is solved, and higher entity positioning accuracy and output quality of relational graph convolution networks are achieved.
Patent Information
- Application Number
- CN202410364106.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-03-28
AI Technical Summary
The existing knowledge graph embedding methods do not target the data characteristics of the architecture field, which leads to the easy neglect of the correlation relationship of the architecture data when positioning the system model data in the global knowledge base, which affects the accuracy of entity positioning.
By obtaining the relationship triplets and attribute triplets of entity nodes from the architecture knowledge graph, fuzzy matching query is performed to determine the entity nodes, and then using the relationship graph convolution network to embed the system knowledge graph and the model data to be located, integrating entity adjacency information and entity relationship information, extracting structural feature information, pre-aligning, and finally using attribute extraction and feature fusion to perform confidence verification to obtain the target entity.
It improves the positioning accuracy of the architecture entity in the global knowledge base, reduces the matching of wrong entity pairs, improves the effect of entity semantic measurement, and further improves the output quality and accuracy of the relationship graph convolution network through iterative training framework.
Smart Images

Figure CN118170926B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software architecture design, and in particular to a method and device for locating system model data in a knowledge graph. Background Art
[0002] With the accumulation of architecture design data, by summarizing the knowledge concepts in the architecture field, building an architecture domain ontology model, extracting and integrating knowledge architecture data from multiple sources, and obtaining a domain knowledge graph, called a global knowledge base. The main application scenario is to use the global knowledge base to mine the architecture for software manufacturing and design experience. This application relies on the system model data and its associations to find its location in the global knowledge base, that is, to find the instance data where the system model data and the entity in the knowledge base refer to the same entity. The accuracy and efficiency of the location directly affect the application effect of the global knowledge base.
[0003] Since the global knowledge base is formed by storing and managing knowledge data from different sources in the form of a knowledge graph and integrating them, and the ontology structure of the architecture graph is stable, when locating the architecture data and its associations in the global knowledge base, the architecture data and its associations are actually aligned as entities with the entities in the knowledge base, that is, entities representing the same object are found in the global knowledge base to form entity pairs, and then the position of the architecture data to be matched in the global knowledge graph is found.
[0004] At present, the existing methods for solving the positioning of system models in system knowledge graphs still rely on traditional query matching. By standardizing and constraining the naming of entities, finding the positioning of system model data in the knowledge graph is essentially entity alignment, finding entities that refer to the same object, that is, the mapping object of the system model data in the knowledge graph. By locating the spatial position of the data in the knowledge graph, it supports the fusion of knowledge graphs, knowledge query, and data relationship verification. The higher the accuracy of entity positioning, the greater the improvement of subsequent functional effects. Among them, entity alignment refers to the discovery of equivalent entities between different graphs, matching and aligning entities that refer to the same object in different graphs, and is applied to various scenarios such as knowledge graph fusion and entity classification. Entity alignment generally refers to ontology alignment and instance alignment. The problem of locating system model data is actually to find the instance position of the system model data referring to the same object in the knowledge base that integrates multiple knowledge, that is, instance alignment. Instance alignment generally depends on the relevant information contained in the knowledge graph entity, including entity name, entity attribute information, graph structure information where the entity is located, entity description information, etc.
[0005] There are currently a variety of entity alignment methods used in different scenarios, such as the method based on similarity calculation, which uses natural language processing to measure instance-related information and calculate the similarity between instances. The instances with the highest similarity are represented as aligned instances. Common implementation technologies include TF-IDF, edit distance, and synonym semantic measurement. There is also a method that uses knowledge representation learning to convert entities in the knowledge graph into low-density vectors. This process is called embedding. The embedded vectors are interactively aligned using aligned entities and mapped to the same space. Then the similarity is measured based on the spatial distance of the vectors. Common implementation technologies include Trans algorithm, GNN, and GCN. In different scenarios of actual applications, due to different application data characteristics, the representation effects of different embedding algorithms vary greatly. Different algorithms need to be implemented according to different data information and characteristics. However, regarding the problem of locating system model data in the architecture knowledge graph, there is currently no algorithm that directly locates entities based on the characteristics of the architecture knowledge graph. The method using similarity measurement is prone to information omission and low semantic measurement effect, which affects the accuracy of entity positioning.
[0006] In summary, the existing knowledge graph embedding methods do not target the data characteristics of the architecture field, but directly locate entities, which easily ignores the association relationship of the architecture data and affects the accuracy of entity location in the global knowledge base. In addition, the existing methods that use entity names for matching are difficult to fully express the semantic information of the architecture data, which also affects the accuracy of entity location in the global knowledge base. Summary of the invention
[0007] Based on this, it is necessary to provide a method and device for locating system model data in a knowledge graph that can improve the accuracy of locating system architecture entities in a global knowledge base in response to the above technical problems.
[0008] The present invention provides a method for locating system model data in a knowledge graph, the method comprising:
[0009] Obtaining relationship triples and attribute triples of entity nodes from the architecture knowledge graph, and performing fuzzy matching query through entity names to determine multiple entities with repeated characters between the entity names;
[0010] The architecture knowledge graph and the model data to be located are embedded through a relational graph convolutional network to map the model data to be located into a low-density vector space, and structural feature information is extracted by integrating entity adjacency information and entity relationship information to obtain pre-aligned entity pairs;
[0011] Extracting attributes from entities in the pre-aligned entity pairs, and fusing features of the pre-aligned entity pairs based on the extracted attribute information to obtain a first entity pair with the highest fused feature ranking;
[0012] Performing confidence checks on entities in the first entity pair to obtain a target entity, and iteratively inputting the target entity into the relationship graph convolutional network to expand a training set, thereby obtaining a spatial position of the model data to be located in the architecture knowledge graph;
[0013] The relationship triples are composed of different architecture entities and entity relationships between the different architecture entities. The architecture entities are the system knowledge of the software in the system global knowledge base, including at least the output relationship, input relationship, associated information and associated system of the software.
[0014] The attribute triplet consists of an architecture entity, an attribute name of the architecture entity and a corresponding attribute value. The target entity is the architecture entity corresponding to the model data to be located in the architecture knowledge graph, and the spatial position of the model data to be located in the architecture knowledge graph is the spatial position of the target entity in the architecture knowledge graph.
[0015] In one embodiment, the acquiring of the relationship triples and attribute triples of the entity nodes from the architecture knowledge graph, and performing a fuzzy matching query through the entity name to determine multiple entities having repeated characters with the entity name, includes:
[0016] Perform fuzzy matching query by entity name and determine whether there is a first entity with the same name in the architecture knowledge graph; if so,
[0017] Directly generate a second entity pair based on the first entity, and input the second entity pair into the relationship graph convolution network; if not, then
[0018] Acquire, from the architecture knowledge graph, a plurality of second entities having repeated characters with the entity name;
[0019] The first entity and the second entity are both system knowledge of the software in a system global knowledge base.
[0020] In one embodiment, the method includes obtaining a relationship triple and an attribute triple of an entity node from the architecture knowledge graph, and performing a fuzzy matching query through the entity name to determine multiple entities having repeated characters with the entity name, and then comprising:
[0021] Sorting the plurality of second entities according to character overlap to obtain second entities sorted from high to low character overlap;
[0022] Selecting a second entity whose character overlap ranking is not less than a first threshold from the second entities ranked from high to low in character overlap;
[0023] A third entity pair is generated based on the second entity whose character overlap ranking is not lower than the first threshold as an input of the relationship graph convolutional network.
[0024] In one embodiment, the architecture knowledge graph and the model data to be located are embedded through a relational graph convolutional network to map the model data to be located to a low-density vector space, and structural feature information is extracted by integrating entity adjacency information and entity relationship information to obtain pre-aligned entity pairs, including:
[0025] Aggregating different entity node information according to different entity relationships, and dividing the entity nodes in the architecture knowledge graph into input and output to generate a structural feature vector;
[0026] Calculating the spatial distance between entities based on the structural feature vector to obtain the entity with the closest spatial distance, wherein the pre-aligned entity pair is composed of the entities with the closest spatial distance;
[0027] The calculation formula of the spatial distance is:
[0028]
[0029] Where D s is the spatial distance, e1 and e2 are two entities, and n is the entity structure feature obtained based on the relationship graph convolution network. and These are the entity feature structures of the two entities e1 and e2. is the dimension, l s is the number of layers of the neural network.
[0030] In one embodiment, extracting attributes from entities in the pre-aligned entity pairs, and fusing features of the pre-aligned entity pairs based on the extracted attribute information to obtain a first entity pair with the highest fused feature ranking, includes:
[0031] Extracting first attribute information of the third entity according to the entity type of the third entity in the pre-aligned entity pair and the attribute feature of the third entity, and constructing a first attribute triplet based on the first attribute information;
[0032] Based on the first attribute triplet, calculating the attribute semantic distance between the third entities, so as to select a fourth entity whose attribute semantic distance is not less than a second threshold from the architecture knowledge graph;
[0033] Wherein, the third entity and the fourth entity are both the system knowledge of the software in the system global knowledge base, and the calculation formula of the attribute semantic distance is:
[0034]
[0035] Where D a (e1, e2) is the attribute semantic distance between the pre-aligned entity pairs, 1≤i≤|A|, D i (e1, ai, e2 are attribute value similarity measures of the pre-aligned entity pair (e1, PreSame, e2), ω i For attribute a i The influence factor of i ≤1.
[0036] In one embodiment, the confidence check of the entities in the first entity pair to obtain the target entity, and iteratively inputting the target entity into the relationship graph convolutional network to expand the training set, and obtaining the spatial position of the model data to be located in the architecture knowledge graph, includes:
[0037] Inputting a second entity pair consisting of the target entities with the highest confidence and the shortest spatial distance among the plurality of first entity pairs into the relationship graph convolution network for iterative calculation to obtain the spatial positions between the target entities in the second entity pair;
[0038] Based on the spatial positions between the target entities, the spatial positions of the model data to be located in the architecture knowledge graph are obtained.
[0039] In one embodiment, the propagation rule of the relational graph convolutional network is:
[0040]
[0041] In the formula, is the structural feature of the relationship propagation node, σ is the activation function, is the structural feature of the entity, is the weight matrix of entity features, is the neighbor node feature, is the weight matrix of the corresponding relationship between neighbor node features, is the regularization constant, is a natural number;
[0042] The calculation formula of the feature fusion is:
[0043] D(e1,e2)=αD s (e1,e2)+(1-α)D a (e1,e2);
[0044] In the formula, D(e1,e2) is the fusion feature, D s (e1, e2) is the vector space distance between the pre-aligned entity pair, D a (e1, e2) is the attribute semantic distance between the pre-aligned entity pair, and α is a weight parameter for adjusting the feature weight of the vector space distance and the attribute semantic distance.
[0045] The present invention also provides a device for locating system model data in a knowledge graph, the device comprising:
[0046] A fuzzy matching module is used to obtain the relationship triples and attribute triples of the entity nodes from the architecture knowledge graph, and perform fuzzy matching query through the entity name to determine multiple entities with repeated characters between the entity names;
[0047] A feature extraction module embeds the architecture knowledge graph and the model data to be located through a relational graph convolutional network to map the model data to be located into a low-density vector space, and extracts structural feature information by integrating entity adjacency information and entity relationship information to obtain pre-aligned entity pairs;
[0048] A feature fusion module, used to extract attributes from entities in the pre-aligned entity pairs, and perform feature fusion on the pre-aligned entity pairs based on the extracted attribute information, so as to obtain a first entity pair with the highest fusion feature ranking;
[0049] A data positioning module, used to perform confidence verification on the entities in the first entity pair to obtain a target entity, and iteratively input the target entity into the relationship graph convolutional network to expand the training set, and obtain the spatial position of the model data to be positioned in the architecture knowledge graph;
[0050] The relationship triples are composed of different architecture entities and entity relationships between the different architecture entities. The architecture entities are the system knowledge of the software in the system global knowledge base, including at least the output relationship, input relationship, associated information and associated system of the software.
[0051] The attribute triplet consists of an architecture entity, an attribute name of the architecture entity and a corresponding attribute value. The target entity is the architecture entity corresponding to the model data to be located in the architecture knowledge graph, and the spatial position of the model data to be located in the architecture knowledge graph is the spatial position of the target entity in the architecture knowledge graph.
[0052] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the method for locating system model data in a knowledge graph as described in any one of the above.
[0053] The present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements a method for locating system model data in a knowledge graph as described in any one of the above.
[0054] The above-mentioned method and device for locating system model data in the knowledge graph extracts the relationship triples and attribute triples of entity nodes from the system architecture knowledge graph, and performs fuzzy matching query through the entity name to determine multiple system architecture entities with repeated characters between the entity names. Subsequently, the system architecture knowledge graph and the model data to be located are embedded through the relationship graph convolution network to map the model data to be located to a low-density vector space, and the corresponding structural feature information is extracted by integrating the entity adjacency information and the entity relationship information to obtain the pre-aligned entity pair. Then, the attributes of the entities in the pre-aligned entity pair are extracted, and the pre-aligned entity pair is feature fused based on the extracted attribute information to obtain the entity pair with the highest fusion feature ranking. Finally, the system architecture entity in the highest ranked entity pair is confidence checked to obtain the target entity with the highest confidence, and the target entity is iteratively input into the relationship graph convolution network to expand the training set, which improves the embedding accuracy of the structural information, and finally obtains the spatial position of the model data to be located in the system architecture knowledge graph, that is, the spatial position of the target entity in the system architecture knowledge graph. This method extracts structural and attribute information from the architecture knowledge graph and performs confidence verification on the generation of entity pairs, which improves the effect of entity semantic measurement and reduces the matching of incorrect entity pairs. In addition, this method uses high-confidence architecture entity pairs as feedback input to the entity training framework, that is, the iterative training framework is used to further improve the output quality and accuracy of the relational graph convolutional network, which is beneficial to subsequent software manufacturing and design. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0056] Figure 1 One of the flow charts of the method for locating system model data in the knowledge graph provided by the present invention;
[0057] Figure 2 A schematic diagram of the process of locating system model data in a system knowledge graph of a method for locating system model data in a knowledge graph in a specific embodiment provided by the present invention;
[0058] Figure 3 A schematic diagram of a fuzzy matching query process of a method for locating system model data in a knowledge graph in a specific embodiment provided by the present invention;
[0059] Figure 4 A schematic diagram of the spatial position of entities with high confidence in a method for locating system model data in a knowledge graph in a specific embodiment provided by the present invention;
[0060] Figure 5 The second flowchart of the method for locating system model data in the knowledge graph provided by the present invention;
[0061] Figure 6 The third flowchart of the method for locating system model data in the knowledge graph provided by the present invention;
[0062] Figure 7 The fourth flowchart of the method for locating system model data in the knowledge graph provided by the present invention;
[0063] Figure 8 The fifth flowchart of the method for locating system model data in the knowledge graph provided by the present invention;
[0064] Fig. 9 The sixth flowchart of the method for locating system model data in the knowledge graph provided by the present invention;
[0065] Fig.10 A schematic diagram of the structure of a positioning device for system model data in a knowledge graph provided by the present invention;
[0066] Fig.11 This is a diagram of the internal structure of the computer device provided by the present invention. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0068] Combine the following Figure 1-Figure 11 The present invention describes the method and device for locating system model data in a knowledge graph.
[0069] like Figure 1 As shown, in one embodiment, a method for locating system model data in a knowledge graph includes the following steps:
[0070] Step S110, obtain the relationship triples and attribute triples of the entity nodes from the architecture knowledge graph, and perform fuzzy matching query through the entity name to determine multiple entities with repeated characters between the entity names.
[0071] Specifically, the server obtains the relationship triples and attribute triples of the entity nodes from the architecture knowledge graph (i.e., the architecture knowledge graph constructed based on the software's output relationship, input relationship, related information, related system and other knowledge information in the system's global knowledge base), and performs fuzzy matching queries through the entity names to determine multiple entities that have repeated characters with the entity names.
[0072] Among them, the relationship triple is composed of different architecture entities and the entity relationships between different architecture entities. The architecture entity is the system knowledge of the software in the system global knowledge base, which at least includes the output relationship, input relationship, associated information and associated system of the software. The attribute triple is composed of the architecture entity, the attribute name of the architecture entity and the corresponding attribute value.
[0073] Combination Figure 2 and Figure 3 As shown, in a specific embodiment, when fuzzy querying, the names of knowledge information architecture entities such as software output relations, input relations, associated information, and associated systems are directly used as query objects for preliminary filtering. If the entity names are consistent, they can be considered to be the same object, and entity pairs are formed with the queried architecture entity as the positioning point. The entity names are sorted according to the number of overlapping characters and recommended to the user. The user can manually align the entities, that is, select objects that can form entity pairs from the candidate items to form entity pairs as the input of R-GCN (relational graph convolutional network).
[0074] In the software manufacturing process, taking the global knowledge base of the information system construction system as an example, the global knowledge base of the system contains ontology metamodel, business data and extracted system knowledge. The ontology metamodel is the basic elements and relationships for building the system architecture framework. The business data mainly comes from asset data, scientific and technological data, equipment database, etc. The extracted system knowledge refers to the system knowledge contained in the knowledge base through relationship extraction and fusion in the process of forming the knowledge graph.
[0075] Step S120, embedding the architecture knowledge graph and the model data to be located through a relational graph convolutional network to map the model data to be located into a low-density vector space, and extracting structural feature information by integrating entity adjacency information and entity relationship information to obtain pre-aligned entity pairs.
[0076] Specifically, the server embeds the architecture knowledge graph and the model data to be located through a relational graph convolutional network to map the model data to be located into a low-density vector space, and extracts the corresponding structural feature information by integrating entity adjacency information and entity relationship information to obtain pre-aligned entity pairs.
[0077] Combination Figure 2 and Figure 3 As shown, in a specific embodiment, a translation model and a deep model are generally used for entity embedding. The translation model captures structural information in the relationship path, such as the Trans family algorithm for vector representation of triples. Since the vector mapping performed by the Trans algorithm only focuses on local structural information and ignores the impact of the overall structure containing semantics on entity alignment, the effect of actual use is difficult to meet expectations. Therefore, it is necessary to use a deep model to mine the overall structural semantics of the graph to be matched, and combine the commonly considered information (entity name, entity description, attribute information, etc.) to measure entity alignment.
[0078] Since the architecture entities with the same meaning in different architecture knowledge graphs have similar adjacency information, the commonly used graph neural network structure embedding method is the graph convolutional neural network (GCN). The basic idea of GCN is to obtain the feature information of each node, including its own features and all neighboring nodes, take the weighted average of all node features (including its own node), and train the obtained feature vector through a neural network. The number of layers of GCN refers to the farthest distance that the node features can be transmitted, that is, in a 1-layer GCN, each node can only obtain information from its neighbor structure, and the process of collecting information for each node is independent and for all nodes at the same time. When another layer is superimposed on the first layer, the process of collecting information is repeated, but this time the neighboring nodes already have their own information, which makes the number of layers the maximum jump that the node can take. However, the larger the number of layers, the better. According to practical experience, GCN generally has 2 or 3 layers to get the best results.
[0079] GCN is oriented to undirected graphs or graphs of the same relationship, but the architecture knowledge graph is a heterogeneous graph with multiple types of structural relationships. Therefore, the impact of different types of structural relationships needs to be considered. For this reason, this embodiment uses a graph convolutional neural network with relational weights, namely, a relational graph convolutional neural network R-GCN.
[0080] In this embodiment, a two-layer relational graph convolutional neural network is used to extract the structural features of the graph to form a structural feature matrix of the entity in the vector space, which not only retains the structural semantic information, but also can obtain a structural feature matrix that can be accurately measured. The input of R-GCN includes the adjacency matrix, degree matrix, and feature matrix of the graph, and outputs the structural feature matrix of the architecture entity. The two-layer R-GCN propagation rule selected in this embodiment is shown as follows:
[0081]
[0082] In the formula, is the structural feature of the relationship propagation node, σ is the activation function, is the structural feature of the entity, is the weight matrix of entity features, is the neighbor node feature, is the weight matrix of the corresponding relationship between neighbor node features, is the regularization constant, is a natural number.
[0083] In this embodiment, different node information is aggregated according to different architecture entity relationships, and the difference between the input in (input) and out (output) of the graph entity node can characterize the relationship direction of the directed graph, and the structural feature vector can be generated after activation of the ReLU module. It can be found from the propagation rule of R-GCN that the node features of each layer are obtained by the node features of the previous layer and the relationship between the nodes. At the same time, the difference from GCN is that the weighted sum of the node's neighbor node features and its own features is used to obtain new features, and the type and direction of the edge are taken into account, so that the extraction and embedding of structural features are more accurate and information loss is reduced.
[0084] In this embodiment, the structural feature matrix of the graph can be obtained after two layers of R-GCN. Since the content of the global knowledge base is relatively stable, the modification is generally local. The result of the global knowledge base after R-GCN conversion can be stored, and there is no need to repeat the conversion when performing entity matching again. Therefore, when locating data, it is only necessary to convert the object to be located, that is, to transform the entity to be aligned through R-GCN, output the structure vector, and the structure vector can be processed only once. After embedding, the existing conventional steps require the structural vectors of the two graphs to be aligned to the same vector space based on the aligned entity pairs because the structural vectors of different knowledge graphs are not in the same space. However, when looking for the location of the graph to be aligned in the global knowledge base, since the model data to be located and the architecture knowledge graph have the same ontology structure, that is, the same knowledge graph, no alignment processing is required after the embedding operation.
[0085] Step S130 , extracting attributes from entities in the pre-aligned entity pairs, and performing feature fusion on the pre-aligned entity pairs based on the extracted attribute information, so as to obtain a first entity pair with the highest fusion feature ranking.
[0086] Specifically, the server extracts attributes from the entities in the pre-aligned entity pairs obtained in step S120, and performs feature fusion on the pre-aligned entity pairs based on the extracted attribute information, thereby obtaining an entity pair with the highest fusion feature ranking, namely, the first entity pair.
[0087] Combination Figure 2 and Figure 3As shown, in a specific embodiment, since it is difficult to completely obtain accurate data positioning based solely on structural information, it is necessary to consider further filtering based on the local attribute semantic information of the data entity to be matched. In the process of attribute extraction, it is first necessary to construct attribute triples and calculate the attribute semantic distance according to the formula. When constructing attribute triples, it is necessary to consider the attributes of the entity. First, it is necessary to construct attribute triples (entity, attribute name, attribute value) based on the basic type of the entity and the attribute features it possesses, using the current method of constructing attribute triples that is commonly used to extract attribute information, for example (drone, maximum pulling force, 1800), and calculate the attribute semantic distance after embedding the attribute features.
[0088] It should be noted that entity data representing the same instance should have similar attribute semantic distances. Entities with the same attribute values can be understood as related entities, indicating that they have the same capabilities and can be in the same position in the global knowledge base. Therefore, related entities can be obtained by screening the attribute semantic distance. Entities generally have multiple attributes, and the degree of discreteness between attributes is relatively large, and the degree of semantic relevance of attributes to entities is also different.
[0089] Step S140, performing confidence check on the entities in the first entity pair to obtain the target entity, and iteratively inputting the target entity into the relational graph convolutional network to expand the training set to obtain the spatial position of the model data to be located in the architecture knowledge graph.
[0090] Among them, the target entity is the architecture entity corresponding to the model data to be located in the architecture knowledge graph, and the spatial position of the model data to be located in the architecture knowledge graph is the spatial position of the target entity in the architecture knowledge graph.
[0091] Specifically, the server performs confidence verification on the entities in the first entity pair to obtain the target entity, and iteratively inputs the target entity into the relational graph convolutional network to expand the training set, thereby improving the accuracy of structural information embedding, and finally obtaining the spatial position of the model data to be located in the architecture knowledge graph.
[0092] Combination Figure 2 and Figure 3 As shown, in a specific embodiment, after characterizing the feature information of both the architecture and the attributes, the combination of the two features is adjusted by using hyperparameters, and the features are fused and sorted to obtain high-scoring entity pairs, which are called pre-aligned entity pairs. All pre-aligned entity pairs constitute a preliminary entity set to be verified. After the pre-aligned entity pairs are verified, high-confidence entity pairs are obtained, and the high-confidence entity pairs are used as the iterative input of R-GCN, which plays a positive role in the generation of the next high-confidence entity pair, thereby improving the quality and efficiency of the pre-aligned entity pairs obtained in the vector space.
[0093] In this embodiment, the feature fusion formula is:
[0094] D(e1,e2)=αD s (e1,e2)+(1-α)D a (e1,e2).
[0095] In the formula, D(e1,e2) is the fusion feature, D s (e1, e2) is the vector space distance between the pre-aligned entity pairs, D a (e1, e2) is the attribute semantic distance of the pre-aligned entity pair, and α is the weight parameter used to adjust the feature weight of the vector space distance and the attribute semantic distance.
[0096] In this embodiment, the weight parameters of the architecture map are adjusted based on multiple experiments and can be determined based on the verification data effect and experience. For example, after calculating the feature fusion value in the pre-aligned entity pair set, the first two pre-aligned entity pairs are obtained after size comparison and sorting, and the top two with the highest scores are verified to obtain high-confidence entity pairs.
[0097] In this embodiment, among the high-confidence entity pairs after feature fusion, it is verified whether the actual distance is the shortest, so as to reduce the iterative introduction of erroneous entity pairs and avoid causing excessive deviation in the implementation effect. The idea of verification is to calculate the relative position of the entity in the vector space as the closest measurement method. The entity pairs that meet this condition are regarded as high-confidence entity pairs, and the high-confidence entity pairs are used as the input of R-GCN for iteration again until the iterative calculation is completed. Figure 4 As shown in the figure, the relative distance between any entity in an entity pair and other entities is greater than the relative distance between the entity pairs. Entities that meet this requirement are considered high-confidence entities. Let entity a have the shortest distance to entity b. If entity a and entity b are a high-confidence entity pair, then the shortest distance to entity b should be entity a. Let entity a have the second shortest distance to entity b'. Except for entity a, entity b has the shortest distance to entity a'. Then the distance D(a,b) from entity a to entity b should be less than the distance D(a,b') from entity a to entity b'. The distance D(b,a) from entity b to entity a should be less than the distance D(b,a') from entity b to entity a'. That is, let γ1=D(a,b')-D(a,b), γ2=D(b,a)-D(b,a'), and γ1≥θ1, γ2≥θ2 should be satisfied. Both θ1 and θ2 are positive numbers. The pre-aligned entity pairs that meet this condition are regarded as high-confidence entity pairs and serve as the learning input of the next iterative R-GCN framework.
[0098] When designing the software manufacturing system architecture (i.e., the design of the system's composition architecture, input-output relationships, interface processes, etc.), the software's input-output relationships are designed, and it is possible to search for certain input data in the global knowledge base, such as the entity name "XX project integration data", locate the entity corresponding to the entity name in the global knowledge base, and use the knowledge information such as the associated system and associated information of the corresponding software in the system's global knowledge base for spatial location. In addition, the system architecture design of software manufacturing is to design knowledge information such as software structure, input-output relationships, interface processes, and system composition.
[0099] In this embodiment, in order to better assist the design work of software manufacturing, the system architecture knowledge graph in the global knowledge base is used. The system architecture knowledge graph includes the ontology metamodel, business data and the extracted system knowledge. The ontology metamodel is the basic elements and relationships of building the system architecture framework. The business data mainly comes from asset data, scientific and technological data, equipment database, etc. The extracted system knowledge refers to the system knowledge contained in the knowledge base through relationship extraction and fusion in the process of forming the knowledge graph. The system architecture knowledge graph is used to find the system architecture entity in the current software manufacturing design (that is, the entity based on software knowledge information defined in the system architecture knowledge graph), and the relevant knowledge information in the global knowledge base is queried according to the entity name involved in the software design process (that is, the knowledge information of the output relationship, input relationship, associated information, associated system, etc. of the corresponding software), and the knowledge information is used to assist the current software manufacturing design, such as system architecture recommendation and completion in software manufacturing design.
[0100] In the global knowledge base of the system, a high-confidence entity pair of the entity to be located is finally obtained based on structural features and attribute features, that is, the position in the global knowledge base of the system, which can obtain a more accurate positioning effect and provide a better foundation for the subsequent functional application of the global knowledge base based on positioning, such as the recommendation of software related information and related systems, system architecture verification and analysis, and other functions.
[0101] The above-mentioned method for locating system model data in the knowledge graph extracts the relationship triples and attribute triples of entity nodes from the system architecture knowledge graph, and performs fuzzy matching query through the entity name to determine multiple system architecture entities with repeated characters between the entity name. Subsequently, the system architecture knowledge graph and the model data to be located are embedded through the relationship graph convolution network to map the model data to be located to a low-density vector space, and the corresponding structural feature information is extracted by integrating the entity adjacency information and the entity relationship information to obtain the pre-aligned entity pair. Then, the attributes of the entities in the pre-aligned entity pair are extracted, and the features of the pre-aligned entity pair are fused based on the extracted attribute information to obtain the entity pair with the highest fusion feature ranking. Finally, the system architecture entity in the highest ranked entity pair is confidence checked to obtain the target entity with the highest confidence, and the target entity is iteratively input into the relationship graph convolution network to expand the training set, which improves the embedding accuracy of the structural information, and finally obtains the spatial position of the model data to be located in the system architecture knowledge graph, that is, the spatial position of the target entity in the system architecture knowledge graph. This method extracts structural and attribute information from the architecture knowledge graph and performs confidence verification on the generation of entity pairs, which improves the effect of entity semantic measurement and reduces the matching of incorrect entity pairs. In addition, this method uses high-confidence architecture entity pairs as feedback input to the entity training framework, that is, the iterative training framework is used to further improve the output quality and accuracy of the relational graph convolutional network, which is beneficial to subsequent software manufacturing and design.
[0102] like Figure 5 As shown, in one embodiment, the method for locating system model data in the knowledge graph provided by the present invention obtains the relationship triples and attribute triples of entity nodes from the system architecture knowledge graph, and performs fuzzy matching query through the entity name to determine multiple entities with repeated characters between the entity names, specifically including the following steps:
[0103] Step S112, perform fuzzy matching query through entity name, and determine whether there is a first entity with the same name in the architecture knowledge graph.
[0104] Specifically, during the fuzzy matching process, the server first performs a fuzzy matching query through the entity name, and determines whether there is an architecture entity with the same name in the architecture knowledge graph, that is, the first entity, which is the system knowledge of the software in the system global knowledge base.
[0105] Step S114: directly generate a second entity pair based on the first entity, and input the second entity pair into the relationship graph convolutional network.
[0106] Specifically, when the judgment result in step S112 is the first entity with the same name in the architecture knowledge graph, the server can directly generate a corresponding entity pair based on the first entity, that is, the second entity pair, and directly use the second entity pair as the input of the relationship graph convolutional network.
[0107] Step S116, obtaining multiple second entities with repeated characters between the entity name and the entity name from the architecture knowledge graph.
[0108] Specifically, when the judgment result in step S112 is that there is no first entity with the same name in the architecture knowledge graph, the server will obtain multiple architecture entities with repeated characters between the entity names from the architecture knowledge graph, that is, second entities, and the second entities are all system knowledge of the software in the system global knowledge base.
[0109] Combination Figure 3 As shown, in a specific embodiment, during the fuzzy matching process, it is first queried whether there are architecture entities with the same name. If so, an entity pair is directly formed and used as the input of R-GCN. If there are no architecture entities with exactly the same name, they are sorted according to the degree of character overlap, and the related architecture entities with repeated characters in the architecture entity names are presented to the user in the form of a list. The user chooses whether to manually bind the correspondence between the two architecture entities according to needs. If there are no related architecture entities with repeated characters in the architecture entity names reaching a threshold, the fuzzy matching operation step is skipped. Because GCN has the characteristics of low dependence on training input, and because it has the same metadata concept, the global knowledge base and the data to be located and their relationship embedding do not need to be mapped to the same space, and can be directly embedded in the vector space.
[0110] like Figure 6 As shown, in one embodiment, the method for locating system model data in a knowledge graph provided by the present invention obtains the relationship triples and attribute triples of entity nodes from the system architecture knowledge graph, and performs fuzzy matching query through the entity name to determine multiple entities with repeated characters between the entity name, and then includes the following steps:
[0111] Step S610 , sorting the plurality of second entities according to the degree of character overlap to obtain second entities sorted from high to low in degree of character overlap.
[0112] Specifically, after the server obtains multiple second entities with repeated characters in the entity name from the architecture knowledge graph, the server sorts the multiple second entities according to the degree of character overlap to obtain second entities sorted from high to low degree of character overlap.
[0113] Step S620 , selecting a second entity whose character overlap ranking is not lower than a first threshold from the second entities whose character overlap ranking is ranked from high to low.
[0114] Specifically, the server selects a second entity whose character overlap ranking is not lower than a set threshold from the second entities whose character overlap rankings are obtained in step S610 and are ranked from high to low.
[0115] Step S630, generating a third entity pair based on the second entity whose character overlap ranking is not lower than the first threshold as an input of the relationship graph convolution network.
[0116] Specifically, the server generates a responsive entity pair, i.e., a third entity pair, based on the second entity whose character overlap ranking is not lower than the first threshold selected in step S620, and uses the third entity pair as an input of the relationship graph convolutional network.
[0117] like Figure 7 As shown, in one embodiment, the method for locating system model data in the knowledge graph provided by the present invention embeds the system architecture knowledge graph and the model data to be located through a relational graph convolutional network to map the model data to be located to a low-density vector space, and extracts structural feature information by integrating entity adjacency information and entity relationship information to obtain pre-aligned entity pairs, which specifically includes the following steps:
[0118] Step S122, aggregate different entity node information according to different entity relationships, and divide the entity nodes in the architecture knowledge graph into input and output to generate a structural feature vector.
[0119] Specifically, during the feature extraction process, the server will first aggregate different entity node information according to different entity relationships, and divide the entity nodes in the architecture knowledge graph into input (in) and output (out) to generate a structural feature vector.
[0120] Step S124, calculating the spatial distance between entities based on the structural feature vector to obtain entities with the shortest spatial distance, and the pre-aligned entity pairs are composed of entities with the shortest spatial distance.
[0121] Specifically, the server calculates the spatial distance between entities in the structure based on the structural feature vector generated in step S122 to obtain entities with the shortest spatial distance, wherein the pre-aligned entity pair is composed of the entities with the shortest spatial distance.
[0122] Combination Figure 2 and Figure 3As shown, in a specific embodiment, the distance between entities in the structural space is calculated using the structural vector, and the entities with the closest spatial distance are regarded as entity pairs pre-aligned according to the graph structural feature information to complete the filtering of structural features.
[0123] The spatial distance D under the structure is obtained by using the ratio of the distance between two entities in the vector space and the dimension. s :
[0124]
[0125] Where D s is the spatial distance, e1 and e2 are two entities, and n is the entity structure feature obtained based on the relationship graph convolution network. and These are the entity feature structures of the two entities e1 and e2. is the dimension, l s is the number of layers of the neural network.
[0126] In this embodiment, the distance formula is used to filter out the entity pairs with the closest distance between the data entity to be located and the entity in the knowledge base in the vector space, that is, a circle range is set (excluding entities that are too far away), and the pre-aligned entity set within the distance range is selected to construct a triple, and PreSame is used to represent the alignment relationship between the pre-aligned entities after structural filtering, that is, (entity, PreSame, entity to be aligned), and the pre-aligned entity set of the data entity to be located is obtained. After the pre-aligned entity set is obtained through R-GCN, there is a correlation of attributes between entities with a real corresponding relationship, that is, the related attribute features should be similar, and having the same attributes represents that the entity has the same ability or has a consistent possibility of playing a role. Therefore, the entity attribute information in the pre-aligned entity pair is fused as an attribute feature with the structural feature, and the pre-aligned entity set is further filtered by attribute features.
[0127] like Figure 8 As shown, in one embodiment, the method for locating system model data in a knowledge graph provided by the present invention extracts attributes of entities in pre-aligned entity pairs, and performs feature fusion on the pre-aligned entity pairs based on the extracted attribute information to obtain a first entity pair with the highest fusion feature ranking, specifically comprising the following steps:
[0128] Step S132: extract first attribute information of the third entity according to the entity type of the third entity in the pre-aligned entity pair and the attribute feature of the third entity, and construct a first attribute triplet based on the first attribute information.
[0129] Specifically, during the attribute extraction process, the server will extract the attribute information of the third entity (i.e., the first attribute information) based on the entity type and attribute characteristics of the entity in the pre-aligned entity pair (i.e., the third entity), and construct a corresponding attribute triplet (i.e., the first attribute triplet) based on the attribute information.
[0130] Step S134, based on the first attribute triplet, calculate the attribute semantic distance between the third entities to select a fourth entity whose attribute semantic distance is not less than the second threshold from the architecture knowledge graph.
[0131] Specifically, the server calculates the attribute semantic distance between the third entities based on the first attribute triplet constructed in step S132, and then selects entities (i.e., the fourth entity) whose attribute semantic distance is not less than the set expectation from the architecture knowledge graph.
[0132] The third entity and the fourth entity are both the system knowledge or knowledge information of the software in the system global knowledge base.
[0133] Combination Figure 2 and Figure 3 As shown, in a specific embodiment, the attribute semantic distance D between entities is obtained by constructing a measure method of attribute similarity. a (e1, e2), let the pre-aligned entity pair be (e1, PreSame, e2), the attribute semantic distance D between entity e1 and entity e2 a The formula for (e1,e2) is:
[0134]
[0135] Where D a (e1,e2) is the attribute semantic distance, 1≤i≤|A|, D i (e1,a i ,e2) is the attribute value similarity measure of the pre-aligned entity pair (e1, PreSame, e2), ω i For attribute a i The influence factor of i ≤1.
[0136] In this embodiment, D i (e1,a i ,e2) is as follows:
[0137]
[0138] In the formula, V(e1,a i ) and V(e2,a i ) represents the attribute a of the i-th entity pair (e1, PreSame, e2)i The attribute value set, V(e1,a i ) and V(e2,a i ) represents the attribute a of the i-th pre-aligned entity pair (e1, PreSame, e2) i The Jaccard coefficient of multiple attributes is used as the semantic association value. For the semantic association degree of a single attribute, the absolute value of the difference of the same attribute is used as the semantic association value.
[0139] In this embodiment, the impact factor ω i The larger the value, the higher the semantic relevance of the attribute to the entity. For example, the influence factor of the same attribute is 1, and the influence factor of completely different attributes is 0. The determination of the attribute influence factor is obtained by comparing and calculating different attributes. The formula is as follows:
[0140]
[0141] In the formula, count(w i ,a i ) is attribute a i Repeated semantic metric value, count(a i ) is attribute a i Complete semantic metric value, which can be customized based on application experience data in a manually annotated manner.
[0142] like Fig. 9 As shown, in one embodiment, the method for locating system model data in the knowledge graph provided by the present invention performs confidence verification on the entities in the first entity pair to obtain the target entity, and iteratively inputs the target entity into the relationship graph convolutional network to expand the training set, and obtains the spatial position of the model data to be located in the system architecture knowledge graph, which specifically includes the following steps:
[0143] Step S142: input a second entity pair consisting of target entities with the highest confidence and the shortest spatial distance among the multiple first entity pairs into a relational graph convolutional network for iterative calculation to obtain the spatial positions between the target entities in the second entity pair.
[0144] Specifically, in the process of positioning the model data to be positioned, the server first uses the entity pair consisting of the target entities with the highest confidence and the shortest spatial distance among multiple first entity pairs (i.e., the second entity pair) as the input of the relationship graph convolution network for iterative calculation, and then obtains the spatial position between the target entities in the second entity pair.
[0145] Step S144, based on the spatial position between the target entities, obtain the spatial position of the model data to be located in the architecture knowledge graph.
[0146] Specifically, the server obtains the spatial position of the model data to be located in the architecture knowledge graph based on the spatial position between the target entities determined in step S142. That is, the spatial position between the target entities in the second entity pair is the spatial position of the model data to be located in the architecture knowledge graph.
[0147] The following is a description of the device for locating the system model data in the knowledge graph provided by the present invention. The device for locating the system model data in the knowledge graph described below and the method for locating the system model data in the knowledge graph described above can be referenced to each other.
[0148] like Fig.10 As shown, in one embodiment, a device for locating system model data in a knowledge graph includes a fuzzy matching module 1010 , a feature extraction module 1020 , a feature fusion module 1030 and a data locating module 1040 .
[0149] The fuzzy matching module 1010 is used to obtain the relationship triples and attribute triples of the entity nodes from the architecture knowledge graph, and perform fuzzy matching queries through the entity names to determine multiple entities with repeated characters between the entity names.
[0150] The feature extraction module 1020 is used to embed the architecture knowledge graph and the model data to be located through a relational graph convolutional network to map the model data to be located into a low-density vector space, and extract structural feature information by integrating entity adjacency information and entity relationship information to obtain pre-aligned entity pairs.
[0151] The feature fusion module 1030 is used to extract attributes from entities in the pre-aligned entity pairs, and perform feature fusion on the pre-aligned entity pairs based on the extracted attribute information to obtain a first entity pair with the highest fusion feature ranking.
[0152] The data positioning module 1040 is used to perform confidence verification on the entities in the first entity pair to obtain the target entity, and iteratively input the target entity into the relationship graph convolutional network to expand the training set to obtain the spatial position of the model data to be located in the architecture knowledge graph.
[0153] The relationship triples are composed of different architecture entities and entity relationships between different architecture entities. The architecture entity is the system knowledge of the software in the system global knowledge base, including at least the output relationship, input relationship, associated information and associated system of the software.
[0154] The attribute triplet consists of an architecture entity, the attribute name of the architecture entity, and the corresponding attribute value. The target entity is the architecture entity corresponding to the model data to be located in the architecture knowledge graph, and the spatial position of the model data to be located in the architecture knowledge graph is the spatial position of the target entity in the architecture knowledge graph.
[0155] In this embodiment, the fuzzy matching module of the positioning device of the system model data in the knowledge graph provided by the present invention is specifically used for:
[0156] Perform fuzzy matching query by entity name and determine whether there is a first entity with the same name in the architecture knowledge graph. If so,
[0157] Directly generate a second entity pair based on the first entity, and input the second entity pair into the relational graph convolutional network.
[0158] A plurality of second entities having repeated characters with the entity name are obtained from the architecture knowledge graph.
[0159] The first entity and the second entity are both system knowledge of the software in the system global knowledge base.
[0160] In this embodiment, the device for locating system model data in a knowledge graph provided by the present invention further includes a character overlap verification module for:
[0161] The plurality of second entities are sorted according to the character overlap, to obtain second entities sorted from high to low character overlap.
[0162] A second entity whose character overlap ranking is not lower than a first threshold is selected from the second entities ranked from high to low in character overlap.
[0163] A third entity pair is generated based on the second entity whose character overlap ranking is not lower than the first threshold, and is used as an input of the relational graph convolutional network.
[0164] In this embodiment, the device for locating system model data in the knowledge graph and the feature extraction module provided by the present invention are specifically used for:
[0165] Aggregate different entity node information according to different entity relationships, and divide the entity nodes in the architecture knowledge graph into input and output to generate a structural feature vector;
[0166] Calculate the spatial distance between entities based on the structural feature vector to obtain the entity with the closest spatial distance, and the pre-aligned entity pair is composed of the entities with the closest spatial distance;
[0167] The calculation formula of spatial distance is:
[0168]
[0169] Where D s is the spatial distance, e1 and e2 are two entities, and n is the entity structure feature obtained based on the relationship graph convolution network. and These are the entity feature structures of the two entities e1 and e2. is the dimension, l s is the number of layers of the neural network.
[0170] In this embodiment, the device for locating system model data in the knowledge graph provided by the present invention, the feature fusion module is specifically used for:
[0171] Extracting first attribute information of the third entity according to the entity type of the third entity in the pre-aligned entity pair and the attribute feature of the third entity, and constructing a first attribute triplet based on the first attribute information;
[0172] Based on the first attribute triplet, calculating the attribute semantic distance between the third entities, so as to select a fourth entity whose attribute semantic distance is not less than the second threshold from the architecture knowledge graph;
[0173] Among them, the third entity and the fourth entity are both the system knowledge of the software in the system global knowledge base, and the calculation formula of the attribute semantic distance is:
[0174]
[0175] Where D a (e1, e2) is the attribute semantic distance between the pre-aligned entity pairs, 1≤i≤|A|, D i (e1,a i ,e2) is the attribute value similarity measure of the pre-aligned entity pair (e1, PreSame, e2), ω i For attribute a i The influence factor of i ≤1.
[0176] In this embodiment, the data location module of the system model data location device in the knowledge graph provided by the present invention is specifically used for:
[0177] A second entity pair consisting of target entities with the highest confidence and the shortest spatial distance among the multiple first entity pairs is input into a relational graph convolutional network for iterative calculation to obtain the spatial positions between the target entities in the second entity pair.
[0178] Based on the spatial positions between target entities, the spatial positions of the model data to be located in the architecture knowledge graph are obtained.
[0179] In this embodiment, the present invention provides a positioning device for system model data in a knowledge graph, and the propagation rule of the relational graph convolutional network is:
[0180]
[0181] In the formula, is the structural feature of the relationship propagation node, σ is the activation function, is the structural feature of the entity, is the weight matrix of entity features, is the neighbor node feature, is the weight matrix of the corresponding relationship between neighbor node features, is the regularization constant, is a natural number;
[0182] The calculation formula for feature fusion is:
[0183] D(e1,e2)=αD s (e1,e2)+(1-α)D a (e1,e2);
[0184] In the formula, D(e1,e2) is the fusion feature, D s (e1, e2) is the vector space distance between the pre-aligned entity pairs, D a (e1, e2) is the attribute semantic distance between the pre-aligned entity pairs, and α is the weight parameter used to adjust the feature weight of the vector space distance and the attribute semantic distance.
[0185] Fig.11 An example of a physical structure diagram of an electronic device is shown. The electronic device may be a smart terminal, and its internal structure diagram may be as follows: Fig.11 As shown. The electronic device includes a processor, an internal memory and a network interface connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for locating system model data in a knowledge graph is implemented, and the method includes:
[0186] Obtain relationship triples and attribute triples of entity nodes from the architecture knowledge graph, and perform fuzzy matching queries through entity names to identify multiple entities with repeated characters between entity names;
[0187] The architecture knowledge graph and the model data to be located are embedded through the relational graph convolutional network to map the model data to be located into a low-density vector space, and the structural feature information is extracted by integrating the entity adjacency information and the entity relationship information to obtain the pre-aligned entity pairs;
[0188] Extracting attributes from entities in the pre-aligned entity pairs, and fusing features of the pre-aligned entity pairs based on the extracted attribute information to obtain a first entity pair with the highest fused feature ranking;
[0189] The entities in the first entity pair are confidence checked to obtain the target entity, and the target entity is iteratively input into the relational graph convolutional network to expand the training set to obtain the spatial position of the model data to be located in the architecture knowledge graph.
[0190] The relationship triples are composed of different architecture entities and entity relationships between different architecture entities. The architecture entity is the system knowledge of the software in the system global knowledge base, including at least the output relationship, input relationship, associated information and associated system of the software.
[0191] The attribute triplet consists of an architecture entity, the attribute name of the architecture entity, and the corresponding attribute value. The target entity is the architecture entity corresponding to the model data to be located in the architecture knowledge graph, and the spatial position of the model data to be located in the architecture knowledge graph is the spatial position of the target entity in the architecture knowledge graph.
[0192] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present invention, and does not constitute a limitation on the electronic device to which the scheme of the present invention is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0193] On the other hand, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the method for locating the above-mentioned system model data in the knowledge graph.
[0194] In another aspect, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, the method for locating the system model data in the knowledge graph is implemented.
[0195] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory.
[0196] By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0197] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0198] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A method for locating system model data in a knowledge graph, characterized in that: The method comprises: Obtaining relationship triples and attribute triples of entity nodes from the architecture knowledge graph, and performing fuzzy matching query through entity names to determine multiple entities with repeated characters between the entity names; The architecture knowledge graph and the model data to be located are embedded through a relational graph convolutional network to map the model data to be located into a low-density vector space, and structural feature information is extracted by integrating entity adjacency information and entity relationship information to obtain pre-aligned entity pairs; Extracting attributes from entities in the pre-aligned entity pairs, and fusing features of the pre-aligned entity pairs based on the extracted attribute information to obtain a first entity pair with the highest fused feature ranking; Performing confidence checks on entities in the first entity pair to obtain a target entity, and iteratively inputting the target entity into the relationship graph convolutional network to expand a training set, thereby obtaining a spatial position of the model data to be located in the architecture knowledge graph; The relationship triples are composed of different architecture entities and entity relationships between the different architecture entities. The architecture entities are the system knowledge of the software in the system global knowledge base, including at least the output relationship, input relationship, associated information and associated system of the software. The attribute triplet consists of an architecture entity, an attribute name of the architecture entity, and a corresponding attribute value; the target entity is the architecture entity corresponding to the model data to be located in the architecture knowledge graph; the spatial position of the model data to be located in the architecture knowledge graph is the spatial position of the target entity in the architecture knowledge graph; The architecture knowledge graph and the model data to be located are embedded and processed by the relational graph convolution network to map the model data to be located to a low-density vector space, and structural feature information is extracted by integrating entity adjacency information and entity relationship information to obtain pre-aligned entity pairs, including: Aggregating different entity node information according to different entity relationships, and dividing the entity nodes in the architecture knowledge graph into input and output to generate a structural feature vector; Calculating the spatial distance between entities based on the structural feature vector to obtain the entity with the closest spatial distance, wherein the pre-aligned entity pair is composed of the entities with the closest spatial distance; The calculation formula of the spatial distance is: ; In the formula, is the spatial distance, e1 and e2 are two entities, and n is the entity structure feature obtained based on the relationship graph convolution network. and These are the entity feature structures of the two entities e1 and e2. For the dimension, is the number of layers of the neural network; The extracting attributes of the entities in the pre-aligned entity pairs and fusing features of the pre-aligned entity pairs based on the extracted attribute information to obtain a first entity pair with the highest fused feature ranking includes: Extracting first attribute information of the third entity according to the entity type of the third entity in the pre-aligned entity pair and the attribute feature of the third entity, and constructing a first attribute triplet based on the first attribute information; Based on the first attribute triplet, calculating the attribute semantic distance between the third entities, so as to select a fourth entity whose attribute semantic distance is not less than a second threshold from the architecture knowledge graph; Wherein, the third entity and the fourth entity are both the system knowledge of the software in the system global knowledge base, and the calculation formula of the attribute semantic distance is: ; In the formula, is the attribute semantic distance between the pre-aligned entity pairs, , A is the total number of attributes, For the pre-aligned entity pair The attribute value similarity measure of , PreSame is the entity to be aligned, For attributes The impact factor of .
2. The method for locating system model data in a knowledge graph according to claim 1, characterized in that: The step of obtaining the relationship triples and attribute triples of the entity nodes from the architecture knowledge graph and performing fuzzy matching query through the entity names to determine multiple entities with repeated characters between the entity names includes: Perform fuzzy matching query by entity name and determine whether there is a first entity with the same name in the architecture knowledge graph; if so, Directly generate a second entity pair based on the first entity, and input the second entity pair into the relationship graph convolution network; if not, then Acquire, from the architecture knowledge graph, a plurality of second entities having repeated characters with the entity name; The first entity and the second entity are both system knowledge of the software in a system global knowledge base.
3. The method for locating system model data in a knowledge graph according to claim 2, characterized in that: The step of obtaining the relationship triples and attribute triples of the entity nodes from the architecture knowledge graph and performing fuzzy matching query through the entity names to determine multiple entities with repeated characters between the entity names and the entity names, and then comprising: Sorting the plurality of second entities according to character overlap to obtain second entities sorted from high to low character overlap; Selecting a second entity whose character overlap ranking is not less than a first threshold from the second entities ranked from high to low in character overlap; A third entity pair is generated based on the second entity whose character overlap ranking is not lower than the first threshold as an input of the relationship graph convolutional network.
4. The method for locating system model data in a knowledge graph according to claim 1, characterized in that: The confidence check is performed on the entities in the first entity pair to obtain a target entity, and the target entity is iteratively input into the relationship graph convolutional network to expand the training set, and the spatial position of the model data to be located in the architecture knowledge graph is obtained, including: Inputting a second entity pair consisting of the target entities with the highest confidence and the shortest spatial distance among the plurality of first entity pairs into the relationship graph convolution network for iterative calculation to obtain the spatial positions between the target entities in the second entity pair; Based on the spatial positions between the target entities, the spatial positions of the model data to be located in the architecture knowledge graph are obtained.
5. The method for locating system model data in a knowledge graph according to claim 1, characterized in that: The propagation rule of the relational graph convolutional network is: ; In the formula, is the structural feature of the relationship propagation node, is the activation function, is the structural feature of the entity, is the weight matrix of entity features, is the neighbor node feature, is the weight matrix of the corresponding relationship between neighbor node features, is the regularization constant, is a natural number, R is a real number set, r is the number of the neighbor node feature, and Respectively represent the numbers of entity nodes and their corresponding relationship propagation nodes; The calculation formula of the feature fusion is: ; In the formula, To fusion features, is the vector space distance between the pre-aligned entity pairs, is the attribute semantic distance between the pre-aligned entity pairs, is the weight parameter used to adjust the feature weight of the vector space distance and the attribute semantic distance.
6. A device for locating system model data in a knowledge graph, characterized in that: A method for locating system model data in a knowledge graph according to any one of claims 1 to 5, the device comprising: A fuzzy matching module is used to obtain the relationship triples and attribute triples of the entity nodes from the architecture knowledge graph, and perform fuzzy matching query through the entity name to determine multiple entities with repeated characters between the entity names; A feature extraction module embeds the architecture knowledge graph and the model data to be located through a relational graph convolutional network to map the model data to be located into a low-density vector space, and extracts structural feature information by integrating entity adjacency information and entity relationship information to obtain pre-aligned entity pairs; A feature fusion module, used to extract attributes from entities in the pre-aligned entity pairs, and perform feature fusion on the pre-aligned entity pairs based on the extracted attribute information, so as to obtain a first entity pair with the highest fusion feature ranking; A data positioning module, used to perform confidence verification on the entities in the first entity pair to obtain a target entity, and iteratively input the target entity into the relationship graph convolutional network to expand the training set, and obtain the spatial position of the model data to be positioned in the architecture knowledge graph; The relationship triples are composed of different architecture entities and entity relationships between the different architecture entities. The architecture entities are the system knowledge of the software in the system global knowledge base, including at least the output relationship, input relationship, associated information and associated system of the software. The attribute triplet consists of an architecture entity, an attribute name of the architecture entity and a corresponding attribute value. The target entity is the architecture entity corresponding to the model data to be located in the architecture knowledge graph, and the spatial position of the model data to be located in the architecture knowledge graph is the spatial position of the target entity in the architecture knowledge graph.
7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Cross-language knowledge graph entity alignment method based on GCN twinning network
CN110472065A
Internet food entity alignment method and system based on graph neural network
CN113342809A
Collaborative relation graph-based recommendation method and related device
CN114461929A
Method and device for positioning system model data in knowledge graph
CN117786132A