Neural architecture search-based knowledge graph link prediction method

By converting the knowledge graph link prediction task into image classification task and automatically searching for the CNN architecture using particle swarm algorithm, the problem of high design cost of deep neural network architecture in the existing technology is solved, and automated search and architecture optimization of the link prediction model are realized.

WO2025112062A1PCT designated stage expired Publication Date: 2025-06-05UNIV OF ELECTRONICS SCI & TECH OF CHINA +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/135972
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing knowledge graph link prediction methods rely on artificially designed deep neural network architectures, which are costly and inefficient, making it difficult to support high design costs in most scenarios.

Method used

The network architecture search method based on evolutionary algorithm is adopted to convert the knowledge graph link prediction task into image classification task, and the optimized CNN architecture is automatically searched using particle swarm algorithm (PSO).

Benefits of technology

Automatic search of link prediction models is realized, reducing dependence on expert design and manual parameter adjustment, improving architectural search efficiency, and generating a network architecture that is more in line with the knowledge of existing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023135972_05062025_PF_FP_ABST
    Figure CN2023135972_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the field of knowledge graph reasoning, and disclosed is a neural architecture search-based knowledge graph link prediction method. By incorporating the characteristics of existing mature networks, certain constraints are imposed on changes in network search, for example, the sequential order of a convolution kernel layer and a pooling layer in a same cell is constrained, and a parallel / serial structure of a network layer in a same cell is also constrained by distance, such that a found architecture is more in line with existing empirical cognition, thereby reducing the probability of searching for poor networks, saving the computational overhead, and improving the search efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

A knowledge graph link prediction method based on network architecture search Technical Field

[0001] The present invention relates to the field of knowledge graph reasoning, and in particular to a knowledge graph link prediction method based on network architecture search. Background Art

[0002] The knowledge graph, a concept formally proposed by Google in 2012, is a graphical data structure used to represent and organize knowledge. It represents real-world entities, concepts, and their relationships in a graph format, providing crucial semantic information for machine understanding and reasoning. Knowledge graphs have since found widespread application in intelligent question-answering, natural language processing, and personalized recommendations. A core task of knowledge graphs is reasoning—inferring the relationships between nodes in the graph, primarily in link prediction, which involves predicting missing entities / relationships or implicit relationships within the knowledge graph.

[0003] Currently, there are several major categories of link prediction methods:

[0004] Similarity-based methods: These methods predict links based on the similarity of entity attributes or relationships. Common similarity calculation methods include cosine similarity, Euclidean distance, and Jaccard similarity. This method is relatively simple but performs poorly.

[0005] Path-based methods: These methods consider direct or indirect paths between entities and discover connections between them through graph traversal or graph neural networks. Common examples include random walk-based algorithms and graph neural network-based methods.

[0006] Representation learning-based methods: These methods transform the link prediction problem into a similarity calculation in a vector space by learning representation vectors for entities and relations. Typical methods include knowledge representation learning models such as TransE, TransR, and TransD.

[0007] Deep learning-based methods: These methods use deep learning models to solve link prediction problems, including convolutional neural networks, recurrent neural networks, attention mechanisms, and other technologies.

[0008] Among the above methods, those based on deep neural networks are the most widely adopted. Currently, the neural network-based methods for link prediction include those based on convolutional neural networks (ConvE, InteractE) and those based on graph neural networks (GNN, GCN). Knowledge graphs consist of nodes (entities) and edges (relations). Deep neural network-based methods first process the data in the knowledge graph into triplets (entity1, relation, entity2), representing a pair of instances in the graph. The neural network is then trained to output a predicted object. Currently, various deep neural networks are being applied to link prediction, most of which use convolutional neural networks (CNNs). The ConvE model employs a single-layer CNN to capture the deep connections between entities in the graph. InteractE and graph convolutional neural networks (GCNs) also employ CNNs. However, the CNN designs in these models are relatively simple and do not fully utilize their capabilities. There is still room for improvement in CNN-based inference models. By improving the existing architecture, a more robust CNN-based inference (link prediction) model can be achieved.

[0009] However, building deep neural network architectures typically requires meticulous manual design. Examples include the mature Inception, Yolo, and XCeption series. Figure 1 shows the Inception series, a superior architecture developed after a significant amount of time, computational resources, and human resources. However, in most scenarios, designing an excellent, targeted architecture at such a high cost is difficult.

[0010] In recent years, research has begun to address the high cost of designing network architectures. This research area, known as Network Architecture Search (NAS), can automatically find optimal architectures for specific scenarios. Currently, the most common NAS research is focused on image classification tasks. NAS uses reinforcement learning or evolutionary algorithms to search for network architectures, but its application to knowledge graph link prediction is still relatively new, and the use of evolutionary algorithms in other research problems is also at a relatively early stage.

[0011] The input of the link prediction task can be regarded as a 1-channel image through matrix transformation, and the image classification task can use NAS to automatically build a model. The present invention is a link prediction method based on network architecture search combined with evolutionary algorithm. Technical issues

[0012] The present invention provides a knowledge graph link prediction method based on network architecture search, which can be well applied to the problem of knowledge graph link prediction. The link prediction task of a knowledge graph is essentially a classification problem. Given an input entity entity1 and a relation relation, another entity entity2 is predicted, requiring that entity1 and entity2 have a relation relation. The present invention combines the input (entity1 and relation) into a two-dimensional matrix through entity and relationship embedding and splicing operations. The two-dimensional matrix can also be regarded as a one-channel image representation. The basis of the present invention is to regard the link prediction task as an image classification task and use a network architecture search based on an evolutionary algorithm to obtain an excellent CNN-based network for the link prediction task. Technical Solutions

[0013] According to a knowledge graph link prediction method based on network architecture search of the present invention, the method comprises the following steps:

[0014] Step 1: Link prediction in deep learning has two input data: entity e1 and relation r. Entity e1 and relation r are respectively embedded through two embedding matrices, Entity Embedding and Relation Embedding, to become two vectorized e_embed and r_embed. The entity e1 includes: the name of the person and the place where it appears, and the relation r is the connection between people.

[0015] Step 2: Convert the vectorized e_embed and r_embed into matrices to obtain matrixed e_embed and r_embed;

[0016] Step 3: Concatenate the matrixed e_embed and r_embed to obtain a graph with the number of channels x height x width of 1 x H x W.

[0017] Step 4: Perform a convolution operation with an output channel number of C on the graph obtained in step 3 to obtain a graph of size C x H x W.

[0018] Step 5: The graph obtained in step 4 is passed through N processing modules in sequence. Each processing module includes a Normal module and a Reduction module in sequence. Padding operations are added to the Normal module to ensure that the input and output shapes of each processing module remain unchanged.

[0019] Step 6: Expand the graph output from step 5 into a one-dimensional vector; then pass it through the fully connected layer FC to convert it into an embedding vector; then multiply it by the transpose of the entity embedding matrix to obtain the probability or score of the predicted entity e2 corresponding to each entity category. The one with the largest probability or score is the prediction result;

[0020] The N processing modules and the fully connected layer FC form a network, and the parameter calculation method of the network is:

[0021] Step B1: fix all parameter settings and operation set settings in the network; initialize the population, where each individual in the population is a network parameter;

[0022] Step B2: Calculate the fitness of each individual. The maximum fitness value among the individuals is the global optimal solution of the current population.

[0023] Step B2: Encode the network parameters using an adjacency matrix. Each individual includes the current position and evolution speed. The current position represents the structure and node parameters of the network.

[0024] Step B4: Update the evolution speed of each individual according to the individual's fitness.

[0025] For individual i, the speed The update formula is as follows:

[0026]

[0027] in, represents the inertia weight of the individual motion, and is the acceleration factor, and is a random number between 0 and 1, represents the individual's historical optimal solution, represents the global optimal solution, Indicates the location of each individual;

[0028]

[0029] represents the current fitness of individual i, represents the historical optimal fitness of individual i in the first k rounds, represents the historical optimal fitness of individual i in the first k rounds;

[0030] Step B5: Update the individual position according to the updated speed;

[0031]

[0032] Step B6: Decode the new individual position and return to step B3 until the set number of cycles is reached, and the global optimal solution is obtained as the output result.

[0033] Furthermore, the decoding method in step B6 is:

[0034] (1) Perform topological sorting on the adjacency matrix to obtain a node calculation sequence;

[0035] (2) Determine whether each node is connected, and perform operations and calculations on the connected nodes in the order of the obtained calculation sequence. Beneficial effects

[0036] The link prediction task is transformed into a graph classification task. At the same time, combined with the network architecture search method for CNN, the automated search of link prediction models is proposed and implemented for the first time, no longer relying on expert design and manual parameter adjustment.

[0037] To solve the problem of network architecture search, the relatively superior PSO (Particle Swarm Optimization) was used instead of the commonly used (GA) genetic algorithm to improve the architecture search efficiency. A network encoder and decoder based on the adjacency matrix was proposed and implemented to cooperate with the PSO algorithm.

[0038] Combining the characteristics of existing mature networks, certain constraints are imposed on changes in network search. For example, the order of convolution kernel layers and pooling layers in the same cell is constrained, and the parallel / serial structure of network layers in the same cell is also subject to distance constraints. This makes the search architecture more consistent with existing empirical knowledge, reduces the probability of searching for poor networks, saves computational overhead, and improves search efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic diagram of the Inception V4 model.

[0040] Figure 2 is a schematic diagram of network architecture search.

[0041] Figure 3 shows a Normal Cell with 5 nodes.

[0042] FIG4 is a schematic diagram of a disconnected network.

[0043] FIG5 is a flow chart of the particle swarm algorithm in the present invention.

[0044] Figure 6 shows the batch randomly generated network coding (1).

[0045] Figure 7 shows the batch randomly generated network coding (2).

[0046] Figure 8 shows the link prediction architecture based on network architecture search.

[0047] Figure 9 shows the shape transformation of the embedding vector.

[0048] Figure 10 shows the concatenated entity and relationship representations.

[0049] FIG11 shows the output prediction entity e2. Best Mode for Carrying Out the Invention

[0050] A knowledge graph of club member relationships is established; entities include e1: person, club, and place; relationship types include: member relationship, family relationship, classmate relationship, similar club, and fellow-township relationship; entity e1 represents a member, and relationship r represents a friend relationship: family relationship, classmate relationship, fellow-township relationship, member relationship, etc., and the final predicted output e2 is obtained by inputting e1 and r. The other entity with which the e1 member has a family relationship / friend relationship / member relationship / fellow-township relationship may be a member or a club; the present invention is further explained in conjunction with the accompanying drawings.

[0051] Step 1 first sets a fixed architecture, which is equivalent to an abstract network structure that determines the search range of the subsequent specific network. As shown in Figure 1-2, the fixed network architecture set by the present invention consists of several parts: (1) The Normal Cell module has multiple nodes and edges, where each node represents the output or output of a basic operation (Op), and each edge represents a basic operation (Conv3x3, Conv5x5, MaxPool, AvgPool, Skip jump link, None empty operation, etc.). (2) The Reduction Cell module has a fixed Pool layer inside and a specified step size of 2, which is used to reduce the shape of the data. (3) The FC module is used to output the corresponding category.

[0052] The number and arrangement of modules in the entire fixed architecture are fixed, but the operations within each module are variable. The entire architecture will present different effects due to different combinations of internal modules.

[0053] Step 2 requires setting the module (Cell) parameters and operation (Op) set. Figure 1-3 shows the internal details of a Normal Cell. The module contains multiple nodes, and there are edges between the nodes. Each edge has a corresponding type / operation (Conv3x3, Conv5x5, MaxPool, AvgPool, Skip link, None, etc.). Step 2 requires setting the number of nodes in each Normal Cell and the types of edges / operations. The greater the number of nodes and operation types, the more possible combinations there are, and the larger the model's search space is, but the corresponding search overhead will also increase. Figure 1-3 shows a 5-node Normal Cell.

[0054] Step 3 requires encoding and decoding the specific network and designing the corresponding encoder and decoder. This invention uses an adjacency matrix for encoding and decoding. Assuming that there are M nodes in each cell, the adjacency matrix will be a three-dimensional matrix (Matrix) of size M x M x OpSize (OpSize is the size of the operation set). Each value in the matrix corresponds to the probability of the edge between two nodes belonging to a certain type (Op). The following three-dimensional matrix represents the encoding in Figure 1-3. The first element value in the first row and second column is 1, indicating that there is an edge of type Op 1 (Conv3x3) between the node numbered 1 and the node numbered 2, and so on.

[0055]

[0056] Furthermore, the selection of edges draws on the design experience of mature models. The farther the nodes are from each other, the lower the probability of connecting them, which helps maintain the local serial structure. There are also constraints on the order of pooling and convolutional layers. In the same cell, the convolutional layer is more likely to come before the pooling layer.

[0057]

[0058] distance(a, b) represents the distance between nodes a and b, and prob(i) represents the probability of using operation i.

[0059] A node can have edges with multiple other nodes. The idea of ​​the encoder is to store the network structure through the adjacency matrix. The design of the decoder is more complex than the encoding area. It is necessary to restore an adjacency matrix back to a network that can actually be used for training and testing. There are two main steps: (1) Topologically sort the adjacency matrix to obtain a node calculation sequence. Taking the Normal Cell in Figure 1-3 as an example, the node calculation sequence obtained is as follows:

[0060]

[0061] Because the network may be disconnected during the evolution process, it is necessary to determine whether it is connected (Figure 1-4 shows the disconnected situation), and then perform operations and calculations according to the order in the above calculation sequence order for each node.

[0062] The evolutionary algorithm used in step 4 is the Particle Swarm Optimization (PSO), an optimization algorithm based on swarm intelligence, inspired by the behavior of natural groups such as flocks of birds and schools of fish. In PSO, candidate solutions are called "particles," which move through the solution space, searching for the optimal solution based on their own experience and that of the group. In this invention, the individuals in the PSO correspond to network codes. By initializing the network code (Matrix), a group of individuals is generated.

[0063] The steps of the particle swarm algorithm in the present invention are shown in Figures 1-5.

[0064] Figures 1-6 and 1-7 show a batch of randomly initialized network codes / individuals (when the number of nodes in a cell is specified to be 8). In the particle swarm algorithm, each individual i has a state / position X i and speed V i , the network coding corresponds to the position of the individual, and the initialization speed is also required.

[0065] In step 5, the fitness of each individual needs to be calculated in the particle swarm algorithm. The present invention regards the quality of the network as the fitness of the individual. It needs to be decoded into a network that can be trained and operated first (decoding details are in step 3). Because it is a link prediction task, indicators such as Hits@1, Hits@3, Hits@10 are used as fitness, and the fitness is obtained after training for a certain number of Epochs. After initialization, the fitness of the first round is obtained, and the maximum fitness of all individuals is selected as the global optimal fitness (gBest). The fitness of each individual is used as its own historical optimal fitness (pBest). In the early stage, the fitness is mainly trained with a small number of Epochs, mainly to see the trend of change. In the later stage, in order to calculate the fitness more accurately, the number of training rounds needs to be increased.

[0066] Step 6: The core idea of ​​the particle swarm algorithm is to find the optimal solution through group movement. After obtaining the fitness in the previous step, the speed of each individual needs to be updated according to the fitness.

[0067] For individual i, the speed V i The update formula is as follows:

[0068]

[0069] Where w represents the inertia weight of the individual motion, c1 and c2 are acceleration coefficients, and rand1 and rand2 are random numbers between 0 and 1.

[0070] A larger w value accelerates the search for network architectures, while a smaller value facilitates a more detailed search. C1 represents the influence of an individual's own historical optimal solution, while c2 represents the influence of the individual's global optimal solution. Larger values ​​for these two factors lead to faster convergence in the early stages of the search, while smaller values ​​for these two factors focus more on local search. rand1 and rand2 introduce random perturbations to the search, helping to maintain algorithm diversity and avoid falling into local optimal solutions. Simultaneously, the global optimal solution (gBest) and the individual's historical optimal solution (pBest) are updated. Fitness(i) represents the current fitness of individual i, and pBest(i, k) represents the historical optimal fitness of individual i over the previous k rounds.

[0071]

[0072] In step 7, the individual position needs to be updated according to the adjusted individual speed.

[0073] For individual i, position X i The updated announcement is as follows:

[0074]

[0075] Updating individual locations means updating the corresponding network modules.

[0076] Step 8: Repeat steps 5, 6, and 7 for a specified number of rounds to finally obtain the result, and select the global optimal solution as the search result.

[0077] Step 9: Decode the global optimal solution to obtain the network model. The link prediction model architecture is shown in Figure 1-8. Compared to the architecture in Figure 1-2, it shows more details in the input and output processing. At the input, two outputs are embedded and concatenated, and at the output, the entity is restored.

[0078] The process of applying the model to the link prediction task is as follows:

[0079] (1) Link prediction in deep learning has two input data: entity e1 and relation r. In order to vectorize the two inputs, they need to pass through two embedding matrices, Entity Embedding and Relation Embedding, respectively, to become two vectorized e_embed and r_embed.

[0080] (2) The vectorized entities and relationships (e_embed and r_embed) need to be reshaped, as shown in Figure 1-9.

[0081] (3) In order to use convolutional neural networks to extract the potential connections (features) between entities and relationships, the two need to be spliced ​​together, as shown in Figure 1-10. After splicing, a 1-channel graph is formed, and the shape of the graph is 1 x H x W (number of channels x height x width).

[0082] (4) After obtaining the 1-channel image, in order to extract features from multiple dimensions, a convolution operation with an output channel number of C is used (here C is set to 32).

[0083] (5) The C x H x W graph is obtained in the previous step, and then passes through N Normal Cells and Reduction Cells. In order to keep the input and output shapes of the data compatible in the calculation of the entire architecture, the Normal Cell adds padding to ensure that the shape remains unchanged after passing through each Normal Cell. However, the height and width of the graph will be reduced after passing through the Reduction Cell.

[0084] (6) To obtain the result entity e2 of the link prediction, three operations are required. First, the graph (shape C x H x W) is transformed into a one-dimensional vector. The specific transformation process is shown in the two upper subgraphs in Figure 1-11. Then, it passes through the fully connected layer, which converts this one-dimensional vector into an embedding vector (representing the embedding representation of the final predicted entity e2). Finally, it is multiplied by the transpose of the entity embedding matrix to obtain the probability / score of e2 corresponding to each entity category. Each element value in the final vector in Figure 1-11 represents each entity. Entity 8 has the highest probability, so the final output e2 of the input e1 and r is entity 8. It is considered that e1 and entity 8 have a relationship r, and the prediction result is obtained.

Claims

1. A knowledge graph link prediction method based on network architecture search, the method comprising the following steps; Step 1: There are two pieces of input data for link prediction in deep learning: entity e1 and relation r; the entity e1 and the relation r are respectively passed through two embedding matrices, Entity Embedding and Relation Embedding, to become two vectorized e_embed and r_embed; The entity e1 includes: person's name, place of appearance, and the relation r is the connection between people; Step 2: Convert the vectorized e_embed and r_embed into matrices to obtain matrixized e_embed and r_embed; Step 3: Concatenate the matrixized e_embed and r_embed to obtain a graph with the number of channels x height x width of 1 x H x W; Step 4: Perform a convolution operation with an output channel number of C on the graph obtained in Step 3 to obtain a graph of C x H x W; Step 5: The graph obtained in Step 4 is successively passed through N processing modules. Each processing module sequentially includes a Normal module and a Reduction module. A padding operation is added in the Normal module to keep the input and output shapes of each processing module unchanged; Step 6: Unfold the graph output in Step 5 into a one-dimensional vector; then pass it through a fully connected layer FC to convert it into an embedding vector; then multiply it by the transpose of the entity embedding matrix to obtain the probability or score corresponding to each entity category of the predicted entity e2, and the one with the largest probability or score is the prediction result; The N processing modules and the fully connected layer FC form a network, and the parameter calculation method of this network is: Step B1; Fix all parameter settings and operation set settings in the network; initialize the population, and each individual in the population is a network parameter; Step B2: Calculate the fitness of each individual, and the maximum fitness value in the individual is the global optimal solution of the current population; Step B2: Encode the network using the adjacency matrix method for the network parameters. Each individual includes the current position and the evolution speed, and the current position represents the structure and node parameters of the network; Step B4: Update the evolution speed of each individual according to the fitness of the individual, For individual i, the speed The update formula is as follows: ; Among them, Represents the inertia weight of individual motion, And is the acceleration coefficient, and is a random number between 0 and 1, Represents the historical optimal solution of an individual, Indicates the global optimal solution, represents the position of each individual; ; Denote the current fitness of individual i, Denote the historical best fitness of individual \(i\) in the first \(k\) rounds. represents the historical optimal fitness of individual i in the previous k rounds; Step B5: Update the individual position according to the updated speed; ; Step B6: Decode the new individual position, and return to Step B3. After reaching the set number of loops, obtain the global optimal solution as the output result.

2. A knowledge graph link prediction method based on network architecture search according to claim 1, characterized in that, the decoding method in Step B6 is: (1) Perform a topological sort on the adjacency matrix to obtain a calculation sequence of nodes; (2) Judge whether each node is connected, and perform operation calculations on the connected nodes in sequence according to the order in the obtained calculation sequence.

Citation Information

Patent Citations

  • Knowledge-driven business operation graph construction method

    CN112507136A

  • Link prediction method based on knowledge graph embedding

    CN113360286A

  • Knowledge graph link prediction method

    CN116757283A

  • Knowledge graph link prediction method based on network architecture search

    CN117591720A