Lightweight super-relation knowledge graph link prediction method based on reversible neural network

By introducing a lightweight super-relational knowledge graph link prediction method of reversible neural networks, the problems of excessive resource consumption and long training time in the prior art are solved, and efficient link prediction and good compatibility with additional information are achieved.

CN120338026APending Publication Date: 2025-07-18SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510415244.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing neural network methods consume too much resource, long training time, and difficult to deal with a large amount of additional information in hyperrelational knowledge graph link prediction.

Method used

The lightweight super-relational knowledge graph link prediction method based on reversible neural network is adopted, and feature space conversion is performed through the reversible transformation module, combining the graph fact embedding module and the noise decoupling module to reduce training resource consumption and improve training efficiency.

Benefits of technology

While maintaining high performance, it significantly reduces training resource consumption and time, improves compatibility with hyperrelational facts, especially when handling more additional information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338026A_ABST
    Figure CN120338026A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight hyperrelational knowledge graph link prediction method based on a reversible neural network. The method comprises the following steps: firstly, building and training a link prediction network comprising a training data preprocessing module, a graph fact embedding module, a reversible transformation module, a noise decoupling module and a decoder module; in the link prediction process, data are firstly changed into training examples and labels through the training data preprocessing module, and atlas fact features are obtained after the data pass through the atlas fact embedding module; noise-containing potential codes are obtained through the forward process of the reversible transformation module, noise-removed potential codes are obtained through the noise decoupling module, and finally complete atlas fact features are obtained through the reverse process of the reversible transformation module, so that missing elements are obtained. Compared with the prior art, the super relation knowledge graph link prediction method has the advantages that resources and time consumed by training are reduced while link prediction performance is kept, and the super relation knowledge graph link prediction method has better compatibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge graph in natural language processing, and specifically relates to a lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network. Background Art

[0002] A knowledge graph is a way to represent and store knowledge through a graph data structure, and uses triples (h, r, t) to describe specific knowledge. With the development of knowledge graph technology, it has become more convenient to obtain and construct knowledge graphs, making the application scenarios of knowledge graphs very extensive. However, although there are already a large number of entity and relationship data in the current knowledge graph, due to reasons such as limited data sources for constructing knowledge graphs and defects in construction technologies, there must be a certain degree of link missing problems in knowledge graphs.

[0003] A knowledge graph that describes knowledge in the form of triples faces many limitations such as being unable to contain context information and being difficult to capture the hierarchy between entities when representing real-world information. To solve these limitations, hyper-relation knowledge graphs effectively expand the dimension of knowledge representation by introducing additional information on the basis of triples, thereby accurately representing real-world information. Inevitably, hyper-relation knowledge graphs also have link missing problems. Therefore, in recent years, researchers have conducted extensive research on the link missing problems in hyper-relation knowledge graphs. Among these works, the link prediction methods implemented using neural network technology stand out with their excellent performance and become the current best-performing technical solutions.

[0004] Currently, the existing hyper-relation knowledge graph link prediction methods using neural networks are mainly divided into two categories. One category is the method based on convolutional neural networks, which mainly extracts the fusion features between triples and additional information by splicing triples and additional information into two-dimensional matrices respectively and then using convolution and fully connected operations; the other category is the method based on graph neural networks and Transformer architectures. This type of method uses the message passing mechanism of graph neural networks to gradually update node features to capture the complex graph structure relationships in hyper-relation knowledge graphs, and then uses Transformer as a decoder to predict the missing elements. However, although neural network-based methods can usually achieve excellent prediction performance, whether using convolutional neural networks or graph neural networks, they all face the problems of excessive consumption of training resources and long training time due to complex model structures and a large number of parameters, and these methods cannot handle hyper-relation facts with more additional information well. Summary of the Invention

[0005] The main objective of the present invention is to overcome the drawbacks and deficiencies of the prior art and provide a lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network. By introducing a reversible neural network, while retaining the high-performance advantages of the neural network, it effectively reduces the consumption of training resources and significantly shortens the training time required. The reversible neural network uses reversible transformation operations, enabling it to not store the activation values of all intermediate layers during the backpropagation process, thereby significantly reducing the consumption of video memory resources and also reducing the time overhead of data transmission and memory access operations. At the same time, the method proposed in the present invention uses the embedding vector summation operation of the graph fact embedding module, making the number of feature channels of the obtained graph fact features independent of the logarithm of the additional information of the hyper-relation facts, thus enabling the processing of hyper-relation facts with more additional information.

[0006] To achieve the above objective, the present invention adopts the following technical solutions:

[0007] A lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network, the lightweight hyper-relation knowledge graph link prediction method comprising the following steps:

[0008] S1. Construct a lightweight hyper-relation knowledge graph link prediction network based on a reversible neural network, the hyper-relation knowledge graph link prediction network including a training data preprocessing module, a graph fact embedding module, a reversible transformation module, a noise decoupling module, and a decoder module connected in sequence; wherein, the reversible transformation module is composed of a plurality of reversible transformation layers connected in sequence;

[0009] S2. Input hyper-relation knowledge graph data in a specific domain, perform data preprocessing operations on the hyper-relation knowledge graph data in the specific domain through the training data preprocessing module to obtain training instances and corresponding labels y, and perform label smoothing operations on the labels y using label smoothing techniques to obtain smoothed labels y s ; the hyper-relation knowledge graph data in the specific domain comes from medical knowledge graphs for disease diagnosis, financial knowledge graphs for risk assessment, educational knowledge graphs for personalized learning, and industrial knowledge graphs for fault diagnosis; the above hyper-relation knowledge graph data in the specific domain is a knowledge graph constructed by human experts using data acquisition devices, computing devices, storage devices, and data annotation tools;

[0010] S3. Perform embedding operations on the training instances through the graph fact embedding module to obtain incomplete graph fact features

[0011] S4. Perform feature space transformation operations on the incomplete graph fact features through the forward process of the reversible transformation module to obtain latent encodings in the latent space

[0012] S5. Potential coding through noise decoupling module The noise in the image is decoupled to obtain the denoised latent code

[0013] S6. Potential encoding through the inverse process of the reversible transformation module Perform feature space transformation to obtain complete graph fact features

[0014] S7. Complete graph fact features Factual characteristics of incomplete maps Subtract to get missing element features Then the decoder module is used to decode the missing element features. Perform a decoding operation to obtain a probability distribution P, and obtain the missing elements of the link prediction based on the probability distribution;

[0015] S8. Train a lightweight hyper-relational knowledge graph link prediction network based on a reversible neural network through hyper-relational knowledge graph data in a specific field, and iterate the training in an end-to-end manner. In each iteration, the probability distribution P and the smooth label y are used to predict the link prediction network. s Calculate the loss value and minimize the loss function to train the network, while updating the parameter values in the network through back propagation until the network converges;

[0016] S9. Input the incomplete hyper-relational knowledge graph into the trained hyper-relational knowledge graph link prediction network, and obtain the missing elements in the incomplete hyper-relational knowledge graph according to the output probability distribution.

[0017] Furthermore, the working process of the training data preprocessing module is as follows:

[0018] The hyper-relational knowledge graph data of a specific field is randomly divided into training set, validation set and test set in a ratio of 8:1:1, and the graph facts in the training set are masked to obtain training instances. And the corresponding label y, and use the label smoothing operation to smooth the label y to obtain the smoothed label y s; To ensure the generalization ability and stability of network training, a stratified sampling method is used in the random partitioning operation to ensure that the proportions of different types of graph facts in the training set, validation set, and test set are consistent. The mask operation is designed to guide the network to learn the semantic features and context relationships of each entity and relationship by randomly masking each element in the facts. The purpose of the label smoothing operation is to alleviate the problem of network overfitting. By introducing an appropriate smoothing factor ∈′ for the labels, the network can be effectively prevented from being too sensitive to a single label;

[0019] Among them, the hyper-relation knowledge graph consists of multiple triple facts and multiple hyper-relation facts, and the hyper-relation fact is expressed as where (h, r, t) represents the main triple, h represents the head entity, r represents the relation connecting the head entity and the tail entity, and t represents the tail entity. represents m pairs of additional information, and q i represents the relation of the i-th pair of additional information, and v i represents the entity of the i-th pair of additional information; h, t, v1, …, v m ∈ ε, r, q1, …, ε is the set of all entities, is the set of all relations; when m = 0, the hyper-relation fact is simplified to the triple fact Therefore, the triple fact and the hyper-relation fact are collectively referred to as graph facts

[0020] Among them, the mask operation process is as follows: For a hyper-relation fact with m pairs of additional information Mask each element in the hyper-relation fact with the mask token [MASK] to create (2m + 3) training instances for the hyper-relation fact ; For a triple fact Mask each element in the triple fact with the mask token [MASK] to create 3 training instances for the triple fact ; Create a label y for each training instance, that is, define where according to the serial number of the masked element in the hyper-relation knowledge graph data, the corresponding position of the label y is set to 1, and the other positions of the label y are set to 0; For the label y, there is where |ε| represents the number of entities, represents the number of relations;

[0021] Among them, the process of the label smoothing operation is as follows: For the t-th element y in the label y t , if the position corresponding to the t-th element is the position of the target entity, then set the value of this element to y t = 1 - ∈′, and set the values of the elements in other positions to where the label smoothing rate Thus, convert the label y to the smoothed label y s .

[0022] Furthermore, the working process of the graph fact embedding module is as follows:

[0023] Given the entity set ε and the relationship set of the hyper-relation knowledge graph data in a specific domain Generate the entity embedding matrix E and the relationship embedding matrix R respectively, and randomly initialize the entity embedding matrix E and the relationship embedding matrix R, where where d is the dimension of the embedding vector and also the number of feature channels; Obtain the embedding vectors of each element in the incomplete graph fact from the entity embedding matrix E and the relationship embedding matrix R, and obtain the feature vector of the incomplete graph fact through the element-wise summation operation In the initialization process of the embedding matrix, the normal distribution is usually used for initialization to ensure that the initial vectors of each entity and relationship have a good distribution. At the same time, in order to enhance the learning ability of the network, a pre-trained embedding matrix can be introduced in the initialization process to optimize and update the embedding matrix. The design of the element-wise summation operation aims to aggregate the features of different entities and relationships, so as to obtain a compact representation of the entire graph fact and further improve the expression effect and discrimination ability of the embedding vector;

[0024] Among them, the process of the element-wise summation operation is as follows: Assume that the incomplete graph fact is missing an entity, there are m pairs of additional information, and there are u relationships. The embedding vectors of the u relationships are represented as (r1, r2, ···, r u ); Then there will be u entities, and the embedding vectors of the u entities are represented as (e1, e2, ···, e u ), where Obtain the incomplete graph fact feature by adding each embedding vector element-wise The element-wise summation operation is defined as follows:

[0025]

[0026] where m = 2u.

[0027] Furthermore, the working process of the forward process of the reversible transformation module is as follows:

[0028] Given the incomplete spectral features The forward process of the reversible transformation module performs a feature space transformation on the incomplete spectral features where the reversible transformation module contains l reversible transformation forward layers. The first reversible transformation forward layer of the reversible transformation module performs a feature layer forward encoding operation on the incomplete spectral features to obtain the intermediate feature f′1. The second reversible transformation forward layer of the reversible transformation module performs a feature layer forward encoding operation on the intermediate feature f′1 to obtain the intermediate feature f′2, and so on. After the feature layer forward encoding operation of the l-th reversible transformation forward layer, the latent encoding in the latent space is obtained In this process, the latent encoding in the latent space contains the complete feature information corresponding to the incomplete spectral features while fully decoupling and reconstructing different dimensions of the features; by introducing multiple levels of reversible transformation forward layers, the network can gradually extract and fuse more abstract and meaningful feature representations, providing a more robust foundation for subsequent reverse reconstruction and feature enhancement;

[0029] Each reversible transformation forward layer contains p reversible transformation forward blocks. The operation process expression of the q-th reversible transformation forward block is The operation process expression of a single reversible transformation forward layer is: Id represents the identity function, and G represents a continuous function represents the input feature, and * represents function composition; the feature layer forward encoding operation of the s-th reversible transformation forward layer is denoted as f′ s = F θ (f′ s-1 ).

[0030] Furthermore, the working process of the noise decoupling module is as follows:

[0031] Given the latent encoding in the latent space In this encoding, the noise feature encoding and the denoised feature encoding are separated onto different feature channels, where d = a + b; by performing a noise decoupling operation on the latent encoding , the noise feature encoding Z n is removed from the latent encoding to obtain the denoised latent encoding By separating the latent encoding into the noise feature encoding Z n and the denoised feature encoding Zc , the network can more effectively decouple the useful information and noise information in the input features. The core idea of this separation mechanism is to concentrate the inevitable noise information in the original features into the Z n part, while preserving the features with more semantic information in the Z c part;

[0032] The above noise decoupling operation is used to perform a zeroing operation on specific channels of the feature vector. For a given input vector and the corresponding mask vector m ∈ {0, 1} d , according to the values at each position in the mask vector , it is determined whether the values at the corresponding positions in the input vector v are retained or set to zero, and the value v' of the r-th channel in the input vector v' is obtained r = m r ·v r , where r = 1, 2, ……, d, m r represents the value of the r-th channel of the vector m, and v r represents the value of the r-th channel of the vector v.

[0033] Furthermore, the working process of the reverse process of the reversible transformation module is as follows:

[0034] Given the denoised latent code The reverse process of the reversible transformation module performs a feature space transformation on the denoised latent code . Among them, the reversible transformation module also contains l reversible transformation reverse layers. The first reversible transformation reverse layer of the reversible transformation module performs a feature layer reverse encoding operation on the denoised latent code to obtain the intermediate feature f1. The second reversible transformation reverse layer of the reversible transformation module performs a feature layer reverse encoding operation on the intermediate feature f1 to obtain the intermediate feature f2, and so on. After the feature layer reverse encoding operation of the l-th reversible transformation reverse layer, the feature vector of the complete graph fact is obtained In the operation of the entire reverse process, the denoised latent code is decoded layer by layer, and finally the feature vector of the complete graph fact is restored Due to the design of the reversible transformation module, this process can ensure the precise retention and reconstruction of information during the encoding and decoding processes. In particular, through the layer-by-layer feature space transformation in the reverse process, the network can gradually recover more detailed and accurate feature information without introducing additional noise. This mechanism not only improves the quality of feature reconstruction but also ensures the efficiency and interpretability of the entire network when processing feature transformation;

[0035] Each inverse transformation layer also contains p inverse transformation blocks, and the operation process expression of the t-th inverse transformation block is A single inverse transformation layer is defined as a combination: Id represents the identity function, and G represents a continuous function. represents the input feature, and * represents function composition; the feature layer inverse encoding operation of the o-th inverse transformation layer is expressed as

[0036] Furthermore, the working process of the decoder module is as follows:

[0037] Given the complete graph fact output by the invertible transformation module and its feature vector Subtract it from the feature vector of the input incomplete graph fact to obtain the feature vector of the missing element Use a multi-layer perceptron as the decoder to obtain the feature vector of the missing element The probability distribution P on the entire entity set and relationship set, and obtain the elements required for link prediction according to the probability distribution P; by comparing the feature vector of the complete graph fact with the feature vector of the input incomplete graph fact perform a difference calculation, that is The network can effectively extract the part of the difference caused by information loss in the input features. This difference feature vector

[0038] contains important information about the missing element, and by comparing with the complete feature vector, the network can better focus on the recovery and reconstruction of the missing information;

[0039]

[0040] where, MLP: represents the multi-layer perceptron, the matrix Q is obtained by vertically concatenating the entity embedding matrix E and the relationship embedding matrix R, is obtained after the softmax operation, and the softmax operation refers to the operation of converting each element in a vector into a probability distribution. The probability distribution P represents the similarity probability with each element in the hyper-relation knowledge graph for obtaining the link prediction result, |ε| represents the number of entities, represents the number of relationships.

[0041] Furthermore, the loss function is defined as follows:

[0042]

[0043] Among them, is the value at the v-th position of the smoothed label y s and p v is the value at the v-th position of the probability distribution P output by the decoder module.

[0044] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0045] 1. Low loss of extracted information. The reversible transformation module of the reversible neural network used in the present invention has the characteristic of lossless information. In the training process of conventional neural networks, due to operations such as pooling and dimensionality reduction, information is inevitably lost. However, the mapping of each layer of the reversible transformation module is bijective. Through the construction of reversible layers, it is ensured that each calculation can be completely restored without losing any information.

[0046] 2. Less video memory occupied. The reversible transformation module of the reversible neural network used in the present invention has the characteristic of high memory efficiency. In traditional neural networks, backpropagation needs to save the intermediate activation values of each layer for use when calculating gradients. This saving mechanism will occupy a large amount of memory. However, the reversible transformation module can reconstruct the intermediate activation values through reverse calculation during backpropagation without explicitly storing the intermediate activation values, thus occupying less video memory.

[0047] 3. Short required training time. The reversible transformation module of the reversible neural network used in the present invention has the characteristic of high training efficiency. In ordinary neural networks, in order to improve the expression ability, the depth or width of the network is usually increased, which will lead to a sharp increase in the number of parameters. However, the reversible transformation module can share some network parameters through its unique architecture design, avoiding repeated redundant calculations, thereby improving the training efficiency of the network and greatly shortening the training time.

[0048] 4. Better compatibility with hyper-relation facts. The graph fact embedding module proposed in the present invention obtains the graph fact features by summing the embedding vectors of each element in the graph fact. In this way, the number of feature channels of the graph fact features is independent of the logarithm of the additional information of the hyper-relation facts. Therefore, for hyper-relation facts containing a lot of additional information, the method proposed in the present invention can also be fully compatible. Description of the Drawings

[0049] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0050] Figure 1 It is a flowchart of the lightweight hyper-relation knowledge graph link prediction method disclosed in the present invention, which includes a training process and a testing process;

[0051] Figure 2 It is a schematic diagram of the composition of the reversible transformation module in the lightweight hyper-relation knowledge graph link prediction method disclosed in the present invention. Detailed implementation manners

[0052] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0053] When "embodiment" is mentioned in the present application, it means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments.

[0054] Embodiment 1

[0055] This embodiment discloses a lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network, which specifically includes the following steps:

[0056] S1. Construct a lightweight hyper-relation knowledge graph link prediction network based on a reversible neural network. The hyper-relation knowledge graph link prediction network includes a training data preprocessing module, a graph fact embedding module, a reversible transformation module, a noise decoupling module, and a decoder module connected in sequence;

[0057] Among them, the reversible transformation module is composed of 5 reversible transformation layers connected in sequence;

[0058] S2. Input the hyper-relation knowledge graph data of a specific domain, and perform data preprocessing operations on the hyper-relation knowledge graph data of the specific domain through a training data preprocessing module to obtain training instances and the corresponding label y, and use label smoothing technology to perform label smoothing operation on label y to obtain the smoothed label y s ;

[0059] Among them, the label smoothing rate ∈′ is 0.2.

[0060] S3. Perform embedding operations on the training instances through a graph fact embedding module to obtain incomplete graph fact features

[0061] Among them, the incomplete graph fact features have a dimension of 512.

[0062] S4. Perform feature space transformation operations on the incomplete graph fact features through the forward process of a reversible transformation module to obtain the latent encoding in the latent space

[0063] Among them, the forward process of the reversible transformation module is as Figure 2 shown, and the latent encoding has a dimension of 512.

[0064] S5. Perform noise decoupling operations on the noise in the latent encoding through a noise decoupling module to obtain the denoised latent encoding

[0065] Among them, the dimension of the noise feature encoding is 180, and the denoised latent encoding has a dimension of 512.

[0066] S6. Perform feature space transformation operations on the latent encoding through the reverse process of a reversible transformation module to obtain the complete graph fact features

[0067] Among them, the reverse process of the reversible transformation module is as Figure 2 shown, and the complete graph fact features has a dimension of 512.

[0068] S7. Subtract the complete graph fact features from the incomplete graph fact features to obtain the missing element features Then perform decoding operations on the missing element features through a decoder module to obtain the probability distribution P, and obtain the missing elements for link prediction according to the probability distribution;

[0069] Among them, the dimension of the missing element feature is 512, and the dimension of the probability distribution P is 35017.

[0070] S8. Train a lightweight hyper-relational knowledge graph link prediction network based on a reversible neural network through domain-specific hyper-relational knowledge graph data, and perform iterative training in an end-to-end manner. During each iteration, calculate the loss value by using the probability distribution P and the smoothed label y s and minimize the loss function to train the network. At the same time, update the parameter values in the network through backpropagation until the network converges;

[0071] Among them, the domain-specific hyper-relational knowledge graph data uses the WikiPeople dataset proposed by Guan et al. in the paper "Link Prediction on N-ary Relational Data". In this dataset, the number of entities is 34839, the number of relations is 178, and the number of facts is 369866, and a total of 1017462 training instances can be generated.

[0072] S9. Input the incomplete hyper-relational knowledge graph data into the trained hyper-relational knowledge graph link prediction network, and obtain the missing elements in the incomplete hyper-relational knowledge graph according to the output probability distribution.

[0073] The loss function is defined as follows:

[0074]

[0075] Among them, is the value at the v-th position of the smoothed label y s , and p v is the value at the v-th position of the probability distribution P output by the decoder module.

[0076] During the training process, the initial value of the learning rate is 0.01. In the first 10 rounds, the learning rate linearly increases from 10 -6 to 0.01. In the next 90 rounds, it linearly decreases from 0.01 to 10 -6 , and a total of 100 rounds of training are performed.

[0077] The training process of the entire network is as Figure 1 shown, and the hardware parameters of the training platform used are shown in Table 1, and the software parameters are shown in Table 2. After the network training is completed, refer to the Figure 1 test process in it to give the incomplete hyper-relational knowledge graph as input for hyper-relational knowledge graph link prediction to obtain the link prediction result.

[0078] Table 1. Hardware Environment Parameter Table

[0079] Hardware Name Hardware Model and Parameters CPU AMD EPYC 9654 RAM Samsung 60G Hard Disk Kingston 1TB GPU Nvidia GeForce RTX 4080 16G

[0080] Table 2. Software Environment Parameter Table

[0081] Software Version Operating System Ubuntu 20.04 LTS Python 3.8.10 PyTorch 1.11.0 Graphics Card Driver NVIDIA Graphics Card Driver 560.35.03 CUDA 11.3 CUDNN 8.9.2.26

[0082] Table 3 shows the quantitative comparison results of the current popular methods HINGE, StarE, HAHE and the present invention on the test set of the WikiPeople dataset. The average metrics are MRR, Hit@1, and Hit@10, and the larger the values, the better. The training time of the network on the training set of the WikiPeople dataset is also compared, and the shorter the time, the better. It can be seen from the table that the method disclosed in the present invention exceeds the existing methods in terms of MRR, Hit@1, and Hit@10 metrics, proving that the link prediction effect of the present invention is very good; in terms of the training time metric, the training time required by the present invention is the shortest, proving that the present invention is a lightweight link prediction method.

[0083] Table 3. Comparison Table of the Hyper-relation Knowledge Graph Link Prediction Method Disclosed in the Present Invention with Existing Results

[0084]

[0085] Example 2

[0086] This example discloses a lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network, which specifically includes the following steps:

[0087] S1. Construct a lightweight hyper-relation knowledge graph link prediction network based on a reversible neural network. This hyper-relation knowledge graph link prediction network includes a training data preprocessing module, a graph fact embedding module, a reversible transformation module, a noise decoupling module, and a decoder module that are sequentially connected in series;

[0088] Among them, the reversible transformation module is composed of 4 sequentially connected reversible transformation layers;

[0089] S2. Input the hyper-relation knowledge graph data of a specific domain, perform data preprocessing operations on the hyper-relation knowledge graph data of the specific domain through the training data preprocessing module to obtain training instances and the corresponding label y, and use the label smoothing technique to perform label smoothing operations on the label y to obtain the smoothed label y s ;

[0090] Among them, the label smoothing rate ∈′ is 0.2.

[0091] S3. Perform embedding operations on the training instances through the graph fact embedding module to obtain incomplete graph fact features

[0092] Among them, the factual features of the incomplete graph have a dimension of 512.

[0093] S4. Perform a feature space transformation operation on the factual features of the incomplete graph through the forward process of the reversible transformation module to obtain a latent encoding in the latent space

[0094] Among them, the forward process of the reversible transformation module is as Figure 2 shown, and the latent encoding has a dimension of 512.

[0095] S5. Perform a noise decoupling operation on the noise in the latent encoding through the noise decoupling module to obtain a denoised latent encoding

[0096] Among them, the dimension of the noise feature encoding is 180, and the denoised latent encoding has a dimension of 512.

[0097] S6. Perform a feature space transformation operation on the latent encoding through the reverse process of the reversible transformation module to obtain the complete graph factual features

[0098] Among them, the reverse process of the reversible transformation module is as Figure 2 shown, and the complete graph factual features has a dimension of 512.

[0099] S7. Subtract the complete graph factual features from the incomplete graph factual features to obtain the missing element features Then, perform a decoding operation on the missing element features through the decoder module to obtain a probability distribution P, and obtain the missing elements for link prediction according to the probability distribution;

[0100] Among them, the missing element features have a dimension of 512, and the probability distribution P has a dimension of 35017.

[0101] S8. Train a lightweight hyper-relation knowledge graph link prediction network based on a reversible neural network with data from a specific domain's hyper-relation knowledge graph, and perform iterative training in an end-to-end manner. In each iteration, calculate the loss value by using the probability distribution P and the smoothed label y s and minimize the loss function to train the network. At the same time, update the parameter values in the network through backpropagation until the network converges;

[0102] Among them, the hyper-relational knowledge graph data in a specific domain uses the WikiPeople dataset proposed by Guan et al. in the paper "Link Prediction on N-ary Relational Data". In this dataset, the number of entities is 34,839, the number of relationships is 178, and the number of facts is 369,866. A total of 1,017,462 training instances can be generated.

[0103] S9. Input the incomplete hyper-relational knowledge graph data into the trained hyper-relational knowledge graph link prediction network, and obtain the missing elements in the incomplete hyper-relational knowledge graph according to the output probability distribution.

[0104] The loss function is defined as follows:

[0105]

[0106] Among them, is the value at the v-th position of the smoothed label y s , and p v is the value at the v-th position of the probability distribution P output by the decoder module.

[0107] During the training process, the initial value of the learning rate is 0.01. In the first 10 rounds, the learning rate linearly increases from 10 -6 to 0.01. In the next 90 rounds, it linearly decreases from 0.01 to 10 -6 . A total of 100 rounds of training are performed.

[0108] Table 4 shows the quantitative comparison results of the current popular methods HINGE, StarE, HAHE and the present invention on the test set of the WikiPeople dataset. The average metrics are MRR, Hit@1 and Hit@10, and the larger the values, the better. The training time of the network on the training set of the WikiPeople dataset is also compared, and the shorter the time, the better. It can be seen from the table that the method disclosed in the present invention exceeds the existing methods in terms of MRR, Hit@1 and Hit@10 metrics, proving that the link prediction effect of the present invention is very good; in terms of the training time metric, the training time required by the present invention is the shortest, proving that the present invention is a lightweight link prediction method.

[0109] Table 4. Comparison table of the hyper-relational knowledge graph link prediction method disclosed in the present invention and the existing results

[0110]

[0111] It should be noted that, for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, some steps can be performed in other sequences or simultaneously.

[0112] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network, characterized in that The lightweight hyper-relation knowledge graph link prediction method includes the following steps: S1. Construct a lightweight hyper-relation knowledge graph link prediction network based on a reversible neural network. This hyper-relation knowledge graph link prediction network includes a training data preprocessing module, a graph fact embedding module, a reversible transformation module, a noise decoupling module, and a decoder module that are sequentially connected in order; among them, the reversible transformation module consists of multiple reversible transformation layers that are sequentially connected in order; S2. Input the hyper-relation knowledge graph data of a specific domain, and perform data preprocessing operations on the hyper-relation knowledge graph data of the specific domain through a training data preprocessing module to obtain training instances and the corresponding label y, and use label smoothing technology to perform label smoothing operations on the label y to obtain the smoothed label y s ; The hyper-relation knowledge graph data of the specific domain comes from medical knowledge graphs for disease diagnosis, financial knowledge graphs for risk assessment, educational knowledge graphs for personalized learning, and industrial knowledge graphs for fault diagnosis; S3. Perform embedding operations on the training instances through the graph fact embedding module to obtain incomplete graph fact features S4. Through the forward process of the reversible transformation module, the incomplete graph fact features are subjected to a feature space transformation operation to obtain a latent encoding in the latent space S5. Perform noise decoupling operation on the noise in the latent encoding to obtain denoised latent encoding S6. Perform a feature space transformation operation on the latent encoding through the reverse process of the reversible transformation module to obtain the complete graph fact feature S7. Subtract the complete graph fact features from the incomplete graph fact features to obtain the missing element features Then, perform a decoding operation on the missing element features to obtain the probability distribution P, and obtain the missing elements for link prediction according to the probability distribution; S8. Train a lightweight hyper-relation knowledge graph link prediction network based on a reversible neural network with hyper-relation knowledge graph data in a specific domain, and perform iterative training in an end-to-end manner. In each iteration, calculate the loss value using the probability distribution P and the smoothed label y s and minimize the loss function to train the network. At the same time, update the parameter values in the network through backpropagation until the network converges; S9. Input the incomplete hyper-relation knowledge graph into the trained hyper-relation knowledge graph link prediction network, and obtain the missing elements in the incomplete hyper-relation knowledge graph according to the output probability distribution.

2. The lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network according to claim 1, wherein The working process of the training data preprocessing module is as follows: randomly divide the obtained hyper-relation knowledge graph data in a specific domain according to the ratio of 8:1:1 to obtain a training set, a validation set, and a test set, and perform a masking operation on the graph facts in the training set to obtain training instances and the corresponding label y, and perform label smoothing operation on the label y using label smoothing operation to obtain the smoothed label y s ; Among them, the hyper-relation knowledge graph consists of multiple triple facts and multiple hyper-relation facts, and the hyper-relation fact is expressed as where \((h, r, t)\) represents the main triple, \(h\) represents the head entity, \(r\) represents the relation connecting the head entity and the tail entity, and \(t\) represents the tail entity, represents \(m\) pairs of additional information, \(q\) i represents the relation of the \(i\)-th pair of additional information, \(v\) i represents the entity of the \(i\)-th pair of additional information; \(h, t, v_1, \ldots, v\) m \(\in \varepsilon\), \(\varepsilon\) is the set of all entities, is the set of all relations; when \(m = 0\), the hyper-relation fact is simplified to a triple fact Therefore, the triple fact and the hyper-relation fact are collectively referred to as graph facts Among them, the mask operation process is as follows: For a hyper-relation fact with m pairs of additional information Mask each element in the hyper-relation fact using the mask token [MASK], and create (2m + 3) training instances for the hyper-relation fact ; For a triple fact Mask each element in the triple fact using the mask token [MASK], and create 3 training instances for the triple fact ; Create a label y for each training instance, that is, define where according to the serial number of the masked element in the hyper-relation knowledge graph data, set the corresponding position of the label y to 1, and set the other positions of the label y to 0; For the label y, there is where |ε| represents the number of entities, represents the number of relations; Among them, the label smoothing operation process is as follows: For the t-th element y in the label y t , if the position corresponding to the t-th element is the target entity position, then set the element value to y t = 1 - ∈′, and set the element values at other positions to where the label smoothing rate Thus, the label y is converted into a smoothed label y s .

3. The lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network according to claim 1, characterized in that The working process of the graph fact embedding module is as follows: Given the entity set ε and the relation set of the hyper-relation knowledge graph data in a specific domain Generate an entity embedding matrix E and a relation embedding matrix R respectively, and randomly initialize the entity embedding matrix E and the relation embedding matrix R, where where d is the dimension of the embedding vector and also the number of feature channels; obtain the embedding vectors of the elements in the incomplete graph fact from the entity embedding matrix E and the relation embedding matrix R according to the index, and obtain the feature vector of the incomplete graph fact through the element-wise summation operation Among them, the element-wise summation operation process is as follows: Assume that the incomplete graph fact is missing one entity, there are m pairs of additional information, and there are u relationships. The embedding vectors of the u relationships are represented as (r1, r2, ···, r u ); then there will be u entities, and the embedding vectors of the u entities are represented as (e1, e2, ···, e u ), where the incomplete graph fact feature is obtained by adding each embedding vector element-wise The element-wise summation operation is defined as follows: wherein m = 2u.

4. The lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network according to claim 1, wherein The workflow of the forward process of the reversible transformation module is as follows: Given the incomplete spectral features The forward process of the reversible transformation module performs feature space transformation on the incomplete spectral features where the reversible transformation module includes l reversible transformation forward layers. The first reversible transformation forward layer of the reversible transformation module performs feature layer forward encoding operation on the incomplete spectral features to obtain the intermediate feature f′1. The second reversible transformation forward layer of the reversible transformation module performs feature layer forward encoding operation on the intermediate feature f′1 to obtain the intermediate feature f′2, and so on. After the feature layer forward encoding operation of the l-th reversible transformation forward layer, the latent code in the latent space is obtained Each invertible transformation forward layer contains p invertible transformation forward blocks, and the operation process expression of the q-th invertible transformation forward block is The operation process expression of a single invertible transformation forward layer is: F θ (x) = F p (x) * F p-1 (x) * … * F1(x), where Id represents the identity function and G represents a continuous function, represents the input feature, and * represents function composition; the forward encoding operation of the feature layer of the s-th invertible transformation forward layer is denoted as f′ s = F θ (f′ s-1 ).

5. The lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network according to claim 1, wherein The working process of the noise decoupling module is as follows: Given a latent encoding in a latent space In this encoding, the noise feature encoding and the denoised feature encoding are separated onto different feature channels, where d = a + b; By performing a noise decoupling operation on the latent encoding the noise feature encoding Z n is removed from the latent encoding to obtain the denoised latent encoding The above noise decoupling operation is used to perform a zeroing operation on specific channels of the eigenvector, for a given input vector and the corresponding mask vector m ∈ {0, 1} d , according to the values at each position in the mask vector , it is determined whether the values at the corresponding positions in the input vector v are retained or set to zero, and the value v' of the r-th channel in the input vector v' is obtained r = m r · v r , where r = 1, 2, ……, d, m r represents the value of the r-th channel of the vector m, and v r represents the value of the r-th channel of the vector v.

6. The lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network according to claim 4, wherein The working process of the reverse process of the reversible transformation module is as follows: Given the latent code of denoising The reverse process of the reversible transformation module for the latent code of denoising Performs feature space transformation, where the reversible transformation module also includes l reversible transformation reverse layers. The first reversible transformation reverse layer of the reversible transformation module performs Feature layer reverse coding operation on the latent code of denoising to obtain intermediate feature f1. The second reversible transformation reverse layer of the reversible transformation module performs feature layer reverse coding operation on intermediate feature f1 to obtain intermediate feature f2, and so on. After the feature layer reverse coding operation of the l-th reversible transformation reverse layer, the feature vector of the complete graph fact is obtained Each inverse transformation reverse layer also contains p inverse transformation reverse blocks, and the operation process expression of the t-th inverse transformation reverse block is A single inverse transformation reverse layer is defined as a combination: Id represents the identity function, G represents a continuous function, represents the input feature, * represents function composition; the feature layer inverse encoding operation of the o-th inverse transformation reverse layer is expressed as 7. The lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network according to claim 1, wherein The working process of the decoder module is as follows: Given the complete graph output by the reversible transformation module The eigenvector of Compare it with the incomplete graph facts of the input The eigenvector of Subtract and get the eigenvector of the missing element Use a multi-layer perceptron as a decoder to obtain the feature vector of the missing element The probability distribution P on the entire entity set and relationship set is used to obtain the elements required for link prediction; The probability distribution P is defined as follows: Among them, denotes a multi-layer perceptron. The matrix Q is obtained by vertically concatenating the entity embedding matrix E and the relation embedding matrix R. is obtained after the softmax operation. The softmax operation refers to the operation of converting each element in a vector into a probability distribution. The probability distribution P represents the similarity probability with each element in the hyper-relation knowledge graph and is used to obtain the link prediction result. |ε| represents the number of entities, and represents the number of relations.

8. The lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network according to claim 7, wherein The loss function is defined as follows: Among them, is the value at the v-th position of the smoothed label y s , and p v is the value at the v-th position of the probability distribution P output by the decoder module.

9. The lightweight hyper-relation knowledge graph link prediction method based on a reversible neural network according to claim 1, characterized in that The hyper-relation knowledge graph data in a specific domain is a knowledge graph constructed by human experts using data acquisition devices, computing devices, storage devices, and data annotation tools.