Time sequence knowledge graph completion method based on multivariate feature coding and double-layer convolution
By using multivariate feature coding and bilayer convolution methods in the timing knowledge graph, the problem of incomplete entity knowledge representation and insufficient fusion of time information are solved, and better knowledge representation and link prediction effects are achieved.
Patent Information
- Application Number
- CN202510038175.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing graph neural network model has incomplete entity knowledge representation in the timing knowledge graph, and the fusion and interaction between time information and entity and relationship information is insufficient, which affects the link prediction effect of the model.
The time series knowledge graph completion method based on multivariate feature encoding and bilayer convolution is adopted to integrate the multivariate features of the entity through multivariate feature encoding, and the interactive features of entities, relationships and time are extracted using bilayer convolution to improve the expression of time information in link prediction.
It effectively improves the embedding of the knowledge representation of entities, enhances the fusion and interaction of time information with entity and relationship information, and improves the effect of link prediction.
Smart Images

Figure CN119962644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and knowledge graphs, and in particular to a temporal knowledge graph completion method based on multi-feature encoding and double-layer convolution. Background Art
[0002] With the rapid development of Internet technology and artificial intelligence, the speed of information generation and updating is accelerating, and the explosive growth of data has become a normal state. Knowledge graphs can effectively organize, store and understand information, and have become a major form of data storage. Many large open knowledge graphs have been launched, such as Freebase, Wikidata and WordNet, and have been widely used in intelligent question answering, personalized recommendation and semantic search.
[0003] Since the current construction method of knowledge graph is mainly manual or semi-automatic, it has serious knowledge missing problems, which in turn leads to poor performance in downstream tasks. In order to solve the problem of knowledge missing, knowledge graph completion methods have emerged. At present, knowledge graph completion methods are mainly based on static knowledge graphs. However, static knowledge graphs ignore the real-time changes and updates of data, and there is a certain information lag, so time series knowledge graphs have emerged. Time series knowledge graphs introduce time information on the basis of static graphs, and can be dynamically expanded and modified in time according to needs to adapt to the needs of scenarios such as time-sensitive event prediction and query. For time series knowledge graphs, the same entity exists at different timestamps, and the same entity at different timestamps has multiple static neighborhood information. Therefore, how to effectively integrate entity time information and neighborhood information is the key to knowledge representation in time series knowledge graphs.
[0004] The existing graph neural network model methods mainly focus on the effective fusion of neighborhood static information at the same timestamp and entity and time information at different timestamps as the final representation of entity information. Although this method effectively considers the temporal global characteristics of entities and the local characteristics of entities, it ignores the static neighborhood information of the same entity at non-target timestamps and the invariant static information of entities at the entire timestamp. In link prediction, the deep fusion of entities, relations, and time is crucial to the final effect of the model. The existing model first converts the quadruple with time information into a triple through simple vector fusion, and then uses the static knowledge graph completion method to achieve it. Although most models fuse time information with entity and relationship information, their fusion interaction is insufficient. Summary of the invention
[0005] In order to overcome the technical defects of the existing graph neural network model method, such as incomplete entity knowledge representation and insufficient time information fusion, the present invention provides a temporal knowledge graph completion method based on multi-feature encoding and double-layer convolution.
[0006] The present invention provides a temporal knowledge graph completion method based on multivariate feature coding and double-layer convolution, the steps are:
[0007] Step S1, performing data preprocessing on the time series knowledge graph, the time series knowledge graph includes multiple quadruple groups, and the data preprocessing is to expand the number of given quadruple groups;
[0008] Step S2: regard the temporal knowledge graph preprocessed in step S1 as a global graph, and fuse the multi-dimensional features of the entity as the final embedding vector of the entity;
[0009] For different representations of the same entity at different timestamps, the embedding vector s of the head entity, the initial embedding vector o of the tail entity and the timestamp embedding vector t are combined to obtain the embedding vector s of the head entity at timestamp t. t and the embedding vector o of the tail entity at timestamp t t , as shown below:
[0010] s t =s+tW t (1)
[0011] o t =o+tW t (2)
[0012] In formula (1) and formula (2), is the time-varying parameter matrix;
[0013] For different representations of the same entity at different timestamps, the temporal global information is obtained by time-level attention fusion. The embedding vector s of the head entity at timestamp t is t The weight coefficient is a t , a t The calculation formula is:
[0014]
[0015] In formula (3), is the weight parameter, a t ∈R;
[0016] Aggregate the embedding vectors of the same entity at different timestamps to obtain the temporal global information s′, and the calculation formula of s′ is:
[0017] s′=∑ t a t s t (4);
[0018] The information of the quadruple in the second layer of the global graph is represented as The calculation formula is:
[0019] In formula (5), r is the relationship vector;
[0020] The weight coefficient of the entity neighborhood information of the quadruple in the second layer of the global graph is b s , b s The calculation formula is:
[0021]
[0022] In formula (6), N s represents the neighborhood relationship of entity s in the global graph and the set of adjacent entities, represents the weight parameter, b s ∈R;
[0023] Aggregate the neighborhood information of the second layer of the global graph and obtain the entity global information s″, the calculation formula of s″ is:
[0024]
[0025] Embed the head entity at timestamp t into vector s t As local information, the initial embedding vector s of the head entity is taken as static information, then the final embedding vector of entity s is expressed as
[0026] For the relationship r in the quadruple, as time changes, the connected head entity and tail entity are changing, but the relationship expressed between the head entity and the tail entity remains unchanged. Therefore, for the relationship r, the relationship initial embedding vector r is used to represent the relationship feature;
[0027] Step S3: embed the final embedding vector of entity s The relationship initial embedding vector r and the timestamp embedding vector t are reshaped into a two-dimensional matrix, and a double-layer convolution is used to extract the interactive features of the three:
[0028] The final embedding vector of entity s The initial embedding vector r of the relationship and the timestamp embedding vector t are concatenated to obtain a matrix vector; a cross-shaped convolution kernel is used to integrate time information with entities and relationships; and a circular convolution is used to extract the interactive features of entities, relationships, and time.
[0029] Step S4: Obtain the prediction vector of the output tail entity through matrix reshaping and linear layer to complete link prediction.
[0030] Preferably, the expansion of the number of given quadruple in step S1 includes adding a reverse relationship quadruple to each quadruple and adding a self-loop relationship pointing to itself to each entity.
[0031] Preferably, the final embedding vector of entity s is The initial embedding vector r of the relation and the timestamp embedding vector t are concatenated to make the final embedding vector of entity s Each dimension value of the initial embedding vector r of the relationship is surrounded by the dimension values of the timestamp embedding vector t. The concatenated matrix vector is M a ∈R m×n , where m×n=4d 0 And m = 2k 1 、n=4k 2 , k 1 , k 2 Is a positive integer; for the concatenated matrix vector M a , use a cross-shaped convolution kernel of size 5 to convolve it twice, that is, the matrix vector M a The upper and left sides of are padded with zeros, and then a cross convolution kernel with a hop count of 2 is used for convolution, and the output is Similarly, for M a The lower and right sides of the convolution are padded with zeros, and the output Finally, use line break splicing M 1 、M 2 get The matrix M of the cross convolution output b , use a convolution kernel of size 7×7 to perform circular convolution on it, and get the output matrix M c .
[0032] Preferably, M 1 and M 2 The splicing method adopted is chessboard splicing of fused time vectors.
[0033] Preferably, in step S4, the matrix M output by the circular convolution is c Adjust its shape to fit the input requirements of the linear layer, input the adjusted matrix into the linear layer for link prediction, and output the prediction vector of the tail entity.
[0034] Compared with the prior art, the technical solution provided by the present invention has the following technical effects:
[0035] In response to the problem of incomplete knowledge representation of entities in existing graph neural network model methods, the present invention proposes a temporal knowledge graph completion method based on multi-feature encoding and double-layer convolution, which integrates four types of information: temporal global information, entity global information, entity local information and static information as the final embedding vector of the entity, effectively improving the knowledge representation embedding of the entity; in order to solve the problem of insufficient fusion and interaction of time information with entity and relationship information, the present invention proposes a method for interactive feature reshaping of entity, relationship and time embedding vectors, and uses cross convolution to fuse time information, effectively improving the expression of time information in link prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0038] Figure 1 It is a flowchart of a method for completing a temporal knowledge graph based on multi-feature coding and double-layer convolution in an embodiment of the present invention;
[0039] Figure 2 This is a model overall framework diagram of a temporal knowledge graph completion method based on multi-feature coding and double-layer convolution described in an embodiment of the present invention;
[0040] Figure 3 A schematic diagram of converting a temporal knowledge graph described in an embodiment of the present invention into a global graph;
[0041] Figure 4 A schematic diagram of vector splicing processing in an embodiment of the present invention;
[0042] Figure 5 A schematic diagram of cross convolution processing in an embodiment of the present invention;
[0043] Figure 6 Schematic diagram of circular convolution processing in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] In order to more clearly understand the above-mentioned purposes, features and advantages of the present invention, the scheme of the present invention will be further described below. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict. In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only part of the embodiments of the present invention, not all of the embodiments.
[0045] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0046] In one embodiment, Figure 1 and Figure 2 As shown, a temporal knowledge graph completion method based on multi-feature encoding and double-layer convolution is disclosed, and the steps are as follows:
[0047] Step S1: preprocessing the time series knowledge graph. The time series knowledge graph includes multiple quadruple groups. Data preprocessing is to expand the number of given quadruple groups. In order to better learn the positive and negative relationships, expanding the number of given quadruple groups includes adding a negative relationship quadruple group to each quadruple group and adding a self-loop relationship pointing to each entity.
[0048] Step S2: regard the temporal knowledge graph preprocessed in step S1 as a global graph, and fuse the multi-dimensional features of the entity as the final embedding vector of the entity;
[0049] like Figure 3 As shown, the temporal knowledge graph is regarded as a global graph, for the i-th entity s i , whose first layer is the i-th entity s i Entity representation at different timestamps, the second layer is the i-th entity s i One-hop neighborhood entity representation at different timestamps;
[0050] In order to effectively fuse the i-th entity s i The global information representation of is represented by using time-level attention to fuse entity representations at different timestamps (the first layer of the global graph) to generate temporal global information, and the entity-level attention to fuse the global one-hop neighborhood entity information (the second layer of the global graph) to generate entity global information, as follows.
[0051] For different representations of the same entity at different timestamps, the embedding vector s of the head entity, the initial embedding vector o of the tail entity and the timestamp embedding vector t are combined to obtain the embedding vector s of the head entity at timestamp t. t and the embedding vector o of the tail entity at timestamp t t , as shown below:
[0052] s t =s+tW t (1)
[0053] o t =o+tW t (2)
[0054] In formula (1) and formula (2), is the time-varying parameter matrix;
[0055] For different representations of the same entity at different timestamps, the temporal global information is obtained by time-level attention fusion. The embedding vector s of the head entity at timestamp t is t The weight coefficient is a t , a t The calculation formula is:
[0056]
[0057] In formula (3), is the weight parameter, a t ∈R;
[0058] Aggregate the embedding vectors of the same entity at different timestamps to obtain the temporal global information s′, and the calculation formula of s′ is:
[0059] s′=∑ t a t s t (4);
[0060] For the one-hop neighborhood information of entity s at different timestamps, the global neighborhood information is fused through entity-level attention; the information of the quadruple in the second layer of the global graph is represented as The calculation formula is:
[0061]
[0062] In formula (5), r is the relationship vector;
[0063] The weight coefficient of the entity neighborhood information of the quadruple in the second layer of the global graph is b s , b s The calculation formula is:
[0064]
[0065] In formula (6), N s represents the neighborhood relationship of entity s in the global graph and the set of adjacent entities, represents the weight parameter, b s ∈R;
[0066] Aggregate the neighborhood information of the second layer of the global graph and obtain the entity global information s″, the calculation formula of s″ is:
[0067]
[0068] For the input query quadruple (s, r, ?, t), the local information of the timestamp t where the entity s is located plays a key role in locating the query of the tail entity, so the head entity at timestamp t is embedded in the vector s t As local information, in the temporal knowledge graph, for entity s, although its neighborhood information changes over time, for the entity itself, some attribute feature information does not change over time. Therefore, the initial embedding vector s of the head entity is taken as static information, and the final embedding vector of entity s is expressed as
[0069]
[0070] For the relationship r in the quadruple, as time changes, the connected head entity and tail entity change, but the relationship between the head entity and the tail entity remains unchanged. For example, for the relationship "husband and wife", it always indicates that the head entity and the tail entity are in a husband and wife relationship. Therefore, for the relationship r, the relationship initial embedding vector r is used to represent the relationship feature.
[0071] Step S3: embed the final embedding vector of entity s The initial relationship embedding vector r and the timestamp embedding vector t are reshaped into a two-dimensional matrix, and a double-layer convolution is used to extract the interactive features of the three.
[0072] like Figure 4 As shown, the final embedding vector of entity s The initial embedding vector r of the relation and the timestamp embedding vector t are concatenated to make the final embedding vector of entity s Each dimension value of the initial embedding vector r of the relationship is surrounded by the dimension values of the timestamp embedding vector t, so that during convolution, the entity and relationship features can have greater interaction with the time features; Figure 3 In, i Represents the dimension value in the entity embedding vector, r i Represents the dimension value in the relation embedding vector, t i Represents the dimension value in the time embedding vector; the concatenated matrix vector is M a ∈R m×n , where m×n=4d 0 And m = 2k 1 、n=4k 2 , k 1 , k 2 is a positive integer;
[0073] like Figure 5 As shown, the concatenated matrix vector M a , use a cross-shaped convolution kernel of size 5 to convolve it twice, that is, the matrix vector M a The upper and left sides of are padded with zeros, and then a cross convolution kernel with a hop count of 2 is used for convolution, and the output is Matrix Elements and They represent the dimension values of the entity embedding vector and the relationship embedding vector after integrating the time dimension. Similarly, for M a The lower and right sides of the convolution are padded with zeros, and the output Finally, use line break splicing M 1 、M 2 get M 1 and M 2 The splicing method used is chessboard splicing of fused time vectors; Figure 6 As shown, the matrix M of the cross convolution output b , use a convolution kernel of size 7×7 to perform circular convolution on it, and get the output matrix M c ; Using cross convolution, not only can time information be effectively integrated into entity embedding vectors and relationship embedding vectors, but also from the perspective of the number of dimensions, one entity or relationship dimension value is integrated with four time dimension values, effectively reducing the number of dimensions of the timestamp embedding vector;
[0074] Step S4: Obtain the prediction vector of the output tail entity through matrix reshaping and linear layer to complete link prediction; for the matrix M output by circular convolution c Adjust its shape to fit the input requirements of the linear layer, input the adjusted matrix into the linear layer for link prediction, and output the prediction vector of the tail entity.
[0075] The above is only a specific implementation of the present invention, which enables those skilled in the art to understand or implement the present invention. Although detailed descriptions are given with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments, and they should all be covered by the protection scope of the claims.
Claims
1. A temporal knowledge graph completion method based on multivariate feature coding and double-layer convolution, characterized in that: The steps are: Step S1, performing data preprocessing on the time series knowledge graph, the time series knowledge graph includes multiple quadruple groups, and the data preprocessing is to expand the number of given quadruple groups; Step S2: regard the temporal knowledge graph preprocessed in step S1 as a global graph, and fuse the multi-dimensional features of the entity as the final embedding vector of the entity; For different representations of the same entity at different timestamps, the embedding vector s of the head entity, the initial embedding vector o of the tail entity and the timestamp embedding vector t are combined to obtain the embedding vector s of the head entity at timestamp t. t and the embedding vector o of the tail entity at timestamp t t , as shown below: s t =s+tW t (1) oh t =o+tW t (2) In formula (1) and formula (2), is the time-varying parameter matrix; For different representations of the same entity at different timestamps, the temporal global information is obtained by time-level attention fusion. The embedding vector s of the head entity at timestamp t is t The weight coefficient is a t , a t The calculation formula is: In formula (3), is the weight parameter, a t ∈R; Aggregate the embedding vectors of the same entity at different timestamps to obtain the temporal global information s′, and the calculation formula of s′ is: s′=∑ t a t s t (4); The information of the quadruple in the second layer of the global graph is represented as The calculation formula is: In formula (5), r is the relationship vector; The weight coefficient of the entity neighborhood information of the quadruple in the second layer of the global graph is b s , b s The calculation formula is: In formula (6), N s represents the neighborhood relationship of entity s in the global graph and the set of adjacent entities, represents the weight parameter, b s ∈R; Aggregate the neighborhood information of the second layer of the global graph and obtain the entity global information s″, the calculation formula of s″ is: Embed the head entity at timestamp t into vector s t As local information, the initial embedding vector s of the head entity is taken as static information, then the final embedding vector of entity s is expressed as For the relationship r in the quadruple, as time changes, the connected head entity and tail entity are changing, but the relationship expressed between the head entity and the tail entity remains unchanged. Therefore, for the relationship r, the relationship initial embedding vector r is used to represent the relationship feature; Step S3: embed the final embedding vector of entity s The initial relationship embedding vector r and the timestamp embedding vector t are reshaped into a two-dimensional matrix, and a double-layer convolution is used to extract the interactive features of the three. The final embedding vector of entity s The initial embedding vector r of the relationship and the timestamp embedding vector t are concatenated to obtain a matrix vector; a cross-shaped convolution kernel is used to integrate time information with entities and relationships; and a circular convolution is used to extract the interactive features of entities, relationships, and time. Step S4: Obtain the prediction vector of the output tail entity through matrix reshaping and linear layer to complete link prediction.
2. The method for completing the temporal knowledge graph based on multivariate feature coding and double-layer convolution according to claim 1 is characterized in that: In step S1, the expansion of the number of given quadruple includes adding a reverse relationship quadruple to each quadruple and adding a self-loop relationship pointing to itself to each entity.
3. The method for completing the temporal knowledge graph based on multivariate feature coding and double-layer convolution according to claim 1 is characterized in that: In step S2, the final embedding vector of entity s is The initial embedding vector r of the relation and the timestamp embedding vector t are concatenated to make the final embedding vector of entity s Each dimension value of the initial embedding vector r of the relationship is surrounded by the dimension values of the timestamp embedding vector t. The concatenated matrix vector is M a ∈R m×n , where m×n=4d0 and m=2k1、n=4k2, k1、k2 are positive integers; for the concatenated matrix vector M a , use a cross-shaped convolution kernel of size 5 to convolve it twice, that is, the matrix vector M a The upper and left sides of are padded with zeros, and then a cross convolution kernel with a hop count of 2 is used for convolution, and the output is Similarly, for M a The lower and right sides of the convolution are padded with zeros, and the output Finally, we use line breaks to concatenate M1 and M2 to get The matrix M of the cross convolution output b , use a convolution kernel of size 7×7 to perform circular convolution on it, and get the output matrix M c .
4. The method for completing the temporal knowledge graph based on multivariate feature coding and double-layer convolution according to claim 3 is characterized in that: The splicing method used by M1 and M2 is chessboard splicing of fused time vectors.
5. The method for completing the temporal knowledge graph based on multivariate feature coding and double-layer convolution according to claim 1 is characterized in that: In step S4, the matrix M output by the circular convolution is c Adjust its shape to fit the input requirements of the linear layer, input the adjusted matrix into the linear layer for link prediction, and output the prediction vector of the tail entity.