Graph feature completion method based on course learning

Through the graph feature completion method based on course learning, node completion difficulty is calculated, node scheduling and weighted directed diffusion propagation are carried out from easy to difficult, and combined with the channel mutual information correlation propagation module, problems such as high feature missing rate, differential neglect, correlation neglect, initialization dependence and difficulty in expansion are solved in the existing technology, and efficient and scalable feature completion effect is achieved.

CN119942162APending Publication Date: 2025-05-06XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510035818.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When the prior art processes the missing features in graph data, it is difficult to deal with high feature missing rates, ignore the differences in the completion of different nodes, ignore the correlation between feature dimensions, rely on the attribute initialization method, and is difficult to expand to large graphs.

Method used

The graph feature completion method based on course learning is adopted, and the node completion difficulty is calculated through a multi-view difficulty measurer. The continuous diffusion scheduler is used to schedule nodes from easy to difficult, and weighted directed diffusion propagation is performed. The feature optimization is performed in combination with the channel mutual information correlation propagation module.

Benefits of technology

Effectively deal with graph data with high feature missing rates, achieve differentiated completion, consider the correlation between feature dimensions, avoid noise interference from attribute initialization, and improve the scalability and overall prediction performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942162A_ABST
    Figure CN119942162A_ABST
Patent Text Reader

Abstract

The invention discloses a graph feature completion method based on course learning. The method comprises the following steps of: 1, calculating the completion difficulty of each node in a given undirected graph through a multi-view difficulty measurer, evaluating the completion difficulty of each node in the graph and scoring; step 2, according to a linear or geometric scheduling function, continuously introducing nodes with missing features into a diffusion process, and dynamically generating an undirected subgraph structure of each round of diffusion; 3, converting symmetric propagation between nodes of the undirected subgraph structure into a directed propagation mode related to the direction, and gathering complemented or known features of adjacent nodes at different diffusion levels to obtain a preliminary complemented feature matrix; and step 4, further optimizing the diffusion completion feature matrix through a channel mutual information correlation propagation module, finally outputting a reconstructed feature matrix, and sending the reconstructed feature matrix into a downstream GNNs for training related tasks. According to the method, the expression ability of the model is enhanced, and the performance of the network on downstream tasks is finally optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of completing missing features in graph data, and specifically relates to a graph feature completion method based on curriculum learning. Background Art

[0002] The core idea of ​​graph neural networks is to capture the dependencies between graph nodes through message passing, aggregate the features of a node with the features of its neighbors, and update the representation of the node. In the process of message passing, complete node features play an important role in the performance of GNNs. In recent years, in response to feature missing in graph data, K-nearest neighbor completion, graph autoencoder, generative adversarial network and other methods have been used to deal with the attribute completion problem of incomplete graphs and improve the downstream task performance of GNNs. These solutions can be divided into GNN variant models and feature completion models according to the final output of the model. The GNN variant model is used to directly use some known features for end-to-end task prediction and output the prediction results. The feature completion model first completes the information and outputs the completed feature matrix, which is combined with GNNs for downstream tasks. Although the existing feature completion technology solutions have alleviated the impact of feature missing to a certain extent, they still have the following major defects:

[0003] ① It is difficult to cope with the situation of high feature missing rate: Existing feature completion methods include graph-related learning-based methods and graph-independent pre-filling methods, such as simple strategies such as setting features to 0, Gaussian distribution random values ​​or global mean. However, these two methods perform poorly under high feature missing rate and are prone to distribution bias or noise, which affects the completion effect. Learning-based methods complete features by mining the relationship between known features and graph structure information, such as using models such as autoencoders and generative networks to infer missing feature values. They show good performance when the feature missing rate is low, but when the acquisition cost of node features is high, that is, there is a high feature missing rate, the available known features will be greatly reduced, resulting in the model being unable to effectively learn the true pattern of feature distribution, which in turn produces bias and causes a significant decline in model performance. The graph-independent pre-filling method is simple and easy to implement, but it may introduce a lot of noise when the feature missing rate is high. This noise will interfere with GNNs' aggregation of node and neighbor features, resulting in the model being unable to accurately capture the relationship between nodes. For example, if a feature value is missing for most nodes, simply setting it to 0 or filling it with the global mean may cause the feature values ​​of all nodes to converge, destroying the differences between nodes and affecting subsequent task processing of GNNs.

[0004] ② Ignoring the differences in the completion of different nodes: Existing completion methods usually adopt a unified processing method when dealing with feature-missing nodes, lacking a differentiated strategy, that is, performing a globally consistent completion operation on the missing features of all nodes, ignoring the differences in completion difficulty between different nodes and local context information (such as neighbor features and graph structure), resulting in less than ideal completion results. In addition, most completion methods fail to fully consider the correlation between node features to enhance the overall completion efficiency, but simply put easy-to-complete and difficult-to-complete features on an equal footing. This may result in the fact that in the completion process, features that are easy to complete are not fully utilized, while features that are difficult to complete cannot be effectively processed, which in turn affects the performance of the overall network.

[0005] ③ Ignoring the correlation between feature dimensions: In graph data, the features of nodes are usually composed of multiple dimensions, each dimension represents different attributes or information, and there may be different relationships between these feature dimensions, including but not limited to linear, nonlinear, conditional dependence and other correlations. For example, the journal impact factor and the number of citations in the citation network are usually positively correlated; the relationship between gene expression levels and protein interaction strength in biological networks may be affected by multiple factors and have nonlinear relationships. Existing methods such as traditional interpolation methods usually focus on a single dimension and independently complete each feature dimension, ignoring the relationship between different dimensions, resulting in incomplete completion results and affecting the performance of subsequent tasks.

[0006] ④ Dependence on attribute initialization method: Most completion schemes rely on the initialization method of missing attributes, such as using feature mean initialization or establishing a structural feature learning model, using the learned structural features as the initial values ​​of the missing features, and completing the features based on the filled initial values. However, the difference in attribute initialization methods does not bring significant performance improvements in most cases, but may affect the accuracy of the model due to the introduction of noise. Taking the end-to-end partial convolution GCN as an example, by ignoring the missing attributes and directly performing the missing part of the graph convolution operation, the negative impact of attribute initialization is successfully avoided, thereby achieving better performance than some mainstream models such as those based on Gaussian mixture models to fill missing data and graph autoencoders in multiple graph learning tasks.

[0007] ⑤ Difficult to scale to large graphs: As the size of graphs increases, the computational complexity and scalability of feature completion models become an important challenge, especially in large-scale graphs, when the feature missing rate is high, existing completion methods may face problems such as insufficient computing resources or slow convergence. For example, in social networks, the number of nodes may reach millions or even more. When the feature missing rate is high, training a model that can effectively complete the missing features may consume a lot of computing resources and time, limiting the performance of the network on large-scale graph datasets.

[0008] Although existing missing feature completion methods have made some progress in improving the generalization of GNNs, they still have problems such as decreased model performance, inability to fully utilize effective information, neglect of feature relevance, reliance on initialization methods, high complexity and difficulty in expansion when data is highly missing. These shortcomings limit the completion effect of the model and the final performance of GNNs on downstream tasks. Summary of the invention

[0009] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a graph feature completion method based on curriculum learning, which uses mutual information to capture the correlation between node features and perform feature optimization, aiming to improve the model's adaptability to the scale of graph data and the degree of information missing, and gradually complete the completion process from easy to difficult, complete the completion in an efficient and scalable manner for downstream tasks, enhance the expressiveness of the model, and ultimately optimize the performance of the network on downstream tasks.

[0010] In order to achieve the above object, the technical solution adopted by the present invention is:

[0011] A graph feature completion method based on curriculum learning includes the following steps;

[0012] Step 1: Calculate the completion difficulty of each node in a given undirected graph through a multi-view difficulty meter, and evaluate and score the completion difficulty of each node in the graph by combining the feature information and structural complexity of the node neighborhood; thus serving the node scheduling process from easy to difficult in step 2;

[0013] Step 2: For the score, the continuous diffusion scheduler module continuously introduces nodes with missing features into the diffusion process according to the linear or geometric scheduling function, and dynamically generates an undirected subgraph structure for each round of diffusion. These subgraph structures will be used in the subgraph weighted directed diffusion process in step 3, and will be hierarchically expanded until all missing nodes are introduced;

[0014] Step 3: Perform weighted directed diffusion propagation on the undirected subgraph structure, convert the symmetric propagation between nodes of the undirected subgraph structure into a directed propagation method related to direction, and aggregate the completion or known features of neighboring nodes at different diffusion levels; when all missing nodes are introduced into the subgraph and feature diffusion is completed, a preliminary completion feature matrix is ​​obtained; and it is passed to the inter-channel correlation propagation module of step 4;

[0015] Step 4: The diffusion completion feature matrix is ​​further optimized through the channel mutual information correlation propagation module to improve the quality of completion, and finally the reconstructed feature matrix is ​​output and sent to the downstream GNNs for training related tasks.

[0016] In step 1, a multi-perspective difficulty meter is constructed from two perspectives: feature and structure, to evaluate the completion difficulty of each node in the graph;

[0017] Considering the completion difficulty index from the perspective of features and structure, the completion difficulty D(v) of node v is defined as:

[0018] D(v)=D f (v)+β·D s (v)

[0019] Where β is the control structure difficulty D s (v) Weight hyperparameters are adjusted according to the requirements of specific tasks.

[0020] D f (v) is the completion difficulty score.

[0021] The characteristic angle is specifically:

[0022] For a given node v, define the k-order neighborhood is the set of all nodes that are at most k steps away from node v, by calculating its k-order neighborhood The number of initially known feature nodes n k (v) is used to measure the amount of neighborhood feature information, namely:

[0023]

[0024] Weighted aggregation sampling is used for neighborhood information of different orders. The weight of each order is the inverse of the total number of neighborhood nodes. As the neighborhood range expands, the weight of the amount of information gradually decreases, making the generated course sequence more consistent with the learning logic. The total feature information I(v) of node v is expressed as:

[0025]

[0026] Where K is the maximum considered neighborhood order.

[0027] According to the above information I(v), define the completion difficulty score D of node v f (v) is:

[0028]

[0029] Among them, I min and I max are the minimum and maximum values ​​of all node feature information, respectively, which are used for normalization. When the amount of neighborhood information I(v) is larger, the completion difficulty score D of the feature angle is f (v)The lower.

[0030] The structural angle is specifically:

[0031] The node degree d(v) and local clustering coefficient C(v) directly measure the closeness of a node's neighborhood. A lower degree or clustering coefficient usually means a higher completion difficulty. Define the structural difficulty score D s (v) Taking into account the above factors:

[0032]

[0033] When the node degree d(v) and clustering coefficient C(v) are smaller, the completion difficulty score D from the structural perspective is s The closer (v) is to 1, the more difficult it is to complete. min and d max are the minimum and maximum values ​​of all node degrees, respectively, C min and C max are the minimum and maximum values ​​of the local clustering coefficients of all nodes, respectively, and are used for normalization processing; ω1 and ω2 are weight parameters, satisfying ω1+ω2=1. Their values ​​are adjusted according to actual conditions to balance the influence of degree and clustering coefficient.

[0034] The step 2 is specifically as follows:

[0035] For each feature channel, based on the completion difficulty scores of all nodes in the channel calculated in step 1, these nodes are sorted in ascending order according to the difficulty scores from low to high;

[0036] Subsequently, a continuous diffusion scheduler is used to gradually introduce a certain proportion of nodes into the current subgraph according to the scoring order, so as to generate the subgraph structure to be diffused in each round, and generate a diffusion course from easy to difficult for each feature channel, that is, first introduce nodes with lower completion difficulty, and then gradually introduce nodes with higher difficulty.

[0037] Specifically, firstly, all missing nodes are sorted in ascending order according to the difficulty of completing the nodes under a given channel. According to the total number of diffusion layers L, a scheduling function is used to design the proportion of unknown nodes that need to be introduced in each diffusion layer λ l (0<λ l ≤1), that is, in the diffusion of the first layer, the simplest λ is introduced l The missing nodes of the proportion are diffused and supplemented; let λ0 be the initial proportion of missing nodes introduced during the initial diffusion, and the introduction proportion of missing nodes in the subsequent hierarchical diffusion considers two scheduling functions: linear scheduling and geometric scheduling:

[0038] Linear Scheduling:

[0039] Geometry Scheduling:

[0040] Linear scheduling increases the difficulty of completing nodes at a constant scheduling rate, while geometric scheduling introduces simple nodes at a slower pace to fully complete simple nodes.

[0041] Initially, all known nodes and missing nodes with a ratio of λ0 are used as the initial graph structure for feature diffusion propagation. After that, the missing nodes are continuously introduced into the diffusion process by using the scheduling function, and the graph structure is hierarchically expanded until all missing nodes are introduced.

[0042] The step 3 is specifically as follows: in the step 2, the scheduling function is used to control the speed of introducing nodes, and nodes are gradually added to the subgraph according to a certain ratio, so as to expand the subgraph and dynamically construct the subgraph structure to be diffused; on the subgraph after each round of expansion in the step 2, the feature weighted diffusion process based on the confidence of the completed features in the step 3 is executed;

[0043] Define the confidence ξ of node v for the completion feature of channel d v,d as follows:

[0044] ξ v,d =α D(v) , 0<α<1

[0045] According to the definition of the confidence function and the difficulty of completing missing nodes, the difficulty of completing a known node D(v) = 0, and the corresponding confidence is 1. The difficulty of completing an unknown node D(v) > 0, and the corresponding confidence is ξ v,d <1.

[0046] The weighted adjacency matrix under each channel is designed according to the feature confidence of the node. The connection weight between nodes is defined as the relative confidence of the two nodes, that is, from the node v i Passed to v j The weight of node v i The confidence level ξ i,d With node v j The confidence level ξ j,d The ratio of , thus constructing a weighted adjacency matrix for the N nodes of channel d As shown below:

[0047]

[0048] Among them, ξ i,d and j,d are the feature confidences of node i and node j in channel d, respectively, and A i,j is the adjacency relationship between nodes i and j in the graph adjacency matrix, A i,j =1 indicates that there is an edge connection between nodes i and j, A i,j = 0 means there is no connection between nodes, and the obtained represents the edge weight of the channel d that spreads from node i to node j.

[0049] Assume that there are k known feature nodes and u missing feature nodes in channel d. When diffusing, use the weighted adjacency matrix W (d) As the diffusion matrix, keep the known features of the current channel unchanged, that is, the block matrix corresponding to the known feature part of the diffusion matrix and Replaced with the unit matrix, corresponding to the unit matrix I in the upper half of the diffusion matrix kk and 0 ku , diffuse to the missing nodes according to different diffusion strengths, and the process of feature iterative diffusion is as follows:

[0050]

[0051] in, and represents the weighted adjacency matrix W of channel d (d) The block matrix corresponding to the u missing feature nodes in x (d) (n) is the characteristic column vector x composed of the eigenvalues ​​of all nodes in channel d (d) The feature vector obtained after n steps of iterative diffusion of the feature;

[0052] Perform feature diffusion propagation as shown above for each channel to obtain the initial completed feature vector after diffusion under each channel According to the node arrangement order of the original feature matrix X, the channel feature vectors are concatenated in channel order to obtain the reconstructed feature matrix

[0053] The step 4 is specifically as follows:

[0054] Assume that each node has F channels. After completing the weighted diffusion process of all channels in step 3, a preliminary complete feature matrix is ​​obtained. in is the completed feature vector composed of all the eigenvalues ​​of the completed channels of the i-th node. The eigenvalue of the i-th node on the d-th feature channel is

[0055] In order to analyze the correlation between the feature channels of the node and promote the propagation of these correlations between channels, the following steps are used to construct a channel correlation graph for node i to propagate the correlation between feature channels:

[0056] Estimate probability distribution: Use kernel density estimation (KDE) or Gaussian mixture model to estimate the joint probability density function and marginal probability density function of each pair of channels;

[0057] Calculate mutual information: For a pair of feature channel variables X and Y, the mutual information I(X; Y) can be defined as follows:

[0058]

[0059] Wherein p(x, y) is the joint probability distribution between channels, p(x) and p(y) are the marginal probability distributions of channel X and channel Y respectively; according to the joint probability density function and the marginal probability density function, the mutual information I(X; Y) between each pair of channels is calculated according to the mutual information calculation formula;

[0060] Construct the mutual information matrix: organize the mutual information values ​​between all channels into an F×F mutual information matrix M, where the (i, j) element of the mutual information matrix M is the channel f i With channel f j The mutual information value I(f i ;f j ),Right now

[0061] Construct a channel correlation graph: The mutual information matrix M is regarded as a weighted fully connected graph, where each node represents a channel and the weight of the edge is the corresponding mutual information value, and a channel correlation graph G is constructed. f =(V f , E f , W f ),in is the channel set, E f is the fully connected edge between channels, W f is the weighted adjacency matrix of the channel correlation graph, and the weight is the calculated mutual information value

[0062] Based on the constructed channel correlation graph, the completed feature matrix is ​​used to propagate the correlation between channels of each node, and the mean of each channel is used to eliminate the influence of the channel extreme value, so that the features between different channels can be propagated and optimized according to their correlation;

[0063] For each node i, its feature on the dth channel Update refinement is performed using the following propagation formula:

[0064]

[0065] where μ f is the feature mean of channel f.

[0066] Beneficial effects of the present invention:

[0067] Effectively deal with graph data with high feature missing rate:

[0068] The present invention learns from known features from the perspective of improving graph smoothness. Aiming at the common local homogeneity assumption in graph tasks, the Dirichlet energy value of nodes with unknown features is minimized, known feature values ​​are diffused and propagated, and feature values ​​that can improve the effect of GNNs downstream tasks are generated, rather than reconstructing the original features themselves. Differential completion is performed for each node, and hierarchical diffusion completion is performed by making full use of the known features under each channel, which helps GNNs to infer unknown nodes by aggregating the features of neighbors under high missing rates. Compared with the prior art that uses models such as autoencoders and generative networks to infer missing feature values, and only considers the similarity between the completed features and the original features, the present invention focuses on the performance contribution of the completed information to downstream tasks. Even when the feature missing rate is as high as 99.5%, downstream tasks can be completed more effectively, improving the overall prediction performance of the network and the adaptability to feature missing.

[0069] Course learning strategies differentiate to complete missing information:

[0070] This invention introduces the easy-to-difficult course learning strategy into the graph learning problem with missing data for the first time. By designing a multi-perspective difficulty meter, the difficulty of completing node information is measured from the feature and structure perspectives, and a continuous diffusion scheduler is designed to perform differentiated and hierarchical completion of nodes. Compared with the prior art that performs a global unified completion operation on all nodes and lacks differentiated processing using the relationship between node features, this invention performs progressive completion from the perspective of completion difficulty, alleviating the impact of nodes that are more difficult to complete on the completion process, thereby improving the overall completion effect.

[0071] Remove the interference of attribute initialization:

[0072] The present invention minimizes the energy of missing feature nodes in an energy function that characterizes the smoothness of the graph, generates a differential diffusion equation on the graph, and uses a numerical iterative scheme to solve it through Euler discretization. In the solution process, the conversion relationship between the graph Laplacian matrix and the adjacency matrix under the graph structure is fully utilized, and finally a recursive diffusion relationship that is independent of the initial value of the missing attribute is obtained. Compared with most existing technologies that require subsequent feature completion based on a certain initialization scheme, the hierarchical diffusion of the present invention is a diffusion completion method based on known features that is independent of the initialization scheme, successfully avoiding the noise interference that may be caused by attribute initialization, and removing the model's dependence on the initialization of missing attributes.

[0073] Consider the correlation between attributes:

[0074] The present invention captures the interdependence between channels through the mutual information between the channels of the nodes, constructs a channel correlation graph, and propagates on it according to the correlation. The existing technology usually processes the feature information of each channel in isolation, without considering the possible correlation between different attribute dimensions. The present invention uses the mutual information between channels to further optimize the features after diffusion completion, thereby improving the completion quality and enhancing the expression and generalization capabilities of the model.

[0075] Efficient and scalable without parameter completion:

[0076] The hierarchical iterative diffusion feature completion method of known features proposed in the present invention performs weighted feature diffusion on each channel respectively. In order to ensure the scalability of the method, when solving the energy minimization problem, the gradient flow equation is used instead of the Laplace matrix inversion operation with high computational complexity, so that the method can be expanded to large graphs with millions of nodes. In addition, the proposed scheme iteratively diffuses known features based on the completion difficulty, which is a parameter-free completion scheme, significantly improves the convergence speed of the model, and reduces the time and computational cost required for completion. Compared with some technologies that face insufficient computing resources on large graphs, the present invention enables the model to efficiently complete completion on large graphs and achieve good performance on downstream tasks.

[0077] In summary, the present invention solves the problems of existing missing feature completion methods, such as decreased model performance, homogenized processing of all nodes, reliance on initialization methods, high complexity and difficulty in expansion, etc., when information is highly missing, by designing known feature energy diffusion based on energy minimization, introducing a course learning strategy from easy to difficult, and a channel correlation propagation module based on mutual information. Ultimately, the technical solution of the present invention can improve the adaptability of the model to the scale of graph data and the degree of information missing, complete the completion for downstream tasks in an efficient and scalable manner, and improve the convergence speed of the model and the overall prediction performance of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 It is the overall system framework diagram of the present invention.

[0079] Figure 2 Schematic diagram of the multi-viewing angle difficulty measuring device of the present invention.

[0080] Figure 3 This is a schematic diagram of the continuous diffusion scheduler of the present invention.

[0081] Figure 4 This is a schematic diagram of weighted directed diffusion of the present invention.

[0082] Figure 5 Schematic diagram of channel mutual information correlation propagation of the present invention. DETAILED DESCRIPTION

[0083] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0084] A graph feature completion method based on course learning. In order to facilitate a more intuitive understanding of the structure, working principle and interaction between the modules of the entire invention, the technical solution is described in detail as follows in conjunction with the accompanying drawings.

[0085] Structure Description

[0086] Attached Figure 1 The overall framework diagram of the completion method proposed in the present invention is shown, including the information transmission path and interaction relationship between each module. It can be seen from the figure that the core structure of the entire invention is mainly divided into multi-perspective difficulty measurement and progressive diffusion propagation framework, and the framework can be subdivided into the following four modules: multi-perspective difficulty measurer, continuous diffusion scheduler, weighted directed diffusion module of subgraphs to be diffused, and channel mutual information correlation propagation module. Each module complements each other and constitutes the overall framework of the incomplete graph information completion method proposed in the present invention.

[0087] The completion process can be divided into the process of the multi-perspective difficulty meter comprehensively generating node difficulty scores from the feature perspective and the structural perspective, and then obtaining feature confidence scores; the process of the continuous diffusion scheduler gradually generating subgraphs to be diffused from easy to difficult; the process of weighted directed diffusion propagation of features on the subgraphs; and the process of mutual information propagation between all channels after the diffusion of each channel is completed. Finally, the completed feature matrix is ​​sent to the downstream GNNs for task training.

[0088] The detailed process of the proposed method is described as follows.

[0089] For the input undirected graph data V is the graph node set, E is the edge index, and X is the node feature matrix composed of F channels. For each feature channel f, all nodes are divided into a known feature node set V according to whether the node feature under this channel exists. k and missing feature node set V u , where V = V k ∪V u For each missing feature node v∈V in the missing feature node set u , use the feature difficulty meter to calculate the amount of known feature information in the 1st to Kth order neighborhood of each missing feature node, and perform weighted aggregation on the neighborhood information of different orders to obtain the feature difficulty score D of each missing node f (v) Use the node degree meter to calculate the degree and local clustering coefficient of each missing feature node, and perform weighted aggregation on the obtained structural information to generate the structural difficulty score D of each missing node. s(v) The generated feature difficulty score and structure difficulty score are combined to obtain the completion difficulty score D of each missing feature node. u (v) = D f (v)+β·D s (v) The proportion of the structural difficulty score is controlled by the hyperparameter β. Define the difficulty score D of all known feature nodes k (v) is 0, then the difficulty score list corresponding to the node in the graph is D(v) = D k (v)∪D u (v). Arrange all nodes in order of difficulty score from low to high, and calculate the feature confidence ξ(v) of each node based on the difficulty score of each node under the channel. Use the continuous diffusion scheduler to introduce the missing nodes to be diffused in this round from the rearranged missing node list, and generate the subgraph G to be diffused in this round according to the edge index between the nodes. l , and generate the weighted adjacency matrix W of the subgraph to be diffused according to the obtained node feature confidence ξ(v) l , perform weighted feature diffusion on the subgraph to be diffused. When all missing nodes in the channel are introduced into the subgraph and feature diffusion is completed, the complete features of all nodes in the channel are obtained. When all channels complete the above diffusion and completion process, the reconstructed node feature matrix is ​​obtained according to The eigenvalues ​​of each channel after completion of the middle node calculate the mutual information between each pair of channels, propagate the mutual information between channels, and output the final feature matrix after propagation between channels Sent to downstream GNNs for task training.

[0090] ① Multi-perspective difficulty meter module: Each node in the graph data may have multiple different attributes, which can be regarded as multiple feature channels. It is assumed that the number of feature channels of all nodes in the graph remains the same. The task of this module is to divide the nodes into known feature nodes and missing feature nodes for each feature channel according to the missing attributes of the nodes in the feature channel, measure the node completion difficulty under the channel from the feature perspective based on the richness of the known feature information around the node and the node structure perspective reflected by the node connectivity and local clustering coefficient, and generate the completion difficulty score of each node under the channel.

[0091] In the attached Figure 2In the figure, you can see the calculation process of the multi-view difficulty meter. For each feature channel, the feature difficulty meter examines the 1st to Kth order neighborhood information of all nodes under the channel, and calculates the feature difficulty score for each node based on the known feature information of each order neighborhood. At the same time, the structural difficulty meter obtains the degree and local clustering coefficient of each node based on the current subgraph structure, normalizes these structural information, and then weighted aggregates them to obtain the structural difficulty score of each node. Combining the scores of the feature difficulty meter and the structural difficulty meter, a comprehensive completion difficulty score is generated for each node under the feature channel, and all nodes are sorted in ascending order from low to high completion difficulty according to the comprehensive score.

[0092] ② Continuous Diffusion Scheduler Module: This module is based on the difficulty score obtained by the multi-view difficulty meter. It continuously introduces nodes with missing features into the diffusion process according to a linear or geometric scheduling function, dynamically generates a subgraph structure for each round of diffusion, and expands it hierarchically until all missing nodes are introduced.

[0093] In the attached Figure 3 The figure shows the scheduling process of the continuous diffusion scheduler under a certain channel. According to the scheduling rules, the proportion of missing nodes in each round of diffusion is obtained, and missing feature nodes with higher completion difficulty scores are continuously added to the subgraph in proportion, and the subgraph to be diffused in each round is dynamically expanded until the subgraph is expanded to be consistent with the original graph nodes. For missing nodes of different difficulty levels, the more difficult the completion difficulty is, the later they are introduced into the diffusion process.

[0094] ③ Subgraph weighted directed diffusion module: This module performs weighted directed diffusion propagation on the current subgraph according to each round of subgraphs to be diffused generated by the continuous diffusion scheduler, transforms the symmetric propagation of the undirected graph into a direction-dependent directed propagation method, and aggregates the completion or known features of neighboring nodes at different diffusion levels.

[0095] In the attached Figure 4 The process of weighted directed diffusion based on the node difficulty score is shown in Figure 2. The node difficulty scores after ascending order are converted into corresponding node feature confidence scores according to the conversion rules. The higher the completion difficulty, the lower the feature confidence score of the corresponding node. Different directions of propagation are performed based on the relative feature confidence between node pairs, so that the propagation weight in a certain direction between node pairs is the inverse of the weight in the opposite direction. The degree of information diffusion between nodes is associated with the completion difficulty to ensure the effectiveness of propagation in different directions.

[0096] ④ Channel mutual information correlation propagation module: After independently performing the hierarchical diffusion completion of the above modules on each feature channel, this module is designed to consider the feature correlation between different channels based on the mutual information between channels, and further optimize the obtained diffusion completion feature matrix to improve the quality of completion.

[0097] In the attached Figure 5 The specific process of channel correlation propagation based on mutual information is shown in Figure 1. The mutual information between all channels of the node is calculated, and a channel correlation graph is constructed for each node. The nodes in the graph represent channels, and the weights of the edges are the mutual information values ​​between channels. This optimizes the feature propagation between channels based on mutual information, and obtains the refined feature values ​​of each channel of each node, which are sent to the downstream GNNs for training related tasks.

[0098] Principle

[0099] This invention diffuses features by minimizing the Dirichlet energy in the graph, and combines it with the strategy of course learning to design a multi-perspective difficulty measurer to divide missing nodes into different levels from the perspective of completion difficulty, and diffuses and completes the missing information in a layered and iterative manner, thereby alleviating the adverse effects of node features that are more difficult to complete. The diffusion strength between nodes is designed based on the difficulty of completion, thereby improving the overall prediction performance of the network and the adaptability to feature missingness in an efficient and scalable manner.

[0100] ① Feature diffusion principle: Feature diffusion is the core algorithm basis of the entire method of the present invention. By using the Dirichlet energy to quantify the smoothness of the graph data, based on the smoothness assumption in the graph neural network, the Dirichlet energy of the missing nodes is minimized to diffuse the known features, generate the differential diffusion equation on the graph, and discretize and solve it to obtain the feature diffusion algorithm. On the basis of the principle of the feature diffusion algorithm, the difficulty of completing each node is further considered. To this end, a multi-perspective difficulty measurer and a continuous diffusion scheduler module are introduced. These two modules work closely together to gradually incorporate the nodes into the subgraph to be diffused in order from easy to difficult, forming a progressive subgraph feature diffusion process. In addition, the diffusion matrix in the feature diffusion algorithm is optimized according to the completion difficulty score of the node, and different diffusion levels are used to aggregate the known features or reconstruct features of neighboring nodes, thereby forming a subgraph weighted directed diffusion module. The following is a specific description of the feature diffusion principle.

[0101] In most graph data, connected nodes tend to have similar features or labels, which is also the basis for GNNs to effectively learn and propagate information. Graph neural networks usually assume that the signals on the graph are smooth, that is, the feature differences between adjacent nodes are small. This smooth assumption enables GNNs to avoid excessive feature fluctuations when propagating information on the graph structure, thereby improving the generalization ability of the model. Dirichlet energy is a commonly used indicator to measure the smoothness of graph signals, which reflects the changes in node features between its neighbors.

[0102] For the feature diffusion process in the subgraph weighted diffusion module, the principle of feature diffusion is explained as follows. For the undirected subgraph G = (V, E) generated by the continuous diffusion scheduler, its adjacency matrix is ​​expressed as N is the total number of nodes in the current subgraph, and the node feature matrix is ​​expressed as Where F is the dimension of the node feature. is the normalized adjacency matrix of the graph, is the graph Laplace matrix obtained by normalizing the graph adjacency matrix. The Dirichlet energy on the graph is in the form of a quadratic form of the graph Laplace matrix. For a subgraph G, its Dirichlet energy function l(x, G) is defined as follows:

[0103]

[0104] in, Represents the normalized adjacency matrix The value at (i, j), x i and x j Represent the feature vectors of node i and node j respectively.

[0105] From the definition of the Dirichlet energy function, it can be seen that if the Dirichlet energy value is low, it means that the signals of neighbors are more similar. Therefore, in order to improve the smoothness of the graph signal, it is necessary to minimize the Dirichlet energy value of nodes with unknown features. Based on this, a feature diffusion method based on minimizing the Dirichlet energy is proposed to iteratively diffuse known features and improve the smoothness of the graph signal, which helps GNNs to infer unknown nodes by aggregating the features of neighbors under high missing rates, thereby improving the overall prediction performance of the network and the adaptability to feature missing.

[0106] Since some node features are missing in the graph, the present invention forms a set of known feature node sets for each feature channel by combining the nodes with known features in each channel. The known node set under the dth feature channel is defined as The set of unknown nodes is According to whether the eigenvalue in channel d exists, the node feature x in channel d, the graph adjacency matrix A and the graph Laplacian matrix Δ are rearranged as follows:

[0107]

[0108] Among them, x, x k and x u is a column vector representing the characteristic signal of the graph node, x k is the characteristic column vector composed of all known node eigenvalues ​​under this channel, x uis the feature column vector corresponding to all nodes with missing features. The adjacency matrix A and the graph Laplacian matrix Δ are arranged in blocks according to the arrangement of the feature vector x. The block matrix A kk Represents the adjacency relationship between known nodes in the adjacency matrix, the block matrix A ku Represents the adjacency relationship between known nodes and unknown nodes, the block matrix A uk is the adjacency relationship between unknown nodes and known nodes, the block matrix A uu is the adjacency relationship between unknown nodes. Similarly, the block matrix Δ kk , Δ ku , Δ uk , Δ uu A kk , A ku , A uk and A uu The computed Laplacian block matrix.

[0109] The Dirichlet energy minimization problem can be solved by taking the derivative of the energy function and setting the derivative to 0 to obtain the eigenvalue x that minimizes the energy. By taking the derivative of the Dirichlet energy, we can get that its derivative is the graph Laplace matrix multiplied by the eigenvector of the node, that is Directly setting the derivative to 0 requires the inversion of the Laplace matrix, which has a high computational complexity and cannot be extended to large graphs with a large number of nodes. Therefore, we consider the gradient flow equation of the energy function instead. The gradient flow equation can be viewed as an initial condition with a feature vector x(0) with missing attributes (the first k rows represent known features and the last u rows represent missing features), and a boundary condition with known features x k Keeping the differential diffusion equation unchanged, the propagation form of the gradient flow is as follows:

[0110]

[0111] Since only the unknown part of the feature vector needs to be diffused during the diffusion process, and the value of the known feature remains unchanged, that is, the gradient corresponding to the known feature during the diffusion process is 0, the gradient flow of the known feature in the above formula is set to 0, as shown in the following formula:

[0112]

[0113] in Represents the energy gradient flow corresponding to the known feature nodes, represents the gradient flow corresponding to the missing feature node, x k is the known feature column vector composed of all known node features, x u (t) is the column vector corresponding to the missing feature node, Δ uk and Δ uuis a block matrix obtained by rearranging the Laplace matrix.

[0114] The solution to the equation is x u =lim t→∞ x u (t) is the node missing feature that minimizes the Dirichlet energy. When solving, the Euler method is used to discretize and solve the above gradient flow differential equation according to the numerical iteration scheme. The linear interpolation between each test point under the fixed step size (t = hk, h>0, k = 1, 2, ...) in the gradient descent is used to approximate the gradient flow. After discretizing the solution domain, it can be converted into a recursive expression as follows:

[0115]

[0116] Where x(n) is the characteristic column vector obtained after n steps of iteration of the eigenvalues ​​of all nodes, h is the step size used in Euler discretization, I is the unit matrix, Δ uk and Δ uu is a block matrix obtained by rearranging the Laplace matrix.

[0117] When the step size, i.e., the learning rate, h = 1, the above formula can be rearranged as:

[0118]

[0119] Since the graph Laplacian matrix Δ and the normalized adjacency matrix There exists , so the normalized adjacency matrix It can be expressed as follows:

[0120]

[0121] Using the definition of the Laplace matrix, the above iterative formula diffuses the small block -Δ in the matrix uk and I-Δ uu Use the normalized adjacency matrix corresponding to the block and Replace the representation and rewrite the iterative formula to obtain the final feature diffusion formula:

[0122]

[0123] When the step size h = 1, the derived final recursive expression can converge to its closed-form solution, and the steady-state solution does not depend on the initialization value x of the missing feature. u (0), so the missing features can be initialized to arbitrary values, removing the model’s dependency on the initialization of missing attributes.

[0124] Proof of convergence of recursive relation:

[0125] The recursive expression of feature diffusion is the core principle of the feature diffusion process in the subgraph weighted diffusion module. The final feature diffusion formula is obtained by discretizing and solving the Dirichlet energy gradient flow equation corresponding to the missing feature nodes:

[0126]

[0127] Among them, x(n) is the feature column vector composed of all node features obtained after n steps of iteration, I is the unit matrix, and is the normalized adjacency matrix after rearrangement The block matrix corresponding to the missing feature nodes in .

[0128] In order to illustrate the effectiveness and rationality of the feature diffusion process in the subgraph weighted diffusion module, the following is a detailed proof of the convergence of the obtained feature diffusion formula.

[0129] The feature column vector x in the feature diffusion recursive expression is represented in blocks according to whether the node feature exists, as shown in the following formula:

[0130]

[0131] in, Represents the known feature column vector corresponding to all known feature nodes after n steps of iteration, is the feature column vector corresponding to all missing feature nodes after n steps of iteration, I k , 0 ku is the block matrix in the identity matrix corresponding to the known feature node, and is the A corresponding to the missing feature nodes in the rearranged adjacency matrix A uk and A uu The resulting normalized adjacency matrix is ​​a block matrix.

[0132] The first k rows of the feature vector are the result of the diffusion of the known features. The known features remain unchanged during the diffusion process, which is recorded as Therefore, the convergence of the entire expression lies in the convergence of the next u rows, that is, The convergence of . By expanding the recursive relationship corresponding to the missing features and finding the limit to obtain its steady-state value, the process is as follows:

[0133]

[0134] Take the limit on both sides:

[0135] Considering When considering the convergence of Convergence of . If the submatrix of the normalized adjacency matrix is the convergence matrix, that is The first item Converges and converges to 0. For a square matrix A, its spectral radius is the maximum value of the modulus (absolute value) of all eigenvalues, that is, the spectral radius of square matrix A The spectral radius is less than 1, which is a necessary and sufficient condition for the matrix to converge. Therefore, it can be calculated by The spectral radius is used to judge the convergence. Since the eigenvalue range of the normalized graph Laplacian matrix Δ is [0, 2], and the normalized adjacency matrix It can be concluded that its eigenvalue range is between [-1, 1], so the spectral radius of the normalized adjacency matrix is And obviously the adjacency matrix submatrix and Therefore, there is Depend on Knowable Matrix converges, so we have

[0136] Consider the second term The convergence of . is a common ratio of the submatrix determinant The geometric series of It can be seen that all its eigenvalues ​​are less than 1, so the matrix Reversible, and So the geometric series converges and converges to And by Available So the second series Finally converges to

[0137] In summary, the recursive relationship corresponding to the missing features Converges, and eventually converges to As shown below:

[0138]

[0139] If the Euler method is not used to discretize the gradient flow equation, the derivative is directly used to solve the minimum value of the energy corresponding to the missing feature, that is, Solve the problem. The solution process is as follows:

[0140]

[0141] Let Δ uk x k +Δ uux u =0, we get Since there is a The relationship between , so the normalized adjacency matrix can be expressed as follows:

[0142]

[0143] From the above formula, we can know Substitute it into The solution to the Dirichlet energy minimization problem using derivatives is This is consistent with the convergence result of the above recursive relation.

[0144] The final convergence result is consistent with the result obtained by directly performing a more complex inversion operation on the gradient flow equation, which illustrates the effectiveness of solving the problem using the Euler discretization method, and further demonstrates the interpretability and efficiency of using the diffusion equation for feature diffusion propagation.

[0145] According to the derived recursive expression, the diffusion propagation process based on Dirichlet energy minimization can be summarized as follows: the node eigenvector is left-multiplied by the diffusion matrix (i.e., the normalized adjacency matrix). Since the upper block matrix of the original adjacency matrix is ​​not a unit matrix, in order to ensure consistency with the result of the recursive expression, the known eigenvalues ​​are reset to the initial values ​​after each diffusion step to simulate the effect of using a unit matrix. This feature reset operation also serves as a diffusion constraint to prevent over-smoothing, which makes the final node eigenvector meaningless.

[0146] According to the above feature diffusion reconstruction step based on known features, the propagated completed feature matrix is ​​obtained and sent to GNN to complete the training of downstream tasks. The reconstruction step is based on Dirichlet energy minimization to generate a diffusion differential equation on the graph. Discretizing and solving the differential equation can obtain a simple, fast, and scalable feature diffusion propagation iterative algorithm.

[0147] ②Course learning strategy: According to the effectiveness of the feature diffusion process to reconstruct missing features, the advantages of energy diffusion propagation and curriculum learning are combined to design a multi-view difficulty meter module and a continuous diffusion scheduler module. When performing feature diffusion, the multi-view difficulty meter measures the difficulty of node completion based on the amount of effective information around the missing node and the complexity of the node's neighborhood. The continuous diffusion scheduler imitates the progressive learning strategy of the human learning process from simple to abstract and complex courses, and gradually introduces the node extension subgraph structure in the order of completion difficulty from easy to difficult, first diffusing on easy data, and then gradually expanding to more difficult data. Combined with the node multi-view difficulty meter and diffusion scheduler, hierarchical diffusion propagation is carried out to alleviate the adverse effects of nodes that are more difficult to complete, thereby improving the model's feature completion effect on graph data and its performance on downstream tasks.

[0148] The training strategy of curriculum learning from easy to difficult has been proven to improve the generalization ability and robustness of the model, so it is widely used in various scenarios such as computer vision and natural language processing. However, the use of curriculum learning strategies has not been investigated in the field of graph information completion of missing data. The existing technical solutions treat all nodes equally and randomly complete the missing nodes, which may lead to suboptimal feature representation. Therefore, this technical design gradually introduces nodes for diffusion completion according to the difficulty of node completion from simple to difficult, and uses the feature information and structural information around the nodes to involve a multi-perspective difficulty meter to measure the difficulty of node completion. On this basis, a continuous diffusion scheduler is used to continuously select appropriate nodes in each layer of diffusion to feed into the hierarchical feature diffusion model for completion. The more difficult the nodes are to complete, the later they are introduced into the diffusion process, reducing their adverse effects in the initial completion.

[0149] Multi-perspective difficulty meter description: The multi-perspective difficulty meter is used to calculate the completion difficulty of each node in a given graph. The completion difficulty will combine the feature information and structural complexity of the node neighborhood. The multi-perspective difficulty meter is constructed from both feature and structure perspectives to evaluate the completion difficulty of each node in the graph.

[0150] I. Feature perspective: First, consider how to identify difficult nodes from a feature perspective. Generally speaking, neighborhood aggregation in GNNs benefits from the homogeneity of the graph, and energy diffusion follows a similar principle. The closer to the heat source, the greater the degree of energy diffusion. If a node's neighborhood features are rich and highly correlated with its own features, the completion difficulty of the node is low; conversely, if the node's neighbor features are sparse or irrelevant, the completion difficulty is high. Therefore, the degree of completion is considered based on the amount of known feature information in the node's k-order neighborhood. Nodes with greater information content are regarded as simple samples and given lower difficulty scores, while nodes with less information content are regarded as difficult samples and given higher difficulty scores.

[0151] For a given node v, define the k-order neighborhood is the set of all nodes that are at most k steps away from node v, by calculating its k-order neighborhood The number of initially known feature nodes n k (v) is used to measure the amount of neighborhood feature information, namely:

[0152]

[0153] Considering that the diffusion degree will gradually decrease as the neighborhood expands, in order to better measure the node difficulty, weighted aggregation sampling is used for neighborhood information of different orders. The weight of each order is the inverse of the total number of neighborhood nodes. As the neighborhood range expands, the weight of the amount of information gradually decreases, making the generated course sequence more consistent with the learning logic. Therefore, the total feature information amount I(v) of node v can be expressed as:

[0154]

[0155] Where K is the maximum considered neighborhood order.

[0156] According to the above information I(v), define the completion difficulty score D of node v f (v) is:

[0157]

[0158] Among them, I min and I max are the minimum and maximum values ​​of all node feature information, respectively, which are used for normalization. When the amount of neighborhood information I(v) is larger, the difficulty of completing the feature angle D f (v)The lower.

[0159] II. Structural perspective: In graph data, the structural characteristics of the graph, such as the connectivity, community affiliation, and global position of nodes, have an important impact on the node completion process. Nodes with high structural complexity are usually located in key positions of the network, and may involve multiple communities or be in complex subgraph patterns, which increases the difficulty of completion. By analyzing the degree, clustering coefficient, and centrality indicators (such as betweenness centrality and closeness centrality) of the nodes, we can have a more comprehensive understanding of the role and environment of the nodes in their graph structure. The measurement of structural information helps to identify those nodes that are topologically difficult to complete, thereby guiding the diffusion process to prioritize the diffusion of simpler nodes, gradually transition to more complex nodes, and improve the overall completion effect.

[0160] The node degree d(v) and local clustering coefficient C(v) can directly measure the closeness of a node's neighborhood. A lower degree or clustering coefficient usually means a higher completion difficulty. For example, isolated nodes or weakly connected nodes lack effective neighborhood information compared to other nodes and are more difficult to propagate and complete. Define the structural difficulty score D s (v) Taking into account the above factors:

[0161]

[0162] When the node degree d(v) and clustering coefficient C(v) are smaller, the completion difficulty score D from the structural perspective is s The closer (v) is to 1, the more difficult it is to complete. min and d max are the minimum and maximum values ​​of all node degrees, respectively, C min and C max are the minimum and maximum values ​​of the local clustering coefficients of all nodes, respectively, and are used for normalization. ω1 and ω2 are weight parameters, satisfying ω1+ω2=1. Their values ​​are adjusted according to actual conditions to balance the influence of degree and clustering coefficient.

[0163] Considering the completion difficulty index from the perspective of features and structure, the completion difficulty D(v) of node v is finally defined as:

[0164] D(v)=D f (v)+β·D s (v)

[0165] Where β is the control structure difficulty D s (v) Weight hyperparameters are adjusted according to the needs of specific tasks. For example, the value of β can be appropriately reduced in feature-rich graph data, while the value of β can be increased in structure-complex graph data to improve the impact of structural factors. Continuous Diffusion Scheduler Description:

[0166] After measuring the completion difficulty of each node, a continuous diffusion scheme is designed to generate a diffusion course from easy to difficult for each feature channel.

[0167] Specifically, firstly, all missing nodes are sorted in ascending order according to the difficulty of completing the nodes under a given channel. According to the total number of diffusion layers L, a scheduling function is used to design the proportion of unknown nodes that need to be introduced in each diffusion layer λ l (0<λ l ≤1), that is, in the first layer diffusion, the simplest λ is introduced l The proportion of missing nodes is diffused and completed. Let λ0 be the initial proportion of missing nodes introduced during the initial diffusion. The proportion of missing nodes introduced in the subsequent hierarchical diffusion considers two scheduling functions: linear scheduling and geometric scheduling.

[0168] Linear Scheduling:

[0169] Geometry Scheduling:

[0170] Linear scheduling increases the difficulty of completing nodes at a constant scheduling rate, while geometric scheduling introduces simple nodes at a slower pace to fully complete simple nodes. Initially, all known nodes and missing nodes with a λ0 ratio are used as the initial graph structure for feature diffusion propagation. After that, the missing nodes are continuously introduced into the diffusion process by using the scheduling function, and the graph structure is hierarchically expanded until all missing nodes are introduced. For missing nodes of different difficulty levels, the more difficult the completion difficulty is, the later they are introduced into the diffusion process.

[0171] ③ Weighted diffusion of completion difficulty: In the diffusion process of each layer, the normalized adjacency matrix is ​​used as the diffusion matrix in the recursive relationship obtained based on energy minimization above, and the known features under each feature channel are diffused and propagated independently. The missing features are reconstructed by aggregating the known or completed features of the neighboring nodes under the channel. The normalized adjacency matrix is ​​a symmetric matrix, and the message transmission between two nodes is carried out with the same intensity regardless of the diffusion direction. Therefore, the node completion difficulty obtained by the multi-view difficulty meter in the course learning will be re-weighted as the channel adjacency matrix of the diffusion matrix, so that the nodes with lower completion difficulty in this layer transmit more information outward, while the nodes with higher completion difficulty diffuse less information outward and receive more information.

[0172] Since the known feature nodes themselves have available known information, there is no need to complete them, so the completion difficulty is 0, while the completion difficulty of the remaining missing nodes is greater than 0. The nodes with lower difficulty are easier to complete, and the confidence of the completed features is higher. As the completion difficulty increases, the confidence of the completed features will gradually decrease. Therefore, the confidence of node v for the completed features of channel d is defined as v,d as follows:

[0173] ξ v,d =α D(v) , 0<α<1

[0174] According to the definition of the confidence function and the difficulty of completing missing nodes, the difficulty of completing a known node D(v) = 0, and the corresponding confidence is 1, while the difficulty of completing an unknown node D(v) > 0, so the corresponding confidence ξ v,d <1. The weighted adjacency matrix under each channel is designed according to the feature confidence of the node. The connection weight between nodes is defined as the relative confidence of the two nodes, that is, from the node v i Passed to v j The weight of node v iThe confidence level ξ i,d With node v j The confidence level ξ j,d The ratio of , thus constructing a weighted adjacency matrix for the N nodes of channel d As shown below:

[0175]

[0176] Among them, ξ i,d and j,d are the feature confidences of node i and node j in channel d, respectively, and A i,j is the adjacency relationship between nodes i and j in the graph adjacency matrix, A i,j =1 indicates that there is an edge connection between nodes i and j, A i,j = 0 means there is no connection between nodes, and the obtained represents the edge weight of the channel d that spreads from node i to node j.

[0177] Compared with the symmetric adjacency matrix A used before, the current weighted adjacency matrix W is an asymmetric matrix, that is, the information diffusion intensity between nodes is related to the direction. The message propagation intensity from high-confidence nodes to low-confidence nodes is large, while the message transmission amount from low-confidence nodes to high-confidence nodes in the opposite direction will be small. By designing the propagation weight based on the completion difficulty, the unified propagation method of the undirected graph is transformed into propagation of different intensities on the directed graph.

[0178] Assume that there are k known feature nodes and u missing feature nodes in channel d. During diffusion, the weighted adjacency matrix W ( d ) as the diffusion matrix, keeping the known features of the current channel unchanged, that is, the block matrix corresponding to the known feature part of the diffusion matrix and Replaced with the unit matrix, corresponding to the unit matrix I in the upper half of the diffusion matrix kk and 0 ku , diffuse to the missing nodes according to different diffusion strengths, and the process of feature iterative diffusion is as follows:

[0179]

[0180] in, and represents the weighted adjacency matrix W of channel d (d) The block matrix corresponding to the u missing feature nodes in x (d) (n) is the characteristic column vector x composed of the eigenvalues ​​of all nodes in channel d (d) The feature column vector is obtained after n steps of iterative diffusion of the feature.

[0181] Perform feature diffusion propagation as shown above for each channel to obtain the initial completed feature vector after diffusion under each channel According to the node arrangement order of the original feature matrix X, the channel feature vectors are concatenated in channel order to obtain the reconstructed feature matrix

[0182] ④ Channel mutual information correlation propagation: After weighted diffusion of each channel, the feature matrix reconstructed based on the difficulty of completing each node is obtained. However, the attribute features of the nodes in the graph data do not exist in isolation, and may influence and interact with each other. Mutual information is a measure of the degree of mutual dependence between random variables, which can capture the dependence between different channels. Compared with the correlation coefficient that only measures the linear relationship between variables, mutual information can capture a wider range of complex nonlinear relationships, thereby reflecting more general dependencies, and has greater flexibility and applicability. Therefore, based on the mutual information between each feature channel, the possible intrinsic connection between different channels is considered, and the correlation propagation between channels is performed to further improve the quality of completion and enhance the expression and generalization capabilities of the model.

[0183] In graph data, different channels of each node can be regarded as different random variables. By calculating the mutual information between channels, the amount of information directly shared by the variables can be described from the perspective of information theory, and the dependencies between them can be quantified. This dependency can help identify which channels have strong correlations, so that the model can better utilize these dependencies for feature propagation and optimization;

[0184] For a pair of feature channel variables X and Y, the mutual information, (X; Y) can be defined as follows:

[0185]

[0186] Where p(x, y) is the joint probability distribution between channels, and p(x) and p(y) are the marginal probability distributions of channel X and channel Y, respectively. The larger the mutual information, the stronger the dependency between the two variables; conversely, the smaller the value of the mutual information, the weaker the dependency between the two variables. When two variables are completely independent, the mutual information is 0. For continuous variables, the mutual information is replaced by the double integral as follows:

[0187]

[0188] Where p(x, y) is the joint probability density function of the current variables X and Y, and p(x) and p(y) are the marginal probability density functions of X and Y, respectively. Mutual information is the intrinsic dependence between the joint distribution of X and Y relative to the joint distribution assuming that X and Y are independent. Assume that each node has F channels, and the eigenvalue of the i-th node on the d-th channel is In order to carry out inter-channel propagation, the following steps are used to construct a channel correlation graph for correlation propagation:

[0189] Estimating probability distribution: Since the probability density function of continuous features is difficult to estimate accurately, the joint probability density function and marginal probability density function of each pair of channels are estimated using kernel density estimation (KDE) or Gaussian mixture model.

[0190] Calculating mutual information: According to the joint probability density function and the marginal probability density function, the mutual information I(X;Y) between each pair of channels is calculated according to the mutual information definition formula.

[0191] Construct the mutual information matrix: organize the mutual information values ​​between all channels into an F×F mutual information matrix M, where the (i, j) element of the mutual information matrix M is the channel f i With channel f j The mutual information value I(f i ;f j ),Right now

[0192] Construct a channel correlation graph: The mutual information matrix M is regarded as a weighted fully connected graph, where each node represents a channel and the weight of the edge is the corresponding mutual information value. Construct a channel correlation graph G f =(V f , E f , W f ),in is the channel set, E f is the fully connected edge between channels, W f is the weighted adjacency matrix of the channel correlation graph, and the weight is the calculated mutual information value

[0193] Based on the constructed channel correlation graph, the completed feature matrix is ​​used to propagate the correlation between channels of each node, and the mean of each channel is used to eliminate the influence of channel extreme values, so that the features between different channels can be propagated and optimized according to their correlation.

[0194] For each node i, its feature value on the dth channel Update refinement is performed using the following propagation formula:

[0195]

[0196] where μ f is the feature mean of channel f.

[0197] Based on the characteristic values ​​of each channel that are diffused and completed based on the completion difficulty, message propagation is performed among all channels for each node according to the channel propagation weight, so as to obtain the refined characteristic values ​​of each channel of each node, and the refined characteristic vectors obtained by each node are The final attribute completion matrix is ​​obtained by splicing according to the node arrangement order And sent to downstream GNNs for training of related tasks.

[0198] Action relationship description

[0199] The action relationship between modules is mainly reflected in the transmission of information flow and the interaction between modules:

[0200] The multi-perspective difficulty meter module extracts and analyzes the feature information and structural information of the nodes under each channel, and generates the final completion difficulty score of each node based on the scores obtained by the comprehensive feature and structural difficulty meters. After sorting all the scores in ascending order of difficulty, it outputs the difficulty score sequence of all nodes from easy to difficult for the continuous diffusion scheduler module to call and use.

[0201] The continuous diffusion scheduler module selects missing nodes in proportion from easy to difficult based on the input difficulty score according to the linear or geometric scheduling function, generates the subgraph structure to be diffused in each round, and passes it cyclically to the subgraph weighted directed diffusion module.

[0202] The subgraph weighted directed diffusion module converts the difficulty score output by the multi-view difficulty meter into the node feature confidence, and performs weighted directed diffusion propagation on the current subgraph to be diffused provided by the continuous diffusion scheduler, so that the feature reaches the convergence value through multiple iterations. Diffusion completion is performed on each subgraph in each round until the subgraph to be diffused has diffused to the same level as the original graph node, and the diffused features of all nodes under this channel are obtained.

[0203] After performing the hierarchical diffusion completion of the above modules independently on all channels, the preliminary completed feature matrix is ​​output to the channel mutual information correlation propagation module. The channel mutual information correlation propagation module considers the feature correlation between different channels, defines the mutual information between channels to describe the correlation between different channels, further propagates and optimizes the diffusion completion feature matrix, and finally outputs the refined feature matrix.

[0204] ① Alternatives to feature diffusion: In the feature diffusion stage based on energy minimization, the known features of each channel are iteratively diffused to different degrees, and finally converged to the steady-state value that minimizes the energy of the missing features. In addition to diffusing all known features, you can also consider using the low-order structural information of the graph to analyze the position and role of the nodes in the graph, and diffuse the nodes of different status separately. For example, nodes in the same community often have similar features. The community structure is used to diffuse and fill in the communities to improve the node differences between communities.

[0205] In addition, feature propagation diffusion can not only be performed locally, but also combine some global information. Through multi-scale feature fusion, the model can capture the feature information of nodes at different levels. For example, local propagation can capture the relationship between a node and its direct neighbors, while global propagation can capture the position and role of the node in the entire graph. This multi-level feature fusion enables the model to understand the feature distribution of nodes more comprehensively, thereby better coping with a high proportion of missing data.

[0206] ②Alternative solutions that spread gradually from easy to difficult:

[0207] For the course learning strategy from easy to difficult, the multi-perspective difficulty meter based on features and structures in the present invention can be replaced by measurement rules from other different angles, for example, using information entropy, label information, and task-specific goals to quantify the difficulty of nodes or edges, or designing a self-supervised learning method to measure the difficulty of nodes, giving priority to completing simple nodes and gradually expanding to complex nodes. In addition to the gradual introduction of nodes, it is also possible to consider the dependencies between nodes from the perspective of edges and gradually integrate more edges for diffusion.

[0208] For the continuous diffusion scheduler, in addition to the linear and geometric scheduling functions in the present invention, an adaptive scheduling mechanism can also be considered to dynamically adjust the course speed according to the current diffusion state of the model, replacing human drive with data drive, so that the model can dynamically adapt to the training state according to feedback, and achieve more flexible scheduling control.

[0209] ③Alternatives to weighted diffusion:

[0210] In the present invention, the node completion difficulty obtained by the feature and structure multi-view difficulty meter is used as the credibility of the node feature, and the relative credibility between the node pairs is used as the diffusion weight to perform different degrees of diffusion propagation between the known feature nodes and the reconstructed feature nodes and the reconstructed nodes. In addition to using the completion difficulty as the credibility, it is also possible to consider using weighted diffusion based on information entropy and attention mechanism. In this method, information entropy is used to evaluate the information uncertainty of the features between nodes, and the attention mechanism is used to guide the propagation of information between node pairs, thereby improving the completion effect.

[0211] ④Alternatives to inter-channel correlation:

[0212] When considering the correlation between different attributes, in addition to measuring the correlation through mutual information in information theory, other measurement methods in statistical methods can also be considered, such as Pearson correlation coefficient, covariance matrix, distance correlation, kernel function, etc., to capture different types of relationships between random variables. In addition, a multi-head attention mechanism can be introduced to dynamically weight different channel information, and different types of dependencies between channels can be captured through multiple attention heads, thereby obtaining richer feature representations.

[0213] like Figure 1 As shown in FIG. 1 , it is the overall framework diagram of the completion method proposed in the present invention, which includes the information transmission path and interaction relationship between each module. The arrows in the figure indicate the whole process of the incomplete graph with missing information under a certain channel, from the multi-view difficulty meter generating node difficulty scores to generating feature confidence scores, the continuous diffusion scheduler progressively generating diffusion subgraphs from easy to difficult, the weighted directed diffusion propagation of subgraph features, and the mutual information propagation between all channels after the diffusion of each channel is completed, and finally sending the completed feature matrix to the downstream GNNs.

[0214] Figure 2 The pseudo code of the detailed process of the completion method proposed in the present invention is shown. Lines 4-20 in the pseudo code describe the process of measuring the node difficulty score from the feature perspective and the structural perspective; lines 26-30 record the process of the continuous diffusion scheduler gradually generating the subgraph to be diffused from easy to difficult; lines 31-38 record the process of weighted directed diffusion on each round of subgraph to be diffused, and lines 35-38 iteratively diffuse the feature matrix during the diffusion process, and restore the known features unchanged after each iteration, simulating the recursive process of the derived diffusion recursion; lines 41-44 record the process of feature optimization by propagating mutual information between channels, and finally output the completed feature matrix, which is sent to the downstream GNNs for task training.

[0215] Figure 3The calculation process of the multi-perspective difficulty meter is demonstrated. Taking the node difficulty measurement under a specific channel as an example, it is assumed that the known feature nodes under this channel are 2, 5, 10, and 14, and the channel features of the remaining nodes are missing. Assuming that the hyperparameter K=3, the node completion difficulty is measured from the feature perspective and the structural perspective. For each unknown node, the known feature information in the 1st to 3rd order neighborhood is calculated and summarized, and the feature difficulty score is obtained according to the calculation rules of the feature difficulty meter. In addition, according to the degree and local clustering coefficient of each missing feature node, it is input into the structural difficulty meter to obtain the structural difficulty score of each node. The figure below shows the features and structural related information of nodes 9, 13, and 16. The completion difficulty of all nodes is obtained by combining the feature difficulty score and the structural difficulty score, and they are sorted from easy to difficult as shown in the figure below.

[0216] Figure 4 The scheduling process of the continuous diffusion scheduler is demonstrated. Taking the case of missing node features under a specific channel as an example, it is assumed that the known feature nodes under this channel are 2, 5, 10, and 14, and the channel features of the remaining nodes are all missing. After obtaining the completion difficulty score of each node from the multi-view difficulty meter, it sorts them from easy to difficult. According to the linear or geometric scheduling function used, the corresponding proportion of nodes are selected from the sorted missing feature node list from easy to difficult, and the subgraph structure corresponding to each round of diffusion is dynamically generated until all the nodes in the original graph are included in the subgraph, and the scheduling process ends.

[0217] Figure 5 The process of weighted directed diffusion based on the node difficulty score is shown. Taking the node difficulty score under channel d as an example, it is assumed that the known feature nodes under this channel are 2, 5, 10, and 14, and the channel features of the remaining nodes are missing. After the completion difficulty score of each node is obtained by the multi-view difficulty meter, the corresponding node feature confidence score ξ is generated according to the conversion rule v,d , the higher the completion difficulty of the node, the lower the corresponding feature confidence. According to each round of the to-be-diffused subgraph generated by the continuous scheduler, weighted directed diffusion propagation is performed on the current subgraph. The propagation weight is the relative feature confidence between the node pairs. The propagation weight in a certain direction between the node pairs is the inverse of the weight in the reverse direction, so that the message propagation intensity from the high-confidence node to the low-confidence node is large, while the message transmission amount in the opposite direction from the low-confidence node to the high-confidence node will be very small. By designing the propagation weight through the completion difficulty, the unified propagation method of the undirected graph is transformed into propagation of different intensities between node pairs on the directed graph.

[0218] The process of channel correlation propagation based on mutual information. Assume that there are 12 nodes in the graph, and each node has 5 feature channels. After the progressive weighted diffusion of each feature channel of all nodes is completed, the initial completed feature matrix is ​​obtained, as shown in the feature matrix on the left in the figure below. Calculate the mutual information between each channel, and construct a channel correlation fully connected graph for each node based on the corresponding channel mutual information. Each node in the channel graph represents a channel, and the weight of the edge is the mutual information value between channels. The 12 nodes in the figure below correspond to 12 channel fully connected graphs. Each node performs feature propagation optimization based on mutual information between channels on its own channel correlation graph, so as to obtain the refined feature values ​​of each channel of each node, which corresponds to the feature matrix on the right in the figure below. The feature matrix generated after channel correlation propagation is sent to the downstream GNNs for training related tasks.

Claims

1. A graph feature completion method based on curriculum learning, characterized in that: The steps include: Step 1: Calculate the completion difficulty of each node in a given undirected graph through a multi-view difficulty meter, and evaluate and score the completion difficulty of each node in the graph by combining the feature information and structural complexity of the node neighborhood; Step 2: for the score, continuously introduce nodes with missing features into the diffusion process according to a linear or geometric scheduling function through a continuous diffusion scheduler module, and dynamically generate an undirected subgraph structure for each round of diffusion; Step 3: Perform weighted directed diffusion propagation on the undirected subgraph structure, convert the symmetric propagation between nodes of the undirected subgraph structure into a directed propagation method related to direction, and aggregate the completion or known features of neighboring nodes at different diffusion levels; when all missing nodes are introduced into the subgraph and feature diffusion is completed, a preliminary completion feature matrix is ​​obtained; Step 4: The diffusion completion feature matrix is ​​further optimized through the channel mutual information correlation propagation module, and finally the reconstructed feature matrix is ​​output and sent to the downstream GNNs for training related tasks.

2. A graph feature completion method based on curriculum learning according to claim 1, characterized in that: In step 1, a multi-perspective difficulty meter is constructed from two perspectives: feature and structure, to evaluate the completion difficulty of each node in the graph; Considering the completion difficulty index from the perspective of features and structure, the completion difficulty D(v) of node v is defined as: D(v)=D f (v)+β·D s (v) Where β is the control structure difficulty D s (v) Weight hyperparameters, which are adjusted according to the requirements of specific tasks; D f (v) is the completion difficulty score.

3. A graph feature completion method based on curriculum learning according to claim 2, characterized in that: The characteristic angle is specifically: For a given node v, define the k-order neighborhood is the set of all nodes that are at most k steps away from node v, by calculating its k-order neighborhood The number of initially known feature nodes n k (v) is used to measure the amount of neighborhood feature information, namely: Weighted aggregation sampling is used for neighborhood information of different orders. The weight of each order is the inverse of the total number of neighborhood nodes. The total feature information I(v) of node v is expressed as: Where K is the maximum considered neighborhood order; According to the total feature information I(v), define the completion difficulty score D of node v f (v) is: Among them, I min and I max are the minimum and maximum values ​​of all node feature information, respectively, which are used for normalization. When the amount of neighborhood information I(v) is larger, the completion difficulty score D of the feature angle is f (v)The lower.

4. A graph feature completion method based on curriculum learning according to claim 2, characterized in that: The structural angle is specifically: The node degree d(v) and local clustering coefficient C(v) directly measure the closeness of a node’s neighborhood and define the structural difficulty score D s (v) Taking into account the above factors: When the node degree d(v) and clustering coefficient C(v) are smaller, the completion difficulty score D from the structural perspective is s The closer (v) is to 1, the more difficult it is to complete. min and d max are the minimum and maximum values ​​of all node degrees, respectively, C min and C max are the minimum and maximum values ​​of the local clustering coefficients of all nodes, respectively, and are used for normalization processing; ω1 and ω2 are weight parameters, satisfying ω1+ω2=1. Their values ​​are adjusted according to actual conditions to balance the influence of degree and clustering coefficient.

5. The graph feature completion method based on curriculum learning according to claim 2, characterized in that: The step 2 is specifically as follows: For each feature channel, based on the completion difficulty scores of all nodes in the channel calculated in step 1, these nodes are sorted in ascending order according to the difficulty scores from low to high; Subsequently, a continuous diffusion scheduler is used to gradually introduce a certain proportion of nodes into the current subgraph according to the scoring order, so as to generate the subgraph structure to be diffused in each round, and generate a diffusion course from easy to difficult for each feature channel, that is, first introduce nodes with lower completion difficulty, and then gradually introduce nodes with higher difficulty.

6. A graph feature completion method based on curriculum learning according to claim 5, characterized in that: First, all missing nodes are sorted in ascending order according to the difficulty of completing the nodes under a given channel. According to the total number of diffusion layers L, a scheduling function is used to design the proportion of unknown nodes λ1 that needs to be introduced in each diffusion layer, that is, in the diffusion of the lth layer, the simplest λ l The missing nodes of the proportion are diffused and supplemented; let λ0 be the initial proportion of missing nodes introduced during the initial diffusion, and the introduction proportion of missing nodes in the subsequent hierarchical diffusion considers two scheduling functions: linear scheduling and geometric scheduling: Linear Scheduling: Geometry Scheduling: Linear scheduling increases the difficulty of completing nodes at a constant scheduling rate, while geometric scheduling introduces simple nodes at a slower pace to fully complete simple nodes. Initially, all known nodes and missing nodes with a ratio of λ0 are used as the initial graph structure for feature diffusion propagation. After that, the missing nodes are continuously introduced into the diffusion process by using the scheduling function, and the graph structure is hierarchically expanded until all missing nodes are introduced.

7. A graph feature completion method based on curriculum learning according to claim 6, characterized in that: The step 3 is specifically as follows: Define the confidence ξ of node v for the completion feature of channel d v,d as follows: x v,d =a D(v) ,0<α<1 According to the definition of the confidence function and the difficulty of completing missing nodes, the difficulty of completing a known node D(v) = 0, and the corresponding confidence is 1. The difficulty of completing an unknown node D(v)>0, and the corresponding confidence is ξ v,d <1; The weighted adjacency matrix under each channel is designed according to the feature confidence of the node. The connection weight between nodes is defined as the relative confidence of the two nodes, that is, from the node v i Passed to v j The weight of node v i The confidence level ξ i,d With node v j The confidence level ξ j,d The ratio of , thus constructing a weighted adjacency matrix for the N nodes of channel d As shown below: Among them, ξ i,d and j,d are the feature confidences of node i and node j in channel d, respectively, and A i,j is the adjacency relationship between nodes i and j in the graph adjacency matrix, A i,j =1 indicates that there is an edge connection between nodes i and j, A i,j = 0 means there is no connection between nodes, and the obtained represents the edge weight of the channel d that spreads from node i to node j.

8. The method for graph feature completion based on curriculum learning according to claim 7, characterized in that: Assume that there are k known feature nodes and u missing feature nodes in channel d. When diffusing, use the weighted adjacency matrix W( d ) as the diffusion matrix, keeping the known features of the current channel unchanged, that is, the block matrix corresponding to the known feature part of the diffusion matrix and Replaced with the unit matrix, corresponding to the unit matrix I in the upper half of the diffusion matrix kk and 0 ku , diffuse to the missing nodes according to different diffusion strengths, and the process of feature iterative diffusion is as follows: in, and represents the weighted adjacency matrix W of channel d (d) The block matrix corresponding to the u missing feature nodes in x (d) (n) is the characteristic column vector x composed of the eigenvalues ​​of all nodes in channel d (d) The feature vector obtained after n steps of iterative diffusion of the feature; Perform feature diffusion propagation as shown above for each channel to obtain the initial completed feature vector after diffusion under each channel According to the node arrangement order of the original feature matrix X, the channel feature vectors are concatenated in channel order to obtain the reconstructed feature matrix 9. A graph feature completion method based on curriculum learning according to claim 8, characterized in that: The step 4 is specifically as follows: Assume that each node has F channels, based on the preliminary completed feature matrix in is the completed feature vector composed of all the eigenvalues ​​of the completed channels of the i-th node. The eigenvalue of the i-th node on the d-th feature channel is The following steps are used to construct a channel correlation graph for node i to propagate the correlation between feature channels; Estimate probability distribution: Use kernel density estimation (KDE) or Gaussian mixture model to estimate the joint probability density function and marginal probability density function of each pair of channels; Calculate mutual information: For a pair of feature channel variables X and Y, the mutual information I(X; Y) can be defined as follows: Wherein p(x, y) is the joint probability distribution between channels, p(x) and p(y) are the marginal probability distributions of channel X and channel Y respectively; according to the joint probability density function and the marginal probability density function, the mutual information I(X; Y) between each pair of channels is calculated according to the mutual information calculation formula; Construct the mutual information matrix: organize the mutual information values ​​between all channels into an F×F mutual information matrix M, where the (i, j) element of the mutual information matrix M is the channel f i With channel f j The mutual information value I(f i ;f j ),Right now Construct a channel correlation graph: The mutual information matrix M is regarded as a weighted fully connected graph, where each node represents a channel and the weight of the edge is the corresponding mutual information value. Construct a channel correlation graph Gf = (V f , E f , W f ),in is the channel set, E f is the fully connected edge between channels, W f is the weighted adjacency matrix of the channel correlation graph, and the weight is the calculated mutual information value 10. A graph feature completion method based on curriculum learning according to claim 9, characterized in that: Based on the constructed channel correlation graph, the completed feature matrix is ​​used to propagate the correlation between channels of each node, and the mean of each channel is used to eliminate the influence of the channel extreme value, so that the features between different channels can be propagated and optimized according to their correlation; For each node i, its feature on the dth channel Update refinement is performed using the following propagation formula: where μ f is the feature mean of channel f.