Graph self-supervised learning method based on probability distribution alignment and information entropy maximization

By adopting the methods of probability distribution alignment and information entropy maximization in graph self-supervised learning, the problems of model collapse and reduced representation diversity are solved, and better graph structure information capture and model performance improvement are achieved.

CN120146140APending Publication Date: 2025-06-13江苏群博智能技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510280334.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In practical applications, graph self-supervised learning methods face problems such as model collapse, reduced representation diversity and insufficient capture of graph structure information.

Method used

A graph self-supervised learning method based on probability distribution alignment and information entropy maximization is adopted to improve the expression ability and information carrying capacity of the model through graph data augmentation processing, cross-view probability distribution alignment and information entropy maximization optimization objective function.

Benefits of technology

It effectively integrates node characteristics and topological dependencies, comprehensively captures the intrinsic relationships between nodes, improves the diversity of representations, avoids model crash problems, and improves the robustness and performance of graph self-supervised learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146140A_ABST
    Figure CN120146140A_ABST
Patent Text Reader

Abstract

The invention discloses a graph self-supervised learning method based on probability distribution alignment and information entropy maximization. The method comprises the following steps: performing graph data enhancement processing on a to-be-processed image under double view angles; performing feature extraction processing on a double-view enhancement processing result; carrying out condition distribution construction based on a node representation layer and a structure layer on a dual-view node representation result of feature extraction; performing cross-view alignment on the double-view multi-layer condition distribution result; setting an information entropy maximization optimization objective function of the graph convolutional neural network; establishing an overall optimization function based on the alignment objective function and the information entropy maximization optimization objective function, and based on the optimization model; according to the method, distribution of the nodes in a topological space and a representation space can be modeled through probability, alignment is carried out based on JS divergence, and the internal relation between the nodes is comprehensively captured; through maximizing the information entropy of the graph neural network parameter matrix, the expression capability and the information bearing capability of the network are significantly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of self-supervised learning, and particularly to a graph self-supervised learning method based on probability distribution alignment and information entropy maximization. Background Art

[0002] Graph Self-Supervised Learning (GSSL), as a powerful technical means, has significant advantages in mining and utilizing large-scale unlabeled graph data; this technology has shown broad application prospects in multiple fields such as recommendation systems, molecular property prediction, and traffic flow prediction; its core lies in constructing auxiliary tasks to automatically mine self-supervised signals from unlabeled data, thereby completing the training of graph models, effectively reducing the dependence on manually labeled data and reducing labor costs.

[0003] Among many graph self-supervised learning methods, the graph self-supervised learning method based on multi-view learning has attracted much attention due to its excellent performance; this type of method generates multiple views of the same graph data through means such as data augmentation, and uses Graph Neural Networks (GNNs) to perform feature extraction and representation learning on these views respectively; in the representation space, the representations of the same node under different views are aligned by minimizing their differences under a specific metric (such as Euclidean distance or cosine similarity); this alignment strategy aims to enhance the consistency of representations under different views, mine shared features among different views, and thus improve the effect of representation learning.

[0004] However, although the graph self-supervised learning method based on multi-view learning has many theoretical advantages, it still faces many challenges in practical applications; among them, the model collapse problem is particularly prominent, which is manifested as the node representations shrinking to a single constant value or a certain low-dimensional subspace of the entire space, resulting in limited expressive ability of the model, significantly reduced representation diversity, and thus generating degenerate solutions.

[0005] To address the model collapse problem, existing research has proposed various strategies; negative sampling and asymmetric architecture design are two commonly used methods; for negative sampling, it increases the learning difficulty of the model by introducing negative samples, thereby improving the distinguishability of representations; for asymmetric architecture design, it changes the model structure, such as using different encoders to process data from different views, to avoid the model falling into trivial solutions; however, these methods still have many limitations in practical applications; the performance of the method based on negative sampling is often positively correlated with the number of negative samples, resulting in a high dependence on computing resources such as video memory; while the strategy based on asymmetric architecture usually has a long convergence time, affecting the training efficiency.

[0006] In addition, the operations of these methods are mostly limited to the feature space and do not explicitly focus on the optimization of the neural network parameter space, while the optimization of the parameter space is also crucial for improving the model performance. Traditional alignment methods, such as directly minimizing the Euclidean distance between the representations of the same node from different perspectives in the space, although achieving strict point-by-point alignment, ignore the topological dependencies between the nodes in the graph and are difficult to comprehensively capture the graph structure information, thus limiting the performance of graph self-supervised learning.

[0007] In summary, graph self-supervised learning has significant advantages in mining and utilizing large-scale unlabeled graph data, but still faces challenges such as model collapse, reduced representation diversity, and insufficient capture of graph structure information in practical applications. Summary of the Invention

[0008] The purpose of the present invention is to provide a graph self-supervised learning method based on probability distribution alignment and information entropy maximization, so as to solve all or one of the above problems existing in the prior art.

[0009] To solve the above technical problems, the specific technical solutions of the present invention are as follows: The present invention provides a graph self-supervised learning method based on probability distribution alignment and information entropy maximization, including the following steps: Graph data augmentation processing: Identify the image to be processed for the graph self-supervised learning task, and perform graph data augmentation processing on the image to be processed from two augmented perspectives to obtain the double-perspective augmentation processing result; Cross-perspective probability distribution alignment: Perform feature extraction processing on the double-perspective augmentation processing result to obtain the double-perspective node representation result; for the double-perspective node representation result, centered on the graph nodes, construct conditional distributions based on the node representation level and the structure level to obtain the double-perspective multi-level conditional distribution result; perform cross-perspective alignment on the double-perspective multi-level conditional distribution result to obtain the alignment objective function; Maximize the information entropy of the model parameters: Set the information entropy maximization optimization objective function for the graph convolutional neural network; Overall optimization: Establish an overall optimization function based on the alignment objective function and the information entropy maximization optimization objective function, and optimize the graph self-supervised learning task of the model based on the overall optimization function.

[0010] As an improved solution, the graph data augmentation processing from the two augmented perspectives both includes: data augmentation processing based on the node feature level and data augmentation processing based on the graph topological structure level; The double-perspective augmentation processing result includes: the first-perspective augmentation processing result and the second-perspective augmentation processing result; The dual-view node representation results include: the first-view node representation results and the second-view node representation results; The dual-view multi-level conditional distribution results include: the first-view multi-level conditional distribution results and the second-view multi-level conditional distribution results; both the first-view multi-level conditional distribution results and the second-view multi-level conditional distribution results are composed of the structural-level conditional distribution and the node representation-level conditional distribution under the corresponding view.

[0011] As an improved solution, the data augmentation processing based on the node feature level includes: Adopt the node feature masking strategy and randomly add Gaussian noise to the node feature matrix of the image to be processed.

[0012] As an improved solution, the data augmentation processing based on the graph topology level includes: Randomly add or remove edges between the nodes of the image to be processed.

[0013] As an improved solution, the feature extraction processing of the dual-view augmentation processing results to obtain the dual-view node representation results includes: Adopt a graph convolutional neural network as the backbone network; Based on the backbone network, perform feature extraction and representation learning on the first-view augmentation processing results to obtain the first-view node representation results; Based on the backbone network, perform feature extraction and representation learning on the second-view augmentation processing results to obtain the second-view node representation results.

[0014] As an improved solution, for the dual-view node representation results, centered on the graph nodes, construct the conditional distribution based on the node representation level and the structural level to obtain the dual-view multi-level conditional distribution results, including: Adopt the global graph diffusion matrix to construct the global topological dependence relationship between the graph nodes, and establish the conditional distribution based on the structural level centered on the graph nodes to obtain the structural-level conditional distribution; Establish the conditional distribution based on the node representation level centered on the graph nodes to obtain the node representation-level conditional distribution.

[0015] As an improved solution, the cross-view alignment of the dual-view multi-level conditional distribution results to obtain the alignment objective function includes: Adopt the JS divergence as the metric criterion, and construct the alignment objective function for the structural-level conditional distribution and the node representation-level conditional distribution under the dual-view based on the first-view multi-level conditional distribution results and the second-view multi-level conditional distribution results: Iteratively optimize the alignment objective function, and achieve cross-perspective alignment between the conditional distribution at the structural level and the conditional distribution at the node representation level based on the optimized alignment objective function.

[0016] As an improved solution, the alignment objective function includes: ; Wherein, refers to the alignment objective function, is the conditional distribution at the structural level under the dual perspectives, is the conditional distribution at the node representation level under the first perspective, is the conditional distribution at the node representation level under the second perspective, i is the row and node number of the normalized diffusion matrix, j is the column of the normalized diffusion matrix, and N is the number of nodes. refers to the alignment of probability distributions with the JS divergence as the measurement criterion.

[0017] As an improved solution, the information entropy maximization optimization objective function for the graph convolutional neural network includes: Normalize the hierarchical parameter matrix of the graph convolutional neural network along the row direction to obtain a normalized parameter matrix; Ignore the constant term in the Gaussian information entropy expression and construct an information entropy maximization constraint objective for the normalized parameter matrix; Integrate the information entropy maximization optimization objective function according to all the layers of the graph convolutional neural network according to the information entropy maximization constraint objective.

[0018] As an improved solution, the overall optimization function includes: ; Wherein, is the overall optimization function, is the alignment objective function, is the information entropy maximization optimization objective function, is the balance factor.

[0019] The beneficial effects of the technical solution of the present invention are: The graph self-supervised learning method based on probability distribution alignment and information entropy maximization described in the present invention can effectively fuse node features and topological dependencies by probabilistically modeling the distributions of nodes in the topological space and the representation space and aligning them based on the JS divergence, comprehensively capturing the internal relationships between nodes; the advantages of the symmetry and boundedness of the JS divergence improve the accuracy of the distribution similarity evaluation, and the introduction of the global diffusion matrix effectively characterizes the long-distance topological dependencies between nodes; by maximizing the information entropy of the graph neural network parameter matrix, the expression ability and information-bearing capacity of the network are significantly enhanced, effectively avoiding the problem of model collapse, improving the diversity of representations, and opening up a new way for the robustness and performance improvement of graph self-supervised learning starting from the model parameter space. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a schematic flowchart of the graph self-supervised learning method based on probability distribution alignment and information entropy maximization described in Embodiment 1 of the present invention; Figure 2 It is a detailed schematic flowchart of the graph self-supervised learning method based on probability distribution alignment and information entropy maximization described in Embodiment 1 of the present invention; Figure 3 It is a schematic diagram of the overall framework of graph self-supervised learning of the graph self-supervised learning method based on probability distribution alignment and information entropy maximization described in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The following will elaborate on the preferred embodiments of the present invention in conjunction with the drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0023] In the description of the present invention, it should be noted that the embodiments described are some embodiments of the present invention, rather than all embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0024] In the description of the present invention, it should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this document are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or equipment.

[0025] In the description of the present invention, it should be noted that: (1) Alignment: It refers to optimizing the objective to make the representations of the same node as close as possible under different perspectives, thereby promoting their consistency in the representation space.

[0026] (2) Backbone network: It is a low-level network used for feature extraction and representation learning.

[0027] (3) Model collapse: It means that during the self-supervised training process, the expressive ability of the model drops significantly, and its output representation collapses to a constant point or a low-dimensional subspace in the representation space.

[0028] (4) Neural information entropy: It refers to the information entropy of the parameter matrix in the neural network, which is used to measure the information-carrying capacity of the network parameters and can be used as a metric for the network's expressive ability.

[0029] (5) Divergence: It is a mathematical tool used to quantify the degree of difference between two probability distributions, providing a quantitative description of the inconsistency between distributions.

[0030] (6) Graph convolutional network: It is a type of neural network model that can directly perform feature extraction and representation learning on graph-structured data. Its core idea is to fuse the information of neighbor nodes into the representation of the target node through the neighborhood aggregation mechanism.

[0031] (7) In information theory, information entropy is an important metric for measuring the amount of information, describing the complexity and diversity of the information source. Specifically, the larger the information entropy, the higher the uncertainty of the system and the richer the potential information. Based on the Gaussian distribution assumption, the entropy of the random variable is expressed as: ; where is the covariance matrix of the variable , and is the determinant of the matrix.

[0032] Embodiment. This embodiment provides a graph self-supervised learning method based on probability distribution alignment and information entropy maximization, as follows Figures 1 to 3 shown, including the following steps: The core function of this method is as follows: Aiming at the problem of insufficient utilization of structural information in view alignment in existing graph self-supervised learning tasks, an innovative method is proposed to perform probability modeling by combining node representations and topological structures, and alignment is achieved by minimizing the JS divergence, effectively capturing the relationship between graph structures and node features, thereby improving the modeling and generalization performance; At the same time, this method introduces the maximization of neural information entropy to solve the problem of model collapse, enhances the information-bearing and expression capabilities, and avoids model collapse. The specific steps of this method are as follows: S100. Graph data augmentation processing: In this step, in response to the graph self-supervised learning task, the image to be processed is confirmed, and graph data augmentation processing is performed on the image to be processed from two augmented perspectives, including: S101. Define the image to be processed as which has nodes, each node contains a -dimensional feature vector, and the feature vectors of all nodes form the node feature matrix ; The adjacency matrix is used to describe the topological connection relationship between nodes, represents the diagonal node degree matrix.

[0033] S102. Define two augmented perspectives, including: perspective A and perspective B.

[0034] S103. Graph data augmentation processing, including: data augmentation at the node feature level and data augmentation at the graph topological structure level, specifically as follows: S1031. For the node feature level: Adopt the node feature masking strategy, randomly add Gaussian noise to the node feature matrix, and thus achieve data augmentation. The formal expression of this process is as follows: ; where is the Gaussian noise matrix, and each element in the Gaussian noise matrix is sampled from a Gaussian distribution with a mean of 0 and a standard deviation of .

[0035] S1032. For the graph topological structure level: Randomly add or remove edges between nodes, and thus achieve data augmentation. The formal expression of this process is as follows: ; where is the position indicator matrix, is a matrix with all element values being 1.

[0036] S104. Finally, in each training step of the graph self-supervised learning task, different augmentation functions are adopted and perform image enhancement processing for enhanced view A and enhanced view B respectively, and finally obtain two views of the original image and .

[0037] S200. Cross-view probability distribution alignment: In this step, based on the graph neural network, feature extraction is performed on the two enhanced views respectively to generate the corresponding node representations; for the node representations of the two enhanced views, with each node as the center, conditional distribution construction at the node representation level and conditional distribution construction at the structural level based on the global graph diffusion matrix are carried out; then, based on the JS divergence, cross-view alignment is performed on the two types of conditional distributions centered on each node, and further the internal correlation between the node representation and the structure is mined, including: S201. Use a graph convolutional neural network with parameter sharing as the backbone network to perform feature extraction and representation learning on the two views and respectively, as follows: S2011. The th layer of the backbone network is expressed as: ; where, refers to the adjacency matrix with self-loops, refers to the diagonal matrix of the matrix , refers to the learnable parameter matrix, is instantiated as the node feature matrix .

[0038] S2012. Finally, the node representations under the two views are respectively expressed as: and .

[0039] S202. Use the global graph diffusion matrix to construct the global topological dependence relationship between nodes, and then comprehensively capture the potential long-distance correlation information between nodes in the graph, and finally obtain the conditional distribution at the structural level centered on nodes, as follows: S2021. Use personalized PageRank diffusion to construct the diffusion matrix, expressed as: ; where, refers to the diffusion matrix, refers to the restart probability in the random walk.

[0040] S2022. Normalize each row of the diffusion matrix so that the sum of each row of the diffusion matrix is 1, and obtain the normalized diffusion matrix ; where, each row of instantiates a conditional distribution, and describes a conditional distribution centered on a certain node The probability distribution of each node centered on .

[0041] S2023. Based on , obtain the conditional distribution at the structural level centered on the node which is expressed as: ; where is the element in the i-th row and j-th column of

[0042] S203. Centered on the node , construct its conditional distribution at the representational level, specifically as follows: S2031. Based on the affinity between node representations, across perspectives, construct the conditional distributions and and

[0043] S2032. ; where , and respectively refer to node representations, and t is a hyperparameter that controls the shape of the distribution; the above conditional distributions reflect the feature similarity under cross-perspective conditions by measuring the relative distances of nodes in the feature space.

[0044] S204. Use the JS divergence as the metric criterion to align the conditional distributions at the structural level and the conditional distributions at the node representational level under the dual perspectives, and then deeply explore the internal relationship between the structure and the representation, specifically as follows: S2041. Construct the alignment objective function for the conditional distributions at the structural level and the conditional distributions at the node representational level under the dual perspectives: .

[0045] S2042. Continuously optimize the above objective function to achieve the cross-perspective alignment of the conditional distribution at the structural level and the conditional distribution at the node representational level.

[0046] It should be noted that through the above steps, the internal relationship between features and structures is effectively revealed, thereby enhancing the comprehensive representation ability of multivariate information in graph data. Additionally, in this embodiment, the JS divergence is selected as the main criterion for measuring the distribution difference during the conditional distribution alignment process. Optionally, in other cases, other metrics can be considered as the measurement standard for distribution alignment according to specific task requirements and data characteristics (such as the total variation distance or the Wasserstein distance, etc.). Moreover, different metrics have unique advantages in different application scenarios (such as stronger robustness or more intuitive geometric interpretations, thus providing more flexibility and adaptability for distribution alignment).

[0047] S300, Maximize the information entropy of model parameters: In this step, an optimization objective of maximizing the information entropy of the graph convolutional neural network is set. Based on this optimization objective, during the model training phase, the parameter matrices in the neural network are extracted layer by layer, the corresponding information entropy is calculated, and the information entropy of all layers is accumulated and optimized, thereby improving the expression ability of the model, including: S301, For the parameter matrix of the th layer of the graph convolutional neural network (i.e., the hierarchical parameter matrix), perform normalization processing on it along the row direction so that each column of the parameter matrix follows a distribution with a mean of 0 and a standard deviation of 1, obtaining the normalized matrix .

[0048] S302, Ignore the constant term in the Gaussian information entropy expression, and construct the information entropy maximization constraint objective for the parameter matrix of the th layer of the graph convolutional neural network as follows: .

[0049] S303, Based on the information entropy maximization constraint objective, considering all layers comprehensively, obtain the final information entropy maximization optimization objective as follows: ; where represents the total number of layers of the neural network.

[0050] It should be noted that through the above steps, the information entropy of the parameter matrices of all layers of the graph convolutional neural network is maximized, thereby enhancing the information-bearing capacity of the graph convolutional neural network, improving the expression ability and feature learning effect of the model, and effectively preventing the model from collapsing.

[0051] S400, Overall optimization: In this step, based on the probability distribution alignment objective function of representation and structure and the information entropy maximization optimization objective, construct the overall objective function as: ; where represents a balance factor; finally, optimize the graph self-supervised learning task of the model based on this overall objective function.

[0052] It should be noted that the present invention maximizes the information entropy of the neural network parameters to improve the information capacity and expression ability of the model, thereby enhancing the model's ability to characterize complex data patterns; optionally, the channels of the parameter matrix can be decoupled to reduce the correlation between channels, thereby reducing information redundancy, and also improving the information capacity and expression ability of the model.

[0053] It should be noted that the above examples are only for explaining the present invention and cannot limit the protection scope of the present invention.

[0054] Different from the prior art, the present application adopts a graph self-supervised learning method based on probability distribution alignment and information entropy maximization, which can effectively integrate node features and topological dependencies and fully capture the intrinsic relationship between nodes through the distribution of probabilistic modeling nodes in topological space and representation space, and align based on JS divergence; the accuracy of distribution similarity evaluation is improved based on the symmetry and boundedness advantages of JS divergence, and the introduction of the global diffusion matrix effectively characterizes the long-distance topological dependency between nodes; by maximizing the information entropy of the graph neural network parameter matrix, the network's expressive power and information carrying capacity are significantly enhanced, the model collapse problem is effectively avoided, the diversity of representation is improved, and starting from the model parameter space, a new path is opened up for the robustness and performance improvement of graph self-supervised learning.

[0055] It should be understood that in the various embodiments of this document, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this document.

[0056] It should also be understood that in the embodiments of this article, the term "and / or" is only a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0057] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this article.

[0058] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0059] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connection.

[0060] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the solutions of the embodiments herein.

[0061] In addition, in each of the embodiments herein, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0062] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution herein, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each of the embodiments herein. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0063] The above are only embodiments of the present invention, and do not thus limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present invention.

Claims

1. A graph self-supervised learning method based on probability distribution alignment and information entropy maximization, characterized in that: The following steps are involved: Graph data enhancement processing: confirming the image to be processed for the graph self-supervised learning task, performing graph data enhancement processing under two enhanced perspectives on the image to be processed, and obtaining a dual-perspective enhancement processing result; Cross-view probability distribution alignment: performing feature extraction processing on the dual-view enhancement processing result to obtain a dual-view node representation result; For the dual-perspective node representation results, taking the graph node as the center, conditional distribution construction based on the node representation level and the structure level is performed to obtain a dual-perspective multi-level conditional distribution result; cross-perspective alignment is performed on the dual-perspective multi-level conditional distribution result to obtain an alignment objective function; Maximizing model parameter information entropy: setting an information entropy maximization optimization objective function for the graph convolutional neural network; Overall optimization: establishing an overall optimization function based on the alignment objective function and the information entropy maximization optimization objective function, and optimizing the graph self-supervised learning task of the model based on the overall optimization function.

2. The graph self-supervised learning method based on probability distribution alignment and information entropy maximization according to claim 1 is characterized in that: The graph data enhancement processing under the two enhancement perspectives includes: data enhancement processing based on the node feature level and data enhancement processing based on the graph topology structure level; The dual-viewing angle enhancement processing result includes: a first-viewing angle enhancement processing result and a second-viewing angle enhancement processing result; The dual-view node characterization results include: a first-view node characterization result and a second-view node characterization result; The dual-perspective multi-level conditional distribution results include: a first-perspective multi-level conditional distribution result and a second-perspective multi-level conditional distribution result; the first-perspective multi-level conditional distribution result and the second-perspective multi-level conditional distribution result are both composed of the structural level conditional distribution and the node representation level conditional distribution under the corresponding perspectives.

3. The graph self-supervised learning method based on probability distribution alignment and information entropy maximization according to claim 2 is characterized in that: The data enhancement processing based on the node feature level includes: A node feature mask strategy is adopted to randomly add Gaussian noise to the node feature matrix of the image to be processed.

4. The graph self-supervised learning method based on probability distribution alignment and information entropy maximization according to claim 2 is characterized in that: The data enhancement processing based on the graph topology structure level includes: Edges are randomly added or removed between nodes of the image to be processed.

5. The graph self-supervised learning method based on probability distribution alignment and information entropy maximization according to claim 2 is characterized in that: The step of performing feature extraction processing on the dual-view enhancement processing result to obtain a dual-view node representation result includes: Use graph convolutional neural network as the backbone network; Performing feature extraction and characterization learning on the first perspective enhancement processing result based on the backbone network to obtain the first perspective node characterization result; Based on the backbone network, feature extraction and characterization learning are performed on the second perspective enhancement processing result to obtain the second perspective node characterization result.

6. The graph self-supervised learning method based on probability distribution alignment and information entropy maximization according to claim 2 is characterized in that: The dual-perspective node representation result is constructed based on the node representation level and the structure level conditional distribution with the graph node as the center, and the dual-perspective multi-level conditional distribution result is obtained, including: A global graph diffusion matrix is ​​used to construct a global topological dependency relationship between the graph nodes, and a conditional distribution based on the structure level is established with the graph nodes as the center to obtain the conditional distribution at the structure level; A conditional distribution based on the node representation level is established with the graph node as the center to obtain the conditional distribution at the node representation level.

7. The graph self-supervised learning method based on probability distribution alignment and information entropy maximization according to claim 2 is characterized in that: The cross-perspective alignment of the dual-perspective multi-faceted conditional distribution results to obtain an alignment objective function includes: Using JS divergence as the measurement criterion, an alignment objective function of the conditional distribution at the structure level and the conditional distribution at the node representation level under dual perspectives is constructed based on the multi-level conditional distribution results of the first perspective and the multi-level conditional distribution results of the second perspective: The alignment objective function is iteratively optimized, and based on the optimized alignment objective function, cross-perspective alignment of the structure-level conditional distribution and the node representation-level conditional distribution is achieved.

8. The graph self-supervised learning method based on probability distribution alignment and information entropy maximization according to claim 2 is characterized in that: The alignment objective function includes: ; in, refers to the alignment objective function, is the structural level conditional distribution under dual perspectives, is the conditional distribution of the node representation level under the first perspective, is the conditional distribution of the node representation level under the second perspective, i is the row of the normalized diffusion matrix and the node number, j is the column of the normalized diffusion matrix, N is the number of nodes, Refers to the probability distribution alignment using JS divergence as the metric.

9. The graph self-supervised learning method based on probability distribution alignment and information entropy maximization according to claim 2, characterized in that: The step of setting an information entropy maximization optimization objective function for the graph convolutional neural network includes: Normalize the layer parameter matrix of the graph convolutional neural network along the row direction to obtain a normalized parameter matrix; Ignoring the constant term of the Gaussian information entropy expression, constructing an information entropy maximization constraint objective about the normalized parameter matrix; According to all the layers of the graph convolutional neural network, the information entropy maximization optimization objective function is integrated according to the information entropy maximization constraint objective.

10. The graph self-supervised learning method based on probability distribution alignment and information entropy maximization according to claim 1, characterized in that: The overall optimization function comprises: ; in, is the overall optimization function, is the alignment objective function, The objective function is optimized for maximizing the information entropy, is the balancing factor.