Positive sample expansion graph comparative learning method based on soft clustering
Through the structure enhancement and fine-grained feature enhancement guided by Jaccard coefficient, combined with fuzzy C-mean clustering to screen high confidence nodes, optimize positive sample quality and quantity, the problems of semantic consistency and insufficient training signals in traditional graph comparison learning are solved, and the performance of the model in node classification and link prediction tasks is improved.
Patent Information
- Application Number
- CN202510417502.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional graph comparison learning models lack semantic consistency in the data augmentation process, stochastic enhancement strategies lead to performance degradation, positive sample division violates the assumption of homogeneity, and training signal noise affects the model effect.
Structural enhancement and fine-grained feature enhancement based on Jaccard coefficients are used, and high confidence nodes are screened in combination with the fuzzy C-means clustering algorithm, positive sample quality and quantity are optimized through multi-task loss function, and edge prediction task supplementary graph topology information is introduced.
Improve the robustness and accuracy of graph representation learning, and enhance the generalization ability of the model in node classification and link prediction tasks.
Smart Images

Figure CN120277411A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields such as graph neural networks, and particularly relates to a positive sample augmentation graph contrast learning method based on soft clustering. Background Art
[0002] The training process of traditional graph contrast learning models can generally be divided into three modules, namely data augmentation, pretext task, and contrast objective.
[0003] The purpose of the data augmentation module is to generate multiple augmented views to provide rich training samples. In this process, it should be ensured that the views of the same sample remain semantically invariant before and after augmentation. However, most graph contrast learning models use random augmentation methods to obtain these augmented views. Although this method has achieved good results to a certain extent, the random augmentation strategy obviously cannot guarantee the semantic invariance of the samples before and after augmentation, which may lead to a decline in model performance, as shown in the attached instructions. Figure 1 As shown. Therefore, if the data augmentation strategy is not reasonable enough, it may generate insufficient augmented views, mislead the encoder training, and even lead to problems such as overconfidence and overfitting.
[0004] In graph neural networks, the homophily hypothesis is an important concept. It means that in a graph, the connection relationship between nodes usually reflects their similarity. In other words, nodes with an edge relationship in the graph usually have the same label. However, in the positive and negative sample division of traditional graph contrast learning, only the same nodes in the contrast views are regarded as positive samples, while the remaining nodes are regarded as negative samples. This way of dividing positive and negative samples obviously violates the homophily hypothesis, generates incorrect training signals, and ultimately deviates from the training goal of maximizing the consistency of positive sample pairs and minimizing the consistency of negative sample pairs. Therefore, finding more positive samples for the anchor node can improve the training effect of the model, as shown in the attached instructions. Figure 2 As shown. Therefore, how to find the correct positive samples becomes a relatively key problem.
[0005] For the contrast objective, most graph contrast learning only relies on the contrast objective to provide training signals for model training. However, due to a certain degree of deviation in the positive and negative sample division strategy in the contrast objective, the training signals provided by the contrast objective may be noisy, thus affecting the training. Summary of the Invention
[0006] Considering that traditional graph contrast learning can generally be divided into three modules: data augmentation, pretext task, and contrast objective, but most methods often adopt random data augmentation strategies, thus unable to effectively guarantee the semantic consistency before and after augmentation, and the single positive sample division method in the pretext task often violates the homophily hypothesis, thereby affecting the training effect and leading to a decline in model performance.
[0007] Aiming at the defects and deficiencies of the existing technologies, the present invention provides a positive sample augmentation graph contrast learning method based on soft clustering, and its core design points include:
[0008] A prior knowledge-driven data augmentation mechanism: quantifying the structural similarity between nodes through the Jaccard coefficient, dynamically guiding the perturbation and supplement of edges, introducing potential similar edges while retaining key topological relationships; adopting an element-level random masking strategy to achieve fine-grained diversity enhancement in the feature dimension and avoid the loss of semantic information caused by dimension-level masking.
[0009] Positive sample screening by collaborating soft clustering and structural constraints: combining the fuzzy C-means (FCM) algorithm to generate a node membership matrix, quantifying the confidence of multi-category membership, and innovatively introducing the membership entropy index to screen high-confidence nodes; based on the homogeneity hypothesis, dynamically augmenting positive samples consistent with the pseudo-labels of target nodes in the first-order neighbors, breaking through the limitations of traditional single positive sample division.
[0010] A multi-dimensional training signal fusion framework: jointly optimizing the contrast loss, membership entropy optimization loss, and edge prediction loss, strengthening the ability to distinguish positive and negative samples through contrast learning, improving the confidence of soft clustering by minimizing the membership entropy, and supplementing the graph topology modeling ability with the edge structure prediction task to form a model training paradigm with multi-angle collaborative supervision.
[0011] Through the collaboration of the above technologies, the present invention effectively solves the problems of inaccurate semantic enhancement, insufficient positive samples, and single supervision signal in traditional methods, and significantly improves the robustness and accuracy of graph representation learning in tasks such as node classification and link prediction.
[0012] The technical solution specifically adopted by the present invention to solve its technical problems is:
[0013] A positive sample augmentation graph contrast learning method based on soft clustering, including:
[0014] Data augmentation: performing structural augmentation and feature augmentation on the original graph to generate enhanced views;
[0015] The structural augmentation guides edge perturbation through the structural similarity between nodes, retains key edges and supplements potential similar edges;
[0016] The feature augmentation generates diverse feature combinations through fine-grained masking;
[0017] Dynamic positive sample augmentation:
[0018] Calculating the membership of nodes to each category based on the fuzzy clustering algorithm, and screening a high-confidence node set;
[0019] Combining the graph structure constraints and the first-order neighbor relationship to augment positive samples for each target node;
[0020] Multi - task joint training: Jointly optimize the contrast loss, clustering uncertainty loss, and edge prediction loss to train the graph neural network model.
[0021] Among them, the fuzzy clustering algorithm preferably adopts fuzzy C - means clustering (FCM).
[0022] Furthermore, the structure enhancement is achieved through the following steps:
[0023] Calculate the structural similarity of node pairs, and use the Jaccard coefficient to measure the overlap degree of the node neighbor sets;
[0024] Dynamically set the edge removal probability according to the structural similarity, and retain the edges with high similarity;
[0025] Perform edge addition operations on node pairs that are not adjacent but have a structural similarity exceeding a preset threshold.
[0026] Furthermore, the screening of the high - confidence node set is achieved through the uncertainty of the node membership distribution; the uncertainty is calculated by membership entropy, and the top K - proportion nodes with the lowest membership entropy values are selected during screening.
[0027] Furthermore, the dynamic augmentation of positive samples needs to meet the following conditions:
[0028] The candidate node is the first - order neighbor of the target node;
[0029] The candidate node belongs to the high - confidence node set;
[0030] The candidate node has the same pseudo - label as the target node, and the pseudo - label is assigned according to the maximum membership degree.
[0031] Furthermore, the loss function of the multi - task joint training includes:
[0032] The contrast loss, calculated based on the augmented positive sample set;
[0033] The membership entropy loss, minimizing the uncertainty of node clustering;
[0034] The edge prediction loss, optimizing the edge prediction task through binary cross - entropy.
[0035] Furthermore, the feature enhancement is achieved through the following steps:
[0036] Element - level random mask generation: Generate a random mask vector with the same dimension as the feature for each node. Each element in the mask vector independently retains the original feature value with probability p or sets it to zero with probability 1 - p;
[0037] Feature mask operation: Multiply the random mask vector and the node features element - by - element to generate fine - grained diverse feature combinations.
[0038] Further, the trigger threshold for the edge addition operation is thr, and an edge is added when the Jaccard coefficient of the node pair exceeds this threshold.
[0039] Further, the edge removal probability is dynamically set according to the Jaccard coefficient of the node pair. The larger the Jaccard coefficient, the lower the edge removal probability, and the negative correlation between the removal probability and the Jaccard coefficient is controlled by a hyperparameter.
[0040] Further, the pseudo-label assignment rule is the category corresponding to the highest membership degree of the node pair clustering center.
[0041] And, an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the steps of the method described above are implemented.
[0042] A non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described above are implemented.
[0043] Compared with the prior art, the present invention and its preferred solutions have the following beneficial effects:
[0044] Improve the semantic consistency of the enhanced view: Through the structure enhancement strategy guided by the Jaccard coefficient, dynamically retain key edges and supplement potentially similar edges, reduce the semantic damage of the graph structure caused by random perturbations; combine element-level feature masks to generate fine-grained and diverse feature combinations, avoid information loss caused by dimension-level masking, and enhance the rationality and effectiveness of view generation.
[0045] Optimize the quality and quantity of positive samples: Quantify the node membership degree distribution based on the fuzzy clustering algorithm, screen the high-confidence node set through membership entropy, and constrain the pseudo-label consistency in the first-order neighbors to achieve the dynamic and reliable expansion of positive samples, and alleviate the problems of insufficient positive samples and interference from noise samples in traditional methods.
[0046] Strengthen multi-angle supervision signals: Optimize the task through membership entropy to reduce clustering uncertainty and improve the reliability of soft clustering results; introduce an edge prediction task to explicitly model the graph structure information, which forms a complement with the contrastive loss and enhances the model's comprehensive understanding ability of the graph topology and node features.
[0047] Improve the model generalization performance: The multi-task joint training framework integrates the collaborative supervision of feature contrast, clustering optimization, and structure prediction, breaks through the limitations of a single contrast signal, and enables the model to show stronger robustness and generalization in tasks such as node classification and link prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments:
[0049] Figure 1 It is a schematic diagram of randomly enhanced destruction of view semantics in the prior art.
[0050] Figure 2 It is a schematic diagram of the comparison relationship between positive and negative sample nodes in the prior art.
[0051] Figure 3 It is a structural diagram of the positive sample expansion model based on soft clustering in an embodiment of the present invention. Specific embodiments
[0052] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0053] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0054] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0055] To solve the problems existing in the prior art, an embodiment of the present invention proposes a positive sample expansion graph contrast learning model based on soft clustering. This model adopts a structure enhancement guided by prior knowledge and a fine-grained feature enhancement strategy to obtain a more reasonable diversity-enhanced view. At the same time, through the soft clustering algorithm and the constraint of first-order neighbors, the number of expanded positive samples is screened from both the feature and structure levels to improve the model training effect. Finally, an additional training task is added based on the contrast objective to supplement multi-angle training information.
[0056] As Figure 3 shown, this model includes the following design processes:
[0057] 1. Designed a structure enhancement strategy based on the Jaccard coefficient and a fine-grained feature enhancement strategy. The importance of edges in the graph is measured by the Jaccard coefficient to guide the edge perturbation process, thereby retaining the important structure of the original view; by reducing the granularity of the feature mask, a more diverse feature composition is generated to improve the diversity of the enhanced view.
[0058] 2. A positive sample augmentation method based on soft clustering is proposed. The Fuzzy C-Means (FCM) algorithm is used to estimate the confidence level of nodes, and combined with the clustering results, augmented positive samples are selected from the first-order neighbors of the nodes to increase the positive sample information and improve the model performance.
[0059] 3. A multi-angle training signal supplementation strategy is proposed. Two additional tasks are added based on the contrast task: the membership entropy optimization task and the edge prediction task. The training signals are supplemented from the clustering perspective and the edge perspective to improve the training quality.
[0060] The following introduces the specific technology of the embodiments of the present invention in conjunction with the accompanying drawings:
[0061] 1. Model overview:
[0062] First, through a data augmentation module based on Jaccard coefficient structure enhancement and fine-grained feature enhancement, an enhanced view is obtained and input into the same encoder for feature encoding to obtain node representations. Then, FCM soft clustering is performed on the original view to obtain a membership matrix. Pseudo-labels are assigned to the nodes according to the membership matrix, and the membership entropy is calculated at the same time. A high-confidence node set is sampled according to the membership entropy. Subsequently, for each node, nodes that belong to the high-confidence node set and have the same pseudo-label are selected from its first-order neighbors and added to its positive sample set. Finally, the contrastive objective is optimized according to the positive and negative sample relationships, and the edge prediction loss and the membership entropy clustering loss are added to supplement additional training signals.
[0063] 2. Design of the data augmentation module:
[0064] Since random data augmentation may cause damage to key structures, resulting in semantic changes before and after augmentation, in this part, the present invention designs an augmentation method based on Jaccard coefficient structure enhancement and fine-grained feature enhancement, making the obtained enhanced view more reasonable and diverse.
[0065] Edge perturbation is selected as the structure enhancement method for structure enhancement. The Jaccard coefficient is used to score each edge on the graph, and the importance of the edge is measured according to the score to retain more important edges. The Jaccard coefficient is used to compare the similarity and difference between finite sample sets. The larger the Jaccard coefficient value, the higher the sample similarity.
[0066] Definition 1 Given two nodes v i , v j , the Jaccard coefficient is defined as the intersection of the neighbor sets of node v i and the neighbor sets of node v j , and the neighbor sets of node v i and the neighbor sets of node v jThe ratio of the union of the neighbor sets is as follows:
[0067]
[0068] where N(v i ) represents the neighbor set of node v i . When the value of J(v i , v j ) is large, it means that the neighbor overlap degree between node v i and node v j is high, indicating that node v i and node v j have relatively similar properties. Therefore, the edge connection relationship between them should be removed with a smaller probability. Conversely, when the value of J(v i , v j ) is small, the edge connection relationship between them should be removed with a larger probability. Therefore, the calculation formula for the edge removal probability is as follows:
[0069]
[0070] where p e is a hyperparameter used to control the edge perturbation probability. In addition to removing the existing edge connection relationships, the present invention also adds edges to nodes with a high Jaccard coefficient but no edge connection between two points to obtain a more reasonable view. The specific method is as follows:
[0071]
[0072] where thr represents the Jaccard coefficient threshold for adding edge connection relationships.
[0073] The feature enhancement selects the feature mask as the feature enhancement method. Most of the feature masks perform the masking operation in units of the dimension of the feature, which cannot provide good diversity. Therefore, in the present invention, the masking granularity is reduced from the dimension to the element to generate a more diverse feature combination, thereby helping the model to better capture local features. The specific method is as follows:
[0074]
[0075] where M i represents the mask vector randomly generated for node v i , and f represents the dimension of the node feature.
[0076] 3. Positive sample augmentation module based on soft clustering:
[0077] Since the traditional contrast mode of single positive sample division will generate incorrect training signals, in order to alleviate this phenomenon, the present invention reasonably expands the positive sample set. In order to select the correct positive samples, the present invention restricts the confidence level of the samples from two perspectives: feature and structure, so as to improve the correctness of the samples.
[0078] From the perspective of features, a clustering method is used to estimate the probability that a sample is a correct positive sample. However, traditional clustering methods (such as K-means) can only assign a hard label to each node, and they have no way to measure the confidence level of the assignment. Therefore, to solve this problem, the present invention uses the FCM algorithm to replace the traditional clustering algorithm. The FCM clustering algorithm is a type of soft clustering method. The clustering result it calculates is not either this or that, but is expressed as the membership degree to each class. The closer a sample is to a certain class, the higher the membership degree to that class, and the sum of the membership degrees of each sample is 1. For a data set with n nodes: X = {x1, x2,..., x n}, it is specified that the data set is divided into k classes. After iterative calculation by the FCM algorithm, the membership degree of each node to each class and the clustering center of each class can be obtained. The following are the objective function and constraints of the FCM algorithm:
[0079]
[0080] Among them, U is the membership degree matrix, C is the clustering center, u ij represents the membership degree of node v i to the clustering center c j , m is the fuzzy index, d ij represents the distance between node v j and the clustering center c j , and generally the Euclidean distance is used. To make the objective function F obtain the minimum value, the Lagrange multiplier method is used for the objective function under the condition of satisfying the constraints, and the membership degree matrix U and the clustering center C are obtained. The calculation formulas are as follows:
[0081]
[0082] In order to evaluate the quality of soft clustering, the present invention utilizes the concept of membership entropy:
[0083]
[0084] The membership entropy EN i represents node v iThe uncertainty of the membership distribution. A node with a high EN value represents that it is close to multiple cluster centroids, indicating that its clustering confidence is low; on the contrary, a node with a low EN value has a higher confidence in the clustering. In addition, the present invention selects the nodes with the lowest membership entropy in a ratio of κ to add to the high-confidence node set O. Secondly, according to the membership matrix U, the present invention also assigns a pseudo label to each node according to its maximum membership, and obtains a pseudo label matrix P.
[0085] From a structural perspective, based on the homogeneity assumption, the present invention chooses to select positive samples only from the first-order neighbors to improve accuracy. i For example, its expanded positive sample set is:
[0086] Pos i = {v j ∈O∩N(v i )|P(v i )=P(v j )}, (11)
[0087] This formula represents the node v i Positive sample set Pos i Any node v in j , three conditions must be met: node v j is a high confidence node, node v j is node v i The first-order neighbors and node v i and node v j have the same pseudo-label.
[0088] 4. Multi-angle training signal supplement module:
[0089] The goal of contrastive learning is to narrow the distance between positive sample pairs while expanding the distance between negative sample pairs. The traditional single positive sample contrast loss is:
[0090]
[0091] After the present invention screens and obtains the expanded positive sample set of each node, the comparison target is:
[0092]
[0093] The overall contrast loss is:
[0094]
[0095] To provide supplementary signals from multiple perspectives to guide the training of the model, the present invention supplements two additional tasks: the membership entropy optimization task and the edge prediction task. The membership entropy is used to measure the uncertainty of node soft clustering. By minimizing the membership entropy loss, the model can improve the soft clustering results of nodes. The edge prediction loss, from the perspective of edge structure, enables the model to mine the structural information in the graph. The corresponding losses are as follows:
[0096]
[0097] where m is the number of sampled edges, σ(·) is the sigmoid function, y i is the true label, and z i is the predicted score. Therefore, the overall loss is:
[0098] L = L con + αL en + βL link , (17)
[0099] where α and β are hyperparameters used to balance different loss terms.
[0100] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically for loading and executing one or more instructions in the computer storage medium to implement the above method.
[0101] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the above method. The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electro-magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0102] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those with ordinary skills in the field to which the present invention belongs. The "first", "second", and similar terms used in the present invention do not represent any order, quantity, or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0103] As described above, these are only the preferred embodiments of the present invention, and the present invention is not limited to other forms. Any person skilled in the relevant art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, as long as it does not depart from the technical content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
[0104] This patent is not limited to the above-mentioned best implementation mode. Anyone inspired by this patent can derive other various forms of a positive sample augmentation graph contrast learning method based on soft clustering. All equivalent changes and modifications made within the scope of the patent application of this invention shall fall within the scope covered by this patent.
Claims
1. A positive sample augmentation graph contrastive learning method based on soft clustering, characterized in that Including: Data augmentation: Perform structural enhancement and feature enhancement on the original graph to generate enhanced views; The structural enhancement guides edge perturbation through the structural similarity between nodes, retains key edges and supplements potential similar edges; The feature enhancement generates diverse feature combinations through fine-grained masks; Positive sample dynamic augmentation: Calculate the membership degree of nodes to each category based on the fuzzy clustering algorithm, and screen the high-confidence node set; Combine graph structure constraints and first-order neighbor relationships to augment positive samples for each target node; Multi-task joint training: Jointly optimize the contrast loss, clustering uncertainty loss and edge prediction loss to train the graph neural network model.
2. The positive sample augmented graph contrast learning method based on soft clustering according to claim 1, wherein: The structural enhancement is achieved through the following steps: Calculate the structural similarity of node pairs, and use the Jaccard coefficient to measure the overlap degree of node neighbor sets; Dynamically set the edge removal probability according to the structural similarity, and retain the high-similarity edges; Perform edge addition operations on node pairs that are not adjacent but have a structural similarity exceeding a preset threshold.
3. A positive sample augmentation graph contrastive learning method based on soft clustering according to claim 1, characterized in that: The screening of the high-confidence node set is achieved through the uncertainty of the node membership distribution; the uncertainty is calculated by membership entropy, and the top K percentage of nodes with the lowest membership entropy value are selected during screening.
4. The positive sample augmented graph contrast learning method based on soft clustering according to claim 1, wherein: The positive sample dynamic augmentation needs to meet the following conditions: The candidate node is a first-order neighbor of the target node; The candidate node belongs to the high-confidence node set; The candidate node has the same pseudo label as the target node, and the pseudo label is assigned according to the maximum membership degree.
5. The positive sample augmented graph contrast learning method based on soft clustering according to claim 1, wherein: The loss function of the multi-task joint training includes: Contrast loss, calculated based on the augmented positive sample set; Membership entropy loss, minimizing the uncertainty of node clustering; Edge prediction loss, optimizing the edge prediction task through binary cross-entropy.
6. The positive sample augmented graph contrast learning method based on soft clustering according to claim 1, wherein: The feature enhancement is achieved through the following steps: Element-level random mask generation: Generate a random mask vector with the same dimension as the feature for each node, and each element in the mask vector independently retains the original feature value with probability p or sets it to zero with probability 1−p; Feature mask operation: Multiply the random mask vector and the node feature element-wise to generate fine-grained diverse feature combinations.
7. A positive sample augmentation graph contrast learning method based on soft clustering according to claim 2, characterized in that: The trigger threshold for the edge addition operation is thr, and an edge is added when the Jaccard coefficient of the node pair exceeds this threshold.
8. A positive sample augmentation graph contrast learning method based on soft clustering according to claim 2, characterized in that: The edge removal probability is dynamically set according to the Jaccard coefficient of the node pair, where the larger the Jaccard coefficient, the lower the edge removal probability, and the negative correlation between the removal probability and the Jaccard coefficient is controlled by hyperparameters.
9. A positive sample augmentation graph contrast learning method based on soft clustering according to claim 4, characterized in that: The pseudo label assignment rule is the category corresponding to the highest membership degree of the node pair to the clustering center.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the method according to any one of claims 1-9 are implemented.
Citation Information
Cited By
Risk control training sample optimization method and device and electronic equipment
CN121117592A
Multi-granularity image contrast learning method and device, equipment, medium and program product
CN122510012A