A text network graph classification method based on multi-modal contrastive learning
Through the multimodal comparison learning method, the topological structure and node text of graph data are encoded and fused respectively, solving the problem of failing to effectively utilize multimodal features of graph data in the existing technology, and achieving a more efficient graph-level classification effect.
Patent Information
- Application Number
- CN202211065236.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-09-01
AI Technical Summary
The existing graph data classification methods fail to effectively mine the topology structure and node information of the graph, and comparison learning fails to simultaneously utilize multimodal features in graph data analysis.
The multimodal contrast learning method is used to extract and preprocess the topological structure and node text in the graph data respectively. The graph neural network and large-scale pre-trained language model are used as encoders. The encoder is trained through the comparative learning framework, the feature cross-pooling matrix is calculated and the attention fusion is performed to obtain the cross-modal common feature vector, and the final input classifier is used for graph-level classification.
It improves the performance of graph-level classification tasks, reduces dependence on label data, enhances feature representation ability, improves classification accuracy and recall, and is suitable for multi-modal and small-sample graph data classification tasks.
Smart Images

Figure CN115526236B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-modal network graph data classification, and particularly relates to a network graph classification algorithm based on multi-modal contrast learning. Background Art
[0002] A graph is a data structure for modeling nodes and relationships. Network data based on nodes and relationships is ubiquitous in human society. Therefore, graph data can be used to model and define data in a large number of different fields, and thus has a powerful ability to abstract and represent real-world data. Social network data in social science, physical systems and protein structures in natural science, and knowledge graphs, etc. can all be modeled and represented using graph data. The ubiquity of graph data has attracted wide attention and application in the field of machine learning. As a unique non-Euclidean data structure, the analysis of graph data focuses on tasks such as graph classification, node classification, link prediction, and clustering. Currently, in the field of deep learning, graph neural networks (GNNs) have been proposed and widely applied and developed. GNNs summarize and aggregate node information in graph data based on the topological structure of the graph through a message passing mechanism, and can capture graph structure information at different depths through different levels of aggregation. Due to its convincing performance, GNNs have recently become a widely used graph analysis method.
[0003] Contrastive learning is an effective self-supervised model training paradigm. Due to its lack of dependence on training sample labels and good generalization ability, it has received attention in multiple fields of deep learning. When applying contrastive learning to graph data analysis in the past, the common approach was to regard the graph data as a whole, and the text or image content contained in the nodes as part of the graph node information, and perform contrastive learning on the whole graph data as one modality, failing to simultaneously mine the features of the graph in terms of topological structure and node information. Summary of the Invention
[0004] The object of the present invention is to overcome the existing deficiencies and provide a network graph classification algorithm based on multi-modal contrast learning.
[0005] To achieve the object of the present invention, the following technical solutions are provided:
[0006] A text network graph classification method based on multi-modal contrast learning, the steps of which are as follows:
[0007] S1: For each text network graph data in the multi-modal network graph data set to be classified, extract the two modalities of data, namely the topological structure in the graph and the text in the nodes respectively. The extracted data is classified by modality and saved in dictionary format; then preprocess each modality of data to meet the input requirements of the corresponding modality encoder;
[0008] S2: Select matching encoders for the topological structure modality and the text modality respectively and train them separately using a contrastive learning framework; based on the trained encoders, perform feature encoding on each modality data preprocessed in S1 to obtain the feature vectors of each modality data in each text network graph data, so as to obtain the feature representations of the text network graph data in different modalities;
[0009] S3: For each text network graph data, after aligning the feature vectors of the corresponding two modality data to the same dimension, calculate the Cartesian product of the two to obtain a feature cross matrix, perform horizontal max pooling on the feature cross matrix to obtain a first feature vector, perform vertical max pooling on the feature cross matrix to obtain a second feature vector, and splice the first feature vector and the second feature vector and then reduce the dimension back to the same dimension, so as to obtain a cross-modal common feature vector;
[0010] S4: For each text network graph data, standardize the feature vectors of the two modality data and the cross-modal common feature vector, then use the attention mechanism to calculate the attention weights for the three feature vectors, and perform weighted fusion on the three feature vectors according to the attention weights to obtain the final graph-level feature, and then input it into the classifier to obtain the classification label of each text network graph data in the multi-modal network graph dataset.
[0011] Preferably, in the step S1, for the multi-modal network graph dataset, each text network graph data therein is extracted, saved and preprocessed according to S11 to S14:
[0012] S11: For each text network graph data G i Assign a unique identifier, i = 1, 2, …, N, where N is the scale of the multi-modal network graph dataset; establish a graph data dictionary for storing different modality data in each text network graph data;
[0013] S12: Number the nodes included in each text network graph data, and store the adjacency relationship of each pair of nodes in the graph in the form of an ordered pair list according to the relationship information in the text network graph data, so as to extract the topological structure modality data of the text network graph data, and store it in the corresponding graph data dictionary according to the unique identifier of the text network graph data in S11;
[0014] S13: For each text network graph data, according to the numbering of the nodes when extracting the topological structure information, extract the content text data in the nodes in sequence, and store it in the corresponding graph data dictionary according to the unique identifier of the text network graph data in S11;
[0015] S14: Preprocess different modal data in the graph data dictionary respectively to form structured data adapted to the required input encoder; where:
[0016] For topological structure modal data, it is necessary to define Graph class data, and store the node labels and the node adjacency relationships saved in the form of an ordered pair list respectively;
[0017] For text modal data, first disassemble and segment the text sequence and standardize it, and map the word and character encoding to a numerical value according to the vocabulary, so as to process the text sequence into a numerical vector.
[0018] Preferably, for text modal data, the Tokenize tool function is used to disassemble and segment the text sequence and standardize it.
[0019] Preferably, the specific method of step S2 is as follows:
[0020] S21: Select the graph neural network GCN as the encoder for the topological structure modality, and select the text pre-training model BERT as the encoder for the text modality;
[0021] S22: Set up a contrastive learning framework for different modalities respectively. For topological structure modal data, the SimGRACE contrastive learning framework is adopted, while for text modal data, the SimCSE contrastive learning framework is adopted; for the encoder of each data modality, after constructing positive and negative samples for the training data in batches and inputting them into the encoder, calculate the contrastive learning loss according to the corresponding contrastive learning framework, and update the model parameters of the encoder through the backward propagation of the neural network until all the training data participates in the training, which is regarded as completing one epoch; according to the situation of the contrastive learning loss decline, set an early stopping strategy. After completing the specified number of epoch training times, obtain the encoder trained based on contrastive learning.
[0022] S23: Based on the encoder corresponding to each data modality after training, encode the structured data of the corresponding modality preprocessed in S1 to obtain the feature vectors of the corresponding modality data; each text network graph data respectively obtains the feature vectors of the topological structure modal data and the text modal data.
[0023] Preferably, in the SimGRACE contrastive learning framework, two graph encoders are used to encode the same data during the training process; during the training process with batch as the data input, the feature vectors obtained by encoding the same graph data twice in each batch are used as positive samples, and the feature vectors obtained by encoding other graph data within the batch are used as negative samples; for the two encoders used in the training, a base encoder needs to be initialized first, and on the basis of copying the parameters of the base encoder, random perturbations based on the original parameter Gaussian distribution are added to obtain the parameters of the other encoder.
[0024] Preferably, in the SimCSE contrastive learning framework, the same sample is input into the encoder twice to obtain the positive samples for contrastive learning.
[0025] Preferably, the specific method of step S3 is as follows:
[0026] S31: For the feature vectors of the topological structure modality data and the text modality data in each text network graph data, align the dimensions of the two feature vectors to a unified dimension;
[0027] S32: Perform the Cartesian vector product on the two aligned feature vectors to obtain the feature cross matrix M; then perform max pooling on the rows of the feature cross matrix M to obtain the first feature vector, and perform max pooling on the columns of the feature cross matrix M to obtain the second feature vector, so as to extract all the information that is important in both modalities;
[0028] S33: Concatenate the first feature vector and the second feature vector obtained by the two max poolings and use linear mapping to reduce the dimension to a unified dimension to obtain the cross-modal common feature vector.
[0029] Preferably, the unified dimension is set to 64, 128, or 768.
[0030] Preferably, the specific method of step S4 is as follows:
[0031] S41: For each text network graph data, standardize the feature vectors of the two modalities and the cross-modal common feature vector, and then input them into the attention mechanism together. The attention mechanism calculates the weights of the three feature vectors to obtain the attention weights of the three feature vectors;
[0032] S42: For each text network graph data, perform weighted fusion on the three feature vectors according to the attention weights calculated in S41, and use the weighted fusion vector as the final graph-level feature representation;
[0033] S43: For each text network graph data in the multi-modal network graph dataset, input the corresponding graph-level feature representation into a linear classifier to obtain the corresponding graph classification result.
[0034] Preferably, the text network graph data is rumor propagation tree data. Each piece of rumor propagation tree data includes a seed node and interaction nodes. The seed node is the original information, and the interaction nodes are the forwards and comments based on the original information. Each node contains text content related to the original information; the classification label corresponding to the rumor propagation tree data is the label indicating whether the original information is a rumor.
[0035] The beneficial effects of the present invention compared with the prior art:
[0036] For the scenario of multi-modal network graph data classification, the present invention fully considers the information of the topological structure of the graph data itself and the node feature information, and adopts a contrastive learning training method. Based on the task of individual discrimination, a contrastive learning loss is designed and constructed, reducing the model's dependence on labeled data. At the same time, the present invention innovatively proposes to extract cross-modal common features by taking the Cartesian product of the feature representations of different modalities, enhancing the graph-level feature representation and effectively improving the model performance on the graph-level classification task. In addition, based on the contrastive learning module in the present invention, the encoder can effectively learn the internal features of a large amount of unlabeled data; combined with the specific graph classification task, by fine-tuning the encoder with a small amount of labeled samples in a task-oriented manner, good classification performance can be achieved. The present invention has improvements in the common graph classification evaluation indicators such as Accuracy, Precision, Recall, and F1 Score, and is simple and easy to operate, with a flexible model framework. The present invention can provide demonstration and reference for other graph data classification tasks with multi-modal features and few-shot graph data classification tasks. Description of the Drawings
[0037] Figure 1 It is a flowchart of the text network graph classification method based on multi-modal contrastive learning;
[0038] Figure 2 It is a framework diagram of the text network graph classification method based on multi-modal contrastive learning containing text content according to the embodiment. Detailed Embodiments
[0039] The present invention will be further described and explained below with reference to the drawings and specific embodiments.
[0040] As Figure 1 shown, it is a flowchart of a network graph classification algorithm based on multi-modal contrastive learning provided in a preferred embodiment of the present invention. Its main steps include 4 steps, namely S1 to S4:
[0041] S1: For each text network graph data in the multi-modal network graph data set to be classified, extract the two types of modal data, namely the topological structure in the graph and the text in the nodes, respectively. The extracted data is classified by modality and saved in dictionary format; then, preprocess each type of modal data to meet the input requirements of the corresponding modal encoder.
[0042] S2: Select matching encoders for the topological structure modality and the text modality respectively, and train them using the contrastive learning framework respectively; based on the trained encoders, perform feature encoding on each type of modal data preprocessed in S1 to obtain the feature vectors of each type of modal data in each text network graph data, so as to obtain the feature representations of the text network graph data in different modalities.
[0043] S3: For each text network graph data, after aligning the feature vectors of the corresponding two types of modal data to the same dimension, calculate the Cartesian product of the two to obtain a feature cross matrix. Perform horizontal max pooling on the feature cross matrix to obtain the first feature vector, and perform vertical max pooling on the feature cross matrix to obtain the second feature vector. Concatenate the first feature vector and the second feature vector and then reduce the dimension back to the same dimension, so as to obtain a cross-modal common feature vector.
[0044] S4: For each text network graph data, standardize the feature vectors of the two types of modal data and the cross-modal common feature vector, then use the attention mechanism to calculate the attention weights for the three feature vectors, and perform weighted fusion on the three feature vectors according to the attention weights to obtain the final graph-level feature, and then input it into the classifier to obtain the classification label of each text network graph data in the multi-modal network graph data set.
[0045] The specific implementation methods and their effects of S1 to S4 in this embodiment are described in detail below.
[0046] In the present invention, the specific implementation method of step S1 is as follows:
[0047] For the multi-modal network graph data set, extract, save, and preprocess each text network graph data according to S11 to S14:
[0048] S11: For each text network graph data, assign a unique identifier in combination with the actual background meaning or order of the graph data, and denote each text network graph data as i = 1, 2,..., N, where N is the scale of the multi-modal network graph data set, that is, the total number of text network graph data in the data set. Establish a graph data dictionary for storing different modal data in each text network graph data, and each graph data dictionary is associated with the aforementioned unique identifier.
[0049] S12: Extract the topological structure modal information from the graph data by combining the relationship information in the original multi-modal network graph data. The specific approach is as follows: Label each node included in the text network graph data, and store the adjacency relationship of each pair of nodes in the graph in the form of an ordered pair list according to the relationship information in the text network graph data, so as to extract the topological structure modal data of the text network graph data, and store the unique identifier of the text network graph data in the corresponding graph data dictionary according to S11. At this time, the text network graph data can be represented as G i ={T:t i}, i = 1, 2, …, N, where T is the key name representing the topological structure, and t i represents the topological structure information of graph G i , specifically an ordered pair list representing the adjacency relationship.
[0050] S13: Extract the text modal information from the graph data by combining the node information in the original multi-modal network graph data. The specific approach is as follows: For each text network graph data, extract the content text data in the nodes in sequence according to the labels of the nodes when extracting the topological structure information, and store the unique identifier of the text network graph data in the corresponding graph data dictionary according to S11. At this time, the text network graph data can be represented as G i ={T:t i , D:d i}, i = 1, 2, …, N, where D is the key name representing the text information, and d i represents the text content of graph G i .
[0051] S14: Perform preprocessing on different modal data in the graph data dictionary respectively to make it form structured data adapted to the required input encoder; among them:
[0052] For the topological structure modal data, it is necessary to define Graph class data, and store the node labels and the node adjacency relationships saved in the form of an ordered pair list respectively. In this embodiment, the Python toolkit PytorchGeometric (hereinafter referred to as PyG) can be used to transpose the nodes and the node adjacency relationships stored in the form of an ordered pair list and store them as the Graph class data predefined by PyG.
[0053] For the text modal data, first disassemble, segment, and standardize the text sequence, and map the word and character encoding to a numerical value according to the vocabulary, so as to process the text sequence into a numerical vector. In this embodiment, the Tokenize tool function in the Transformer toolkit released by the HuggingFace open source community can be used to implement the disassembling and standardizing of the text, and map the words and characters to numerical values according to the vocabulary, so as to process the text sequence into a numerical vector.
[0054] If the preprocessing of different modality data in the graph data dictionary is denoted as a function Then the text network graph data can be represented as Denotes the graph nodes and edges stored in the Graph class data, Denotes the graph G i The numerical vector of character numbers after text preprocessing of And are the corresponding keys respectively.
[0055] In the present invention, the specific implementation method of the above step S2 is as follows:
[0056] S21: Select appropriate encoders for different modality data. In this example, the graph neural network ResGCN is used as the encoder for the graph topology modality data, and large-scale pre-trained language models such as BERT are used as the encoders for the text modality data. ResGCN is a graph neural network model based on GCN with residual connections added, and has stronger encoding ability compared to GCN and others. BERT is a large-scale pre-trained language model proposed by Google in 2018, and has been widely used due to its performance in representing text semantics and compatibility with multiple downstream task scenarios.
[0057] S22: Construct positive and negative samples for different modality and design a specific contrastive learning framework. For the graph topology, the SimGRACE contrastive learning framework is adopted, and two graph encoders are used to encode the same data during the training process. During the training process with data input in batches, the feature vectors obtained by encoding the same graph data twice in each batch are used as positive samples, and the features obtained by encoding other data in the batch are used as negative samples. For the two encoders used in the training, a base encoder needs to be initialized first, and at the same time, based on the copy of the base encoder parameters, random perturbations based on the original parameter Gaussian distribution are added to obtain the parameters of the other encoder. For the text modality data, the SimCSE contrastive learning framework is adopted. Since the Dropout mechanism in the text encoder BERT is random, there is no need to set two encoders, and only by inputting the same sample into the encoder twice can two text feature representations that are positive samples of each other be obtained.
[0058] Based on the above steps, calculate the similarity of the positive and negative samples obtained within the same batch respectively, and calculate the loss of contrastive learning. In this example, the cosine similarity calculation method is used for vector similarity calculation. Based on the above framework, the contrastive learning loss function can be summarized as:
[0059]
[0060] Where, l iDenote the loss of the $i$-th sample. Denote two graph data feature representations obtained by encoding the same data with two different encoders, that is, two positive samples; $\tau$ represents the temperature parameter that controls the magnitude of the contrast loss. In the example, this temperature parameter is a hyperparameter and needs to be adjusted by grid search in combination with the data performance.
[0061] Construct positive and negative samples from the training data in batches and input them into the encoder. Calculate the loss of the data within the batch based on the defined contrast learning loss and update the model parameters through the backward propagation of the neural network until all data participate in the training, which is regarded as completing one epoch. According to the decrease of the contrast learning loss and referring to the contrast learning pre-training parameter settings, set the maximum number of training epochs to 200 to obtain the encoder trained based on contrast learning.
[0062] S23: Based on the encoder corresponding to each data modality after training, encode the preprocessed structured data of the corresponding modality in S1 to obtain the feature vectors of the corresponding modality data. Each text network graph data respectively obtains the feature vectors of the topological structure modality data and the feature vectors of the text modality data
[0063] In the present invention, the specific implementation method of the above step S3 is as follows:
[0064] S31: For the feature vectors of the topological structure modality data and the feature vectors of the text modality data in each text network graph data, check whether the dimensions of the feature vectors of different modalities are aligned to a unified dimension. If not, the dimensions of the two feature vectors need to be aligned to a unified dimension. In this embodiment, the feature vectors of the topological structure modality data and the feature vectors of the text modality data are both set to a unified dimension of 768, that is, the feature representation dimension is unified as $d = 768$.
[0065] S32: Perform the Cartesian vector product on the two aligned feature vectors to obtain the feature cross matrix $M$; then perform max pooling on the rows of the feature cross matrix $M$ to obtain the first feature vector, and perform max pooling on the columns of the feature cross matrix $M$ to obtain the second feature vector. In this embodiment, denote the feature cross matrix as $M$ 768*768 , for the matrix $M$ 768*768 Perform max pooling on the rows respectively, and then perform max pooling on the columns to obtain two vectors with a length of 768 dimensions, so as to extract all the information that is important in both modalities.
[0066] S33: Concatenate the first feature vector and the second feature vector obtained by the two max poolings and use linear mapping to reduce the dimension to a unified dimension to obtain the cross-modal common feature vector. In this embodiment, denote the cross-modal feature representation as is a 768-dimensional vector.
[0067] In the present invention, the specific implementation method of the above step S4 is as follows:
[0068] S41: For each text network graph data, after performing 0-1 standardization on the feature vectors of the two modalities and the cross-modal common feature vectors respectively, merge them to obtain Input h i into the attention mechanism, and calculate the weights of the three feature vectors through the attention mechanism to obtain the attention weights of the three feature vectors, denoted as
[0069] S42: For each text network graph data, perform weighted fusion on the three feature vectors according to the attention weights calculated in S41, and use the weighted fusion vector as the final graph-level feature representation, denoted as
[0070] S43: For each text network graph data in the multi-modal network graph dataset, input the corresponding graph-level feature representation into the linear classifier to obtain the corresponding graph classification result.
[0071] Next, based on the above-described embodiment method, it is applied to a specific example to demonstrate its effect. In this embodiment, the text network graph data it targets is rumor propagation tree data. Each piece of rumor propagation tree data includes a seed node and interaction nodes. The seed node is the original information, and the interaction nodes are the forwards and comments based on this original information. Each node contains text content related to the original information. Therefore, in this embodiment, essentially a rumor recognition method based on multi-modal contrast learning is provided, and its final recognition result is the classification label corresponding to the rumor propagation tree data, that is, the label indicating whether the original information is a rumor. The specific process of this method is as described above. The difference is that the input data and output labels are specified, so it will not be fully elaborated here. Below, mainly its specific parameter settings and implementation effects are shown.
[0072] Embodiment
[0073] Next, taking the publicly available Weibo rumor dataset as an example, the present invention will be specifically described. The specific steps are as follows:
[0074] 1) The publicly available Weibo rumor dataset, Weibo Dataset, is adopted, and Python is used to conduct preliminary analysis and cleaning of the data. This dataset is collected based on the rumor information reported by the Sina Community Management Center, with a total of 2,313 rumor microblogs and 2,351 non-rumor microblogs collected, and the forwarding microblog information of the corresponding microblogs is also collected. By interpreting the data structure of the original data, 4,664 rumor propagation tree data are parsed. Each piece of rumor propagation tree data contains the original microblog information as the seed node; at the same time, it contains the interaction information for the original microblog information, specifically including the forwarding of this microblog and interaction nodes such as secondary forwarding and comments based on the first-level forwarding; at the same time, each node contains the text content related to the original microblog information.
[0075] 2) According to the aforementioned step S1, for the 4,664 microblog propagation tree structure data in the publicly available Weibo rumor dataset Weibo Dataset, a unique identifier is defined according to the rumor event ID, and each graph data is denoted as G i , i = 1, 2, …, 4,664. Combining the forwarding information in the rumor propagation tree, each pair of adjacency relationships in the graph is stored in the form of an ordered pair list, so as to extract the topological structure of the graph data; at the same time, using the node information in the rumor propagation tree data, the text content in the node data is extracted. For the extracted topological structure data, the Python toolkit PytorchGeometric (hereinafter referred to as PyG) is used to transpose the nodes and the node adjacency relationships stored in the ordered pair list and store them as the Graph class data predefined by PyG; for the extracted text data, based on the Tokenize tool function in the Transformer toolkit released by the HuggingFace open source community, the text is disassembled and standardized, and the words and characters are mapped to numerical values according to the vocabulary table, so as to process the text sequence into a numerical vector; the preprocessing is denoted as Then there is represents the graph nodes and edges stored in the Graph class data, represents the graph G i The character number vector after text preprocessing of.
[0076] 3) Select appropriate encoders for data of different modalities according to the aforementioned S2, and pre-train the encoders in the way of contrastive learning. In this example, use the graph neural network ResGCN as the encoder for graph topology structure data, and use large-scale pre-trained language models such as BERT as the encoder for text vector data. For the graph topology structure, adopt the SimGRACE contrastive learning framework; for data in the text modality, adopt the SimCSE contrastive learning framework. In this example, the loss calculation of contrastive learning adopts the cosine similarity calculation method. The contrastive learning temperature parameter is set to 0.001 with reference to previous research work. Set the maximum number of training epochs of the encoder to 200. Based on the trained encoder, encode the original modality data to obtain the feature vectors of the corresponding modalities. Denote the feature representation of the graph topology structure modality as The feature representation of the text modality as
[0077] 4) According to the aforementioned S3, fuse the feature representations of the topology structure information and text information obtained by encoding with the encoder. The graph node feature representation and text feature representation are set to 768 dimensions, and the aligned topology structure feature vector and text feature vector are used to perform the Cartesian vector product to extract cross-modal common features, and the cross-modal feature representation is normalized from 0 to 1. Denote the cross-modal feature representation as
[0078] 5) According to the steps of the aforementioned S4, after normalizing the extracted cross-modal common features and the original features of different modalities respectively, merge them to obtain Calculate the weights of the three parts of feature vectors through the attention mechanism, weight the features, and use the weighted vector as the final graph-level feature representation. Denote the graph-level feature representation as Input it into the linear classifier to obtain the graph classification result.
[0079] Name the graph data classification method of the present invention as MMCLGC. Compared with the original classical graph data classification method GCN (Kipf, Thomas N., and Max Welling. "Semi-supervised classification with graph convolutional networks." arXiv preprint arXiv:1609.02907 (2016)) under the small-sample setting, multiple indicators of its model recognition performance have been improved. The following Table 1 shows the model performance data under the setting that 1% of the training data has known labels. The comprehensive index F1 of the classification task has increased by 8.26%.
[0080] Table 1
[0081]
[0082] In order to further analyze the effects of each step in the MMCLGC method proposed by the present invention and its impact on the reconstruction result, a comparative experiment with different practices was designed by adjusting the steps and experimental parameters. The specific scheme and test results are shown in Table 2, where the test parameters represent the steps executed.
[0083] Table 2
[0084]
[0085] Among them, the parameters of Test 1 are the same as those of the classical GCN, and the parameters of Test 4 are the same as those of the MMCLGC proposed by the present invention. The order of the test parameters is the same as the order of the operation process in the test. It can be analyzed from the specific test parameters and test results that: comparing Test 1 and Test 3, the contrastive learning of the topological structure data effectively utilizes the information of the unlabeled training data. Comparing Test 2 and Test 4, the contrastive learning of the text data effectively utilizes the information of the unlabeled training data. Comparing Test 5 with Test 3 and Test 4, the multi-modal feature fusion can utilize the topological structure information and text information simultaneously, indicating that multi-modal learning can effectively aggregate different modal information and improve the expression ability of the data. Comparing Test 6 and Test 5, on the basis of the feature fusion of Test 5, Test 6 uses the Cartesian product to extract the common features of different modalities, effectively enhancing the expression ability of different modal features during fusion.
[0086] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A text network graph classification method based on multi-modal contrastive learning, characterized in that, The steps are as follows: S1: For each text network graph data in the multi-modal network graph dataset to be classified, extract the two types of modal data, namely the topological structure in the graph and the text in the nodes respectively. After the extracted data is classified by modality, it is saved in dictionary format; then preprocess each type of modal data to meet the input requirements of the corresponding modal encoder; S2: Select matching encoders for the topological structure modality and the text modality respectively, and train them using the contrastive learning framework respectively; Based on the trained encoders, perform feature encoding on each type of modal data preprocessed in S1 to obtain the feature vectors of each type of modal data in each text network graph data, so as to obtain the feature representations of the text network graph data in different modalities; S3: For each text network graph data, after aligning the feature vectors of the corresponding two types of modal data to the same dimension, calculate the Cartesian product of the two to obtain a feature cross matrix. Perform horizontal max pooling on the feature cross matrix to obtain a first feature vector, and perform vertical max pooling on the feature cross matrix to obtain a second feature vector. Concatenate the first feature vector and the second feature vector and then re-reduce the dimension to the said same dimension, so as to obtain a cross-modal common feature vector; S4: For each text network graph data, standardize the feature vectors of the two types of modal data and the cross-modal common feature vector, then use the attention mechanism to calculate the attention weights for the three feature vectors, and perform weighted fusion on the three feature vectors according to the attention weights to obtain the final graph-level feature and then input it into the classifier to obtain the classification label of each text network graph data in the multi-modal network graph dataset.
2. The text network graph classification method based on multi-modal contrastive learning according to claim 1, wherein In the step S1, for the multi-modal network graph dataset, each text network graph data is extracted, saved and preprocessed according to S11~S14: S11: For each text network graph data G i Assign a unique identifier, i = 1, 2, …, N, where N is the scale of the multimodal network graph data set; establish a graph data dictionary for storing different modal data in each text network graph data; S12: Number the nodes included in each text network graph data, and store the adjacency relationship of each pair of nodes in the graph in the form of an ordered pair list according to the relationship information in the text network graph data, so as to extract the topological structure modal data of the text network graph data, and store the unique identifier of the text network graph data in the corresponding graph data dictionary according to S11; S13: For each text network graph data, extract the content text data in the nodes in order according to the node numbers when extracting the topological structure information, and store the unique identifier of the text network graph data in the corresponding graph data dictionary according to S11; S14: Perform preprocessing on different modal data in the graph data dictionary respectively to make it form structured data adapted to the required input encoder; among them: For the topological structure modal data, it is necessary to define Graph class data, and store the node numbers and the node adjacency relationships saved in the form of an ordered pair list respectively; For the text modal data, first disassemble and segment the text sequence and standardize it, and map the word and character encoding to a numerical value according to the vocabulary, so as to process the text sequence into a numerical vector.
3. The text network graph classification method based on multi-modal contrastive learning according to claim 1, characterized in that, For the text modal data, disassemble and segment the text sequence and standardize it based on the Tokenize tool function.
4. The method for classifying text network diagrams based on multi-modal contrastive learning according to claim 1, wherein The specific method of step S2 is as follows: S21: Select the graph neural network GCN as the encoder for the topological structure modality, and select the text pre-trained model BERT as the encoder for the text modality; S22: Set up contrastive learning frameworks for different modalities respectively. Among them, for the topological structure modality data, adopt the SimGRACE contrastive learning framework, and for the text modality data, adopt the SimCSE contrastive learning framework; for the encoders of each data modality, after constructing positive and negative samples from the training data in batches and inputting them into the encoder, calculate the contrastive learning loss according to the corresponding contrastive learning framework, and update the model parameters of the encoder through backpropagation of the neural network until all training data participates in the training, which is regarded as completing one epoch; according to the situation of the decrease in the contrastive learning loss, set an early stopping strategy. After completing the specified number of epoch training times, obtain the encoder trained based on contrastive learning; S23: Based on the encoders corresponding to each data modality after training, encode the corresponding modality structured data preprocessed in S1 to obtain the feature vectors of the corresponding modality data; each text network graph data respectively obtains the feature vectors of the topological structure modality data and the text modality data.
5. The text network graph classification method based on multi-modal contrastive learning according to claim 4, characterized in that In the SimGRACE contrastive learning framework, two graph encoders are used to encode the same data during the training process; during the training process with batch as the data input, the feature vectors obtained by encoding the same graph data in each batch twice are used as positive samples, and the feature vectors obtained by encoding other graph data in the batch are used as negative samples; for the two encoders used in the training, a base encoder needs to be initialized first, and at the same time, on the basis of copying the parameters of the base encoder, add random perturbations based on the original parameter Gaussian distribution to obtain the parameters of the other encoder.
6. The method for classifying text network diagrams based on multi-modal contrastive learning according to claim 4, characterized in that In the SimCSE contrastive learning framework, inputting the same sample into the encoder twice obtains the positive samples for contrastive learning.
7. The method for classifying text network diagrams based on multi-modal contrastive learning according to claim 1, wherein The specific method of step S3 is as follows: S31: For the feature vectors of the topological structure modality data and the text modality data in each text network graph data, align the dimensions of the two feature vectors to a unified dimension; S32: Perform the Cartesian vector product on the two aligned feature vectors to obtain the feature cross matrix M; then perform max pooling on the row vectors of the feature cross matrix M to obtain the first feature vector, and perform max pooling on the column vectors of the feature cross matrix M to obtain the second feature vector, so as to extract all the information that is important in both modalities; S33: Concatenate the first feature vector and the second feature vector obtained by the two max poolings and use linear mapping to reduce the dimension to a unified dimension to obtain the cross-modal common feature vector.
8. The text network graph classification method based on multi-modal contrastive learning according to claim 7, characterized in that, The unified dimension is set to 64, 128, or 768.
9. The method for classifying text network diagrams based on multi-modal contrastive learning according to claim 1, wherein, The specific method of step S4 is as follows: S41: For each text network graph data, standardize the feature vectors of the two modality data and the cross-modal common feature vector, and then input them into the attention mechanism together. Through the attention mechanism, calculate the weight of the three feature vectors to obtain the attention weights of the three feature vectors; S42: For each text network graph data, perform weighted fusion on the three feature vectors according to the attention weights calculated in S41, and use the vector after weighted fusion as the final graph-level feature representation; S43: For each text network graph data in the multimodal network graph dataset, input the corresponding graph-level feature representation into a linear classifier to obtain the corresponding graph classification result.
10. The text network graph classification method based on multi-modal contrast learning according to claim 1, wherein, The text network graph data is rumor propagation tree data. Each piece of rumor propagation tree data contains a seed node and interaction nodes. The seed node is the original information, and the interaction nodes are the forwards and comments based on the original information. Each node contains text content related to the original information; the classification label corresponding to the rumor propagation tree data is the label indicating whether the original information is a rumor.