Text Classification Method Based on Self-Supervised Dual-Granularity Multi-Graph Learning

Through the self-supervised two-grained multi-graph learning method, a multi-graph data set is constructed and self-supervised learning is performed, which solves the problems of intrinsic structural relationships and fine-grained annotation in text classification, and accurately predicts labels on coarse and fine-grained sizes.

CN116401361BActive Publication Date: 2025-07-18NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310038679.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-13
Publication Date
2025-07-18
Estimated Expiration
2043-01-13

AI Technical Summary

Technical Problem

Existing text classification methods are difficult to effectively capture the internal structural relationships and fine-grained labeling information of each part of the text, and cannot make accurate label predictions at the same time on coarse and fine-grained size.

Method used

By using the self-supervised two-particle multi-graph learning method, a multi-graph data set is constructed, and a graph encoder is used to learn graph representations using the enhanced encoder and graph encoder. Combining the multi-head self-attention mechanism and packet-level graph generation, package-level comparison loss and graph-level comparison loss are designed to realize the graph-graph learning mechanism, and self-supervised learning package representation and graph representation.

Benefits of technology

Effectively retain the internal structural relationship and global structural relationship of text data, and can simultaneously predict labels on coarse and fine-grained size, reduce dependence on label information, and improve the accuracy of text classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116401361B_ABST
    Figure CN116401361B_ABST
Patent Text Reader

Abstract

The present invention provides a text classification method based on self-supervised dual-granularity multi-graph learning, which relates to the technical field of text classification. The method first obtains the original text dataset and the corresponding label set, and performs data preprocessing on the original text dataset to obtain a multi-graph dataset; then uses an enhanced encoder to enhance the graph data, and uses a graph encoder to learn the enhanced graph representation; then applies the multi-head self-attention mechanism to the graph representation to learn the context information between the graphs in the graph bag, generates a bag-level graph, and uses a bag encoder to learn the bag representation through the bag-level graph; then simultaneously learns the graph representation and the bag representation through a graph-graph learning mechanism, and designs a bag-level contrast loss and a graph-level contrast loss as loss functions to self-supervisedly learn the bag representation and the graph representation; finally, for the text classification task to be classified, uses the learned bag representation and graph representation to simultaneously perform label prediction on the text to be classified at a coarse-grained and fine-grained level, thereby realizing text classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text classification, and in particular to a text classification method based on self-supervised dual-granularity multi-graph learning. Background Art

[0002] With the continuous development of the Internet, the text data on the network is increasing day by day. If these data can be effectively classified, it is more conducive to mining valuable information from them. Therefore, the management and integration of text data are very important. Text classification refers to the automatic classification and annotation of a text set (or other entities or objects) by a computer according to a certain classification system or standard.

[0003] The key to the text classification problem lies in extracting feature representations from text data that can express text information as much as possible. Traditional text classification methods are mainly divided into two types, namely machine learning-based text classification methods and deep learning-based text classification methods. The common idea of machine learning-based text classification methods is to use feature engineering for text representation, and then classify through classifiers such as Support Vector Machine (SVM), Naive Bayes (NB), and K-Nearest Neighbor (KNN). Compared with machine learning-based text classification methods, deep learning-based text classification methods usually use their own network structures to automatically learn the feature representations of data, without manually obtaining text features through feature engineering. Common deep learning-based text classification models include Convolutional Neural Networks (CNNs) models, Recurrent Neural Networks (RNNs) models, etc. For example, the TextCNN model regards a sentence as a feature matrix composed of multiple word vectors, performs convolution on it using convolutional kernels of different sizes, and then extracts features from the convolved results through a pooling layer. Liu et al. used an RNN model for classification in a text classification task, regarded the text as a time series, and combined context information to learn the feature representation of the text. Yang et al. combined an RNN model with an attention mechanism, represented the text as a hierarchical structure of "word - sentence - text", and learned based on different attention weights.

[0004] In recent years, graph neural networks have received extensive attention, and some methods for text classification using graph neural networks have emerged. For example, the TextGCN model first applied the graph convolutional network to the text classification task, regarding both words and texts as nodes, constructing an undirected weighted graph, and learning text embeddings and word embeddings. Huang et al. proposed the Text-level GCN model, which constructs each text as a directed graph, uses the globally shared node feature matrix and edge weight matrix for learning, and adopts the message passing mechanism for updating. Yuan et al. adopted the G-ATT model, used the attention mechanism for text sentiment analysis, constructed an undirected dependency tree for each sentence to describe the dependency relationship between words and grammar, and performed fusion through the memory network.

[0005] Traditional machine learning-based text classification methods and deep learning-based text classification methods usually adopt the single-instance or multi-instance learning framework, that is, it is assumed that each text data can be represented as one or more feature vectors (instances) through feature engineering or feature learning, and it is assumed that this one or more feature vectors (instances) can contain the key information of the text, and this key information is sufficient to distinguish it from text data of other categories. However, text data often has complex semantic features. Representing it only with data in the Euclidean space such as feature vectors cannot accurately express the structural information of the text data and the structural relationship between contexts, resulting in information loss. In addition, traditional machine learning-based models and deep learning-based models such as CNN usually need to utilize the translational invariance of the Euclidean space to effectively obtain the feature information of the data. For non-Euclidean data such as text data, this often limits the expressive power of the neural network and cannot express complex semantic information.

[0006] In recent years, the proposed text classification methods based on graph neural networks adopt the single-graph learning framework, which is an extension of the single-instance learning framework, using a single graph to represent instead of a single instance (feature vector). That is, it is assumed that each text data can be represented as a graph-structured data. However, the single-graph learning framework represents each text in the form of a single graph, with each part represented as a node and in the form of a feature vector. This way can only represent the global structural relationship between parts of the text, but cannot represent the internal structural relationship of each part of the text and cannot capture fine-grained structural information. Taking the text sentiment classification task as an example, the single-graph learning framework usually represents a text in the form of a single graph, with each sentence (paragraph) as a node and represented by a feature vector. This representation form often ignores the internal structural relationship of each sentence (paragraph) and can only capture the global structural relationship of the whole paragraph (whole article).

[0007] In addition, in existing text classification tasks, due to the ambiguity of fine-grained (e.g., each paragraph in an article) text annotation information, most classification methods can only perform coarse-grained (e.g., the whole article) text annotation and cannot perform fine-grained text annotation. This is because in text classification tasks, most datasets can only contain the annotation information of each article and do not contain the annotation information of each natural paragraph inside the article. Therefore, most existing text classification methods can only predict the annotation information of the whole article, while the annotation information of each natural paragraph inside the article is ambiguous and cannot be predicted. However, in real life, the annotation information of different natural paragraphs in an article is also very important. For example, in text sentiment classification tasks, different natural paragraphs may express different emotions, and predicting the emotions of each natural paragraph in the article is beneficial to tasks such as mental health assessment and human-computer interaction. Therefore, for text classification tasks, it is very important to perform both coarse-grained and fine-grained label prediction based on coarse-grained annotation information. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a text classification method based on self-supervised dual-grained multi-graph learning to achieve automatic classification of texts in view of the above-mentioned deficiencies of the prior art.

[0009] To solve the above technical problem, the technical solution adopted by the present invention is as follows: A text classification method based on self-supervised dual-grained multi-graph learning includes the following steps:

[0010] Step 1: Obtain the original text dataset and the corresponding label set;

[0011] Step 2: Perform data preprocessing on the original text dataset to obtain the multi-graph data structure corresponding to the original text dataset, that is, a graph package, to form a multi-graph dataset;

[0012] Extract the relevance between keywords in the text; then use the keywords of each text as nodes and the relevance between keywords as the edge weight value to construct a graph, and remove the edges less than a given threshold based on the threshold, and set the edge weight values greater than or equal to the given threshold to 1 to form an undirected graph; represent each text in the original text dataset as a multi-graph structure to form a graph package B = {g1, g2,..., g n}, where g1, g2,..., g n represent multiple graphs formed by the selected texts;

[0013] Step 3: Use an enhanced encoder to enhance the graph data and use a graph encoder to learn the enhanced graph representation;

[0014] Perform two data augmentation operations on all graph data in the multi-graph dataset using an augmentation encoder. Since each graph in the graph package undergoes two data augmentations, two sets of augmented multi-graph packages will be obtained respectively. The augmentation encoder for data augmentation of graph data is shown in the following formula:

[0015] f aug (G) = {V, ε; ∈} (1)

[0016] Among them, f aug represents the augmentation encoder, G = (V, ε) represents the graph, V is the vertex set, v p ∈V, containing the attribute information of each node; ε is the edge set, (v p , v q ) ∈ ε, p ≠ q; ∈ is the augmentation method; the augmented graph is denoted as The graph package is denoted as:

[0017]

[0018] Among them, represents the augmented graph package;

[0019] To effectively retain the structural information of the graph data in the multi-graph dataset, a graph encoder is used to learn the augmented graph representation, as shown in the following formula:

[0020]

[0021] Among them, is the graph representation generated by the graph encoder, in the form of a vector, containing all the node attribute information and the internal structural relationship in the augmented graph; H l represents the augmented graph The node representations of all nodes in the l-th layer of the neural network in are represented, and the node representation of each layer is updated through the function f genc based on the node representation of the previous layer. Initially, H 0 is the attribute set of all nodes in the augmented graph, Λ is the number of layers of the neural network to be learned; f genc is a learnable function for updating the node representation of each layer, Among them, W l is the learnable weight matrix, represents the adjacency matrix after adding self-connections, A is the adjacency matrix, I is the identity matrix, is obtained from the degree matrix; f p is the pooling function for obtaining the graph representation by taking the mean of the learned node representations;

[0022] Step 4: Apply the multi-head self-attention mechanism to the graph representations to learn the context information among the graphs in the graph bag. The graph representations containing context information are connected based on similarity to generate a bag-level graph, and a bag encoder is used to learn the bag representation through the bag-level graph;

[0023] Apply the multi-head self-attention mechanism to each enhanced graph representation in the bag to obtain graph representations containing context information. The formula of the multi-head self-attention mechanism is as follows:

[0024]

[0025] where W′ is a learnable parameter, and head m represents the result of the m-th self-attention head, and m represents the number of defined self-attention heads; head m is expressed as:

[0026]

[0027] where, represents the learnable parameter in the m-th head, and d k is the hidden layer dimension; is formed by concatenating the graph representations in the graph bag, represents the bag containing context information learned after passing through the multi-head self-attention mechanism, and is composed of multiple graph representations containing context information after passing through the multi-head self-attention mechanism concatenated together, represents concatenation, and n represents the number of graphs in the graph bag;

[0028] Adopt a graph generation method, use each graph in the graph bag as a node, and use the correlation between the graph representations containing context information among the graphs as the edge weight, and the graph representation of each graph as the node attribute value. Based on a threshold, a bag-level graph of the graph bag is formed. The generation method of the bag-level graph is as follows:

[0029]

[0030] where I[·] is an indicator function, and the result is 1 when the content in I[·] is greater than 0, otherwise 0; μ is a threshold used to remove edges with low correlation between graphs, is the cosine similarity, used to measure the similarity between any two graphs in the graph bag, is the weight value between graph i and graph j generated based on the threshold and cosine similarity, serving as the adjacency matrix of the bag-level graph; the generated bag-level graph has each graph in the graph bag as a node, with the graph representation as the node attribute, as the adjacency matrix, that is, the generated bag-level graph

[0031] To obtain the vector representation of the multi-graph package containing the global structural relationships between the graphs in the package based on the generated package-level graph, a package encoder is set as shown in the following formula:

[0032]

[0033] where is the package representation generated by the package encoder, and f norm is the regularization function used to regularize the data; f benc is a learnable function that uses the graph convolution operator to update the node representations at each layer, and f benc is expressed as where W l is the learnable weight matrix, σ is the activation function, and initially H 0 is the set of attributes of all nodes in the package-level graph;

[0034] Step 5: Simultaneously learn the graph representation and the package representation through the graph-graph learning mechanism, and effectively retain the context information and global structural relationships between the graphs in the graph package;

[0035] The package encoder and the graph encoder are used to simultaneously learn the package representation and the graph representation, that is, the package-level graph and the nodes in the package-level graph are learned simultaneously, forming a graph-graph learning mechanism. This learning mechanism can effectively learn the package representation and the graph representation, while retaining the context information and global structural relationships between the individual graphs in the package, which is beneficial to the coarse-grained and fine-grained classification tasks of the multi-graph learning problem;

[0036] Step 6: Design the package-level contrastive loss and the graph-level contrastive loss as loss functions, and self-supervisedly learn the package representation and the graph representation on the premise of ensuring package-level invariance and graph-level invariance;

[0037] To adopt the mechanism of contrastive learning and self-supervisedly learn the package representation and the graph representation, it is necessary to ensure package-level invariance and graph-level invariance; for this purpose, the package-level contrastive loss and the graph-level contrastive loss are designed as loss functions, as shown in the following formula:

[0038]

[0039]

[0040] where I g ={1...2n}, n represents the number of graphs in the dataset, and I b ={1...2N}, N represents the number of packages in the dataset; A g (i)=I g \{i}, A b (i)=I b \{i}; is a graph contrastive loss function for ensuring graph-level invariance, is a bag contrastive loss function for ensuring bag-level invariance; sim(·) and simb(·) are functions for measuring the similarity between two representations, which can be expressed as sim(z1, z2) = exp(z1·z1 / τ) and simb(Z1, Z2) = exp(cos(Z1, Z2) / τ) respectively, where τ is a temperature parameter and cos(Z1, Z2) is a cosine function, f proj is a projection network f proj (x) = σ(f norm (ω x +b));

[0041] Then the loss function of the text classification method based on self-supervised dual-granularity multi-graph learning is expressed as

[0042] Step 7: For the text classification task to be classified, use the bag representation and graph representation learned in Step 6 to simultaneously perform label prediction on the text to be classified at both the coarse-grained and fine-grained levels, realizing text classification.

[0043] The beneficial effects of adopting the above technical solutions are as follows: The text classification method based on self-supervised dual-granularity multi-graph learning provided by the present invention models the text classification problem as a multi-graph learning problem and adopts a multi-graph learning framework, which can effectively retain the internal structural relationships of each part and the global structural relationships between parts in the text data, making it easier to express text information and effectively retaining the coarse-grained text information and fine-grained text information. At the same time, a self-supervised learning paradigm is adopted, and the training process is carried out in a self-supervised manner based on contrastive learning without the need for additional annotation information. In addition, based on self-supervised contrastive representation learning, a graph-graph learning mechanism is proposed, which can simultaneously self-supervised learn graph representations (fine-grained) and bag representations (coarse-grained) in each iteration, so that the downstream task can perform label classification at both the coarse-grained and fine-grained levels using only a small number of bag labels (coarse-grained labels). Description of the Drawings

[0044] Figure 1 is a flowchart of the text classification method based on self-supervised dual-granularity multi-graph learning provided by an embodiment of the present invention;

[0045] Figure 2 is a schematic diagram of converting a text data set into a multi-graph data set provided by an embodiment of the present invention;

[0046] Figure 3 is a schematic diagram of the graph-graph learning mechanism provided by an embodiment of the present invention;

[0047] Figure 4This is the text classification diagram provided by the embodiments of the present invention. Among them, (a) is the original scientific and technological paper text, (b) is the reference of the scientific and technological paper, (c) is the annotated scientific and technological paper text, and (d) is the annotated reference. Detailed implementation manners

[0048] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0049] In this embodiment, a text classification method based on self-supervised dual-granularity multi-graph learning, as Figure 1 shown, includes the following steps:

[0050] Step 1: Obtain the original text dataset and the corresponding label set;

[0051] This embodiment takes the bibliographic data in the field of computer science from the DBLP database as an example. Among them, each record contains information such as abstract, author, year, title, and references. The paper data published in three fields of artificial intelligence (AI), computer vision (CV), and data mining (DB) is selected to form the original text dataset. Based on the text information of each paper, the scientific research field to which the paper belongs is predicted.

[0052] Step 2: Perform data preprocessing on the original text dataset to obtain the multi-graph data structure corresponding to the original text dataset, that is, the graph package, and form a multi-graph dataset;

[0053] Figure 2 This is the schematic diagram of converting the DBLP dataset provided by this embodiment into a multi-graph dataset. The method of converting the selected DBLP dataset into a multi-graph dataset is as Figure 2 shown, and can be specifically described as: First, select a paper from the formed DBLP text literature dataset. This paper usually contains information such as abstract, abstract keywords, text, references, and authors. Since in a scientific and technological paper, the keywords in the abstract of this paper can usually accurately describe the content and the field of this paper. At the same time, the other papers cited in the references also have a certain ability to describe the field of this scientific and technological paper. Therefore, we use the keywords in the abstracts of each paper and its references as nodes, and the association relationship between the keywords as edges to construct an undirected graph. In this way, not only can the information of the keywords themselves be modeled, but also the structural relationship between different keywords in the abstract can be modeled, which can better express the field information of the text literature.

[0054] Specifically, the E-FCM algorithm is used to extract the correlations between keywords in each abstract; then, taking the keywords of each abstract as nodes and the correlations between keywords as the edge weight values, a graph is constructed, and edges with weights less than a given threshold are removed based on the threshold, and the edge weight values greater than or equal to the given threshold are all set to 1 to form an undirected graph; each abstract of a literature and its references can be transformed into a graph, so a literature is represented as a multi-graph structure, forming a graph package. Therefore, the text literature dataset constructed based on the DBLP database can be modeled as a multi-graph dataset, that is, each paper can be represented as a graph package B = {g1, g2,..., g n}, where g1 represents the graph formed by the abstract of the selected paper, and g2...g n represents the graph formed by the abstracts of the references of the selected paper.

[0055] Step 3: Use an augmentation encoder to augment the graph data and use a graph encoder to learn the augmented graph representation;

[0056] In order to better learn the feature representation of text literature data from the modeled multi-graph dataset as much as possible without using label information. Adopting the idea of contrastive learning, regarding the multi-graph package corresponding to each paper as a separate category, learning the representation information that can distinguish all multi-graph packages as much as possible. In addition, for the fine-grained text classification task (that is, predicting which category the paper and its references in the multi-graph package formed by each paper belong to), it is also necessary to regard each graph (abstract) as a separate category and learn the representation information that can distinguish all graphs as much as possible. Therefore, as Figure 3 shown, two data augmentation operations are performed on all graph data in the multi-graph dataset using an augmentation encoder. Since each graph in the graph package has been augmented twice, two sets of augmented multi-graph packages will be obtained respectively; the augmentation encoder for data augmentation of graph data is shown in the following formula:

[0057] f aug (G) = [V, ε; ∈} (1)

[0058] where f aug represents the augmentation encoder, G = (V, ε) represents the graph, V is the vertex set, v p ∈ V, containing the attribute information of each node; ε is the edge set, (v p , v q ) ∈ ε, p ≠ q; ∈ is the augmentation method, including deleting some nodes, deleting some edges, masking the attributes of some nodes or edges, etc.; the augmented graph is represented as The graph package is represented as:

[0059]

[0060] where Denote the augmented graph packet;

[0061] To effectively retain the structural information of graph data in a multi-graph dataset, a graph encoder is used to learn the augmented graph representation, as shown in the following formula:

[0062]

[0063] where is the graph representation generated by the graph encoder, in the form of a vector, containing all the node attribute information and the internal structural relationships in the augmented graph; H l denotes the augmented graph represents the node representations of all nodes in the augmented graph at the l-th layer of the neural network. The node representations of each layer are updated through the function f genc based on the node representations of the previous layer. Initially, H 0 is the set of attributes of all nodes in the augmented graph, Λ is the number of layers of the neural network to be learned; f genc is a learnable function for updating the node representations of each layer, where l is the learnable weight matrix, denotes the adjacency matrix after adding self-connections, A is the adjacency matrix, I is the identity matrix, is obtained from the degree matrix; f p is the pooling function for obtaining the graph representation by taking the mean of the learned node representations;

[0064] That is, the learning process in formula (3) can be described as iteratively learning the information of each keyword in the abstract Λ times through the function f genc to obtain the vector representation of each keyword. Then, through the pooling function f p , the vector representation of this abstract is obtained based on the learned vector representations of the keywords and the structural information between the keywords

[0065] Step 4: Apply the multi-head self-attention mechanism to the graph representation to learn the context information between the graphs in the graph packet. The graph representations containing context information are connected based on similarity to generate a packet-level graph, and a packet encoder is used to learn the packet representation through the packet-level graph;

[0066] Since there is a certain connection between a paper and its references, this connection helps in learning the packet representation of the paper. Therefore, to effectively retain the relationship between the paper and its references, the multi-head self-attention mechanism is applied to the augmented graph representations of each graph in the packet to obtain the graph representations containing context information; the formula of the multi-head self-attention mechanism is as shown in the following formula:

[0067]

[0068] Among them, W′ is a learnable parameter, and head m represents the result of the m-th self-attention head, where m represents the number of defined self-attention heads. Different self-attention heads can describe the relationship between the paper and its references from different perspectives. head m is expressed as:

[0069]

[0070] Among them, represents the learnable parameter in the m-th head, and d k is the hidden layer dimension; is formed by splicing the graphs in the graph package, then can also be expressed as represents the package containing context information learned after the multi-head self-attention mechanism, which is composed of multiple graphs containing context information after the multi-head self-attention mechanism spliced together, represents splicing, and n represents the number of graphs in the graph package;

[0071] Due to the representation form of vectors, the global structural relationship existing between the paper and its references cannot be effectively retained. Therefore, in the way of graph generation, each abstract in the paper package is used as a node, and the graph representation containing context information between the abstracts the correlation between them is used as the edge weight, and the vector representation of each paper is used as the node attribute value. Based on the threshold, the package-level graph of the paper is formed. The generation method of the paper package-level graph is shown in the following formula:

[0072]

[0073] Among them, I[·] is the indicator function, which is 1 when the content in [·] is greater than 0, and 0 otherwise; μ is the threshold, which is used to remove the edges with low correlation between the abstracts, is the cosine similarity, which is used to measure the similarity between any two abstracts in the paper package, is the weight value between abstract i and abstract j generated based on the threshold and cosine similarity, which is used as the adjacency matrix of the package-level graph; the generated package-level graph uses each abstract in the paper package as a node and the graph representation of the abstract as the node attribute, as the adjacency matrix, that is, the generated package-level graph

[0074] In order to obtain the vector representation of the paper package containing the global structural relationship between the abstracts based on the generated package-level graph, a package encoder is set, as shown in the following formula:

[0075]

[0076] Among them, is the packet representation generated by the packet encoder, and f norm is a regularization function used to regularize the data; f benc is a learnable function that uses a graph convolutional operator to update the node representation of each layer, and f benc is expressed as Among them, W l is a learnable weight matrix, σ is an activation function, and initially H 0 is the set of attributes of all nodes in the packet-level graph;

[0077] Step 5: Simultaneously learn the graph representation and the packet representation through the graph-graph learning mechanism as shown in Figure 3 and effectively retain the context information and global structural relationships between the graphs in the graph packet;

[0078] As described in Steps 3 and 4, the packet-level graph takes the graphs in the graph packet as nodes, and the global structural relationships containing context information between the graphs as edges. That is, the graph packet is represented in the form of a packet-level graph, and its nodes themselves are also graphs; the packet encoder and the graph encoder are used to simultaneously learn the packet representation and the graph representation, that is, the packet-level graph and the nodes in the packet-level graph (the graphs in the packet) are simultaneously learned, forming a graph-graph learning mechanism. This learning mechanism can effectively learn the packet representation and the graph representation, while retaining the context information and global structural relationships between the individual graphs in the packet, which is beneficial to the coarse-grained (packet-level) and fine-grained (graph-level) classification tasks of the multi-graph learning problem;

[0079] The two enhanced graph representations generated by the enhanced encoder are respectively trained using graph encoders, multi-head self-attention mechanisms, graph generation mechanisms, and packet encoders with the same structure but different parameters to generate graph (abstract) representations and packet (paper) representations. Since the packet-level graph generated by using the graph generation mechanism is a graph-structured data with each abstract in the paper packet as nodes and the global structural relationships containing context information between the abstracts as edges, and each abstract in the paper is represented as a graph-structured data with keywords as nodes and local structural relationships between the keywords as edges, that is, a graph-graph structure where the nodes within the graph are themselves graphs. In each iteration, the graph representation (abstract) and the packet representation (paper) are simultaneously learned, forming a graph-graph learning mechanism. This learning mechanism can effectively learn the graph representation and the packet representation, while effectively retaining the internal structural relationships between the individual keywords within the abstract and the global structural relationships between the individual abstracts within the paper, which is beneficial to the coarse-grained and fine-grained classification tasks of the text classification task modeled based on multi-graph learning.

[0080] Step 6: Design the bag-level contrastive loss and the graph-level contrastive loss as loss functions, and self-supervised learn the bag representation and the graph representation while ensuring bag-level invariance and graph-level invariance;

[0081] To adopt the mechanism of contrastive learning and self-supervised learn the bag representation and the graph representation, it is necessary to ensure bag-level invariance and graph-level invariance; for this purpose, the bag-level contrastive loss and the graph-level contrastive loss are designed as loss functions, as shown in the following formula:

[0082]

[0083]

[0084] where I g ={1...2n}, n represents the number of graphs in the dataset, I b ={1...2N}, N represents the number of bags in the dataset; A g (i)=I g \{i}, A b (i)=I b \{i}; is the graph contrastive loss function for ensuring graph-level invariance, is the bag contrastive loss function for ensuring bag-level invariance; sim(·) and simb(·) are functions for measuring the similarity between two representations, which can be expressed as sim(z1, z2)=exp(z1·z1 / τ) and simb(Z1, Z2)=exp(cos(Z1, Z2) / τ) respectively, where τ is a temperature parameter and cos(Z1, Z2) is a cosine function, f proj is a projection network f proj (x)=σ(f norm (ωx + b));

[0085] Then, the loss function of the text classification method based on self-supervised dual-grained multi-graph learning is expressed as

[0086] Step 7: For the text classification task to be classified, use the bag representation and the graph representation learned in Step 6 to simultaneously perform label prediction on the text to be classified at both the coarse-grained and fine-grained levels, and achieve text classification.

[0087] The processes described in Steps 1-6 are all the training stages of the self-supervised learning task, which are used to self-supervised learn the graph representation and the bag representation; then, the learned graph representation and bag representation can be used to perform label prediction at both the bag level and the graph level by using models such as SVM, KNN, and multi-layer perceptron as classifiers, and only using a small amount of bag label data (10% - 20%).

[0088] In this embodiment, the method of the present invention is used to annotate the scientific and technological papers and their references shown in Figure 4 (a) and Figure 4 (b). The paper is represented as a multi-graph package composed of graphs transformed from the abstract and the reference abstract; Figure 4 (c) and Figure 4 (d) are scientific and technological papers classified by the method of the present invention; among them, the scientific and technological paper package is labeled with two tags: artificial intelligence (AI) and computer vision (CV), and the graphs in the package (paper abstract and reference abstract) are also respectively labeled with two tags: artificial intelligence (AI) and computer vision (CV).

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.

Claims

1. A text classification method based on self-supervised dual-granularity multi-graph learning, characterized in that: It includes the following steps: Step 1: Obtain the original text dataset and the corresponding label set; Step 2: Perform data preprocessing on the original text dataset to obtain the multi-graph data structure corresponding to the original text dataset, i.e., the graph bag, and form a multi-graph dataset; Step 3: Use the augmentation encoder to augment the graph data and use the graph encoder to learn the augmented graph representation; Step 4: Apply the multi-head self-attention mechanism to the graph representation to learn the context information between the graphs in the graph bag. The graph representations containing context information are connected based on similarity to generate the bag-level graph, and the bag encoder is used to learn the bag representation through the bag-level graph; In the way of graph generation, each graph in the graph package is used as a node, and the correlation between the graph representations with context information between the graphs is used as the edge weight. The graph representation of each graph is the node attribute value. Based on the threshold, the package-level graph of the graph package is constructed. The generation method of the package-level graph is shown in the following formula: The correlation between them is used as the edge weight, and the graph representation of each graph is the node attribute value. Based on the threshold, the package-level graph of the graph package is constructed. The generation method of the package-level graph is shown in the following formula: where, I[·] is an indicator function that results in 1 when the content in [·] is greater than 0, and 0 otherwise; μ is a threshold used to remove edges with relatively low correlation between graphs. is the cosine similarity, which is used to measure the similarity between any two graphs in the graph bag. is the weight value between graph i and graph j generated based on the threshold and cosine similarity, serving as the adjacency matrix of the bag-level graph; the generated bag-level graph has each graph in the graph bag as a node, and is represented by as the node attribute. is composed of the adjacency matrix, that is, the generated bag-level graph. In order to obtain the vector representation of the multi-graph bag containing the global structural relationship between the graphs in the bag based on the generated bag-level graph, a bag encoder is set as shown in the following formula: Among them, is the packet representation generated by the packet encoder, f norm is a regularization function used to regularize the data; f benc is a learnable function that uses the graph convolutional operator to update the node representations at each layer, f benc is expressed as where, W l is the learnable weight matrix, σ is the activation function, and initially H 0 is the set of attributes of all nodes in the packet-level graph; f p is the pooling function used to obtain the graph representation by taking the mean of the learned node representations; H l represents the enhanced graph represents the node representations of all nodes in the neural network at the l-th layer, and the node representations at each layer are updated by the function f genc based on the node representations of the previous layer; represents the adjacency matrix after adding self-connections, A is the adjacency matrix, I is the identity matrix, is composed of to obtain the degree matrix; Step 5: Simultaneously learn the graph representation and the bag representation through the graph-graph learning mechanism, and effectively retain the context information and global structural relationship between the graphs in the graph bag; Step 6: Design the bag-level contrast loss and the graph-level contrast loss as loss functions, and self-supervisedly learn the bag representation and the graph representation on the premise of ensuring bag-level invariance and graph-level invariance; Step 7: Use the bag representation and the graph representation learned in Step 6 for the text classification task to be classified, and perform label prediction on the text to be classified simultaneously at the coarse-grained and fine-grained levels to achieve text classification.

2. The text classification method based on self-supervised dual-granularity multi-graph learning according to claim 1, characterized in that: The specific method of Step 2 is as follows: Extract the relevance between keywords in the text; then construct a graph with the keywords of each text as nodes and the relevance between keywords as the edge weight values, and remove the edges with weights less than the given threshold based on the threshold, and set the edge weight values greater than or equal to the given threshold to 1 to form an undirected graph; represent each text in the original text dataset as a multi-graph structure, and form a graph package B = {g1, g2, …, g n}, where g1, g2, …, g n represent multiple graphs composed of the selected texts.

3. The text classification method based on self-supervised dual-granularity multi-graph learning according to claim 2, wherein: The specific method of Step 3 is as follows: Perform two data augmentation operations on all the graph data in the multi-graph dataset using the augmentation encoder. Since each graph in the graph bag is augmented twice, two sets of augmented multi-graph bags will be obtained respectively. The augmentation encoder for augmenting the graph data is as shown in the following formula: f aug (G) = {V, ε; ∈} (1) Among them, f aug represents the enhanced encoder, G = (V, ε) represents a graph, V is the vertex set, and v p ∈ V, which contains the attribute information of each node; ε is the edge set, (v p , v q ) ∈ ε, p ≠ q; ∈ is the enhancement method; the enhanced graph is represented as The graph package is represented as: Among them, represents the increased map package; In order to effectively retain the structural information of the graph data in the multi-graph dataset, use the graph encoder to learn the augmented graph representation as shown in the following formula: Among them, is the graph representation generated by the graph encoder, which is in the form of a vector and contains all the node attribute information and the internal structural relationships in the enhanced graph; H l represents the enhanced graph is the node representation of all nodes in the enhanced graph at the l-th layer of the neural network. The node representation of each layer is updated through the function f genc based on the node representation of the previous layer. Initially, H 0 is the set of attributes of all nodes in the enhanced graph, Λ is the number of layers of the neural network to be learned; f genc is a learnable function used to update the node representation of each layer. Among them, W l is the learnable weight matrix. represents the adjacency matrix after adding self-connections. A is the adjacency matrix, I is the identity matrix. is obtained by The degree matrix; f p is the pooling function, which is used to obtain the graph representation by taking the mean of the learned node representations.

4. The text classification method based on self-supervised dual-granularity multi-graph learning according to claim 3, characterized in that: The specific method of Step 4 is as follows: Apply the multi-head self-attention mechanism to the augmented graph representations in the bag to obtain the graph representation containing context information. The formula of the multi-head self-attention mechanism is as shown in the following formula: where W′ is a learnable parameter, head m represents the result of the m-th self-attention head, and m represents the number of defined self-attention heads; head m is expressed as: Among them, represents the learnable parameters in the m-th head, and d k is the hidden layer dimension; It is formed by splicing the diagrams in the diagram package, representing the package containing context information learned after the multi-head self-attention mechanism, and is composed of multiple diagrams containing context information after the multi-head self-attention mechanism formed by splicing indicating splicing, where n represents the number of diagrams in the diagram package 5. The text classification method based on self-supervised dual-granularity multi-graph learning according to claim 4, wherein: Step 5 uses the bag encoder and the graph encoder to simultaneously learn the bag representation and the graph representation, that is, the bag-level graph and the nodes in the bag-level graph are learned simultaneously, forming a graph-graph learning mechanism. This learning mechanism can effectively learn the bag representation and the graph representation, and at the same time retain the context information and global structural relationship between the graphs in the bag, which is beneficial to the coarse-grained and fine-grained classification tasks of the multi-graph learning problem.

6. The text classification method based on self-supervised dual-granularity multi-graph learning according to claim 5, wherein: The bag-level contrast loss and the graph-level contrast loss in Step 6 are as shown in the following formula: Among them, I g = {1...2n}, where n represents the number of graphs in the dataset, and I b = {1...2N}, where N represents the number of packages in the dataset; A g (i) = I g \{i}, and A b (i) = I b \{i}; is a graph contrast loss function for ensuring graph-level invariance, is a package contrast loss function for ensuring package-level invariance; sim(·) and simb(·) are functions for measuring the similarity between two representations, which can be expressed as sim(z1, z2) = exp(z1·z1 / τ) and simb(Z1, Z2) = exp(cos(Z1, Z2) / τ) respectively, where τ is a temperature parameter and cos(Z1, Z2) is a cosine function, f proj is a projection network f proj (x) = σ(f norm (ωx + b)); The loss function of the text classification method based on self-supervised dual-granularity multi-graph learning is expressed as 7. The text classification method based on self-supervised dual-granularity multi-graph learning according to claim 2, wherein: In Step 2, the E-FCM algorithm is used to extract the correlation between the keywords in the text.

Citation Information

Patent Citations

  • Graph neural network travel package recommendation method based on multi-task self-encoding

    CN115017405A

  • Method and apparatus for generating inferential question on basis of low labeled resource

    WO2022036616A1