Multi-modal data-oriented graph neural network classification method

Through the adaptive edge generation mechanism and GraphSAGE graph neural network, multimodal data correlation is dynamically captured and semantic and attribute graphs are constructed, which solves the problems of information interference and insufficient graph structure expression in multimodal data processing, and realizes efficient multi-task classification.

CN120296595APending Publication Date: 2025-07-11KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510359472.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has information interference problems in multimodal data processing, the graph structure expression ability is insufficient, and the computing efficiency and model robustness need to be improved, making it difficult to effectively process large-scale multimodal data and multitasking classification.

Method used

The graph structure is constructed through an adaptive edge generation mechanism, dynamically capture the correlation between different modal data, and a pre-trained language model is used to convert text and image data into embedded vectors, construct semantic and attribute graphs, and train them in combination with GraphSAGE graph neural network to optimize node and edge classification losses.

Benefits of technology

It significantly improves the accuracy and robustness of multimodal data fusion, improves the classification accuracy and scalability of the model, adapts to the needs of large-scale multimodal data processing, and improves the overall performance of multitasking classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296595A_ABST
    Figure CN120296595A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to a multi-modal data-oriented graph neural network classification method, which comprises the following steps of: forming a cross-modal relation graph by multi-modal data by utilizing an attribute relation and a semantic relation of the multi-modal data, and jointly learning multi-modal node representation through a multi-layer graph neural network to obtain a multi-modal relation graph; and a model is trained by utilizing multi-modal node classification and node pair classification loss. According to the method, the incidence relation between the samples is effectively utilized through the graph neural network, the problem that modal information in the data samples is lost is relieved, and compared with an existing multi-modal data classification model, the classification performance is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a classification method of a graph neural network for multi-modal data. Background Art

[0002] With the rapid development of information technology, multi-modal data has been widely used in many fields. Multi-modal data refers to data that contains various different forms of information, such as images, texts, audios, videos, etc. The diversity of this data brings new opportunities for information processing, but also brings many challenges. There are significant differences in feature representation, data scale, and semantic information among different modal data, resulting in the difficulty of effectively processing multi-modal data with traditional single-modal processing methods. In addition, the fusion of multi-modal data needs to consider the correlation and complementarity between different modalities, and existing fusion methods often have problems such as information interference and low efficiency.

[0003] As a powerful deep learning model, the graph neural network (GNN) can effectively process graph-structured data and capture the complex relationships between nodes. It shows unique advantages in multi-modal data processing and can model the associations between different modalities through graph structures to achieve efficient fusion and interaction of information. In intelligent traffic monitoring, the graph neural network can fuse various modal data such as images, point clouds, and sounds, thereby improving the system's adaptability to complex traffic environments.

[0004] However, despite the great potential of the graph neural network in multi-modal data processing, there are still some deficiencies in the existing technologies. During the multi-modal feature fusion process, the problem of information interference may lead to poor fusion effects, thereby affecting the performance of the classification model. In addition, when dealing with large-scale multi-modal data, the computational efficiency and model robustness of existing methods still need to be improved. Especially for the construction of graphs and the generation of edges, existing technologies often lack effective mechanisms to dynamically capture the complex relationships between modalities, resulting in insufficient expressive power of graph structures. At the same time, in the multi-task classification scenario, how to make full use of the advantages of graph structures to simultaneously process multiple tasks is also an urgent problem to be solved.

[0005] To overcome the limitations of the prior art, the present invention proposes a graph neural network classification model for multi-modal data. Through innovative graph structure design and fusion mechanisms, this model can effectively solve the information interference problem in multi-modal data fusion and significantly improve the accuracy and robustness of classification. Specifically, in the process of graph construction, the present invention dynamically captures the correlation between different modal data through an adaptive edge generation mechanism, thereby constructing a graph structure with stronger expressive power. In addition, the present invention also proposes a multi-task classification framework that can simultaneously process multiple classification tasks, make full use of the information interaction in the graph structure, and improve the overall performance of the model. At the same time, the present invention optimizes the model efficiency and scalability, enabling it to better meet the processing requirements of large-scale multi-modal data. Summary of the Invention

[0006] The object of the present invention is to provide a classification method for a graph neural network for multi-modal data, which can efficiently fuse multi-modal data and jointly optimize multi-task classification by constructing a graph structure through node pairs, significantly improving the classification accuracy and robustness of the model.

[0007] To achieve the above technical objectives and reach the above technical effects, the present invention is realized through the following technical solutions:

[0008] A classification method for a graph neural network for multi-modal data, comprising the following steps:

[0009] S1: Using a pre-trained language model, convert multi-modal data into embedding vectors, while capturing text semantic information and visual information of images for subsequent graph structure construction;

[0010] S2: Use the embedding vectors of text and images as node features of the graph;

[0011] S3: Construct a semantic graph by calculating the semantic similarity of multi-modal node embedding vectors, including text-text, image-image, and text-image similarities. When the similarity of a node pair exceeds the threshold, generate cross-modal semantic edges. At the same time, based on the inherent properties of multi-modal data, use multi-modal nodes of the same sample or text keyword co-occurrence nodes to construct attribute edges to form an attribute graph. Finally, merge the semantic graph and the attribute graph into a unified graph structure, where all nodes and edges share the same type and feature space, providing a structured input for the graph neural network;

[0012] S4: Input the graph constructed in step S3 into the graph neural network, train the graph structure data, and simultaneously optimize the node classification loss and the edge classification loss;

[0013] S5: Evaluate the model performance on the test set and output a classification report and a confusion matrix.

[0014] Advantages of the present invention:

[0015] In the present invention, a pre-trained language model (AltClip) is used to convert text and image data into embedding vectors, achieving efficient fusion of multi-modal information. The text data is converted into semantic vectors through the text encoder of AltClip, and the image data is converted into visual vectors through the image encoder. Both are mapped into the same high-dimensional vector space, and L2 normalization is performed to ensure the unity of the vector lengths. The design of the AltClip model enables the embedding vectors of text and image to be represented in the same feature space, thus achieving deep fusion of cross-modal information. This unified representation method not only retains the semantic information of the text and the visual features of the image, but also enables direct comparison and interaction between different modal data. The calculation of text-image similarity is achieved through dot product or cosine similarity, providing a basis for the construction of subsequent semantic graphs. In addition, text and image describe the same object from semantic and visual perspectives respectively, and the combination of the two can make up for the deficiencies of single-modal information. The text can provide detailed context information, while the image can intuitively display visual features. The fusion of the two enables the model to understand the data more comprehensively. The calculation and storage of the embedding vectors are saved through dictionaries and disks, avoiding repeated calculations, especially suitable for the processing of large-scale multi-modal data, significantly improving the calculation efficiency, and providing efficient data support for the construction of subsequent graph structures and graph neural networks.

[0016] In the present invention, through the construction of semantic graphs and attribute graphs, accurate construction and information enrichment of the graph structure are achieved. The semantic graph is generated by calculating the similarity of node embedding vectors (such as text-image, text-text, image-image similarity), and the attribute graph is generated based on node attributes (such as sample attribution, hash tags, entity co-occurrence relationships). Finally, the two are merged into a unified graph structure. By setting a similarity threshold (such as τ = 0.8), cross-modal semantic edges are generated, which can accurately capture the semantic correlations between text-text, image-image, and text-image. When the similarity between the text description and the image content exceeds the threshold, cross-modal semantic edges are generated, which provides rich semantic relationship information for the graph neural network. Attribute edges are constructed based on node attributes, further enriching the information content of the graph structure. Attribute edges are generated between multi-modal nodes of the same sample, which can reflect the relevance of the same object in different modalities; the introduction of hash tags and entity co-occurrence relationships adds more diverse information dimensions to the graph structure. Merging the semantic graph and the attribute graph into a unified graph structure enables all nodes and edges to share the same type and feature space. This unified design not only simplifies the input complexity of the graph neural network, but also improves the interpretability and operability of the model. In the GraphSAGE graph neural network, the unified graph structure can support efficient node feature update and neighbor aggregation operations.

[0017] The present invention uses the GraphSAGE graph neural network to train graph-structured data. Through sampling and aggregation of neighbor nodes, node features are dynamically updated, and node classification and edge classification tasks are optimized through the cross-entropy loss function. GraphSAGE can dynamically update node features by sampling and aggregating neighbor nodes, thereby capturing local and global graph structure information. For a certain node, the features of its neighbor nodes are integrated through an aggregation function (such as mean aggregation, maximum aggregation), and combined with the features of the current node for non-linear transformation to generate updated node features. This enables the model to better adapt to the non-linear relationships of complex multi-modal data. By optimizing the node classification loss and the edge classification loss simultaneously, the model can not only improve the accuracy of node classification but also enhance the ability to capture the relationships between nodes. In the node classification task, the model classifies by predicting node labels; in the edge classification task, the model classifies by predicting the category of node pairs. By fully utilizing the information in the graph structure, the overall performance is significantly improved. Experimental results show that the present invention performs excellently in both node classification and multi-modal node pair classification tasks. In the node classification task, the overall accuracy of the model reaches 0.8571, and the F1 score is at a relatively high level in multiple categories; in the multi-modal node pair classification task, the overall accuracy of the model is as high as 0.9131, and the F1 score reaches 0.9455 in the "infrastructure and utility damage" category. This proves the strong performance and generalization ability of the model in classification tasks.

[0018] The present invention has been experimentally verified on the CrisisMMD dataset, evaluating the performance of the model on node classification and edge classification tasks, and comprehensively evaluated through precision, recall, F1-score, and confusion matrix. The CrisisMMD dataset covers multimodal data of various natural disaster events, with high representativeness and challenges. The experimental results show that the model performs excellently on this dataset, demonstrating its feasibility and effectiveness in practical applications. In the humanitarian classification task, the model can accurately classify categories such as affected individuals and infrastructure damage, providing precise information support for humanitarian relief work. The model shows high accuracy and F1-score on the test set, and the confusion matrix also shows that the model can correctly classify most categories, with only a few samples misclassified. This indicates that the model has strong generalization ability and good category discrimination ability on unseen data. In the "non-humanitarian" category, the F1-score of the model reaches 0.8844, which can effectively distinguish informative and non-informative data, providing strong support for the screening and decision-making of disaster information. Compared with the baseline model, the technical solution of the present invention has significant improvements in classification accuracy and F1-score. In the humanitarian classification task, the accuracy and F1-score of the baseline model are only 78.4% and 78.3% respectively, while the accuracy and F1-score of the present solution reach 0.8571 and 0.8571 respectively. This further proves the technical advantages of the present solution in practical applications.

[0019] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is an overall framework diagram of a classification model of a graph neural network for multimodal data provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] Embodiment 1

[0024] A classification method of a graph neural network for multi-modal data according to this embodiment includes the following steps:

[0025] S1: Using a pre-trained language model, convert multi-modal data into embedding vectors, while capturing text semantic information and visual information of images for subsequent construction of graph structures;

[0026] S2: Use the embedding vectors of text and images as node features of the graph;

[0027] S3: Construct a semantic graph by calculating the semantic similarity of multi-modal node embedding vectors, including text-text, image-image, and text-image similarities. When the similarity of a node pair exceeds the threshold, generate cross-modal semantic edges. At the same time, based on the inherent properties of multi-modal data, use multi-modal nodes or text keyword co-occurrence nodes of the same sample to construct attribute edges to form an attribute graph. Finally, merge the semantic graph and the attribute graph into a unified graph structure, where all nodes and edges share the same type and feature space, providing a structured input for the graph neural network;

[0028] S4: Input the graph constructed in step S3 into the graph neural network, train the graph structure data, and optimize the node classification loss and edge classification loss at the same time;

[0029] S5: Evaluate the model performance on the test set and output a classification report and a confusion matrix.

[0030] In this embodiment, the specific content of S1 includes:

[0031] For text data, first preprocess it to remove URL links, non-alphanumeric characters, and extra spaces. The preprocessed text sequence is denoted as T. Subsequently, use the text encoder in the AltClip model to encode the sequence T and convert it into an embedding vector. This process is achieved through the internal mechanism of the text encoder, which realizes the conversion of the semantic information of the text into a vector representation in a high-dimensional space. The generated text embedding vector will be stored in a dictionary for subsequent use.

[0032] For image data, first load the image file and convert it to the RGB format to ensure the unity and compatibility of the image data. The input image is denoted as I, and then use the image encoder in the AltClip model to encode the image I and convert it into an embedding vector. The image encoder can extract the visual features in the image and convert them into a vector representation in a high-dimensional space. The generated image embedding vector is also stored in the above dictionary, jointly constituting the embedding representation of multi-modal data with the text embedding vector.

[0033] To improve computational efficiency and avoid repeated calculations on the same data, the generated embedding vectors are not only stored in a dictionary but also saved to disk. In subsequent processing, the pre-computed embedding vectors can be directly loaded from disk without having to perform the encoding calculations again.

[0034] In this embodiment, S2 specifically includes:

[0035] When constructing the node feature representations of the graph neural network, the embedding vectors of each node are designed to have the same dimension D. These embedding vectors, serving as the feature representations of the nodes, are systematically organized into a matrix X. Each row of matrix X corresponds to the embedding vector of a node, so its shape is N×D, where N represents the total number of nodes and D is the dimension of the embedding vector.

[0036] Taking the structure of matrix X as an example, its first row corresponds to the embedding vector of the first node (which can be text or image), the second row corresponds to the embedding vector of the second node, and so on. Through this matrix-based organization method, the features of all nodes are efficiently integrated to form a compact and easily manipulable feature matrix.

[0037] In this embodiment, S3 specifically includes:

[0038] When constructing the topological structure of the graph neural network, the generation of cross-modal semantic edges depends on the similarity measurement between node feature vectors. First, perform L2 normalization on the feature vectors of all nodes to ensure that their lengths are uniformly 1. Subsequently, generate a similarity matrix by calculating the dot product between the normalized feature vectors to quantify the semantic correlation between nodes. Based on the set similarity threshold τ = 0.8, when the similarity between any two nodes exceeds this threshold, an edge will be generated between them. Finally, by extracting the upper triangular part of the similarity matrix, the set of edges that meet the conditions is determined, thus realizing the construction of cross-modal semantic edges and further constructing the semantic graph.

[0039] The construction of attribute edges is based on the metadata information of the nodes, covering sample attribution, hash tags, and entity co-occurrence relationships. For each sample m in the sample set M, an edge is generated between any two nodes in the node set N m contained in it; for each label t in the hash tag set T, an edge is generated between any two nodes in the corresponding node set N t ; for each entity e in the entity set E, an edge is generated between any two nodes in the corresponding node set N e to construct the attribute graph. By organically combining the cross-modal semantic graph and the attribute graph, a complete graph topological structure is constructed.

[0040] In this embodiment, S4 specifically includes:

[0041] We use the GraphSAGE graph neural network to process graph-structured data. First, we sample and aggregate the features of neighboring nodes, and then perform a non-linear transformation in combination with the features of the current node to dynamically update the features of the current node. For node v, its feature update formula is:

[0042]

[0043] For the node classification task, we use the cross-entropy loss function to measure the difference between the prediction result and the true label; the edge classification task also uses the cross-entropy loss function and integrates the node classification loss and the edge classification loss by weighted summation to obtain the total loss function.

[0044] Finally, we use a fully connected layer to map the finally obtained node feature representation to the class scores, and then use the Softmax function to convert these scores into a probability distribution, making the sum of the probabilities of each class equal to 1. That is:

[0045]

[0046] In this embodiment, the S5 specifically includes:

[0047] When evaluating the model performance on the test set, we first load the model onto the specified device and switch to the evaluation mode, and then perform forward propagation on the test set data to generate prediction results. For the node classification and edge classification tasks, we calculate the precision, recall, and F1 score for each class respectively, and generate a classification report to comprehensively evaluate the model performance. At the same time, we visually display the relationship between the model prediction results and the true labels through a confusion matrix, and analyze the performance of the model on different classes.

[0048] Embodiment 2

[0049] As described in this embodiment, a classification method for a graph neural network for multi-modal data, as Figure 1 shown, includes the following steps:

[0050] S1: Using a pre-trained language model, convert the pre-processed multi-modal data into embedding vectors, while capturing both the text semantic information and the visual information of the images, for subsequent graph structure construction.

[0051] S1.1: We perform data preprocessing on the original multi-modal data, use the AltClip pre-trained language model to convert the processed text sequences and RGB format files into embedding vectors, and the generated text and image embedding vectors are integrated into a dictionary and finally saved to disk to avoid repeated calculations.

[0052] S2: Use the embedding vectors of the text and images as the node features of the graph, and organize all the embedding vectors into an N×D tensor.

[0053] S3: Determine the connection relationships between nodes by calculating the similarities between the embedding vectors. A cross-modal semantic edge is generated between node pairs whose similarity exceeds the threshold, thus constructing a semantic graph; use the specific attributes of the nodes to define the relationships between nodes to construct attribute edges, thus constructing an attribute graph; merge the semantic graph and the attribute graph into a unified graph structure, where all nodes and edges share the same type and feature space, providing a structured input for the graph neural network.

[0054] S3.1: We calculate the cosine similarity between the embedding vectors to determine which nodes have cross-modal semantic edges, generate attribute edges between nodes according to the similarity threshold τ = 0.8, and finally generate the edge indices.

[0055] S4: Feed the homogeneous graph constructed in S3 into the graph neural network, train the graph structure data, optimize the node classification loss and the edge classification loss simultaneously, and finally use the classification layer to output the results.

[0056] S4.1: In the node pair classification task, the features of each pair of nodes are represented by concatenating the features of its source node and target node, and the feature representations of the nodes and edges are allowed to share the learned features. This representation method can capture the context information of the edges. We use the output of the fully connected layer for classification to predict the category of each node pair.

[0057] S4.2: For the node classification task, we directly use the feature representation and label of the node, use the output of the fully connected layer for classification, and predict the category of each node.

[0058] S5: We set the maximum number of training epochs to 200 and use an early stopping mechanism, that is, during the training process, if the validation loss (eval_loss) does not improve within 3 consecutive epochs, the training will stop early. After the training ends, load the best model (the model with the smallest validation loss). When evaluating the model performance on the test set, the model is first loaded onto the specified device and switched to the evaluation mode, and then the test set data is propagated forward to generate the prediction results. For the node classification and edge classification tasks, calculate the precision, recall, and F1-score respectively, and generate a classification report. At the same time, intuitively display the relationship between the prediction results and the true labels through the confusion matrix, and analyze the performance of the model on different categories.

[0059] In the experiment, we used the CrisisMMD dataset (version v1), which is a multimodal dataset containing tweets and related images collected from Twitter. The data is sourced from seven natural disaster events that occurred in 2017. The dataset is annotated with three tasks: informativeness classification (judging whether a tweet or image is useful for humanitarian aid), humanitarian classification (classifying tweets or images into eight humanitarian categories, such as injured or dead people, infrastructure damage, etc.), and damage severity classification (only applicable to images). In the baseline experiment, only the first two tasks were used, and data with inconsistent text and image labels was filtered to ensure consistent labels for each tweet-image pair. Data statistics show that there are 11,400 texts and 12,708 images in the informativeness classification task, among which there are 7,632 informative tweets and 8,431 informative images respectively; there are 7,216 texts and 8,079 images in the humanitarian classification task, and the categories include affected individuals, rescue work, infrastructure damage, etc.

[0060] In the baseline model, the multimodal model used concatenates the feature vectors of text and image to form a shared representation layer and performs classification through Softmax. During the training process, the Adam optimizer is used for the text modality with a learning rate of 0.01 and a maximum number of training epochs of 50; the Adam optimizer is used for the image modality with an initial learning rate of 10-6 and a maximum number of training epochs of 1000, and no hyperparameter tuning is performed on the model.

[0061] The experimental results of the baseline model show that in the humanitarian classification task, the accuracy and F1 score of the multimodal model are only 78.4% and 78.3% respectively, and the number of misclassifications is relatively large. The confusion matrix is as follows:

[0062]

[0063] In the present invention, we also use the CrisisMMD dataset (version v1) to implement the node classification and multimodal node pair classification tasks. Below, we will show the experimental results of the model with the best performance:

[0064] Node classification:

[0065]

[0066] Multimodal node pair classification:

[0067]

[0068]

[0069] Confusion matrix:

[0070]

[0071] From the above output results, we can clearly and intuitively see that the model performs excellently in both node classification and edge classification tasks, demonstrating strong performance and generalization ability. In the node classification task, the overall accuracy of the model reaches 0.8571, and the F1 scores of most categories are relatively high. Especially in the "not_humanitarian" category, the F1 score reaches 0.8844, indicating that the model can effectively distinguish different types of nodes. In the multi-modal node pair classification task, the model's performance is even more prominent, with an overall accuracy as high as 0.9131, and the F1 scores of all categories are at a relatively high level. Among them, the F1 scores of the "infrastructure_and_utility_damage" category and the "not_humanitarian" category reach 0.9455 and 0.9246 respectively, indicating that the model can accurately capture the relationships between nodes. In addition, the model shows high accuracy and F1 scores on the test set, and the confusion matrix also shows that the model can correctly classify most categories, with only a few samples misclassified, which further proves that the model has strong generalization ability and good category discrimination ability on unseen data.

[0072] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A classification method for graph neural networks for multi-modal data, characterized in that, It includes the following steps: S1: Utilize a pre-trained language model to convert multi-modal data into embedding vectors, capturing both text semantic information and visual information of images for subsequent graph structure construction; S2: Use the embedding vectors of text and images as node features of the graph; S3: Construct a semantic graph by calculating the semantic similarity of multi-modal node embedding vectors, including text-text, image-image, and text-image similarities; when the similarity of a node pair exceeds the threshold, generate cross-modal semantic edges; meanwhile, based on the inherent properties of multi-modal data, construct attribute edges using multi-modal nodes of the same sample or co-occurring nodes of text keywords to form an attribute graph; finally, merge the semantic graph and the attribute graph into a unified graph structure, where all nodes and edges share the same type and feature space, providing structured input for the graph neural network; S4: Input the graph constructed in step S3 into the graph neural network to train the graph structure data while optimizing the node classification loss and the edge classification loss; S5: Evaluate the model performance on the test set and output a classification report and a confusion matrix.

2. The classification method of the graph neural network for multi-modal data according to claim 1, characterized in that: The specific steps of step S1 include the following sub-steps: S1.1: Process the text data, remove URLs and non-alphanumeric characters, and use a multi-modal prediction training language encoder to convert the cleaned text into an embedding vector. The generation of the embedding vector can be expressed as: e T = f T (T); S1.2: Load the image and convert it to the RGB format. Use the multi-modal prediction training language encoder to convert the image into an embedding vector, and the generation of its embedding vector can be expressed as: e I = f I (I); S1.3: The generated embedding vectors will be stored in a dictionary and finally saved to disk to avoid repeated calculations; Among them, T is the preprocessed text sequence, and f T (T) is the text encoder, and e T is the embedding vector of the text; I is the input image, and f I (I) is the image encoder, and e I is the embedding vector of the image.

3. The classification method of the graph neural network for multi-modal data according to claim 1, wherein: The specific steps of step S2 include: regarding the embedding vectors of text and images as N nodes, and their embedding vectors are E1, E2, ..., E N ; the node feature matrix X is an N×D matrix, where D is the dimension of the embedding vector; the generation formula of the node feature matrix is: X = [E1, E2, ..., E N T .​ 4. The classification method of the graph neural network for multimodal data according to claim 1, wherein: The steps of step S3 include constructing cross-modal semantic edges and constructing attribute edges, specifically including the following sub-steps: S3.1: Normalize the node features: S3.2: Calculate the similarity matrix S between the normalized node features: S3.3: According to the set similarity threshold τ, determine which node pairs have edges; if the similarity between two nodes is greater than τ, generate an edge between these two nodes; that is: E = triu(S > τ, diagonal = 1); S3.4: Extract the indices of the edges from the Boolean matrix E, which indicate which nodes have edges between them; i.e., Edge index = nonzero(E) T ; S3.5: When each sample contains multiple nodes, edges are generated for the nodes in the same sample. Edges from the same sample are represented as: S3.6: Generate edges for nodes with the same hash label based on the hash label information of each node, expressed as: S3.7: Capture the co-occurrence relationships of these entity information among different nodes according to the entity information of each node, and generate edges for them, expressed as: Among them, the normalization function uses L2 normalization, triu is the upper triangular function, the similarity threshold τ is explicitly defined as 0.8 in this paper, diagonal = 1 indicates calculation starting from the main diagonal, M is the message set, N m is the set of nodes in message m, T is the set of hash tags, N t is the set of nodes with hash tag t, E is the set of entities, N e is the set of nodes with entity e.

5. The classification method of the graph neural network for multi-modal data according to claim 1, characterized in that: The specific steps of step S4 include the following sub-steps: S4.1: The graph neural network used is GraphSage, which updates the features of the current node by aggregating the features of neighboring nodes: For the node classification task, the cross-entropy loss function is used: The node pair classification loss also uses the cross-entropy loss function, i.e.: S4.2: Weighted sum of the node classification loss and the node pair classification loss to obtain the total loss; where σ is a non - linear activation function, is the feature representation of node v at the l - th layer, N(v) is the set of neighbor nodes of node v, and W (l) is a learnable weight matrix; υ is the set of all nodes in the graph, C is the number of classes, and y vc is the true label of node v. If node v belongs to class c, then y vc = 1, otherwise y vc = 0; is the probability that the model predicts node v belongs to class c; ε is the set of all node pairs in the graph, C is the number of classes, and y uv,c is the true label of the node pair (u, v). If the node pair (u, v) belongs to class c, then y uv,c = 1, otherwise y uv,c = 0; is the probability that the model predicts the node pair (u, v) belongs to class c; S4.3: Use a fully connected layer to map the finally obtained node feature representation h v (L) to class scores, and then convert these scores into a probability distribution through the Softmax function, such that the sum of probabilities for each class is 1; that is: where z v is the class score vector of node v, W (C) is the weight matrix of the classification layer, b (C) is the bias term of the classification layer, is the feature representation of node v in the last layer of the GNN; is the probability that node v belongs to class c, z vc is the score of node v for class c, C is the number of classes, z vk is the score of node v for the k-th class.

Citation Information

Cited By

  • Multi-modal content conflict resolution method based on modal time sequence dependence modeling

    CN121116225A

  • Multi-modal fusion analysis-based bid and string bid duplicate checking detection method and system

    CN121686503A