Graph neural network embedding method combining heterogeneous graph and multi-modal data, medium and equipment
By introducing a modal adaptive fusion mechanism and differentiated edge weight allocation strategy into graph neural networks, the problem of information integration imbalance and aggregation strategy in heterogeneous graph multimodal data processing is solved, and the model's processing capability of multimodal data and the performance of downstream tasks is improved.
Patent Information
- Application Number
- CN202510679749.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the existing technology deals with heterogeneous graph multimodal data, the lack of a dynamic weight allocation mechanism leads to imbalance in information integration between modals, the singleization of aggregation strategies caused by heterogeneous edge type differences leads to inefficient information fusion, and the limited inference accuracy caused by insufficient modeling of multimodal global associations in graph neural networks.
A graph neural network embedding method combining heterogeneous graphs and multimodal data is proposed. Through the modal adaptive fusion mechanism, multimodal data is mapped to the same potential space, dynamically adjust the weights of different modes, and differentiated weighting functions are defined according to heterogeneous edge types, and node embedding is updated through multi-layer graph neural network propagation.
Adaptive fusion of multimodal data is realized, the model captures dynamic correlation relationships between modes is improved, the model model modeling ability of complex heterogeneous relationships is enhanced, and the downstream task performance is improved.
Smart Images

Figure CN120197646A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-modal graph machine learning, and particularly relates to a graph neural network embedding method combining heterogeneous graphs and multi-modal data, which is particularly applicable to recommendation system scenarios such as e-commerce, video platforms, and social media content recommendations, as well as other scenarios that require multi-modal information processing. Background Art
[0002] With the rapid development of information technology, the diversity and complexity of data have increased significantly. Especially driven by big data and artificial intelligence technologies, multi-modal data such as text, images, and audio have become key information carriers. In practical applications, data usually presents as a heterogeneous graph structure containing multiple types of nodes, edges, and multi-modal features, which makes how to efficiently process such complex data structures an important challenge in the field of machine learning, especially more prominent in recommendation system scenarios such as e-commerce, video platforms, and social media content recommendations.
[0003] With the explosive growth of multi-source heterogeneous data, the existing graph neural network methods have the following technical problems: (1) Lack of adaptive processing ability for dynamically evolving multi-modal data, and it is difficult to capture the dynamically changing correlation relationships between modalities in real time; (2) The existing methods fuse multi-modal features too simply and cannot make full use of the complementary information between different modalities, resulting in unsatisfactory fusion effects; (3) When dealing with heterogeneous graph structures, there is a lack of differential weight assignment strategies for different types of edges, which affects the model's ability to model complex relationships.
[0004] The above problems seriously restrict the application effects of the existing methods in complex multi-modal scenarios. Summary of the Invention
[0005] The following problems still need to be solved urgently in the prior art: The lack of a dynamic weight assignment mechanism for heterogeneous graph multi-modal data leads to unbalanced integration of information between modalities, the single aggregation strategy caused by differences in heterogeneous edge types results in low information fusion efficiency, and the limited reasoning accuracy caused by the insufficient modeling of multi-modal global associations by graph neural networks.
[0006] Aiming at the problems existing in the prior art, the purpose of the present invention is to propose a graph neural network embedding method combining heterogeneous graphs and multi-modal data.
[0007] The technical solution to achieve the above object of the present invention is: A graph neural network embedding method combining heterogeneous graphs and multi-modal data, comprising the following steps: Step 1: Initialize the features of nodes and edges in the heterogeneous graph. Initialize the node features using text pre-training models, image convolutional networks, or audio processing models according to the node types, and assign weights and encode features for semantic edges, image edges, or social edges according to the edge types; Step 2: Map the multi-modal data to the same latent space through a modality adaptive fusion mechanism, and dynamically adjust the weights of different modalities. The modality adaptive fusion mechanism includes constructing a shared embedding space and a multi-head attention mechanism; Step 3: Aggregate node information weighted based on heterogeneous edge types, propagate and update node embeddings through a multi-layer graph neural network, and define different weighting functions according to edge types during the aggregation process; Step 4: Apply the final node embeddings to downstream tasks and optimize the model through task-related loss functions.
[0008] Further, in the modality adaptive fusion mechanism, constructing a shared embedding space includes mapping the features of text, image, and audio modalities to a latent space of the same dimension. The multi-head attention mechanism generates dynamic weights by calculating the cosine similarity between modalities and performs weighted fusion on multi-modal features.
[0009] Further, the weight assignment and feature encoding for edges include: encoding semantic relationships for semantic edges using a text representation model, calculating association weights for image edges through feature similarity, and assigning dynamic weights for social edges based on a user behavior model.
[0010] Further, in the propagation of the multi-layer graph neural network, each layer of aggregation operation combines edge weights and node features, retains the initial features through residual connections, and updates node embeddings layer by layer using the ReLU activation function.
[0011] Further, in the multi-head attention mechanism, the calculation method of modality weights is: input the features of each modality into the self-attention layer, normalize the similarity scores through the Softmax function, generate dynamic weights, and then perform weighted summation to obtain the fused embedding.
[0012] Further, the weighted aggregation includes: using cosine similarity to weighted aggregate the neighbor node information for semantic edges, and weighted aggregating the neighbor node information for social edges based on behavior frequencies. After aggregation, update the node embeddings through a non-linear activation function.
[0013] Further, the downstream tasks include node classification or link prediction. For node classification, calculate the classification probability through the embedding vector and optimize it using cross-entropy loss. For link prediction, calculate the association score between nodes through the embedding vector and optimize it using contrastive loss.
[0014] Furthermore, the task-related loss function includes: introducing a modality alignment loss in the multi-modal fusion stage to constrain the consistency of different modality embedding spaces, and introducing an edge type discrimination loss in the heterogeneous edge aggregation stage to enhance the representation ability of edge features.
[0015] According to another aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the steps described in any one of the above methods are implemented.
[0016] According to another aspect of the present invention, there is provided an electronic device, which includes a processor, a memory, and a computer program stored on the memory, and the processor is configured to implement the steps described in any one of the above methods when executing the computer program.
[0017] The present invention proposes a graph neural network embedding method for heterogeneous graph and multi-modal data fusion. Compared with traditional methods, the present invention realizes the following technical effects through a modality adaptive fusion mechanism and a differential edge weight allocation strategy: (1) Through a dynamic weight allocation mechanism, the adaptive fusion of multi-modal data is realized, and the model's ability to capture the dynamic correlation relationship between modalities is improved; (2) The multi-head attention mechanism is adopted for modality fusion, which can adaptively identify and strengthen the contributions of key modalities while suppressing noise interference; (3) Through the edge type differentiation processing strategy, the model's ability to model complex heterogeneous relationships is enhanced, and the performance of downstream tasks is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings incorporated into the specification and constituting a part of the specification illustrate embodiments of the present invention and, together with the related written description, are used to explain the principles of the present invention. In these drawings, like reference numerals are used to represent like elements. The drawings in the following description are some embodiments of the present invention, not all embodiments. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 Shows a flowchart of a graph neural network embedding method for combining a heterogeneous graph and multi-modal data provided by an embodiment of the present invention.
[0020] Figure 2 Shows a block diagram of a computer device according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined arbitrarily with each other.
[0022] An embodiment of the present invention provides a graph neural network embedding method that combines heterogeneous graphs and multimodal data. The execution steps can be more intuitively understood through the Figure 1 flowchart of the graph neural network embedding method that combines heterogeneous graphs and multimodal data provided in the appendix, including: S1: Initialize the features of the nodes and edges in the heterogeneous graph. According to the node types, use text pre-training models, image convolutional networks, or audio processing models to initialize the node features respectively, and assign weights and perform feature encoding on semantic edges, image edges, or social edges according to the edge types. This step constructs initial feature representations with semantic distinctiveness and modal adaptability for the nodes and edges in the heterogeneous graph, thereby providing high-quality information input for subsequent graph neural network learning.
[0023] Among them, according to the node types, a pre-training model adapted to the modality is used to generate initial feature vectors, and the specific implementation is as follows: For text nodes such as documents and comments, use a pre-trained language model (such as BERT, GPT, etc.) to convert the text data into a vector representation of a fixed dimension. For the feature initialization of text-type nodes, take the node containing the text content "A knowledge graph is a graph data structure" as an example, and use the pre-trained language model BERT-base for feature extraction. Specifically, input the text into the encoding layer of the model, and extract the last hidden state vector corresponding to the [CLS] token as the initial node feature, generating a 768-dimensional feature representation: v text = [0.245, 0.654, -0.123,..., 0.432] (a vector of length 768), and this vector will be used as the text feature embedding of the node.
[0024] For image nodes such as product pictures and user avatars, use a convolutional neural network (CNN) to extract image features and obtain the image vector of the node. For the feature initialization of image-type nodes, take the node containing a computer picture as an example, and use the pre-trained ResNet-50 model for feature extraction. Specifically, input the image into the convolutional neural network layer of the model, and after global average pooling processing, extract the 2048-dimensional feature vector output by the last fully connected layer to generate the initial embedding representation of the node: vimage =[0.521, 0.397, -0.219, …, 0.123] (a vector of length 2048), and this vector will be used as the image feature embedding of the node.
[0025] For audio nodes such as speech and audio content, use a sound processing model to extract audio features and convert them into low-dimensional vectors. For the feature initialization of audio type nodes, take the node containing an audio segment (such as a piano piece segment) as an example, and use a pre-trained WaveNet or 3D-CNN model for feature extraction, and reduce the dimension to the dimension aligned with the text features through pooling operations.
[0026] Regarding the multi-modal characteristics of the heterogeneous graph, through type-based feature initialization, the independence and compatibility of different modal data in the embedding space can be ensured. Moreover, in view of the problems such as low computational efficiency and insufficient feature expression ability in the processing of heterogeneous data in industrial scenarios, in this embodiment, the feature initialization scheme for practical applications can maintain a low computational overhead when processing a large amount of heterogeneous data through an improved modal adaptive preprocessing module. Especially applied to the product recommendation system suitable for e-commerce platforms, in the processing of tens of millions of product data per day, text features are extracted through a domain-adapted BERT variant model, and at the same time a lightweight CNN network is used to process product images, realizing a feature initialization response in milliseconds.
[0027] For each node type in the heterogeneous graph, the present invention defines an independent initial embedding space to store node features. Based on the node type and modal characteristics, independent initial embedding spaces are allocated for each node type in the heterogeneous graph to store their modal-specific feature vectors. By separating the embedding spaces of different modalities, feature confusion caused by the heterogeneity of the heterogeneous graph is avoided, and a basic data structure support is provided for subsequent modal adaptive fusion.
[0028] In one implementation, optionally, the weight assignment and feature encoding of the edges in this step include: encoding the semantic relationship of the semantic edges using a text representation model, calculating the association weight of the image edges through feature similarity, and assigning dynamic weights to the social edges based on the user behavior model.
[0029] For the feature initialization of semantic edges, take the edge connecting the "user" node and the "movie" node with the semantic label "like" as an example, and use a pre-trained Word2Vec model for feature encoding. Specifically, input the semantic label into the model to generate a 300-dimensional word vector as the edge feature representation: v semantic=[0.324, -0.512, 0.249, …, 0.675] (a vector of length 300). This vector will be used as the initial embedding feature of the semantic edge for weight calculation and relationship modeling in subsequent heterogeneous edge weighted aggregation.
[0030] Design differentiated weight assignment and feature encoding strategies for different edge types (semantic edges, image edges, social edges) to distinguish the differentiated impacts of text, image, and social relationships, so as to achieve accurate modeling of heterogeneous relationships and multi-modal information fusion, and improve the accuracy of downstream tasks.
[0031] S2: Map multi-modal data to the same latent space through a modality adaptive fusion mechanism, and dynamically adjust the weights of different modalities. The modality adaptive fusion mechanism includes constructing a shared embedding space and a multi-head attention mechanism. This step solves the semantic alignment and dynamic collaboration problems of multi-modal data in heterogeneous spaces through a unified modality adaptive fusion framework. Using the multi-head attention mechanism can not only dynamically weigh the influence of different modalities on node representations according to task requirements, but also adaptively strengthen the semantic contributions of key modalities and suppress noise interference, thereby achieving cross-modal semantic alignment and enhanced complementarity.
[0032] In one implementation, optionally, in the modality adaptive fusion mechanism of this step, constructing a shared embedding space includes mapping the features of text, image, and audio modalities to a latent space of the same dimension. The multi-head attention mechanism generates dynamic weights by calculating the cosine similarity between modalities and performs weighted fusion on multi-modal features. When processing different modality data, a method that can effectively measure the similarity degree of feature vectors is required. In this embodiment, cosine similarity is selected because it has the characteristics of high computational efficiency and numerical stability, and is especially suitable for processing high-dimensional feature vectors. In practical applications, this calculation method can accurately capture the semantic associations between different modalities while keeping the computational cost within an acceptable range.
[0033] In this embodiment, by learning a shared embedding space E shared , all modality data share a unified representation in this space, eliminating the distribution differences between modalities and also optimizing the heterogeneous node aggregation speed. Further, by using the multi-head attention mechanism to fuse different modality information, the complementary features between modalities can be effectively identified. Finally, the embedding representation v of each node i contains information from all modalities. Specifically: , where represents text features, represents image features, represents sound features.
[0034] For the feature fusion of multimodal nodes, taking the "movie" node that includes the text modality ("science fiction movie") and the image modality (movie poster) as an example, the following operations are implemented: (1) Generate text embedding vectors through the BERT model: v movie_text =[0.245, 0.654, -0.123, …, 0.432], Generate image embedding vectors through the ResNet-50 model: v movie_image =[0.521, 0.397, -0.219, …, 0.123].
[0035] (2) Generate dynamic weights by calculating the cosine similarity between modalities. Generating dynamic weights by calculating the cosine similarity between modalities can achieve cross-modal semantic alignment and enhanced noise robustness. At the same time, the complementarity of multi-granularity features is strengthened through multi-head parallel computing, ultimately improving the inference performance and generalization ability of the model in complex scenarios.
[0036] In this embodiment, the modality attention weights are calculated through the self-attention mechanism , that is: , where represents the features of the modality, represents the features of the modality, text and image represent different modalities, and Similarity(·, ·) calculates the similarity between the two modalities. In this embodiment, for the sake of simplifying the example, it is set that the calculated weights (α text = 0.6, α image = 0.4) are used as the fusion coefficients. In actual business scenarios, the contributions of different modalities to the final decision are often dynamically changing. For example, in e-commerce recommendations, for clothing products, image features may be more important than text descriptions; while for electronic products, the parameter information in the text description may be more valuable. Therefore, in this embodiment, an adaptive attention weight calculation mechanism is introduced to enable the model to automatically adjust the importance of each modality according to the specific scenario, thereby improving the adaptability of the model in complex business scenarios.
[0037] (3) Fuse the text and image modality embeddings by weighted average according to the weights. Generate a unified embedding representation: , The obtained fusion vector: , Finally, v movie_fusedAs the multi-modal joint embedding of nodes, it is input into the subsequent graph neural network for heterogeneous information propagation and downstream task inference.
[0038] In one implementation, optionally, in the multi-head attention mechanism, the modal weights are calculated as follows: Each modal feature is input into the self-attention layer, the similarity scores are normalized by the Softmax function to generate dynamic weights, and then the weighted sum is obtained to get the fused embedding. Taking the same "movie" node as an example, the cosine similarity between the text and image modalities is calculated as the attention score through the multi-head attention mechanism. The cosine similarity formula can be used to calculate their similarity: , where v1 and v2 respectively represent a modal feature.
[0039] Assume: , and then, calculate the attention weights and : , .
[0040] Furthermore, the fused embedding obtained by weighted sum based on the dynamic weights is: , S3: Based on heterogeneous edge types, weighted aggregation of node information is performed, and the node embeddings are propagated and updated through a multi-layer graph neural network. Different weighted functions are defined according to the edge types during the aggregation process.
[0041] After completing the feature initialization of nodes and edges (S1) and multi-modal fusion (S2), this step performs weighted aggregation of node information based on heterogeneous edge types and realizes iterative update of node embeddings through a multi-layer graph neural network. The specific implementation process is as follows: (1) Aggregate the neighbor information of each node. The aggregation operation not only considers the features of the nodes but also the edge type information. For the edges in the heterogeneous graph, a weighted aggregation strategy is adopted, and the contribution degree of neighbor nodes is adjusted according to the weights of different edge types.
[0042] In one implementation, optionally, the weighted aggregation includes: using cosine similarity to weighted aggregate the neighbor node information for semantic edges, and using behavior frequency to weighted aggregate the neighbor node information for social edges. After aggregation, the node embeddings are updated through a non-linear activation function to achieve adaptive optimization of the edge weights, improve the model's ability to capture dynamic relationships, enable the model to capture user interest changes in a timely manner, and effectively solve the problem that traditional static feature extraction schemes cannot adapt to the dynamic changes of the business environment.
[0043] For node i and its set of neighbor nodes N (i), the aggregation operation combines the node's own features and the edge type weights, and the calculation formula is: , Among them, is the feature of node j in the upper layer, indicating that node j belongs to the neighbor nodes of node i.
[0044] In the social network scenario, taking the information interaction between user node i and movie node j as an example, the initial feature of node i represents the preference intensity for science fiction and action movies, and the initial feature of node j reflects the distribution of science fiction and action attributes of the movie. The two nodes are connected by edge type 1 (user click behavior), and the edge weight indicates the high attention of the user to the movie.
[0045] Heterogeneous edge weighted aggregation and node embedding update: , During the aggregation process, first, the self - feature of node i is linearly transformed through the weight matrix W = [0.5, 0.5] to obtain , and at the same time, the feature of node j is scaled by the edge weight to ; then the two are added and non - linearly mapped through the ReLU activation function, and finally the updated embedding feature of node i is generated. This process shows that the user's preference for science fiction movies has increased from the initial 0.2 to 0.28, and the preference for action movies has increased from 0.4 to 0.5, intuitively reflecting the technical advantage of the graph neural network in dynamically mining user interests by weighted aggregation of neighbor information.
[0046] In another social network scenario, heterogeneous edge type differential processing is further introduced to show how user node i aggregates multi - type information and updates the embedding feature through semantic edges and social edges at the same time. The specific implementation is as follows: Semantic edge: Reflects content relevance (such as the interaction between user and movie), and uses the relevance in text representation for weighting. The "click behavior" relationship (edge type 1) between user i and movie j, with weight = 0.5 (calculated based on the text relevance between movie content and user preference); Social edge: Reflects the influence between users. Based on the social relationship of nodes, a higher weight is given to the social edge. The "follow relationship" (edge type 2) between user i and user k, with weight = 0.3 (calculated based on historical interaction frequency).
[0047] The aggregated information will be updated with different weights according to the edge type: , Known neighbor node features: , representing the attribute distribution of node j (e.g., for the movie "Science Fiction" attribute 0.4, "Action" attribute 0.6).
[0048] , representing the social influence characteristics of user k (e.g., user activity 0.7, dissemination power 0.8).
[0049] Updated characteristics of node i: , , The above heterogeneous edge aggregation results integrate content attributes (node ) and social behaviors (node ), enhancing the comprehensiveness of the embedding representation. In a multi-layer graph neural network, information updates node representations through multiple rounds of propagation. The aggregation in each layer not only considers direct neighbors but may also consider node information at further levels, gradually strengthening the global association between nodes. In one implementation, optionally, in the multi-layer graph neural network propagation in this step, the aggregation operation in each layer combines edge weights and node features, and retains the initial features through residual connections, and uses the ReLU activation function to update node embeddings layer by layer, thereby preventing the loss of feature information in the deep network and enhancing the model stability.
[0050] Specifically, the update of node embeddings in each layer: , Among them, is the weight matrix of the layer, is the weighted coefficient between nodes, which is the weighted coefficient between nodes and can also be called the edge weight value. Assume that in the aforementioned scenario, a two-layer graph neural network is used for information propagation, and the embeddings of node i and its neighbor node j are respectively: The propagation result of the first layer , The propagation result of the first layer of neighbor node j , Assume that the weight matrix of the first layer is , and the weight matrix of the second layer is . Node is updated to: , Specific calculation: , , The final embedding of node i is: 。
[0051] In this embodiment, the node embeddings are dynamically updated through the weighted aggregation and non-linear activation of two-layer graph neural networks, and the initial features are optimized through two propagation iterations into , fully demonstrating the differential edge type weights (such as edge weights ) and the synergistic effect of multi-layer information propagation (weight matrices , ) on node representation learning in the heterogeneous graph network.
[0052] In this embodiment, the update of node features is a process of progressive optimization. Considering the complexity of node relationships in practical applications, single-layer updates are difficult to fully utilize the structural information in the network. By adopting a multi-layer update mechanism, it is ensured that each layer integrates different types of information: the first layer mainly integrates the features of directly connected nodes, and the deeper layers can obtain a wider range of structured information. This hierarchical update mechanism enables the model to more comprehensively understand the role of nodes in the network, thereby providing more accurate prediction results.
[0053] S4: Apply the final node embeddings to downstream tasks and optimize the model through task-related loss functions.
[0054] By applying the node embeddings updated by the multi-layer graph neural network to downstream tasks (such as node classification, graph classification, link prediction, etc.), for example, in a recommendation system, use the final node embeddings for user and item matching scoring, calculate the prediction error based on the task results, and adjust the weight matrices and edge type weights in reverse, enabling the model to be automatically optimized from the original data to the task output throughout the process.
[0055] In one implementation, optionally, the downstream tasks in this step include node classification or link prediction. Node classification calculates the classification probability through the embedding vector and optimizes it using cross-entropy loss. Link prediction calculates the association score between nodes through the embedding vector and optimizes it using contrastive loss. For example, in the node classification task, the final node embeddings can be input into a classifier to generate the node category probability distribution (such as user interest classification, commodity type recognition). Calculate the prediction deviation through the classification error loss function (such as cross-entropy loss), and backpropagate to optimize the graph neural network weights and edge type weight parameters, enabling the embedding features to better distinguish different category nodes. The link prediction task can calculate the association degree score between nodes based on the node embeddings (such as the user-commodity interaction possibility), and optimize the model through the association degree contrastive loss function (such as margin loss), making the embedding similarity of connected node pairs higher than that of unconnected node pairs, thereby improving the accuracy of relationship prediction.
[0056] In one implementation, optionally, the task-related loss function in this step includes: introducing a modality alignment loss in the multi-modal fusion stage to constrain the consistency of different modality embedding spaces, and introducing an edge type discrimination loss in the heterogeneous edge aggregation stage to enhance the representation ability of edge features. For example, in a node classification task, taking user node i as an example, its true label is y i = 1 (science fiction enthusiast), and the model prediction probability = 0.8. Calculate using the cross-entropy loss function: , , Optimize the model through backpropagation, minimize the loss function to obtain the final node embedding.
[0057] To achieve the co-embedding of heterogeneous graphs and multi-modal data, this embodiment uses a graph neural network framework (such as PyTorch Geometric) to process the node / edge type division and topological propagation of heterogeneous graphs, and at the same time combines a deep learning framework (such as PyTorch) to extract multi-modal features such as text and images. Based on PyTorch and PyTorch Geometric, heterogeneous graph data and multi-modal data (text and images) can be fused, and graph neural network embedding can be performed. By combining heterogeneous graph data and multi-modal data (text and images), and performing node embedding and classification tasks through a graph neural network (GNN), this implementation framework enhances the feasibility of the technical solution.
[0058] Generally speaking, in the embodiments of the present invention, an innovative multi-modal adaptive fusion mechanism is used to dynamically adjust the weights of different modalities, combined with a differential weighted aggregation strategy for heterogeneous edge types, and a deep information propagation technology is adopted to enhance the global correlation modeling ability, effectively solving the problems of unbalanced modal information fusion, single-edge type processing, and insufficient long-distance node interaction modeling in traditional graph neural networks in multi-modal scenarios. This method provides an efficient technical means for processing complex heterogeneous graphs and multi-modal data, significantly improving the quality of node embedding representations, and can be widely applied to recommendation system scenarios such as e-commerce, video platforms, and social media content recommendations, thus promoting the technological development of related fields and improving the inference accuracy of practical applications. The technical solutions of the present invention have been fully verified in the actual business environment. Taking an e-commerce platform as an example, this solution effectively solves the problem of insufficient understanding of user interests by traditional recommendation systems by integrating product text and image information and user behavior sequences. In the A / B test of an e-commerce platform, the use of the calculation process implemented by the present invention improved the recommendation accuracy by more than 15% compared with the non-use. In the social media scenario, the application of the multi-modal fusion framework significantly improved the accuracy of content understanding, and the heterogeneous relationship modeling helped discover more potential user interest points, solving the problem of content homogenization in information flow recommendations. These practices show that the present invention is not only innovative at the theoretical level, but also provides a complete solution for industrial-level heterogeneous data processing.
[0059] The embodiments of the present invention also provide a readable storage medium, on which a computer program is stored, wherein when the program is executed by a processor, the steps described in any one of the above methods are implemented.
[0060] The embodiments of the present invention also provide an electronic device, as Figure 2 shown, the electronic device 200 includes a processor 201, a memory 202, and a computer program stored on the memory. The processor 201 is used to implement the steps described in any one of the above methods when executing the computer program.
[0061] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A graph neural network embedding method combining heterogeneous graphs and multi-modal data, characterized in that, It includes the following steps: Step 1: Initialize the features of nodes and edges in the heterogeneous graph. Initialize the node features using text pre-training models, image convolutional networks, or audio processing models according to the node types, and assign weights and encode features for semantic edges, image edges, or social edges according to the edge types; Step 2: Map the multi-modal data to the same latent space through a modality adaptive fusion mechanism, and dynamically adjust the weights of different modalities. The modality adaptive fusion mechanism includes constructing a shared embedding space and a multi-head attention mechanism; Step 3: Aggregate node information weighted based on heterogeneous edge types, and propagate and update node embeddings through a multi-layer graph neural network. Different weighting functions are defined according to the edge types during the aggregation process; Step 4: Apply the final node embeddings to downstream tasks, and optimize the model through task-related loss functions.
2. The method according to claim 1, wherein In the modality adaptive fusion mechanism of Step 2, constructing a shared embedding space includes mapping the features of text, image, and audio modalities to a latent space of the same dimension. The multi-head attention mechanism generates dynamic weights by calculating the cosine similarity between modalities and performs weighted fusion on the multi-modal features.
3. The method according to claim 1, wherein The weight assignment and feature encoding for edges in Step 1 include: encoding semantic relationships for semantic edges using a text representation model, calculating association weights for image edges through feature similarity, and assigning dynamic weights for social edges based on a user behavior model.
4. The method according to claim 1, characterized in that, In the propagation of the multi-layer graph neural network in Step 3, each layer of aggregation operation combines edge weights and node features, and retains the initial features through residual connections. The ReLU activation function is used to update node embeddings layer by layer.
5. The method according to claim 2, wherein In the multi-head attention mechanism, the modality weights are calculated as follows: Input the features of each modality into the self-attention layer, normalize the similarity scores through the Softmax function, generate dynamic weights, and then perform weighted summation to obtain the fused embedding.
6. The method according to claim 1, characterized in that, The weighted aggregation in Step 3 includes: using cosine similarity to weight and aggregate the neighbor node information for semantic edges, and aggregating the neighbor node information based on the behavior frequency for social edges. After aggregation, the node embeddings are updated through a non-linear activation function.
7. The method according to claim 1, wherein The downstream tasks in Step 4 include node classification or link prediction. For node classification, calculate the classification probability through the embedding vector and optimize it using cross-entropy loss. For link prediction, calculate the association score between nodes through the embedding vector and optimize it using contrastive loss.
8. The method according to claim 1, characterized in that, The task-related loss functions in Step 4 include: introducing a modality alignment loss in the multi-modal fusion stage to constrain the consistency of different modality embedding spaces, and introducing an edge type discrimination loss in the heterogeneous edge aggregation stage to enhance the representational ability of edge features.
9. A computer-readable storage medium, characterized in that, A computer program is stored, and when the computer program is executed, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-8 as described above.
Citation Information
Patent Citations
Image text matching method based on graph neural network
CN117035008A
Dynamic data pipeline construction method based on artificial intelligence and multi-modal data processing
CN119830200A
News recommendation method and system based on heterogeneous graph neural network and fusing multi-behavior and multi-modal data
CN119848329A
Heterogeneous tree graph neural network for label prediction
US20240330679A1
Cited By
Resource recall method and device, server and storage medium
CN120448643A
Cancer gene identification method based on multiplexing heterogeneous graph neural network
CN120895103A
Method for predicting curative effect of image heterogeneity region fusion technology based on graph network
CN121459054A
Power system optimization method and device based on physical constraint heterogeneous graph, medium and equipment
CN121983946A