A Multimodal Feature Fusion Method, System, Device, Medium and Program Product

The proposed method leverages graph neural networks and large language models to enhance multi-modal feature fusion by addressing redundancy and conflict in diverse data modalities, resulting in improved feature representation and task accuracy.

CN120068009BActive Publication Date: 2025-07-15SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510553653.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-15
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with redundancy and conflicts between different modes in multimodal feature fusion, resulting in inconsistency of information and inaccurate feature representation, affecting the performance of downstream tasks.

Method used

The large language model is used for semantic enhancement, and the multimodal data is characterized by aggregation of features by combining graph neural networks. By sparse and normalizing the similarity matrix, the relationship between different modal features is captured and meaningful features are extracted.

Benefits of technology

The accuracy and robustness of multimodal feature fusion is improved, and information redundancy and conflict are effectively alleviated, and the accuracy and robustness of downstream tasks are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068009B_ABST
    Figure CN120068009B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal feature fusion method, system, device, medium and program product, which relates to the technical field of feature fusion, and includes: encoding and semantic enhancement of multi-modal data to obtain multi-modal original features and multi-modal semantic features; calculating an original feature similarity matrix and a semantic feature similarity matrix, and performing sparsification processing and normalization processing; according to the normalized original feature similarity matrix and semantic feature similarity matrix, aggregating the original features and semantic features, and obtaining a product set ID embedding feature according to the interaction between the user and the product; extracting modal preferences from the product set ID embedding feature, and combining the product set original features and the product set semantic features to obtain the overall multi-modal original features, multi-modal semantic features and ID features, and splicing the three to obtain multi-modal fusion features. It solves the problems of information redundancy and conflict of multi-modal data, captures the relationship between different modal features, and improves the fusion effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of feature fusion, and particularly to a multi-modal feature fusion method, system, device, medium and program product. Background Art

[0002] Traditional feature fusion methods mostly rely on the processing of single-modal data, such as images, texts or other single types of data. These methods usually encounter problems such as data inconsistency, feature redundancy and information conflict. Especially when facing data from different sources, how to effectively integrate various modal features to achieve efficient information extraction and processing is an important challenge in the current technology.

[0003] Currently, the applications of multi-modal feature fusion have been widely penetrated into various fields, such as speech processing, medical image analysis, etc. Existing methods usually rely on pre-trained models or specific encoding networks, and by extracting features of different modalities, map them into a shared representation space.

[0004] Although these methods have achieved certain success in dealing with the complementarity of different types of data, they still face the problem of how to balance and optimize the cooperation and conflict between different modalities. Especially when there are redundant or irrelevant features, how to effectively eliminate this information and extract meaningful features is still a difficult point that needs to be solved urgently in the current technology.

[0005] For example, in the fusion of medical images and diagnostic texts, the images provide visual information of diseases, and the texts provide more detailed medical record descriptions. If redundant features cannot be accurately distinguished, it may lead to inaccurate feature representations after fusion, and further affect the performance of downstream tasks.

[0006] Similarly, in the product recommendation of e-commerce platforms, the images provide the appearance information of products, and the texts describe product attributes and user experiences. If effective feature fusion cannot be achieved, it is easy to introduce information redundancy or semantic inconsistency, thus affecting the recommendation effect.

[0007] Therefore, in the research of multi-modal feature fusion, how to enhance the semantic understanding of the fused features and optimize redundant information is one of the core issues. Summary of the Invention

[0008] To solve the above problems, the present invention proposes a multi-modal feature fusion method, system, device, medium and program product, which performs semantic enhancement on multi-modal data, deeply mines the potential semantic information in different modalities, and solves the problems of information redundancy and conflict existing in multi-modal data; uses a graph neural network to aggregate the original features and semantic features, captures the relationships between different modal features, and improves the feature fusion effect.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] In a first aspect, the present invention provides a multi-modal feature fusion method, including:

[0011] Encoding the multi-modal data of the obtained commodity to obtain multi-modal raw features, and performing modal semantic enhancement according to the prompt words of each modality to obtain multi-modal semantic features;

[0012] According to the multi-modal raw features and multi-modal semantic features, calculate the raw feature similarity matrix and semantic feature similarity matrix between different commodities in each modality, and only retain the top K similarities for each commodity, thereby obtaining a sparsified raw feature similarity matrix and semantic feature similarity matrix, and perform normalization processing on them respectively;

[0013] According to the normalized raw feature similarity matrix and semantic feature similarity matrix, aggregate the raw features and semantic features of each modality to obtain the raw features of the commodity set and the semantic features of the commodity set, and obtain the commodity set ID embedding feature according to the interaction between the user and the commodity;

[0014] Extract the modal preference from the commodity set ID embedding feature, combine the raw features of the commodity set and the semantic features of the commodity set in each modality to obtain the overall multi-modal raw features, multi-modal semantic features and ID features, and splice the three to obtain the multi-modal fusion feature, so as to predict the interaction score between the user and the commodity according to the multi-modal fusion feature and recommend commodities to the user.

[0015] As an alternative implementation, after obtaining the multi-modal raw features and multi-modal semantic features, perform linear projection on the multi-modal raw features and multi-modal semantic features to obtain the th commodity's raw projection feature and semantic projection feature in modality , and perform interaction feature purification processing on them respectively to obtain the th commodity's raw purification feature and semantic purification feature in modality , respectively:

[0016] ; ;

[0017] wherein, are all trainable parameter matrices; is a bias vector; is an element-wise product; is a Sigmoid activation function; is the Embedding of the product ID of each product.

[0018] As an alternative embodiment, the elements in the original feature similarity matrix and the semantic feature similarity matrix between different products in each modality are respectively:

[0019] ; ;

[0020] The elements in the sparsified original feature similarity matrix and the semantic feature similarity matrix are respectively:

[0021] ;

[0022] ;

[0023] Among them, is the original feature similarity between product and product in modality , is the semantic feature similarity between product and product in modality ; is the original feature embedding representation of product in modality , is the original feature embedding representation of product in modality , is the semantic feature embedding representation of product in modality , is the semantic feature embedding representation of product in modality ; is the original feature edge weight between product and product in modality ; is the semantic feature edge weight between product and product in modality ; is the original feature similarity between product and product in modality , is the semantic feature similarity between product and product in modality ; is a set of products.

[0024] As an alternative implementation, the normalized original feature similarity matrix and the semantic feature similarity matrix are respectively:

[0025] ; ;

[0026] wherein, is the sparsified original feature similarity matrix in modality ; is the sparsified semantic feature similarity matrix in modality ; is the degree matrix of , is the degree matrix of .

[0027] As an alternative implementation, the original features of the commodity set and the semantic features of the commodity set are respectively:

[0028] ; ;

[0029] ; ;

[0030] wherein, is the original feature of the commodity set obtained by aggregation through the graph neural network from the perspective of the original features in the -th commodity modality ; is the semantic feature of the commodity set obtained by aggregation through the graph neural network from the perspective of the semantic features in the -th commodity modality ; is the number of layers of the graph neural network; and are respectively the original purified features of the -th commodity modality in the -th layer and the -th layer; and are respectively the semantic purified features of the -th commodity modality in the -th layer and the -th layer; is the total number of layers;

[0031] Commodity set ID embedding feature is:

[0032] ;

[0033] ;

[0034] Among them, is the user-product interaction adjacency matrix; is the degree matrix of matrix ; and are respectively the product ID embeddings of the th layer and the th layer of the th product.

[0035] As an alternative implementation, extract the original preference and semantic preference of modality from the product set ID embedding features :

[0036] ; ;

[0037] The overall multi-modal original features and multi-modal semantic features are respectively:

[0038] ; ;

[0039] Among them, M is the total number of modality sets; is the product set original feature; is the product set semantic feature; is a trainable parameter matrix, is the bias vector; is the Sigmoid activation function;

[0040] Divide the multi-modal fusion feature into user representation and product representation , and use the inner product to determine the interaction score between the user and the product.

[0041] Second aspect, the present invention provides a multi-modal feature fusion system, including:

[0042] An enhancement module, configured to encode the obtained multi-modal data of the product to obtain multi-modal original features, and perform modality semantic enhancement according to the prompt words of each modality to obtain multi-modal semantic features;

[0043] The similarity processing module is configured to calculate the original feature similarity matrix and the semantic feature similarity matrix between different commodities in each modality according to the multimodal original features and the multimodal semantic features, and only retain the first K similarities for each commodity, thereby obtaining the sparse original feature similarity matrix and the semantic feature similarity matrix, and performing normalization processing respectively;

[0044] An aggregation module is configured to aggregate the original features and semantic features of each modality according to the normalized original feature similarity matrix and the semantic feature similarity matrix to obtain the original features of the product set and the semantic features of the product set, and obtain the embedded features of the product set ID according to the interaction between the user and the product;

[0045] The fusion module is configured to extract modal preferences from the embedded features of the product set ID, combine the original features of the product set under each modality and the semantic features of the product set to obtain the overall multimodal original features, multimodal semantic features and ID features, and concatenate the three to obtain multimodal fusion features, so as to predict the interaction score between users and products based on the multimodal fusion features and make product recommendations to users.

[0046] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in the first aspect is performed.

[0047] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, wherein when the computer instructions are executed by a processor, the method described in the first aspect is performed.

[0048] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in the first aspect.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] The present invention proposes a multimodal feature fusion method based on a large language model, introduces the semantic enhancement of the large language model into the multimodal feature fusion, and performs semantic enhancement processing of the large language model on multimodal data such as images and texts. It can deeply mine the potential semantic information in different modalities such as images and texts, improve the synergy between the modalities, enhance the feature representation capability, solve the information redundancy in the multimodal data itself and the conflict between multimodal information and behavioral information, and effectively improve the accuracy and robustness of downstream tasks.

[0051] The present invention aggregates the original features and semantic features using a graph neural network, which can better capture the relationships between different modal features, improve the feature fusion effect, and make the final multi-modal features more accurate and effective.

[0052] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0054] Figure 1 It is a flowchart of the multi-modal feature fusion method provided in Embodiment 1 of the present invention;

[0055] Figure 2 It is a schematic diagram of the multi-modal feature fusion method provided in Embodiment 1 of the present invention;

[0056] Figure 3 It is a diagram of the large language model enhancement module provided in Embodiment 1 of the present invention;

[0057] Figure 4 It is a structural diagram of the multi-modal feature fusion system provided in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] The present invention will be further described below in conjunction with the drawings and embodiments.

[0059] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0060] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those clearly listed steps or units, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0061] Under the condition of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0062] Embodiment 1

[0063] With the emergence of large language models (LLMs), they have shown excellent capabilities in processing semantic information, cross-modal learning, etc. Using large language models for semantic mining can not only further improve the quality of feature representation, but also effectively alleviate the problems of information redundancy and conflict. Through the semantic understanding and cross-modal learning capabilities of large language models, effective features can be more accurately identified and interfering features can be removed, providing a new solution for multi-modal feature fusion. Therefore, this embodiment combines the advantages of large language models and realizes more efficient feature fusion through semantic enhancement methods.

[0064] This embodiment relies on the Key Laboratory of Artificial Intelligence Application for People's Livelihood Services in Shandong Province and the Future Industry Laboratory of Shandong Province Universities for General Artificial Intelligence, and provides a multi-modal feature fusion method based on large language models, such as Figure 1 shown, including:

[0065] Encode the multi-modal data of the obtained commodities to obtain multi-modal original features, and perform modal semantic enhancement according to the prompt words of each modality to obtain multi-modal semantic features;

[0066] According to the multi-modal original features and multi-modal semantic features, calculate the original feature similarity matrix and semantic feature similarity matrix between different commodities in each modality, and only keep the top K similarities for each commodity, thereby obtaining a sparsified original feature similarity matrix and semantic feature similarity matrix, and perform normalization processing on them respectively;

[0067] According to the normalized original feature similarity matrix and semantic feature similarity matrix, aggregate the original features and semantic features of each modality to obtain the original features of the commodity set and the semantic features of the commodity set, and obtain the ID embedding features of the commodity set according to the interaction between the user and the commodity;

[0068] Extract the modal preferences from the ID embedding features of the commodity set, combine the original features of the commodity set and the semantic features of the commodity set in each modality to obtain the overall multi-modal original features, multi-modal semantic features and ID features, and splice the three to obtain multi-modal fusion features, so as to predict the interaction score between the user and the commodity according to the multi-modal fusion features and recommend commodities to the user.

[0069] The following combines Figure 1 - Figure 2 to elaborate on the method of this embodiment in detail.

[0070] S1: Obtain The original multimodal data of the commodity is preprocessed to ensure that the formats of different modal data are consistent and have processing capabilities.

[0071] Among them, the multimodal data includes, but is not limited to, different types of data such as images, texts, and audios. In this embodiment, the image modality and the text modality are taken as examples, and the expansion of more modal data is realized based on this.

[0072] Among them, the preprocessing includes data cleaning and denoising, removing error items and noises in the data; for image modality data, a Gaussian filter is used to remove noises, and for text modality data, it is cleaned by methods such as removing stop words and normalizing vocabulary.

[0073] Thus, after preprocessing, the multimodal data of the commodity is obtained , where represents the multimodal data of the th commodity, and respectively represent the image modality data and the text modality data of the th commodity.

[0074] S2: Based on the multimodal data, through the large language model, guide the semantic understanding process of multimodal data such as images and texts, and help remove redundant information and enhance the semantic relevance between different modalities.

[0075] Specifically, it includes:

[0076] S2-1: Use the pre-trained model to encode the data of each modality to obtain the original features of the modality ; among them, the modality includes the image modality and the text modality. For the image modality, the ResNet (Residual Neural Network) model is used, and for the text modality, the Distil BERT model (Distil Bidirectional Encoder Representations from Transformers) is used for feature extraction;

[0077] (1);

[0078] Among them, represents the original features of the modality , is the pre-trained model of the modality .

[0079] S2-2: AsFigure 3 As shown, by designing specific prompts for each modality and leveraging large language models, multi-modal data such as images and texts are subjected to modality semantic enhancement to generate semantic summary information. , and the semantic summary information is used as new input, and the pre-trained model is used again for encoding to obtain multi-modal semantic features, thereby extracting rich semantic information from multi-modal data such as images and texts through the powerful semantic understanding and cross-modal learning capabilities of the large semantic model.

[0080] (2);

[0081] (3);

[0082] Among them, represents the semantic summary information, represents the semantic features of modality , represents the large language model, is the prompt for modality , is the pre-trained model for modality .

[0083] Among them, the prompt refers to the text prompt input to the large language model. The design of the prompt is usually set manually based on experience and can be understood as a kind of parameter. By specifying different prompts, the large language model gives different results, thereby affecting the model performance.

[0084] S2-3: To ensure the effective fusion of different modality features in a unified embedding space, in this embodiment, a Multilayer Perceptron (MLP) is used to perform linear projection processing on the original features and semantic features of modality respectively, to obtain the original projection feature and semantic projection feature of the th product under modality :

[0085] (4);

[0086] (5);

[0087] Among them, are all trainable parameter matrices, are all bias vectors.

[0088] S2-4: To remove information conflicts and enhance the synergistic effect among multimodal features, this embodiment introduces an interactive feature purifier. The original projection feature and semantic projection feature of the th product in modality are processed by the interactive feature purifier to obtain the original purified feature and semantic purified feature of the th product in modality respectively, which are: in modality as follows: and respectively;

[0089] (6);

[0090] (7);

[0091] where and are both interactive feature purifiers; are both trainable parameter matrices; is a bias vector; represents element-wise product; represents the Sigmoid activation function; represents the product ID embedding of the th product. It should be understood that this product ID embedding represents the initial feature of the product after removing multimodal data.

[0092] Through the above process, this embodiment completes the feature initialization process before the propagation of the graph neural network, ensuring that multimodal features are effectively enhanced in a unified embedding representation space.

[0093] S3: For multimodal data embeddings and product ID embeddings, this embodiment proposes a multi-view graph neural network (GNN) aggregation process. The views of each modality data are divided into the original perspective and the semantic perspective to enhance the model's ability to capture potential multi-dimensional information in the modality data, aggregate the semantic features and original features of different modalities, and utilize the high-order relationship capture ability of the graph neural network among multimodal data to model the complex relationships among different modality features and obtain more accurate feature representations.

[0094] Furthermore, graph neural network aggregation is to perform feature aggregation on the nodes in the graph structure through graph convolution operations and efficiently integrate the features using the relationship information between nodes.

[0095] Specifically:

[0096] S3-1: Establish a similarity matrix for different feature perspectives, namely the original feature similarity matrix and the semantic feature similarity matrix, by calculating the original feature similarity and semantic feature similarity in each modality.

[0097] Then, among the multi-modal original features and multi-modal semantic features, the original feature similarity and semantic feature similarity between product and product are respectively:

[0098] (8);

[0099] (9);

[0100] Wherein, represents the original feature similarity between product and product and product in modality represents the semantic feature similarity between product and product and product in modality is the original feature embedding representation of product in modality is the original feature embedding representation of product is the semantic feature embedding representation of product in modality is the semantic feature embedding representation of product in modality is the semantic feature embedding representation of product in modality is the semantic feature embedding representation of product in modality in modality

[0101] S3-2: To capture the most critical features in the neighborhood, in this embodiment, KNN (K-Nearest Neighbor) sparsification is performed on the dense graph. For each product , only the first edges with the greatest similarity are retained, thereby obtaining the sparsified original feature similarity matrix and the semantic feature similarity matrix , and the elements in the sparsified original feature similarity matrix and the semantic feature similarity matrix are respectively:

[0102] (10);

[0103] (11);

[0104] Among them, represents the modality the original feature edge weight between commodities and commodity ; represents the modality the semantic feature edge weight between commodities and commodity ; is the original feature similarity between commodity in modality and commodity , is the semantic feature similarity between commodity in modality and commodity ; represents the commodity set.

[0105] Among them, according to the similarity (original feature similarity / semantic feature similarity) between commodity and commodity , a similarity matrix (original feature similarity matrix / semantic feature similarity matrix) for all commodities is obtained. This similarity matrix is dense because there is a similarity between each commodity and other commodities. Regarding this similarity matrix as an adjacency matrix, a commodity similarity graph is obtained. The weight of the edge in the commodity similarity graph is the similarity between two adjacent commodities. Similarly, the commodity similarity graph is also dense. KNN sparsification means that for each commodity, only the K edges with the maximum similarity are selected, and the weights of the others are set to 0, that is, this edge is removed, thereby sparsifying the dense graph and at the same time retaining the most critical similarity information.

[0106] S3-3: Normalize the sparsified original feature similarity matrix and semantic feature similarity matrix to obtain the normalized original feature similarity matrix and semantic feature similarity matrix to alleviate the possible problem of gradient explosion;

[0107] (12);

[0108] (13);

[0109] Among them, represents 's degree matrix, represents 's degree matrix.

[0110] S3-4: Based on the normalized original feature similarity matrix and semantic feature similarity matrix, use the GCN network to propagate the multi-modal original features and multi-modal semantic features of all products in the corresponding similarity matrix:

[0111] (14);

[0112] (15);

[0113] (16);

[0114] (17);

[0115] Among them, represents the original feature of the product set obtained after aggregation by the graph neural network from the perspective of the original features in the th product modality ; represents the semantic feature of the product set obtained after aggregation by the graph neural network from the perspective of semantic features in the th product modality ; represents the number of layers of the graph neural network; and are respectively the original purified features of the th layer and the th layer in the th product modality ; and are respectively the semantic purified features of the th layer and the th layer in the th product modality ; is the total number of layers.

[0116] S3-5: From the perspective of product ID embedding, in this embodiment, recursively propagate the collaborative signal in the user-product interaction graph to obtain the product set ID embedding feature of the th product:

[0117] (18);

[0118] (19);

[0119] Among them, is the user-product interaction adjacency matrix; is the degree matrix of the matrix ; and They are respectively the -layer and the -layer embedding of the th product's product ID; is the total number of layers.

[0120] S4: After enhancement by the large language model and aggregation by the multi-view graph neural network, the th product modality original features of the product set and the semantic features of the product set are obtained, as well as the ID embedding features of the product set ; Extract the original preference of modality and the semantic preference of modality :

[0121] (20);

[0122] (21);

[0123] Among them, is a trainable parameter matrix, is a bias vector.

[0124] Combining the original features of the product set and the semantic features of the product set, further derive the overall multi-modal original features and multi-modal semantic features , and represent the ID embedding features of the product set as ID features :

[0125] (22);

[0126] (23);

[0127] Among them, M is the total number of modality sets.

[0128] In addition, in order to maximize the mutual information of the above features, this embodiment designs two self-supervised auxiliary tasks and gives two contrast loss functions:

[0129] (24);

[0130] (25);

[0131] Among them, represents the contrast loss function between the original features and the semantic features, Represents the contrast loss function between the ID embedding feature and two types of auxiliary features (Side); is the temperature coefficient of the contrast loss function; is the multi-modal raw feature embedding of the product ; is the multi-modal semantic feature embedding of the product ; is the multi-modal raw feature embedding of the product ; is the multi-modal semantic feature of the product ; is the ID embedding feature of the product ; is the ID embedding feature of the product ;

[0132] Finally, concatenate , and to obtain the multi-modal fusion feature enhanced by the large language model , completing the feature fusion process:

[0133] (26).

[0134] S5: Apply the multi-modal fusion feature to the downstream task; in this embodiment, the multi-modal fusion feature is applied to the product recommendation system in the e-commerce platform.

[0135] Specifically:

[0136] First, use MLP to perform feature transformation on , and then divide into user representation and product representation according to the category:

[0137] (27);

[0138] Among them, is the embedding representation of the user in the multi-modal fusion feature ; is the embedding representation of the product in the multi-modal fusion feature ;

[0139] Then, use the inner product to determine the interaction score between the user i and the product :

[0140] (28);

[0141] Finally, select the top X items with the highest interaction scores and recommend them to user u.

[0142] In this embodiment, in terms of model optimization, the BPR loss function (The BPR loss function (Bayesian Personalized Ranking loss) is a loss function for user personal preferences in a recommendation system) is used to reconstruct the historical data, which gives priority to the cases where the observed item scores are higher.

[0143] Furthermore, this embodiment combines two contrast loss functions and an L2 regularization term for joint optimization, so that the overall loss function of the model is:

[0144] (30);

[0145] Where, and are responsible for adjusting the proportion of the contrast task, is responsible for adjusting the proportion of the L2 regularization term.

[0146] In the recommendation task, this embodiment conducts experiments using an infant product dataset on a certain platform. This dataset provides visual modality features and text modality features of the products, and uses a 5-core setting of users and products to filter the original data. For the visual modality, this embodiment uses 2048-dimensional features obtained from the Residual Network (ResNet) structure. For the text modality, this embodiment uses 384-dimensional features obtained from the distilled BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture) model. In addition, the dataset is divided into a training set, a validation set, and a test set according to the ratio of 6:2:2.

[0147] The experimental parameter settings adopted in this embodiment are shown in Table 1; where, represents the model learning rate, and respectively represent the number of GCN aggregation layers of the user-item interaction graph and the item feature graph, represents the number of neighbors of the K-Nearest Neighbor (KNN) algorithm, represents the temperature coefficient of the contrast loss term, and represent the weights of the two contrast loss terms, represents the proportion of the L2 regularization term.

[0148] Table 1 Experimental parameter settings;

[0149] 。

[0150] In this embodiment, the recommendation performance of different models is evaluated by indicators such as Recall@K (recall rate) and NDCG@K (Normalized Discounted Cumulative Gain) (K = 10, 20). The method of this embodiment is compared with models such as BPR (Bayesian Personalized Ranking), LightGCN (Light Graph Convolution Network), VBPR (Visual Bayesian Personalized Ranking), MMGCN (Multi-modal Graph Convolutional Network), DualGNN (Dual Graph Neural Network), FREEDOM (Freezing and Denoising Graph Structures for Multimodal Recommendation), MGCN (Multi-View Graph Convolutional Network), DRAGON (Dyadic Relations Augmentation with Graphs for Multimodal Recommendation), and SMORE (Spectrum-based Modality Representation Fusion Graph Convolutional Network). The experimental results are shown in Table 2.

[0151] Table 2 Performance comparison of different models;

[0152] 。

[0153] As can be seen from Table 2, the model of this embodiment outperforms other methods in all metrics, indicating that the multimodal feature fusion method based on large language models proposed in this embodiment can effectively improve the recommendation performance in downstream tasks of the recommendation system. Specifically, in this embodiment, high-quality semantic features are generated for multimodal data by using large language models, further alleviating the conflict between multimodal information and behavioral information. At the same time, this embodiment also retains the original features and introduces two self-supervised tasks to maximize the mutual information between the original features, semantic features, and behavioral features.

[0154] Furthermore, compared with the matrix factorization-based method, the graph convolutional neural network-based method is more vulnerable to information conflicts. For example, VBPR outperforms BPR by directly concatenating modal and behavioral features, while the performance of MMGCN is worse than that of LightGCN. This is mainly because the message propagation mechanism in the graph convolutional neural network-based method continuously propagates noise and contaminates the representations of all users and items, thus exacerbating the conflict between multimodal information and behavioral information.

[0155] In addition, compared with the MGCN and SMORE models that focus on solving information conflicts, the model of this embodiment also achieves further performance improvement, which benefits from the semantic features generated by large language models, effectively alleviating redundancy and noise in multimodal data, thus significantly improving the performance of the model in complex scenarios. Through the in-depth understanding and semantic alignment of multimodal data, the potential associations between different modalities can be automatically captured, and information can be fused in a more precise manner, avoiding the common modality conflict problems in traditional methods.

[0156] It should be noted that the acquisition of all data is based on compliance with laws and regulations and user consent, and the data is legally applied.

[0157] Embodiment 2

[0158] As Figure 4 shown, this embodiment provides a multimodal feature fusion system, including:

[0159] An enhancement module, configured to encode the acquired multimodal data of items to obtain multimodal original features, and perform modal semantic enhancement according to the prompt words of each modality to obtain multimodal semantic features;

[0160] A similarity processing module, configured to calculate the original feature similarity matrix and semantic feature similarity matrix between different items in each modality according to the multimodal original features and multimodal semantic features, and only retain the top K similarities for each item, thereby obtaining a sparsified original feature similarity matrix and semantic feature similarity matrix, and performing normalization processing respectively;

[0161] An aggregation module, configured to aggregate the original features and semantic features of each modality according to the normalized original feature similarity matrix and semantic feature similarity matrix to obtain the original features of the product set and the semantic features of the product set, and obtain the product set ID embedding features according to the interaction between the user and the product;

[0162] A fusion module, configured to extract modality preferences from the product set ID embedding features, combine the original features of the product set and the semantic features of the product set under each modality to obtain the overall multi-modal original features, multi-modal semantic features and ID features, and splice the three to obtain multi-modal fusion features, so as to predict the interaction score between the user and the product according to the multi-modal fusion features and recommend products to the user.

[0163] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.

[0164] In more embodiments, there is also provided:

[0165] An electronic device, including a memory and a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be repeated here.

[0166] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0167] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0168] A computer-readable storage medium, used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.

[0169] The method in Embodiment 1 can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can be located in mature storage media in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage media is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0170] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method described in Embodiment 1.

[0171] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed. The machine-executable instructions for program modules can be executed within local or distributed devices. In a distributed device, program modules can be located in local and remote storage media.

[0172] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, so that when the program code is executed by the computer or other programmable data processing devices, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0173] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc. Examples of signals can include electrical, optical, radio, sound, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0174] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0175] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they do not limit the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A multimodal feature fusion method, characterized in that including: Encoding the multi-modal data of the obtained commodities to obtain multi-modal original features, and using a large language model to perform modal semantic enhancement according to the prompt words of each modality to obtain multi-modal semantic features; According to the multi-modal original features and multi-modal semantic features, calculate the original feature similarity matrix and semantic feature similarity matrix between different commodities in each modality, and only retain the top K similarities for each commodity, thereby obtaining a sparsified original feature similarity matrix and semantic feature similarity matrix, and perform normalization processing on them respectively; Based on the normalized original feature similarity matrix and semantic feature similarity matrix, aggregate the original features and semantic features of each modality based on a graph neural network to obtain the original features of the commodity set and the semantic features of the commodity set, and obtain the commodity set ID embedding features according to the interaction between the user and the commodity; Extract the modality preference from the commodity set ID embedding features, combine the original features of the commodity set and the semantic features of the commodity set in each modality to obtain the overall multi-modal original features, multi-modal semantic features and ID features, and concatenate the three to obtain multi-modal fusion features, so as to predict the interaction score between the user and the commodity according to the multi-modal fusion features and recommend commodities to the user.

2. The multimodal feature fusion method according to claim 1, wherein After obtaining the multi-modal raw features and multi-modal semantic features, perform a linear projection on the multi-modal raw features and multi-modal semantic features to obtain the raw projection feature of the th commodity in the modality and the semantic projection feature . Then, perform an interactive feature purification process on each of them to obtain the raw purified feature of the th commodity in the modality and the semantic purified feature , which are respectively: ; ; Among them, are all trainable parameter matrices; is a bias vector; is an element-wise product; is the Sigmoid activation function; is the commodity ID embedding of the 3. A multimodal feature fusion method according to claim 1, characterized in that, The elements in the original feature similarity matrix and semantic feature similarity matrix between different commodities in each modality are respectively: ; ; The elements in the sparsified original feature similarity matrix and semantic feature similarity matrix are respectively: ; ; Among them, is the modal the original feature similarity between the products and the products ; is the modal the semantic feature similarity between the products and the products ; is the original feature embedding representation of the product in the modal ; is the original feature embedding representation of the product in the modal ; is the semantic feature embedding representation of the product in the modal ; is the semantic feature embedding representation of the product in the modal ; is the original feature edge weight between the product and the product in the modal ; is the semantic feature edge weight between the product and the product in the modal ; is the original feature similarity between the product and the product in the modal ; is the semantic feature similarity between the product and the product in the modal ; is the set of products.

4. A multimodal feature fusion method according to claim 1, wherein The normalized original feature similarity matrix and the semantic feature similarity matrix are respectively as follows: ; ; Among them, is the original feature similarity matrix sparsified in modality ; is the semantic feature similarity matrix sparsified in modality ; is the degree matrix of , is the degree matrix of .

5. A multimodal feature fusion method according to claim 1, characterized in that, The original features of the commodity set and the semantic features of the commodity set are respectively: ; ; ; ; Among them, is the th commodity modality in which the original feature of the commodity set obtained by aggregation through the graph neural network from the perspective of the original features; is the th commodity modality in which the semantic feature of the commodity set obtained by aggregation through the graph neural network from the perspective of the semantic features; is the number of layers of the graph neural network; and are respectively the th commodity modality in which the original purification features of the th layer and the th layer; and are respectively the th commodity modality in which the semantic purification features of the th layer and the th layer; is the total number of layers; Commodity set ID embedding feature is as follows: ; ; Among them, is the user-product interaction adjacency matrix; is the degree matrix of the matrix ; and are the product ID embeddings of the -th and -th products of the -th layer, respectively.

6. A multimodal feature fusion method according to claim 1, wherein, Extract the modality from the commodity set ID embedding feature of the original preference and semantic preference : ; ; Overall multimodal raw features and multimodal semantic features are respectively as follows: ; ; where M is the total number of modal sets; is the original feature of the product set; is the semantic feature of the product set; is a trainable parameter matrix, is a bias vector; is the Sigmoid activation function; Divide the multi-modal fusion features into user representations and item representations , and use the inner product to determine the interaction score between the user and the item.

7. A multimodal feature fusion system, characterized in that, including: An enhancement module configured to encode the multi-modal data of the obtained commodities to obtain multi-modal original features, and use a large language model to perform modal semantic enhancement according to the prompt words of each modality to obtain multi-modal semantic features; A similarity processing module configured to calculate the original feature similarity matrix and semantic feature similarity matrix between different commodities in each modality according to the multi-modal original features and multi-modal semantic features, and only retain the top K similarities for each commodity, thereby obtaining a sparsified original feature similarity matrix and semantic feature similarity matrix, and perform normalization processing on them respectively; An aggregation module configured to aggregate the original features and semantic features of each modality based on a graph neural network according to the normalized original feature similarity matrix and semantic feature similarity matrix to obtain the original features of the commodity set and the semantic features of the commodity set, and obtain the commodity set ID embedding features according to the interaction between the user and the commodity; A fusion module configured to extract the modality preference from the commodity set ID embedding features, combine the original features of the commodity set and the semantic features of the commodity set in each modality to obtain the overall multi-modal original features, multi-modal semantic features and ID features, and concatenate the three to obtain multi-modal fusion features, so as to predict the interaction score between the user and the commodity according to the multi-modal fusion features and recommend commodities to the user.

8. An electronic device, characterized in that, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in any one of claims 1-6 is completed.

9. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by a processor, the method according to any one of claims 1-6 is completed.

10. A computer program product, characterized in that, Comprising a computer program, when the computer program is executed by a processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Graph Transform article recommendation method based on multi-modal semantic fusion

    CN118657560A

  • Unsupervised cross-modal retrieval method and system based on hypergraph convolution, medium and equipment

    CN118916497A