Multi-modal feature fusion method, system and device, medium and program product

By introducing large language models into multimodal feature fusion for semantic enhancement, and using graph neural network to aggregate features, the problems of information redundancy and conflict in multimodal data are solved, and more accurate and effective feature fusion is achieved, improving the performance of downstream tasks.

CN120068009AActive Publication Date: 2025-05-30SHANDONG UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510553653.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

When existing multimodal feature fusion methods deal with the complementarity of different types of data, it is difficult to balance and optimize the coordination and conflict between different modes. Especially when there are redundant or irrelevant features, how to effectively eliminate this information and extract meaningful features is still a difficult point that current technology needs to solve urgently.

Method used

A multimodal feature fusion method based on large language model is adopted to enhance the multimodal data by semantic enhancement, and aggregation of original features and semantic features is used to capture the relationship between different modal features and improve the feature fusion effect.

Benefits of technology

It effectively improves the semantic understanding ability of multimodal features, optimizes redundant information, improves the accuracy and robustness of feature representation, and thus improves the performance of downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068009A_ABST
    Figure CN120068009A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal feature fusion method, system and device, a medium and a program product, and relates to the technical field of feature fusion, and the method comprises the steps: carrying out the coding and semantic enhancement of multi-modal data, and obtaining a multi-modal original feature and a multi-modal semantic feature; calculating an original feature similarity matrix and a semantic feature similarity matrix, and performing sparse processing and normalization processing; according to the normalized original feature similarity matrix and semantic feature similarity matrix, original features and semantic features are aggregated, and commodity set ID embedded features are obtained according to interaction between users and commodities; and extracting modal preference from the commodity set ID embedded features, combining the commodity set original features and the commodity set semantic features to obtain overall multi-modal original features, multi-modal semantic features and ID features, and splicing the three features to obtain multi-modal fusion features. The problems of information redundancy and conflict of multi-modal data are solved, the relation between different modal features is captured, and the fusion effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of feature fusion, and particularly to a multi-modal feature fusion method, system, device, medium and program product. Background Art

[0002] Traditional feature fusion methods mostly rely on the processing of single-modal data, such as images, texts or other single types of data. These methods usually encounter problems such as data inconsistency, feature redundancy and information conflict. Especially when facing data from different sources, how to effectively integrate various modal features to achieve efficient information extraction and processing is an important challenge in the current technology.

[0003] At present, the application of multi-modal feature fusion has been widely penetrated into various fields, such as speech processing, medical image analysis, etc. Existing methods usually rely on pre-trained models or specific coding networks. By extracting features of different modalities, they are mapped into a shared representation space.

[0004] Although these methods have achieved certain success in dealing with the complementarity of different types of data; they still face the problem of how to balance and optimize the cooperation and conflict between different modalities. Especially when there are redundant or irrelevant features, how to effectively eliminate this information and extract meaningful features is still a difficult point that needs to be solved urgently in the current technology.

[0005] For example, in the fusion of medical images and diagnostic texts, the images provide visual information of diseases, and the texts provide more detailed medical record descriptions. If redundant features cannot be accurately distinguished, it may lead to inaccurate feature representations after fusion, thereby affecting the performance of downstream tasks.

[0006] Similarly, in the product recommendation of e-commerce platforms, the images provide the appearance information of products, and the texts describe product attributes and user experiences. If effective feature fusion cannot be achieved, it is easy to introduce information redundancy or semantic inconsistency, thus affecting the recommendation effect.

[0007] Therefore, in the research of multi-modal feature fusion, how to enhance the semantic understanding of fused features and optimize redundant information is one of the core issues. Summary of the Invention

[0008] To solve the above problems, the present invention proposes a multi-modal feature fusion method, system, device, medium and program product, which performs semantic enhancement on multi-modal data, deeply mines the potential semantic information in different modalities, and solves the problems of information redundancy and conflict existing in multi-modal data; uses a graph neural network to aggregate the original features and semantic features, captures the relationships between different modal features, and improves the feature fusion effect.

[0009] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a multi-modal feature fusion method, including: Encoding the multi-modal data of the obtained commodity to obtain multi-modal original features, and performing modal semantic enhancement according to the prompt words of each modality to obtain multi-modal semantic features; According to the multi-modal original features and multi-modal semantic features, calculate the original feature similarity matrix and semantic feature similarity matrix between different commodities in each modality, and only keep the top K similarities for each commodity, thereby obtaining a sparsified original feature similarity matrix and semantic feature similarity matrix, and performing normalization processing respectively; According to the normalized original feature similarity matrix and semantic feature similarity matrix, aggregate the original features and semantic features of each modality to obtain the original features of the commodity set and the semantic features of the commodity set, and obtain the commodity set ID embedding feature according to the interaction between the user and the commodity; Extract the modal preference from the commodity set ID embedding feature, combine the original features of the commodity set and the semantic features of the commodity set in each modality to obtain the overall multi-modal original features, multi-modal semantic features and ID features, and splice the three to obtain the multi-modal fusion feature, so as to predict the interaction score between the user and the commodity according to the multi-modal fusion feature and recommend commodities to the user.

[0010] As an alternative implementation, after obtaining the multi-modal original features and multi-modal semantic features, perform linear projection on the multi-modal original features and multi-modal semantic features to obtain the th commodity's original projection feature under modality and semantic projection feature , and perform interaction feature purification processing respectively to obtain the th commodity's original purification feature under modality and semantic purification feature , respectively: ; ; wherein, are all trainable parameter matrices; is a bias vector; is an element-wise product; is a Sigmoid activation function; is the commodity ID embedding of the th commodity.

[0011] As an alternative implementation, the elements in the original feature similarity matrix and semantic feature similarity matrix between different commodities in each modality are respectively: ; ; The elements in the sparsified original feature similarity matrix and the semantic feature similarity matrix are respectively: ; ; Among them, is the original feature similarity between product and product in modality , is the semantic feature similarity between product and product in modality ; is the original feature embedding representation of product in modality , is the original feature embedding representation of product in modality , is the semantic feature embedding representation of product in modality ; is the semantic feature embedding representation of product in modality ; is the original feature edge weight between product and product in modality ; is the semantic feature edge weight between product and product in modality ; is the original feature similarity between product and product in modality , is the semantic feature similarity between product and product in modality ; is the set of products.

[0012] As an alternative implementation, the normalized original feature similarity matrix and the semantic feature similarity matrix are respectively: ; ; Among them, is in modality ​​​The original feature similarity matrix after medium sparsification; is the modality The semantic feature similarity matrix after medium sparsification; is the degree matrix of is the degree matrix of

[0013] As an alternative implementation, the original features of the product set and the semantic features of the product set are respectively: ; ; ; ; where is the th product modality in the original feature perspective, the original features of the product set obtained after aggregation by the graph neural network; is the th product modality in the semantic feature perspective, the semantic features of the product set obtained after aggregation by the graph neural network; is the number of layers of the graph neural network; and are respectively the th product modality in the th layer and the th layer of the original purified features; and are respectively the th product modality in the th layer and the th layer of the semantic purified features; is the total number of layers; Product set ID embedding features are: ; ; where is the user-product interaction adjacency matrix; is the degree matrix of the matrix ; and are respectively the th layer and the th layer of the th product's product ID embedding.

[0014] As an alternative implementation, extract the modality from the product set ID embedding features ​ Original preference and semantic preference : ; ; Overall multimodal original features and multimodal semantic features are respectively: ; ; where M is the total number of modal sets; is the original feature of the commodity set; is the semantic feature of the commodity set; is the trainable parameter matrix, is the bias vector; is the Sigmoid activation function; The multimodal fusion features are divided into user representation and commodity representation , and the inner product is used to determine the interaction score between the user and the commodity.

[0015] In a second aspect, the present invention provides a multimodal feature fusion system, including: An enhancement module, configured to encode the obtained multimodal data of the commodity to obtain multimodal original features, and perform modal semantic enhancement according to the prompt words of each modality to obtain multimodal semantic features; A similarity processing module, configured to calculate the original feature similarity matrix and semantic feature similarity matrix between different commodities in each modality according to the multimodal original features and multimodal semantic features, and only retain the top K similarities for each commodity, thereby obtaining a sparsified original feature similarity matrix and semantic feature similarity matrix, and respectively performing normalization processing; An aggregation module, configured to aggregate the original features and semantic features of each modality according to the normalized original feature similarity matrix and semantic feature similarity matrix to obtain the original features of the commodity set and the semantic features of the commodity set, and obtain the ID embedding feature of the commodity set according to the interaction between the user and the commodity; A fusion module, configured to extract the modal preference from the ID embedding feature of the commodity set, combine the original features of the commodity set and the semantic features of the commodity set in each modality to obtain the overall multimodal original features, multimodal semantic features and ID features, splice the three to obtain multimodal fusion features, so as to predict the interaction score between the user and the commodity according to the multimodal fusion features and recommend commodities to the user.

[0016] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in the first aspect is performed.

[0017] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, wherein when the computer instructions are executed by a processor, the method described in the first aspect is performed.

[0018] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in the first aspect.

[0019] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes a multimodal feature fusion method based on a large language model, introduces the semantic enhancement of the large language model into the multimodal feature fusion, and performs semantic enhancement processing of the large language model on multimodal data such as images and texts. It can deeply mine the potential semantic information in different modalities such as images and texts, improve the synergy between the modalities, enhance the feature representation capability, solve the information redundancy in the multimodal data itself and the conflict between multimodal information and behavioral information, and effectively improve the accuracy and robustness of downstream tasks.

[0020] The present invention utilizes graph neural networks to aggregate original features and semantic features, which can better capture the relationship between features of different modalities, improve the feature fusion effect, and make the final multimodal features more accurate and effective.

[0021] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0023] Figure 1 A flow chart of a multimodal feature fusion method provided in Example 1 of the present invention; Figure 2 Schematic diagram of the multimodal feature fusion method provided in Example 1 of the present invention; Figure 3 A diagram of a large language model enhancement module provided in Example 1 of the present invention; Figure 4This is the structural diagram of the multi-modal feature fusion system provided in Embodiment 2 of the present invention. Detailed implementation manners

[0024] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0025] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0026] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0028] Embodiment 1 With the emergence of large language models (LLMs), they have shown excellent capabilities in processing semantic information, cross-modal learning, etc. Using large language models for semantic mining can not only further improve the quality of feature representation, but also effectively alleviate the problems of information redundancy and conflict. Through the semantic understanding and cross-modal learning capabilities of large language models, effective features can be more accurately identified, and interfering features can be eliminated, providing a new solution for multi-modal feature fusion. Therefore, this embodiment combines the advantages of large language models and realizes more efficient feature fusion through semantic enhancement methods.

[0029] This embodiment relies on the Key Laboratory of Artificial Intelligence Application for People's Livelihood Services in Shandong Province and the Future Industry Laboratory of General Artificial Intelligence in Shandong Province's institutions of higher learning to provide a multi-modal feature fusion method based on large language models, as Figure 1 shown, including: Encoding the multi-modal data of the obtained commodity to obtain multi-modal original features, and performing modal semantic enhancement according to the prompt words of each modality to obtain multi-modal semantic features; Calculate the original feature similarity matrix and semantic feature similarity matrix between different commodities for each modality based on the multi-modal original features and multi-modal semantic features, and only retain the top K similarities for each commodity, thereby obtaining the sparsified original feature similarity matrix and semantic feature similarity matrix, and perform normalization processing on them respectively; Aggregate the original features and semantic features of each modality based on the normalized original feature similarity matrix and semantic feature similarity matrix to obtain the original features of the commodity set and the semantic features of the commodity set, and obtain the commodity set ID embedding features according to the interaction between the user and the commodity; Extract the modality preference from the commodity set ID embedding features, combine the original features of the commodity set and the semantic features of the commodity set for each modality, obtain the overall multi-modal original features, multi-modal semantic features and ID features, and splice the three to obtain the multi-modal fusion features, so as to predict the interaction score between the user and the commodity based on the multi-modal fusion features and recommend commodities to the user.

[0030] Next, in combination with Figure 1 - Figure 2 Elaborate on the method of this embodiment in detail.

[0031] S1: Obtain The original multi-modal data of the commodities and perform preprocessing to ensure that the formats of different modality data are consistent and have processing capabilities.

[0032] Among them, the multi-modal data includes but is not limited to different types of data such as images, texts, and audios. In this embodiment, the image modality and text modality are taken as examples, and the extension of more modality data is realized based on this.

[0033] Among them, the preprocessing includes data cleaning and denoising processing to remove the error items and noises in the data; for the image modality data, use a Gaussian filter to remove the noises, and for the text modality data, clean it by methods such as removing stop words and normalizing vocabulary.

[0034] Thus, after preprocessing, obtain The multi-modal data of the commodities , where Represents the multi-modal data of the th commodity, And Respectively represent the image modality data and text modality data of the th commodity.

[0035] S2: Based on the multi-modal data, through the large language model, guide the semantic understanding process of multi-modal data such as images and texts, and help remove redundant information and enhance the semantic relevance between different modalities.

[0036] Specifically include: S2-1: Use the pre-trained model to encode the data of each modality to obtain the original features of the modality ; where the modality includes the image modality and the text modality. For the image modality, use the ResNet (Residual Neural Network) model, and for the text modality, use the DistilBERT model (DistilBidirectional Encoder Representations from Transformers) for feature extraction; (1); where represents the original features of the modality , is the pre-trained model of the modality .

[0037] S2-2: As Figure 3 shown, by designing specific prompts for each modality, use the large language model to enhance the modality semantics of multi-modal data such as images and texts, and generate semantic summary information , and use the semantic summary information as the new input, and use the pre-trained model to encode again to obtain multi-modal semantic features, so as to extract rich semantic information in multi-modal data such as images and texts through the powerful semantic understanding and cross-modal learning capabilities of the large semantic model.

[0038] (2); (3); where represents the semantic summary information, represents the semantic features of the modality , represents the large language model, is the prompt of the modality , is the pre-trained model of the modality .

[0039] Among them, the prompt refers to the text prompt input to the large language model. The design of the prompt is usually set manually according to experience and can be understood as a parameter. By specifying different prompts, the large language model gives different results, thereby affecting the model performance.

[0040] S2-3: To ensure the effective fusion of different modal features in a unified embedding space, this embodiment uses a Multilayer Perceptron (MLP) to perform linear projection processing on the original features and semantic features of each modality respectively, obtaining the original projection feature and semantic projection feature of the -th commodity in modality : : (4); (5); Among them, are all trainable parameter matrices, and are all bias vectors.

[0041] S2-4: To remove information conflicts and enhance the synergistic effect between multi-modal features, this embodiment introduces an interactive feature purifier. After the original projection feature and semantic projection feature of the -th commodity in modality are processed by the interactive feature purifier, the obtained original purified feature and semantic purified feature of the -th commodity in modality are respectively: (6); (7); Among them, and are both interactive feature purifiers; are all trainable parameter matrices; is a bias vector; represents element-wise multiplication; represents the Sigmoid activation function; represents the commodity ID embedding of the -th commodity. It should be understood that this commodity ID embedding represents the initial features of the commodity after removing multi-modal data.

[0042] Through the above process, this embodiment completes the feature initialization process before the propagation of the graph neural network, ensuring the effective enhancement of multi-modal features in a unified embedding representation space.

[0043] S3: For multi-modal data embedding and product ID embedding, this embodiment proposes a multi-view based Graph Neural Network (GNN) aggregation process. The views of each type of modal data are divided into the Original perspective and the Semantic perspective to enhance the model's ability to capture potential multi-dimensional information in the modal data, aggregate the semantic features and original features of different modalities, and utilize the GNN's ability to capture high-order relationships between multi-modal data to model the complex relationships between different modal features and obtain more accurate feature representations.

[0044] Furthermore, the GNN aggregation is to perform feature aggregation on the nodes in the graph structure through graph convolution operations, and efficiently integrate the features using the relationship information between nodes.

[0045] Specifically: S3-1: By calculating the original feature similarity and semantic feature similarity for each modality, establish similarity matrices for different feature perspectives, namely the original feature similarity matrix and the semantic feature similarity matrix; Then, among the multi-modal original features and multi-modal semantic features, the and original feature similarity and semantic feature similarity between products are respectively: (8); (9); Where, represents the original feature similarity between product and product in modality , represents the semantic feature similarity between product and product in modality ; is the original feature embedding representation of product in modality , is the original feature embedding representation of product in modality , is the semantic feature embedding representation of product in modality , is the semantic feature embedding representation of product in modality .

[0046] S3-2: To capture the most critical features among neighbors, in this embodiment, KNN (K-Nearest Neighbor) sparsification is performed on the dense graph. For each commodity , only the first edges with the greatest similarity are retained, thereby obtaining the sparsified original feature similarity matrix and the semantic feature similarity matrix . The elements in the sparsified original feature similarity matrix and the semantic feature similarity matrix are respectively: (10); (11); Among them, represents the original feature edge weight between commodity in modality and commodity ; represents the semantic feature edge weight between commodity in modality and commodity ; is the original feature similarity between commodity in modality and commodity , is the semantic feature similarity between commodity in modality and commodity ; represents the commodity set.

[0047] Among them, according to the similarity (original feature similarity / semantic feature similarity) between commodity and commodity , a similarity matrix (original feature similarity matrix / semantic feature similarity matrix) for all commodities is obtained. This similarity matrix is dense because there will be a similarity between each commodity and other commodities. Regarding this similarity matrix as an adjacency matrix, a commodity similarity graph is obtained. The weight of the edge in the commodity similarity graph is the similarity between two adjacent commodities. Similarly, the commodity similarity graph is also dense. KNN sparsification means that for each commodity, only the K edges with the greatest similarity are selected, and the weights of the others are set to 0, that is, this edge is removed, thereby sparsifying the dense graph and simultaneously retaining the most critical similarity information.

[0048] S3-3: Normalize the sparsified original feature similarity matrix and semantic feature similarity matrix to obtain the normalized original feature similarity matrix and the semantic feature similarity matrix , to alleviate the possible gradient explosion problem; (12); (13); Among them, represents 's degree matrix, represents 's degree matrix.

[0049] S3-4: According to the normalized original feature similarity matrix and semantic feature similarity matrix, adopt the GCN network to propagate the multimodal original features and multimodal semantic features of all commodities in the corresponding similarity matrix: (14); (15); (16); (17); Among them, represents the th commodity modality in the original feature perspective, the original feature of the commodity set obtained after aggregation by the graph neural network; represents the th commodity modality in the semantic feature perspective, the semantic feature of the commodity set obtained after aggregation by the graph neural network; represents the number of layers of the graph neural network; and are respectively the th commodity modality in the th layer and the th layer of the original purification feature; and are respectively the th commodity modality in the th layer and the th layer of the semantic purification feature; is the total number of layers.

[0050] S3-5: Starting from the perspective of commodity ID embedding, in this embodiment, the collaborative signal is recursively propagated in the user-commodity interaction graph to obtain the commodity set ID embedding feature of the th commodity: (18); (19); Among them, is the user-product interaction adjacency matrix; is the matrix 's degree matrix; and are respectively the -th layer and the -th layer's -th product ID embedding; is the total number of layers.

[0051] S4: After enhancement by the large language model and aggregation by the multi-view graph neural network, the -th product modality 's original features of the product set and semantic features of the product set , as well as the ID embedding features of the product set ; Extract the original preference of modality and semantic preference of modality : (20); (21); Among them, is the trainable parameter matrix, is the bias vector.

[0052] Combining the original features of the product set and the semantic features of the product set, further derive the overall multi-modal original features and multi-modal semantic features , and represent the ID embedding features of the product set as ID features : (22); (23); Among them, M is the total number of modality sets.

[0053] In addition, in order to maximize the mutual information of the above features, this embodiment designs two self-supervised auxiliary tasks and gives two contrastive loss functions: (24); (25); Among them, represents the contrastive loss function between the original features and the semantic features, represents the contrastive loss function between the ID embedding features and two types of auxiliary features (Side); is the temperature coefficient of the contrast loss function; is the commodity 's multi-modal raw feature embedding; is the commodity 's multi-modal semantic feature embedding; is the commodity 's multi-modal raw feature embedding; is the commodity 's multi-modal semantic feature; is the commodity 's ID embedding feature; is the commodity 's ID embedding feature.

[0054] Finally, concatenate , and to obtain the multi-modal fusion feature enhanced by the large language model , completing the feature fusion process: (26).

[0055] S5: Apply the multi-modal fusion feature to the downstream task; in this embodiment, the multi-modal fusion feature is applied to the commodity recommendation system in the e-commerce platform.

[0056] Specifically: First, use MLP to perform feature transformation, and then divide into user representation and commodity representation according to the category: (27); Among them, is the embedding representation of user in the multi-modal fusion feature ; is the embedding representation of commodity in the multi-modal fusion feature ;

[0057] Then, use the inner product to determine the interaction score between user i and commodity : (28); Finally, select the top X commodities with the highest interaction scores and recommend them to user u.

[0058] In this embodiment, in terms of model optimization, the BPR loss function The (Bayesian Personalized Ranking loss) is a loss function for user personal preferences in a recommendation system. It reconstructs historical data, giving priority to cases with higher observed item scores.

[0059] Furthermore, this embodiment combines two contrast loss functions and an L2 regularization term for joint optimization, making the overall loss function of the model as follows: (30); Where, and are responsible for adjusting the proportion of the contrast task, is responsible for adjusting the proportion of the L2 regularization term.

[0060] In the recommendation task, this embodiment uses a dataset of infant and toddler products on a certain platform for experiments. This dataset provides visual modality features and text modality features of products, and filters the original data using a 5-core setting of users and products. For the visual modality, this embodiment uses 2048-dimensional features obtained from a Residual Network (ResNet) structure. For the text modality, this embodiment uses 384-dimensional features obtained from a distilled BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture) model. In addition, the dataset is divided into a training set, a validation set, and a test set in a ratio of 6:2:2.

[0061] The experimental parameter settings used in this embodiment are shown in Table 1; where, represents the model learning rate, and respectively represent the number of GCN aggregation layers of the user-product interaction graph and the product feature graph, represents the number of neighbors of the K-Nearest Neighbor (KNN) algorithm, represents the temperature coefficient of the contrast loss term, and represent the weights of the two contrast loss terms, represents the proportion of the L2 regularization term.

[0062] Table 1 Experimental parameter settings; .

[0063] In this embodiment, the recommendation performance of different models is evaluated by metrics such as Recall@K (recall rate) and NDCG@K (Normalized Discounted Cumulative Gain) (K = 10, 20). The method of this embodiment is compared with models such as BPR (Bayesian Personalized Ranking), LightGCN (Light Graph Convolution Network), VBPR (Visual Bayesian Personalized Ranking), MMGCN (Multi-modal Graph Convolutional Network), DualGNN (Dual Graph Neural Network), FREEDOM (Freezing and Denoising Graph Structures for Multimodal Recommendation), MGCN (Multi-View Graph Convolutional Network), DRAGON (Dyadic Relations Augmentation with Graphs for Multimodal Recommendation), and SMORE (Spectrum-based Modality Representation Fusion Graph Convolutional Network). The experimental results are shown in Table 2.

[0064] Table 2 Performance comparison of different models; 。

[0065] As can be seen from Table 2, the model of this embodiment outperforms other methods in all metrics, indicating that the proposed multi-modal feature fusion method based on large language models can effectively improve the recommendation performance in downstream tasks of the recommendation system. Specifically, in this embodiment, high-quality semantic features are generated for multi-modal data by using large language models, further alleviating the conflict between multi-modal information and behavioral information. At the same time, this embodiment also retains the original features and introduces two self-supervised tasks to maximize the mutual information between the original features, semantic features, and behavioral features.

[0066] Furthermore, compared with the matrix factorization-based method, the graph convolutional neural network-based method is more vulnerable to information conflicts. For example, VBPR outperforms BPR by directly concatenating modal and behavioral features, while the performance of MMGCN is worse than that of LightGCN. This is mainly because the message propagation mechanism in the graph convolutional neural network-based method continuously spreads noise and contaminates the representations of all users and items, thus exacerbating the conflict between multi-modal information and behavioral information.

[0067] In addition, compared with the MGCN and SMORE models that focus on solving information conflicts, the model of this embodiment also achieves further performance improvement, which benefits from the semantic features generated by the large language model, effectively alleviating redundancy and noise in multi-modal data, and thus significantly improving the performance of the model in complex scenarios. Through the in-depth understanding and semantic alignment of multi-modal data, the potential associations between different modalities can be automatically captured, and information can be fused in a more precise manner, avoiding the common modality conflict problems in traditional methods.

[0068] It should be noted that the acquisition of all data is based on compliance with laws and regulations and user consent, and the data is legally applied.

[0069] Embodiment 2 As Figure 4 shown, this embodiment provides a multi-modal feature fusion system, including: An enhancement module, configured to encode the multi-modal data of the acquired item to obtain multi-modal raw features, and perform modal semantic enhancement according to the prompt words of each modality to obtain multi-modal semantic features; A similarity processing module, configured to calculate the raw feature similarity matrix and semantic feature similarity matrix between different items in each modality according to the multi-modal raw features and multi-modal semantic features, and only retain the top K similarities for each item, thereby obtaining a sparsified raw feature similarity matrix and semantic feature similarity matrix, and performing normalization processing respectively; An aggregation module, configured to aggregate the raw features and semantic features of each modality according to the normalized raw feature similarity matrix and semantic feature similarity matrix to obtain the item set raw features and item set semantic features, and obtain the item set ID embedding features according to the interaction between the user and the item; A fusion module, configured to extract the modal preference from the item set ID embedding features, combine the item set raw features and item set semantic features in each modality to obtain the overall multi-modal raw features, multi-modal semantic features and ID features, splice the three to obtain the multi-modal fusion features, so as to predict the interaction score between the user and the item according to the multi-modal fusion features and recommend items to the user.

[0070] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0071] In more embodiments, there is also provided: An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.

[0072] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0073] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0074] A computer-readable storage medium for storing computer instructions, which, when executed by the processor, complete the method described in Embodiment 1.

[0075] The method in Embodiment 1 can be directly embodied as being executed by a hardware processor, or completed by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0076] A computer program product, including a computer program, which, when executed by the processor, implements the method described in Embodiment 1.

[0077] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the processes / methods described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed among the program modules. The machine-executable instructions for the program modules can be executed within local or distributed devices. In a distributed device, the program modules can be located in local and remote storage media.

[0078] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0079] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc. Examples of signals can include electrical, optical, radio, sound, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0080] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0081] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A multimodal feature fusion method, characterized in that: include: The obtained multimodal data of the goods are encoded to obtain the multimodal original features, and the modal semantics are enhanced according to the prompt words of each modality to obtain the multimodal semantic features; According to the multimodal original features and multimodal semantic features, the original feature similarity matrix and semantic feature similarity matrix between different products in each modality are calculated, and only the first K similarities are retained for each product, thereby obtaining the sparse original feature similarity matrix and semantic feature similarity matrix, and normalizing them respectively; According to the normalized original feature similarity matrix and semantic feature similarity matrix, the original features and semantic features of each modality are aggregated to obtain the original features of the product set and the semantic features of the product set. According to the interaction between users and products, the embedded features of the product set ID are obtained. Modal preferences are extracted from the embedded features of the product set ID. The original features of the product set and the semantic features of the product set under each modality are combined to obtain the overall multimodal original features, multimodal semantic features and ID features. The three are concatenated to obtain multimodal fusion features. The interaction score between users and products is predicted based on the multimodal fusion features, and products are recommended to users.

2. A multimodal feature fusion method as claimed in claim 1, characterized in that: After obtaining the multimodal original features and multimodal semantic features, linear projection is performed on the multimodal original features and multimodal semantic features to obtain the first Products in modal The original projection characteristics under and semantic projection features , and perform interactive feature purification respectively to obtain Products in modal The original purification characteristics under and semantic purification features , respectively: ; ; in, All are trainable parameter matrices; is the bias vector; is the element-wise product; is the Sigmoid activation function; For the The product ID of each product is embedded.

3. A multimodal feature fusion method as claimed in claim 1, characterized in that: The elements in the original feature similarity matrix and semantic feature similarity matrix between different products in each modality are: ; ; The elements in the sparse original feature similarity matrix and the semantic feature similarity matrix are: ; ; in, For modal Medium Products and products The original feature similarity of For modal Medium Products and products The similarity of semantic features; For modal Medium Products The original feature embedding representation of For modal Medium Products The original feature embedding representation of For modal Medium Products The semantic feature embedding representation of For modal Medium Products The semantic feature embedding representation of For modal Medium Products and products The original feature edge weights between ; For modal Medium Products and products The semantic feature edge weight between them; For modal Medium Products and products The original feature similarity of For modal Medium Products and products The similarity of semantic features; A collection of goods.

4. A multimodal feature fusion method as claimed in claim 1, characterized in that: Normalized original feature similarity matrix and the semantic feature similarity matrix They are: ; ; in, For modal The original feature similarity matrix is ​​sparsely populated; For modal The sparse semantic feature similarity matrix in ; for The degree matrix of for The degree matrix of .

5. A multimodal feature fusion method as claimed in claim 1, characterized in that: The original features of the product set and the semantic features of the product set are: ; ; ; ; in, For the Product modal In the figure, the original features of the product set are obtained after aggregation by the graph neural network from the perspective of the original features; For the Product modal In the figure, the semantic features of the product set are obtained after aggregation by the graph neural network from the perspective of semantic features; is the number of layers of the graph neural network; and Respectively Product modal In Layer and The original cleansing characteristics of the layer; and Respectively Product modal In Layer and Semantic purification features of layers; is the total number of layers; Product collection ID embedding feature for: ; ; in, is the user-item interaction adjacency matrix; For the matrix The degree matrix of and Respectively Layer and Layer The product ID of each product is embedded.

6. A multimodal feature fusion method as claimed in claim 1, characterized in that: Embed features from product collection ID Extract the modal Original preference and semantic preference : ; ; Overall multimodal raw features and multimodal semantic features They are: ; ; Where M is the total number of modal sets; The original features of the product set; is the semantic feature of the product set; is the trainable parameter matrix, is the bias vector; is the Sigmoid activation function; Divide multimodal fusion features into user representations And product display , the inner product is used to determine the interaction score between users and items.

7. A multimodal feature fusion system, characterized in that: include: The enhancement module is configured to encode the acquired multimodal data of the commodity to obtain multimodal original features, and perform modal semantic enhancement according to the prompt words of each modality to obtain multimodal semantic features; The similarity processing module is configured to calculate the original feature similarity matrix and the semantic feature similarity matrix between different commodities in each modality according to the multimodal original features and the multimodal semantic features, and only retain the first K similarities for each commodity, thereby obtaining the sparse original feature similarity matrix and the semantic feature similarity matrix, and performing normalization processing respectively; An aggregation module is configured to aggregate the original features and semantic features of each modality according to the normalized original feature similarity matrix and the semantic feature similarity matrix to obtain the original features of the product set and the semantic features of the product set, and obtain the embedded features of the product set ID according to the interaction between the user and the product; The fusion module is configured to extract modal preferences from the embedded features of the product set ID, combine the original features of the product set under each modality and the semantic features of the product set to obtain the overall multimodal original features, multimodal semantic features and ID features, and concatenate the three to obtain multimodal fusion features, so as to predict the interaction score between users and products based on the multimodal fusion features and make product recommendations to users.

8. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 6 is completed.

9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 6.

10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Visual perception recommendation method and system based on cross-modal semantic reasoning and fusion

    CN114936901A

  • Graph Transform article recommendation method based on multi-modal semantic fusion

    CN118657560A

  • Unsupervised cross-modal retrieval method and system based on hypergraph convolution, medium and equipment

    CN118916497A

  • Multi-modal recommendation method and system for enhancing user representation through graph convolutional neural network

    CN119202398A

  • Visual text interaction-oriented multi-modal data fusion method and system

    CN119203021A