A Multimodal Collaborative Recommendation Method, System, Device, Medium and Program Product

Feature representation is enhanced through large language models and modal purification mechanisms, combined with multi-view interactive modeling and dynamic modal preference gating mechanisms, the problems of semantic information mining and noise filtering in multi-modal recommendations are solved, and personalized feature fusion and recommendation performance improvements are achieved.

CN120070010BActive Publication Date: 2025-07-15SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510541169.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-15
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing multimodal recommendation method is difficult to fully mine the rich semantic information of multimodal data, fails to effectively filter noise, and is difficult to adapt to the user's modal preference differences in different scenarios.

Method used

Through large language models, enhance user portraits and product attribute characteristics, introduce modal reference vectors for purification, build a multi-view interaction model and user-product two-part diagram, and combine differential perception attention mechanisms and dynamic modal preference gating mechanisms to integrate features.

Benefits of technology

It improves feature expression capabilities, effectively filters noise, realizes comprehensive modeling of user preferences and personalized features, significantly improving the performance of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070010B_ABST
    Figure CN120070010B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal collaborative recommendation method, system, device, medium and program product, which relates to the technical field of multi-modal recommendation, and includes: using a large language model to enhance the features of user portraits and commodity attributes, and obtaining a more comprehensive and accurate text modal representation by generating richer semantic descriptions; introducing a modal reference vector, calculating the similarity between modal features and the modal reference vector to evaluate the importance of different modalities, thereby filtering modal noise, constructing a user-commodity bipartite graph and a commodity-commodity similarity graph based on modal features, and capturing user preferences from user-commodity interactions and commodity-commodity semantic relationships; by integrating a difference-aware attention mechanism and a dynamic modal preference gating mechanism, realizing fine-grained fusion of multi-modal features, enabling the model to adaptively adjust the importance of each modality according to different scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal recommendation, and in particular to a multimodal collaborative recommendation method, system, device, medium and program product. Background Art

[0002] With the continuous development of e-commerce platforms, recommendation systems play an increasingly important role in helping users discover relevant content. Traditional collaborative filtering methods mainly rely on historical user-item interaction data to generate recommendation results. However, it is often difficult to accurately grasp the deep preferences of users based solely on interaction data. In recent years, the widespread availability of multimodal data has prompted researchers to explore integrating multiple modal information such as images and text into recommendation frameworks to improve accuracy.

[0003] Existing multimodal recommendation methods still face the following limitations:

[0004] First, relying solely on limited raw features will hinder the full mining and utilization of rich semantic information in multimodal data.

[0005] Secondly, failure to effectively filter noise in multimodal data may lead to degraded recommendation performance. For example, background and brightness interference in visual data and redundant information in text data will significantly affect the quality of feature learning.

[0006] Thirdly, simple modal feature fusion strategies are difficult to solve the problem of different users' preferences for different modalities. Studies have shown that users' modal preferences vary significantly in different scenarios. For example, when browsing clothing products, users may rely more on visual features, while when buying books, they pay more attention to text descriptions. Although some existing methods have adopted an adaptive fusion mechanism based on behavior perception, these feature fusion methods still use a unified strategy and cannot adapt to such dynamic preference characteristics.

[0007] Therefore, how to effectively use large language models to enhance feature expression, how to deal with complex multi-source modal noise, and how to achieve personalized feature fusion are key issues that need to be urgently addressed in the field of multimodal recommendation. Summary of the invention

[0008] In order to solve the above problems, the present invention proposes a multimodal collaborative recommendation method, system, device, medium and program product, which enhance the features of user portraits and product attributes and effectively improve the feature expression ability; through multi-view interactive modeling and similarity-based modal purification mechanism, the noise problem in multimodal data is effectively solved, and comprehensive modeling of user preferences is achieved at the same time; by integrating the difference-aware attention mechanism and the dynamic modal preference gating mechanism, the fine-grained fusion of multimodal features is achieved, and the importance of each modality of different products can be adaptively adjusted according to different scenarios.

[0009] In order to achieve the above object, the present invention adopts the following technical solution:

[0010] In a first aspect, the present invention provides a multimodal collaborative recommendation method, comprising:

[0011] Generate user portrait representation based on user historical behavior data, and generate text modal features based on product original attributes;

[0012] Calculate the similarity between the visual modal features and textual modal features of the product and the reference vector of the corresponding modality, obtain the purification weight according to the similarity, weight each modal feature based on the purification weight to obtain the multimodal purified feature, and obtain the fused modal purified feature after secondary purification of the multimodal purified feature according to the cross-modal similarity;

[0013] After integrating the user portrait representation with the text modal features, a modal similarity graph is constructed based on the purified features of different products in each modality to determine the multimodal features. A user-product bipartite graph is constructed based on the user's historical behavior data, and the user-product interaction features are obtained after multi-layer graph convolution and cross-layer aggregation.

[0014] The gating vector is calculated according to the user-product interaction characteristics, and the attention weight is calculated according to the difference between the multimodal features and the fused modal purification features. The multimodal features and the fused modal purification features are weighted, and the weighted modal features are aggregated based on the gating vector to obtain the final fused features to be tested, so as to obtain the product recommendation results for the user.

[0015] As an optional implementation, the process of generating a user portrait representation based on the user's historical behavior data includes:

[0016] Convert user u's historical interaction records into structured text : ;in, Indicates the product ID. Indicates the interaction timestamp;

[0017] Use large language models to generate user profile features for structured text : ;

[0018] The user portrait features are linearly projected to obtain the user portrait representation : ;in, is the learnable transformation matrix, is the bias vector;

[0019] The process of generating text modal features based on the original attributes of the product includes:

[0020] Using a large language model, based on the original attributes of the product to generate new attributes of the product : ; : ;

[0021] By combining the original attributes of the product and the new attributes of the product to obtain enhanced text modality features : ;

[0022] As an alternative implementation, the process of obtaining multi-modal purification features and fusion-modal purification features includes:

[0023] Mapping the modality features to a unified space using a linear transformation to obtain the linearly transformed modality features : ; is the weight matrix; is the modality feature of the product before linear transformation ; is the bias vector;

[0024] Calculating the cosine similarity between the linearly transformed modality features and the corresponding modality reference vectors : ; ;

[0025] Obtaining the purification weights based on the cosine similarity : ;

[0026] Multiplying the linearly transformed modality features and the purification weights to obtain the multi-modal purification features : , , represent the visual modality and the text modality respectively;

[0027] Calculating the cross-modal similarity to perform secondary purification on the multi-modal purification features to obtain the fusion-modal purification features ; ; ; ; where, is the purified visual feature; is the purified text feature; is the sigmoid activation function; is the purified cross-modal weight.

[0028] As an alternative implementation, the integration process of the user portrait representation and the text modal features is as follows: where: is the user embedding; is the learnable fusion weight; is the user portrait representation; is the interaction feature aggregated from the text modal feature embedding;

[0029] Modal similarity graph is where: and are the purified features of commodity and commodity in the modality ; , representing the visual, text, and fusion modalities respectively;

[0030] Multi-modal features is where: represents the purified feature; is the modal similarity graph of the corresponding modality; , representing the visual modality, text modality, and fusion modality respectively.

[0031] As an alternative implementation, an adjacency matrix A is created according to the number of users and the number of commodities. Based on the user historical behavior data, each interaction record of user-commodity in the adjacency matrix A is set to 1 at the corresponding position, and thus a user-commodity bipartite graph is obtained. Based on this, the normalized adjacency matrix of the user-commodity bipartite graph , , is the degree matrix of the user-commodity bipartite graph;

[0032] The process of obtaining the interaction feature of user-commodity through multi-layer graph convolution cross-layer aggregation is as follows: ; where: is the number of layers of the graph convolutional network, is the total number of layers of the graph convolutional network, is the node feature vector of the -th layer; is the node feature vector of the

[0033] As an alternative embodiment, the process of weighting the multi-modal features and the fused-modal purification features is as follows: the visual-modal features and the text-modal features in the multi-modal features are concatenated with the fused-modal purification features and then given attention weights to obtain weighted visual features and weighted text features;

[0034] According to the user-item interaction features The calculated gating vector is :

[0035] ;

[0036] In the formula: , and are the multi-layer perceptrons corresponding to the visual modality, the text modality, and the fused modality respectively; is the feature concatenation operation; is the sigmoid activation function;

[0037] The gating vector is decomposed into the gating vectors corresponding to the visual modality, the text modality, and the fused modality. The weighted visual features, the weighted text features, and the fused-modal purification features are weighted based on their respective corresponding gating vectors and then averaged for normalization, thereby completing the aggregation. The fused features after weighted aggregation are concatenated with the user-item interaction features to obtain the final fused features to be measured. The fused features to be measured are separated into user embedding and item embedding . Calculate the inner product of the user embedding and the item embedding as the user's interest score for the item, and thus obtain the item recommendation result for the user.

[0038] In a second aspect, the present invention provides a multi-modal collaborative recommendation system, including:

[0039] A feature enhancement module, configured to generate a user portrait representation according to user historical behavior data and generate text-modal features according to the original attributes of items;

[0040] A modality purification module, configured to calculate the similarity between the visual-modal features and the text-modal features of an item and the reference vectors of the corresponding modalities, obtain the purification weights according to the similarities, weight each modality feature based on the purification weights to obtain multi-modal purification features, and perform secondary purification on the multi-modal purification features according to the cross-modal similarity to obtain fused-modal purification features;

[0041] The multi-view interaction modeling module is configured to integrate the user portrait representation with the text modal features, and then construct a modal similarity graph based on the purified features of different products in each modality to determine the multi-modal features; construct a user-product bipartite graph based on the user's historical behavior data, and obtain the user-product interaction features after multi-layer graph convolution and cross-layer aggregation;

[0042] The feature aggregation module is configured to calculate the gating vector based on the user-product interaction features, calculate the attention weight based on the difference between the multimodal features and the fused modal purification features, weight the multimodal features and the fused modal purification features, aggregate the obtained weighted modal features based on the gating vector, and obtain the final fused features to be tested, so as to obtain the product recommendation results for the user.

[0043] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in the first aspect is performed.

[0044] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, wherein when the computer instructions are executed by a processor, the method described in the first aspect is performed.

[0045] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in the first aspect.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] The present invention proposes a multimodal collaborative recommendation method, system, device, medium and program product, which utilizes a large language model to enhance the features of user portraits and product attributes, obtains more comprehensive and accurate text modal features by generating richer semantic descriptions, and effectively improves the feature expression capability.

[0048] The present invention proposes a multimodal collaborative recommendation method, system, device, medium and program product, which effectively solve the noise problem in multimodal data through multi-view interaction modeling and a modal purification mechanism based on similarity, and realize comprehensive modeling of user preferences at the same time; wherein, based on multi-view interaction modeling, a user-item bipartite graph and an item-item similarity graph based on modal purification features are constructed to capture the user's complex preference patterns from two perspectives: user-item interaction and item-item semantic relationship; through a similarity-aware modal purification mechanism, the modal features are denoised, and by introducing a learnable modal reference vector, the similarity between the modal features and the modal reference vector is calculated to realize the evaluation of the importance of different modalities.

[0049] The present invention provides a multi-modal collaborative recommendation method, system, device, medium and program product, and proposes feature fusion based on multi-modal adaptation. By integrating a difference-aware attention mechanism and a dynamic modal preference gating mechanism, fine-grained fusion of multi-modal features is achieved, and the importance of each modality of different commodities can be adaptively adjusted according to different scenarios. Among them, the difference-aware attention mechanism evaluates the unique contribution of each modality by calculating the difference between the modal feature and the fused feature; the dynamic modal preference gating mechanism adjusts the importance of each modality in different scenarios based on collaborative feature learning to achieve personalized feature fusion, significantly improving the performance of the recommendation system.

[0050] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0052] Figure 1 It is a principle flow chart of the multi-modal collaborative recommendation method provided in Embodiment 1 of the present invention;

[0053] Figure 2 It is a schematic diagram of the structure principle of the multi-modal collaborative recommendation system provided in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following will further illustrate the present invention in conjunction with the drawings and embodiments.

[0055] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0056] It should be noted that the terms used herein are only for describing specific embodiments, and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "comprise" and any variation are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0057] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0058] Example 1

[0059] This embodiment relies on the Shandong Provincial Key Laboratory of Artificial Intelligence Applications in People's Livelihood Services and the Shandong Provincial Higher Education Future Industry Laboratory of General Artificial Intelligence to provide a multimodal collaborative recommendation method based on large language models (LLMs) enhancement;

[0060] like Figure 1 As shown, including:

[0061] Generate user portrait representation based on user historical behavior data, and generate text modal features based on product original attributes;

[0062] Calculate the similarity between the visual modal features and textual modal features of the product and the reference vector of the corresponding modality, obtain the purification weight according to the similarity, weight each modal feature based on the purification weight to obtain the multimodal purified feature, and obtain the fused modal purified feature after secondary purification of the multimodal purified feature according to the cross-modal similarity;

[0063] After integrating the user portrait representation with the text modal features, a modal similarity graph is constructed based on the purified features of different products in each modality to determine the multimodal features. A user-product bipartite graph is constructed based on the user's historical behavior data, and the user-product interaction features are obtained after multi-layer graph convolution and cross-layer aggregation.

[0064] The gating vector is calculated according to the user-product interaction characteristics, and the attention weight is calculated according to the difference between the multimodal features and the fused modal purification features. The multimodal features and the fused modal purification features are weighted, and the weighted modal features are aggregated based on the gating vector to obtain the final fused features to be tested, so as to obtain the product recommendation results for the user.

[0065] The multimodal collaborative recommendation method of this embodiment is described in detail below.

[0066] S1: Enhance the features of user portraits and product attributes using large language models, and obtain a more comprehensive and accurate text modality representation by generating richer semantic descriptions.

[0067] Specifically, it includes:

[0068] S1-1: Obtain the historical behavior data of user u, where the historical behavior data is the interaction between user u and product at a certain moment; through user behavior sequence conversion, convert the historical interaction records of user u into structured text : :

[0069] ;

[0070] Among them, represents the product ID, represents the interaction timestamp.

[0071] S1-2: Use a large language model to enhance the structured text, and generate enhanced user portrait features through the semantic understanding ability of the large language model:

[0072] ;

[0073] In the formula: represents the user portrait feature.

[0074] S1-3: Obtain the final user portrait representation by linearly projecting the user portrait feature :

[0075] ;

[0076] In the formula: is a learnable transformation matrix, is a bias vector.

[0077] S1-4: For product attribute enhancement, use a large language model to generate new attributes of product based on the original attributes of product ; the new attributes of product are supplementary descriptive text and category information, in order to obtain a more rich and accurate text modality representation; product attributes such as product descriptions, titles, comments, etc.

[0078] Among them, the new attributes of product are ; in the formula: represents the original attributes of product ;

[0079] By combining products Original attributes and generated products Text modality features with enhanced new attributes :[[]] ; Among them, represents the text modality; represents the product text modality features.

[0080] It can be understood that after obtaining the historical behavior data of the user and the original attributes of the product, preprocessing operations such as data cleaning, missing value filling, and feature standardization can be performed on them.

[0081] S2: To effectively process the noise information in multi-modal data, this embodiment proposes a modality purification mechanism based on similarity perception. By introducing a learnable modality reference vector, the similarity between each modality feature and the modality reference vector of the corresponding modality is calculated to achieve the evaluation of the importance of each modality feature and noise filtering.

[0082] Specifically, it includes:

[0083] S2-1: Introduce the modality reference vector , among them, , respectively represent the visual modality and the text modality.

[0084] S2-2: To effectively compare each modality feature with the corresponding modality reference vector, the modality feature is mapped to a unified space through a linear transformation to obtain the linearly transformed modality feature :[[]]

[0085] ;

[0086] Among them, is the weight matrix; is the product before linear transformation modality features; is the bias vector.

[0087] S2-3: Calculate the cosine similarity between the linearly transformed modality feature and the corresponding modality reference vector :[[]] .

[0088] S2-4: Perform a sigmoid transformation on the cosine similarity to obtain the purification weight :[[]]

[0089] ;

[0090] Among them, is the sigmoid activation function.

[0091] S2-5: Weight the modal features through the soft purification mechanism, that is, multiply the linearly transformed modal features and the purification weights element-wise to perform feature purification and obtain the purified features :

[0092] .

[0093] S2-6: To further improve the quality of the purified features, this embodiment also utilizes the complementarity between different modalities to perform secondary purification on the purified features by calculating the cross-modal similarity to obtain the fused modal purified features ;

[0094] ;

[0095] ;

[0096] ;

[0097] where is the purified visual feature; is the purified text feature; is the sigmoid activation function; is the purified cross-modal weight.

[0098] In this embodiment, the process of secondary purification based on cross-modal similarity has two important functions: (1) helping to identify features that are important in different modalities and providing an additional noise filtering layer; (2) promoting the fusion of complementary information between different modalities to form a more comprehensive and robust feature representation.

[0099] S3: Based on multi-view interaction modeling. To comprehensively capture user preferences, it is necessary to consider not only direct user-item interactions but also the semantic relationships between items. Specifically: construct a user-item bipartite graph and an item-item similarity graph based on the purified modal features; model user preferences from two perspectives: user-item interaction and item-item semantic relationship; from the user-item perspective, construct a user-item bipartite graph and capture high-order connection patterns through a graph convolutional network; from the item-item perspective, construct an item-item similarity graph based on the purified features of the modality, that is, based on the purified visual features , the purified text features and the fused modal purified features Construct the commodity-commodity similarity graph for each corresponding modality respectively, and obtain the semantic association between commodities through feature propagation. For the text modality, additionally integrate the user portrait features enhanced by the large language model.

[0100] Specifically, it includes:

[0101] S3-1: In the commodity-commodity view, for each modality , based on the purified features (i.e., purified visual features , purified text features and fused modality purified features ), construct the modality similarity graph :

[0102] ;

[0103] Where: and are the purified features of commodity and commodity in modality ; , representing the visual, text, and fused modalities respectively.

[0104] Meanwhile, to reduce noise and computational cost, in the modality similarity graph, only the top edges with the maximum similarity for each commodity are retained.

[0105] For each modality (representing the visual, text, and fused modalities respectively), through the corresponding modality similarity graph for feature propagation, obtain the enhanced commodity multi-modal features of the similarity graph after graph convolution on the purified features , , :

[0106] ;

[0107] Where: represents the purified features, including purified visual features , purified text features and fused modality purified features ; is the modality similarity graph of the corresponding modality;

[0108] Among them, in the enhanced commodity multi-modal features , different from the visual modality and the fused modality, the text modality features also need to integrate the user portrait representation enhanced by the large language model, that is, calculate the text modality features through adaptive fusionUser embedding :

[0109] ;

[0110] where: is the learnable fusion weight; is the user profile representation enhanced by the large language model; represents the interaction features aggregated from the text modality feature embeddings. The text modality features of the product such as title, attributes, reviews, etc.) are converted into high-dimensional vector representations (i.e., embeddings), and these vectors are aggregated into comprehensive interaction features to capture the multi-dimensional semantic information of the product.

[0111] S3-2: In the user-product view, create an adjacency matrix A of size , where is the number of users, is the number of products; According to the user historical behavior data, traverse each interaction record of user-product , and set the corresponding position in the adjacency matrix A to 1, indicating that there is an interaction between user and product . Thus, a user-product bipartite graph is obtained; and construct the normalized adjacency matrix of the user-product bipartite graph:

[0112] ;

[0113] where: is the adjacency matrix of the user-product bipartite graph, is the degree matrix of the user-product bipartite graph.

[0114] The vector representation obtained by concatenating or combining the features of users and products is the initial node feature . This combination method allows the model to consider the features of both users and products simultaneously during the graph convolution process, thus more comprehensively capturing the interaction relationship between users and products. The initial node feature is:

[0115] ;

[0116] where is the initial feature vector of the user; is the initial feature vector of the product.

[0117] To capture high-order connection patterns, message passing is performed through multi-layer graph convolution:

[0118] ;

[0119] In the formula: represents the number of layers of the graph convolutional network, is the node feature vector of the -th layer; is the node feature vector of the

[0120] By cross-layer aggregation, information of different interaction orders is retained, and finally the user-item interaction feature is obtained:

[0121] ;

[0122] Among them, is the total number of layers of the graph convolutional network.

[0123] S4: Feature fusion based on multi-modal adaptation. By integrating the difference-aware attention mechanism and the dynamic modality preference gating mechanism, fine-grained fusion of multi-modal features is achieved, and the importance of each modality of different items can be adaptively adjusted according to different scenarios. Among them, the difference-aware attention mechanism evaluates the unique contribution of each modality by calculating the difference between the purified visual feature and the purified text feature and the fused modality purified feature ; the dynamic modality preference gating mechanism adjusts the importance of each modality under different scenarios based on collaborative feature learning to achieve personalized feature fusion.

[0124] Specifically, it includes:

[0125] S4-1: To solve the problem that traditional feature fusion methods are difficult to capture the complex complementary relationship between different modalities, this embodiment proposes a difference-aware attention mechanism:

[0126] ;

[0127] In the formula: represents the initial fusion feature, that is, the fused modality purified feature ; is the enhanced multi-modal feature, ; is the difference vector.

[0128] The difference vector is converted into an attention score through a modality-specific multi-layer perceptron (MLP):

[0129] ;

[0130] 。

[0131] Among them, is the mapping result of the difference vector by a modality-specific multi-layer perceptron (MLP), and is the intermediate value for generating the attention score .

[0132] An additive attention mechanism is adopted to obtain weighted modality features:

[0133] .

[0134] This design creates direct paths for both shared information (through ) and modality-specific information (through ).

[0135] S4-2: To capture user modality preferences, a dynamic modality preference gating mechanism based on collaborative embedding is introduced to calculate the gating vector :

[0136] ;

[0137] In the formula: represents the collaborative embedding feature of user-item interaction obtained in step S3-2; , and are the multi-layer perceptrons corresponding to the visual modality, text modality, and fusion modality respectively; represents the feature concatenation operation; is the sigmoid activation function.

[0138] By learning the collaborative signal, the gating mechanism can identify and adapt to user preference patterns in different scenarios.

[0139] ;

[0140] In the formula: The operation decomposes the unified gating vector into modality-specific gating vectors, , and correspond to the importance weights of the visual modality, text modality, and fusion modality respectively, that is, the modality-specific gating vectors obtained by decomposition, to achieve fine-grained control of the contribution degree of each modality.

[0141] The final fusion feature is obtained through weighted aggregation:

[0142] ;

[0143] In the formula: and respectively represent the weighted visual features and weighted text features processed by the attention mechanism; is the initial fusion feature; it is divided by 3 for normalization to ensure the stability of training and maintain the proportional relationship of relative importance.

[0144] By introducing residual connections to maintain collaborative and content-based signals, a fused representation is obtained :

[0145] .

[0146] Finally, the fused representation is separated into user embeddings and item embeddings :

[0147] .

[0148] Among them, is the number of users; is the number of items;

[0149] In this embodiment, the ultimate goal is to perform product recommendations. Then, by calculating the inner product of the user embedding and the item embedding as the interest score of the user for the product. This method assumes that the higher the similarity between the user embedding and the item embedding, the greater the user's interest in the product. Finally, several products with the highest interest scores are selected and recommended to the user.

[0150] In this embodiment, a dataset from a real e-commerce scenario is used for verification. Specifically, two categories of a platform's product review dataset are used: baby products (Baby) and sports and outdoors (Sports). The basic statistical information is shown in Table 1.

[0151] Table 1 Dataset statistical information;

[0152] .

[0153] Model evaluation uses common recommendation system evaluation metrics, including R@20 (Recall@20, the recall rate of the top 20 retrieval results) and N@20 (NDCG@20, which refers to evaluating the quality of the top 20 retrieval results when using the NDCG (Normalized Discounted Cumulative Gain) metric).

[0154] Table 2 presents the performance comparison results between the method of this embodiment and existing methods such as BPR (Bayesian Personalized Ranking), LightGCN (Light Graph Convolution Network), LATTICE (Lattice Long Short-Term Memory), BM3 (Bootstrap Latent Representations for Multi-modal Recommendation), FREEDOM (Freezing and Denoising Graph Structures for Multimodal Recommendation), VBPR (Visual Bayesian Personalized Ranking), MMGCN (Multi-modal Graph Convolutional Network), and MGCN (Multi-View Graph Convolutional Network).

[0155] Table 2 Comparison of experimental results;

[0156] 。

[0157] Based on the results in Table 2, it can be seen that the multi-modal collaborative recommendation method proposed in this embodiment performs best in all evaluation metrics. Compared with existing methods, this embodiment obtains richer semantic representations through large language model feature enhancement, effectively filters noise information through similarity-aware modal purification, and achieves more accurate feature fusion through multi-modal adaptive feature aggregation, thus significantly improving the performance of the recommendation system.

[0158] It should be noted that the acquisition of all data is based on compliance with laws and regulations and user consent, and the data is legally applied.

[0159] Embodiment 2

[0160] As Figure 2 shown, this embodiment provides a multi-modal collaborative recommendation system, including:

[0161] A feature enhancement module, configured to generate a user profile representation according to user historical behavior data and generate text modal features according to the original attributes of commodities;

[0162] The modality purification module is configured to calculate the similarity between the visual modality features and text modality features of the product and the reference vector of the corresponding modality, obtain the purification weight according to the similarity, weight each modality feature based on the purification weight to obtain the multimodal purification feature, and obtain the fusion modality purification feature after secondary purification of the multimodal purification feature according to the cross-modality similarity;

[0163] The multi-view interaction modeling module is configured to integrate the user portrait representation with the text modal features, and then construct a modal similarity graph based on the purified features of different products in each modality to determine the multi-modal features; construct a user-product bipartite graph based on the user's historical behavior data, and obtain the user-product interaction features after multi-layer graph convolution and cross-layer aggregation;

[0164] The feature aggregation module is configured to calculate the gating vector based on the user-product interaction features, calculate the attention weight based on the difference between the multimodal features and the fused modal purification features, weight the multimodal features and the fused modal purification features, aggregate the obtained weighted modal features based on the gating vector, and obtain the final fused features to be tested, so as to obtain the product recommendation results for the user.

[0165] In this embodiment, a preprocessing module is also included, which is configured to obtain the user's historical behavior data and the original attributes of the goods, and then perform preprocessing operations such as data cleaning, missing value completion and feature standardization.

[0166] In this embodiment, after the feature aggregation module, the purpose is to predict the user's interest score for the product in order to make product recommendations. Then, the user embedding and product embedding are obtained according to the final fusion feature to be tested, and the inner product of the user embedding and the product embedding is calculated as the user's interest score for the product. Finally, several products with the highest interest scores are selected and recommended to the user.

[0167] It should be noted that the above modules correspond to the steps described in Example 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.

[0168] In further embodiments, there is also provided:

[0169] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in Embodiment 1 is performed. For the sake of brevity, it will not be described in detail here.

[0170] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0171] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0172] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.

[0173] The method in Embodiment 1 can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0174] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented.

[0175] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules may be combined or divided as needed. The machine-executable instructions for program modules may be executed locally or within a distributed device. In a distributed device, program modules may be located in local and remote storage media.

[0176] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to the processors of general-purpose computers, special-purpose computers, or other programmable data processing devices, such that when the program codes are executed by the computers or other programmable data processing devices, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0177] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0178] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0179] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A multi-modal collaborative recommendation method, characterized in that, include: Generate user portrait representation based on user historical behavior data, and generate text modal features based on product original attributes; Calculate the similarity between the visual modal features and textual modal features of the product and the reference vector of the corresponding modality, obtain the purification weight according to the similarity, weight each modal feature based on the purification weight to obtain the multimodal purified feature, and obtain the fused modal purified feature after secondary purification of the multimodal purified feature according to the cross-modal similarity; After integrating the user portrait representation with the text modal features, a modal similarity graph is constructed based on the purified features of different products in each modality to determine the multimodal features. A user-product bipartite graph is constructed based on the user's historical behavior data, and the user-product interaction features are obtained after multi-layer graph convolution and cross-layer aggregation. The gating vector is calculated according to the user-product interaction characteristics, and the attention weight is calculated according to the difference between the multimodal features and the fused modal purification features. The multimodal features and the fused modal purification features are weighted, and the weighted modal features are aggregated based on the gating vector to obtain the final fused features to be tested, so as to obtain the product recommendation results for the user.

2. The multimodal collaborative recommendation method according to claim 1, characterized in that The process of generating a user portrait representation based on user historical behavior data includes: Convert the historical interaction record of user u into structured text : ; where represents the product ID represents the interaction timestamp; Generating user portrait features from structured text using large language models : ; The user portrait features are linearly projected to obtain the user portrait representation : ; where is a learnable transformation matrix, is a bias vector; The process of generating text modal features based on the original attributes of the product includes: Using large language models, based on the original attributes of the product to generate new attributes of the product : ; ​​ By combining the original attributes of the product and the new attributes of the product to obtain enhanced text modal features : .

3. The multimodal collaborative recommendation method according to claim 1, characterized in that The process of obtaining multi-modal purification features and fusion modal purification features includes: Map the modal features to a unified space by linear transformation to obtain the modal features after linear transformation : ; is the weight matrix; is the modal feature of the product before linear transformation ; is the bias vector; Calculate the cosine similarity between the modal characteristics after linear transformation and the corresponding modal reference vectors and : ; Obtain the purification weight according to the cosine similarity : ; The linearly transformed modal features and the purification weights are multiplied to obtain the multi-modal purification features : , , represent the visual modality and the text modality respectively; Calculate cross-modal similarity to perform secondary purification on the multi-modal purification features and obtain the fused modal purification features ; ; ; ; where is the purified visual feature; is the purified text feature; is the sigmoid activation function; is the purified cross-modal weight.

4. The multimodal collaborative recommendation method according to claim 1, characterized in that: The integration process of the user portrait representation and the text modality features is as follows: ; where: is the user embedding; is the learnable fusion weight; is the user portrait representation; is the interaction feature aggregated from the text modality feature embedding; Modal similarity graph is ; where: and are the purification characteristics of product and product in the mode ; Multi-modal features is ; where: represents the purification feature; is the modal similarity graph of the corresponding modality; , respectively represent the visual modality, the text modality, and the fusion modality.

5. The multimodal collaborative recommendation method according to claim 1, characterized in that: Create an adjacency matrix A based on the number of users and the number of items. According to the user's historical behavior data, set the corresponding position of each interaction record of user-item in the adjacency matrix A to 1, thereby obtaining a user-item bipartite graph, and based on this, determine the normalized adjacency matrix of the user-item bipartite graph , , is the degree matrix of the user-item bipartite graph; The interaction features of users and commodities are obtained through multi-layer graph convolution cross-layer aggregation The process is as follows: ; ; In the formula: is the number of layers of the graph convolutional network, is the total number of layers of the graph convolutional network, is the layer node feature vector; is layer node feature vector.

6. A multimodal collaborative recommendation method as claimed in claim 1, characterized in that: The process of weighting the multimodal features and the fused modal purification features is as follows: the visual modal features and text modal features in the multimodal features are concatenated with the fused modal purification features and then the attention weights are assigned to obtain the weighted visual features and weighted text features; According to the user-item interaction characteristics The calculated gating vector is : ; Wherein: , and are multi-layer perceptrons corresponding to the visual modality, the text modality, and the fusion modality, respectively; is a feature concatenation operation; is a sigmoid activation function; Decompose the gating vector into gating vectors corresponding to the visual modality, the text modality, and the fusion modality. After weighting the weighted visual features, the weighted text features, and the purified features of the fusion modality based on their respective corresponding gating vectors and then averaging them for normalization, aggregation is thus completed. Concatenate the weighted aggregated fusion features with the user-item interaction features to obtain the final fusion features to be measured. Separate the fusion features to be measured into user embeddings and item embeddings . Calculate the inner product of the user embeddings and the item embeddings as the user's interest score for the item, and thus obtain the item recommendation results for the user.

7. A multimodal collaborative recommendation system, characterized in that, include: A feature enhancement module is configured to generate a user portrait representation based on the user's historical behavior data and to generate a text modality feature based on the original attributes of the product; The modality purification module is configured to calculate the similarity between the visual modality features and text modality features of the product and the reference vector of the corresponding modality, obtain the purification weight according to the similarity, weight each modality feature based on the purification weight to obtain the multimodal purification feature, and obtain the fusion modality purification feature after secondary purification of the multimodal purification feature according to the cross-modality similarity; The multi-view interaction modeling module is configured to integrate the user portrait representation with the text modal features, and then construct a modal similarity graph based on the purified features of different products in each modality to determine the multi-modal features; construct a user-product bipartite graph based on the user's historical behavior data, and obtain the user-product interaction features after multi-layer graph convolution and cross-layer aggregation; A feature aggregation module, configured to calculate a gating vector according to user-item interaction features, calculate an attention weight based on the difference between multi-modal features and fusion modal purification features, thereby weighting the multi-modal features and the fusion modal purification features, aggregating the obtained weighted modal features based on the gating vector to obtain the final fusion feature to be measured, and thereby obtaining a commodity recommendation result for the user.

8. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in any one of claims 1-6 is completed.

9. A computer-readable storage medium, characterized in that, It is used to store computer instructions. When the computer instructions are executed by the processor, the method described in any one of claims 1-6 is completed.

10. A computer program product, characterized in that, It includes a computer program. When the computer program is executed by the processor, the method described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Alzheimer's disease language detection classification system and method based on a feature purification network

    CN113961700A

  • E-commerce platform commodity recommendation method and system based on user preference analysis

    CN119398864A