A commodity recommendation method for multi-modal data

By constructing a multimodal correlation diagram and using LightGCN and hierarchical comparison learning methods, the problem of difficult to determine modal fusion weights, user interest offsets and preference noise in product recommendations is solved, and higher quality project and user preference modeling is achieved, significantly improving the accuracy of recommendations.

CN119831711BActive Publication Date: 2025-06-20MEIYI XINRUI (BEIJING) TECHNOLOGY DEVELOPMENT CO LTD

Patent Information

Application Number
CN202510329033.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-20
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The prior art faces problems in product recommendations that difficult to determine modal fusion weights, preference noise caused by user interest offsets, and lack of mining and alignment of explicit and implicit preferences of users, which limits the improvement of recommendation quality.

Method used

By constructing user-project association diagrams, project modal association diagrams, and user modal association diagrams, and using LightGCN for feature convolution, the project and user embedding are generated. Implement hierarchical comparison learning between the project and the user in each round of feature convolution to improve modal feature alignment capabilities and user preference consistency.

Benefits of technology

It effectively solves the problems of modal imbalance and user interest offset, improves the mining and alignment capabilities of users' explicit and implicit preferences, improves the quality of project and user preference modeling, and significantly improves the accuracy of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119831711B_ABST
    Figure CN119831711B_ABST
Patent Text Reader

Abstract

The present invention discloses a commodity recommendation method for multi-modal data, which relates to the technical field of commodity recommendation. The method includes constructing a user-item association graph, an item-modal association graph, and a user-modal association graph; constructing a modal intersection association graph of items and a modal intersection association graph of users; using LightGCN to perform feature convolution on the association graphs to generate embeddings for items and users, and implementing hierarchical contrast learning for items and users in each round of feature convolution; through a prediction model with joint training of multi-dimensional losses, using the form of inner product as the score of the user for the item to perform recommendation prediction; when performing convolution operation on the user-item association graph, the loss function is optimized by user preference to increase the consistency of explicit preference and implicit preference. The model of the present invention has strong scene generalization ability, effectively improves the accuracy of item and user preference modeling, and improves the recommendation quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of product recommendation, and in particular to a product recommendation method for multi-modal data. Background Art

[0002] Recommendation systems are widely applied to various e-commerce platforms. By using the historical interaction information between users and products (such as click, browse, purchase, favorite, etc.), the interest points of users are predicted, and potential products to be purchased are recommended to users.

[0003] In recent years, technicians have found that various modal information such as text, images, and videos contained in product displays all have an important impact on the recommendation quality. Different modal information has its own unique advantages in displaying product features. For example, text information is mainly used to describe the functions and materials of products, image information can well display the appearance and colors of products, and videos can more conveniently express the usage methods of products. Different user groups will show different sensitivities to different modal features of products. The formation of user preferences is often affected by the combined action of various modal information of products. Therefore, recommendation considering multiple modal information of products (i.e., multi-modal recommendation) has become a current hot technology.

[0004] In multi-modal recommendation, high-quality item and user preference modeling is the key to improving the recommendation quality. In terms of item modeling (i.e., the embedded representation of items), the current popular technology is to use an item-item association graph to assist in improving the quality of item embedding. Such an approach usually fuses the embedded representation of items in the item-item association graph in different modalities with the item embedding in the user-item graph to obtain the final item embedding.

[0005] When modeling user preferences, mainly based on obtaining the item embedding, convolution operations are performed in the user-item graph to obtain the user's embedding. At the same time, in order to improve the recommendation quality, when training the recommendation model, a joint loss function is usually constructed in a way of modal alignment and user preference enhancement to improve the recommendation quality.

[0006] Rich modal information provides useful help for users to comprehensively understand products. The above popular practices have improved the quality of item and user preference modeling, but there are still three technical problems in the recommendation process, which limit the improvement of the recommendation quality:

[0007] First, when using an item-item association graph to assist in generating item embedding, there is a problem that it is difficult to determine the modal fusion weight, which affects the quality of item embedding generation. For the modal fusion weight, whether it is a simple averaging or the method of adjusting hyperparameters, it will cause modal imbalance during item embedding, easily lead to overfitting of the training item membership scenario, and thus the trained model lacks the ability of scenario generalization.

[0008] Second, users may have interest deviation due to the display of a certain modality information of the project, thus generating preference noise and reducing the quality of user preference modeling.

[0009] Third, the existing technical solutions mainly enhance user preferences from the perspectives of interest denoising and diversity, lacking the mining and alignment of explicit and implicit preferences of users. Summary of the Invention

[0010] In order to overcome the above problems existing in the prior art, the present invention proposes a commodity recommendation method for multi-modal data.

[0011] The technical solution adopted by the present invention to solve its technical problems is: a commodity recommendation method for multi-modal data, including the following steps:

[0012] Step 1, construct a user-item association graph, an item modality association graph, and a user modality association graph;

[0013] Step 2, construct a modality intersection association graph of items and a modality intersection association graph of users;

[0014] Step 3, use LightGCN to perform feature convolution on the association graphs obtained in Step 1 and Step 2 to generate embeddings for items and users, and implement hierarchical contrast learning of items and users in each round of feature convolution;

[0015] Step 4, through a prediction model with joint training of multi-dimensional losses, use the form of inner product as the score of the user for the item to perform recommendation prediction;

[0016] In the above-mentioned item modality association graph in Step 1, it includes a text modality association graph of items and a visual modality association graph of items, and the user modality association graph includes a text modality association graph of users and a visual modality association graph of users;

[0017] When performing convolution operation on the user-item association graph in Step 3, the consistency of explicit and implicit preferences is increased through a user preference optimization loss function.

[0018] For the above-mentioned commodity recommendation method for multi-modal data, the construction process of the user-item association graph in Step 1 specifically includes: taking all users and items as nodes of the user-item association graph, and if there is an interaction relationship between a user and an item, establish an edge between the two, and finally obtain the user-item association graph 。

[0019] For the above-mentioned commodity recommendation method for multi-modal data, the construction process of the item modality association graph in Step 1 specifically includes: taking all item nodes as nodes of the association graph;

[0020] For the item and , its similarity on modality m is defined as the cosine similarity between and , denoted as , where and are the embedding representations of item and on modality m respectively; meanwhile, construct a set of top-k similar items of item on modality m;

[0021] For each item node, establish an associated edge with the nodes in the set of top-k similar items on modality to obtain the item modality association graph under modality .

[0022] The above-mentioned method for recommending products for multi-modal data, the construction process of the user modality association graph in step 1 specifically includes: taking all user nodes as the nodes of the association graph;

[0023] Calculate the initial preference embedding of the user on modality m and freeze it;

[0024] For users and , their modality preference similarity on modality m is defined as the cosine similarity between and , denoted as , where and are the embedding representations of users and on modality m; meanwhile, construct a set of top-k similar items of user on modality m;

[0025] For each user node, establish an associated edge with the nodes in the set of top-k similar items on modality m to obtain the user modality association graph under modality m.

[0026] The above-mentioned method for recommending products for multi-modal data, the construction of the modality intersection association graph of items in step 2 specifically includes: taking all item nodes as the nodes of the association graph; if two nodes and both have edges in the text modality association graph and the visual modality association graph of the item, then establish an edge in the modality intersection association graph of the item; finally, the modality intersection association graph of the item can be obtained;

[0027] Constructing the modal intersection correlation graph of users specifically includes: taking all user nodes as the nodes of the correlation graph; if two nodes and , and there are edges in both the text modal correlation graph and the visual modal correlation graph of the user, then an edge is established in the modal intersection correlation graph of the user; finally, the modal intersection correlation graph of the user can be obtained .

[0028] In the above-mentioned commodity recommendation method for multi-modal data, the user embedding generation process in step 3 includes: convolving the item embedding of the th layer through the user-item correlation graph to obtain the embedding , sending into the modal intersection correlation graph of the user for convolution to obtain the embedding , and taking the mean of and as the final item embedding of the th layer ;

[0029] In the item embedding generation process in step 3, it includes: convolving the user embedding of the th layer through the user-item correlation graph to obtain the embedding , sending into the modal intersection correlation graph of the item for convolution to obtain the embedding , and taking the mean of and as the final user embedding of the th layer ;

[0030] For the obtained user embeddings and item embeddings of layer L, the average pooling method is respectively used to obtain the final user embedding matrix and the item embedding matrix .

[0031] In the above-mentioned commodity recommendation method for multi-modal data, the hierarchical contrast learning of users in step 3 specifically includes: performing contrast learning on the user embeddings after convolution in the user-item correlation graph and the modal intersection correlation graph of the user. The specific formula is:

[0032] ;

[0033] where represents the hierarchical contrast loss of the user embeddings obtained in the user-item correlation graph and the modal intersection correlation graph of the user, and represents the The user obtained by layer convolution embedding, represents the temperature coefficient, represents the layer convolution in the user-item association graph of the user embedding, represents the layer convolution in the user-item association graph of the user represents the layer, represents a user in the user set represents another user in the user set;

[0034] For the contrastive learning of the user's modal intersection association graph and the user embedding after convolution of the user's different modal association graphs, the specific formula is:

[0035] ;

[0036] Among them, represents the hierarchical contrast loss of the embedding obtained by the user in the user's modal intersection association graph and the association graph of each modality. m represents a certain modality, represents a user in the user set, represents the modality set, represents the layer convolution in the user's m-modal association graph of the user represents the layer convolution in the user's modal intersection association graph of the user represents the layer convolution in the user's m-modal association graph of the user

[0037] The above-mentioned commodity recommendation method for multi-modal data, the hierarchical contrast learning of the item in step 3 specifically includes: performing contrastive learning on the item embedding after convolution in the user-item association graph and the item's modal intersection association graph, and the specific formula is:

[0038] ;

[0039] Among them, represents the hierarchical contrast loss of the item embedding obtained in the user-item association graph and the item's modal intersection association graph, represents the Items obtained by layer convolution Embedding Indicates the temperature coefficient Indicates in the user-item association graph at the Items obtained by layer convolution Embedding of Indicates in the user-item association graph at the Items obtained by layer convolution Embedding of Indicates the item set, L indicates the number of convolution layers Indicates the Layer;

[0040] For the item modal intersection association graph and the item embeddings after convolution of the item's different modal association graphs, contrastive learning is performed. The specific formula is:

[0041] ;

[0042] Wherein, Indicates the hierarchical contrast loss of the item embeddings obtained in the item modal intersection association graph and the item's association graph in each modality. m represents a certain modality Indicates an item in the item set Indicates another item in the item set Indicates the modality set Indicates in the item m-modal association graph at the Items obtained by layer convolution Embedding of Indicates in the item modal intersection association graph at the Items obtained by layer convolution Embedding of Indicates in the item m-modal association graph at the Items obtained by layer convolution Embedding of

[0043] The beneficial effect of the present invention is that the present invention patent constructs a multi-modal recommendation method, which uses the modal intersection graph and hierarchical contrast learning to solve the problems of modal imbalance and user interest deviation existing in item embedding in the prior art, and solves the problem of lack of mining and alignment of user explicit preferences and implicit preferences in the prior art by improving the consistency of user explicit preferences and implicit preferences. The recommendation method model of the present invention has strong scene generalization ability, improves the quality of item and user preference modeling, and effectively improves the accuracy of recommendation. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Is the schematic flow diagram of the present invention;

[0045] Figure 2It is a schematic diagram of the user embedding generation process of the present invention;

[0046] Figure 3 It is a schematic diagram of the item embedding generation process of the present invention;

[0047] Figure 4 It is a schematic diagram of the loss calculation processes in the user-level contrastive learning of the present invention;

[0048] Figure 5 It is a schematic diagram of the loss calculation processes in the item-level contrastive learning of the present invention;

[0049] Figure 6 It is a schematic diagram of the user preference optimization loss calculation of the present invention;

[0050] Figure 7 It is a schematic diagram of the relationship between the prediction model and the loss of the present invention. Detailed implementation manners

[0051] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific implementation manners. This embodiment discloses a commodity recommendation method for multi-modal data. According to the actual scenario of the recommendation, the text modality (t modality) and visual modality information (v modality) of the item are considered. Among them, the text modality features can be obtained from the literal introduction of the item, and the visual modality features are obtained from the image of the item. The present invention uses the special term "item" in the recommendation system to represent the recommended commodity.

[0052] In this embodiment, for the item text and visual modality in the recommendation system (denoted as ), the representation of the item in this modality is extracted through a pre-trained model , and the modality feature matrix of the item is , where the visual features are extracted by CNN and the text features are extracted by sentence-transformers, being the dimension of the modality features.

[0053] As Figure 1 shown, it specifically includes the following steps:

[0054] Step 1: Construct three types of association graphs, namely, user-item association graph, item modality association graph, and user modality association graph.

[0055] 1.1 The user-item association graph is established based on the interaction relationship between the user and the item. The specific steps are as follows:

[0056] (1) All users and items are used as the nodes of the user-item graph

[0057] (2) If there is an interaction relationship between the user and the item, an edge is established between the two.

[0058] Based on the above steps, construct a user-project interaction graph according to the user-project interaction history. This graph is formally defined as a bipartite graph , where:

[0059] (1) is the set of users, represents the th user.

[0060] (2) is the set of projects, represents the th project.

[0061] (3) is the set of edges, representing the interaction relationship between users and projects.

[0062] For graph , set its corresponding user-project interaction matrix as . If user interacts with project , then . Conversely, . Calculate the user-project normalized adjacency matrix as shown in formulas (1)-(2):

[0063] (1) (2);

[0064] Where, is the normalized user-project adjacency matrix, is the normalized project-user adjacency matrix, is the number of neighbors of user u, is the user-project interaction matrix, is the project 's number of neighbors;

[0065] 1.2 The modal association graph of projects includes the text modal association graph of projects and the visual modal association graph of projects. For the association graph corresponding to any one of the modalities ( is t or v), the construction steps are as follows:

[0066] (1) Take all project nodes as the nodes of the association graph.

[0067] (2) For projects and , their similarity on modality m is defined as the cosine similarity between and , denoted as . Based on this, construct a project The set of top-k similar items on modality m, see formula (3):

[0068] (3).

[0069] (3) For each item node, establish the associated edges between it and the nodes in the set of top-k similar items on modality to obtain the item-item association graph under modality .

[0070] Based on the above steps, the modality association graph of items can be formally defined as where

[0071] (1) is the set of items, represents the th item.

[0072] (2) , represents the corresponding modality of the item.

[0073] (3) is the set of edges, representing the association relationship between items under modality m.

[0074] In the modality association graph of items under modality , each node has established edges with the k nodes with the largest similarity to it on modality . Denote the adjacency matrix of as , and the normalized adjacency matrix is .

[0075] 1.3 The modality association graph of users includes the text modality association graph of users and the visual modality association graph of users. For any modality ( is or ) the corresponding association graph, and the construction steps are as follows

[0076] (1) Take all user nodes as the nodes of the association graph.

[0077] (2) First, calculate the initial preference embedding of user on modality m and freeze it, see formula (4):

[0078] (4);

[0079] where is in the adjacency matrix , user The value of the adjacency matrix corresponding to each item is the embedding matrix of the item in modality m.

[0080] (3) For a user and , the modality preference similarity on modality m is defined as the cosine similarity between , denoted as . Meanwhile, construct a top-k similar item set of user on modality m, as shown in formula (5):

[0081] (5);

[0082] where represents the set of the top k items with the highest similarity for user , and k represents the number of items.

[0083] (4) For each user node, establish associated edges with the nodes in the top-k similar item set on modality m, and the user-user association graph under modality m can be obtained.

[0084] Based on the above steps, the modality association graph of users can be formally defined as , where:

[0085] (1) is the set of users, represents the th user.

[0086] (2) , represents the corresponding modality of the item.

[0087] (3) is the set of edges, representing the association relationship between users under modality m.

[0088] In the modality user association graph under modality m, each node has established edges with the top k nodes with the highest similarity on modality m. Denote the adjacency matrix as , and the normalized adjacency matrix is .

[0089] Step 2: Construct the modality intersection association graphs of users and items respectively.

[0090] 2.1 Construct the modality intersection association graph of items. The construction method of the modality intersection association graph of items is as follows:

[0091] (1) Take all item nodes as the nodes of the association graph;

[0092] (2) If two nodes and both have edges in the text modality association graph and the visual modality association graph of the project, then an edge is established in the modality intersection association graph of the project.

[0093] Based on the above steps, the modality intersection association graph of the project can be formally defined as , and the corresponding adjacency matrix is , and the corresponding normalized adjacency matrix is , where:

[0094] (1) is the set of projects, and represents the th project.

[0095] (2) is the set of edges, representing the association relationship between projects.

[0096] 2.2 Construct the modality intersection association graph of the user. The construction method of the user's modality intersection association graph is as follows:

[0097] (1) Take all user nodes as the nodes of the association graph.

[0098] (2) If two nodes and both have edges in the text modality association graph and the visual modality association graph of the user, then an edge is established in the user's modality intersection association graph.

[0099] Based on the above steps, the user intersection association graph can be formally defined as , and the corresponding adjacency matrix is , and the normalized adjacency matrix is , where:

[0100] is the set of users, and represents the th user.

[0101] is the set of edges, representing the association relationship between users.

[0102] 3. Perform feature convolution using LightGCN under the multi-dimensional association graph to generate embeddings for projects and users.

[0103] 3.1 The user embedding of the th layer is generated as follows. First, the project embedding of the th layer is first passed through ​​Perform convolution to obtain an embedding , and then are respectively sent into the user's t-modal correlation graph , the user's v-modal correlation graph and the user's intersection correlation graph to perform convolution and respectively obtain embeddings , and , where , are respectively matrices composed of , , and are used in the hierarchical contrast learning shown in Section 4.1 later. Then, the mean of and is used as the final user embedding of the th layer . The calculation methods of each embedding are shown in Formulas (6)-(10), and the specific process is as shown in Figure 2 .

[0104] (6);

[0105] (7);

[0106] (8);

[0107] (9);

[0108] (10).

[0109] 3.2 The generation of the item embedding of the th layer is as follows. The user embedding of the rd is first convolved through to obtain an embedding , and then is respectively sent into the item's t-modal correlation graph , the item's v-modal correlation graph and the item's intersection correlation graph to perform convolution and respectively obtain embeddings , and , where , are respectively matrices composed of , , and are used in the hierarchical contrast learning shown in Section 4.2 later. Then, the mean of and is used as the final item embedding of the th layer , the respective embedding calculation methods are shown in Formulas (11)-(15), and the specific process is as follows Figure 3 .

[0110] (11);

[0111] (12);

[0112] (13);

[0113] (14);

[0114] (15).

[0115] 3.3 For the obtained L-layer user and item embeddings , , the average pooling method is adopted to obtain the final user embedding matrix and item embedding matrix , and the calculation formulas are shown in (16)(17):

[0116] (16);

[0117] (17).

[0118] 4. Implement embedding contrast for items and users respectively when performing feature convolution aggregation at different levels using hierarchical contrast learning.

[0119] 4.1 Implement hierarchical contrast learning for users in each round of feature convolution. Specific approach: Use Formula (18) to perform contrast learning on the user embeddings after graph and graph and intersection graph convolution, so as to improve the consistency of the aggregated features of users in the user-item interaction graph and the user modality intersection graph . At the same time, use Formula (19) to perform contrast learning on the user embeddings after , , graph convolution, so as to enhance the consistency of the aggregated features of users in different modality association graphs. The above-mentioned hierarchical contrast learning for users can effectively solve the problem of user preference noise caused by interest deviation during user modeling (solving the second technical problem). The schematic diagram of each loss calculation process is as Figure 4 shown:

[0120] (18);

[0121] Among them, Represents the hierarchical contrast loss of the user embeddings obtained from the user-item association graph and the user's modal intersection association graph. Represents the user obtained from the -th layer convolution in the user modal intersection association graph embedding. Represents the temperature coefficient. Represents the user obtained from the -th layer convolution in the user-item association graph embedding. Represents the user obtained from the -th layer convolution in the user-item association graph embedding. Represents the user set, L represents the number of convolution layers. Represents the -th layer. Represents a user in the user set Represents another user in the user set;

[0122] (19);

[0123] Among them, Represents the hierarchical contrast loss of the embeddings obtained by the user in the user's modal intersection association graph and the user in the association graph of each modality. m represents a certain modality. Represents a user in the user set. Represents the modality set. Represents the user obtained from the -th layer convolution in the user m-modal association graph embedding. Represents the user obtained from the -th layer convolution in the user's modal intersection association graph embedding. Represents the user obtained from the -th layer convolution in the user m-modal association graph embedding, where and are respectively the embeddings corresponding to in the embedding matrix , . According to the values t and v of the modality m, can correspond to and , and and can be obtained from step 3.1.

[0124] (20).

[0125] 4.2 Implement hierarchical contrastive learning of items in each round of feature convolution. Specific approach: Use formula (21) to perform contrastive learning on the item embeddings after intersection graph convolution in Figure and to enhance the consistency of the aggregated features of items in the user-item interaction graph and the intersection graph of item modalities . At the same time, use formula (22) to perform contrastive learning on the item embeddings after intersection, , , graph convolution to enhance the consistency of the aggregated features of items in different modality association graphs. The above hierarchical contrastive learning of items can effectively improve the multi-modal feature alignment ability during item modeling and alleviate the problem of modal feature imbalance in item embeddings caused by modal feature fusion weights (solving the first technical problem). The schematic diagram of each loss calculation process is as Figure 5 shown.

[0126] (21);

[0127] where, represents the hierarchical contrastive loss of the item embeddings obtained from the user-item association graph and the intersection association graph of item modalities, represents the item obtained from the layer convolution in the intersection association graph of item modalities, represents the temperature coefficient, represents the item obtained from the layer convolution in the user-item association graph, represents the item obtained from the layer convolution in the user-item association graph, represents the item set, L represents the number of convolution layers, represents the layer.

[0128] (22);

[0129] where, represents the hierarchical contrastive loss of the embeddings obtained from the intersection association graph of item modalities and the association graph of each modality of the item, m represents a certain modality, represents an item in the item set, represents another item in the item set, represents the modality set, represents the item obtained from the Embedding of represents the item obtained by convolution at the layer in the modal intersection correlation graph of the project, represents the item obtained by convolution at the layer in the modal correlation graph of item m, where and are the embeddings corresponding to item and item in the item embedding matrix . According to the values t and v of modal m, can correspond to and , and and can be obtained from step 3.2.

[0130] (23).

[0131] 5. Use the loss function shown in formula (24) to improve the consistency between user explicit preferences and implicit preferences, thereby improving the quality of user modeling.

[0132] In the user-item graph, through convolution operations, explicit preferences can be obtained from the item nodes directly associated with the user; implicit preferences can be obtained from the nodes that are farther away from the user (not directly associated). Implicit preferences mean the preferences that the user may generate in the future. Increasing the consistency between explicit preferences and implicit preferences is beneficial to more accurately obtaining user preferences (solving the 3rd technical problem). The loss calculation process is shown as Figure 6 shown.

[0133] Designed the following user preference optimization loss function, where is the consistency difference loss between explicit preferences and implicit preferences, see formula (24):

[0134] (24).

[0135] 6. Establish a prediction model with multi-dimensional loss joint training, and use the model shown in formula (25) for recommendation prediction. The overall loss of the model is shown in formula (27).

[0136] After obtaining the final embedding representation of the user item, use the form of inner product as the user's score for the item, see formula (25):

[0137] (25);

[0138] where represents the user 's score for item The predicted preference score, e i represents the final item embedding, e u represents the final user embedding.

[0139] In the training of the model, Bayesian Personalized Ranking Loss (BPR) was selected as the loss function, which makes the user's score for the positive item higher than that for the negative item. See Equation (26):

[0140] (26);

[0141] where is the set of training instances, and each triple satisfies that the user u has interacted with the item and has not interacted with the item and is the sigmoid function.

[0142] The overall loss function of this embodiment consists of the above four parts: , , , , see Equation (27):

[0143] + (27);

[0144] where , , γ are hyperparameters represents the total hierarchical contrast loss of items, represents the total hierarchical contrast loss of users. The relationship between the prediction model and the loss is as Figure 7 shown.

[0145] Experimental comparison:

[0146] Datasets: Three publicly available datasets, Baby, Sports, and Clothing, from the Amazon e-commerce platform were selected. Each dataset contains text and visual features. For the text features and visual features of items, they were respectively encoded into 384-dimensional text embeddings and 4096-dimensional visual embeddings.

[0147] Implementation details:

[0148] Unify the embedding vector spaces of users and projects, with the unified dimension being 64. This model is built based on the PyTorch framework, and the Adam optimizer is used to optimize the parameters. During training, the learning rate is set to 1e-3, the mini-batch size for all datasets is set to 2048, and the number of training iterations is set to 1000 rounds. To prevent the model from overfitting, an early stopping strategy is introduced. If the validation set does not improve on R@20 for 20 consecutive iterations, the training will automatically terminate.

[0149] Comparison methods:

[0150] (1) VBPR: By combining the visual features extracted by CNN with matrix factorization technology, the user preferences in the visual dimension are obtained.

[0151] (2) MMGCN: Introduce a multi-modal graph convolutional network to enhance user representation through information exchange between users and projects, thereby capturing fine-grained cross-modal preferences.

[0152] (3) LightGCN: Simplify the structure of the recommendation model by removing the non-linear activation layer and feature transformation layer of the original GCN.

[0153] (4) GRCN: This method introduces a graph-optimized convolutional network (GRCN) that adaptively adjusts the structure of the interaction graph according to the training status of the model, effectively extracting information signals about user preferences from the optimized graph.

[0154] (5) LATTICE: This method learns the item-item relationship graph for each modality, aggregates them into a latent item graph, and injects high-order affinity using graph convolution, thereby improving the accuracy of collaborative filtering recommendations.

[0155] (6) DualGNN: This method designs a multi-modal representation learning module to explicitly simulate the user's attention to different modalities and inductively learn multi-modal user preferences.

[0156] (7) FREEDOM: This method introduces a degree-sensitive edge pruning technique to eliminate potential noisy edges during the graph sampling process. It simultaneously freezes the item-item graph and denoises the user-item interaction graph to achieve multi-modal recommendations.

[0157] (8) SLMRec: This method designs data augmentation based on multi-modal content to generate multiple views of a single item and uses contrastive learning to distinguish the views of the item from other views, thereby extracting additional supervision signals.

[0158] (9) POWEREC: This method proposes a prompt-based weak modality enhanced multi-modal recommendation framework. It utilizes different modality prompts to simulate user interests and enhances the learning of user preferences in modalities with less reliable predictions.

[0159] (10) LGMRec: This method proposes a local and global graph learning guided multi-modal recommender (LGMRec) to jointly simulate local and global user interests. The global embedding is combined with two decoupled local embeddings to improve the accuracy and robustness of recommendations.

[0160] (11) SynerGraph: This method improves the modeling of user preferences by integrating knowledge with item information. Additionally, it develops a filter to remove noise in various types of data, thereby enhancing the reliability of recommendations.

[0161] Among the above comparison methods, classic methods over the years and the latest methods in recent years were selected for comparison. Among them, VBPR and MMGCN are from 2019, LightGCN and GRCN are from 2020, LATTICE and DualGNN are from 2021, SLMRec is from 2022, FREEDOM is from 2023, POWEREC, LGMRec, and SynerGraph are from 2024. Experimental comparisons were conducted, and the comparison results are shown in Tables 1 - 3.

[0162] As can be seen from Tables 1 - 3, in the four evaluation metrics of R@10, R@20, N@10, and N@20, the method of this embodiment is superior to the above baseline methods. Among them, %Improv1 is the percentage improvement compared to the sub-optimal baseline, and %Improv2 is the percentage improvement relative to the average value of all baselines. From the comparison data, it can be seen that the commodity recommendation method of this embodiment is significantly superior to the comparison methods.

[0163] Table 1

[0164]

[0165] Table 2

[0166]

[0167] Table 3

[0168]

[0169] The above embodiments are only exemplary embodiments of the present invention and are not used to limit the present invention. Those skilled in the art can make various modifications or equivalent replacements to the present invention within the essence and protection scope of the present invention, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present invention.

Claims

1. A commodity recommendation method for multimodal data, characterized in that: The steps include: Step 1, constructing a user-item association graph, an item modal association graph, and a user modal association graph; Step 2, constructing a modal intersection association graph of the project and a modal intersection association graph of the user; Step 3: Use LightGCN to perform feature convolution on the association graph obtained in steps 1 and 2 to generate embeddings for items and users, and implement hierarchical comparative learning of items and users in each round of feature convolution; Step 4: Use the prediction model with multi-dimensional loss joint training to make recommendation predictions by taking the inner product as the user's rating of the item; In the step 1, the project modality association diagram includes a text modality association diagram of the project and a visual modality association diagram of the project, and the user modality association diagram includes a text modality association diagram of the user and a visual modality association diagram of the user; When performing the convolution operation on the user-item association graph in step 3, the consistency of explicit preference and implicit preference is increased by optimizing the loss function through user preference; The process of constructing the user modality association graph in step 1 specifically includes: taking all user nodes as nodes of the association graph; Calculate the user’s initial preference embedding on modality m and freeze it; For Users and , its modality preference similarity on modality m is defined as and The cosine similarity between ,in and Is a user and Embedding representation on modality m; at the same time, construct a user The set of top-k similar items on modality m; For each user node, establish an associated edge between it and the nodes in the top-k similarity item set on modality m, and obtain the user modality association graph under modality m; The process of constructing the modal intersection association graph of the project in step 2 specifically includes: taking all project nodes as nodes of the association graph; if two nodes and , if there is an edge in both the text modal association graph of the project and the visual modal association graph of the project, then an edge is established in the modal intersection association graph of the project; the modal intersection association graph of the project is obtained ; Constructing the user's modal intersection association graph specifically includes: taking all user nodes as nodes of the association graph; if two nodes and , if there is an edge in both the user's text modal association graph and the user's visual modal association graph, then an edge is established in the user's modal intersection association graph; the user's modal intersection association graph is obtained .

2. The method for recommending products based on multimodal data according to claim 1, characterized in that: The construction process of the user-item association graph in step 1 specifically includes: taking all users and items as nodes of the user-item association graph, and if there is an interactive relationship between the user and the item, an edge is established between the two, and finally the user-item association graph is obtained. .

3. The method for recommending products based on multimodal data according to claim 1, characterized in that: The process of constructing the project modal association graph in step 1 specifically includes: taking the nodes of all projects as the nodes of the association graph; For Projects and , and its similarity on mode m is defined as and The cosine similarity of ,in and The projects are and Embedding representation on modality m; at the same time, build a project The set of top-k similar items on modality m; For each project node, establish the The associated edges between the nodes in the top-k similarity item set are obtained Next project modal association diagram .

4. The method for recommending products based on multimodal data according to claim 1, characterized in that: The user embedding generation process in step 3 includes: Layer item embedding Convolution is performed through the user-item association graph to obtain the embedding ,Will The modal intersection graph sent to the user Perform convolution to get embedding ,Will and The mean value of The final item embedding of the layer ; The project embedding generation process in step 3 includes: User embedding Convolution is performed through the user-item association graph to obtain the embedding ,Will The modal intersection association graph of the input project is convolved to obtain the embedding ,Will and The mean value of End-user embedding of layers ; For the obtained Layer item embedding and user embedding, respectively, use average pooling to obtain the final user embedding matrix and the item embedding matrix .

5. The method for recommending products based on multimodal data according to claim 1, characterized in that: The hierarchical comparative learning of users in step 3 specifically includes: performing comparative learning on the user embedding after convolution of the user-item association graph and the user's modal intersection association graph, and the specific formula is: ; in, represents the hierarchical contrast loss of the embedding obtained by the user in the user-item association graph and the user's modal intersection association graph, Indicates the first The user obtained by layer convolution Embed, represents the temperature coefficient, Indicates the first The user obtained by layer convolution Embed, Indicates the first The user obtained by layer convolution Embed, represents the user set, L represents the number of convolution layers, Indicates layer, Represents a user in the user collection Represents another user in the user set; The user embedding after convolution of the user's modal intersection association graph and the user's different modal association graph is compared and learned. The specific formula is: ; in, represents the hierarchical contrast loss of the user embedding obtained in the user's modal intersection association graph and the user association graph of each modality, m represents a certain modality, Represents a user in the user collection, represents a set of modes, Indicates the first The user obtained by layer convolution The embedding Indicates the first The user obtained by layer convolution The embedding Indicates the first The user obtained by layer convolution Embedding.

6. The method for recommending products based on multimodal data according to claim 1, characterized in that: The hierarchical comparative learning of the items in step 3 specifically includes: comparative learning of the item embedding after convolution of the user-item association graph and the modal intersection association graph of the items, and the specific formula is: ; in represents the hierarchical contrast loss of item embeddings obtained in the user-item association graph and the modal intersection association graph of items, Indicates the modal intersection association diagram of the project. Items obtained by layer convolution Embed, represents the temperature coefficient, Indicated in the user-item association graph Items obtained by layer convolution The embedding Indicated in the user-item association graph Items obtained by layer convolution The embedding represents the item set, L represents the number of convolution layers, Indicates layer; The modal intersection association graph of the project and the project embedding after convolution of the different modal association graphs of the project are compared and learned. The specific formula is: ; in, represents the hierarchical contrast loss of the item embedding obtained in the modal intersection association graph of the item and the item association graph of each modality, m represents a certain modality, Represents an item in a collection of items, Represents another item in the item collection, represents a set of modes, Indicates the first Items obtained by layer convolution The embedding Indicates the modal intersection association diagram of the project. Items obtained by layer convolution The embedding Indicates the first Items obtained by layer convolution Embedding.

Citation Information

Patent Citations

  • Multi-relation-based graph convolution collaborative filtering recommendation method, system and equipment

    CN116541612A

  • Multi-modal recommendation method and system based on intra-modal and inter-modal comparative learning

    CN118760802A

Cited By

  • A predictive analysis method and platform for quality control

    CN120410715B