Multi-modal recommendation method and device, electronic equipment, storage medium and program product
By weighted fusion of project description pictures and text, computing user preference scores with user interaction historical data, and inputting multimodal preference optimization recommendation model, the problems of low efficiency of multimodal information fusion and insufficient user preference modeling are solved, and more accurate recommendation results are achieved.
Patent Information
- Application Number
- CN202510070891.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, multimodal information fusion efficiency and insufficient user preference modeling lead to poor recommendation results.
By determining the project description picture and text, weighted fusion is performed to generate overall modal feature expression; combining the interactive historical data between the user and the project, calculate the user preference score; input the user preference score into the pre-trained multimodal preference optimization recommendation model, optimize the score, and finally generate the recommendation result.
Improve the efficiency of multimodal information fusion, capture and model user preferences more accurately, thereby improving recommendation results.
Smart Images

Figure CN120196796A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of multimodal data processing, and in particular, to a multimodal recommendation method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] This section aims to provide background or context for the embodiments of the present disclosure stated in the claims. The description herein is not admitted to be prior art merely by virtue of its inclusion in this section.
[0003] Multimodal recommendation is an advanced recommendation system technology that comprehensively understands and models the characteristics of users and items by integrating multiple data modalities (such as text, images, videos, audio, etc.); this method not only utilizes traditional user behavior data (such as clicks, purchases, ratings, etc.), but also combines rich multimodal content information, thereby being able to capture user preferences and items.
[0004] However, in the related art, there are problems of low multimodal information fusion efficiency and insufficient user preference modeling. Summary of the Invention
[0005] In view of this, an object of the present disclosure is to provide a multimodal recommendation method, apparatus, electronic device, storage medium, and program product, which at least to some extent solve one of the technical problems in the related art.
[0006] Based on the above object, in the first aspect of the exemplary embodiments of the present disclosure, a multimodal recommendation method is provided, which is applied to a server, and the method includes:
[0007] Determine a project description picture and a project description text, and perform weighted fusion on the project description picture and the project description text to obtain an overall modal feature representation;
[0008] Determine the interaction history data between the user and the item, and based on the overall modal feature representation and the interaction history data, obtain a user preference score;
[0009] Input the user preference score into a pre-trained multimodal preference optimization recommendation model to obtain an optimized user preference score;
[0010] Based on the optimized user preference score, obtain a recommendation result.
[0011] Based on the same inventive concept, in the second aspect of the exemplary embodiments of the present disclosure, a multimodal recommendation apparatus is provided, including:
[0012] A feature representation determination module, configured to determine a project description picture and a project description text, and perform weighted fusion on the project description picture and the project description text to obtain an overall modal feature representation;
[0013] A preference score determination module, configured to determine historical interaction data between a user and an item, and obtain a user preference score based on the overall modal feature representation and the historical interaction data;
[0014] An optimized preference score determination module, configured to input the user preference score into a pre-trained multi-modal preference optimization recommendation model to obtain an optimized user preference score;
[0015] A recommendation result determination module, configured to obtain a recommendation result based on the optimized user preference score.
[0016] Based on the same inventive concept, a third aspect of the exemplary embodiments of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.
[0017] Based on the same inventive concept, a fourth aspect of the exemplary embodiments of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the first aspect.
[0018] Based on the same inventive concept, a fifth aspect of the exemplary embodiments of the present disclosure provides a computer program product including computer program instructions, which when running on a computer, cause the computer to execute the method described in the first aspect.
[0019] As can be seen from the above, the multi-modal recommendation method, apparatus, electronic device, storage medium, and program product provided by the embodiments of the present disclosure include:
[0020] Determine a project description picture and a project description text, perform weighted fusion on the project description picture and the project description text to obtain an overall modal feature representation; determine historical interaction data between a user and an item, and obtain a user preference score based on the overall modal feature representation and the historical interaction data; input the user preference score into a pre-trained multi-modal preference optimization recommendation model to obtain an optimized user preference score; obtain a recommendation result based on the optimized user preference score. The present disclosure can improve the efficiency of multi-modal information fusion and more accurately capture and model user preferences. Description of the Drawings
[0021] To more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following descriptions are only embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 Schematic diagram of the application scenario of the multi-modal recommendation method provided by an exemplary embodiment of the present disclosure;
[0023] Figure 2 Schematic flow diagram of a multi-modal recommendation method provided by an exemplary embodiment of the present disclosure;
[0024] Figure 3 Schematic structural diagram of a multi-modal recommendation device provided by an exemplary embodiment of the present disclosure;
[0025] Figure 4 Schematic diagram of the hardware structure of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners
[0026] It can be understood that before using the technical solutions disclosed in the embodiments of the present application, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0027] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that executes the technical solution of the present application according to the prompt message.
[0028] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window. The prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0029] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present application. Other ways that meet relevant laws and regulations can also be applied to the implementation manner of the present application.
[0030] It can be understood that the data involved in the technical solution of the present application (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0031] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and thereby implement the present disclosure, rather than limiting the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.
[0032] In this document, it should be understood that the number of any element in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0033] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the embodiments of the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. The article "a" or "an" before an element does not exclude the existence of multiple such elements.
[0034] The principles and spirit of the present disclosure will be elaborated below with reference to several representative embodiments of the present disclosure.
[0035] As described in the background art, in the related art, there are problems of low multi-modal information fusion efficiency and insufficient user preference modeling. Specifically, in the related art, by designing a modal perception structure learning layer, an item-item relationship graph is learned from multi-modal features, and the graph structures of multiple modalities are aggregated to construct a potential multi-modal item graph. Subsequently, high-order item relationships are injected into the item representation through graph convolution operations, thereby enhancing the item embedding. Finally, these enhanced item representations can be combined with existing collaborative filtering methods to achieve more accurate recommendations.
[0036] However, it performs relatively poorly in terms of efficient fusion of multimodal information and dynamic knowledge updating, fails to fully utilize the multimodal information of the items, and cannot effectively solve the forgetting problem, which affects the generalization ability of the model. In addition, the model mainly focuses on the relationship between items and relatively little modeling of user preferences.
[0037] In related technologies, a frozen project-project relationship graph is constructed and a graph neural network is used to embed projects. Then, the user-project interaction graph is denoised by using a degree-sensitive edge pruning technique to remove possible noise edges. Ultimately, the optimized user and project embedding representations can not only efficiently complete personalized recommendation tasks, but also use these embedding representations to accurately predict user preferences and build a more accurate recommendation system.
[0038] However, this method does not discuss computational efficiency and scalability in detail, nor does it provide detailed hyperparameter tuning strategies or automated tuning methods.
[0039] In related technologies, internal and external relationships between users and projects are captured by constructing user-user graphs and project-project graphs. The specific steps include: constructing a user-project bipartite graph for each modality and learning modality-specific representations; using the attention splicing method to fuse multimodal information to ensure complementarity without losing unimodal information; constructing user co-occurrence graphs and project semantic graphs to capture internal relationships. DRAGON performs graph learning on user-project heterogeneous and homogeneous graphs, applies LightGCN to perform graph convolution operations on the user-project bipartite graphs of each modality, learns modality-specific representations, and performs graph convolution operations on user co-occurrence graphs and project semantic graphs to capture internal relationships.
[0040] However, this method relies too much on a single attention splicing fusion method, which limits the full utilization of multimodal information. Although user-user and item-item graphs are constructed, the impact of high-order relationships on recommendation performance is not fully explored.
[0041] In related technologies, a behavior-guided purifier uses behavior information to filter out noise in modal features, avoid modal noise contamination, and thus improve the purity of features. Next, high-order collaborative signals and semantic-related signals are captured in the user-item view and the item-item view, respectively, to enhance the distinguishability of features and ensure that the features not only contain the interactive information between users and items, but also capture the semantic associations between items. Then, a behavior-aware fuser is designed to adaptively fuse modal features based on the user's behavioral characteristics, ensuring that the user's preferences for different modalities are considered during the fusion process, thereby more comprehensively modeling user preferences. In addition, a self-supervised auxiliary task is introduced to maximize the mutual information between the fused multimodal features and the behavioral features, so as to simultaneously capture complementary and supplementary preference information, further improving the recommendation performance of the model.
[0042] However, this method relies too much on the purification process of behavioral information, resulting in inevitable information loss. At the same time, it does not fully consider the dynamic preference changes of users in different situations, and the stability and robustness of the effect need to be improved.
[0043] In related technologies, user embedding features, item embedding features, and multi-modal features of items (including text embedding features and image embedding features) are extracted from the interaction data between users and items. Then, a graph neural network encoder is used for interaction-based collaborative filtering and modality-based collaborative filtering, and the local embedding vectors of users and items are fused. At the same time, noise in the multi-modal features is filtered by a feature filter, and the global embedding vectors of users and items are generated by combining a learnable transformation vector and an adjacency matrix. Then, the local embedding vectors and the global embedding vectors are fused to obtain the final embedding vectors of users and items. Finally, the preference score of the user for the item is calculated based on the final embedding vector of the user and the final embedding vector of the item, and a recommendation result is generated.
[0044] However, this method has a single model method and is not effective on sparse graphs (with fewer redundant features).
[0045] In related technologies, a multi-modal user-item bipartite graph is obtained and decomposed into multiple single-modal user-item bipartite graphs. Then, a graph attention network is used to calculate the user embedding and item embedding of each single modality. Then, an MR-GNN model is constructed, which includes a GNN model corresponding to each modality. During the training process, for each modality, the final loss function is determined based on the user embedding and item embedding of that modality, and the GNN model, user embedding parameters, and item embedding parameters corresponding to each modality are updated until the MR-GNN model converges to obtain a trained MR-GNN model. Finally, target items are recommended to the target user based on the trained MR-GNN model.
[0046] However, this method only uses serial or average integration methods for single-modal embedding integration, does not fully explore the impact of other integration strategies (such as weighted integration) on the recommendation performance, and does not consider the dynamic importance changes of different modality features.
[0047] To solve the above problems, the present disclosure provides a multi-modal recommendation method, device, electronic device, storage medium, and program product solution, which specifically includes:
[0048] Determine the project description picture and the project description text, perform weighted fusion on the project description picture and the project description text to obtain an overall modal feature representation; determine the interaction history data between the user and the project, and obtain a user preference score based on the overall modal feature representation and the interaction history data; input the user preference score into a pre-trained multi-modal preference optimization recommendation model to obtain an optimized user preference score; obtain a recommendation result based on the optimized user preference score. The present disclosure integrates the project description picture and text, uses weighted fusion technology to generate an overall modal feature representation, and combines the interaction history data between the user and the project to calculate the user preference score. Subsequently, input these preference scores into a pre-trained multi-modal preference optimization recommendation model to further optimize the scores, and finally generate accurate recommendation results. Based on this, it can solve the problems of low multi-modal information fusion efficiency and insufficient user preference modeling in the prior art.
[0049] After introducing the basic principle of the present disclosure, various non-limiting embodiments of the present disclosure will be specifically introduced below.
[0050] Reference Figure 1 , which is a schematic diagram of an application scenario of the multi-modal recommendation method provided by an exemplary embodiment of the present disclosure.
[0051] In this application scenario, it includes a terminal device 101 and a server 102. Among them, the terminal device 101 and the server 102 can be connected through a wired or wireless communication network to achieve data interaction.
[0052] The terminal device 101 can be an electronic device near the user side with data transmission, multimedia input / output functions, including but not limited to a desktop computer, a mobile phone, a mobile computer, a tablet computer, a media player, a smart wearable device, a personal digital assistant (PDA), or other electronic devices capable of implementing the above functions. The electronic device can include a processor and a display screen with touch input function. The display screen is used to present a graphical user interface, and the graphical user interface can display an application interface. The processor is used to process application data, generate a graphical user interface, and control the display of the graphical user interface on the display screen.
[0053] The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0054] In some exemplary embodiments, the multimodal recommendation method may run on the terminal device 101 or the server 102.
[0055] When the multimodal recommendation method runs on the server 102, the server 102 is used to provide multimodal recommendation services to the users of the terminal device 101.
[0056] The terminal device 101 determines the project description picture and the project description text, and the terminal device 101 transmits the project description picture and the project description text to the server 102;
[0057] The server 102 receives the project description picture and the project description text transmitted from the terminal device 101, and the server 102 performs weighted fusion on the project description picture and the project description text to obtain an overall modal feature representation;
[0058] The server 102 determines the interaction history data between the user and the project, and the server 102 obtains a user preference score based on the overall modal feature representation and the interaction history data;
[0059] The server 102 inputs the user preference score into a pre-trained multimodal preference optimization recommendation model to obtain an optimized user preference score;
[0060] After the server 102 obtains a recommendation result based on the optimized user preference score, the server 102 transmits the recommendation result to the terminal device 101.
[0061] It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0062] Refer to Figure 2 , for the multimodal recommendation method, the method includes the following steps:
[0063] Step S210: Determine the project description picture and the project description text, and perform weighted fusion on the project description picture and the project description text to obtain an overall modal feature representation.
[0064] Specifically, the project description picture refers to:
[0065] Image data related to the project, and these images can be product display pictures, product usage scenario pictures, activity posters, etc. These pictures convey features such as the appearance, style, and usage method of the project through visual information.
[0066] Specifically, the project description text refers to:
[0067] Text data related to the project, which can be detailed descriptions of products, user reviews, product specifications, usage instructions, etc. These texts convey information such as the functions, features, and advantages of the project through semantic information.
[0068] In the above exemplary embodiment, the project description picture and the project description text are introduced. Next, the method for obtaining the overall modal feature representation is introduced:
[0069] In this exemplary embodiment, the weighted fusion of the project description picture and the project description text is performed to obtain the overall modal feature representation, including:
[0070] Feature extraction is performed on the project description picture and the project description text to obtain the initial visual feature and the initial text feature; the initial visual feature and the initial text feature are concatenated to obtain a joint feature vector; the joint feature vector is mapped to obtain the visual modal space and the text modal space; based on the initial visual feature and the visual modal space, a visual attention matrix is obtained; based on the initial text feature and the text modal space, a text attention matrix is obtained; the initial visual feature is weighted based on the visual attention matrix to obtain a weighted visual matrix; the initial text feature is weighted based on the text attention matrix to obtain a weighted text matrix; the initial visual feature is combined with the weighted visual matrix to obtain a visual feature representation; the initial text feature is combined with the weighted text matrix to obtain a text feature representation; the visual feature representation and the text feature representation are concatenated to obtain the overall modal feature representation.
[0071] Specifically, when implementing, the method for feature extraction of the project description picture and the project description text to obtain the initial visual feature and the initial text feature:
[0072] The image processing unit is used to extract feature information from the description pictures of the commodities. The specific processing flow is as follows: Each picture is uniformly adjusted to a resolution of 224×224; normalization operation is performed on each pixel of the image, scaled from [0, 255] to [0, 1] for easy processing; the obtained tensor is input into the ResNet-50 model, and the output of the 2048-dimensional global average pooling (GAP) layer is extracted as the initial visual feature.
[0073] The original text embedding unit is mainly used to obtain the text features of the descriptive text of each item (such as the title and product description). The main process is as follows: delete irrelevant information such as special characters and html tags; use unified lowercase letter replacement and word segmentation operations to ensure consistent input format; for text with a length exceeding 128 tokens, retain the first 128 tokens and truncate the excess; input the processed text into the Sentence-BERT model, and extract the output at the [CLS] token position as the text embedding of the item, generating a 384-dimensional original text feature vector.
[0074] When specifically implemented, the method of concatenating the initial visual feature and the initial text feature to obtain a joint feature vector is as follows:
[0075] Concatenate the processed image and text features to obtain a joint feature vector J, whose dimension is the sum of the dimensions of the image and text features:
[0076] J = [X v ; X t
[0077] Where X v and X t respectively represent the preliminary features of the image and text, and are represented in the form of a matrix.
[0078] When specifically implemented, the method of mapping the joint feature vector to obtain the visual modality space and the text modality space is as follows:
[0079] The process of mapping the joint feature vector is realized through matrix multiplication:
[0080] v proj = J · W jv , t proj = J · W jt
[0081] Respectively obtain the mapped visual feature v proj and text feature t proj .
[0082] Where, and are both learnable weight matrices, which are respectively used to map the joint feature to the visual and text modality spaces.
[0083] When specifically implemented, based on the initial visual feature and the visual modality space, obtain a visual attention matrix; based on the initial text feature and the text modality space, obtain a text attention matrix in the following way:
[0084] Calculate the correlation between the image feature and the text feature and the joint feature respectively to obtain the visual attention matrix Cv and the text attention matrix C t :
[0085]
[0086] Wherein, tanh is the hyperbolic tangent activation function, which is used to limit the result within (0,1) while maintaining the ability of non-linear change; and respectively represent the transposed matrices of the originally mapped feature matrices.
[0087] In specific implementation, the initial visual features are weighted based on the visual attention matrix to obtain a weighted visual matrix; the initial text features are weighted based on the text attention matrix to obtain a weighted text matrix in the following way:
[0088] Through the attention matrices C v and C t weight the original image features X v and the text features X t to obtain the weighted matrices H v and H t , and the specific operations are as follows:
[0089] H v = ReLU(C v ·X v ·W cv ), H t = ReLU(C t ·X t ·W ct )
[0090] Wherein, and are learnable weight matrices for linear transformation; the Relu function is an activation function for non-linear transformation.
[0091] In specific implementation, the initial visual features are combined with the weighted visual matrix to obtain a visual feature representation; the initial text features are combined with the weighted text matrix to obtain a text feature representation in the following way:
[0092] By combining the weighted image H v and the text features H t with the original features, the feature representation is further enhanced:
[0093] X' v = H v +(X v ·W hv ), X' t = H t+(X t ·W ht )
[0094] wherein, and are learnable weight matrices for combining the original features with the weighted features.
[0095] Wherein, based on the initial visual features and the visual modality space, a visual attention matrix is obtained, based on the initial text features and the text modality space, a text attention matrix is obtained, based on the visual attention matrix, the initial visual features are weighted to obtain a weighted visual matrix, based on the text attention matrix, the initial text features are weighted to obtain a weighted text matrix, combining the initial visual features with the weighted visual matrix to obtain a visual feature representation, and combining the initial text features with the weighted text matrix to obtain a text feature representation. The above steps are a complete iteration process, and this process needs to be recursively iterated multiple times.
[0096] After each iteration, the updated features X′ v and X′ t will be used as the input for the next round and become the new original features to participate in the next round of operations:
[0097] X v = X′ v , X t = X′ t
[0098] until the preset number of loops R is reached. Each iteration is continuously adjusting and optimizing the feature representation, enabling the model to better focus on the key features.
[0099] Specifically in implementation, the visual feature representation and the text feature representation are concatenated to obtain the overall modality feature expression:
[0100] After R rounds of loops, the obtained embeddings are concatenated:
[0101] X ati = [X′ v ; X′ t
[0102] This feature representation contains information from both the image and text modalities, has been optimized by the attention mechanism, and has stronger expressive power. After that, the visual features and text features of the items are no longer distinguished, and only this overall multi-modal feature is used to participate in subsequent operations.
[0103] Step S220: Determine the interaction history data between the user and the project, and obtain the user preference score based on the overall modal feature expression and the interaction history data.
[0104] Specifically, the interaction history data between the user and the project refers to:
[0105] User-project interaction data is usually stored in the form of logs or tables, recording the interaction behaviors (such as clicks, purchases, collections, etc.) between each user and different projects.
[0106] In the above exemplary embodiment, the interaction history data between the user and the project is introduced. Next, the method for obtaining the user preference score is introduced:
[0107] In this exemplary embodiment, obtaining the user preference score based on the overall modal feature expression and the interaction history data includes:
[0108] Based on the interaction history data, construct an interaction matrix; denoise the interaction matrix to obtain a denoised interaction matrix; normalize the denoised interaction matrix to obtain a normalized interaction matrix; determine the user neighbor set and the project neighbor set of the interaction history data, and obtain the final user embedding based on the normalized interaction matrix and the user neighbor set; obtain the project embedding based on the normalized interaction matrix and the project neighbor set; obtain the final project embedding based on the overall modal feature expression and the project embedding; obtain the user preference score based on the final user embedding and the final project embedding.
[0109] Specifically, the method for constructing an interaction matrix based on the interaction history data:
[0110] The data is processed by binarization to generate an interaction matrix:
[0111]
[0112] It is a sparse matrix of N u ×N i where N u is the number of users, and N i is the number of projects.
[0113] Specifically, the method for denoising the interaction matrix to obtain a denoised interaction matrix:
[0114] Perform graph aggregation on the obtained interaction matrix R, and in each graph aggregation process, randomly remove each edge with a certain probability to prevent overfitting and excessive dependence on certain high-frequency edges.
[0115] Specifically, for the connecting edge between project i and project j, the probability of removal each time is:
[0116]
[0117] where deg u and deg i represent the node degrees of user u and project i respectively.
[0118] Next, the interaction matrix R will randomly sample the edges according to the value of P ij to generate a denoised interaction matrix R ′ :
[0119]
[0120] In specific implementation, the way to normalize the denoised interaction matrix to obtain the normalized interaction matrix is as follows:
[0121] Perform normalization operations to ensure the balance of edge weights and prevent the excessive influence of nodes with higher degrees:
[0122]
[0123] where D is the degree matrix, and the diagonal elements are the degrees of the nodes; R ′ is the interaction matrix after denoising operation; is the normalized matrix.
[0124] In specific implementation, the way to determine the user neighbor set and project neighbor set of the interaction historical data, and obtain the final user embedding based on the normalized interaction matrix and the user neighbor set, and obtain the project embedding based on the normalized interaction matrix and the project neighbor set is as follows:
[0125] Use the processed interaction graph for graph aggregation to obtain the embedding expression:
[0126]
[0127] where represents the neighbor set of user u, represents the neighbor set of project i. After multiple rounds of information aggregation (the number of iterations is determined by the hyperparameter L), the final user embedding and project embedding will be used as the final user embedding, while will further participate in subsequent operations as the calculation input for the final project embedding expression. For specific details, refer to step S3.
[0128] Note that for each layer of graph aggregation, the original interaction matrix R needs to be denoised and normalized again, that is, each edge will be resampled to determine whether to remove it. This can increase the generalization ability of the model and make the propagation paths of each layer not exactly the same.
[0129] In specific implementation, based on the overall modal feature representation and the item embedding, the final item embedding is obtained; based on the final user embedding and the final item embedding, the method for obtaining the user preference score is as follows:
[0130] As a specific embodiment, the user and item feature solving module uses the original user-item interaction information and the cross-modal attention network to construct two graphs and perform graph convolutional propagation respectively to obtain the embedded user and item embedding expressions.
[0131] The first graph is the user-item interaction graph, and the current feature expressions of the user and the item are obtained through graph aggregation. Specifically,
[0132] The user updates the feature through the items with interaction records:
[0133]
[0134] In specific implementation, based on the normalized interaction matrix and the item neighbor set, the item embedding is obtained; based on the overall modal feature representation and the item embedding, the final item embedding is obtained; based on the final user embedding and the final item embedding, the method for obtaining the user preference score is as follows:
[0135] Continuing with the above exemplary embodiment, the item updates the feature through the users with interaction records:
[0136]
[0137] Among them, R is the user-item interaction matrix, which is a binary matrix - the value at the corresponding position is 1 if there is an interaction, otherwise it is 0.
[0138] Considering the utilization of the multi-modal features of the item (here considering images and texts), the multi-modal features are passed through the cross-modal attention network (CMAN) to obtain the fused modal feature representation:
[0139] X att = CMAN(X v , X t )
[0140] Among them, X v is the original visual feature matrix of the item, X tis the original text feature matrix of the project (obtained in the data processing module); CMAN (Cross-Modal Attention Network) is the multi-modal fusion network proposed in the present invention, aiming to capture the interaction relationship between visual and text features through the attention mechanism and obtain a stronger overall modal feature representation. This part is described in detail on the technical side.
[0141] Next, the project modal features obtained from the first graph convolution are further processed, the inner product between them is calculated to obtain the similarity graph of the projects, and the final project representation is obtained by performing graph aggregation on the similar projects of each project again:
[0142]
[0143] where h j is the project modal feature obtained through graph aggregation; X att,k is the fused modal feature of project k; S jk is the similarity weight between project j and project k, calculated from the fused modal feature.
[0144] Using the final user and project feature representations obtained in the previous stage, the preference score of the user for each project is calculated by the way of inner product:
[0145] f ij = u i ·h j
[0146] where f ij is the preference score of user i for project j, used to measure the size of user i's purchase intention for project j; u i is the final embedded representation of the user; h j is the final embedded representation of the project.
[0147] In the above exemplary embodiment, the method for obtaining the user preference score is introduced. Next, the method for obtaining the final project embedding is specifically introduced:
[0148] In this exemplary embodiment, based on the overall modal feature representation and the project embedding, the final project embedding is obtained, including:
[0149] Based on the overall modal feature representation, a similarity matrix is constructed; the similarity matrix is denoised to obtain the denoised similarity matrix; the denoised similarity matrix is normalized to obtain the normalized similarity matrix; the project embedding and the normalized similarity matrix are subjected to graph aggregation to obtain the final project embedding.
[0150] During specific implementation, based on the overall modal feature representation, construct a similarity matrix; the method for denoising the similarity matrix to obtain the denoised similarity matrix is as follows:
[0151] Calculate the similarity between projects through modal features and construct a similarity matrix.
[0152] Calculate through cosine similarity:
[0153]
[0154] Among them, and are respectively and 's L2 norm; S ij is the similarity matrix.
[0155] During specific implementation, the method for normalizing the denoised similarity matrix to obtain the normalized similarity matrix is as follows:
[0156] Denoise and normalize the constructed similarity matrix, and retain key information.
[0157] Specifically, sort all similarity values of each project i, and retain the top K most relevant projects; then divide the similarity values into four levels and assign different weights:
[0158]
[0159] Among them, top-K(S i ) represents the top K most relevant projects of project i.
[0160] Next, in the same operation as step S22, we normalize the denoised similarity matrix to ensure consistent data scale:
[0161]
[0162] Among them, D is a diagonal matrix,
[0163] During specific implementation, the method for graph aggregation of the project embedding and the normalized similarity matrix to obtain the final project embedding is as follows:
[0164] Perform graph aggregation on the normalized project-project similarity graph. Specifically:
[0165]
[0166] Among them, is the embedding of item i in the l-th layer, is the set of neighbors of item i.
[0167] After multi-layer information aggregation, the final expression of the item embedding is obtained, where the number of layers for graph aggregation is determined by a preset parameter N.
[0168] It should be noted that the initial value of h is obtained in the item embedding based on the normalized interaction matrix and the item neighbor set. The item embedding obtained by performing graph aggregation on S for L times is used as the initial value for this graph aggregation:
[0169]
[0170] Through such a setting, the modal information of items can be effectively integrated and utilized into the entire recommendation system to make better personalized recommendations.
[0171] As of the current step, the final user embedding is obtained and item embedding
[0172] Step S230: Input the user preference score into a pre-trained multi-modal preference optimization recommendation model to obtain an optimized user preference score.
[0173] In this exemplary embodiment, the multi-modal preference optimization recommendation model is trained by the following method:
[0174] Construct a sample set including a number of samples; wherein, the samples include: sample data and label data; the sample data includes training user preference scores; the label data includes training optimized user preference scores; input the sample data into the multi-modal preference optimization recommendation model to be trained to obtain predicted data output by the model, where the predicted data includes the user preference prediction score output by the model; determine the difference between the predicted data and the label data; based on the difference, update the parameters of the multi-modal preference optimization recommendation model to be trained through a loss function until the difference between the predicted data and the label data is minimized to obtain the pre-trained multi-modal preference optimization recommendation model.
[0175] In specific implementation, a sample set including a number of samples is constructed; wherein, the samples include: sample data and label data; the sample data includes user preference scores for training; the label data includes optimized user preference scores for training; the sample data is input into the multi-modal preference optimization recommendation model to be trained, and predicted data output by the model is obtained, wherein the predicted data includes user preference prediction scores output by the model; the difference between the predicted data and the label data is determined; based on the difference, the parameters of the multi-modal preference optimization recommendation model to be trained are updated through a loss function until the difference between the predicted data and the label data is minimized, and the pre-trained multi-modal preference optimization recommendation model is obtained in the following manner:
[0176] Construct a sample set containing a number of samples, and each sample includes sample data and label data:
[0177] Sample data: User preference scores f for training s (u i ,v j ), and these preference scores are calculated by the dot product of user embedding u i and item embedding v j .
[0178] Label data: Optimized user preference scores for training, and these optimized preference scores are obtained through a certain optimization method (such as the InfoBPR loss function based on mutual information).
[0179] Define the InfoBPR loss function based on mutual information for optimizing user preference scores:
[0180]
[0181] where E is the set of user-item interactions; denotes a set of negative samples v uniformly sampled from the item set V k ; f s (u i ,v j ) is the dot product of the final embeddings of user u i and item v j , representing the preference score of user u i for item v j .
[0182] Input the sample data into the multi-modal preference optimization recommendation model to be trained, and the model will output predicted data, where the predicted data includes user preference prediction scores output by the model
[0183] Calculate the predicted data output by the model and the label data f s(u i , v j ) The difference between. This difference is usually measured by a loss function, such as the InfoBPR loss function based on mutual information.
[0184] Based on the calculated difference, update the parameters of the multi-modal preference optimization recommendation model to be trained through the loss function. The update process usually uses the gradient descent method or other optimization algorithms, and the goal is to minimize the difference between the predicted data and the label data; until the difference between the predicted data and the label data of the model reaches the minimum, the finally obtained model is the pre-trained multi-modal preference optimization recommendation model.
[0185] Step S240, obtain a recommendation result based on the optimized user preference score.
[0186] In this exemplary embodiment, obtaining a recommendation result based on the optimized user preference score includes:
[0187] Perform a descending order sorting on the optimized user preference scores to obtain a sorted item list; select from the sorted item list based on a set threshold to obtain the recommendation result.
[0188] Specifically, when implementing, the method of performing a descending order sorting on the optimized user preference scores to obtain a sorted item list:
[0189] For each user u i , sort the optimized preference scores of all items in descending order. This means that items with high preference scores will be ranked in the front, while items with low preference scores will be ranked in the back; thus obtaining a sorted item list.
[0190] Specifically, when implementing, the method of selecting from the sorted item list based on a set threshold to obtain the recommendation result:
[0191] Select the top Top-K items from the sorted item list and recommend them to the user:
[0192]
[0193] These items are predicted to be the items that the user is most likely to be interested in, so they are ranked at the top of the recommendation list.
[0194] Based on the above exemplary embodiment, it can be seen that the present disclosure integrates the modal information of items into two graph structures through a cross-modal attention network mechanism, and through two graph aggregations, graph denoising, similarity grading and other fine processing, effectively improves the performance of the recommendation system in multi-modal recommendation tasks. Specifically:
[0195] The performance of the solution of the present invention on the multimodal recommendation task on the Baby and Clothing datasets is shown in Table 1. Among them, Recall@K represents the proportion of items actually selected by users among the first K items recommended by the system. In other words, Recall@K measures the ability of the recommendation system to cover the items actually selected by users in the first K recommendations. The higher the Recall@K, the greater the probability that the recommendation system will hit the items actually selected by users in the first K recommendations. NDCG@K is an indicator that comprehensively considers the recommendation effect. It not only evaluates whether the recommendation system successfully hits the items preferred by users, but also emphasizes whether the order of these items in the recommendation list is reasonable. The higher the NDCG@K value, the more it means that the system can not only accurately identify the items preferred by users, but also place these items in a better position in the recommendation list.
[0196] The effects of different technical solutions in the present invention are compared, as shown in Table 1:
[0197] Table 1 Performance comparison of multimodal recommendation methods
[0198]
[0199] It can be seen that the complete method proposed in this invention has achieved the best results on all evaluation data sets. In addition, the following findings are made:
[0200] Removing the cross-modal attention mechanism, that is, directly averaging the embeddings of different modalities, does not produce ideal results. Considering that different modal features have different initial dimensions and different importance on different data sets, it is not appropriate to directly assign weights to different modalities. In contrast, the cross-modal attention mechanism can dynamically adjust the weights of different modalities and has better results.
[0201] Only performing graph aggregation once, that is, not using modal information but relying solely on the interaction records between users and projects, has the worst effect; this shows that the multimodal characteristics of the project have played a considerable role, and trying to mine the information in the modal characteristics is of great significance to improving the effect.
[0202] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.
[0203] Note that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0204] Based on the same inventive concept, corresponding to any of the above-described method embodiments, the present disclosure also provides a multimodal recommendation device.
[0205] Referring to Figure 3 , the multimodal recommendation device includes:
[0206] A feature expression determination module 310, configured to determine an item description picture and an item description text, and perform weighted fusion on the item description picture and the item description text to obtain an overall modal feature expression;
[0207] A preference score determination module 320, configured to determine historical interaction data between a user and an item, and obtain a user preference score based on the overall modal feature expression and the historical interaction data;
[0208] An optimized preference score determination module 330, configured to input the user preference score into a pre-trained multimodal preference optimization recommendation model to obtain an optimized user preference score;
[0209] A recommendation result determination module 340, configured to obtain a recommendation result based on the optimized user preference score.
[0210] In this exemplary embodiment, the feature expression determination module 310 is specifically configured to:
[0211] Determine the project description picture and the project description text, extract features from the project description picture and the project description text to obtain initial visual features and initial text features; splice the initial visual features and the initial text features to obtain a joint feature vector; map the joint feature vector to obtain a visual modality space and a text modality space; based on the initial visual features and the visual modality space, obtain a visual attention matrix; based on the initial text features and the text modality space, obtain a text attention matrix; weight the initial visual features based on the visual attention matrix to obtain a weighted visual matrix; weight the initial text features based on the text attention matrix to obtain a weighted text matrix; combine the initial visual features with the weighted visual matrix to obtain a visual feature representation; combine the initial text features with the weighted text matrix to obtain a text feature representation; splice the visual feature representation and the text feature representation to obtain the overall modality feature expression.
[0212] In this exemplary embodiment, the preference score determination module 320 is specifically configured to:
[0213] Based on the interaction history data, construct an interaction matrix; denoise the interaction matrix to obtain a denoised interaction matrix; normalize the denoised interaction matrix to obtain a normalized interaction matrix; determine the user neighbor set and the project neighbor set of the interaction history data, and based on the normalized interaction matrix and the user neighbor set, obtain a final user embedding; based on the normalized interaction matrix and the project neighbor set, obtain a project embedding; based on the overall modality feature expression, construct a similarity matrix; perform denoising processing on the similarity matrix to obtain a denoised similarity matrix; perform normalization processing on the denoised similarity matrix to obtain a normalized similarity matrix; perform graph aggregation on the project embedding and the normalized similarity matrix to obtain the final project embedding; based on the final user embedding and the final project embedding, obtain the user preference score.
[0214] In this exemplary embodiment, the optimized preference score determination module 330 is specifically configured to:
[0215] Input the user preference score into a pre-trained multi-modal preference optimization recommendation model to obtain an optimized user preference score. Among them, the multi-modal preference optimization recommendation model is trained by the following method: construct a sample set including a number of samples. Among them, the sample includes: sample data and label data. The sample data includes the user preference score for training. The label data includes the optimized user preference score for training. Input the sample data into the multi-modal preference optimization recommendation model to be trained to obtain the predicted data output by the model. Among them, the predicted data includes the user preference prediction score output by the model. Determine the difference between the predicted data and the label data. Based on the difference, update the parameters of the multi-modal preference optimization recommendation model to be trained through a loss function until the difference between the predicted data and the label data is minimized, and obtain the pre-trained multi-modal preference optimization recommendation model.
[0216] In this exemplary embodiment, the recommendation result determination module 340 is specifically configured to:
[0217] Sort the optimized user preference scores in descending order to obtain a sorted item list. Select from the sorted item list based on a set threshold to obtain the recommendation result.
[0218] For the convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0219] The device in the above embodiment is used to implement the corresponding multi-modal recommendation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0220] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the multi-modal recommendation method in any of the above embodiments.
[0221] Figure 4 Figure 19 shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0222] The processor 1010 can be implemented in the form of a general - purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0223] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0224] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0225] The communication interface 1040 is used to connect to the communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or can also achieve communication through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).
[0226] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0227] It should be noted that although the above - mentioned device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above - mentioned device may also only include the components necessary to implement the solution of the embodiments of this specification and does not necessarily include all the components shown in the figure.
[0228] The electronic device of the above embodiment is used to implement the corresponding multimodal recommendation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0229] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the multimodal recommendation method described in any of the foregoing embodiments.
[0230] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0231] The above non-transitory computer-readable storage medium can be any available medium or data storage device accessible by a computer, including but not limited to magnetic memory (such as floppy disks, hard disks, magnetic tapes, magneto-optical discs (MO), etc.), optical memory (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memory (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid state drives (SSD)), etc.
[0232] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the multimodal recommendation method described in any of the above exemplary method embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0233] Based on the same inventive concept, corresponding to the multimodal recommendation method described in any of the above embodiments, the present disclosure also provides a computer program product including computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of the computer to cause the computer and / or the processor to execute the multimodal recommendation method. Corresponding to the execution subject of each step in each embodiment of the multimodal recommendation method, the processor executing the corresponding step can belong to the corresponding execution subject.
[0234] The computer program product of the above embodiments is used to cause the computer and / or the processor to execute the multi-modal recommendation method described in any one of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0235] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to as "circuit", "module", or "system" in this article. In addition, in some embodiments, the present disclosure can also be implemented in the form of a computer program product in one or more computer-readable media, which contain computer-readable program code.
[0236] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive examples) of the computer-readable storage medium can include, for example: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0237] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0238] The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0239] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0240] It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, thereby producing a machine, such that the computer program instructions executed by the computer or other programmable data processing apparatus create means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0241] These computer program instructions may also be stored in a computer-readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instruction means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0242] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide a process for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0243] In addition, although the operations of the method of this disclosure are depicted in the figures in a particular order, this is not required or implied to perform these operations in that particular order, or to perform all of the illustrated operations to achieve the desired result. On the contrary, the steps depicted in the flowchart may be reordered. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step and performed, and / or one step may be decomposed into multiple steps and performed.
[0244] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0245] It should be noted that although several modules or units of devices for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0246] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; under the concept of the present application, the technical features between the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.
[0247] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the devices may be shown in block diagram form in order not to make the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be completely within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0248] Although the present application has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0249] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present application.
[0250] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of various aspects does not mean that the features in these aspects cannot be combined for benefits. Such division is only for convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures and functions.
Claims
1. A multimodal recommendation method, characterized in that: include: Determine a project description picture and a project description text, and perform weighted fusion on the project description picture and the project description text to obtain an overall modal feature expression; Determine historical interaction data between the user and the item, and obtain a user preference score based on the overall modal feature expression and the historical interaction data; Inputting the user preference score into a pre-trained multimodal preference optimization recommendation model to obtain an optimized user preference score; Based on the optimized user preference scores, a recommendation result is obtained.
2. The method according to claim 1, characterized in that The weighted fusion of the project description image and the project description text to obtain the overall modal feature expression includes: Extracting features from the project description image and the project description text to obtain initial visual features and initial text features; Concatenating the initial visual features and the initial text features to obtain a joint feature vector; Mapping the joint feature vector to obtain a visual modal space and a textual modal space; Based on the initial visual features and the visual modal space, obtaining a visual attention matrix; Based on the initial text features and the text modal space, a text attention matrix is obtained; weighting the initial visual features based on the visual attention matrix to obtain a weighted visual matrix; Weighting the initial text features based on the text attention matrix to obtain a weighted text matrix; Combining the initial visual feature with the weighted visual matrix to obtain a visual feature representation; Combining the initial text features with the weighted text matrix to obtain a text feature representation; The visual feature representation and the text feature representation are concatenated to obtain the overall modal feature expression.
3. The method according to claim 1, characterized in that: The obtaining of a user preference score based on the overall modal feature expression and the interaction history data includes: Based on the interaction history data, construct an interaction matrix; denoise the interaction matrix to obtain a denoised interaction matrix; Normalizing the denoised interaction matrix to obtain a normalized interaction matrix; Determine a user neighbor set and an item neighbor set of the interaction history data, and obtain a final user embedding based on the normalized interaction matrix and the user neighbor set; Obtaining an item embedding based on the normalized interaction matrix and the item neighbor set; Obtaining a final item embedding based on the overall modality feature expression and the item embedding; The user preference score is obtained based on the final user embedding and the final item embedding.
4. The method according to claim 3, characterized in that: The obtaining of a final item embedding based on the overall modality feature expression and the item embedding comprises: Based on the overall modal feature expression, a similarity matrix is constructed; and the similarity matrix is denoised to obtain a denoised similarity matrix; Normalizing the denoised similarity matrix to obtain a normalized similarity matrix; Performing graph aggregation on the item embedding and the normalized similarity matrix to obtain the final item embedding.
5. The method according to claim 1, characterized in that The method further comprises training the multimodal preference optimization recommendation model by: Constructing a sample set including several samples; wherein the samples include: sample data and label data; the sample data includes training user preference scores; the label data includes training optimized user preference scores; Inputting the sample data into the multimodal preference optimization recommendation model to be trained to obtain prediction data output by the model, wherein the prediction data includes the user preference prediction score output by the model; Determining a difference between the predicted data and the label data; Based on the difference, the parameters of the multimodal preference optimization recommendation model to be trained are updated through a loss function until the difference between the predicted data and the label data is minimized, thereby obtaining the pre-trained multimodal preference optimization recommendation model.
6. The method according to claim 1, characterized in that The obtaining of recommendation results based on the optimized user preference scores includes: Sorting the optimized user preference scores in descending order to obtain a sorted project list; The ranked item list is selected based on a set threshold to obtain the recommendation result.
7. A multimodal recommendation device, characterized in that: include: A feature expression determination module is configured to determine a project description image and a project description text, and perform weighted fusion on the project description image and the project description text to obtain an overall modal feature expression; a preference score determination module, configured to determine the interaction history data between the user and the item, and obtain the user preference score based on the overall modal feature expression and the interaction history data; an optimized preference score determination module, configured to input the user preference score into a pre-trained multimodal preference optimization recommendation model to obtain an optimized user preference score; The recommendation result determination module is configured to obtain a recommendation result based on the optimized user preference score.
8. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 6.
Citation Information
Cited By
Recommendation information generation method and device, equipment, medium and program product
CN121365168A