Multi-modal collaborative recommendation method, system and device, medium and program product
Feature expression is enhanced through large language models, combined with multi-view interactive modeling and modal purification mechanism, and integrated the differential perception attention mechanism and dynamic modal preference gating mechanism, solving the problems of semantic information utilization, noise filtering and modal preference adaptability in multi-modal recommendations, achieving efficient and personalized multi-modal recommendation effects.
Patent Information
- Application Number
- CN202510541169.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing multimodal recommendation methods are difficult to effectively utilize the rich semantic information in multimodal data, cannot effectively filter noise, and are difficult to adapt to the user's modal preference differences in different scenarios.
Feature enhancement of user portraits and product attributes through large language models, combining multi-view interactive modeling and similarity-based modal purification mechanisms, integrating the difference-perceptual attention mechanism and dynamic modal preference gating mechanism to achieve fine-grained fusion and personalized recommendations for multi-modal features.
It significantly improves feature expression capabilities, effectively filters noise in multimodal data, can adaptively adjust the importance of each modal according to different scenarios, and significantly improves the performance of the recommendation system.
Smart Images

Figure CN120070010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal recommendation, and in particular to a multimodal collaborative recommendation method, system, device, medium and program product. Background Art
[0002] With the continuous development of e-commerce platforms, recommendation systems play an increasingly important role in helping users discover relevant content. Traditional collaborative filtering methods mainly rely on historical user-item interaction data to generate recommendation results. However, it is often difficult to accurately grasp the deep preferences of users based solely on interaction data. In recent years, the widespread availability of multimodal data has prompted researchers to explore integrating multiple modal information such as images and text into recommendation frameworks to improve accuracy.
[0003] Existing multimodal recommendation methods still face the following limitations: First, relying solely on limited raw features will hinder the full mining and utilization of rich semantic information in multimodal data.
[0004] Secondly, failure to effectively filter noise in multimodal data may lead to degraded recommendation performance. For example, background and brightness interference in visual data and redundant information in text data will significantly affect the quality of feature learning.
[0005] Thirdly, simple modal feature fusion strategies are difficult to solve the problem of different users' preferences for different modalities. Studies have shown that users' modal preferences vary significantly in different scenarios. For example, when browsing clothing products, users may rely more on visual features, while when buying books, they pay more attention to text descriptions. Although some existing methods have adopted an adaptive fusion mechanism based on behavior perception, these feature fusion methods still use a unified strategy and cannot adapt to such dynamic preference characteristics.
[0006] Therefore, how to effectively use large language models to enhance feature expression, how to deal with complex multi-source modal noise, and how to achieve personalized feature fusion are key issues that need to be urgently addressed in the field of multimodal recommendation. Summary of the invention
[0007] In order to solve the above problems, the present invention proposes a multimodal collaborative recommendation method, system, device, medium and program product, which enhance the features of user portraits and product attributes and effectively improve the feature expression ability; through multi-view interactive modeling and similarity-based modal purification mechanism, the noise problem in multimodal data is effectively solved, and comprehensive modeling of user preferences is achieved at the same time; by integrating the difference-aware attention mechanism and the dynamic modal preference gating mechanism, the fine-grained fusion of multimodal features is achieved, and the importance of each modality of different products can be adaptively adjusted according to different scenarios.
[0008] In order to achieve the above object, the present invention adopts the following technical solution: In a first aspect, the present invention provides a multimodal collaborative recommendation method, comprising: Generate user portrait representation based on user historical behavior data, and generate text modal features based on product original attributes; Calculate the similarity between the visual modal features and textual modal features of the product and the reference vector of the corresponding modality, obtain the purification weight according to the similarity, weight each modal feature based on the purification weight to obtain the multimodal purified feature, and obtain the fused modal purified feature after secondary purification of the multimodal purified feature according to the cross-modal similarity; After integrating the user portrait representation with the text modal features, a modal similarity graph is constructed based on the purified features of different products in each modality to determine the multimodal features. A user-product bipartite graph is constructed based on the user's historical behavior data, and the user-product interaction features are obtained after multi-layer graph convolution and cross-layer aggregation. The gating vector is calculated according to the user-product interaction characteristics, and the attention weight is calculated according to the difference between the multimodal features and the fused modal purification features. The multimodal features and the fused modal purification features are weighted, and the weighted modal features are aggregated based on the gating vector to obtain the final fused features to be tested, so as to obtain the product recommendation results for the user.
[0009] As an optional implementation, the process of generating a user portrait representation based on the user's historical behavior data includes: Convert user u's historical interaction records into structured text : ;in, Indicates the product ID. Indicates the interaction timestamp; Use large language models to generate user profile features for structured text : ; The user portrait features are linearly projected to obtain the user portrait representation : ;in, is the learnable transformation matrix, is the bias vector; The process of generating text modal features based on the original attributes of the product includes: Using large language models, based on commodity The original properties , generate goods New properties of : ; By combining products The original attributes and products The text modal features with enhanced new attributes : .
[0010] As an alternative implementation, the process of obtaining the multi-modal purification features and the fusion-modal purification features includes: Mapping the modal features to a unified space by linear transformation to obtain the linearly transformed modal features : ; is the weight matrix; is the commodity before linear transformation of the modal features; is the bias vector; Calculating the cosine similarity between the linearly transformed modal features and the corresponding modal reference vectors : ; Obtaining the purification weight according to the cosine similarity : ; Multiplying the linearly transformed modal features and the purification weight to obtain the multi-modal purification features : , , respectively represent the visual modality and the text modality; Calculating the cross-modal similarity , to perform secondary purification on the multi-modal purification features to obtain the fusion-modal purification features ; ; ; ; where, is the purified visual feature; is the purified text feature; is the sigmoid activation function; is the purified cross-modal weight.
[0011] As an alternative implementation, the integration process of the user portrait representation and the text modal features is: ; In the formula: is the user embedding; is the learnable fusion weight; is the user portrait representation; is the interaction feature aggregated from the text modal feature embedding; Modal similarity graph is ; In the formula: and For a commodity and the commodity in the modality Purification features; , representing the visual, text, and fusion modalities respectively; Multimodal features are ; where: represents the purification feature; is the modality similarity graph of the corresponding modality; , represent the visual modality, text modality, and fusion modality respectively.
[0012] As an alternative implementation, an adjacency matrix A is created based on the number of users and the number of commodities. According to the user historical behavior data, each interaction record of user - commodity in the adjacency matrix A is set to 1 at the corresponding position, thus obtaining a user - commodity bipartite graph. Based on this, the normalized adjacency matrix of the user - commodity bipartite graph , , is the degree matrix of the user - commodity bipartite graph; The process of obtaining the interaction features of user - commodity through multi - layer graph convolution cross - layer aggregation is as follows: ; ; where: is the number of layers of the graph convolutional network, is the total number of layers of the graph convolutional network, is the layer node feature vector; is layer node feature vector.
[0013] As an alternative implementation, the process of weighting the multimodal features and the fusion modality purification features is as follows: The visual modality features and text modality features in the multimodal features are concatenated with the fusion modality purification features and then given attention weights to obtain weighted visual features and weighted text features; According to the interaction features of user - commodity The calculated gating vector is : ; where: , and are the multi - layer perceptrons corresponding to the visual modality, text modality, and fusion modality respectively; is the feature concatenation operation; is the sigmoid activation function; The gating vector Decompose it into gating vectors corresponding to the visual modality, text modality, and fusion modality. After weighting the weighted visual features, weighted text features, and fusion modality purification features based on their respective corresponding gating vectors and then averaging them for normalization, aggregation is thus completed. Concatenate the weighted and aggregated fusion features with the user-item interaction features to obtain the final fusion features to be measured. Separate the fusion features to be measured into user embeddings and item embeddings . Calculate the inner product of the user embedding and the item embedding as the user's interest score for the item, and thus obtain the item recommendation result for the user.
[0014] In a second aspect, the present invention provides a multi-modal collaborative recommendation system, including: A feature enhancement module, configured to generate a user portrait representation according to user historical behavior data and generate text modality features according to the original item attributes; A modality purification module, configured to calculate the similarity between the visual modality features and text modality features of an item and the reference vectors of the corresponding modalities, obtain purification weights according to the similarity, weight each modality feature based on the purification weights to obtain multi-modal purification features, and perform secondary purification on the multi-modal purification features according to the cross-modal similarity to obtain fusion modality purification features; A multi-view interaction modeling module, configured to integrate the user portrait representation and the text modality features, construct a modality similarity graph based on the purification features of different items in each modality, and thus determine multi-modal features; construct a user-item bipartite graph according to user historical behavior data, and obtain user-item interaction features after multi-layer graph convolution cross-layer aggregation based on this; A feature aggregation module, configured to calculate gating vectors according to the user-item interaction features, calculate attention weights according to the differences between the multi-modal features and the fusion modality purification features, weight the multi-modal features and the fusion modality purification features accordingly, aggregate the obtained weighted modality features based on the gating vectors to obtain the final fusion features to be measured, and thus obtain the item recommendation result for the user.
[0015] In a third aspect, the present invention provides an electronic device, including a memory and a processor, as well as computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.
[0016] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.
[0017] In a fifth aspect, the present invention provides a computer program product, including a computer program which, when executed by a processor, implements the method described in the first aspect.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a multi-modal collaborative recommendation method, system, device, medium and program product. By using a large language model, it enhances the features of user portraits and commodity attributes, obtains more comprehensive and accurate text modal features by generating richer semantic descriptions, and effectively improves the feature expression ability.
[0019] The present invention proposes a multi-modal collaborative recommendation method, system, device, medium and program product. Through multi-view interaction modeling and a similarity-based modal purification mechanism, it effectively solves the noise problem in multi-modal data and simultaneously realizes a comprehensive modeling of user preferences. Among them, based on multi-view interaction modeling, a user-item bipartite graph and an item-item similarity graph based on modal purification features are constructed to capture the complex preference patterns of users from two perspectives: user-item interaction and item-item semantic relationship. Through a similarity-aware modal purification mechanism, noise reduction processing is performed on modal features, and by introducing a learnable modal reference vector, the similarity between modal features and the modal reference vector is calculated to evaluate the importance of different modalities.
[0020] The present invention proposes a multi-modal collaborative recommendation method, system, device, medium and program product, which proposes feature fusion based on multi-modal adaptation. By integrating a difference-aware attention mechanism and a dynamic modal preference gating mechanism, fine-grained fusion of multi-modal features is achieved, and the importance of each modality of different commodities can be adaptively adjusted according to different scenarios. Among them, the difference-aware attention mechanism evaluates the unique contribution of each modality by calculating the difference between modal features and fusion features; the dynamic modal preference gating mechanism adjusts the importance of each modality in different scenarios based on collaborative feature learning to achieve personalized feature fusion, significantly improving the performance of the recommendation system.
[0021] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0023] Figure 1It is the principle flow chart of the multi-modal collaborative recommendation method provided in Embodiment 1 of the present invention; Figure 2 It is the schematic diagram of the structural principle of the multi-modal collaborative recommendation system provided in Embodiment 2 of the present invention. Specific embodiments
[0024] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0025] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0026] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "comprise" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0028] Embodiment 1 This embodiment relies on the Key Laboratory of Artificial Intelligence Application for People's Livelihood Services in Shandong Province and the Future Industry Laboratory of Shandong Province for General Artificial Intelligence, and provides a multi-modal collaborative recommendation method enhanced based on large language models (LLMs); As Figure 1 shown, it includes: Generate a user portrait representation according to the user's historical behavior data, and generate text modality features according to the original attributes of the commodity; Calculate the similarity between the visual modality features and text modality features of the commodity and the reference vectors of the corresponding modalities, obtain the purification weights according to the similarity, weight each modality feature based on the purification weights to obtain multi-modal purification features, and perform secondary purification on the multi-modal purification features according to the cross-modal similarity to obtain fusion modality purification features; After integrating the user portrait representation with the text modal features, a modal similarity graph is constructed based on the purified features of different products in each modality to determine the multimodal features. A user-product bipartite graph is constructed based on the user's historical behavior data, and the user-product interaction features are obtained after multi-layer graph convolution and cross-layer aggregation. The gating vector is calculated according to the user-product interaction characteristics, and the attention weight is calculated according to the difference between the multimodal features and the fused modal purification features. The multimodal features and the fused modal purification features are weighted, and the weighted modal features are aggregated based on the gating vector to obtain the final fused features to be tested, so as to obtain the product recommendation results for the user.
[0029] The multimodal collaborative recommendation method of this embodiment is described in detail below.
[0030] S1: Use large language models to enhance the features of user portraits and product attributes, and obtain more comprehensive and accurate text modality representation by generating richer semantic descriptions.
[0031] Specifically include: S1-1: Obtain the historical behavior data of user u. The historical behavior data is the user u's Moments and Products Through user behavior sequence conversion, the historical interaction records of user u are converted into structured text : ; in, Indicates the product ID. Indicates the interaction timestamp.
[0032] S1-2: Use a large language model to enhance the structured text and generate enhanced user profile features through the semantic understanding ability of the large language model: ; Where: Represents user portrait features.
[0033] S1-3: Use linear projection to obtain the final user portrait representation : ; Where: is the learnable transformation matrix, is the bias vector.
[0034] S1-4: For product attribute enhancement, use a large language model based on product The original attributes of the generated product New attributes of goods The new attributes are supplementary descriptive texts and category information to obtain a richer and more accurate text modality representation; product attributes such as product descriptions, titles, reviews, etc.
[0035] Among them, the product The new attributes are ; where: Represents the original attributes of the product ; By combining the original attributes of the product and the generated new attributes of the product Enhanced text modality features are obtained : ; among them, Represents the text modality; Represents the text modality features of the product ;
[0036] It can be understood that after obtaining the historical behavior data of the user and the original attributes of the product, preprocessing operations such as data cleaning, missing value filling, and feature standardization can be performed on them.
[0037] S2: To effectively process the noise information in multi-modal data, this embodiment proposes a modality purification mechanism based on similarity perception. By introducing learnable modality reference vectors, the similarity between each modality feature and the modality reference vector of the corresponding modality is calculated to achieve the evaluation of the importance of each modality feature and noise filtering.
[0038] Specifically, it includes: S2-1: Introduce modality reference vectors , where , Represent the visual modality and the text modality respectively.
[0039] S2-2: To achieve an effective comparison between each modality feature and the corresponding modality reference vector, the modality feature is mapped to a unified space through a linear transformation to obtain the linearly transformed modality feature : ; Among them, Is the weight matrix; Is the modality feature of the product Before linear transformation; Is the bias vector.
[0040] S2-3: Calculate the cosine similarity between the linearly transformed modality feature and the corresponding modality reference vector : ;
[0041] S2-4: Perform sigmoid transformation on the cosine similarity to obtain the purification weight : ; wherein, is the sigmoid activation function
[0042] S2-5: Weight the modal features through the soft purification mechanism, that is, multiply the linearly transformed modal features and the purification weight element-wise to perform feature purification and obtain the purified features : .
[0043] S2-6: To further improve the quality of the purified features, this embodiment also utilizes the complementarity between different modalities, calculates the cross-modal similarity , and performs secondary purification on the purified features to obtain the fused modal purified features ; ; ; ; wherein, is the purified visual feature; is the purified text feature; is the sigmoid activation function; is the purified cross-modal weight
[0044] In this embodiment, the process of secondary purification based on cross-modal similarity has two important functions: (1) helping to identify features that are important in different modalities and providing an additional noise filtering layer; (2) promoting the fusion of complementary information between different modalities to form a more comprehensive and robust feature representation
[0045] S3: Based on multi-view interaction modeling. To comprehensively capture user preferences, not only direct user-item interactions but also semantic relationships between items need to be considered. Specifically: construct a user-item bipartite graph and an item-item similarity graph based on the purified modal features; model user preferences from two perspectives of user-item interactions and item-item semantic relationships; from the user-item perspective, construct a user-item bipartite graph and capture high-order connection patterns through a graph convolutional network; from the item-item perspective, construct an item-item similarity graph based on the purified features of the modality, that is, based on the purified visual features , purified text features and fused modal purified features Construct the commodity-commodity similarity graph for each corresponding modality respectively, and obtain the semantic association between commodities through feature propagation. For the text modality, additionally integrate the user portrait features enhanced by the large language model.
[0046] Specifically, it includes: S3-1: In the commodity-commodity view, for each modality , based on the purified features (i.e., purified visual features , purified text features and fused modality purified features ), construct the modality similarity graph : ; Where: and are the purified features of commodity and commodity in modality ; , representing the visual, text, and fused modalities respectively.
[0047] Meanwhile, to reduce noise and computational cost, in the modality similarity graph, only the top edges with the maximum similarity are retained for each commodity.
[0048] For each modality (representing the visual, text, and fused modalities respectively), through the corresponding modality similarity graph for feature propagation, obtain the enhanced commodity multi-modal features of the purified features after graph convolution on the modality similarity graph , : : ; Where: represents the purified features, including purified visual features , purified text features and fused modality purified features ; is the modality similarity graph of the corresponding modality; Among them, in the enhanced commodity multi-modal features , different from the visual modality and the fused modality, the text modality features also need to integrate the user portrait representation enhanced by the large language model, that is, calculate the user embedding of the text modality features through adaptive fusion: ; Where: is a learnable fusion weight; is the user profile representation enhanced by the large language model; represents the interaction features aggregated from the text modality feature embeddings. The text modality features of the product such as title, attributes, reviews, etc.) are converted into high-dimensional vector representations (i.e., embeddings), and these vectors are aggregated into comprehensive interaction features to capture the multi-dimensional semantic information of the product.
[0049] S3-2: In the user-product view, create an adjacency matrix A of size , where is the number of users, is the number of products; According to the user's historical behavior data, traverse each interaction record of the user-product , and set the corresponding position in the adjacency matrix A to 1, indicating that there is an interaction between user and product . Thus, a user-product bipartite graph is obtained; and construct the normalized adjacency matrix of the user-product bipartite graph: ; In the formula: is the adjacency matrix of the user-product bipartite graph, is the degree matrix of the user-product bipartite graph.
[0050] The vector representation obtained by concatenating or combining the features of the user and the product is the initial node feature . This combination method allows the model to consider the features of both the user and the product simultaneously during the graph convolution process, thereby more comprehensively capturing the interaction relationship between the user and the product. The initial node feature is: ; where is the initial feature vector of the user; is the initial feature vector of the product.
[0051] To capture the high-order connection patterns, message passing is performed through multi-layer graph convolution: ; In the formula: represents the number of layers of the graph convolutional network, is the node feature vector of the th layer; is the node feature vector of the th layer.
[0052] Through cross-layer aggregation, information of different interaction orders is retained, and finally the user-item interaction features are obtained. : ; where, is the total number of layers of the graph convolutional network.
[0053] S4: Feature fusion based on multi-modal adaptation. By integrating the difference-aware attention mechanism and the dynamic modal preference gating mechanism, fine-grained fusion of multi-modal features is achieved, and the importance of each modality of different items can be adaptively adjusted according to different scenarios. Among them, the difference-aware attention mechanism calculates the purified visual features and the purified text features to evaluate the unique contributions of each modality by the difference from the fused modality purified features ; the dynamic modal preference gating mechanism adjusts the importance of each modality in different scenarios based on collaborative feature learning to achieve personalized feature fusion.
[0054] Specifically, it includes: S4-1: To solve the problem that traditional feature fusion methods are difficult to capture the complex complementary relationships between different modalities, this embodiment proposes a difference-aware attention mechanism: ; In the formula: represents the initial fusion feature, that is, the fused modality purified feature ; is the enhanced multi-modal feature, ; is the difference vector.
[0055] The difference vector is converted into an attention score through a modality-specific multi-layer perceptron (MLP, Multilayer Perceptron): ; .
[0056] where, is the mapping result of the difference vector through a modality-specific multi-layer perceptron (MLP), and is the intermediate value for generating the attention score .
[0057] The additive attention mechanism is adopted to obtain the weighted modality features: .
[0058] This design is for sharing information (through ) and modality-specific information (through ) have all created direct paths.
[0059] S4-2: To capture users' modal preferences, a dynamic modal preference gating mechanism based on collaborative embedding is introduced to calculate the gating vector : ; In the formula: represents the collaborative embedding feature of user-item interaction obtained in step S3-2; , and are the multi-layer perceptrons corresponding to the visual modality, text modality, and fusion modality respectively; represents the feature concatenation operation; is the sigmoid activation function.
[0060] By learning the collaborative signal, the gating mechanism can identify and adapt to users' preference patterns in different scenarios.
[0061] ; In the formula: The operation decomposes the unified gating vector into modality-specific gating vectors, , and correspond to the importance weights of the visual modality, text modality, and fusion modality respectively, that is, the modality-specific gating vectors obtained by decomposition, to achieve fine-grained control of the contribution degree of each modality.
[0062] The final fusion feature is obtained through weighted aggregation: ; In the formula: and represent the weighted visual feature and weighted text feature processed by the attention mechanism respectively; is the initial fusion feature; divided by 3 for normalization to ensure the stability of training and maintain the proportional relationship of relative importance.
[0063] By introducing residual connections to maintain collaborative and content-based signals, the fusion representation is obtained: .
[0064] Finally, the fusion representation is separated into the user embedding and the item embedding : .
[0065] Among them, is the number of users; is the number of items; In this embodiment, the ultimate goal is to perform product recommendations. Then, by calculating the inner product of the user embedding and the product embedding as the interest score of the user in the product. This method assumes that the higher the similarity between the user embedding and the product embedding, the greater the user's interest in the product. Finally, several products with the highest interest scores are selected and recommended to the user.
[0066] In this embodiment, a dataset from a real e-commerce scenario is used for verification. Specifically, two categories of product review datasets on a certain platform are used: baby products (Baby) and sports and outdoors (Sports). The basic statistical information is shown in Table 1.
[0067] Table 1 Dataset statistical information; .
[0068] The model evaluation uses common evaluation metrics for recommendation systems, including R@20 (Recall@20, the recall rate of the top 20 retrieval results) and N@20 (NDCG@20, which refers to evaluating the quality of the top 20 retrieval results when using the NDCG (Normalized Discounted Cumulative Gain) metric).
[0069] Table 2 presents the performance comparison results between the method of this embodiment and existing methods such as BPR (Bayesian Personalized Ranking), LightGCN (Light Graph Convolution Network), LATTICE (Lattice Long Short-Term Memory), BM3 (Bootstrap Latent Representations for Multi-modal Recommendation), FREEDOM (Freezing and Denoising Graph Structures for Multimodal Recommendation), VBPR (Visual Bayesian Personalized Ranking), MMGCN (Multi-modal Graph Convolutional Network), and MGCN (Multi-View Graph Convolutional Network).
[0070] Comparison of experimental results in Table 2; 。
[0071] Based on the results in Table 2, it can be seen that the multi-modal collaborative recommendation method proposed in this embodiment performs best in all evaluation metrics. Compared with existing methods, this embodiment obtains richer semantic representations through large language model feature enhancement, effectively filters noise information through similarity-aware modal purification, and achieves more accurate feature fusion through multi-modal adaptive feature aggregation, thus significantly improving the performance of the recommendation system.
[0072] It should be noted that the acquisition of all data is based on compliance with laws, regulations, and user consent, and the data is legally applied.
[0073] Embodiment 2 As Figure 2 shown, this embodiment provides a multi-modal collaborative recommendation system, including: A feature enhancement module, configured to generate a user portrait representation according to user historical behavior data and generate text modal features according to the original attributes of commodities; The modality purification module is configured to calculate the similarity between the visual modality features and text modality features of the product and the reference vector of the corresponding modality, obtain the purification weight according to the similarity, weight each modality feature based on the purification weight to obtain the multimodal purification feature, and obtain the fusion modality purification feature after secondary purification of the multimodal purification feature according to the cross-modality similarity; The multi-view interaction modeling module is configured to integrate the user portrait representation with the text modal features, and then construct a modal similarity graph based on the purified features of different products in each modality to determine the multi-modal features; construct a user-product bipartite graph based on the user's historical behavior data, and obtain the user-product interaction features after multi-layer graph convolution and cross-layer aggregation; The feature aggregation module is configured to calculate the gating vector based on the user-product interaction features, calculate the attention weight based on the difference between the multimodal features and the fused modal purification features, weight the multimodal features and the fused modal purification features, aggregate the obtained weighted modal features based on the gating vector, and obtain the final fused features to be tested, so as to obtain the product recommendation results for the user.
[0074] In this embodiment, a preprocessing module is also included, which is configured to obtain the user's historical behavior data and the original attributes of the goods, and then perform preprocessing operations such as data cleaning, missing value completion and feature standardization.
[0075] In this embodiment, after the feature aggregation module, the purpose is to predict the user's interest score for the product in order to make product recommendations. Then, the user embedding and product embedding are obtained according to the final fusion feature to be tested, and the inner product of the user embedding and the product embedding is calculated as the user's interest score for the product. Finally, several products with the highest interest scores are selected and recommended to the user.
[0076] It should be noted that the above modules correspond to the steps described in Example 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.
[0077] In further embodiments, there is also provided: An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in Embodiment 1 is performed. For the sake of brevity, it will not be described in detail here.
[0078] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0079] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0080] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.
[0081] The method in Embodiment 1 can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0082] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented.
[0083] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to execute the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules may be combined or divided as needed. The machine-executable instructions for program modules may be executed locally or within a distributed device. In a distributed device, program modules may be located in local and remote storage media.
[0084] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program code is executed by the computer or other programmable data processing devices, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.
[0085] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that the device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.
[0086] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0087] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A multimodal collaborative recommendation method, characterized in that: include: Generate user portrait representation based on user historical behavior data, and generate text modal features based on product original attributes; Calculate the similarity between the visual modal features and textual modal features of the product and the reference vector of the corresponding modality, obtain the purification weight according to the similarity, weight each modal feature based on the purification weight to obtain the multimodal purified feature, and obtain the fused modal purified feature after secondary purification of the multimodal purified feature according to the cross-modal similarity; After integrating the user portrait representation with the text modal features, a modal similarity graph is constructed based on the purified features of different products in each modality to determine the multimodal features. A user-product bipartite graph is constructed based on the user's historical behavior data, and the user-product interaction features are obtained after multi-layer graph convolution and cross-layer aggregation. The gating vector is calculated according to the user-product interaction characteristics, and the attention weight is calculated according to the difference between the multimodal features and the fused modal purification features. The multimodal features and the fused modal purification features are weighted, and the weighted modal features are aggregated based on the gating vector to obtain the final fused features to be tested, so as to obtain the product recommendation results for the user.
2. A multimodal collaborative recommendation method as claimed in claim 1, characterized in that: The process of generating a user portrait representation based on user historical behavior data includes: Convert user u's historical interaction records into structured text : ;in, Indicates the product ID. Indicates the interaction timestamp; Use large language models to generate user profile features for structured text : ; The user portrait features are linearly projected to obtain the user portrait representation : ;in, is the learnable transformation matrix, is the bias vector; The process of generating text modal features based on the original attributes of the product includes: Using large language models, based on commodity The original properties , generate goods New properties of : ; By combining products The original attributes and products New properties of the enhanced text modal feature : .
3. A multimodal collaborative recommendation method as claimed in claim 1, characterized in that: The process of obtaining multi-modal purification features and fusion modal purification features includes: The modal features are mapped to the unified space using linear transformation to obtain the modal features after linear transformation : ; is the weight matrix; is the product before linear transformation The modal characteristics of is the bias vector; Calculate the modal characteristics after linear transformation and the corresponding modal reference vector Cosine similarity between : ; Get the purification weight according to the cosine similarity : ; The modal characteristics after linear transformation and purification weights Multiply to get multimodal purification features : , , Representing visual modality and textual modality respectively; Calculating cross-modal similarity , to purify the multimodal features Secondary purification to obtain fusion modal purification features ; ; ; ;in, To purify visual features; To purify text features; is the sigmoid activation function; To purify cross-modal weights.
4. The multimodal collaborative recommendation method according to claim 1, characterized in that: The integration process of user portrait representation and text modality features is as follows: ; Where: Embed for users; is the learnable fusion weight; Represents the user portrait; is the interactive feature obtained by aggregation from text modality feature embedding; Modal Similarity Graph for ; Where: and For goods and products In modal Purification features in Multimodal features for ; Where: Indicates purification characteristics; is the modal similarity graph of the corresponding modality; , They represent visual modality, textual modality and fusion modality respectively.
5. The multimodal collaborative recommendation method according to claim 1, characterized in that: Create an adjacency matrix A based on the number of users and products. According to the user's historical behavior data, set the corresponding position of each user-product interaction record in the adjacency matrix A to 1, thereby obtaining a user-product bipartite graph. Based on this, determine the standardized adjacency matrix of the user-product bipartite graph , , is the degree matrix of the user-item bipartite graph; The user-product interaction features are obtained through cross-layer aggregation of multi-layer graph convolutions. The process is: ; ; Where: is the number of layers of the graph convolutional network, is the total number of layers of the graph convolutional network, For the The node feature vector of the layer; for The node feature vector of the layer.
6. A multimodal collaborative recommendation method as claimed in claim 1, characterized in that: The process of weighting the multimodal features and the fused modal purification features is as follows: the visual modal features and text modal features in the multimodal features are concatenated with the fused modal purification features and then the attention weights are assigned to obtain the weighted visual features and weighted text features; Based on the user-product interaction characteristics The calculated gating vector is : ; Where: , and They are the multi-layer perceptrons corresponding to the visual modality, text modality and fusion modality respectively; It is a feature concatenation operation; is the sigmoid activation function; The gate vector Decompose into gate vectors corresponding to visual modality, textual modality and fusion modality, weight the weighted visual features, weighted textual features and fusion modality purification features based on their corresponding gate vectors and then average them for normalization, thus completing the aggregation, concatenating the weighted aggregated fusion features with the user-product interaction features to obtain the final fusion features to be tested, and separate the fusion features to be tested into user embedding and product embedding , calculate user embedding and product embedding The inner product of is used as the user's interest score for the product, thereby obtaining the product recommendation result for the user.
7. A multimodal collaborative recommendation system, characterized in that: include: A feature enhancement module is configured to generate a user portrait representation based on the user's historical behavior data and to generate a text modality feature based on the original attributes of the product; The modality purification module is configured to calculate the similarity between the visual modality features and text modality features of the product and the reference vector of the corresponding modality, obtain the purification weight according to the similarity, weight each modality feature based on the purification weight to obtain the multimodal purification feature, and obtain the fusion modality purification feature after secondary purification of the multimodal purification feature according to the cross-modality similarity; The multi-view interaction modeling module is configured to integrate the user portrait representation with the text modal features, and then construct a modal similarity graph based on the purified features of different products in each modality to determine the multi-modal features; construct a user-product bipartite graph based on the user's historical behavior data, and obtain the user-product interaction features after multi-layer graph convolution and cross-layer aggregation; The feature aggregation module is configured to calculate the gating vector based on the user-product interaction features, calculate the attention weight based on the difference between the multimodal features and the fused modal purification features, weight the multimodal features and the fused modal purification features, aggregate the obtained weighted modal features based on the gating vector, and obtain the final fused features to be tested, so as to obtain the product recommendation results for the user.
8. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 6 is completed.
9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 6.
10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Alzheimer's disease language detection classification system and method based on a feature purification network
CN113961700A
E-commerce platform commodity recommendation method and system based on user preference analysis
CN119398864A
Electric power material recommendation method based on intention perception knowledge graph
CN119577242A
Graph model-based short video recommendation method, intelligent terminal and storage medium
WO2021179640A1
Cited By
Recommendation method and system based on hierarchical collaborative attention and large language model
CN120563210A
Personalized intelligent recommendation method and system based on large language model
CN120873294A
Well engineering-oriented multi-modal abstract generation method and device and storage medium
CN121256057A