Cross-modal collaborative recommendation method and system based on large language model
By adopting a cross-modal collaborative recommendation method based on large language models in the recommendation system, the problems of data sparseness, cold start and information asymmetry between modes in the prior art are solved, and a higher recommendation accuracy and personalization are achieved.
Patent Information
- Application Number
- CN202510144944.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
The existing recommendation technology faces data sparsity, cold start problems, and information asymmetry between modals, resulting in insufficient accuracy and personalization of recommendation results.
The cross-modal collaborative recommendation method based on large language models is adopted. By inputting user identity information and project sets into the pre-trained recommendation model, the click-through rate of the project is predicted, and through feature mapping and feature fusion technology, semantic knowledge and collaborative signals are effectively fused to improve the model's understanding ability.
It improves the accuracy and personalization of large language models in project recommendations, effectively solves the problems of data sparseness and cold start, and balances the information differences between different modes.
Smart Images

Figure CN120067464A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a cross-modal collaborative recommendation method and system based on a large language model. Background Art
[0002] In recent years, with the rapid development of deep learning technology, especially the rise of large-scale pre-trained models, recommendation technology has ushered in new development opportunities. Existing recommendation methods can be divided into collaborative filtering recommendation technology, multimodal recommendation technology and large model recommendation technology. Among them, collaborative filtering, as a widely used recommendation technology, has the core idea of predicting the content that users may be interested in by analyzing the user's historical behavior data. Its basic assumption is that if two users have shown similar preferences in the past, they may be interested in the same or similar content in the future. Collaborative filtering technology occupies an important position in the recommendation system because of its simplicity and effectiveness. Especially in the case of large-scale user groups and diversified content, collaborative filtering can effectively provide personalized recommendation services. However, collaborative filtering also faces some challenges and development bottlenecks. The first is the data sparsity problem, that is, the user-item rating matrix is often very sparse, which will lead to a decrease in the accuracy of the recommendation results. The second is the cold start problem, that is, how to provide effective recommendations for new users or new items because there is a lack of sufficient historical data to calculate similarity.
[0003] Cross-modal recommendation technology is a recommendation technology that combines multiple data forms (such as text, images, audio and video, etc.). It aims to provide users with richer, more accurate and personalized recommendation results by comprehensively utilizing information from different modalities. In terms of specific implementation, the core problem that cross-modal recommendation technology needs to solve first is how to effectively map data from different modalities into the same feature space for consistency comparison and analysis. This usually involves the process of cross-modal feature representation learning, that is, extracting features from data from different modalities through technical means such as deep learning, and mapping these features into the same space through shared encoders or other similarity measurement methods. In addition to feature representation, cross-modal recommendation technology also needs to consider how to make full use of multi-modal information to make decisions during the recommendation process. This often involves the design of cross-modal association analysis and reasoning mechanisms. However, cross-modal recommendation technology still faces many challenges in practical applications, one of the most important challenges being how to deal with information asymmetry between modalities. Since data from different modalities may differ in quality and quantity, how to balance these differences and avoid excessive dominance of one modality in the recommendation results is the key to achieving effective cross-modal recommendation.
[0004] Large model recommendation technology mainly uses the generative prediction ability of large language models (LLMs) to process data in multiple dimensions such as user behavior data, item attribute information, and context environment, so as to achieve more accurate and personalized recommendations. To further improve the recommendation performance of large models, existing work based on large models attempts to integrate collaborative signals into LLM training by converting them into natural language descriptions to capture implicit feedback. However, user-item interactions described in unstructured natural language cannot effectively guide the LLM to understand the correlations hidden in the large, sparse, non-linear, and high-order model data, resulting in very limited performance improvement for these methods. On the other hand, there are also some methods that attempt to directly insert the latent representation of collaborative signals into the prompt template and directly inject the embeddings of collaborative signals together with semantic information embeddings into the input sequence of the LLM to improve the recommendation performance of the model. However, the LLM may not be able to effectively understand the collaborative signals in the above form because semantic knowledge and entity collaborative filtering embeddings come from two different modalities and are in different representation spaces. Summary of the Invention
[0005] In view of the above technical problems, this application provides a cross-modal collaborative recommendation method and system based on a large language model, which effectively utilizes the semantic knowledge and collaborative signals in cross-modal data and improves the item recommendation accuracy of the large language model.
[0006] In a first aspect, an embodiment of this application provides a cross-modal collaborative recommendation method based on a large language model, including:
[0007] Obtain the identity information of the user and the set of items to be recommended;
[0008] Input the identity information and the item set into a preset recommendation model, so that the recommendation model predicts the click-through rate of each item in the item set according to the identity information, and then outputs an item recommendation list according to the click-through rate of each item;
[0009] Among them, the recommendation model is constructed based on a large language model and trained using a training data set, and the training data set is constructed by performing prompt sequence conversion, feature mapping, and feature fusion on several recommendation domain text data and several user-item interaction data based on the large language model and collaborative filtering technology.
[0010] An embodiment of the present application provides a cross-modal collaborative recommendation method based on a large language model. A recommendation model is constructed and trained based on the large language model, and then the recommendation model is used to predict the click-through rate of each item in the item set, and an item recommendation list is output according to the prediction results to achieve personalized recommendation for users. In terms of processing the training data set, prompt sequence conversion can encode user semantic knowledge and collaborative information into the same prompt framework at the same time to better assist model training, and feature mapping and feature fusion will further fuse the prompt sequences in the same prompt framework to effectively capture complementary information between different modalities. Compared with the existing methods that attempt to convert collaborative signals into natural language descriptions, the embodiment of the present application effectively fuses the potential embedding features of semantic knowledge and collaborative signals, two different modality data, through a series of processing of recommendation domain text data and several user-item interaction data, improves the large language model's understanding ability of collaborative signals, and further improves the item recommendation accuracy of the large language model.
[0011] In a possible implementation manner, based on the large language model and collaborative filtering technology, constructing the training data set after performing prompt sequence conversion, feature mapping, and feature fusion on several recommendation domain text data and several user-item interaction data includes:
[0012] Obtain several recommendation domain text data and several user-item interaction data, where the user-item interaction data includes user features, item features, and real interaction data between users and items;
[0013] According to a preset prompt template, convert each of the recommendation domain text data and each of the user-item interaction data into several prompt sequences through natural language encoding and collaborative signal encoding, where the prompt sequences include natural language word segmentation, placeholders for users, and placeholders for items;
[0014] According to the type of each word segmentation in each of the prompt sequences, based on the large language model and collaborative filtering technology, map each of the word segmentations to corresponding potential embedding features, where the potential embedding features are collaborative filtering embedded user features, large model embedded language features, or hybrid embedded item features;
[0015] Perform dimension alignment and feature fusion on each of the potential embedding features to obtain embedding features corresponding to each of the word segmentations;
[0016] According to the order of each word segmentation in each of the prompt sequences, combine the embedding features corresponding to each word segmentation to obtain each embedding sequence corresponding to each of the prompt sequences;
[0017] Construct the training data set according to each of the embedding sequences.
[0018] An embodiment of the present application provides a method for constructing a training dataset. In order to make full use of the generative ability of large language models, the embodiment of the present application first converts recommendation domain text data and user-item interaction data into prompt sequences, and its advantages are mainly in two aspects: one is to encode and transform the recommendation domain text data and user-item interaction data into natural language, and the other is to encode collaborative signals into latent embedding representations, using a single tokenization template to represent the heterogeneous modalities of each item, so as to achieve cross-modal data fusion. Then enter the mapping stage, convert the prompt sequence into a latent embedding, and map each token to a collaborative filtering embedded user feature, a large model embedded language feature, or a hybrid embedded item feature, making data preparations for subsequent feature fusion. Finally, through dimension alignment and feature fusion, the collaborative filtering embedded user features, large model embedded language features, or hybrid embedded item features of different modalities are unified in terms of dimension and nature, obtaining the embedding features corresponding to each of the tokens and constructing a training dataset, realizing cross-modal feature fusion. The constructed training dataset has a better training effect than the existing single-modal training dataset, improving the project recommendation accuracy of large language models.
[0019] Further, mapping each of the tokens to a corresponding latent embedding feature based on the large language model and collaborative filtering technology according to the types of the tokens in each of the prompt sequences includes:
[0020] If the current token is a natural language token, obtain the corresponding large model embedded language feature by looking up through the input embedding of the large language model according to the current token;
[0021] If the current token is a user placeholder, obtain the corresponding collaborative filtering embedded user feature by means of a pre-trained collaborative filtering model according to the current token;
[0022] If the current token is an item placeholder, obtain the corresponding collaborative filtering embedded item feature by means of a pre-trained collaborative filtering model according to the current token, and then obtain the corresponding large model embedded item feature by looking up through the input embedding of the large language model according to the item description information of the item, and combine the collaborative filtering embedded item feature and the large model embedded item feature to construct the hybrid embedded item feature.
[0023] In the embodiments of the present application, different mapping methods are used for different types of word segmentation. Among them, natural language word segmentation can directly obtain the corresponding large model embedded language features through the input embedding lookup of the large language model. The user's placeholders contain the user's semantic knowledge and the user's collaborative signals. However, due to the heterogeneity problem of user text attributes among different data sets, and these inconsistent data features will further affect the user prediction accuracy. Therefore, the embodiments of the present application do not consider using the user's semantic knowledge, but only obtain the corresponding collaborative filtering embedded user features through the collaborative filtering model. Similarly, the project placeholders contain the project's semantic knowledge and the project's collaborative signals, which are respectively extracted as the collaborative filtering embedded project features and the collaborative filtering embedded project features. The embodiments of the present application construct various types of potential embedded features through different mapping methods, greatly improving the richness of the data and providing a data basis for subsequent feature fusion and model training.
[0024] Further, the dimension alignment and feature fusion of each of the potential embedded features to obtain the embedded features corresponding to each of the word segmentations includes:
[0025] Using a preset alignment network model to align each of the collaborative filtering embedded user features with the embedding space of the large language model to obtain each corresponding aligned collaborative filtering embedded user feature;
[0026] Using a preset alignment network model to align each of the collaborative filtering embedded project features in each of the hybrid embedded project features with the embedding space of the large language model to obtain each corresponding aligned collaborative filtering embedded project feature, and then using a preset GATE network model to perform feature fusion on the aligned collaborative filtering embedded project feature and the large model embedded project feature in the same hybrid embedded project feature, so as to obtain each corresponding aligned hybrid embedded project feature;
[0027] Each of the aligned collaborative filtering embedded user features, each of the aligned hybrid embedded project features, and each of the large model embedded language features are the embedded features corresponding to each of the word segmentations.
[0028] In the embodiments of the present application, dimensional alignment and feature fusion are further performed on various types of potential embedding features. The main purpose is to unify data of different modalities into the spatial dimension of the large language model, improving the understanding ability and training effect of the large language model. Specifically, in actual training, since the latent spaces of collaborative filtering embedding features and large model embedding features are not consistent, collaborative filtering embeddings cannot be directly fused with large model embeddings. Therefore, in the embodiments of the present application, an alignment network model is first used to map collaborative embedding features to the embedding space of the large model before feature fusion. Specifically, the collaborative filtering embedding user features and the collaborative filtering embedding item features are respectively aligned with the embedding space of the large language model to obtain the aligned collaborative filtering embedding user features and the aligned collaborative filtering embedding item features. At this time, there are still two types of embedding features for the item, namely the aligned collaborative filtering embedding item features and the large model embedding item features, and these two types of embedding features are still difficult to be directly understood and utilized by the large model due to their different natures. Therefore, in the embodiments of the present application, the GATE network model is used to perform feature fusion on these two types of embedding features to obtain the aligned hybrid embedding item features, realizing the fusion and unification of cross-modal data.
[0029] In one possible implementation manner, constructing and training the recommendation model using the training dataset based on the large language model includes:
[0030] Combining the large language model, a preset alignment network model, a preset GATE network model, a pre-trained collaborative filtering model, and LoRA technology to construct an initial recommendation model, where the initial recommendation model includes at least three types of parameters to be trained, namely the LoRA weight set, the alignment network model training parameters, and the GATE network model training parameters;
[0031] Using the training dataset and a preset loss function to perform two-stage training on the initial recommendation model to obtain the recommendation model;
[0032] Among them, in the first-stage training, input the training dataset into the initial recommendation model, and update the LoRA weight set according to the loss function to obtain the first recommendation model;
[0033] In the second-stage training, use the alignment network model and the GATE network model to reconstruct the training dataset to obtain an updated training dataset; input the updated training dataset into the first recommendation model, and update the alignment network model training parameters and the GATE network model training parameters according to the loss function to obtain the recommendation model.
[0034] The embodiments of the present application provide a training method for a recommendation model, which constructs an initial recommendation model based on the LoRA technology. LoRA can be used as an adapter by training low-rank network weights. Compared with the trainable parameters of the LLM, this method significantly reduces the number of parameters. In addition, in terms of the parameters to be trained, in addition to the LoRA weight set, the embodiments of the present application also need to train the training parameters of the alignment network model and the GATE network model, because the alignment network model and the GATE network model are used for feature fusion of different modality data, so their parameters also need to be trained and adjusted to further improve the prediction accuracy of the recommendation model. During the training process, the recommendation model based on the large model usually trains all parameters in an end-to-end manner for prediction. However, experience shows that the above method may weaken the performance of the actual recommendation model, especially in the cold start situation. Before the end-to-end tuning converges, the LLM may mislead the training divergence of cross-modal attention fusion. Therefore, the embodiments of the present application adopt a two-stage training method to update the LoRA weight set for improving the prediction effect and the training parameters of the alignment network model and the GATE network model for improving the feature fusion effect respectively, and finally obtain the recommendation model, improving the project recommendation accuracy of the large language model.
[0035] Further, the parameters to be trained of the initial recommendation model further include the training parameters of the collaborative filtering model. In the second-stage training, the training parameters of the alignment network model, the GATE network model, and the collaborative filtering model are updated according to the loss function to obtain the recommendation model.
[0036] In the embodiments of the present application, the training parameters of the collaborative filtering model are further increased. In this technical solution, by increasing the number of parameters to be trained, the prediction accuracy of the final recommendation model is further improved.
[0037] In a possible implementation manner, the recommendation model predicts the click-through rate of each item in the item set according to the identity information, and then outputs an item recommendation list according to the click-through rate of each item, including:
[0038] Combining the identity information with each item in the item set respectively to construct a number of user-item combination data;
[0039] Performing prompt sequence conversion, feature mapping, and feature fusion on each of the user-item combination data respectively to construct corresponding embedding sequences;
[0040] Predicting the click-through rate according to each of the embedding sequences to obtain the click-through rate of each item;
[0041] Sorting each item in the item set according to the click-through rate of each item to generate a sorted item list;
[0042] Extract the top N items from the sorted item list, construct and output the item recommendation list, where N is a preset value.
[0043] In a second aspect, correspondingly, an embodiment of the present application provides a cross-modal collaborative recommendation system based on a large language model, including an acquisition module, a recommendation module, and a model training module;
[0044] The acquisition module is used to acquire the identity information of the user and the set of items to be recommended;
[0045] The recommendation module is used to input the identity information and the item set into a preset recommendation model, so that the recommendation model predicts the click-through rate of each item in the item set according to the identity information, and then outputs an item recommendation list according to the click-through rate of each item;
[0046] The model training module is used to construct a recommendation model based on a large language model and train it using a training data set. The training data set is constructed by performing prompt sequence conversion, feature mapping, and feature fusion on a number of recommendation domain text data and a number of user-item interaction data based on the large language model and collaborative filtering technology.
[0047] In a possible implementation manner, constructing the training data set by performing prompt sequence conversion, feature mapping, and feature fusion on a number of recommendation domain text data and a number of user-item interaction data based on the large language model and collaborative filtering technology includes:
[0048] Obtain a number of recommendation domain text data and a number of user-item interaction data, where the user-item interaction data includes user features, item features, and real interaction data between the user and the item;
[0049] According to a preset prompt template, convert each of the recommendation domain text data and each of the user-item interaction data into a number of prompt sequences through natural language encoding and collaborative signal encoding. The prompt sequences include natural language word segmentation, placeholders for users, and placeholders for items;
[0050] According to the type of each word segment in each of the prompt sequences, map each of the word segments to a corresponding latent embedding feature based on the large language model and collaborative filtering technology. The latent embedding feature is a collaborative filtering embedded user feature, a large model embedded language feature, or a hybrid embedded item feature;
[0051] Perform dimension alignment and feature fusion on each of the latent embedding features to obtain an embedding feature corresponding to each of the word segments;
[0052] According to the order of each word segment in each of the prompt sequences, combine the embedding features corresponding to each word segment to obtain each embedding sequence corresponding to each of the prompt sequences;
[0053] Construct the training dataset based on the obtained embedding sequences.
[0054] In a possible implementation manner, the training the recommendation model by constructing and using a training dataset based on a large language model includes:
[0055] Combine the large language model, a preset alignment network model, a preset GATE network model, a pre-trained collaborative filtering model, and LoRA technology to construct an initial recommendation model, where the initial recommendation model includes at least three parameters to be trained, namely a LoRA weight set, alignment network model training parameters, and GATE network model training parameters;
[0056] Use the training dataset and a preset loss function to perform two-stage training on the initial recommendation model to obtain the recommendation model;
[0057] Among them, in the first-stage training, input the training dataset into the initial recommendation model, and update the LoRA weight set according to the loss function to obtain a first recommendation model;
[0058] In the second-stage training, use the alignment network model and the GATE network model to reconstruct the training dataset to obtain an updated training dataset; input the updated training dataset into the first recommendation model, and update the alignment network model training parameters and GATE network model training parameters according to the loss function to obtain the recommendation model. Description of the Drawings
[0059] Figure 1 : It is a schematic flowchart of a cross-modal collaborative recommendation method based on a large language model provided by an embodiment of the present application.
[0060] Figure 2 : It is a schematic flowchart of constructing a training dataset in a cross-modal collaborative recommendation method based on a large language model provided by an embodiment of the present application.
[0061] Figure 3 : It is a schematic overall framework diagram of a cross-modal collaborative recommendation method based on a large language model provided by an embodiment of the present application.
[0062] Figure 4 : It is a schematic structural diagram of a cross-modal collaborative recommendation system based on a large language model provided by an embodiment of the present application. Detailed Embodiments
[0063] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.
[0064] It should be noted that the step numbers in the text are only for the convenience of explaining specific embodiments and do not serve to limit the execution order of the steps. In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0065] Throughout the specification, the large model or large language model described in the present application refers to a deep learning model trained using a large-scale dataset. Such models have millions or more parameters and can exhibit excellent performance when dealing with complex tasks. Through a combination of pre-training and fine-tuning, large models can perform well in a variety of downstream tasks, including natural language processing, computer vision, speech recognition, and other fields. Their core advantage lies in their strong generalization ability, that is, they can maintain stable output quality even on unseen data. In addition, with the improvement of hardware computing power and the development of distributed training technology, the training efficiency of large models has been significantly improved, further promoting their application in all walks of life. For example, in the field of natural language processing, large models based on the Transformer architecture such as BERT, GPT series, etc. have become standard tools for tasks such as text generation, machine translation, and sentiment analysis; in the field of computer vision, models such as Vision Transformer are gradually replacing traditional convolutional neural networks and becoming the new favorites for tasks such as image classification and object detection.
[0066] The collaborative filtering (CF) described in this application is a widely used classical recommendation algorithm. Its basic principle is to analyze the historical behavior data of users, find user groups or item sets with similar preferences, and then recommend items that users may be interested in. Collaborative filtering can be divided into two categories: user-based collaborative filtering (UBCF) and item-based collaborative filtering (IBCF). The former recommends items by calculating the similarity between users, while the latter realizes recommendations by calculating the similarity between items. However, with the growth of data scale, collaborative filtering faces problems such as sparsity and cold start, and needs to combine technical means such as matrix factorization and deep learning to improve the accuracy and real-time performance of recommendations.
[0067] The multi-modal technology described in this application refers to the technical means of integrating multiple data expression forms (such as text, image, audio, video, etc.), aiming to improve the machine's understanding and processing ability of complex information by comprehensively using data of different modalities. The core of this technology lies in cross-modal feature representation learning, that is, mapping different types of media data to the same feature space for consistent analysis and processing. Through the deep learning framework, multi-modal technology can extract meaningful features from various data and generate more comprehensive and rich information representations by fusing these features, thus providing a solid foundation for subsequent applications. For example, in the tasks of image recognition and description generation, the system can not only recognize the objects in the image, but also understand its context meaning and describe it in natural language. Multi-modal technology has been widely used in many fields, including but not limited to intelligent recommendation systems, virtual assistants, medical and health analysis, and autonomous driving. For example, in virtual assistants, multi-modal technology enables the device to simultaneously understand voice commands and gesture actions, providing a more natural human-computer interaction experience. In addition, in the field of medical and health, multi-modal technology can be used to integrate patients' clinical records, imaging data and genetic data to assist doctors in disease diagnosis and treatment plan formulation.
[0068] The embodiments of this application aim to solve the click-through rate (CTR) prediction task in the recommendation system. The goal of CTR prediction is to predict the advertisements or items that the target user may click. CTR prediction is usually formulated as a supervised binary classification task, which usually involves a set of user set U and item set I. Each instance represents a triple (u, i, y) of the interaction probability between the user and the item, where u ∈ U represents the user, i ∈ I represents the item, and y ∈ {0, 1} represents the actual user and item interaction status (i.e., clicked or not clicked). Given the features x of the useru and the feature x of the project i , the goal of the click-through rate prediction model in the recommendation system is to accurately predict the posterior probability P(u|x u ,x i ).
[0069] Classic collaborative filtering recommendation technology aims to predict the possible interaction status of users based on the low-dimensional embeddings of user and project collaborative signals, where the embedding of user u is represented as and the embedding of project i is represented as In most traditional collaborative filtering recommendation models, the above entity embeddings are all calculated and obtained through the representation model ENC, where the trainable parameters of the representation model are Θ E .
[0070]
[0071] Then, the encoded user and project embeddings are injected into the interaction network INT() for prediction, and the prediction result is represented as where Θ T are the trainable parameters of the interaction network. In most existing application models, INT() is a non-linear parameter function. Usually, this model is optimized by minimizing the cross-entropy loss function, which measures the prediction error between the predicted value and the true label y in the training data. However, when using large language models for recommendation, the embodiments of this application need to convert the interaction history between users and projects into a mixed prompt P u,i , and use for prediction, where f is obtained from the LLM4Rec model parameterized by Θ.
[0072] Embodiment 1:
[0073] As Figure 1 shown, Embodiment 1 provides a cross-modal collaborative recommendation method based on a large language model, including steps S1-S2:
[0074] Step S1, obtain the identity information of the user and the set of items to be recommended;
[0075] Step S2, input the identity information and the item set into a preset recommendation model, so that the recommendation model predicts the click-through rate of each item in the item set according to the identity information, and then outputs an item recommendation list according to the click-through rate of each item;
[0076] Among them, the recommendation model is constructed based on a large language model and trained using a training data set, and the training data set is constructed by performing prompt sequence conversion, feature mapping, and feature fusion on a number of text data in the recommendation field and a number of user-item interaction data based on the large language model and collaborative filtering technology.
[0077] An embodiment of the present application provides a cross-modal collaborative recommendation method based on a large language model. A recommendation model is constructed and trained based on the large language model, and then the recommendation model is used to predict the click-through rate of each item in the item set, and an item recommendation list is output according to the prediction result to achieve personalized recommendation for users. In terms of the processing of the training data set, prompt sequence conversion can encode user semantic knowledge and collaborative information into the same prompt framework at the same time to better assist model training, and feature mapping and feature fusion will further fuse the prompt sequences in the same prompt framework to effectively capture complementary information between different modalities. Compared with the existing method that attempts to convert collaborative signals into natural language descriptions, the embodiment of the present application effectively fuses the potential embedding features of semantic knowledge and collaborative signals, two different modality data, through a series of processing of text data in the recommendation field and a number of user-item interaction data, improves the large language model's understanding ability of collaborative signals, and further improves the item recommendation accuracy of the large language model.
[0078] In a preferred embodiment, when the method provided by the embodiment of the present application is applied to the e-commerce field, in step S1, the identity information of the user can be the user's historical purchase records and user ID information, and the item set to be recommended can be different modality data such as text descriptions of various commodities and corresponding commodity pictures.
[0079] In a possible implementation manner, in step S2, the recommendation model predicts the click-through rate of each item in the item set according to the identity information, and then outputs an item recommendation list according to the click-through rate of each item, including:
[0080] Combining the identity information with each item in the item set respectively to construct a number of user-item combination data;
[0081] Performing prompt sequence conversion, feature mapping, and feature fusion on each of the user-item combination data respectively to construct corresponding embedding sequences;
[0082] Predicting the click-through rate according to each of the embedding sequences to obtain the click-through rate of each item;
[0083] Sorting each item in the item set according to the click-through rate of each item to generate a sorted item list;
[0084] Extract the top N items from the sorted item list, construct and output the item recommendation list, where N is a preset value.
[0085] In a possible implementation, based on the large language model and collaborative filtering technology, the training data set is constructed after performing prompt sequence conversion, feature mapping, and feature fusion on several recommended domain text data and several user-item interaction data, as Figure 2 shown, including steps S301 - S306:
[0086] Step S301, obtain several recommended domain text data and several user-item interaction data, where the user-item interaction data includes user features, item features, and real interaction data between users and items;
[0087] Step S302, according to a preset prompt template, convert each of the recommended domain text data and each of the user-item interaction data into several prompt sequences through natural language encoding and collaborative signal encoding, where the prompt sequences include natural language tokenization, placeholders for users, and placeholders for items;
[0088] Step S303, according to the type of each token in each of the prompt sequences, based on the large language model and collaborative filtering technology, map each of the tokens to a corresponding latent embedding feature, where the latent embedding feature is a collaborative filtering embedded user feature, a large model embedded language feature, or a hybrid embedded item feature;
[0089] Step S304, perform dimension alignment and feature fusion on each of the latent embedding features to obtain embedding features corresponding to each of the tokens;
[0090] Step S305, according to the order of each token in each of the prompt sequences, combine the embedding features corresponding to each token to obtain each embedding sequence corresponding to each of the prompt sequences;
[0091] Step S306, construct the training data set according to each of the embedding sequences.
[0092] The embodiment of the present application provides a method for constructing a training data set. In order to make full use of the generative capabilities of a large language model, the embodiment of the present application first converts the recommendation field text data and the user-project interaction data into a prompt sequence. Its advantages are mainly in two aspects: first, the recommendation field text data and the user-project interaction data are encoded and converted into natural language, and second, the collaborative signal is encoded into a potential embedding representation, and a single word segmentation template is used to represent the heterogeneous modality of each project, thereby realizing cross-modal data fusion. Then enter the mapping stage, convert the prompt sequence into a potential embedding, and map each word segmentation into collaborative filtering embedded user features, large model embedded language features or mixed embedded project features, so as to prepare data for subsequent feature fusion. Finally, through dimensional alignment and feature fusion, the collaborative filtering embedded user features, large model embedded language features or mixed embedded project features of different modalities are unified in dimension and nature, and the embedded features corresponding to each of the word segments are obtained and the training data set is constructed, realizing cross-modal feature fusion. The constructed training data set has better training effect than the existing single-modal training data set, and improves the accuracy of project recommendation of the large language model.
[0093] In a preferred embodiment, in step S302, the prompt template is as follows:
[0094] #Question: Users rated the following items highly: u In addition, in the user feature {U i} encodes information about the user’s preferences. After taking all available information, the model predicts whether the user will like the item {I i}.
[0095] #Answer: "Yes" or "No"
[0096] The present invention will use several special participles as placeholders to represent user and item features in the prompt template. In terms of item representation, the present invention uses {I i} to represent the ID (i) of the project entity, and {I u} represents a special item word list that the user (u) has interacted with. In terms of user representation, the present invention adopts {U i} is used as a placeholder to represent the ID (u) of the user entity. Therefore, the prompt template proposed in the present invention mainly consists of the following three parts: natural language task description, user entity and project entity. It should be noted that, unlike the traditional method of combining two different modal data in the prompt Separately, the present invention designs a single word segmentation template to represent the heterogeneous modalities of each item, thereby providing a data basis for subsequent cross-modal data fusion.
[0097] Further, in step S303, mapping each of the word segments to corresponding latent embedding features based on the large language model and collaborative filtering technology according to the types of the word segments in each of the prompt sequences includes:
[0098] If the current word segment is a natural language word segment, obtain the corresponding large model embedded language feature through the input embedding lookup of the large language model according to the current word segment;
[0099] If the current word segment is a user placeholder, obtain the corresponding collaborative filtering embedded user feature through the pre-trained collaborative filtering model according to the current word segment;
[0100] If the current word segment is an item placeholder, obtain the corresponding collaborative filtering embedded item feature through the pre-trained collaborative filtering model according to the current word segment, and then obtain the corresponding large model embedded item feature through the input embedding lookup of the large language model according to the item description information of the item, and combine the collaborative filtering embedded item feature and the large model embedded item feature to construct the hybrid embedded item feature.
[0101] In a preferred embodiment, encode the user feature and two types of item features based on the large language model and collaborative filtering technology to obtain the hybrid embedded item feature and the collaborative filtering embedded user feature. The hybrid embedded item feature is represented as where represents the semantic knowledge embedding of the item text attribute, T i is the total number of word segments in the attributes of item i, and each word segment embedding is a low-dimensional vector, is the collaborative filtering embedded item feature encoded by the pre-trained collaborative filtering model. Similarly, the collaborative filtering embedded user feature is represented as It should be noted that the method proposed in the embodiment of the present application does not include the semantic knowledge of the user. The main reason is that there are heterogeneity problems in user text attributes between different data sets, and these inconsistent data features will further affect the user prediction accuracy.
[0102] Specifically, it includes the following process: First, for the task description and user or item attributes (such as titles, etc.), map the above data to the latent embedding of semantic knowledge for subsequent large model training; then, for the collaborative signals of users and items, map these signals to the latent embedding based on the pre-trained collaborative filtering recommendation model. Assume that the hybrid prompt of user u and item i is P u,iFor each token tk after tokenization, the present invention will obtain its potential embedding through the following rules: If tk is a natural language token, the present invention will directly look up and obtain its corresponding potential embedding through the input embedding of the large model; If tk is a placeholder of the user, then the present invention obtains the potential embedding from the pre-trained collaborative filtering model If tk is a placeholder of the item, then the present invention will obtain embeddings of two modalities, both the embedding of the large model with semantic knowledge and the embedding of the collaborative signal. In addition, the present invention first obtains the collaborative embedding of the item from the collaborative model Then, extract the text attributes of the item from the database and obtain the entity embedding by looking up through the input embedding of the large model Finally, input the above two embeddings into the adaptive fusion strategy in the next part.
[0103] In the above way, each token in the prompt sequence P u,i is mapped to its corresponding potential embedding, and these embeddings will be used as the input sequence of the LLM for model training.
[0104] Further, in step S304, the dimension alignment and feature fusion of each of the potential embedding features to obtain the embedding features corresponding to each of the tokens includes:
[0105] Use a preset alignment network model to align each of the collaborative filtering embedding user features with the embedding space of the large language model to obtain each corresponding aligned collaborative filtering embedding user feature;
[0106] Use a preset alignment network model to align each of the collaborative filtering embedding item features in each of the hybrid embedding item features with the embedding space of the large language model to obtain each corresponding aligned collaborative filtering embedding item feature, and then use a preset GATE network model to perform feature fusion on the aligned collaborative filtering embedding item feature and the large model embedding item feature in the same hybrid embedding item feature, so as to obtain each corresponding aligned hybrid embedding item feature;
[0107] Each of the aligned collaborative filtering embedding user features, each of the aligned hybrid embedding item features, and each of the large model embedding language features are the embedding features corresponding to each of the tokens.
[0108] In the embodiments of the present application, dimensional alignment and feature fusion are further performed on various types of potential embedding features. The main purpose is to unify data of different modalities into the spatial dimension of the large language model, improving the understanding ability and training effect of the large language model. Specifically, in actual training, since the latent spaces of collaborative filtering embedding features and large model embedding features are not consistent, collaborative filtering embeddings cannot be directly fused with large model embeddings. Therefore, in the embodiments of the present application, an alignment network model is first used to map collaborative embedding features to the embedding space of the large model before feature fusion. Specifically, the collaborative filtering embedding user features and the collaborative filtering embedding item features are respectively aligned with the embedding space of the large language model to obtain the aligned collaborative filtering embedding user features and the aligned collaborative filtering embedding item features. At this time, there are still two types of embedding features for the item, namely the aligned collaborative filtering embedding item features and the large model embedding item features. However, due to their different natures, these two types of embedding features are still difficult to be directly understood and utilized by the large model. Therefore, in the embodiments of the present application, a GATE network model is used to perform feature fusion on these two types of embedding features to obtain the aligned hybrid embedding item features, realizing the fusion and unification of cross-modal data.
[0109] In a preferred embodiment, assume a given alignment network model Then the alignment process can be expressed as follows:
[0110]
[0111] where, Θ A is a set of trainable parameters. are the aligned collaborative filtering embedding user features and the aligned hybrid embedding item features after alignment. In actual model training, the alignment network model can usually adopt the classic multi-layer perceptron method.
[0112] However, only aligning the dimensionality of collaborative embeddings and the text embeddings of the large model is not enough, because the latent embedding features of collaborative signals are still difficult to be directly understood and utilized by the large model due to their different natures. For this reason, the embodiments of the present application propose a cross-modal collaborative fusion strategy based on the attention mechanism. Inspired by the cross-gate mechanism, this embodiment trains a learnable GATE network model to fuse two different modalities of embeddings in the same dimension in an adaptive manner. Specifically, for each attribute tokenization of item i, the GATE network model can learn a vector of fusion weights which is defined as:
[0113]
[0114] where, is a set of trainable parameters. Finally, the present invention uses the weight vector α to fuse the large model embedding item features Align and Collaboratively Filter to Embed Project Features The weight vector is adapted to the collaborative signals and semantic knowledge of each project as follows:
[0115]
[0116] Among them, the operator ⊙ represents element-wise multiplication.
[0117] Finally, this embodiment uses the mapped and fused embedding features to form the final embedding sequence and inputs it into the LLM for training or prediction. For each token in the prompt P u,i in this application, the present application replaces the token with the corresponding embedding and follows the same rules as the mapping step. For example, if it is a natural language token, this embodiment replaces it with the LLM embedding; if it is a placeholder for user u, this embodiment uses the aligned collaborative embedding to replace it; if it is a placeholder for item i, this embodiment uses to replace it. These embeddings are concatenated in the order of the tokens in the prompt sequence P u,i .
[0118] In a possible implementation manner, constructing the recommendation model based on the large language model and training it using the training dataset includes:
[0119] Combining the large language model, a preset alignment network model, a preset GATE network model, a pre-trained collaborative filtering model, and LoRA technology to construct an initial recommendation model, where the initial recommendation model includes at least three parameters to be trained, namely the LoRA weight set, the alignment network model training parameters, and the GATE network model training parameters;
[0120] Using the training dataset and a preset loss function to perform two-stage training on the initial recommendation model to obtain the recommendation model;
[0121] Among them, in the first-stage training, input the training dataset into the initial recommendation model and update the LoRA weight set according to the loss function to obtain the first recommendation model;
[0122] In the second-stage training, use the alignment network model and the GATE network model to reconstruct the training dataset to obtain an updated training dataset; input the updated training dataset into the first recommendation model and update the alignment network model training parameters and the GATE network model training parameters according to the loss function to obtain the recommendation model.
[0123] The embodiment of the present application provides a training method for a recommendation model. An initial recommendation model is constructed based on the LoRA technology. LoRA can be used as an adapter by training low-rank network weights. Compared with the trainable parameters of the LLM, this method significantly reduces the number of parameters. In addition, in terms of the parameters to be trained, in addition to the LoRA weight set, the embodiment of the present application also needs to train the training parameters of the alignment network model and the GATE network model, because the alignment network model and the GATE network model are used for feature fusion of different modalities of data, so their parameters also need to be trained and adjusted to further improve the prediction accuracy of the recommendation model. During the training process, the recommendation model based on the large model usually trains all parameters in an end-to-end manner for prediction. However, experience shows that the above method may weaken the performance of the actual recommendation model, especially in the cold start situation. Before the end-to-end tuning converges, the LLM may mislead the training divergence of cross-modal attention fusion. Therefore, the embodiment of the present application adopts a two-stage training method to update the LoRA weight set for improving the prediction effect and the training parameters of the alignment network model and the GATE network model for improving the feature fusion effect respectively, and finally obtains the recommendation model, improving the project recommendation accuracy of the large language model.
[0124] Further, the parameters to be trained of the initial recommendation model further include the training parameters of the collaborative filtering model. In the second-stage training, the training parameters of the alignment network model, the GATE network model, and the collaborative filtering model are updated according to the loss function to obtain the recommendation model.
[0125] In the embodiment of the present application, the training parameters of the collaborative filtering model are further increased. In this technical solution, by increasing the number of parameters to be trained, the prediction accuracy of the final recommendation model is further improved.
[0126] In a preferred embodiment, the training of the recommendation model mainly includes the learning objective and the two-stage training strategy.
[0127] (1) Learning objective. Although the pre-trained LLM can perform many prediction tasks in the zero-shot setting, it is not specifically trained for collaborative signals. Therefore, directly using the fused embedding sequence for recommendation cannot achieve the optimal effect. Therefore, the embodiment of the present application uses the LoRA module to fine-tune the LLM. LoRA can be used as an adapter by training low-rank network weights. Compared with the trainable parameters of the LLM, this method significantly reduces the number of parameters. Among them, the set of trainable parameters of the method proposed by the present invention is Θ = {ΘE, ΘA, ΘG, ΘL}, where ΘL is the LoRA weight set; the fine-tuning of ΘE is optional because it has been pre-trained, and the embodiment of the present application only uses the ENC network in the collaborative model; ΘA and ΘG are the training parameters of the alignment model ALG and the GATE network.
[0128] Traditional click-through rate prediction models can generate a Bernoulli distribution of user-item interaction probabilities, while language models can generate a probability distribution over the entire vocabulary. To align the predictions with traditional click-through rate prediction models and achieve alignment with the output of traditional CTR models, embodiments of this application define probabilities as the metric for positive predictions, as the metric for negative predictions. Here, "Yes" and "No" are two tokens in the LLM vocabulary. For this purpose, embodiments of this application will optimize the LLM from two perspectives. First, embodiments of this application aim to align the predicted labels with the true labels y. Second, the present invention further constrains the relationship between p yes and p no to ensure that p yes > p no when y = 1, and vice versa. Therefore, the present invention optimizes the LLM by minimizing the following loss:
[0129]
[0130] where, represents the classification loss corresponding to the first objective, is the ranking loss corresponding to the second objective. The present invention adopts binary cross-entropy (BCE) loss, adopts Bayesian personalized ranking (BPR) loss. k is a scalar used to balance the weights between the two objectives.
[0131] Two-stage training. Usually, recommendation models based on large models train all parameters in an end-to-end manner for prediction. However, experience shows that the above method may weaken the performance of the actual recommendation model, especially in the cold start situation. Before the end-to-end tuning converges, the LLM may mislead the training divergence of cross-modal attention fusion. Therefore, the present invention adopts a two-stage training strategy, dividing Θ into two non-overlapping subsets Θ1 and Θ2, where Θ1 = {ΘL}, Θ2 = {ΘE, ΘA, ΘG}. We first minimize the objective function and only update Θ1 until it converges, then separately update Θ2 until it also converges, and finally continue to minimize the objective function to obtain the prediction result.
[0132] The principle of two-stage training is as follows: After the first stage, the LLM has adapted to the recommendation task. However, at this time, the collaborative signal is still disturbed by noise because both the ALG and GATE networks are not well trained, thus limiting the recommendation ability. Therefore, in the second stage, the present invention enhances the representation ability of cross-modal signals by focusing on the cross-modal attention collaborative fusion strategy.
[0133] In a preferred embodiment, the training and prediction framework of the recommendation model is as follows Figure 3 shown. Specifically, first, the embodiments of the present application define the recommendation domain tasks to be solved and completed, that is, combining large model technology and traditional collaborative recommendation technology to collaboratively improve the prediction accuracy of click-through rate in the recommendation domain; then, the embodiments of the present application convert the recommendation domain text data and user-item interaction data into a mixed prompt sequence, which encodes structured collaborative signals and semantic knowledge into heterogeneous modal data. Subsequently, the present invention proposes a framework combining large model and traditional collaborative recommendation technology to implement an attention-enabled cross-modal data collaborative fusion strategy, and then feeds the above-encoded heterogeneous modal data into the fine-tuning stage of the large model, including a mapping stage that maps the tokens in the mixed prompt to LLM or CF embeddings accordingly; and a fusion stage that adaptively aligns and fuses heterogeneous embeddings through the GATE network architecture. The above fully fused potential heterogeneous modal data embeddings are injected into the fine-tuning training of the LLM to obtain the final recommendation model.
[0134] The method proposed by the embodiments of the present application is applicable to various recommendation scenarios, including e-commerce, social media, biomedicine and other fields. Taking the e-commerce field as an example, the model of the present invention can process various types of data, including different modal data such as product text descriptions, user historical purchase records, user ID information, and product pictures; then the method proposed by the present invention fuses these cross-modal data through a mixed prompt template strategy for model fine-tuning training; then the model calculates the probability value of each candidate product being clicked, sorts the candidate products according to these probability values, and finally pushes the Top-N products in the sorted list to the target user. The method proposed in this embodiment can also be applied to other fields such as social media and biomedicine, and more accurate user click-through rate prediction and personalized recommendation results can be achieved by fusing different modal data.
[0135] Compared with the prior art, the advantages of a cross-modal collaborative recommendation method based on a large language model proposed by the embodiments of the present application are as follows.
[0136] 1. At a high level, the goal of existing work is to integrate collaborative signals by directly injecting the latent embeddings pre-trained by different collaborative models into the LLM. However, the embodiments of the present application aim to solve the problem that the LLM cannot effectively understand collaborative embeddings by unifying two different modal data into the same space, rather than just introducing a new modal data.
[0137] 2. In existing work, the mixed prompt encodes text and collaborative signals in different fields (tokens) respectively, without considering their fusion. In contrast, the present invention uses a single token to unify the two different modal representations of each item.
[0138] 3. In existing work, the two modalities have not been effectively fused, and the potential embeddings of semantic knowledge (compatible with pre-trained LLMs) and collaborative signals (extracted from pre-trained CF models) have not been directly input into the LLMs. In sharp contrast, the present invention fuses the above two modality data through a trainable cross-gate mechanism.
[0139] In summary, although some existing work has realized the importance of introducing collaborative signals, they only directly inject the pre-trained collaborative signals into the LLMs, ignoring the differences between the two modalities. Therefore, the LLMs may not be able to understand the pre-trained collaborative signals. In contrast, this embodiment uses a cross-modal collaborative fusion strategy based on the attention mechanism to first fully fuse the potential embeddings of these two heterogeneous modalities into the same space, improving the project recommendation accuracy of the large language model and providing a new solution idea for the fusion of the two modalities in the large model recommendation paradigm.
[0140] Embodiment 2:
[0141] As Figure 4 shown, correspondingly, Embodiment 2 provides a cross-modal collaborative recommendation system based on a large language model, including an acquisition module 10, a recommendation module 20, and a model training module 30;
[0142] Among them, the acquisition module 10 is used to acquire the identity information of the user and the set of items to be recommended;
[0143] The recommendation module 20 is used to input the identity information and the set of items into a preset recommendation model, so that the recommendation model predicts the click-through rate of each item in the set of items according to the identity information, and then outputs an item recommendation list according to the click-through rate of each item;
[0144] The model training module 30 is used to construct and train the recommendation model based on the large language model using a training data set, and the training data set is constructed by performing prompt sequence conversion, feature mapping, and feature fusion on a number of recommendation domain text data and a number of user-item interaction data based on the large language model and collaborative filtering technology.
[0145] In a possible implementation manner, constructing the training data set by performing prompt sequence conversion, feature mapping, and feature fusion on a number of recommendation domain text data and a number of user-item interaction data based on the large language model and collaborative filtering technology includes:
[0146] Acquire a number of recommendation domain text data and a number of user-item interaction data, and the user-item interaction data includes user features, item features, and real interaction data between the user and the item;
[0147] According to a preset prompt template, convert each of the recommended domain text data and each of the user-item interaction data into a number of prompt sequences through natural language encoding and collaborative signal encoding. The prompt sequences include natural language word segmentation, placeholders for users, and placeholders for items;
[0148] Based on the types of each word segmentation in each of the prompt sequences, map each of the word segmentations to corresponding latent embedding features based on the large language model and collaborative filtering technology. The latent embedding features are collaborative filtering embedded user features, large model embedded language features, or hybrid embedded item features;
[0149] Perform dimension alignment and feature fusion on each of the latent embedding features to obtain embedding features corresponding to each of the word segmentations;
[0150] According to the order of each word segmentation in each of the prompt sequences, combine the embedding features corresponding to each word segmentation to obtain each embedding sequence corresponding to each of the prompt sequences;
[0151] Construct the training data set based on each of the embedding sequences.
[0152] In a possible implementation manner, the recommended model is constructed based on the large language model and trained using the training data set, including:
[0153] Combine the large language model, a preset alignment network model, a preset GATE network model, a pre-trained collaborative filtering model, and LoRA technology to construct an initial recommended model. The initial recommended model includes at least three parameters to be trained, namely LoRA weight sets, alignment network model training parameters, and GATE network model training parameters;
[0154] Use the training data set and a preset loss function to perform two-stage training on the initial recommended model to obtain the recommended model;
[0155] Among them, in the first-stage training, input the training data set into the initial recommended model, and update the LoRA weight sets according to the loss function to obtain a first recommended model;
[0156] In the second-stage training, use the alignment network model and the GATE network model to reconstruct the training data set to obtain an updated training data set; input the updated training data set into the first recommended model, and update the alignment network model training parameters and GATE network model training parameters according to the loss function to obtain the recommended model.
[0157] An embodiment of the present application provides a cross-modal collaborative recommendation system based on a large language model. A recommendation model is constructed and trained based on the large language model, and then the recommendation model is used to predict the click-through rate of each item in the item set, and an item recommendation list is output according to the prediction result to achieve personalized recommendation for users. In terms of the processing of the training data set, the prompt sequence conversion can encode user semantic knowledge and collaborative information into the same prompt framework at the same time to better assist model training, and feature mapping and feature fusion will further fuse the prompt sequences in the same prompt framework to effectively capture complementary information between different modalities. Compared with the existing method that attempts to convert collaborative signals into natural language descriptions, the embodiment of the present application effectively fuses the potential embedding features of semantic knowledge and collaborative signals, two different modality data, through a series of processing of the text data in the recommendation field and several user-item interaction data, improves the large language model's ability to understand collaborative signals, and further improves the item recommendation accuracy of the large language model.
[0158] The more detailed working principle and step flow of this embodiment can but are not limited to refer to the relevant records of Embodiment 1.
[0159] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A cross-modal collaborative recommendation method based on a large language model, characterized in that: include: Obtain the user's identity information and the set of items to be recommended; Inputting the identity information and the item set into a preset recommendation model, so that the recommendation model predicts the click rate of each item in the item set according to the identity information, and then outputs an item recommendation list according to the click rate of each item; The recommendation model is constructed based on a large language model and trained using a training data set. The training data set is constructed based on the large language model and collaborative filtering technology by performing prompt sequence conversion, feature mapping and feature fusion on a number of recommendation domain text data and a number of user-item interaction data.
2. The cross-modal collaborative recommendation method based on a large language model as claimed in claim 1, characterized in that: Based on the large language model and collaborative filtering technology, a plurality of recommendation domain text data and a plurality of user-item interaction data are subjected to prompt sequence conversion, feature mapping and feature fusion to construct the training data set, including: Acquire a number of recommendation domain text data and a number of user-item interaction data, wherein the user-item interaction data includes user features, item features, and real interaction data between users and items; According to a preset prompt template, each of the recommendation field text data and each of the user-item interaction data is converted into a plurality of prompt sequences through natural language encoding and collaborative signal encoding, wherein the prompt sequence includes natural language segmentation, user placeholders and item placeholders; According to the type of each word segment in each prompt sequence, based on the large language model and collaborative filtering technology, each word segment is mapped to a corresponding potential embedding feature, wherein the potential embedding feature is a collaborative filtering embedded user feature, a large model embedded language feature, or a mixed embedded project feature; Performing dimension alignment and feature fusion on each of the potential embedding features to obtain an embedding feature corresponding to each of the word segments; According to the order of each word segment in each prompt sequence, the embedding features corresponding to each word segment are combined to obtain each embedding sequence corresponding to each prompt sequence; The training data set is constructed and obtained according to the respective embedding sequences.
3. The cross-modal collaborative recommendation method based on a large language model as claimed in claim 2, characterized in that: According to the type of each word segment in each prompt sequence, based on the large language model and collaborative filtering technology, each word segment is mapped to a corresponding potential embedding feature, including: If the current word segmentation is a natural language word segmentation, then according to the natural language word segmentation, the corresponding large model embedding language feature is obtained through the input embedding search of the large language model; If the current segmented word is a placeholder for the user, then the corresponding collaborative filtering embedded user features are obtained through a pre-trained collaborative filtering model according to the placeholder for the user; If the current word segmentation is a placeholder for the project, then based on the placeholder of the project, the corresponding collaborative filtering embedded project features are obtained through the pre-trained collaborative filtering model, and then based on the project description information of the project, the corresponding large model embedded project features are obtained through the input embedding search of the large language model, and the collaborative filtering embedded project features are combined with the large model embedded project features to construct the hybrid embedded project features.
4. The cross-modal collaborative recommendation method based on a large language model as claimed in claim 3, characterized in that: The step of performing dimension alignment and feature fusion on each of the potential embedding features to obtain an embedding feature corresponding to each of the word segments includes: Using a preset alignment network model, each of the collaborative filtering embedded user features is aligned with the embedding space of the large language model to obtain each corresponding aligned collaborative filtering embedded user feature; Using a preset alignment network model, each collaborative filtering embedded project feature in each of the hybrid embedded project features is aligned with the embedding space of the large language model to obtain each corresponding aligned collaborative filtering embedded project feature, and then using a preset GATE network model, the aligned collaborative filtering embedded project feature and the large model embedded project feature in the same hybrid embedded project feature are feature-fused to obtain each corresponding aligned hybrid embedded project feature; Each of the aligned collaborative filtering embedded user features, each of the aligned mixed embedded project features, and each of the large model embedded language features are embedded features corresponding to each of the word segmentations.
5. The cross-modal collaborative recommendation method based on a large language model as claimed in claim 1, characterized in that: The recommendation model is constructed based on a large language model and trained using a training data set, including: In combination with the large language model, the preset alignment network model, the preset GATE network model, the pre-trained collaborative filtering model and the LoRA technology, an initial recommendation model is constructed, wherein the initial recommendation model includes at least three parameters to be trained, namely, a LoRA weight set, an alignment network model training parameter and a GATE network model training parameter; Inputting the training data set into the initial recommendation model, and updating the LoRA weight set according to the loss function to obtain a first recommendation model; Reconstruct the training data set using the aligned network model and the GATE network model to obtain an updated training data set; The updated training data set is input into the first recommendation model, and the alignment network model training parameters and the GATE network model training parameters are updated according to the loss function to obtain the recommendation model.
6. A cross-modal collaborative recommendation method based on a large language model as claimed in claim 5, characterized in that: The parameters to be trained of the initial recommendation model further include collaborative filtering model training parameters; and updating the alignment network model training parameters and the GATE network model training parameters according to the loss function to obtain the recommendation model includes: Update collaborative filtering model training parameters according to the loss function; The recommendation model is obtained through the updated alignment network model training parameters, GATE network model training parameters and collaborative filtering model training parameters.
7. A cross-modal collaborative recommendation method based on a large language model as described in any one of claims 1 to 6, characterized in that: The recommendation model predicts the click rate of each item in the item set according to the identity information, and then outputs an item recommendation list according to the click rate of each item, including: According to the identity information, each item in the item set is combined to construct a plurality of user-item combination data; Each of the user-item combination data is converted into a prompt sequence, mapped, and fused, and then the corresponding embedding sequences are obtained; Perform click rate prediction based on each of the embedded sequences to obtain the click rate of each item; Sorting each item in the item set according to the click rate of each item, and generating a sorted item list; The first N items are extracted from the ranked item list, and the item recommendation list is constructed and output, where N is a preset value.
8. A cross-modal collaborative recommendation system based on a large language model, characterized in that: Includes acquisition module, recommendation module and model training module; The acquisition module is used to acquire the user's identity information and the set of items to be recommended; The recommendation module is used to input the identity information and the item set into a preset recommendation model, so that the recommendation model predicts the click rate of each item in the item set according to the identity information, and then outputs an item recommendation list according to the click rate of each item; The model training module is used to construct the recommendation model based on the large language model and use the training data set for training. The training data set is based on the large language model and collaborative filtering technology, and is constructed by performing prompt sequence conversion, feature mapping and feature fusion on a number of recommendation domain text data and a number of user-item interaction data.
9. A cross-modal collaborative recommendation system based on a large language model as claimed in claim 8, characterized in that: Based on the large language model and collaborative filtering technology, a plurality of recommendation domain text data and a plurality of user-item interaction data are subjected to prompt sequence conversion, feature mapping and feature fusion to construct the training data set, including: Acquire a number of recommendation domain text data and a number of user-item interaction data, wherein the user-item interaction data includes user features, item features, and real interaction data between users and items; According to a preset prompt template, each of the recommendation field text data and each of the user-item interaction data is converted into a plurality of prompt sequences through natural language encoding and collaborative signal encoding, wherein the prompt sequence includes natural language segmentation, user placeholders and item placeholders; According to the type of each word segment in each prompt sequence, based on the large language model and collaborative filtering technology, each word segment is mapped to a corresponding potential embedding feature, wherein the potential embedding feature is a collaborative filtering embedded user feature, a large model embedded language feature, or a mixed embedded project feature; Performing dimension alignment and feature fusion on each of the potential embedding features to obtain an embedding feature corresponding to each of the word segments; According to the order of each word segment in each prompt sequence, the embedding features corresponding to each word segment are combined to obtain each embedding sequence corresponding to each prompt sequence; The training data set is constructed and obtained according to the respective embedding sequences.
10. A cross-modal collaborative recommendation system based on a large language model as claimed in claim 8, characterized in that: The recommendation model is constructed based on a large language model and trained using a training data set, including: In combination with the large language model, the preset alignment network model, the preset GATE network model, the pre-trained collaborative filtering model and the LoRA technology, an initial recommendation model is constructed, wherein the initial recommendation model includes at least three parameters to be trained, namely, a LoRA weight set, an alignment network model training parameter and a GATE network model training parameter; Inputting the training data set into the initial recommendation model, and updating the LoRA weight set according to the loss function to obtain a first recommendation model; The training data set is reconstructed using the alignment network model and the GATE network model to obtain an updated training data set; the updated training data set is input into the first recommendation model, and the alignment network model training parameters and the GATE network model training parameters are updated according to the loss function to obtain the recommendation model.
Citation Information
Cited By
Recommendation method, system and equipment based on large language model and medium
CN120508712A
Recommendation method, system, device and medium based on large language model
CN120508712B
Multi-mode controllable large model identity tag implantation and transmission method and device based on low-rank fusion, and electronic equipment
CN121278693A