Recommendation method and device combining language model and collaborative architecture, medium and equipment

By combining the Bert4Rec architecture, mT5 model and comparison learning technology, the problem of accurate recommendations for users' problems in the existing technology is solved, and a more accurate and personalized recommendation effect is achieved.

CN120179909APending Publication Date: 2025-06-20GUANGDONG ELECTRIC POWER SCI RES INST ENERGY TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510338938.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art cannot effectively combine language models and recommendation systems, resulting in the inability to provide accurate recommendations to user problems.

Method used

A hybrid recommendation model that combines Bert4Rec architecture, mT5 model and comparison learning technology is adopted to obtain user's interactive history information and natural language description of the project, and to use language understanding and collaborative filtering capabilities for in-depth analysis, and optimize prediction accuracy through mask evaluation mechanism.

Benefits of technology

It significantly improves the performance of the recommendation system, provides more accurate and personalized recommendation results, and solves the problem that the existing technology cannot provide accurate recommendations to users' problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179909A_ABST
    Figure CN120179909A_ABST
Patent Text Reader

Abstract

The invention discloses a recommendation method and device combining a language model and a collaborative architecture, a medium and equipment. According to the method, information input by a user is obtained, and the information comprises interaction historical information of the user and natural language description of an item; inputting the information into a trained mixed recommendation model, so that the mixed recommendation model outputs recommendation of the information based on a mask evaluation mechanism; wherein the mixed recommendation model is obtained by training a preset mask language model according to a Bert4Rec architecture, an mT5 model and a contrast learning technology. According to the method combining the multi-modal data and the advanced model architecture, the performance of a recommendation system can be remarkably improved, a more accurate and satisfactory recommendation result is provided for the user, and the problem that accurate recommendation cannot be provided for the problem of the user in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of recommendations that combine language models and collaborative architectures, and particularly to a recommendation method, device, medium, and equipment that combine language models and collaborative architectures. Background Art

[0002] In the digital age, recommendation systems have become a key technology for helping users discover personalized content in a vast amount of information. Although traditional recommendation techniques such as collaborative filtering and content-based recommendation have made progress in some aspects, they still have limitations, especially the problem of data sparsity when dealing with new users or new items, and the difficulty in deeply capturing user preferences and context information.

[0003] With the development of natural language processing technology, especially the emergence of pre-trained language models such as BERT and GPT series, recommendation systems have begun to try to use the semantic understanding ability of these models to improve the relevance and accuracy of recommendations. However, how to effectively combine these language models with recommendation systems, and how to handle the semantic gap between structured and unstructured data, are still challenges in technological development. In addition, although hybrid recommendation techniques improve the accuracy and user satisfaction of recommendations by integrating multiple recommendation strategies, issues such as model complexity, computational efficiency, and real-time performance still need to be solved in practical applications. These problems lead to the inability of existing technologies to provide accurate recommendations for users' questions. Summary of the Invention

[0004] This application provides a recommendation method, device, medium, and equipment that combine language models and collaborative architectures to solve the problem in the prior art that accurate recommendations cannot be provided for users' questions.

[0005] In a first aspect, this application provides a recommendation method that combines a language model and a collaborative architecture, including:

[0006] Obtain the information input by the user, where the information includes the user's interaction history information and the natural language description of the item;

[0007] Input the information into a trained hybrid recommendation model so that the hybrid recommendation model outputs a recommendation for the information based on a masked evaluation mechanism;

[0008] The hybrid recommendation model is obtained by training a preset masked language model according to the Bert4Rec architecture, the mT5 model, and contrastive learning techniques.

[0009] By obtaining the user's interaction history information and the natural language description of the project, this application can comprehensively capture user preferences and project characteristics. These comprehensive information are input into a hybrid recommendation model trained based on the Bert4Rec architecture, the mT5 model, and contrast learning technology. The model can utilize its advanced language understanding and collaborative filtering capabilities to deeply analyze the user's behavior sequence. Combining with the masked evaluation mechanism, when the model predicts the masked project, it can further optimize its prediction accuracy, thereby improving the relevance and personalization of the recommendation. This method of combining multi-modal data and advanced model architectures can significantly improve the performance of the recommendation system, providing more accurate and satisfactory recommendation results for users to solve the problem that the prior art cannot provide accurate recommendations for user problems.

[0010] As a preferred embodiment of the first aspect, the hybrid recommendation model is obtained by training a preset masked language model according to the Bert4Rec architecture, the mT5 model, and contrast learning technology, specifically:

[0011] Obtain the natural language description of the project of the information and the identifier of the project of the information;

[0012] According to the Bert4Rec architecture, perform context feature encoding on the identifier of the project to obtain project ID embeddings;

[0013] According to the mT5 model, perform semantic encoding on the natural language description of the project to obtain project context embeddings;

[0014] Train a preset masked language model according to the project ID embeddings, project context embeddings, and contrast learning technology.

[0015] In this preferred embodiment, by combining the Bert4Rec architecture, the mT5 model, and contrast learning technology, the hybrid recommendation model can effectively extract deep semantic information from the natural language description of the project and capture context features from the identifier of the project. First, the Bert4Rec architecture performs context feature encoding on the identifier of the project to generate project ID embeddings, which helps the model understand the relationships between projects and the user's interaction patterns. Then, the mT5 model performs semantic encoding on the natural language description of the project to generate project context embeddings, enabling the model to grasp the specific content of the project and the possible interest points of the user. Finally, through contrast learning technology, the model learns to distinguish the subtle differences between different projects, further optimizing the representations of the project ID embeddings and project context embeddings. This training method makes the preset masked language model more accurate in predicting the masked project, thereby improving the performance of the recommendation system in dealing with complex user requirements and project characteristics.

[0016] As a preferred embodiment of the first aspect, training the preset masked language model according to the item ID embedding, item context embedding, and contrastive learning techniques specifically includes:

[0017] Aligning the item ID embedding and item context information embedding according to the InfoNCE loss function to obtain an initial training set;

[0018] Training the preset masked language model according to the initial training set and the preset Perceiver network.

[0019] In this preferred embodiment, the present application aligns the item ID embedding and item context information embedding by using the InfoNCE loss function, ensuring that these two different modalities of features can be effectively represented in the same latent feature space, thereby improving the consistency and complementarity of the features. This alignment enables the model to more accurately capture the subtle differences and similarities between items, and thus generate a higher-quality initial training set. Subsequently, using this initial training set and the preset Perceiver network to train the masked language model, the model learns to predict and fill in missing item information when part of the information is masked. This is similar to the masked language model task in the BERT model but extended to the item recommendation field. This training method improves the model's in-depth understanding of item features and enhances its prediction ability when facing incomplete information, ultimately resulting in the recommendation system being able to provide more accurate and personalized recommendation results.

[0020] As a preferred embodiment of the first aspect, training the preset masked language model according to the initial training set and the preset Perceiver network specifically includes:

[0021] Inputting the item ID embedding and item context embedding into the pre-trained Perceiver network to enable the Perceiver network to output a mixed coding representation;

[0022] Training the preset masked language model according to the initial training set and the mixed coding representation to obtain a mixed recommendation model.

[0023] In this preferred embodiment, the present application takes the project ID embedding and the project context embedding as inputs, and utilizes the capabilities of the Perceiver network to fuse these two different modalities of information, generating a hybrid encoded representation. This process enables the model to capture the characteristics of the project in different dimensions, including structured identifier information and rich natural language description information. Subsequently, this hybrid encoded representation and the initial training set are used to train the masked language model, enabling the model to learn how to predict missing project information based on the context when part of the information is masked. This training method improves the model's in-depth understanding of project characteristics, enhances its prediction ability in the face of incomplete information, and ultimately enables the recommendation system to provide more accurate and personalized recommendation results.

[0024] As a preferred embodiment of the first aspect, inputting the information into a pre-trained hybrid recommendation model to enable the hybrid recommendation model to output a recommendation for the information based on a masked evaluation mechanism specifically includes:

[0025] Input the information into a pre-trained hybrid recommendation model, so that the hybrid recommendation model outputs a recommendation for the information according to a preset prompt template, and evaluates the accuracy of the recommendation according to a preset evaluator.

[0026] In this preferred embodiment, the present application inputs the user interaction history information and the natural language description of the project into a pre-trained hybrid recommendation model. This model can utilize its preset prompt templates to generate personalized recommendations, and these templates contain the text category information of the project, which helps to guide the model to more accurately capture the user's interest points. Subsequently, a preset evaluator evaluates the recommendation accuracy based on the consistency between the recommendation result output by the model and the actual behavior of the user. This process not only verifies the relevance of the recommendation but also provides feedback to the model, enabling it to continuously learn and adjust to improve the quality of future recommendations.

[0027] In a second aspect, the present application provides a recommendation device that combines a language model and a collaborative architecture. The device includes an acquisition module and an input / output module;

[0028] The acquisition module is used to acquire the information input by the user, where the information includes the user's interaction history information and the natural language description of the project;

[0029] The input / output module is used to input the information into a pre-trained hybrid recommendation model, so that the hybrid recommendation model outputs a recommendation for the information based on a masked evaluation mechanism;

[0030] Among them, the hybrid recommendation model is obtained by training a preset masked language model according to the Bert4Rec architecture, the mT5 model, and contrast learning techniques.

[0031] This device uses two modules to divide the work and coordinate with each other to output more accurate information recommendations. By obtaining the user's interaction history information and the natural language description of the item, this application can comprehensively capture the user's preferences and item characteristics. These comprehensive information are input into a hybrid recommendation model trained based on the Bert4Rec architecture, mT5 model, and contrastive learning technology. The model can utilize its advanced language understanding and collaborative filtering capabilities to deeply analyze the user's behavior sequence. Combining with the masked evaluation mechanism, when the model predicts the masked item, it can further optimize its prediction accuracy, thereby improving the relevance and personalization of the recommendation. This method of combining multi-modal data and advanced model architectures can significantly improve the performance of the recommendation system, provide more accurate and satisfactory recommendation results for users, and solve the problem that the prior art cannot provide accurate recommendations for the user's problems.

[0032] As a preferred embodiment of the second aspect, the hybrid recommendation model is obtained by training a preset masked language model according to the Bert4Rec architecture, mT5 model, and contrastive learning technology, specifically:

[0033] Obtain the natural language description of the item of the information and the identifier of the item of the information;

[0034] According to the Bert4Rec architecture, perform context feature encoding on the identifier of the item to obtain an item ID embedding;

[0035] According to the mT5 model, perform semantic encoding on the natural language description of the item to obtain an item context embedding;

[0036] Train a preset masked language model according to the item ID embedding, item context embedding, and contrastive learning technology.

[0037] In this preferred embodiment, by combining the Bert4Rec architecture, mT5 model, and contrastive learning technology, this application enables the hybrid recommendation model to effectively extract deep semantic information from the natural language description of the item and capture context features from the identifier of the item. First, the Bert4Rec architecture performs context feature encoding on the identifier of the item to generate an item ID embedding, which helps the model understand the relationships between items and the user's interaction patterns. Then, the mT5 model performs semantic encoding on the natural language description of the item to generate an item context embedding, enabling the model to grasp the specific content of the item and the user's possible points of interest. Finally, through contrastive learning technology, the model learns to distinguish the subtle differences between different items, further optimizing the representations of the item ID embedding and item context embedding. This training method makes the preset masked language model more accurate in predicting the masked item, thereby improving the performance of the recommendation system in dealing with complex user requirements and item characteristics.

[0038] As a preferred embodiment of the second aspect, training a preset masked language model according to the item ID embedding, item context embedding, and contrastive learning techniques specifically includes:

[0039] Aligning the item ID embedding and item context information embedding according to the InfoNCE loss function to obtain an initial training set;

[0040] Training the preset masked language model according to the initial training set and a preset Perceiver network.

[0041] In this preferred embodiment, the present application aligns the item ID embedding and item context information embedding by using the InfoNCE loss function, ensuring that these two different modal features can be effectively represented in the same latent feature space, thereby improving the consistency and complementarity of the features. This alignment enables the model to more accurately capture the subtle differences and similarities between items, and thus generate a higher-quality initial training set. Subsequently, the masked language model is trained using this initial training set and a preset Perceiver network. The model learns to predict and fill in missing item information when part of the information is masked. This is similar to the masked language model task in the BERT model but extended to the field of item recommendation. This training method improves the model's in-depth understanding of item features, enhances its prediction ability in the face of incomplete information, and ultimately enables the recommendation system to provide more accurate and personalized recommendation results.

[0042] In a third aspect, the present application provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the recommendation method combining a language model and a collaborative architecture as described above. Its beneficial effects are the same as those of the recommendation method combining a language model and a collaborative architecture provided in the first aspect of the present application.

[0043] In a fourth aspect, the present application provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the recommendation method combining a language model and a collaborative architecture as described in any item of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic flowchart of an embodiment of the recommendation method combining a language model and a collaborative architecture provided by the present application;

[0045] Figure 2Schematic diagram of an embodiment of a recommendation process combining a language model and a collaborative architecture provided by this application;

[0046] Figure 3 Schematic diagram of an embodiment of a recommendation device combining a language model and a collaborative architecture provided by this application. Detailed implementation manners

[0047] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.

[0048] Embodiment 1

[0049] Please refer to Figure 1 , a recommendation method combining a language model and a collaborative architecture provided by an embodiment of this application.

[0050] In this embodiment, the process of the recommendation method combining a language model and a collaborative architecture in this application is described in detail through steps S01 - S02.

[0051] S01: Obtain the information input by the user, where the information includes the user's interaction history information and the natural language description of the item.

[0052] S02: Input the information into the pre-trained hybrid recommendation model so that the hybrid recommendation model outputs a recommendation for the information based on the masked evaluation mechanism;

[0053] Among them, the hybrid recommendation model is obtained by training a preset masked language model according to the Bert4Rec architecture, the mT5 model, and contrast learning techniques.

[0054] As a preferred embodiment of Embodiment 1, the hybrid recommendation model is obtained by training a preset masked language model according to the Bert4Rec architecture, the mT5 model, and contrast learning techniques, specifically:

[0055] Obtain the natural language description of the item of the information and the identifier of the item of the information;

[0056] According to the Bert4Rec architecture, perform context feature encoding on the identifier of the item to obtain an item ID embedding;

[0057] According to the mT5 model, perform semantic encoding on the natural language description of the item to obtain an item context embedding;

[0058] Train a preset masked language model according to the project ID embedding, project context embedding, and contrastive learning techniques.

[0059] More specifically, the method proposed in this application uses the Bert4Rec architecture to enhance the context features of project IDs. The Bert4Rec architecture is a sequence recommendation model based on Transformer, which is specifically designed to process the interaction sequences between users and projects. It draws on the bidirectional encoder structure of the BERT (Bidirectional Encoder Representations from Transformers) model and models the user's behavior sequence through a masked language model (MLM). During training, Bert4Rec randomly masks certain items in the user interaction sequence and then allows the model to predict these masked items, thereby learning the context information in the user behavior sequence and the relationships between items. This method can effectively capture the long-term dependencies and short-term interest changes of user behavior, providing more accurate prediction capabilities for the recommendation system, especially suitable for scenarios dealing with large-scale user behavior data and complex interaction patterns.

[0060] Sequence recommendation models such as Bert4Rec are usually trained using a masked language model (MLM), similar to how BERT processes language tokens. During model training, a random subset of items in the input sequence is replaced with special masked tokens. Therefore, the goal of the model is to predict these masked items based on the user's context. For example, the input item sequence [ID1, ID2, ID3, ID4, ID5] can be randomly masked to become [ID1, <mask>1,ID3, <mask>2, ID5]. Then, the actual tags of the above mask sequence are:

[0061] <mask>1 = ID2, <mask>2 = ID4;

[0062] The input to the initial Transformer layer can be represented as:

[0063]

[0064] where, e i represents the embedding of item ID i and e <mask>< / mask> represents the masked embedding.

[0065] In addition, the Transformer layer will process the masked sequence and feed the output at the corresponding masked positions into the Softmax function to calculate the probabilities on the item vocabulary. Therefore, the loss of the model is the average negative log-likelihood at all masked positions:

[0066]

[0067] where, S m is the set of masked items, is the actual ID of the corresponding masked item, S' is the input sequence after masking, and P(·) represents the probabilities output by Softmax.

[0068] To integrate the context information of items, this application enhances the item embeddings in the masked sequence:

[0069] [c1, e < mask > , c3, e < mask > , c5];

[0070] where, c i is obtained by combining the item embedding e i and the context embedding t i representing the item metadata. Its context embedding t i is encoded using a pre-trained frozen LLM encoder and a Perceiver model to obtain a fixed-length embedding. The final model will output and project an item embedding consistent with the item embedding to c i . Therefore, this application can use the text as auxiliary information to learn the prediction of masked items, thereby improving the performance of the hybrid recommendation task. In an actual scenario, this application uses the mT5 Base model to encode the text information, matches the dimension of the item embedding through the output projection of the Perceiver network, and finally concatenates all the item embeddings.

[0071] The mT5 Base model is a multilingual pre-trained language model based on the Transformer architecture, specifically designed for handling natural language processing tasks. It is the multilingual version of the T5 (Text-to-Text Transfer Transformer) model, capable of processing text data in multiple languages and performing well in various natural language tasks. mT5 learns rich language features and semantic information through unsupervised pre-training on a large-scale text corpus. In a recommendation system, the mT5 Base model can be used to encode the natural language descriptions of items (such as titles, introductions, user reviews, etc.) to generate item context embeddings, thereby helping the model better understand the semantic features of items and the user's points of interest in the items. This ability enables the recommendation system to more accurately capture the personalized needs of users and provide more relevant and personalized recommendation results.

[0072] In addition, the reason why the Perceiver network is adopted in this application to fuse and encode the item context information and text information is that it can encode a variable-length embedding sequence into a fixed-length output embedding sequence, and realize the joint encoding of different modal data types by concatenating different modal data types in the sequence dimension. Therefore, the Perceiver network provides a natural method for this application to connect different data modal information, thereby effectively improving the data encoding efficiency.

[0073] In this preferred embodiment, by combining the Bert4Rec architecture, the mT5 model, and contrastive learning techniques, the hybrid recommendation model can effectively extract deep semantic information from the natural language descriptions of items and capture context features from the item identifiers. First, the Bert4Rec architecture encodes the context features of the item identifiers to generate item ID embeddings, which helps the model understand the relationships between items and the user's interaction patterns. Then, the mT5 model encodes the semantics of the natural language descriptions of items to generate item context embeddings, enabling the model to grasp the specific content of items and the user's possible points of interest. Finally, through contrastive learning techniques, the model learns to distinguish the subtle differences between different items, further optimizing the representations of the item ID embeddings and item context embeddings. This training method makes the preset masked language model more accurate in predicting masked items, thus improving the performance of the recommendation system in handling complex user needs and item features.

[0074] As a preferred embodiment of Embodiment 1, training the preset masked language model according to the item ID embeddings, item context embeddings, and contrastive learning techniques specifically includes:

[0075] Align the item ID embedding and the item context information embedding according to the InfoNCE loss function to obtain an initial training set;

[0076] Train a preset masked language model according to the initial training set and a preset Perceiver network.

[0077] More specifically, in order to align the item ID embedding and the item context information embedding in the recommendation field, this application adopts contrastive learning technology to better learn the model loss. The motivation for this method is that the representation of those "tail" items with low frequencies in the training data is very limited, and the model can rely on their context representation to learn their embeddings. On the contrary, frequently occurring items can obtain collaborative filtering signals from the training of the sequential recommendation model. By aligning the item ID and context embeddings, this application embeds collaborative filtering and context signals into the same latent space, enabling the context to fill the gap when the collaborative signal is insufficient. Specifically, this application uses the InfoNCE loss function to learn their embedding features, and its definition is as follows.

[0078]

[0079] where τ is the temperature parameter, m is the Softmax supplement term added to each positive sample pair, N is the number of samples, and e i represents the embedding of item ID i , and t i represents the embedding of item ID i , and t j represents the embedding of item ID j in the context of natural language.

[0080] In this preferred embodiment, this application aligns the item ID embedding and the item context information embedding by using the InfoNCE loss function, ensuring that these two different modalities of features can be effectively represented in the same latent feature space, thereby improving the consistency and complementarity of the features. This alignment enables the model to more accurately capture the subtle differences and similarities between items, and thus generate a higher-quality initial training set. Subsequently, using this initial training set and a preset Perceiver network to train the masked language model, the model learns to predict and fill in the missing item information when part of the information is masked. This is similar to the masked language model task in the BERT model, but extended to the item recommendation field. This training method improves the model's in-depth understanding of item features and enhances its prediction ability in the face of incomplete information, ultimately resulting in the recommendation system being able to provide more accurate and personalized recommendation results.

[0081] As a preferred embodiment of the first embodiment, inputting the information into the trained hybrid recommendation model so that the hybrid recommendation model outputs a recommendation for the information based on a masked evaluation mechanism specifically includes:

[0082] Input the information into the trained hybrid recommendation model so that the hybrid recommendation model outputs a recommendation for the information according to a preset prompt template and evaluates the accuracy of the recommendation according to a preset evaluator.

[0083] More specifically, the core component of the hybrid recommendation architecture proposed in this application is that it can effectively utilize text content semantics and collaborative filtering information for model training and recommendation. However, how to efficiently achieve the fusion training between the above two modalities is a huge challenge. Therefore, the model cannot process different modality data simultaneously, resulting in the downstream sequence recommendation task being unable to fully utilize the available information in each modality.

[0084] For this reason, this application proposes a novel model masking evaluation paradigm, that is, by designing an "evaluator" to evaluate whether the hybrid recommendation model can appropriately utilize the above two modality data for recommendation. Specifically, in the user's item sequence S, this application adds the product category S in text form to each item i, providing a generalized prompt template for the model to guide the model to recommend the next item. That is to say, each item i in the item sequence becomes i = (ID, T, S), where S is the evaluation string. When training the model and inferring the masked position, this application only masks ID and the item description T, retaining the sequence encoding of S, so that the model can use S as a prompt to predict the target item (that is, i = (<mask, mask>, S)).

[0085] Next, the key to the recommendation task in this application is to adopt a prompt technique to associate the model evaluator with other semantic signals in the user sequence and be able to adapt to and fuse different modality data. It should be noted that the above task form can be applied to many actual scenarios of recommendation systems, such as recommending items of interest to users based on the user's natural language query or the navigation part of the currently browsed web page. Essentially, the evaluation string can be used as a tool to guide the model prediction to integrate into the user's current context.

[0086] Then, this application integrates the evaluation text into the proposed method, and the evaluation text is an input modality data different from the item context text. This application completes the representation of the item embedding features by encoding each modality separately (with different types of embeddings) and then using a Perceiver network to fuse these two text modalities. Given item i k =(ID k , T k , S k ), which includes an item ID k , the natural language description T k and the evaluation sequence S k , this application will calculate the modified item context representation t k and the enhanced item representation c with the evaluation text k , and their calculation processes are defined as follows.

[0087] T k ′ = LLM(T k ) + θ T ;

[0088] S k ′ = LLM(S k ) + θ S ;

[0089] t k = Linear(Perceiver([S k ′; T k ′]));

[0090] c k = t k + e k ;

[0091] Among them, θ T and θ S are item type embeddings used to label the text and evaluation as different types, ";" represents the concatenation operation, "+" represents the vector addition operation, and e k is the embedding feature of the item ID k . During the training or inference of the masked item process, this application will replace T k ′ and e k with masks at the required positions <mask>Embed and allow the model to utilize S′ as a prompt. By treating the evaluation text as an independent input modality, this application hypothesizes that the model can learn the connection between the evaluation text and the item context. Among them, the evaluation text is applied as an optional additional input to the Perceiver network model of the proposed method.

[0092] Subsequently, this application uses the MLM loss method to train the model. By masking items in the user sequence with a probability of 15%, at least one item is ensured to be masked. For the loss of the objective function, this application will use the weighted sum of the MLM loss and the item ID-text embedding feature contrast loss to obtain, and its definition is as follows.

[0093] L total = αL MLM +(1 - α)L c ;

[0094] Finally, it is found through controlling different parameter weights by the hyperparameter α that when α = 0.5, the model training can obtain the best result. In addition, for the contrast loss, this application sets the temperature parameter τ = 0.3 and the margin term m = 0.2 for comparison and verification.

[0095] In this preferred embodiment, this application inputs the user interaction history information and the natural language description of the item into the trained hybrid recommendation model. This model can utilize its preset prompt templates to generate personalized recommendations. These templates contain the text category information of the item, which helps to guide the model to capture the user's interest points more accurately. Subsequently, the preset evaluator evaluates the recommendation accuracy based on the consistency between the recommendation result output by the model and the user's actual behavior. This process not only verifies the relevance of the recommendation but also provides feedback to the model, enabling it to continuously learn and adjust to improve the quality of future recommendations.

[0096] This application proposes a hybrid recommendation technology framework diagram combining a language model and a collaborative architecture, and its overall framework is as Figure 2 As shown in the figure. The framework mainly includes three stages: (1) Hybrid coding architecture: The present invention uses the mT5 large language model to encode text corpora, and at the same time uses the Bert4Rec collaborative model to perform context embedding encoding on project structured information. Then, the Perceiver network is used to combine the above two text embeddings with ID embedding features to provide a richer hybrid coding representation, thereby providing a data coding basis for model evaluation and training, and further improving the recommendation accuracy. (2) Project ID-text feature contrast learning strategy: The present invention uses the contrast learning method to align the representations of the context embedding and ID embedding of the project, and maps them to the same latent feature space for network training of the model. (3) Model evaluation and training: The present invention designs a mask evaluation mechanism to support users to evaluate the model and provide feedback results to adjust the hybrid recommendation architecture. Finally, the optimal hybrid recommendation model is trained through the network and the final prediction results are generated.

[0097] By obtaining the user's interaction history information and the natural language description of the project, this application can comprehensively capture user preferences and project characteristics. Inputting this comprehensive information into the hybrid recommendation model trained based on the Bert4Rec architecture, mT5 model, and contrast learning technology, the model can utilize its advanced language understanding and collaborative filtering capabilities to deeply analyze the user's behavior sequence. Combining with the mask evaluation mechanism, when the model predicts the masked project, it can further optimize its prediction accuracy, thereby improving the relevance and personalization of the recommendation. This method of combining multi-modal data and advanced model architectures can significantly improve the performance of the recommendation system, provide more accurate and satisfactory recommendation results for users, and solve the problem that the prior art cannot provide accurate recommendations for user problems.

[0098] Embodiment 2

[0099] Please refer to Figure 3 , which is a recommendation device combining a language model and a collaborative architecture provided by an embodiment of this application.

[0100] In this embodiment, the recommendation device combining a language model and a collaborative architecture includes an acquisition module 10 and an input / output module 20.

[0101] The acquisition module 10 is used to acquire the information input by the user, where the information includes the user's interaction history information and the natural language description of the project.

[0102] The input / output module 20 is used to input the information into the trained hybrid recommendation model, so that the hybrid recommendation model outputs a recommendation for the information based on the mask evaluation mechanism;

[0103] Among them, the hybrid recommendation model is obtained by training a preset masked language model according to the Bert4Rec architecture, the mT5 model, and contrast learning techniques.

[0104] As a preferred embodiment of the second embodiment, the hybrid recommendation model is obtained by training a preset masked language model according to the Bert4Rec architecture, the mT5 model, and contrast learning techniques. Specifically:

[0105] Obtain the natural language description of the item of the information and the identifier of the item of the information;

[0106] According to the Bert4Rec architecture, perform context feature encoding on the identifier of the item to obtain an item ID embedding;

[0107] According to the mT5 model, perform semantic encoding on the natural language description of the item to obtain an item context embedding;

[0108] Train a preset masked language model according to the item ID embedding, the item context embedding, and contrast learning techniques.

[0109] More specifically, the method proposed in this application uses the Bert4Rec architecture to enhance the context features of item IDs. The Bert4Rec architecture is a sequence recommendation model based on Transformer, which is specifically designed to process the interaction sequence between users and items. It draws on the bidirectional encoder structure of the BERT (Bidirectional Encoder Representations from Transformers) model and models the user's behavior sequence through the masked language model (MLM). During the training process, Bert4Rec randomly masks some items in the user interaction sequence, and then allows the model to predict these masked items, so as to learn the context information and the relationship between items in the user behavior sequence. This method can effectively capture the long-term dependence and short-term interest changes of user behavior, provide more accurate prediction ability for the recommendation system, and is especially suitable for scenarios dealing with large-scale user behavior data and complex interaction patterns.

[0110] Sequence recommendation models such as Bert4Rec are usually trained using the masked language model (MLM), similar to how BERT processes language tokens. During model training, a random subset of items in the input sequence is replaced with special masked tokens. Therefore, the goal of the model is to predict these masked items based on the user's context. For example, the input item sequence [ID1, ID2, ID3, ID4, ID5] can be randomly masked to [ID1, <mask>1,ID3, <mask>2, ID5]. Then, the actual tags of the above mask sequence are:

[0111] <mask>1 = ID2, <mask>2 = ID4;

[0112] The input to the initial Transformer layer can be represented as:

[0113]

[0114] where, e i represents the embedding of item ID i and e <mask>< / mask> represents the masked embedding.

[0115] In addition, the Transformer layer will process the masked sequence and feed the output corresponding to the masked positions into the Softmax function to calculate the probabilities on the item vocabulary. Therefore, the loss of the model is the average negative log-likelihood of all masked positions:

[0116]

[0117] where, S m is the set of masked items, is the actual ID corresponding to the masked item, S' is the input sequence after masking, and P(·) represents the probabilities output by Softmax.

[0118] To integrate the context information of items, this application enhances the item embeddings in the masked sequence:

[0119] [c1, e <mask>< / mask> , c3, e <mask>< / mask> , c5];

[0120] where, c i is obtained by combining the item embedding e i and the context embedding t i representing the item metadata. Its context embedding t i is encoded using a pre-trained frozen LLM encoder and a Perceiver model to obtain a fixed-length embedding. The final model will output an item embedding consistent with the item embedding and project it into the Transformer-encoded item embedding c i . Therefore, this application can use the text as auxiliary information to learn the prediction of masked items, thereby improving the performance of the hybrid recommendation task. In an actual scenario, this application uses the mT5 Base model to encode the text information, matches the dimension of the item embedding through the output projection of the Perceiver network, and finally concatenates all the item embeddings.

[0121] The mT5 Base model is a multilingual pre-trained language model based on the Transformer architecture, specifically designed for processing natural language processing tasks. It is the multilingual version of the T5 (Text-to-Text Transfer Transformer) model, capable of handling text data in multiple languages and performing well in various natural language tasks. mT5 learns rich language features and semantic information through unsupervised pre-training on a large-scale text corpus. In a recommendation system, the mT5 Base model can be used to encode natural language descriptions of items (such as titles, introductions, user reviews, etc.) to generate item context embeddings, thereby helping the model better understand the semantic features of items and the user's points of interest in the items. This ability enables the recommendation system to more accurately capture the personalized needs of users and provide more relevant and personalized recommendation results.

[0122] In addition, the reason why the Perceiver network is adopted in this application to fuse and encode the item context information and text information is that it can encode a variable-length embedding sequence into a fixed-length output embedding sequence, and by concatenating different modal data types in the sequence dimension, joint encoding of different modal data is achieved. Therefore, the Perceiver network provides a natural method for this application to connect different data modal information, thereby effectively improving the data encoding efficiency.

[0123] In this preferred embodiment, by combining the Bert4Rec architecture, the mT5 model, and contrastive learning techniques, the hybrid recommendation model can effectively extract deep semantic information from the natural language descriptions of items and capture context features from the item identifiers. First, the Bert4Rec architecture encodes the context features of the item identifiers to generate item ID embeddings, which helps the model understand the relationships between items and the user's interaction patterns. Then, the mT5 model encodes the semantics of the natural language descriptions of items to generate item context embeddings, enabling the model to grasp the specific content of items and the user's possible points of interest. Finally, through contrastive learning techniques, the model learns to distinguish the subtle differences between different items, further optimizing the representations of the item ID embeddings and item context embeddings. This training method makes the preset masked language model more accurate in predicting masked items, thus improving the performance of the recommendation system when dealing with complex user needs and item features.

[0124] As a preferred embodiment of the second embodiment, training the preset masked language model according to the item ID embeddings, item context embeddings, and contrastive learning techniques is specifically as follows:

[0125] Align the item ID embedding and the item context information embedding according to the InfoNCE loss function to obtain an initial training set;

[0126] Train a preset masked language model according to the initial training set and a preset Perceiver network.

[0127] More specifically, in order to align the item ID embedding and the item context information embedding in the recommendation field, this application adopts contrastive learning technology to better learn the model loss. The motivation for this method is that the representation of those "tail" items with low frequencies in the training data is very limited, and the model can rely on their context representation to learn their embeddings. On the contrary, frequently occurring items can obtain collaborative filtering signals from the training of the sequential recommendation model. By aligning the item ID and context embeddings, this application embeds collaborative filtering and context signals into the same latent space, enabling the context to fill the gap when the collaborative signal is insufficient. Specifically, this application uses the InfoNCE loss function to learn their embedding features, and its definition is as follows.

[0128]

[0129] where τ is the temperature parameter, m is the Softmax supplement term added to each positive sample pair, N is the number of samples, and e i represents the embedding of item ID i , and t i represents the embedding of item ID i , and t j represents the embedding of item ID j .

[0130] In this preferred embodiment, this application aligns the item ID embedding and the item context information embedding by using the InfoNCE loss function, ensuring that these two different modalities of features can be effectively represented in the same latent feature space, thereby improving the consistency and complementarity of the features. This alignment enables the model to more accurately capture the subtle differences and similarities between items, and then generate a higher-quality initial training set. Subsequently, use this initial training set and a preset Perceiver network to train the masked language model. The model learns to predict and fill in the missing item information when part of the information is masked, which is similar to the masked language model task in the BERT model, but extended to the item recommendation field. This training method improves the model's in-depth understanding of item features and enhances its prediction ability in the face of incomplete information, ultimately resulting in the recommendation system being able to provide more accurate and personalized recommendation results.

[0131] As a preferred embodiment of the second embodiment, inputting the information into the trained hybrid recommendation model so that the hybrid recommendation model outputs a recommendation for the information based on a masked evaluation mechanism specifically includes:

[0132] Input the information into the trained hybrid recommendation model so that the hybrid recommendation model outputs a recommendation for the information according to a preset prompt template and evaluates the accuracy of the recommendation according to a preset evaluator.

[0133] More specifically, the core component of the hybrid recommendation architecture proposed in this application is that it can effectively utilize text content semantics and collaborative filtering information for model training and recommendation. However, how to efficiently achieve the fusion training between the above two modalities is a huge challenge. Therefore, the model cannot process different modality data simultaneously, resulting in the downstream sequence recommendation task being unable to fully utilize the available information in each modality.

[0134] To this end, this application proposes a novel model masking evaluation paradigm, that is, by designing an "evaluator" to evaluate whether the hybrid recommendation model can appropriately utilize the above two modality data for recommendation. Specifically, in the user's item sequence S, this application adds the product category S in text form to each item i, providing the model with a generalized prompt template to guide the model to recommend the next item. That is to say, each item i in the item sequence becomes i = (ID, T, S), where S is the evaluation string. When training the model and inferring the masked position, this application only masks ID and the item description T, retaining the sequence encoding of S, so that the model can use S as a prompt to predict the target item (i.e., i = (<mask, mask>, S)).

[0135] Next, the key to the recommendation task in this application is to adopt a prompting technique to associate the model evaluator with other semantic signals in the user sequence and be able to adapt to and fuse different modality data. It should be noted that the above task form can be applied to many actual scenarios of recommendation systems, such as recommending items of interest to users based on the user's natural language query or the navigation part of the currently browsed web page. Essentially, the evaluation string can be used as a tool to guide the model prediction to integrate into the user's current context.

[0136] Then, this application incorporates the evaluation text into the proposed method, and the evaluation text is an input modality data different from the item context text. This application completes the representation of the item embedding features by encoding each modality separately (with different types of embeddings) and then using the Perceiver network to fuse these two text modalities. Given item i k =(ID k ,T k ,S k ), which includes an item ID k , the natural language description T k and the evaluation sequence S k , this application will calculate the modified item context representation t k and the enhanced item representation c with the evaluation text k , and their calculation processes are defined as follows.

[0137] T k ′ = LLM(T k ) + θ T ;

[0138] S k ′ = LLM(S k ) + θ S ;

[0139] t k = Linear(Perceiver([S′ k ; T′ k ));

[0140] c k = t k + e k ;

[0141] where, θ T and θ S are item type embeddings used to label the text and evaluation as different types, ";" represents the concatenation operation, "+" represents the vector addition operation, and e k is the embedding feature of the item ID k . During the training or inference of the masked item process, this application will replace T k ′ and e k with the masks at the required positions <mask>Embed and allow the model to utilize S′ as a prompt. By treating the evaluation text as an independent input modality, this application assumes that the model can learn the connection between the evaluation text and the item context. Among them, the evaluation text is applied as an optional additional input to the Perceiver network model of the proposed method.

[0142] Subsequently, this application uses the MLM loss method to train the model. By masking items in the user sequence with a probability of 15%, at least one item is ensured to be masked. For the loss of the objective function, this application will use the weighted sum of the MLM loss and the item ID-text embedding feature contrast loss to obtain, and its definition is as follows.

[0143] L total = αL MLM +(1 - α)L c ;

[0144] Finally, it is found by controlling different parameter weights through the hyperparameter α that when α = 0.5, the model training can obtain the best results. In addition, for the contrast loss, this application sets the temperature parameter τ = 0.3 and the margin term m = 0.2 for comparison and verification.

[0145] In this preferred embodiment, this application inputs the user interaction history information and the natural language description of the item into the trained hybrid recommendation model. The model can utilize its preset prompt templates to generate personalized recommendations. These templates contain the text category information of the item, which helps to guide the model to capture the user's interest points more accurately. Subsequently, the preset evaluator evaluates the recommendation accuracy based on the consistency between the recommendation results output by the model and the user's actual behavior. This process not only verifies the relevance of the recommendation but also provides feedback to the model, enabling it to continuously learn and adjust to improve the quality of future recommendations.

[0146] This application proposes a hybrid recommendation technology framework diagram combining a language model and a collaborative architecture, and its overall framework is as Figure 2 As shown in the figure. The framework mainly includes three stages: (1) Hybrid coding architecture: The present invention uses the mT5 large language model to encode text corpora, and at the same time uses the Bert4Rec collaborative model to perform context embedding encoding on project structured information. Then, the Perceiver network is used to combine the above two text embeddings with ID embedding features to provide a richer hybrid coding representation, thereby providing a data coding basis for model evaluation and training, and further improving the recommendation accuracy. (2) Project ID-text feature contrast learning strategy: The present invention uses the contrast learning method to align the representations of the context embedding and ID embedding of the project, and maps them to the same latent feature space for network training of the model. (3) Model evaluation and training: The present invention designs a mask evaluation mechanism to support users in evaluating the model and providing feedback results to adjust the hybrid recommendation architecture. Finally, the optimal hybrid recommendation model is trained through the network and the final prediction results are generated.

[0147] This device uses two modules to divide labor and work in coordination to more accurately output information recommendations. By obtaining the user's interaction history information and the natural language description of the project, this application can comprehensively capture user preferences and project characteristics. Inputting this comprehensive information into the hybrid recommendation model trained based on the Bert4Rec architecture, mT5 model, and contrast learning technology, the model can use its advanced language understanding and collaborative filtering capabilities to deeply analyze the user's behavior sequence. Combined with the mask evaluation mechanism, when the model predicts the masked project, it can further optimize its prediction accuracy, thereby improving the relevance and personalization of the recommendation. This method of combining multi-modal data and advanced model architectures can significantly improve the performance of the recommendation system, provide more accurate and satisfactory recommendation results for users, and solve the problem that the prior art cannot provide accurate recommendations for user problems.

[0148] Embodiment 3:

[0149] The embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium includes a stored computer program, and wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the recommendation method combining a language model and a collaborative architecture as described above;

[0150] Among them, for the recommended method combining a language model and a collaborative architecture, when implemented in the form of software functional units and used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0151] Embodiment 4

[0152] This application provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements any one of the recommended methods combining a language model and a collaborative architecture as described in Embodiment 1.

[0153] For the above specific embodiments, the purpose, technical solution, and beneficial effects of this application have been further described in detail. It should be understood that the above are only specific embodiments of this application and are not used to limit the protection scope of this application. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application should be included in the protection scope of this application.< / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask>

Claims

1. A recommendation method combining a language model and a collaborative architecture, characterized in that: include: Acquire information input by a user, wherein the information includes the user's interaction history information and a natural language description of the project; Inputting the information into a trained hybrid recommendation model so that the hybrid recommendation model outputs a recommendation of the information based on a mask evaluation mechanism; The hybrid recommendation model is obtained by training a preset mask language model based on the Bert4Rec architecture, the mT5 model and the contrastive learning technology.

2. The recommendation method combining language model and collaborative architecture according to claim 1, characterized in that: The hybrid recommendation model is obtained by training the preset mask language model according to the Bert4Rec architecture, the mT5 model and the contrastive learning technology, specifically: obtaining a natural language description of the item of information and an identifier of the item of information; According to the Bert4Rec architecture, context feature encoding is performed on the identifier of the project to obtain a project ID embedding; According to the mT5 model, the natural language description of the project is semantically encoded to obtain the project context embedding; The preset mask language model is trained according to the item ID embedding, item context embedding and contrastive learning techniques.

3. The recommendation method combining language model and collaborative architecture according to claim 2, characterized in that: The preset mask language model is trained according to the project ID embedding, project context embedding and contrastive learning technology, specifically: According to the InfoNCE loss function, the project ID embedding and the project context information embedding are aligned to obtain an initial training set; The preset mask language model is trained based on the initial training set and the preset Perceiver network.

4. The recommendation method combining language model and collaborative architecture according to claim 3, characterized in that: The preset mask language model is trained according to the initial training set and the preset Perceiver network, specifically: Inputting the item ID embedding and the item context embedding into a trained Perceiver network so that the Perceiver network outputs a hybrid coding representation; The preset mask language model is trained according to the initial training set and the hybrid coding representation to obtain a hybrid recommendation model.

5. The recommendation method combining language model and collaborative architecture according to claim 1, characterized in that: The step of inputting the information into a trained hybrid recommendation model so that the hybrid recommendation model outputs a recommendation of the information based on a mask evaluation mechanism is specifically as follows: The information is input into a trained hybrid recommendation model, so that the hybrid recommendation model outputs a recommendation of the information according to a preset prompt template, and the accuracy of the recommendation is evaluated according to a preset evaluator.

6. A recommendation device combining a language model and a collaborative architecture, characterized in that: Including acquisition module and input and output module; The acquisition module is used to acquire information input by the user, wherein the information includes the user's interaction history information and the natural language description of the project; The input-output module is used to input the information into the trained hybrid recommendation model so that the hybrid recommendation model outputs a recommendation of the information based on the mask evaluation mechanism; The hybrid recommendation model is obtained by training a preset mask language model based on the Bert4Rec architecture, the mT5 model and the contrastive learning technology.

7. The recommendation device combining language model and collaborative architecture according to claim 6, characterized in that: The hybrid recommendation model is obtained by training the preset mask language model according to the Bert4Rec architecture, the mT5 model and the contrastive learning technology, specifically: obtaining a natural language description of the item of information and an identifier of the item of information; According to the Bert4Rec architecture, context feature encoding is performed on the identifier of the project to obtain a project ID embedding; According to the mT5 model, the natural language description of the project is semantically encoded to obtain the project context embedding; The preset mask language model is trained according to the item ID embedding, item context embedding and contrastive learning techniques.

8. The recommendation device combining language model and collaborative architecture according to claim 7, characterized in that: The preset mask language model is trained according to the project ID embedding, project context embedding and contrastive learning technology, specifically: According to the InfoNCE loss function, the project ID embedding and the project context information embedding are aligned to obtain an initial training set; The preset mask language model is trained based on the initial training set and the preset Perceiver network.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the recommendation method combining a language model and a collaborative architecture as described in any one of claims 1 to 5.

10. A terminal device, characterized in that: The invention comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the recommendation method combining a language model and a collaborative architecture as described in any one of claims 1 to 5 is implemented.

Citation Information

Cited By

  • Recommendation method, system and equipment based on large language model and medium

    CN120508712A