Recommendation method and device based on large language model
By combining the hybrid semantic ID model with the generative recommendation large language model, and utilizing contrastive learning and reinforcement learning training methods, the problem that the recommendation system is difficult to adapt to changes in user interests in a dynamic environment is solved, achieving higher recommendation accuracy and stability.
Patent Information
- Application Number
- CN202510798226.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Existing recommendation solutions based on large models find it difficult to quickly adapt to changes in user interests and behaviors in dynamic environments, resulting in low accuracy and stability of recommendation results.
A training method combining a hybrid semantic ID model and a generative recommendation large language model is adopted. The large language model is trained through contrastive learning and reinforcement learning to improve the semantic association between user information and item descriptions and enhance the reliability of recommendations.
It improves the consistency, robustness and accuracy of recommendation results, and enhances the accuracy and interpretability of recommendation systems in sequence prediction.
Smart Images

Figure CN120671837A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of recommendation systems, and in particular to a recommendation method and device based on a large language model. Background Art
[0002] Currently, generative recommendation solutions based on large models have become an important research direction in the field of recommendation systems. They aim to combine the natural language understanding and generation capabilities of large models to improve the personalization, explainability, and generalization of recommendations. The more commonly used technical solutions are as follows:
[0003] (1) Generative recommendation technology based on sequential recommendation
[0004] Sequential recommendation technology is a recommendation paradigm that dynamically models user behavior sequences through generative models. Its core concept is to use a user's historical interaction sequences (such as clicks, browsing, and purchases) to directly predict items that the user is likely to interact with in the future, rather than traditional candidate set screening based on matching or ranking. This approach overcomes the limitations of traditional sequential models in modeling long-term dependencies and enables the generation of diverse recommendation results.
[0005] (2) Generative recommendation technology based on semantic ID
[0006] Generative recommendation technology based on semantic IDs is a cutting-edge paradigm that reconstructs item representations through structured semantic identifiers (SemanticIDs) and dynamically generates recommendations using generative models. Its core concept is to encode item semantic information (such as categories, attributes, and content features) into hierarchical, interpretable discrete identifiers (SemanticIDs), replacing traditional continuous embedding vectors or single ID representations. This explicitly models the semantic relationships between items, supports fine-grained reasoning about user interests, and generalizes across domains using generative models. It also captures subtle semantic connections between products.
[0007] As new items continue to emerge, users' interests and behaviors are also changing dynamically. Discrete semantic IDs are not flexible enough in dynamic environments and are difficult to adapt to these changes quickly, which affects the recommendation effect and stability, and the accuracy and practicality of the recommendation results are low. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide a recommendation method and device based on a large language model. From the perspective of practical problems and applications, a hybrid semantic ID based on contrastive learning and a large model training method based on GRPO are adopted to effectively improve the semantic association between user information and item descriptions, as well as the reliability of recommendations generated by the large model.
[0009] In order to solve the above technical problems, the first aspect of the embodiments of the present invention discloses a recommendation method based on a large language model, characterized in that the method includes:
[0010] S1, obtain training data;
[0011] S2, using the training data to train the large language model to obtain an optimized hybrid semantic ID model and a generative recommendation large language model;
[0012] S3, obtain user demand data;
[0013] S4: Process the user demand data using the optimized hybrid semantic ID model and the generative recommendation large language model to obtain recommendation sequence data.
[0014] As an optional implementation, in the first aspect of the embodiment of the present invention, the training process of the large language model using the training data to obtain the optimized hybrid semantic ID model and the generative recommendation large language model includes:
[0015] S21, encoding the training data to obtain first user latent vector data, first item latent vector data, and first semantic ID latent vector data;
[0016] S22, performing vector concatenation processing on the first user latent vector data and the first item latent vector data to obtain first user mixed vector data;
[0017] S23, performing residual quantization processing on the first user mixed vector data to obtain first mixed semantic ID data;
[0018] S24: Use the first training mixed semantic ID data and the first semantic ID latent vector data to train the large model to obtain an optimized mixed semantic ID model and a generative recommendation large language model.
[0019] As an optional implementation manner, in the first aspect of the embodiment of the present invention, the performing residual quantization processing on the first user mixed vector data to obtain first mixed semantic ID data includes:
[0020] S231, preset the total number of residual processing stages N and the residual processing stage i=0;
[0021] S232: When the number of residual processing stages is less than the total number of residual processing stages, performing residual quantization processing on the first user mixed vector data to obtain quantization result data; otherwise, executing S234;
[0022] S233, increasing the residual processing level by 1, and executing S232;
[0023] S234 , encoding and reconstructing the quantization result data to obtain first mixed semantic ID data.
[0024] As an optional implementation manner, in the first aspect of the embodiment of the present invention, performing step-by-step residual quantization processing on the first user mixed vector data to obtain quantization result data includes:
[0025] S2321, preset each level of codebook is {B1, B2, ..., B N}; Initialize the original residual l0 = x0;
[0026] S2322, using the level number i of the quantization processing and the codebook R of the corresponding level, i , performing residual calculation processing on the first user mixed vector data to obtain vector residual result data;
[0027] The residual calculation processing expression is:
[0028] l i =l i-1 -r i ;
[0029]
[0030] Among them, l i represents the residual value of the quantization process at level i; x i represents the i-th mixed vector data of the first user; r i Indicates the difference between the current codebook and the residual l i-1 The closest vector; r represents the codebook B i Contains learnable vectors;
[0031] S2323, performing quantization processing on the vector residual result data to obtain quantized result data;
[0032] The quantization processing expression is:
[0033]
[0034] Among them, L i Represents the quantization result of level i.
[0035] As an optional implementation, in the first aspect of the embodiment of the present invention, the method of training a large model using the first training mixed semantic ID data and the first semantic ID latent vector data to obtain an optimized mixed semantic ID model and a generative recommendation large language model includes:
[0036] S241, based on a contrast loss function, using the first training mixed semantic ID data and the first semantic ID latent vector data to perform contrast training on the mixed semantic ID model to obtain an optimized mixed semantic ID model;
[0037] The contrast loss function expression is:
[0038]
[0039] Among them, L con represents the specific loss; Represents the user's semantic loss; represents the item semantic loss, where U represents the user, I represents the item, and M represents the large language model;
[0040] The contrastive learning expression is:
[0041]
[0042] Where N is the number of positive samples, 2N is the number of samples after enhancement; sim(·,·) is the cosine similarity; z i represents the potential vector data of the first semantic ID i; j represents the jth potential vector data of the first semantic ID; τ represents the temperature parameter used to control the shape of the loss function; k represents the sequence number of the sample; z k represents the kth potential vector data of the first semantic ID;
[0043] S242, parsing the first training mixed semantic ID data to obtain first mixed SID data and first item sequence label mixed SID data;
[0044] S243 : Using the first mixed SID data and the first item sequence label mixed SID data, a large language model is trained to obtain a generative recommendation large language model.
[0045] As an optional implementation, in the first aspect of the embodiment of the present invention, the using the optimized hybrid semantic ID model and the generative recommendation large language model to process the user demand data to obtain recommendation sequence data includes:
[0046] S41, using the optimized hybrid semantic ID model, processing the user demand data to obtain second hybrid semantic ID data;
[0047] S42: Using the generative recommendation large language model, the second mixed semantic ID data and the sequence prediction prompt word data are processed to obtain recommendation sequence data.
[0048] As an optional implementation, in the first aspect of the embodiment of the present invention, the using the optimized hybrid semantic ID model to process the user demand data to obtain the second hybrid semantic ID data includes:
[0049] S411, parsing the user demand data to obtain user ID data, user portrait data, project ID data, project title data, and project description data;
[0050] S412, using the optimized hybrid semantic ID model, processing the user ID data and the user portrait data to obtain second user latent vector data;
[0051] S413, using the optimized hybrid semantic ID model, processing the project ID data, the project title data, and the project description data to obtain second project latent vector data;
[0052] S414: Perform vector concatenation processing on the second user latent vector data and the second item latent vector data to obtain second mixed semantic ID data.
[0053] A second aspect of an embodiment of the present invention discloses a recommendation device based on a large language model, the device comprising: a first acquisition module, a model training module, a second acquisition model and a data processing module;
[0054] The first acquisition module is used to acquire training data;
[0055] The model training module is used to train the large language model using the training data to obtain an optimized hybrid semantic ID model and a generative recommendation large language model;
[0056] The second acquisition module is used to obtain user demand data;
[0057] The data processing module is used to process the user demand data using the optimized encoder model and the generative recommendation large language model to obtain recommendation sequence data.
[0058] A third aspect of the present invention discloses another recommendation device based on a large language model, the device comprising:
[0059] a memory storing executable program code;
[0060] a processor coupled to the memory;
[0061] The processor calls the executable program code stored in the memory to execute part or all of the steps in the large language model-based recommendation method disclosed in the first aspect of the embodiment of the present invention.
[0062] The fourth aspect of the present invention discloses a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer instructions, and when the computer instructions are called, they execute some or all of the steps in the large language model-based recommendation method disclosed in the first aspect of the embodiment of the present invention.
[0063] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0064] In an embodiment of the present invention, a hybrid semantic ID model is adopted to jointly encode user behavior information and project information, so that even when user behavior fluctuates frequently, the generated semantic ID remains stable, thereby ensuring the consistency and robustness of the recommendation results; contrastive learning is used to train the hybrid semantic ID model, which improves the correlation between user and project information and the semantic understanding ability of the encoder, thereby improving the accuracy of the recommendation generation results; reinforcement learning methods are used to train large language models, which improves the accuracy and interpretability of the recommendation system in sequence prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0066] Figure 1 Schematic diagram of a scenario of a recommendation device based on a large language model provided by an embodiment of the present invention;
[0067] Figure 2 This is a flowchart of a recommendation method based on a large language model disclosed in an embodiment of the present invention;
[0068] Figure 3 1 is a schematic diagram of the structure of a recommendation system based on a large language model disclosed in an embodiment of the present invention;
[0069] Figure 4 It is a structural diagram of another recommendation device based on a large language model disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0070] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0071] The terms "first," "second," and so on, in the description and claims of the present invention and the accompanying drawings are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or device.
[0072] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0073] In this application, the word "exemplary" is used to mean "serving as an example, illustration, or illustration." Any embodiment described in this application as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. The following description is given to enable any person skilled in the art to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that one of ordinary skill in the art can recognize that the present application can be implemented without using these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present application with unnecessary details. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in this application.
[0074] It should be noted that since the method of the embodiment of the present application is executed in a computer device, the processing objects of each computer device exist in the form of data or information. For example, time is actually time information. It can be understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, the corresponding data exist for the computer device to process. The details will not be repeated here.
[0075] It should be noted that the large model involved in this application is explained, and the large model refers to an artificial neural network model with a very large number of parameters. In the field of artificial intelligence, a large model generally refers to a model with hundreds of millions to trillions of parameters. The model usually needs to be trained on a large-scale data set and requires a large amount of computing resources to be optimized and adjusted. Large models are generally used to solve complex tasks such as natural language processing, computer vision, and speech recognition. Generative AI is an AI that can create new content and ideas, including conversations, stories, images, videos, and music. In the embodiment of the present application, the large model can be ChatGPT, BERT, XLNet, Zhipu model, Claude, Moonshot AI model, ChatGLM model, Qianyi Tongwen model, MiniMax model, Spark model, Llama model, 360GPT model, Qwen model, Baichuan model, Skylark model, vivoLM model and Wenxin Yiyan and other large-scale language models, which are not limited in the embodiment of the present application.
[0076] The embodiments of the present application provide a recommendation method, apparatus, computer device, and computer-readable storage medium based on a large language model, which are described in detail below.
[0077] See also Figure 1 , Figure 1 This is a scenario diagram of a decision management system provided by an embodiment of the present application. The decision management system based on a large language model may include a computer device 100, in which a recommendation device based on a large language model is integrated, such as Figure 1 Computer equipment in.
[0078] In the embodiments of the present application, the computer device 100 may be an independent server, or a server network or server cluster composed of servers. For example, the computer device 100 described in the embodiments of the present application includes, but is not limited to, a computer, a network host, a single network server, a set of multiple network servers, or a cloud server composed of multiple servers. A cloud server is composed of a large number of computers or network servers based on cloud computing.
[0079] It is understood that the computer device 100 used in the embodiments of the present application can be a device that includes both receiving and transmitting hardware, that is, a device that has receiving and transmitting hardware capable of performing two-way communication over a two-way communication link. Such a device may include: a cellular or other communication device that has a single-line display, a multi-line display, or a cellular or other communication device without a multi-line display. The specific computer device 100 can be a desktop terminal or a mobile terminal. The computer device 100 can also be a mobile phone, a tablet computer, a laptop computer, etc.
[0080] Those skilled in the art will understand that Figure 1 The application environment shown in the figure is only one application scenario of the present application solution and does not constitute a limitation on the application scenario of the present application solution. Other application environments may also include Figure 1 More or fewer computer devices as shown in Figure 1 Only one computer device is shown in the figure. It can be understood that the system can also include one or more other services, which are not limited here.
[0081] In addition, if Figure 1 As shown, the decision management system may further include a memory 200 for storing historical data, such as user demand data, recommendation process data, and recommendation result data.
[0082] It should be noted that Figure 1 The scenario diagram of the question-answering decision system shown is only an example. The recommendation device and scenario based on the large language model described in the embodiment of the present application are intended to more clearly illustrate the technical solution of the embodiment of the present application, and do not constitute a limitation on the technical solution provided by the embodiment of the present application. Ordinary technicians in this field can know that with the evolution of the question-answering decision system and the emergence of new business scenarios, the technical solution provided by the embodiment of the present application is also applicable to similar technical problems.
[0083] This invention discloses a recommendation method and device based on a large language model. This method utilizes knowledge graph data and real-time updated retrieval-enhanced question-and-answer data within a knowledge graph. By integrating knowledge graph question-and-answer with retrieval-enhanced question-and-answer, the method utilizes auxiliary tools to address the need for specialized computing tools to support question-and-answer in decision-making support scenarios. These are described in detail below.
[0084] Example 1
[0085] See also Figure 2 , Figure 2 This is a flow chart of a recommendation method based on a large language model disclosed in an embodiment of the present invention. Figure 2 The described recommendation method based on a large language model is applied to a question-answering decision system, such as a local server or a cloud server for the question-answering decision system, and the embodiment of the present invention does not limit this. Figure 2 As shown, the recommendation method based on the large language model may include the following operations:
[0086] S1, obtain training data;
[0087] It should be noted that the training data is used to train the large language model; it includes training user ID data, training user portrait data, training project ID data, training project title data, training project description data and training project sequence label data;
[0088] It should be noted that, in this embodiment, the training data includes user behavior information and product information; the user behavior information includes, but is not limited to, user search record data, user click data, user purchase history data, and user profile data; the product information includes, but is not limited to, product ID data, product title data, product detailed description information, and user evaluation information;
[0089] S2, using the training data to train the large language model to obtain an optimized hybrid semantic ID model and a generative recommendation large language model;
[0090] S3, obtain user demand data;
[0091] It should be noted that the user demand data includes user ID data, user portrait data, project ID data, project title data and project description data;
[0092] It should be noted that, in this embodiment,
[0093] S4: Process the user demand data using the optimized hybrid semantic ID model and the generative recommendation large language model to obtain recommendation sequence data.
[0094] It can be seen that the implementation of the large language model-based recommendation method described in the embodiment of the present invention uses contrastive learning to train the hybrid semantic ID model and adopts reinforcement learning to train the large language model, which improves the consistency, robustness and accuracy of the recommendation results, thereby improving the accuracy and interpretability of the recommendation system in sequence prediction.
[0095] Optionally, the training process of the large language model using the training data to obtain an optimized hybrid semantic ID model and a generative recommendation large language model includes:
[0096] S21, encoding the training data to obtain first user latent vector data, first item latent vector data, and first semantic ID latent vector data;
[0097] It should be noted that, in this embodiment, the user search record data, user click data, user purchase history data and user portrait data are encoded to obtain the first user potential vector data;
[0098] Encoding the product ID data, product title data, product detailed description information, and user evaluation information to obtain first item latent vector data;
[0099] S22, performing vector concatenation processing on the first user latent vector data and the first item latent vector data to obtain first user mixed vector data;
[0100] It should be noted that the splicing process refers to splicing data in order;
[0101] S23, performing residual quantization processing on the first user mixed vector data to obtain first mixed semantic ID data;
[0102] S24: Use the first training mixed semantic ID data and the first semantic ID latent vector data to train the large model to obtain an optimized mixed semantic ID model and a generative recommendation large language model.
[0103] It can be seen that by implementing the large language model-based recommendation method described in the embodiment of the present invention, the large language model is trained using the training data to obtain an optimized hybrid semantic ID model and a generative recommendation large language model, which improves the consistency, robustness and accuracy of the recommendation results and lays the foundation for improving the accuracy and interpretability of the recommendation system in sequence prediction.
[0104] Optionally, encoding the training data to obtain first user latent vector data, first item latent vector data, and first semantic ID latent vector data includes:
[0105] S211, parsing the training data to obtain training user ID data, training user portrait data, training item ID data, training item title data, training item description data, and training item sequence label data;
[0106] It should be noted that the parsing process refers to extracting corresponding data according to the field type;
[0107] S212, using a text encoder of a large language model, encoding the training user ID data and the training user portrait data to obtain first user latent vector data;
[0108] S213, using a text encoder of a large language model, encoding the training item ID data, the training item title data, and the training item description data to obtain first item latent vector data;
[0109] S214, using a text encoder of a large language model to encode the training item sequence label data to obtain first semantic ID latent vector data;
[0110] It can be seen that by implementing the recommendation method based on the large language model described in the embodiment of the present invention, the training data is encoded and processed to obtain the first user latent vector data, the first item latent vector data and the first semantic ID latent vector data, which provides data support for the subsequent training of the hybrid semantic ID model and the large language model, improves the consistency, robustness and accuracy of the recommendation results, and lays the foundation for improving the accuracy and interpretability of the recommendation system in sequence prediction.
[0111] Optionally, performing residual quantization processing on the first user mixed vector data to obtain first mixed semantic ID data includes:
[0112] S231, preset the total number of residual processing stages N and the residual processing stage i=0;
[0113] S232: When the number of residual processing stages is less than the total number of residual processing stages, performing residual quantization processing on the first user mixed vector data to obtain quantization result data; otherwise, executing S234;
[0114] S233, increasing the residual processing level by 1, and executing S232;
[0115] S234, encoding and reconstructing the quantization result data to obtain first mixed semantic ID data;
[0116] It can be seen that the recommendation method based on the large language model described in the embodiment of the present invention is implemented, the first user mixed vector data is subjected to residual quantization processing to obtain first mixed semantic ID data, and the mixed semantic ID model is used to jointly encode user behavior information and project information. In this way, even when user behavior frequently fluctuates, the generated semantic ID remains stable, thereby ensuring the consistency and robustness of the recommendation results, and laying the foundation for improving the consistency, robustness and accuracy of the recommendation results, and enhancing the accuracy and interpretability of the recommendation system in sequence prediction.
[0117] Optionally, performing step-by-step residual quantization processing on the first user mixed vector data to obtain quantization result data includes:
[0118] S2321, preset each level of codebook is {B1, B2, ..., B N}; Initialize the original residual l0 = x0;
[0119] S2322, using the level number i of the quantization processing and the codebook R of the corresponding level, i , performing residual calculation processing on the first user mixed vector data to obtain vector residual result data;
[0120] The residual calculation processing expression is:
[0121] l i =l i-1 -r i ;
[0122]
[0123] Among them, l i represents the residual value of the quantization process at level i; x i represents the i-th mixed vector data of the first user; r i Indicates the difference between the current codebook and the residual l i-1 The closest vector; r represents the codebook B i Contains learnable vectors;
[0124] S2323, performing quantization processing on the vector residual result data to obtain quantized result data;
[0125] The quantization processing expression is:
[0126]
[0127] Among them, L i Represents the quantization result of level i.
[0128] It can be seen that the large language model-based recommendation method described in the embodiment of the present invention is implemented, and the first user mixed vector data is subjected to step-by-step residual quantization processing to obtain quantization result data, which provides data support for the subsequent training of the mixed semantic ID model and the large language model, improves the consistency, robustness and accuracy of the recommendation results, and lays the foundation for improving the accuracy and interpretability of the recommendation system in sequence prediction.
[0129] Optionally, the step of training a large model using the first training mixed semantic ID data and the first semantic ID latent vector data to obtain an optimized mixed semantic ID model and a generative recommendation large language model includes:
[0130] S241, based on a contrast loss function, using the first training mixed semantic ID data and the first semantic ID latent vector data to perform contrast training on the mixed semantic ID model to obtain an optimized mixed semantic ID model;
[0131] The contrast loss function expression is:
[0132]
[0133] Among them, L con represents contrast loss; Represents the user's semantic loss; represents the item semantic loss, where U represents the user, I represents the item, and M represents the large language model;
[0134] The contrastive learning expression is:
[0135]
[0136] Where N is the number of positive samples, 2N is the number of samples after enhancement; sim(·,-) is the cosine similarity; z i represents the potential vector data of the first semantic ID i; j represents the jth potential vector data of the first semantic ID; τ represents the temperature parameter used to control the shape of the loss function; k represents the sequence number of the sample; z k represents the kth potential vector data of the first semantic ID;
[0137] S242, parsing the first training mixed semantic ID data to obtain first mixed SID data and first item sequence label mixed SID data;
[0138] S243: Using the first mixed SID data and the first item sequence label mixed SID data, a large language model is trained to obtain a generative recommendation large language model;
[0139] It should be noted that, in this embodiment, the large language model is trained based on the GRPO algorithm.
[0140] It can be seen that by implementing the recommendation method based on the large language model described in the embodiment of the present invention, the first user mixed vector data is subjected to step-by-step residual quantization processing to obtain quantization result data, which provides data support for the subsequent training of the hybrid semantic ID model and the large language model, improves the consistency, robustness and accuracy of the recommendation results, and lays the foundation for improving the accuracy and interpretability of the recommendation system in sequence prediction.
[0141] Optionally, the using the first mixed SID data and the first item sequence label mixed SID data to train a large language model to obtain a generative recommendation large language model includes:
[0142] Based on the reward evaluation model, the GRPO algorithm is used to train the large language model using the first mixed SID data and the first item sequence label mixed SID data to obtain a generative recommendation large language model;
[0143] The reward evaluation model expression is:
[0144] J=αJ o +βJ f +μJ c;
[0145]
[0146] Among them, J represents the reward value; α represents the sequential reward coefficient; J o represents the sequential reward value; β represents the format reward coefficient; J f represents the format reward value; μ represents the accuracy reward coefficient; J c represents the accuracy reward value; q represents the index value of the current recommendation in the recommendation sequence; Q represents the number of elements in the recommendation sequence; U q Indicates the current recommended user sequence value; S q Indicates the currently recommended system sequence value; F p Indicates the pth recommended format meets the value; S p represents the p-th recommendation accuracy value;
[0147] It should be noted that, in this embodiment, α is set to 0.3; β is set to 0.2; μ is set to 0.5;
[0148] It should be noted that, in this embodiment, when the format of the recommended item meets the user's requirements, the format reward value Jf is set to 1, otherwise, it is set to 0;
[0149] It should be noted that, in this embodiment, the accuracy reward value J c Set to 1-5 points;
[0150] It should be noted that the GRPO (Group Relative Policy Optimization) algorithm is a reinforcement learning algorithm that eliminates the reliance on a separate value network based on the improvement of traditional proximal policy optimization. Specifically, in each state, GRPO samples a set of candidate outputs from the old policy, and then normalizes the rewards of this set of outputs (calculates the standard score of each output relative to the average reward in the group) to measure the relative advantage of each output. Next, the algorithm uses a clipping mechanism similar to PPO to update the strategy, that is, while keeping the update amplitude not too large, the model is more inclined to generate outputs with higher relative advantages. In order to prevent the new strategy from deviating too much from the reference strategy, a KL divergence penalty term is added. This method not only reduces the training variance and saves computing resources, but also achieves good results in reinforcement learning tuning of large-scale language models;
[0151] It should be noted that in this embodiment, three rewards are defined for reinforcement learning, namely, sequence reward, format reward and accuracy reward; the sequence reward is to improve the accuracy of the large model in predicting the sequence in the sequence recommendation with a sequence; the format reward is used to constrain the semantic ID generated by the large language model to conform to the required format, such as the output result with a thinking process, the semantic ID format is<a_4> etc.; the accuracy reward is the key to improving the prediction accuracy of the large language model; the large language model trained by the GPR0 reinforcement learning algorithm can have stronger reasoning ability, which to a certain extent alleviates the interpretability of the large language model; in addition, the large language model trained based on the reinforcement learning algorithm can obtain better generalization ability and can obtain better stability and robustness in the recommendation system
[0152] It can be seen that the large language model-based recommendation method described in the embodiment of the present invention is implemented, and the first mixed SID data and the first item sequence label mixed SID data are used to train the large language model to obtain a generative recommendation large language model. This lays the foundation for subsequently using the optimized mixed semantic ID model and the generative large language model to obtain recommendation sequence data, improves the consistency, robustness and accuracy of the recommendation results, and lays the foundation for improving the accuracy and interpretability of the recommendation system in sequence prediction.
[0153] Optionally, the utilizing the optimized hybrid semantic ID model and the generative recommendation large language model to process the user demand data to obtain recommendation sequence data includes:
[0154] S41, using the optimized hybrid semantic ID model, processing the user demand data to obtain second hybrid semantic ID data;
[0155] S42: Using the generative recommendation large language model, the second mixed semantic ID data and the sequence prediction prompt word data are processed to obtain recommendation sequence data.
[0156] It can be seen that the implementation of the large language model-based recommendation method described in the embodiment of the present invention uses the generative recommendation large language model to process the user demand data to obtain recommendation sequence data, thereby improving the consistency, robustness and accuracy of the recommendation results, and laying the foundation for improving the accuracy and interpretability of the recommendation system in sequence prediction.
[0157] Optionally, the using the optimized hybrid semantic ID model to process the user demand data to obtain second hybrid semantic ID data includes:
[0158] S411, parsing the user demand data to obtain user ID data, user portrait data, project ID data, project title data, and project description data;
[0159] S412, using the optimized hybrid semantic ID model, processing the user ID data and the user portrait data to obtain second user latent vector data;
[0160] It should be noted that the processing refers to encoding the user ID data and the user portrait data using the optimized hybrid semantic ID model;
[0161] S413, using the optimized hybrid semantic ID model, processing the project ID data, the project title data, and the project description data to obtain second project latent vector data;
[0162] It should be noted that the processing refers to encoding the project ID data, the project title data and the project description data using the optimized hybrid semantic ID model;
[0163] S414: Perform vector concatenation processing on the second user latent vector data and the second item latent vector data to obtain second mixed semantic ID data.
[0164] It can be seen that the large language model-based recommendation method described in the embodiment of the present invention is implemented to process the user demand data to obtain the second mixed semantic ID data, which improves the consistency, robustness and accuracy of the recommendation results and lays the foundation for improving the accuracy and interpretability of the recommendation system in sequence prediction.
[0165] Optionally, the using the generative recommendation large language model to process the second mixed semantic ID data to obtain recommendation sequence data includes:
[0166] S421, obtaining sequence prediction prompt word data;
[0167] S422, concatenating the second mixed semantic ID data and the sequence prediction prompt word data to obtain mixed SID data;
[0168] S423, using the generative recommendation large language model to process the mixed SID data to obtain recommendation sequence data
[0169] It can be seen that the implementation of the large language model-based recommendation method described in the embodiment of the present invention uses the generative recommendation large language model to process the second mixed semantic ID data to obtain recommendation sequence data, thereby improving the consistency, robustness and accuracy of the recommendation results and laying the foundation for improving the accuracy and interpretability of the recommendation system in sequence prediction.
[0170] Alternatively, a recommendation method based on a large language model can be applied to online video platforms to provide users with personalized movie recommendations, including:
[0171] 1. Obtain user information and movie information
[0172] 2. Encode the user information and the movie information to obtain a mixed semantic ID;
[0173] Encoding the user information using a hybrid semantic ID model to obtain a user latent vector;
[0174] It should be noted that the user behavior data includes but is not limited to user viewing history information, rating information and preference tag information;
[0175] Encoding the movie information using a hybrid semantic ID model to obtain a movie latent vector;
[0176] It should be noted that the movie information includes but is not limited to movie ID information, movie title information, plot summary information and actor information;
[0177] Concatenating the user latent vector and the movie latent vector to obtain a mixed vector;
[0178] The mixed vector is processed using the residual quantization algorithm model of RQ-VAE to obtain a mixed semantic ID;
[0179] This ensures that the semantic ID remains stable even if the user's behavior information fluctuates;
[0180] 3. Get the prompt word information;
[0181] 4. Concatenate the mixed semantic ID and the prompt word information to obtain large model input data;
[0182] The mixed semantic ID is concatenated with the sequence prediction prompt word to form the input of the large language model. This step ensures that the generated recommendation sequence not only takes into account the user's historical behavior, but also incorporates the multi-dimensional information of the movie itself.
[0183] 5. Input the large model input data into the large language model to generate recommendation sequence data;
[0184] A large model trained with the GRPO algorithm is used to generate a sequence of predicted movie recommendations. During the training process, sequence rewards (to ensure the order of recommendations is reasonable), format rewards (to ensure the output format meets expectations, such as including the reasoning for the recommendation), and accuracy rewards (to improve the matching accuracy of the recommendations) are defined. After reinforcement learning tuning, the model can infer and generate movie recommendation lists that are both in line with user interests and have a certain degree of explanatory power.
[0185] It can be seen that the implementation of the large language model-based recommendation method described in the embodiment of the present invention uses contrastive learning to train the hybrid semantic ID model using user information and movie information, and uses reinforcement learning to train the large language model, thereby providing users with personalized movie recommendations. This improves the consistency, robustness, and accuracy of the recommendation results, thereby enhancing the accuracy and interpretability of the recommendation system in sequence prediction.
[0186] Optionally, a large language model-based recommendation method is used on e-commerce platforms to recommend suitable products to users based on their recent browsing and purchasing behavior and the characteristics of the products themselves, including:
[0187] Step 1: Obtain user behavior information and product information;
[0188] Step 2: Process the user behavior information and the product information to obtain a mixed semantic ID;
[0189] Using a hybrid semantic ID model, the user behavior information is processed to obtain user latent vector information;
[0190] The user behavior information includes but is not limited to user search history information, user click information, user purchase history information and user profile information;
[0191] Processing the product information using the hybrid semantic ID model to obtain product latent vector information;
[0192] It should be noted that the product information includes but is not limited to product ID information, product title information, product detailed description information and user evaluation information;
[0193] splicing the user latent vector information and the product latent vector information to obtain mixed vector information;
[0194] The mixed vector is processed using the residual quantization algorithm model of RQ-VAE to obtain mixed semantic ID information;
[0195] It should be noted that this method can effectively alleviate the impact of user behavior changes on the stability of semantic IDs;
[0196] Step 3, obtaining scene prompt word information;
[0197] For example, the scene prompt word information is: recommending a continuous sequence of recent hot-selling and well-received products to you;
[0198] Step 4: concatenate the mixed semantic ID information and the scene prompt word information to obtain large language model input data;
[0199] Step 5: Input the large model input data into the large language model to generate recommendation sequence data;
[0200] A large model is trained using a GRPO-based reinforcement learning method, which incorporates sequential rewards (to account for the consistency of user purchase decisions), format rewards (to ensure standardized output structure, such as labeling product categories and recommendation reasons), and accuracy rewards (to ensure that recommendations match user interests). The resulting product recommendation sequence generated by the model not only reflects the user's current shopping preferences but also captures the inherent connections between products, improving overall recommendation effectiveness.
[0201] Step 6: Using real-time user behavior information and real-time product information, the large language model is fed back and iterated to obtain an updated large language model;
[0202] By performing online feedback iteration on the large language model, an updated large language model is obtained to ensure that the recommended content always matches the user's current needs.
[0203] It can be seen that the large language model-based recommendation method described in the embodiment of the present invention uses contrastive learning to train the hybrid semantic ID model based on user behavior information and product information, and uses reinforcement learning to train the large language model. Based on the user's recent browsing and purchasing behavior and the characteristics of the product itself, suitable products are pushed to the user, thereby improving the consistency, robustness and accuracy of the recommendation results, thereby improving the accuracy and interpretability of the recommendation system in sequence prediction.
[0204] Example 2
[0205] See also Figure 3 , Figure 3 Schematic diagram of a recommendation device based on a large language model disclosed in an embodiment of the present invention. Figure 3 The described device can be applied to a question-answering decision system, such as a local server or a cloud server for a question-answering decision system, and the embodiment of the present invention does not limit this. Figure 3 As shown, the apparatus may include: a first acquisition module 101, a model training module 102, a second acquisition module 103 and a data processing module 104;
[0206] The first acquisition module 101 is used to acquire training data;
[0207] The model training module 102 is used to train the large language model using the training data to obtain an optimized hybrid semantic ID model and a generative recommendation large language model;
[0208] The second acquisition module 103, the user acquires user demand data;
[0209] The data processing module 104 is configured to process the user demand data using the optimized encoder model and the generative recommendation large language model to obtain recommendation sequence data.
[0210] Example 3
[0211] See also Figure 4 , Figure 4 Schematic diagram of a recommendation device based on a large language model disclosed in an embodiment of the present invention. Figure 4 The described device can be applied to a multi-sample simulation control management system, such as a local server or cloud server for a multi-sample simulation control system, and the embodiment of the present invention does not limit this. Figure 4 As shown, the device may include:
[0212] A memory 201 storing executable program code;
[0213] a processor 202 coupled to the memory 201;
[0214] The processor 202 calls the executable program code stored in the memory 201 to execute the steps of the recommendation method based on the large language model described in the first embodiment.
[0215] Example 4
[0216] An embodiment of the present invention discloses a computer-readable storage medium storing a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps of the large language model-based recommendation described in the first embodiment.
[0217] Example 5
[0218] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps of the large language model-based recommendation method described in Example 1.
[0219] The device embodiments described above are merely illustrative. Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0220] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, or of course, by means of hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, the storage medium including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0221] Finally, it should be noted that the recommendation method and device based on a large language model disclosed in the embodiments of the present invention only disclose a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A recommendation method based on a large language model, characterized in that: The method comprises: S1, obtain training data; S2, using the training data to train the large language model to obtain an optimized hybrid semantic ID model and a generative recommendation large language model; S3, obtain user demand data; S4: Process the user demand data using the optimized hybrid semantic ID model and the generative recommendation large language model to obtain recommendation sequence data.
2. The recommendation method based on a large language model according to claim 1, characterized in that The method of training the large language model using the training data to obtain an optimized hybrid semantic ID model and a generative recommendation large language model includes: S21, encoding the training data to obtain first user latent vector data, first item latent vector data, and first semantic ID latent vector data; S22, performing vector concatenation processing on the first user latent vector data and the first item latent vector data to obtain first user mixed vector data; S23, performing residual quantization processing on the first user mixed vector data to obtain first mixed semantic ID data; S24: Use the first training mixed semantic ID data and the first semantic ID latent vector data to train the large model to obtain an optimized mixed semantic ID model and a generative recommendation large language model.
3. The recommendation method based on a large language model according to claim 2, characterized in that: The performing residual quantization processing on the first user mixed vector data to obtain first mixed semantic ID data includes: S231, preset the total number of residual processing stages N and the residual processing stage i=0; S232: When the number of residual processing stages is less than the total number of residual processing stages, performing residual quantization processing on the first user mixed vector data to obtain quantization result data; otherwise, executing S234; S233, increasing the residual processing level by 1, and executing S232; S234 , encoding and reconstructing the quantization result data to obtain first mixed semantic ID data.
4. The recommendation method based on a large language model according to claim 3, characterized in that The step of performing step-by-step residual quantization processing on the first user mixed vector data to obtain quantization result data includes: S2321, preset each level of codebook is {B1, B2, ..., B N }; Initialize the original residual l0 = x0; S2322, using the level number i of the quantization processing and the codebook R of the corresponding level, i , performing residual calculation processing on the first user mixed vector data to obtain vector residual result data; The residual calculation processing expression is: l i =l i-1 -r i ; Among them, l i represents the residual value of the quantization process at level i; x i represents the i-th mixed vector data of the first user; r i Indicates the difference between the current codebook and the residual l i-1 The closest vector; r represents the codebook B i Contains learnable vectors; S2323, performing quantization processing on the vector residual result data to obtain quantized result data; The quantization processing expression is: Among them, L i Represents the quantization result of level i.
5. The recommendation method based on a large language model according to claim 2, characterized in that The method of training the large model using the first training mixed semantic ID data and the first semantic ID latent vector data to obtain an optimized mixed semantic ID model and a generative recommendation large language model includes: S241, based on a contrast loss function, using the first training mixed semantic ID data and the first semantic ID latent vector data to perform contrast training on the mixed semantic ID model to obtain an optimized mixed semantic ID model; The contrast loss function expression is: Among them, L con represents contrast loss; Represents the user's semantic loss; represents the item semantic loss, where U represents the user, I represents the item, and M represents the large language model; The contrastive learning expression is: Where N is the number of positive samples, 2N is the number of samples after enhancement; sim(·,·) is the cosine similarity; z i represents the potential vector data of the first semantic ID i; j represents the jth potential vector data of the first semantic ID; τ represents the temperature parameter used to control the shape of the loss function; k represents the sequence number of the sample; z k represents the kth potential vector data of the first semantic ID; S242, parsing the first training mixed semantic ID data to obtain first mixed SID data and first item sequence label mixed SID data; S243 : Using the first mixed SID data and the first item sequence label mixed SID data, a large language model is trained to obtain a generative recommendation large language model.
6. The recommendation method based on a large language model according to claim 1, characterized in that The method of processing the user demand data using the optimized hybrid semantic ID model and the generative recommendation large language model to obtain recommendation sequence data includes: S41, using the optimized hybrid semantic ID model, processing the user demand data to obtain second hybrid semantic ID data; S42: Using the generative recommendation large language model, the second mixed semantic ID data and the sequence prediction prompt word data are processed to obtain recommendation sequence data.
7. The recommendation method based on a large language model according to claim 1, characterized in that The step of processing the user demand data using the optimized hybrid semantic ID model to obtain second hybrid semantic ID data includes: S411, parsing the user demand data to obtain user ID data, user portrait data, project ID data, project title data, and project description data; S412, using the optimized hybrid semantic ID model, processing the user ID data and the user portrait data to obtain second user latent vector data; S413, using the optimized hybrid semantic ID model, processing the project ID data, the project title data, and the project description data to obtain second project latent vector data; S414: Perform vector concatenation processing on the second user latent vector data and the second item latent vector data to obtain second mixed semantic ID data.
8. A recommendation device based on a large language model, characterized in that: The device includes: a first acquisition module, a model training module, a second acquisition module and a data processing module; The first acquisition module is used to acquire training data; The model training module is used to train the large language model using the training data to obtain a generative recommendation large language model; The second acquisition module is used to obtain user demand data; The data processing module is used to process the user demand data using the generative recommendation large language model to obtain recommendation sequence data.
9. A recommendation device based on a large language model, characterized in that: The device comprises: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the recommendation method based on a large language model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which, when called, are used to execute the large language model-based recommendation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Knowledge graph generation type question answering method and system based on large language model
CN117033608A
Recommendation method and device based on large language model, equipment and storage medium
CN117973545A
Generation of explanations with multistep reasoning for ranking in recommender systems
WO2024177735A1