Self-optimization fine tuning method of big language model recommendation system and recommendation system

Through self-distillation technology and course learning strategies, the training process of large language models is optimized, and the knowledge gap and insufficient fit are solved, and more accurate and personalized recommendations are achieved.

CN120338044APending Publication Date: 2025-07-18ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510535319.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There is a knowledge gap in the recommendation system for large language models, which leads to problems such as difficulty in fine-tuning, low degree of fit and homogeneity of recommended content.

Method used

The auxiliary training data set is generated using self-distillation technology, and the training focus is gradually adjusted through the course learning fine-tuning strategy, transfer from the auxiliary training data set to the real data set, dynamically adjust the training difficulty, and use self-distillation to generate diversified candidate outputs, and combine with the course learning to optimize the model generation strategy.

Benefits of technology

It improves the adaptability and personalized recommendation effect of large language models in the recommendation system, reduces homogeneous content, and improves recommendation accuracy and diversity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338044A_ABST
    Figure CN120338044A_ABST
Patent Text Reader

Abstract

The invention discloses a self-optimization fine tuning method for a big language model recommendation system and the recommendation system, and the method comprises the steps: generating an auxiliary training data set through a self-distillation technology, and enabling the auxiliary training data set to generate a plurality of outputs according to the input through a supervised and fine-tuned big language model, and selecting the output closest to the real item from the items to construct. A course learning fine tuning strategy is adopted, training weights of the simple tasks and the difficult tasks are adaptively adjusted according to the current learning state of the large language model, and training focuses are gradually transferred to a real data set from an auxiliary training data set. Through the self-distillation technology, the model generates data closer to recommended field distribution as an intermediate training target, and the field adaptation difficulty is relieved. And through a course learning strategy, the training data difficulty is dynamically adjusted, so that the model gradually adapts to real data distribution. Diversified candidate outputs are generated through self-distillation, and a strategy is generated in combination with a curriculum learning optimization model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of recommendation systems, specifically to a self-optimizing fine-tuning method for a large language model recommendation system and a recommendation system. Background Art

[0002] A recommendation system is a type of information filtering system that screens out information or products that may be of interest to users from a large amount of information by analyzing users' historical behaviors, preferences, and needs. Recommendation systems have wide applications in many fields, such as e-commerce, news push, social networks, movie and music recommendations, etc. The main goal of a recommendation system is to provide personalized content or services to improve user satisfaction and user experience, and at the same time, it can also help enterprises increase sales or user activity.

[0003] In recent years, with the rapid development of artificial intelligence and machine learning technologies, the research and application of recommendation systems have also made remarkable progress. Among them, the emergence and development of large language models (LLMs) have provided new possibilities for the technological progress of recommendation systems. A large language model is a natural language processing (NLP) model based on deep learning that can understand and generate human language, thus enabling more intelligent and personalized recommendations. Large language models, such as OpenAI's GPT series models, are trained with large-scale text data and can generate coherent and contextually appropriate natural language texts. The emergence of these models has brought natural language processing technology to an unprecedented level in terms of text understanding and generation. For recommendation systems, this means being able to better understand users' needs and generate more personalized recommendation results.

[0004] Large language models have rich open-world knowledge and text reasoning capabilities, enabling them to be applied to recommendation systems. However, large language models pre-trained in the open domain lack knowledge in the recommendation field, and directly applying large language models to recommendation systems often does not produce good results. Existing technologies generally use supervised fine-tuning (SFT) to fine-tune large language models to improve the recommendation ability of large language models while maintaining their original open world.

[0005] Specifically, the training data is used for the SFT training of the large language model, where the input consists of the text titles of the user's historical interaction items and an artificially written instruction prompt, and the output label is the text title of the user's next interaction item. During the SFT training process, the goal of the large language model is to increase the output probability of each token of the output label given the tokenized input and the previous tokens of the output label.

[0006] After the large language model (LLM) has undergone supervised fine-tuning (SFT), the problem of hallucinations still exists, generating non-existent items. Therefore, when applied to recommendations, it still needs to go through a grounding step. Existing work calculates the L2 distance between the output items of the LLM and the token embedding vectors of all items, and selects the item with the smallest distance, that is, the item with the highest semantic similarity to the LLM output, as the final recommended item. However, it still has the following drawbacks: 1. There is a gap between the knowledge of the LLM and the recommended knowledge, and directly using SFT to fine-tune the LLM is a very difficult task; 2. SFT will lead to insufficient fitting ability of the LLM to the training data; 3. SFT will lead to homogeneous recommended content generated by the LLM for different users. Summary of the Invention

[0007] In this embodiment, a self-optimizing fine-tuning method, system, electronic device, and storage medium for a large language model recommendation system are provided to solve the problems of difficult fine-tuning, low fitting degree of the LLM to the training data, and homogeneous recommended content caused by the gap between the LLM and the recommended knowledge in the related art.

[0008] In a first aspect, an embodiment of the present invention provides a self-optimizing fine-tuning method for a large language model recommendation system. The self-optimizing fine-tuning method for the large language model recommendation system includes: Using self-distillation technology to generate an auxiliary training dataset, which is constructed by the supervised fine-tuned large language model generating multiple outputs according to the input and selecting the output closest to the real item therefrom; Adopting a curriculum learning fine-tuning strategy to adaptively adjust the training weights of simple tasks and difficult tasks according to the current learning state of the large language model, and gradually shifting the training focus from the auxiliary training dataset to the real dataset.

[0009] In an optional embodiment, the self-distillation technology includes: Generating M outputs for each input by the supervised fine-tuned large language model, and selecting the output with the smallest distance from the real item as the auxiliary training data by calculating the L2 distance of the token embedding vectors.

[0010] In an optional embodiment, the auxiliary training dataset is used to guide the large language model to gradually adapt to the distribution of the real dataset and serves as an intermediate knowledge base.

[0011] In an optional embodiment, the curriculum learning fine-tuning strategy is implemented through the following loss function: ; wherein, is the supervised fine-tuning loss, is the self-distillation fine-tuning loss, is a weight parameter that is dynamically adjusted during the training process.

[0012] In an optional embodiment, the weight parameter has the following calculation formula: ; where is the average L2 distance between the current model output and the true label, represents the average distance at the start of the large language model training, represents a hyperparameter.

[0013] In an optional embodiment, the average L2 distance is calculated by sampling one-tenth of the training data.

[0014] In an optional embodiment, the method further includes: During the training process, gradually decrease the value of the weight parameter so that the training focus of the large language model gradually shifts from the auxiliary training dataset to the real dataset.

[0015] Compared with the prior art, the beneficial effects of the self-optimizing fine-tuning method of the large language model recommendation system of the present invention are as follows: Through the self-distillation technology, the present invention enables the model to generate data closer to the distribution of the recommendation field by itself as an intermediate training target, alleviating the difficulty of domain adaptation. Through the curriculum learning strategy, dynamically adjust the training data difficulty to enable the model to gradually adapt to the real data distribution. Generate diverse candidate outputs through self-distillation and optimize the model generation strategy in combination with curriculum learning. Train in stages (first learn easy samples and then gradually increase the difficulty) to avoid the model falling into local optima prematurely. This method only requires adjusting the training strategy (without modifying the model structure) and supports parameter-efficient fine-tuning techniques such as LoRA.

[0016] In a second aspect, an embodiment of the present invention provides a large language model recommendation system that is trained using the self-optimizing fine-tuning method of the large language model recommendation system described in the first aspect and is used to generate personalized recommendation content. In a third aspect, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a bus, where the processor, the communication interface, and the memory communicate with each other through the bus, and the processor can call the logical instructions in the memory to execute the steps of the method provided in the first aspect.

[0017] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the self-optimizing fine-tuning method of the large language model recommendation system described in the first aspect.

[0018] Compared with the prior art, the beneficial effects of the large language model recommendation system, electronic device and storage medium of the present invention are the same as those of the self-optimizing fine-tuning method of the large language model recommendation system described in the first aspect, so they will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a flowchart of the self-optimizing fine-tuning method of the large language model recommendation system in the embodiment of the present invention; Figure 2 It is a schematic diagram of the self-optimizing fine-tuning method of the large language model recommendation system in the embodiment of the present invention; Figure 3 It is an experimental table of conducting experiments on three subsets of the Amazon review dataset in the embodiment of the present invention; Figure 4 It is a structural block diagram of the electronic device in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application will be described and illustrated below with reference to the accompanying drawings and embodiments.

[0022] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meanings understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "an", "one kind", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variants thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly connected. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0023] In an embodiment of the present invention, a self-optimizing fine-tuning method for a large language model recommendation system is provided. Figure 1 It is a flowchart of the self-optimizing fine-tuning method for the large language model recommendation system of the present invention, as Figure 1 and Figure 2 shown. This process includes the following steps: S100. Use self-distillation technology to generate an auxiliary training data set. The auxiliary training data set is constructed by a large language model that has been supervised and fine-tuned to generate multiple outputs according to the input and select the output closest to the real item from them; It should be noted that the self-distillation technology includes: Generate M outputs for each input through the large language model that has been supervised and fine-tuned, and select the output with the smallest distance from the real item as the auxiliary training data by calculating the L2 distance of the token embedding vectors.

[0024] The auxiliary training data set is used to guide the large language model to gradually adapt to the distribution of the real data set and serve as an intermediate knowledge base.

[0025] It should be added that, given the difficulty of large language models fitting real recommendation datasets, it would be promising to design an alternative method to generate datasets that are easier to learn. Based on this, a method using self-distillation is proposed, on the grounds that the data generated by large language models themselves is inherently closer to the model's internal patterns, making it easier for large language models to learn from it.

[0026] More specifically, the large language model after SFT is adopted to generate M outputs for each input. Among these M outputs, we select the one closest to the real item by calculating the L2 distance of the token embedding vectors. The large language model receives additional supervision signals through this process, and the output closest to the real item among its generated outputs is marked as preferred, even if these outputs have not fully fitted the real data. This is achieved by identifying the output with the minimum distance, as follows: Subsequently, a new dataset is constructed. Compared with the real dataset, this dataset is closer to the distribution of the large language model, making it easier for the model to learn. Although it is not the most accurate, this dataset can still serve as an intermediate knowledge base to guide the large language model to gradually adapt to and learn more complex knowledge. This is similar to the concept in human education where simple intermediate knowledge serves as a foundation to facilitate the learning and understanding of more complex concepts.

[0027] S200. Adopt a curriculum learning fine-tuning strategy to adaptively adjust the training weights of simple tasks and difficult tasks according to the current learning state of the large language model, and gradually shift the training focus from the auxiliary training dataset to the real dataset.

[0028] Furthermore, the original large language model is fine-tuned using the dataset generated by self-distillation. Different from previous work, this dataset is not chosen to further fine-tune the SFT model because the performance of the SFT model does not meet the preset expectations. Instead, a curriculum learning strategy is adopted, allowing the large language model to gradually shift its attention from simple data to more challenging data during the fine-tuning process. This method mimics the process in human education curriculums.

[0029] For the curriculum fine-tuning (CFT) process, we adopt a straightforward method. We regard the entire real dataset as a "more difficult" sample and the dataset obtained from self-distillation as a "more easy" sample. Based on this, we construct the SOFT loss function. Specifically, the curriculum learning fine-tuning strategy is achieved through the following loss function (SOFT loss function): ; where, is the supervised fine-tuning loss, is the self-distillation fine-tuning loss, is the weight parameter, which is dynamically adjusted with the training process.

[0030] Furthermore, the self-distillation fine-tuning loss expression is as follows: ; In the formula, represents the input-output pair taken from the dataset obtained by self-distillation, represents the token length of the output label, represents the probability that the large language model predicts the -th token of the output label given the first tokens of the input and output labels.

[0031] Even further, the calculation formula for the weight parameter is: ; where, is the average L2 distance between the current model output and the true label. It should be noted that the average L2 distance is calculated by sampling one-tenth of the training data (the calculation of the average L2 distance will be elaborated below), represents the average distance at the beginning of the large language model training, represents a hyperparameter. This loss function contains a parameter which determines the weights of the SDFT loss and the SFT loss.

[0032] This method also includes: During the training process, gradually reduce the value of the weight parameter so that the training focus of the large language model gradually shifts from the auxiliary training dataset to the real dataset. As the training progresses, the parameter gradually decreases, causing the large language model to shift its focus from the easier SDFT loss to the more difficult SFT loss.

[0033] For the design of the parameter , utilize the average distance between the output of the current model and the true label. Specifically, before starting each training cycle, first generate an output from the model and calculate the average L2 distance of the token embedding vectors between the model output and the true label: ; where, in the formula represents the token embedding output by the large language model at the -th stage, while represents the token embedding of the real data label. This average distance serves as a measure of the deviation between the current model and the optimal model. As the training progresses and the distance decreases, more difficult tasks should be assigned to the large language model and the weight of should be reduced.

[0034] Then, a simple exponential function is adopted as the progress function, allowing the parameter to be expressed in the form of the calculation formula of the above-mentioned parameter . In this embodiment, in order to reduce the inference time, only one-tenth of the sampled training data is selected to calculate the average distance.

[0035] To verify the effectiveness of this scheme, as Figure 3 shown, experiments are carried out on three subsets of the Amazon review dataset: Toys and Games, Movies and TV, and Kindle Store. These datasets include user behavior sequences and product names, providing a comprehensive view of user-product interactions. To ensure the quality of the datasets, a filtering step is performed, in which users and products with fewer than 10 interactions are deleted. Next, the interaction sequences are sorted in ascending order of timestamp, and each dataset is divided into training, validation, and test sets in a ratio of 8:1:1. For all datasets, 5000 sequences are sampled for testing and 5000 sequences are sampled for validation. For large language models, 10000 sequences are sampled for training, while for traditional models, 100000 sequences are sampled for training. For all large language model-based methods, Llama2-7B is selected as the backbone model, and LoRA is used to fine-tune the large language model on 4 Nvidia A800 GPUs. For traditional recommendation models, SASRec is used as the backbone model.

[0036] Compared with the current state-of-the-art traditional recommendation models and large language recommendation models, SOFT has an average improvement of 50.06% in recommendation accuracy. The hit rate (HR) and normalized discounted cumulative gain (NDCG) are used to measure the accuracy. Among traditional models, SASRec and DORS are used for comparison; among large language models, BIGRec and LLaRA are used as baselines, and three fine-tuning strategies, SFT, SDPO, and SOFT, are compared; additionally, the large language model-enhanced traditional recommendation models LLM-CF and DLLM2Rec are also compared.

[0037] As shown in Table 1 below, SOFT improves the diversity of large language model recommendations. The Self-BLEU and the entropy of the output tokens are used to measure the output diversity, and the output diversities of SOFT, SFT, and the target (true label) are compared; Table 1:

[0038] As shown in Table 2 below, SOFT improves the fitting degree of the large language model on the training data. The hit rate (HR) and mean reciprocal rank (MRR) are used as measurement indicators, and SOFT and SFT are compared; Table 2:

[0039] An embodiment of the present invention also provides a large language model recommendation system, which is trained by using the self-optimizing fine-tuning method of the above large language model recommendation system and is used to generate personalized recommendation content.

[0040] The beneficial effects of the large language model recommendation system of the present invention are the same as the technical effects of the self-optimizing fine-tuning method of the above large language model recommendation system, both of which are used to generate personalized recommendation content more accurately, so they will not be elaborated here.

[0041] Figure 4 It is a structural block diagram of an electronic device provided by an embodiment of the present invention. As Figure 4 shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the following methods: Generate an auxiliary training data set using self-distillation technology. The auxiliary training data set is constructed by the large language model after supervised fine-tuning generating multiple outputs according to the input and selecting the output closest to the real item therefrom; Adopt a curriculum learning fine-tuning strategy to adaptively adjust the training weights of simple tasks and difficult tasks according to the current learning state of the large language model, and gradually shift the training focus from the auxiliary training data set to the real data set.

[0042] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0043] The embodiments of the present invention also provide a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned embodiments.

[0044] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0045] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A self-optimizing fine-tuning method for a large language model recommendation system, characterized in that, The self-optimizing fine-tuning method for the large language model recommendation system includes: Using self-distillation technology to generate an auxiliary training dataset, which is constructed by the large language model after supervised fine-tuning generating multiple outputs for the input and selecting the output closest to the real item therefrom; Adopting a curriculum learning fine-tuning strategy to adaptively adjust the training weights of simple tasks and difficult tasks according to the current learning state of the large language model, and gradually shifting the training focus from the auxiliary training dataset to the real dataset.

2. The self-optimizing fine-tuning method for the large language model recommendation system according to claim 1, wherein The self-distillation technology includes: Generating M outputs for each input through the large language model after supervised fine-tuning, and selecting the output with the smallest distance from the real item as the auxiliary training data by calculating the L2 distance of the token embedding vectors.

3. The self-optimizing fine-tuning method for the large language model recommendation system according to claim 1, wherein, The auxiliary training dataset is used to guide the large language model to gradually adapt to the distribution of the real dataset and serve as an intermediate knowledge base.

4. The self-optimizing fine-tuning method of the large language model recommendation system according to claim 1, characterized in that, The curriculum learning fine-tuning strategy is implemented through the following loss function: ; Among them, is the supervised fine-tuning loss, is the self-distillation fine-tuning loss, is the weight parameter, which is dynamically adjusted with the training process.

5. The self-optimizing fine-tuning method for the large language model recommendation system according to claim 4, characterized in that, The weight parameter has the following calculation formula: ; wherein, is the average L2 distance between the current model output and the true label, represents the average distance at the start of the large language model training, denotes a hyperparameter.

6. The self-optimizing fine-tuning method for the large language model recommendation system according to claim 5, wherein The average L2 distance is calculated by sampling one-tenth of the training data.

7. The self-optimizing fine-tuning method for the large language model recommendation system according to claim 1, characterized in that The method further includes: During the training process, gradually reducing the value of the weight parameter to gradually shift the training focus of the large language model from the auxiliary training dataset to the real dataset.

8. A large language model recommendation system, characterized in that, Training is carried out using the self-optimizing fine-tuning method for the large language model recommendation system according to any one of claims 1-7 for generating personalized recommendation content.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the self-optimizing fine-tuning method for the large language model recommendation system according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the self-optimizing fine-tuning method for the large language model recommendation system according to any one of claims 1 to 7.