A large language model alignment fine-tuning method, system and device for a recommendation system

By obtaining prompt words and training data sets in the recommendation system, filtering and fusion of high-quality knowledge-enhanced texts generated by large language models, building fine-tuning data sets, and aligning and fine-tuning of large language models, solving data sparseness and noise problems and improving recommendation effects.

CN119886255BActive Publication Date: 2025-06-06CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510361196.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-06
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

Due to sparse data and noise generated by large language models, existing recommendation systems are difficult to effectively utilize external knowledge, resulting in poor recommendation results.

Method used

By obtaining the prompt words and training data sets of the targets to be recommended, using the large language model to generate knowledge-enhanced text, filtering high-quality samples, building fine-tuning data sets, aligning and fine-tuning the large language model, fusing feature vectors and embedding vectors, and optimizing the recommended model.

Benefits of technology

It improves the data utilization rate of the recommendation system, reduces the impact of noise, and improves the recommendation effect and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886255B_ABST
    Figure CN119886255B_ABST
Patent Text Reader

Abstract

The present application discloses a method, system and device for aligning and fine-tuning a large language model of a recommendation system. The method converts the key parts of the prompt words and each knowledge-enhanced text into high-dimensional embedding vectors to obtain a first high-dimensional embedding vector and a second high-dimensional embedding vector; screens multiple knowledge-enhanced texts; uses a training data set to train a preset recommendation model to obtain a first embedding vector matrix and a second embedding vector matrix; fuses the high-dimensional embedding vector corresponding to each screened knowledge-enhanced text and the embedding vectors related to the target to be recommended in the second embedding vector matrix; scores each screened knowledge-enhanced text based on the fused feature vector and the first embedding vector matrix corresponding to each screened knowledge-enhanced text; constructs a fine-tuning data set based on the scoring results, and uses the fine-tuning data set to align and fine-tune the large language model. The present application can utilize the extensive external knowledge of the large language model to improve the recommendation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of information recommendation, and in particular to a large language model alignment and fine-tuning method, system and device for a recommendation system. Background Art

[0002] With the development of Internet technology, the rapid rise of major platforms such as e-commerce and social media, users are inevitably faced with the problem of information overload due to the explosive growth of information. In response to this phenomenon, the recommendation system, as a technology that can quickly mine user interests, has become an important part of major network platforms. The recommendation system mines the user's historical behavior data and the feature information of items, and uses algorithm models to recommend personalized content to users, thereby improving user experience and platform efficiency.

[0003] Since the data in the public training data sets of most recommendation systems are relatively sparse, this will lead to the recommendation system failing to effectively utilize other external knowledge. If the recommendation system can obtain comprehensive knowledge about the recommendation target, it will often have better performance. As an emerging technology, large language models have a wide range of application scenarios in recommendation systems. Through self-supervised learning, large language models extract knowledge from massive data and obtain strong external knowledge coverage and knowledge expression capabilities, which provides a way to supplement the background knowledge of recommendation targets. However, it is not a simple task to obtain highly relevant and valuable knowledge from large language models that is highly relevant to the recommendation task. Since the training data sets of large language models cover a wide range and are diverse in form, this means that noise may be introduced into the results generated by the model, which will adversely affect the performance of the recommendation system. Summary of the invention

[0004] This application aims to propose a large language model alignment fine-tuning method, system and device for a recommendation system, which can make full use of the extensive external knowledge of the large language model to improve the recommendation effect.

[0005] In a first aspect, an embodiment of the present application provides a large language model alignment fine-tuning method for a recommendation system, the method comprising:

[0006] Obtain the prompt words and training data set corresponding to the target to be recommended;

[0007] Inputting the prompt word into a large language model to obtain a plurality of knowledge-enhanced texts related to the target to be recommended;

[0008] Converting a key part of the prompt word and each of the knowledge-enhanced texts into a high-dimensional embedding vector to obtain a first high-dimensional embedding vector corresponding to the prompt word and a second high-dimensional embedding vector corresponding to each of the knowledge-enhanced texts, wherein the key part is used to represent information related to the target to be recommended;

[0009] Based on the first high-dimensional embedding vector and the second high-dimensional embedding vector, the multiple knowledge-enhanced texts are screened to obtain multiple screened knowledge-enhanced texts;

[0010] Using the training data set to train a preset recommendation model, obtaining a first embedding vector matrix and a second embedding vector matrix, wherein the first embedding vector matrix is ​​a matrix containing all users and implicit features, and the second embedding vector matrix is ​​a matrix containing all recommended targets and implicit features;

[0011] Fusing the high-dimensional embedding vector corresponding to each of the screened knowledge-enhanced texts with the embedding vectors related to the target to be recommended in the second embedding vector matrix to obtain a fused feature vector corresponding to each of the screened knowledge-enhanced texts;

[0012] Scoring each of the screened knowledge enhancement texts based on the fused feature vector corresponding to each of the screened knowledge enhancement texts and the first embedding vector matrix to obtain a scoring result for each of the screened knowledge enhancement texts;

[0013] According to the scoring results of each of the screened knowledge-enhanced texts, a fine-tuning dataset is constructed, and the large language model is aligned and fine-tuned using the fine-tuning dataset.

[0014] Compared with the prior art, the first aspect of the present application has the following beneficial effects:

[0015] This method obtains the prompt words and training data sets corresponding to the target to be recommended; inputs the prompt words into the large language model to obtain multiple knowledge-enhanced texts related to the target to be recommended, and optimizes the knowledge expression ability of the response text through the prompt words, effectively alleviating the data sparsity problem. By converting the key parts of the prompt words and each knowledge-enhanced text into high-dimensional embedding vectors, the first high-dimensional embedding vector corresponding to the prompt word and the second high-dimensional embedding vector corresponding to each knowledge-enhanced text are obtained. Based on the first high-dimensional embedding vector and the second high-dimensional embedding vector, multiple knowledge-enhanced texts are screened to obtain multiple screened knowledge-enhanced texts, which can remove low-quality samples, retain high-quality samples, and increase the distribution difference of knowledge-enhanced samples in subsequent selections, solving the problem of low efficiency of later fine-tuning. Then, the preset recommendation model is trained using the training data set to obtain the first embedding vector matrix and the second embedding vector matrix. The high-dimensional embedding vector corresponding to each screened knowledge enhancement text is fused with the embedding vector related to the target to be recommended in the second embedding vector matrix to obtain the fused feature vector corresponding to each screened knowledge enhancement text. Based on the fused feature vector and the first embedding vector matrix corresponding to each screened knowledge enhancement text, each screened knowledge enhancement text is scored to obtain the scoring result of each screened knowledge enhancement text. The enhanced knowledge text is scored based on the preset recommendation model to generate a supervision signal, which lays a good data foundation for the subsequent construction of a fine-tuning data set and alignment fine-tuning of the large language model. Finally, according to the scoring results of each screened knowledge enhancement text, a fine-tuning data set is constructed, and the fine-tuning data set is used to align and fine-tune the large language model. The fine-tuning data set is constructed through high-quality knowledge enhancement texts, which can make full use of the extensive external knowledge of the large language model to align and fine-tune the large language model. The large language model after alignment and fine-tuning is used to recommend the target to be recommended, which can improve the recommendation effect.

[0016] In some implementations, the filtering of the plurality of knowledge-enhanced texts based on the first high-dimensional embedding vector and the second high-dimensional embedding vector to obtain a plurality of filtered knowledge-enhanced texts includes:

[0017] Calculating the similarity between the first high-dimensional embedding vector and the second high-dimensional embedding vector, removing the knowledge enhancement texts whose similarity is less than a preset value from the multiple knowledge enhancement texts, and obtaining remaining knowledge enhancement texts;

[0018] Clustering the second high-dimensional embedding vectors corresponding to the remaining knowledge-enhanced texts to obtain clustering results;

[0019] The remaining knowledge enhancement texts are screened according to the clustering results to obtain a plurality of screened knowledge enhancement texts.

[0020] In some implementations, the using the training data set to train a preset recommendation model to obtain a first embedding vector matrix and a second embedding vector matrix includes:

[0021] Build the Bayesian personalized ranking algorithm as a preset recommendation model;

[0022] The preset recommendation model is trained using the training data set to obtain a first embedding vector matrix and a second embedding vector matrix.

[0023] In some implementations, scoring each of the screened knowledge-enhanced texts based on the fused feature vector corresponding to each of the screened knowledge-enhanced texts and the first embedding vector matrix to obtain a scoring result for each of the screened knowledge-enhanced texts includes:

[0024] Inputting the fused feature vector corresponding to each of the screened knowledge-enhanced texts into a multi-layer perceptron to obtain an output result corresponding to each of the screened knowledge-enhanced texts;

[0025] Performing a dot product of the output result and the first embedding vector matrix to obtain a dot product result corresponding to each of the filtered knowledge-enhanced texts;

[0026] According to the dot product result, each of the screened knowledge enhancement texts is scored to obtain a scoring result for each of the screened knowledge enhancement texts.

[0027] In some implementations, constructing a fine-tuning dataset according to the scoring results of each of the screened knowledge-enhanced texts includes:

[0028] Sorting the scoring results of each of the screened knowledge-enhanced texts to obtain sorted scores;

[0029] According to the ranked scores, the screened knowledge-enhanced texts are combined in pairs to construct a fine-tuning dataset.

[0030] In some implementations, before using the fine-tuning dataset to perform alignment fine-tuning on the large language model, the method includes:

[0031] Normalizing the score of each screened knowledge-enhanced text in the fine-tuning dataset into a probability distribution;

[0032] According to the probability distribution, calculating information entropy;

[0033] Calculating a weighting factor according to the information entropy;

[0034] According to the weighting factors, a target optimization function is constructed when aligning and fine-tuning the large language model.

[0035] In some implementations, constructing a target optimization function for fine-tuning the alignment of the large language model based on the weighting factor includes:

[0036] ;

[0037] in, represents the target optimization function, Indicates the current strategy, represents the reference strategy, represents the fine-tuning dataset, Indicates prompt words, Indicates Knowledge enhancement samples, Indicates Knowledge enhancement samples, represents the weighting factor, represents the logistic function, Represents the temperature coefficient.

[0038] In a second aspect, an embodiment of the present application further provides a large language model alignment and fine-tuning system for a recommendation system, the system comprising:

[0039] A data acquisition unit, used to acquire prompt words and training data sets corresponding to the target to be recommended;

[0040] A knowledge-enhanced text acquisition unit, used for inputting the prompt word into a large language model to obtain a plurality of knowledge-enhanced texts related to the target to be recommended;

[0041] A high-dimensional embedding vector conversion unit, used to convert a key part of the prompt word and each of the knowledge-enhanced texts into a high-dimensional embedding vector, to obtain a first high-dimensional embedding vector corresponding to the prompt word and a second high-dimensional embedding vector corresponding to each of the knowledge-enhanced texts, wherein the key part is used to represent information related to the target to be recommended;

[0042] A knowledge enhancement text screening unit, configured to screen the plurality of knowledge enhancement texts based on the first high-dimensional embedding vector and the second high-dimensional embedding vector to obtain a plurality of screened knowledge enhancement texts;

[0043] An embedding vector matrix acquisition unit, used to train a preset recommendation model using the training data set to obtain a first embedding vector matrix and a second embedding vector matrix, wherein the first embedding vector matrix is ​​a matrix containing all users and implicit features, and the second embedding vector matrix is ​​a matrix containing all recommended targets and implicit features;

[0044] An embedding vector fusion unit, used to fuse the high-dimensional embedding vector corresponding to each of the screened knowledge-enhanced texts with the embedding vectors related to the target to be recommended in the second embedding vector matrix to obtain a fused feature vector corresponding to each of the screened knowledge-enhanced texts;

[0045] A knowledge enhancement text scoring unit, used for scoring each of the screened knowledge enhancement texts based on the fused feature vector corresponding to each of the screened knowledge enhancement texts and the first embedding vector matrix, to obtain a scoring result for each of the screened knowledge enhancement texts;

[0046] The large language model alignment and fine-tuning unit is used to construct a fine-tuning dataset according to the scoring results of each of the screened knowledge-enhanced texts, and use the fine-tuning dataset to align and fine-tune the large language model.

[0047] In a third aspect, an embodiment of the present application further provides an electronic device comprising at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the large language model alignment fine-tuning method for a recommendation system as described above.

[0048] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the large language model alignment and fine-tuning method for a recommendation system as described above.

[0049] It can be understood that the beneficial effects of the second to fourth aspects compared with the related art are the same as the beneficial effects of the first aspect compared with the related art. Please refer to the relevant description in the first aspect, and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0051] Figure 1 It is a flowchart of an embodiment of a large language model alignment and fine-tuning method for a recommendation system provided by the present application;

[0052] Figure 2 It is a schematic diagram of the overall process in the best embodiment of the large language model alignment and fine-tuning method for the recommendation system provided by the present application;

[0053] Figure 3It is a structural diagram of an embodiment of a large language model alignment and fine-tuning system for a recommendation system provided by the present application;

[0054] Figure 4 It is a schematic diagram of the structure of an embodiment of the electronic device provided by the present application. DETAILED DESCRIPTION

[0055] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.

[0056] In the description of this application, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0057] In the description of the present application, it should be understood that the descriptions involving orientation, such as the orientation or positional relationship indicated as up, down, etc., are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.

[0058] In the description of this application, it should be noted that, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meaning of the above terms in this application based on the specific content of the technical solution.

[0059] With the development of Internet technology, the rapid rise of major platforms such as e-commerce and social media, users are inevitably faced with the problem of information overload due to the explosive growth of information. In response to this phenomenon, the recommendation system, as a technology that can quickly mine user interests, has become an important part of major network platforms. The recommendation system mines the user's historical behavior data and the feature information of items, and uses algorithm models to recommend personalized content to users, thereby improving user experience and platform efficiency.

[0060] Since the data in the public training data sets of most recommendation systems are relatively sparse, this will lead to the recommendation system failing to effectively utilize other external knowledge. If the recommendation system can obtain comprehensive knowledge about the recommendation target, it will often have better performance. As an emerging technology, large language models have a wide range of application scenarios in recommendation systems. Through self-supervised learning, large language models extract knowledge from massive data and obtain strong external knowledge coverage and knowledge expression capabilities, which provides a way to supplement the background knowledge of recommendation targets. However, it is not a simple task to obtain highly relevant and valuable knowledge from large language models that is highly relevant to the recommendation task. Since the training data sets of large language models cover a wide range and are diverse in form, this means that noise may be introduced into the results generated by the model, which will adversely affect the performance of the recommendation system.

[0061] In order to solve the problem that there is some noise irrelevant to the recommendation system in the response text generated by the large language model, which has an adverse effect on the performance of the recommendation system, the present application proposes a large language model alignment fine-tuning method, system and device for the recommendation system.

[0062] Reference Figure 1 The embodiment of the present application provides a large language model alignment fine-tuning method for a recommendation system, the method comprising the following steps:

[0063] Step S100, obtaining prompt words and training data sets corresponding to the target to be recommended;

[0064] Step S200: input the prompt word into the large language model to obtain multiple knowledge-enhanced texts related to the target to be recommended;

[0065] Step S300: convert the key part of the prompt word and each knowledge-enhanced text into a high-dimensional embedding vector to obtain a first high-dimensional embedding vector corresponding to the prompt word and a second high-dimensional embedding vector corresponding to each knowledge-enhanced text, wherein the key part is used to represent information related to the target to be recommended;

[0066] Step S400: screening multiple knowledge-enhanced texts based on the first high-dimensional embedding vector and the second high-dimensional embedding vector to obtain multiple screened knowledge-enhanced texts;

[0067] Step S500: training a preset recommendation model using a training data set to obtain a first embedding vector matrix and a second embedding vector matrix, wherein the first embedding vector matrix is ​​a matrix containing all users and implicit features, and the second embedding vector matrix is ​​a matrix containing all recommended targets and implicit features;

[0068] Step S600, fusing the high-dimensional embedding vector corresponding to each screened knowledge-enhanced text with the embedding vector related to the target to be recommended in the second embedding vector matrix to obtain a fused feature vector corresponding to each screened knowledge-enhanced text;

[0069] Step S700: scoring each of the screened knowledge enhancement texts based on the fused feature vector and the first embedding vector matrix corresponding to each of the screened knowledge enhancement texts to obtain a scoring result for each of the screened knowledge enhancement texts;

[0070] Step S800: construct a fine-tuning dataset based on the scoring results of each screened knowledge-enhanced text, and use the fine-tuning dataset to align and fine-tune the large language model.

[0071] In this embodiment, by obtaining the prompt words and training data sets corresponding to the target to be recommended; inputting the prompt words into the large language model, multiple knowledge-enhanced texts related to the target to be recommended are obtained, and the knowledge expression ability of the response text is optimized by the prompt words, effectively alleviating the data sparsity problem. By converting the key parts of the prompt words and each knowledge-enhanced text into high-dimensional embedding vectors, the first high-dimensional embedding vector corresponding to the prompt words and the second high-dimensional embedding vector corresponding to each knowledge-enhanced text are obtained, and based on the first high-dimensional embedding vector and the second high-dimensional embedding vector, multiple knowledge-enhanced texts are screened to obtain multiple screened knowledge-enhanced texts, which can remove low-quality samples, retain high-quality samples, and increase the distribution difference of knowledge-enhanced samples in subsequent selections, solving the problem of low efficiency of later fine-tuning. Then, the preset recommendation model is trained using the training data set to obtain the first embedding vector matrix and the second embedding vector matrix. The high-dimensional embedding vector corresponding to each screened knowledge enhancement text is fused with the embedding vector related to the target to be recommended in the second embedding vector matrix to obtain the fused feature vector corresponding to each screened knowledge enhancement text. Based on the fused feature vector and the first embedding vector matrix corresponding to each screened knowledge enhancement text, each screened knowledge enhancement text is scored to obtain the scoring result of each screened knowledge enhancement text. The enhanced knowledge text is scored based on the preset recommendation model to generate a supervision signal, which lays a good data foundation for the subsequent construction of a fine-tuning data set and alignment fine-tuning of the large language model. Finally, according to the scoring results of each screened knowledge enhancement text, a fine-tuning data set is constructed, and the fine-tuning data set is used to align and fine-tune the large language model. The fine-tuning data set is constructed through high-quality knowledge enhancement texts, which can make full use of the extensive external knowledge of the large language model to align and fine-tune the large language model. The large language model after alignment and fine-tuning is used to recommend the target to be recommended, which can improve the recommendation effect.

[0072] The above-mentioned large language model can be an open source lightweight large language model Llama-3.2-3B model, or can be other large language models known to those skilled in the art, and this embodiment does not make any specific limitation.

[0073] The above converting the key part of the prompt word and each knowledge enhancement text into a high-dimensional embedding vector to obtain a first high-dimensional embedding vector corresponding to the prompt word and a second high-dimensional embedding vector corresponding to each knowledge enhancement text can be performed by inputting the prompt word into an mpnet model (i.e., the all-mpnet-base-v2 model), converting the key part of the prompt word into a high-dimensional embedding vector through mpnet model encoding, and obtaining the first high-dimensional embedding vector corresponding to the prompt word; inputting each knowledge enhancement text into the mpnet model, converting each knowledge enhancement text into a high-dimensional embedding vector through mpnet model encoding, and obtaining the second high-dimensional embedding vector corresponding to each knowledge enhancement text.

[0074] The above-mentioned method of screening multiple knowledge enhancement texts based on the first high-dimensional embedding vector and the second high-dimensional embedding vector to obtain multiple screened knowledge enhancement texts can be to calculate the similarity between the first high-dimensional embedding vector and the second high-dimensional embedding vector, and then screen the multiple knowledge enhancement texts according to the similarity to obtain multiple screened knowledge enhancement texts.

[0075] The above-mentioned use of the training data set to train the preset recommendation model to obtain the first embedding vector matrix and the second embedding vector matrix may be to use the Bayesian personalized ranking algorithm as the preset recommendation model, and use the training data set to train the preset recommendation model to obtain the first embedding vector matrix and the second embedding vector matrix. Alternatively, other recommendation algorithms known to those skilled in the art may be used as the preset recommendation model, which is not specifically limited in this embodiment.

[0076] The above-mentioned fusion of the high-dimensional embedding vector corresponding to each filtered knowledge enhancement text and the embedding vector related to the target to be recommended in the second embedding vector matrix can be achieved by using a splicing method to fuse the high-dimensional embedding vector corresponding to each filtered knowledge enhancement text and the embedding vector related to the target to be recommended in the second embedding vector matrix.

[0077] According to the scoring result of each screened knowledge enhancement text, a fine-tuning dataset is constructed. The scoring result of each screened knowledge enhancement text is sorted by score, and then the sorted knowledge enhancement texts are combined in pairs to construct the fine-tuning dataset.

[0078] In some implementations, based on the first high-dimensional embedding vector and the second high-dimensional embedding vector, a plurality of knowledge-enhanced texts are screened to obtain a plurality of screened knowledge-enhanced texts, including:

[0079] Calculate the similarity between the first high-dimensional embedding vector and the second high-dimensional embedding vector, remove the knowledge enhancement texts whose similarity is less than a preset value from the multiple knowledge enhancement texts, and obtain the remaining knowledge enhancement texts;

[0080] Clustering the second high-dimensional embedding vectors corresponding to the remaining knowledge-enhanced texts to obtain clustering results;

[0081] The remaining knowledge enhancement texts are screened according to the clustering results to obtain a plurality of screened knowledge enhancement texts.

[0082] In this embodiment, by calculating the similarity between the first high-dimensional embedding vector and the second high-dimensional embedding vector, the knowledge enhancement texts with similarity less than a preset value are removed from the multiple knowledge enhancement texts to obtain the remaining knowledge enhancement texts, which can eliminate low-quality samples and retain high-quality samples. The second high-dimensional embedding vectors corresponding to the remaining knowledge enhancement texts are clustered to obtain clustering results, and the remaining knowledge enhancement texts are screened according to the clustering results to obtain multiple screened knowledge enhancement texts, which can increase the distribution difference of the knowledge enhancement samples in subsequent selections, thereby solving the problem of low fine-tuning efficiency when the samples are too similar due to direct preference optimization of the large language model.

[0083] In some implementations, a preset recommendation model is trained using a training data set to obtain a first embedding vector matrix and a second embedding vector matrix, including:

[0084] Build the Bayesian personalized ranking algorithm as a preset recommendation model;

[0085] The preset recommendation model is trained using the training data set to obtain a first embedding vector matrix and a second embedding vector matrix.

[0086] In this embodiment, the Bayesian personalized ranking algorithm is constructed as a preset recommendation model; the preset recommendation model is trained using a training data set to obtain a first embedding vector matrix and a second embedding vector matrix. In this way, the purpose of using the Bayesian personalized ranking algorithm is to use a recommendation system to score the response text of the large language model to generate a supervision signal, thereby producing a fine-tuning data set and fine-tuning and aligning the large language model.

[0087] In some implementations, each screened knowledge enhancement text is scored based on the fused feature vector and the first embedding vector matrix corresponding to each screened knowledge enhancement text, and a scoring result of each screened knowledge enhancement text is obtained, including:

[0088] Input the fused feature vector corresponding to each filtered knowledge-enhanced text into a multi-layer perceptron to obtain the output result corresponding to each filtered knowledge-enhanced text;

[0089] Perform a dot product of the output result and the first embedding vector matrix to obtain the dot product result corresponding to each filtered knowledge-enhanced text;

[0090] According to the dot product result, each filtered knowledge enhancement text is scored to obtain a scoring result of each filtered knowledge enhancement text.

[0091] In this embodiment, the fused feature vector corresponding to each screened knowledge-enhanced text is input into a multi-layer perceptron to obtain the output result corresponding to each screened knowledge-enhanced text; the output result is dot-multiplied with the first embedding vector matrix to obtain the dot-multiplication result corresponding to each screened knowledge-enhanced text; and each screened knowledge-enhanced text is scored according to the dot-multiplication result to obtain the scoring result of each screened knowledge-enhanced text. In this way, by scoring and sorting each screened knowledge-enhanced text to generate a supervisory signal, this method can efficiently score the text generated by the large language model, laying a good data foundation for the later construction of a fine-tuning dataset and alignment and fine-tuning of the large language model.

[0092] In some implementations, constructing a fine-tuning dataset based on the scoring results of each screened knowledge-enhanced text includes:

[0093] Sorting the scoring results of each screened knowledge-enhanced text to obtain a sorted score;

[0094] According to the ranked scores, the screened knowledge-enhanced texts are combined in pairs to construct a fine-tuning dataset.

[0095] In this embodiment, the scoring results of each screened knowledge-enhanced text are sorted to obtain the sorted scores; according to the sorted scores, the screened knowledge-enhanced texts are combined in pairs to construct a fine-tuning dataset. In this way, by constructing a fine-tuning dataset with high-quality knowledge-enhanced texts, the extensive external knowledge of the large language model can be fully utilized in the later stage to align and fine-tune the large language model.

[0096] In some implementations, before using the fine-tuning dataset to align and fine-tune the large language model, the method includes:

[0097] Normalize the score of each filtered knowledge-enhanced text in the fine-tuning dataset into a probability distribution;

[0098] According to the probability distribution, calculate the information entropy;

[0099] According to the information entropy, the weighting factor is calculated;

[0100] According to the weighting factors, the objective optimization function for alignment fine-tuning of the large language model is constructed.

[0101] In this embodiment, the scores of each filtered knowledge-enhanced text in the fine-tuning dataset are normalized into probability distribution; information entropy is calculated according to the probability distribution; weighting factors are calculated according to the information entropy; and the target optimization function for alignment fine-tuning of the large language model is constructed according to the weighting factors. In this way, the weighting factors can make the large language model pay more attention to samples with strong discrimination and pay less attention to samples with uniform distribution, thereby improving the quality of alignment fine-tuning.

[0102] In some implementations, constructing a target optimization function for fine-tuning alignment of a large language model based on a weighting factor includes:

[0103] ;

[0104] in, represents the target optimization function, Indicates the current strategy, represents the reference strategy, represents the fine-tuning dataset, Indicates prompt words, Indicates Knowledge enhancement samples, Indicates Knowledge enhancement samples, represents the weighting factor, represents the logistic function (i.e., the logistic function, also the sigmoid function), Represents the temperature coefficient.

[0105] To facilitate understanding by those skilled in the art, a set of best embodiments is provided below:

[0106] The prior art has the following problems:

[0107] 1. Item Content information is too little or missing.

[0108] The item content itself carries too little information, which makes it difficult to capture the fine-grained features of the item. It is impossible to understand new item attributes through the content features of the item content. The model cannot fully capture the content information, and the model is prone to overfitting other unimportant features, resulting in a decrease in generalization ability. Item content refers to some features (Content) extracted for an item (Item). For example, a common public recommendation system dataset such as MovieLens contains data such as movie serial number, movie name, movie category, release time, user serial number, user gender, user age, job type, zip code, user rating of the movie, and timestamp, which is insufficient in describing the content of the movie itself.

[0109] Second, there is a lack of utilization of external knowledge.

[0110] Traditional collaborative filtering algorithms rely only on the interaction data between users and items, but fail to effectively utilize other external information (such as item descriptions, user reviews, and domain knowledge, etc.). This limits the performance of the algorithm in sparse data scenarios and makes it difficult for the recommendation results to be more accurate and diverse.

[0111] 3. There is noise in the response text of the large language model that is irrelevant to the recommendation system.

[0112] Large language models are not naturally trained to perform data augmentation for recommendation systems. When large language models are used to perform data augmentation for recommendation systems, the expected results are poor. It is not easy to obtain highly relevant and valuable knowledge from large language models. Since the training datasets of large language models cover a wide range and are diverse, this means that noise may be introduced into the results generated by the model, which will adversely affect the performance of the recommendation system.

[0113] In order to solve the problems existing in the above-mentioned prior art, this embodiment designs a method for aligning and fine-tuning a large language model for a recommendation system. For the recommendation scenario, a data expansion method based on a prompt template is designed to make full use of the extensive external knowledge of the large language model to enhance the knowledge of the public original data set of the recommendation system. The knowledge-enhanced data set contains more rich semantic information that can characterize user preferences and item characteristics. In order to further improve the recommendation performance, a text quality evaluation method for the recommendation effect of the recommendation system is designed to construct a large language model fine-tuning data set. Based on the DPO principle, by introducing the effect feedback of the collaborative filtering algorithm (such as the BPR algorithm), the large language model is retrained so that it can more effectively mine the knowledge in the recommended text that is beneficial to the recommendation system. This feedback mechanism uses the recommendation effect of the collaborative filtering algorithm as a constraint condition by designing the target optimization function to guide the fine-tuning of the large language model, thereby realizing the dynamic collaborative optimization of the recommendation system and the large language model.

[0114] Reference Figure 2 The technical solution of this embodiment specifically includes the following contents:

[0115] 1. Knowledge-enhanced text generation module.

[0116] For large language models, prompt words are first used to generate response text. For public recommendation system datasets such as MovieLens, they contain data such as movie serial number, movie name, movie category, release time, user serial number, user gender, user age, job type, zip code, user rating of the movie, and timestamp.

[0117] Table 1 Original movie information in MovieLens

[0118]

[0119] In the original dataset MovieLens, as shown in Table 1, the only additional information about movies is the style and the movie name. There is too little content information, so a large language model is needed to enhance the knowledge of the text content.

[0120] For example, for the same movie, given its name, release date and style, the prompt word is , the response text collection is , the number of texts in the response text collection is Positive integer.

[0121] Input the large language model prompt word text into the large language model and run it repeatedly times, we get the value of movie m (i.e. the target to be recommended) Knowledge-enhancing texts, Knowledge-enhanced texts as a knowledge-enhanced text set . Assume that there are Film Projects, Film Projects Each knowledge enhancement text is ,Right now:

[0122] ;

[0123] in, Indicates that the prompt word Input into a large language model.

[0124] For prompt words , introducing the way of thinking chain, compared with directly using prompt words, thinking chain prompt words can significantly improve the performance of the model. Thinking chain prompt words can be constructed in English form, and the content of the constructed prompt words can be changed according to actual conditions, which will not be described in detail in this embodiment.

[0125] 2. Data cleaning module.

[0126] Considering that the large language model generates text by predicting the probability distribution of words, it does not understand or verify the authenticity of the generated content, and sometimes generates answers that are completely different from the user's intention, or even fabricates content that has nothing to do with the question. Therefore, the data cleaning module performs a preliminary scoring of the knowledge-enhanced text generated by the large language model, thereby eliminating low-quality samples and retaining high-quality samples.

[0127] all-mpnet-base-v2 is a sentence embedding vector generation model that can efficiently generate sentence embedding vectors. Do a sample score in advance to ensure that the generated knowledge-enhanced text is relevant to the prompt word. Use all-mpnet-base-v2 to score the key parts of the prompt word. and knowledge-enhanced text To encode, the key part It can be changed according to the actual situation of the prompt word. This embodiment does not make specific restrictions. It can be a movie or other recommendation targets, and the high-dimensional embedding vector corresponding to the knowledge-enhanced text and the prompt word is extracted.

[0128] Assume there are M movie items in total, and the knowledge enhancement text of movie item m is , the above uses the mpnet model to enhance each knowledge text Transformed into a high-dimensional embedding vector with dimension , that is, converted into a high-dimensional embedding vector through the following formula:

[0129] ;

[0130] n ;

[0131] Then the similarity between the high-dimensional embedding vectors of the knowledge-enhanced text and the prompt word is calculated by cosine similarity for:

[0132] ;

[0133] in, represents the all-mpnet-base-v2 model, Indicates that the key part of the word will be prompted The high-dimensional vector (i.e., the first high-dimensional embedding vector) obtained by inputting the mpnet model (i.e., the all-mpnet-base-v2 model) is Indicates knowledge enhancement text Input the high-dimensional vector obtained in the mpnet model (i.e., the second high-dimensional embedding vector).

[0134] The similarity Less than the preset value Knowledge-enhanced text removal, retaining values ​​greater than or equal to the preset value knowledge enhancement text (i.e. the remaining knowledge enhancement text).

[0135] ;

[0136] if , then it means knowledge-enhanced text Reserved; if , it means that the knowledge enhancement text is removed.

[0137] Assume that there are p pieces of knowledge-enhanced text remaining, and their high-dimensional embedding vector set is for:

[0138] ;

[0139] The preset value in this embodiment The value can be 0.4, the default value It can be changed according to actual conditions and is not specifically limited in this embodiment.

[0140] In order to solve the problem that the fine-tuning efficiency of large language models is too low when the samples are too similar, this embodiment uses the idea of ​​clustering to increase the distribution difference of knowledge-enhanced samples in subsequent selections. Specifically:

[0141] The above has obtained p remaining knowledge enhancement texts for movie m, and a high-dimensional embedding vector set of p remaining knowledge enhancement texts , using the idea of ​​hierarchical clustering, it is divided according to the discrimination of Euclidean distance to obtain each cluster, and the distance threshold is set , the distance is less than or equal to The high-dimensional embedding vectors of are grouped into the same cluster. Refers to the method of calculating the distance between clusters (i.e. the Euclidean distance used), that is, calculating the maximum distance of all points between clusters. Clustering is performed in the following way:

[0142] ;

[0143] ;

[0144] ;

[0145] in, Represents the calculated Euclidean distance value, represents the total number of dimensions of the high-dimensional embedding vector, Represents the dimension, Represents a high-dimensional embedding vector No. Dimensions, Represents a high-dimensional embedding vector No. Dimensions, represents hierarchical clustering, Indicates Knowledge-enhancing texts, Indicates Knowledge-enhancing texts, represents the clustering result, Indicates clusters, and the final number of clusters is indivual.

[0146] Then randomly select samples, and finally form knowledge-enhanced texts (i.e., multiple screened knowledge-enhanced texts).

[0147] 3. Feature fusion module.

[0148] For the MovieLens dataset, the matrix after training with the Bayesian personalized ranking algorithm (i.e., the BPR algorithm) is a matrix P and a matrix Q. The P matrix is ​​the User matrix (i.e., the first embedding vector matrix), which contains a matrix of users and implicit features; the Q matrix is ​​the Item matrix (i.e., the second embedding vector matrix), which contains a matrix of implicit features and items (i.e., the recommended targets, which can be the movies exemplified in this embodiment (i.e., including Movie 1, Movie 2, ..., Movie m, ..., Movie M), and the recommended targets can also be music and other products, etc., which are not specifically limited in this embodiment); the R matrix is ​​the User-Item matrix, which consists of Get, that is:

[0149] ;

[0150] in, Indicates user For items The actual rating of Indicates user For items The actual rating of represents the parameter vector of the Bayesian personalized ranking algorithm, represents the regularization coefficient, represents the sigmoid function, Represents the training dataset (i.e., the MovieLens dataset).

[0151] The purpose of using the BPR recommendation algorithm is to use a recommendation system to score the response text of the large language model to generate a supervision signal, thereby creating a fine-tuning dataset and fine-tuning and aligning the large language model.

[0152] The BPR algorithm learns latent features based on the user's implicit feedback (clicks). In the original BPR recommendation algorithm training process, only the user interaction sequence is used, and no other item features are used. Specifically, for the MovieLens dataset, each row is actually an interaction between a user and a movie, as shown in Table 2.

[0153] Table 2 User interaction sequence

[0154]

[0155] The matrix Q does not contain movie content information. Therefore, it is necessary to integrate the movie content into the matrix Q through feature fusion. The trained Q matrix has a dimension of The embedding vector matrix of , specifically:

[0156] ;

[0157] The item matrix Q is the feature vector containing all the movies, where Row indicates the The feature vector of a movie. Represents the embedding vector of movie m in the Q matrix, with the specific dimensions (1, ), the set of knowledge-enhanced texts of known movie m after screening is , a total of h knowledge-enhanced texts. For the knowledge-enhanced text collection , through the mpnet model to enhance the knowledge text collection Each knowledge-enhanced text in is encoded, and the encoding result is a dimension of The embedding vector matrix of , whose text embedding is:

[0158] ;

[0159] in, Represents a dimension The embedding vector space of Indicates The text is input into the pre-trained model mpnet to obtain the embedding vector matrix (i.e. Knowledge Enhanced Text high-dimensional vector).

[0160] Now we need to transform the dimension into (1, )of and the dimensions are (1, )of Perform feature fusion. First, concatenate the two feature vectors to form a fused feature vector. and They come from different feature spaces and may contain complementary information. Through the concatenation operation, they can be merged into a new input vector , so that the model can simultaneously consider and information.

[0161] ;

[0162] A multi-layer perceptron (MLP) maps an input vector to a target output vector through a series of linear transformations and nonlinear activation functions.

[0163] The dimension is ( ), and this embodiment needs to be connected with the matrix P (the dimension of the matrix P is equal to the dimension of the matrix Q) ) to calculate the prediction score, so we need to From the dimension ( ) is converted to dimension. Therefore, this embodiment needs to use MLP to achieve this purpose. Specifically: MLP includes L layers of fully connected layers, and the first layer is mapped as follows:

[0164] ;

[0165] in, is the fused feature vector, is the weight matrix, is the bias vector, the function It's the first layer The activation function is:

[0166] ;

[0167] The activation function is suitable for sparse data, which makes the model less likely to be oversaturated and solves the dead ReLU problem.

[0168] No. The layer mapping is:

[0169] ;

[0170] in, is the weight matrix, is the bias vector, the function It is Layer Activation function.

[0171] Through MLP, From the dimension ( ) is converted to Embedding vector matrix of dimension .

[0172] Then use the mean square error (MSE) as the loss function to minimize and The squared error between them is calculated to make the two as close as possible in the collaborative space.

[0173] ;

[0174] in, is the number of training samples, It is The knowledge-enhanced samples are passed through the MLP The result of layer transformation is Indicates the first The embedding vector of the knowledge-enhanced sample in the Q matrix.

[0175] Therefore, for each knowledge-enhanced text of movie m , the final output result is obtained through MLP, that is, The corresponding item matrix that integrates the Item Content information (i.e. each filtered knowledge-enhanced text corresponding output results).

[0176] 4. Text quality measurement module based on user-item representation.

[0177] The recommended texts generated by the large language model need to be scored and sorted to generate supervision signals. This method can efficiently score the texts generated by the large language model. The specific scoring process is:

[0178] For the large language model, when the same prompt word is input to the same movie item, the response of the large language model is not exactly the same each time, so it needs to be scored. For the h knowledge enhancement samples generated by the same item and the same prompt word, they are scored in the following way.

[0179] For each piece of knowledge enhancement text for movie m In this embodiment, the corresponding item matrix of fused Item Content information is obtained , now we need to To score, :

[0180] ;

[0181] in, is the set of users who have interacted with movie m, is the user matrix (i.e. the first embedding vector matrix), It is an item matrix that integrates Item Content information.

[0182] By calculating the ratings of movie m for all users and taking the average, we can get the knowledge-enhanced text Score (i.e. the scoring result).

[0183] 5. Large language model fine-tuning dataset construction module.

[0184] Based on the scoring and ranking results of h knowledge-enhanced samples of the same item, a large language model fine-tuning dataset is constructed. The score size order of the same movie m is h knowledge-enhanced text collections. That is, according to the sorted scores, Knowledge-enhanced text in Combine two by two to construct data. Assume that there are movies, then there are a total of Piece of data.

[0185] 6. Large language model alignment fine-tuning module.

[0186] Based on the principle of DPO (direct preference optimization) for fine-tuning a large language model, we construct the target optimization function for alignment fine-tuning of a large language model, which specifically includes the following contents:

[0187] The direct preference optimization algorithm uses instruction preference data to fine-tune the large language model, simplifying the reinforcement learning process into a supervised fine-tuning similar to the large language model, thereby improving the speed and stability of training. Its purpose is to maximize the model's reward for input data, that is, to maximize the difference between the preferred and non-preferred sample data of the large language model, and then learn human preferences.

[0188] Now, human preference is actually the recommendation effect feedback of the BPR algorithm. Based on the feedback of the recommendation effect, the large language model can be fine-tuned to remove the noise from the knowledge-enhanced text of the large language model, thereby better facilitating the recommendation of the recommendation system.

[0189] Based on the principle of fine-tuning DPO (direct preference optimization) of the large language model, a weighting function is introduced according to the previous score of the knowledge-enhanced samples, so that when calculating the loss, the score difference is used to calculate the weighted function. and The size of the loss function is dynamically adjusted. and For the constructed A piece of data in a data set.

[0190] Set the weighting factor for the score ,The idea of ​​information entropy is used to weight it. If the score distribution is more concentrated and more discriminative, its weight is higher, otherwise the weight is lower.

[0191] For a given score , normalized to a probability distribution ,Right now:

[0192] ;

[0193] Information Entropy Describes the uniformity of the score distribution, namely:

[0194] ;

[0195] is the weighting factor, namely:

[0196] ;

[0197] is a very small positive integer to prevent H=0;

[0198] Based on the above weighting factors, the objective optimization function is constructed as follows:

[0199] ;

[0200] in, represents the target optimization function, Indicates the current strategy, represents the reference strategy, represents the fine-tuning dataset, Indicates prompt words, Indicates Knowledge enhancement samples, Indicates Knowledge enhancement samples, represents the weighting factor, represents the logistic function, that is, the logistic function, represents the temperature coefficient, where .

[0201] This method can pay more attention to samples with strong discrimination and reduce the attention to samples with uniform distribution, thereby improving the quality of alignment fine-tuning.

[0202] For the large language model, the open source lightweight large language model Llama-3.2-3B model can be selected as the original large language model, and then based on the constructed large language model fine-tuning dataset and target optimization function, the original large language model (ie, Llama-3.2-3B model) is aligned and fine-tuned to generate a large language model that is more suitable for the recommendation system.

[0203] Compared with the prior art, the technical solution of this embodiment has the following advantages:

[0204] (1) This embodiment designs a novel method for fine-tuning a large language model for collaborative filtering scenarios in recommendation systems. This method can fine-tune the large language model to achieve better recommendation effects for recommendation systems based on collaborative filtering.

[0205] This method targets the collaborative filtering scenario of the recommendation system, and uses the extensive external knowledge coverage of the large language model to mine the key knowledge implicit in the recommendation text through carefully designed prompt templates. On this basis, the recommendation data set is expanded to generate a high-quality enhanced data set. Through the design of prompt words, the knowledge expression ability of the recommendation data is further optimized, which effectively alleviates the data sparsity problem and provides richer training data for the training of the collaborative filtering algorithm, thereby better adjusting the large language model and improving the recommendation performance of the large language model.

[0206] (2) Generate supervisory signals for scoring and ranking the generated text of the large language model (LLM) through recommendation feedback. Use DPO to align the LLM with the recommendation system in the recommendation system field, so that the large language model can better output knowledge that meets the needs of the recommendation system.

[0207] (3) Compared with the prior art, this embodiment not only expands the training samples of the recommendation dataset, but also fully captures the use of interaction sequence information between users and items, and deeply mines the content in the knowledge enhancement samples provided by the large language model that is deeply related to the recommendation system.

[0208] Reference Figure 3 The embodiment of the present application also provides a large language model alignment fine-tuning system for a recommendation system, which includes a data acquisition unit 100, a knowledge-enhanced text acquisition unit 200, a high-dimensional embedding vector conversion unit 300, a knowledge-enhanced text screening unit 400, an embedding vector matrix acquisition unit 500, an embedding vector fusion unit 600, a knowledge-enhanced text scoring unit 700 and a large language model alignment fine-tuning unit 800, wherein:

[0209] The data acquisition unit 100 is used to acquire the prompt words and training data sets corresponding to the target to be recommended;

[0210] The knowledge-enhanced text acquisition unit 200 is used to input the prompt word into the large language model to obtain a plurality of knowledge-enhanced texts related to the target to be recommended;

[0211] A high-dimensional embedding vector conversion unit 300 is used to convert the key part of the prompt word and each knowledge-enhanced text into a high-dimensional embedding vector, to obtain a first high-dimensional embedding vector corresponding to the prompt word and a second high-dimensional embedding vector corresponding to each knowledge-enhanced text, wherein the key part is used to represent information related to the target to be recommended;

[0212] A knowledge enhancement text screening unit 400 is used to screen a plurality of knowledge enhancement texts based on the first high-dimensional embedding vector and the second high-dimensional embedding vector to obtain a plurality of screened knowledge enhancement texts;

[0213] The embedding vector matrix acquisition unit 500 is used to train the preset recommendation model using the training data set to obtain a first embedding vector matrix and a second embedding vector matrix, wherein the first embedding vector matrix is ​​a matrix containing all users and implicit features, and the second embedding vector matrix is ​​a matrix containing all recommended targets and implicit features;

[0214] An embedding vector fusion unit 600 is used to fuse the high-dimensional embedding vector corresponding to each screened knowledge-enhanced text with the embedding vector related to the target to be recommended in the second embedding vector matrix to obtain a fused feature vector corresponding to each screened knowledge-enhanced text;

[0215] The knowledge enhancement text scoring unit 700 is used to score each of the filtered knowledge enhancement texts based on the fused feature vector and the first embedding vector matrix corresponding to each of the filtered knowledge enhancement texts to obtain a scoring result for each of the filtered knowledge enhancement texts;

[0216] The large language model alignment and fine-tuning unit 800 is used to construct a fine-tuning dataset based on the scoring results of each filtered knowledge enhancement text, and use the fine-tuning dataset to align and fine-tune the large language model.

[0217] It should be noted that since the large language model alignment and fine-tuning system of a recommendation system in this embodiment and the large language model alignment and fine-tuning method of a recommendation system mentioned above are based on the same inventive concept, the corresponding contents in the method embodiment are also applicable to the system embodiment and will not be described in detail here.

[0218] Reference Figure 4 The embodiment of the present application further provides an electronic device, the electronic device comprising:

[0219] at least one memory;

[0220] at least one processor;

[0221] at least one program;

[0222] The program is stored in the memory, and the processor executes at least one program to implement the large language model alignment and fine-tuning method of the recommendation system implemented in the present disclosure.

[0223] The electronic device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc.

[0224] The electronic device according to the embodiment of the present application is described in detail below.

[0225] The processor 1600 may be implemented by a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present disclosure;

[0226] The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1700 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1700, and the processor 1600 calls and executes the large language model alignment fine-tuning method of the recommendation system of the embodiment of the present disclosure.

[0227] Input / output interface 1800, used to implement information input and output;

[0228] Communication interface 1900, used to realize communication interaction between the device and other devices, which can be realized through wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);

[0229] Bus 2000 , which transmits information between various components of the device (e.g., processor 1600 , memory 1700 , input / output interface 1800 , and communication interface 1900 );

[0230] The processor 1600 , the memory 1700 , the input / output interface 1800 , and the communication interface 1900 are connected to each other in communication within the device via the bus 2000 .

[0231] The embodiment of the present disclosure also provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the large language model alignment and fine-tuning method of the recommendation system.

[0232] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0233] The embodiments described in the embodiments of the present disclosure are intended to more clearly illustrate the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are also applicable to similar technical problems.

[0234] Those skilled in the art will appreciate that the technical solutions shown in the figures do not limit the embodiments of the present disclosure and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0235] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0236] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0237] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0238] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0239] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0240] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0241] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0242] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store programs. The above is a detailed description of the embodiments of the present application in conjunction with the accompanying drawings, but the present application is not limited to the above embodiments. Within the scope of knowledge possessed by ordinary technicians in the relevant technical field, various changes can be made without departing from the purpose of the present application.

[0243] The embodiments of the present application are described in detail above in conjunction with the accompanying drawings, but the present application is not limited to the above embodiments. Various changes can be made within the knowledge scope of ordinary technicians in the relevant technical field without departing from the purpose of the present application.

Claims

1. A large language model alignment and fine-tuning method for a recommendation system, characterized in that: The method comprises: Obtain the prompt words and training data set corresponding to the target to be recommended; Inputting the prompt word into a large language model to obtain a plurality of knowledge-enhanced texts related to the target to be recommended; Converting a key part of the prompt word and each of the knowledge-enhanced texts into a high-dimensional embedding vector to obtain a first high-dimensional embedding vector corresponding to the prompt word and a second high-dimensional embedding vector corresponding to each of the knowledge-enhanced texts, wherein the key part is used to represent information related to the target to be recommended; Based on the first high-dimensional embedding vector and the second high-dimensional embedding vector, the multiple knowledge-enhanced texts are screened to obtain multiple screened knowledge-enhanced texts; Using the training data set to train a preset recommendation model, obtaining a first embedding vector matrix and a second embedding vector matrix, wherein the first embedding vector matrix is ​​a matrix containing all users and implicit features, and the second embedding vector matrix is ​​a matrix containing all recommended targets and implicit features; Fusing the high-dimensional embedding vector corresponding to each of the screened knowledge-enhanced texts with the embedding vectors related to the target to be recommended in the second embedding vector matrix to obtain a fused feature vector corresponding to each of the screened knowledge-enhanced texts; Scoring each of the screened knowledge enhancement texts based on the fused feature vector corresponding to each of the screened knowledge enhancement texts and the first embedding vector matrix to obtain a scoring result for each of the screened knowledge enhancement texts; According to the scoring results of each of the screened knowledge-enhanced texts, a fine-tuning dataset is constructed, and the large language model is aligned and fine-tuned using the fine-tuning dataset.

2. The large language model alignment and fine-tuning method for a recommendation system according to claim 1, characterized in that: The step of screening the plurality of knowledge-enhanced texts based on the first high-dimensional embedding vector and the second high-dimensional embedding vector to obtain a plurality of screened knowledge-enhanced texts includes: Calculating the similarity between the first high-dimensional embedding vector and the second high-dimensional embedding vector, removing the knowledge enhancement texts whose similarity is less than a preset value from the multiple knowledge enhancement texts, and obtaining remaining knowledge enhancement texts; Clustering the second high-dimensional embedding vectors corresponding to the remaining knowledge-enhanced texts to obtain clustering results; The remaining knowledge enhancement texts are screened according to the clustering results to obtain a plurality of screened knowledge enhancement texts.

3. The large language model alignment and fine-tuning method for a recommendation system according to claim 1, characterized in that: The method of using the training data set to train a preset recommendation model to obtain a first embedding vector matrix and a second embedding vector matrix includes: Build the Bayesian personalized ranking algorithm as a preset recommendation model; The preset recommendation model is trained using the training data set to obtain a first embedding vector matrix and a second embedding vector matrix.

4. The large language model alignment and fine-tuning method for a recommendation system according to claim 1, characterized in that: Scoring each of the screened knowledge-enhanced texts based on the fused feature vector corresponding to each of the screened knowledge-enhanced texts and the first embedding vector matrix to obtain a scoring result for each of the screened knowledge-enhanced texts includes: Inputting the fused feature vector corresponding to each of the screened knowledge-enhanced texts into a multi-layer perceptron to obtain an output result corresponding to each of the screened knowledge-enhanced texts; Performing a dot product of the output result and the first embedding vector matrix to obtain a dot product result corresponding to each of the filtered knowledge-enhanced texts; According to the dot product result, each of the screened knowledge enhancement texts is scored to obtain a scoring result for each of the screened knowledge enhancement texts.

5. The large language model alignment and fine-tuning method for a recommendation system according to claim 1, characterized in that: The step of constructing a fine-tuning dataset according to the scoring results of each of the screened knowledge-enhanced texts includes: Sorting the scoring results of each of the screened knowledge-enhanced texts to obtain sorted scores; According to the ranked scores, the screened knowledge-enhanced texts are combined in pairs to construct a fine-tuning dataset.

6. The large language model alignment and fine-tuning method for a recommendation system according to claim 1, characterized in that: Before using the fine-tuning dataset to align and fine-tune the large language model, the method includes: Normalizing the score of each screened knowledge-enhanced text in the fine-tuning dataset into a probability distribution; According to the probability distribution, calculating information entropy; Calculating a weighting factor according to the information entropy; According to the weighting factors, a target optimization function is constructed when aligning and fine-tuning the large language model.

7. The large language model alignment fine-tuning method for a recommendation system according to claim 6, characterized in that: The step of constructing a target optimization function for aligning and fine-tuning the large language model according to the weighting factor includes: ; in, represents the target optimization function, Indicates the current strategy, represents the reference strategy, represents the fine-tuning dataset, Indicates prompt words, Indicates Knowledge enhancement samples, Indicates Knowledge enhancement samples, represents the weighting factor, represents the logistic function, Represents the temperature coefficient.

8. A large language model alignment and fine-tuning system for a recommendation system, characterized in that: The system comprises: A data acquisition unit, used to acquire prompt words and training data sets corresponding to the target to be recommended; A knowledge-enhanced text acquisition unit, used for inputting the prompt word into a large language model to obtain a plurality of knowledge-enhanced texts related to the target to be recommended; A high-dimensional embedding vector conversion unit, used to convert a key part of the prompt word and each of the knowledge-enhanced texts into a high-dimensional embedding vector, to obtain a first high-dimensional embedding vector corresponding to the prompt word and a second high-dimensional embedding vector corresponding to each of the knowledge-enhanced texts, wherein the key part is used to represent information related to the target to be recommended; A knowledge enhancement text screening unit, configured to screen the plurality of knowledge enhancement texts based on the first high-dimensional embedding vector and the second high-dimensional embedding vector to obtain a plurality of screened knowledge enhancement texts; An embedding vector matrix acquisition unit, used to train a preset recommendation model using the training data set to obtain a first embedding vector matrix and a second embedding vector matrix, wherein the first embedding vector matrix is ​​a matrix containing all users and implicit features, and the second embedding vector matrix is ​​a matrix containing all recommended targets and implicit features; An embedding vector fusion unit, used to fuse the high-dimensional embedding vector corresponding to each of the screened knowledge-enhanced texts with the embedding vectors related to the target to be recommended in the second embedding vector matrix to obtain a fused feature vector corresponding to each of the screened knowledge-enhanced texts; A knowledge enhancement text scoring unit, used for scoring each of the screened knowledge enhancement texts based on the fused feature vector corresponding to each of the screened knowledge enhancement texts and the first embedding vector matrix, to obtain a scoring result for each of the screened knowledge enhancement texts; The large language model alignment and fine-tuning unit is used to construct a fine-tuning dataset according to the scoring results of each of the screened knowledge-enhanced texts, and use the fine-tuning dataset to align and fine-tune the large language model.

9. An electronic device, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the large language model alignment fine-tuning method for the recommendation system as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the large language model alignment and fine-tuning method for a recommendation system according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Trusted text semantic detection method based on data enhancement

    CN117875332A

  • Knowledge recommendation method and device for online education platform, computer equipment and medium

    CN119168819A