Model training sample construction method and device, equipment and readable medium

By using the generative model to generate and evaluate recommendation requests and results, the training samples are constructed, which solves the problem of insufficient training data and improves the training efficiency and performance of the conversational recommendation model.

CN120387023APending Publication Date: 2025-07-29JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410116882.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, the number of sample data used to train the recommended model is insufficient and the cost is high, making it difficult to meet the performance requirements of the conversational recommended model.

Method used

The trained generative model is used to generate recommendation requests and results based on the user's context information, and a training sample is constructed through the evaluation results for training the second recommendation model.

Benefits of technology

While maintaining the quality of training samples, it reduces the difficulty of building sample data, solves the data set size and diversity problems, and improves the training efficiency and performance of the recommended model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387023A_ABST
    Figure CN120387023A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training sample construction method and device, equipment and a medium. The method comprises the following steps: generating a first recommendation request based on context information input of a first user by utilizing a trained generative model; generating a first recommendation result at least based on the first recommendation request by using a trained first recommendation model; processing the first recommendation result by using a trained generative model to generate a first evaluation result at least indicating whether the first recommendation result meets an expected target; if the first evaluation result indicates that the first recommendation result meets the expected target, the first recommendation request and the first recommendation result are constructed into a first training sample to be used for training a second recommendation model, and the trained first recommendation model and the trained second recommendation model are used for recommending the same type of content. According to the scheme, the training sample in the recommendation task can be generated by utilizing the generative model, and the construction difficulty of the training sample is reduced while the quality of the training sample is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to methods, apparatuses, devices, and computer-readable storage media for constructing model training samples. Background Art

[0002] In order to enhance the user's life experience, intelligent recommendation technology has emerged. The intelligent recommendation technology can recommend content that meets the user's expectations based on the reference information provided by the user.

[0003] With the development of computer technology, recommendation solutions for providing recommended content to users based on recommendation models have also become increasingly mature. Generally, the performance of a recommendation model is directly related to the quantity and quality of the sample data used when training the recommendation model. Traditionally, the sample data used to train a recommendation model often comes from various publicly available offline data sets (i.e., offline data) and / or online data in actual business scenarios. However, such sample data often has problems such as insufficient quantity and high cost. There is a desire to conveniently and quickly obtain high-quality sample data. Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for constructing model training samples is provided. The method includes: using a trained generative model to generate a first recommendation request based on the context information input of a first user; using a trained first recommendation model to generate a first recommendation result based at least on the first recommendation request; using the trained generative model to generate a first evaluation result for the first recommendation result, the first evaluation result at least indicating whether the first recommendation result meets the expected goal; and if the first evaluation result indicates that the first recommendation result meets the expected goal, constructing the first recommendation request and the first recommendation result as a first training sample for training a second recommendation model, where the trained first recommendation model and the second recommendation model are used to recommend the same type of content.

[0005] In a second aspect of the present disclosure, there is provided an apparatus for constructing model training samples. The apparatus includes: a first generation module configured to generate a first recommendation request based on the context information input of a first user by using a trained generative model; a second generation module configured to generate a first recommendation result at least based on the first recommendation request by using a trained first recommendation model; a third generation module configured to generate a first evaluation result for the first recommendation result based on the first recommendation result by using a trained generative model, the first evaluation result at least indicating whether the first recommendation result meets the expected goal; and a training module configured to, if the first evaluation result indicates that the first recommendation result meets the expected goal, construct the first recommendation request and the first recommendation result into a first training sample for training a second recommendation model, and the trained first recommendation model and the second recommendation model are used to recommend the same type of content.

[0006] In a third aspect of the present disclosure, there is provided an electronic device. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to execute the method of the first aspect of the present disclosure.

[0007] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and it can be executed by a processor to execute the method according to the first aspect of the present disclosure.

[0008] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the following, in conjunction with the drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various implementations of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0010] Figure 1 A schematic diagram showing an example environment in which the embodiments of the present disclosure can be implemented;

[0011] Figure 2 A schematic diagram showing an architecture for training a recommendation model according to some embodiments of the present disclosure;

[0012] Figure 3A A block diagram showing a process of constructing model training samples according to some embodiments of the present disclosure;

[0013] Figure 3B A block diagram showing a process of generating a recommendation request according to some embodiments of the present disclosure;

[0014] Figure 3C A block diagram showing a process of generating a recommendation result according to some embodiments of the present disclosure;

[0015] Figure 4 A flowchart showing a process of constructing a model training sample according to some embodiments of the present disclosure;

[0016] Figure 5 A block diagram showing an apparatus for constructing a model training sample according to some embodiments of the present disclosure; and

[0017] Figure 6 A block diagram showing an electronic device in which one or more embodiments of the present disclosure can be implemented. Detailed Description of Specific Embodiments

[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0019] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0020] As used herein, the term "model" can learn the corresponding association between input and output from training data, so that after training is completed, for a given input, a corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In this article, the "model" can also be referred to as a "machine learning model", a "machine learning network", a "neural network", or a "network", and these terms can be used interchangeably in this article.

[0021] "Neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and generally includes an input layer, an output layer, and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications usually include many hidden layers, thus increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.

[0022] Generally, machine learning can roughly include three stages, namely, a training stage, a testing stage, and an application stage (also called an inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, so as to determine the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values obtained from training and determine the corresponding output.

[0023] It should be noted that in the technical solutions of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0024] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to relevant laws and regulations.

[0025] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information, so that the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.

[0026] As an optional but non-limiting implementation manner, in response to receiving an active request from a user, the manner of sending a prompt message to the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0027] Example Environment and Basic Working Principle

[0028] For a machine learning model, the performance of the model is highly correlated with the training data. Considering the problem of lack of training sample data in the conventional solution when training a conversational recommendation model (hereinafter referred to as the recommendation model). Thus, an embodiment of the present disclosure provides a method for constructing a model training sample. According to an embodiment of the present disclosure, a trained generative model is used to generate a first recommendation request based on the context information input of a first user; a trained first recommendation model is used to generate a first recommendation result based at least on the first recommendation request; the trained generative model is used to process the first recommendation result to generate a first evaluation result at least indicating whether the first recommendation result meets the expected goal; and if the first evaluation result indicates that the first recommendation result meets the expected goal, the first recommendation request and the first recommendation result are constructed as a first training sample for training a second recommendation model for performing the same type of recommendation task as the trained first recommendation model. Thus, by using the generative model to simulate and act as a real user based on the context information of the user to generate training samples for training the recommendation model, the construction difficulty of the training samples can be reduced while maintaining the quality of the training samples.

[0029] Example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0030] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In Figure 1 environment 100, it is desired to train and use such a machine learning model (i.e., model 130), which is configured for a variety of application environments, for example, based on request information entered by a user, recommending corresponding content, etc. For example, when model 130 is a conversational movie recommendation model, movies can be recommended to the user according to the basic information, movie preferences, etc. provided by the user.

[0031] As Figure 1 shown, environment 100 includes a model training system 150. Figure 1The upper part shows the process of the model training phase, and the lower part shows the process of the model application phase. Before training, the parameter values of model 130 can have initial values or can have pre-trained parameter values obtained through a pre-training process. The model 130 can be trained via forward propagation and backpropagation, and during the training process, the parameter values of the model 130 can be updated and adjusted. After training is completed, model 130' can be obtained. At this time, the parameter values of model 130' have been updated, and based on the updated parameter values, model 130 can be used to implement a recommendation task during the model application phase.

[0032] During the model training phase, based on a training data set 110 including a plurality of training samples 112, the model 130 can be trained using a model training system 150. Here, each training data 112 can be in a binary tuple format and include a recommendation request 120 and a recommendation result 122. At this time, the model 130 can be trained using the training samples 112 including the recommendation request 120 and the recommendation result 122. Specifically, the training process can be iteratively executed using a large number of training samples. After training is completed, the model 130 can have the ability to execute a task to be processed. During the model application phase, the model 130' (at this time, the model 130' has trained parameter values) can be used to execute the corresponding task. For example, a recommendation request 142 to be processed by the model can be received, and a corresponding recommendation result 144 can be output.

[0033] In Figure 1 this context, the model training system 150 can include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. The terminal device can refer to any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in a cloud environment, and so on.

[0034] It should be understood that Figure 1 the components and arrangements in the illustrated environment 100 are merely examples, and the computing systems suitable for implementing the exemplary implementations described in the present disclosure can include one or more different components, other components, and / or different arrangements. The implementations of the present disclosure are not limited in this regard.

[0035] As described above, the performance of the recommendation model is directly related to the quantity and quality of the sample data used when training the recommendation model. For example, a recommendation model trained using a smaller quantity of sample data may have performance inferior to that of a recommendation model trained using a larger quantity of sample data.

[0036] A Conversational Recommender System, also known as a Conversational Recommendation Model, is a personalized recommendation system based on a dialogue interaction mode that has emerged in recent years. Different from traditional recommendation systems, it can not only recommend products or services to users based on their historical behaviors and interests, but also obtain real-time feedback information from users through real-time conversations with them, so as to more accurately understand users' needs and preferences.

[0037] The conversational recommendation system can be divided into three modules: an intent understanding module, a recommendation system module, and a reinforcement learning network module. The main functions of the three modules are as follows: (1) The main function of the intent understanding module is to understand the user's intent and transmit the result of the user intent understanding to the reinforcement learning network module; (2) The reinforcement learning module receives the result of the user intent understanding and determines the best strategy for the next step. For example, it may be to invoke the recommendation system module to recommend products or services, or it may be to continue to interact with the user to obtain more information; (3) The function of the recommendation system module is to run a personalized recommendation algorithm based on the input information and return the recommendation result to the user.

[0038] Traditionally, the sample data used to train the recommendation model often comes from various publicly available offline datasets (i.e., offline data) and / or online data in actual business scenarios (such as cutting into online traffic in actual business scenarios to conduct a gray-scale experiment on conversational recommendation to accumulate data). However, such sample data has certain problems. Some problems with the offline dataset include: (1) The dataset size is too small; (2) The dataset is relatively regular and difficult to reflect the actual situation of the real world; (3) The dataset is produced by a limited number of researchers or annotators and cannot adapt to the user diversity in actual business. Problems with the online dataset: (1) The gray-scale experiment of cutting into traffic will have a certain impact on the actual business, the experimental cost is very high, the proportion of gray-scale needs to be strictly controlled, and it is very likely to cause problems such as customer complaints; (2) The gray-scale experiment requires the active participation and interaction of users to accumulate data, and a low user participation rate will further lead to data sparsity. As a result, it is difficult to train an ideal effect for the conversational recommendation model, and the performance of the model is difficult to meet the requirements.

[0039] In response to this, the present disclosure provides an improved solution for constructing model training samples. For ease of understanding, a specific scenario can be used for exemplary illustration. For example, in the movie recommendation scenario. In the movie recommendation scenario, a recommendation model for recommending movies to users can be based on the context information provided by the users as a reference (for example, the historical movie viewing records, historical movie viewing comments, etc. provided by the users for reference), and recommend movies that the users "may be interested in" according to the instructions of the users. When training such a recommendation model, the model training system 150 can use a trained generative model to simulate the decision-making of the users based on the context information of the users, or act as the users, in order to obtain a recommendation request for requesting the recommendation model to recommend movies.

[0040] The trained generative model can be a language model (abbreviated as LM), which can output additional content based on the input content. The trained generative model can be a multimodal model, which can at least process text information, such as generating output text based on the input text. The trained generative model can also be a large language model (abbreviated as LLM), which can have a certain degree of "reasoning" ability after appropriate instruction tuning and chain of thought excitation when the relevant parameter scale reaches a certain order of magnitude. Further, the model training system 150 uses a trained recommendation model (for ease of description, it is referred to as the trained first recommendation model) to generate a recommendation result based at least on the recommendation request. For example, when the recommendation request 120 is "Based on the movies B, C, and D that user A has watched, recommend movies with a warm theme for user A", the corresponding recommendation result 122 can be a movie with a "warm theme" recommended for "user A" (for example, the specific movie name "Movie A").

[0041] Further, the model training system 150 uses the above-mentioned trained generative model again to generate an evaluation result for the recommendation result 122. For example, the evaluation result can indicate whether the recommendation result meets the expected goal. For example, in the above example, the trained generative model can determine whether the recommendation result meets the expected goal based on the degree of match between "Movie A" and the semantic recognition result of the recommendation request 120. Or, it is judged whether "Movie A" is a movie suitable for "user A who has watched movies B, C, and D", whether it is a movie with a "warm theme", etc.

[0042] If the trained generative model indicates in the evaluation result that the recommendation result meets the expected goal, the model training system 150 may choose to construct the recommendation request 120 and the recommendation result 122 into a training sample pair to train the recommendation model that has not been completed (for convenience of description, such a recommendation model may be referred to as the second recommendation model. For example, the second recommendation model may be the above-mentioned model 130). In the embodiments of the present disclosure, the trained first recommendation model and the second recommendation model may be models for recommending the same type of content. For example, both of them are models that can recommend movies based on the movie preference information provided by the user. In addition, although the trained first recommendation model and the second recommendation model may be for recommending the same type of content, the present disclosure does not limit that the trained first recommendation model and the second recommendation model have the same model performance for recommending the same type of content. Or rather, their model performances may be different. For example, the trained first model is a classic recommendation model (for example, a recommendation model for content recommendation based on text processing), while the second recommendation model may be a conversational recommendation model.

[0043] The trained first recommendation model and the second recommendation model can both be used to perform the recommendation task. For example, the trained first recommendation model and the second recommendation model can both perform the movie recommendation task. In some embodiments, the first recommendation model and the second recommendation model may be the same model or different models. For example, the first recommendation model may be a non-conversational recommendation model, and the second recommendation model may be a conversational recommendation model. In the embodiments of the present disclosure, the trained first recommendation model and the second recommendation model are used to recommend the same type of content. In other words, the model output of the trained first recommendation model and the model output of the second recommendation model can indicate the recommendation results of the same type of recommendation task.

[0044] Thus, the model training system 150 can use the trained generative model to simulate the behavior of real users based on the reference information provided by the user to request recommended movies and the recommended movies. The model training system 150 can also use the trained generative model to analyze the recommended results to determine whether the recommended needs of the "simulated real users" are met. In this way, the inference ability of the trained generative model can be used to efficiently obtain training samples similar to those provided by real users to train the second recommendation model. This makes it possible to generate any specified number of training data for any user, solving the problem of the scale of the data set. At the same time, this solution applies the understanding and inference ability of the generative mode (such as LLM), can use real user information as prompt data to guide the interaction of the generative mode, introduces user diversity, and solves the problem of the regularity of the data set to a certain extent. In addition, by applying this solution to collect data, there is no need to conduct a gray-scale experiment on the online cut-in traffic. After the recommendation model is launched, it is still possible to continuously train, improve performance, and reduce the data collection cost.

[0045] For ease of understanding in the following, the example of movie recommendation is also used for explanation. It can be understood that only the movie recommendation scenario is used as an example for exemplary description, and this solution can also be applied to any other appropriate recommendation scenarios.

[0046] Some example embodiments of the present disclosure will be further described below with reference to the accompanying drawings.

[0047] In some embodiments, the model training system 150 can also be composed of a construction part (such as a sample construction subsystem) for constructing training samples (for example, for constructing training sample 112) and a training part (such as a training subsystem) for training the model, thereby realizing the separation of the sample construction and model training links. For example, the model training system 150 can use the construction part therein to use the trained generative model and the first recommendation model to generate training samples in the recommendation task, and use the training part therein to train the second recommendation model based on the training samples. For example, the model training system 150 can be composed of a group of devices, a part of the group of devices is used as the above construction part, and another part of the group of devices is used as the above training part. In some embodiments, the model training system 150 can also use the same device to construct training samples and train the second recommendation model based on the training samples. The present disclosure is not intended to be limited thereto.

[0048] For ease of understanding in the following, it is exemplified that the model training system 150 uses the same device to construct training samples and train the second recommendation model based on the training samples.

[0049] Figure 2A schematic diagram of an architecture 200 for training a recommendation model according to some embodiments of the present disclosure is shown. For ease of discussion, the architecture 200 will be described with reference to Figure 1 the environment 100.

[0050] In an embodiment of the present disclosure, the model training system 150 may pre-maintain context information of a user. In some embodiments, the context information of the user may be obtained in advance through interaction with a real user. For example, the real user is requested to provide relevant data for training the recommendation model. The context information of the user may include, for example, but is not limited to, basic information of the user (e.g., gender of the user), preference information (e.g., specific movie genres liked, such as horror, action, warm-hearted, etc.), and historical behavior of the user (e.g., movies previously watched, historical evaluations filled in for movies, etc.). In some embodiments, the context information of the user may be determined based on the historical context information of historical users associated with the trained first recommendation model 220. For example, in a case where the historical context information of historical user A (e.g., the movie viewing record of historical user A is having watched movies B, C, D, and preferring action movies) has been used to train an untrained first model to obtain the (trained) first model, the historical context information of this historical user A may be collected. Thus, it is possible to use the real data that has previously trained the first model to construct training samples to improve the quality of the training samples. In some embodiments, historical user A may be a user who has used the trained first recommendation model 220 online, thereby providing more real user information.

[0051] Exemplarily, reference may be made to Figure 2 In Figure 2 , the model training system 150 may use the pre-maintained historical context information 243 of historical users to generate a recommendation request.

[0052] It should be understood that in the process of constructing the model training samples provided by the present disclosure, context information of different users may be used (e.g., constructing multiple different training samples). For ease of understanding, an illustration will first be given for one round of constructing a training sample. Correspondingly, the context information of the user used in this round (or, the current round) may be referred to as the context information of the first user, for distinguishing the context information of, for example, the second user used in other rounds. Similarly, when subsequently elaborating on the recommendation request, recommendation result, evaluation result, and training sample involved in this current round, they are also expressed as the first recommendation request, first recommendation result, first evaluation result, and first training sample, respectively, to indicate that these contents are generated and / or used in this current round.

[0053] Continuing to refer to Figure 2, the model training system 150 can utilize the trained generative model 210 to generate a recommendation request (e.g., the first recommendation request) based on the historical context 243 information of historical users.

[0054] In some embodiments, an input template can be configured for the trained generative model 210 to quickly construct an instruction for indicating the actions of the trained generative model 210 by using the input template. In some embodiments, different input templates can be configured for different application scenarios (e.g., a first input template can be configured for the scenario of generating recommendation results, and a second input template can be configured for the scenario of generating evaluation results) to further exert the value of the input template. This will be described in detail later.

[0055] Exemplarily, in Figure 2 , a first input template set 241 and a second input template 242 are also configured for the model training system 150. In some embodiments, the number of input templates configured in the first input template set 241 and the second input template set 242 can be the same.

[0056] Furthermore, after the model training system 150 obtains a recommendation request from the trained generative model 210, it can utilize the trained first recommendation model 220 to process the recommendation request to generate a recommendation result.

[0057] Furthermore, the model training system 150 utilizes the trained generative model 210 again to process the recommendation result to generate an evaluation result. As described above, if the evaluation result indicates that the recommendation result meets the expected goal, it is determined as part of the training sample set 110 for the user to train the second recommendation model 230.

[0058] The following will be combined with Figure 3A to describe the process of constructing the model training sample in detail. Figure 3A FIG. shows a block diagram of a process 300A for constructing a model training sample according to some embodiments of the present disclosure. The process 300A can be implemented at the model training system 150. For ease of discussion, the environment 100 of Figure 1 and the architecture 200 of Figure 2 will be referred to to describe the process 300A.

[0059] In block 310, the model training system 150 utilizes the trained generative model to generate a recommendation request based on the input of the user's context information. In the embodiments of the present disclosure, the model training system 150 can provide the user's context information (e.g., the historical context information 243 of historical users) to the trained generative model 210 to utilize the trained generative model 210 to generate a recommendation request (e.g., the first recommendation request) based on the input of the user's context information.

[0060] In some embodiments, the model training system 150 may utilize a first set of input templates to generate an indication for achieving the above object. Specifically, in the model training system 150, a first set of input templates 241 may be configured for the scenario of generating a recommendation request. The first set of input templates 241 may include multiple input templates (for the convenience of discussion, referred to as the first input templates). The input templates may also be referred to as input prompts. The first input templates may be used to define the description format of the user's context information and the generation indication for the recommendation request. For example, the first input template may be "Please act as User A of a movie recommendation system. Your portrait preference information is [K1], and your behavior sequence in a recent period is [K2]. Now you come to the movie recommendation system again. Please interact reasonably and put forward your demands for this movie recommendation." Thus, when the model training system 150 obtains the user's context information (for example, the context information of the first user), it can quickly and standardly determine the first model input by directly inserting the context information into the first input template (for example, inserting the content in the determined context information of the first user into at least one of the positions of "K1" and "K2") so as to use the first model input to instruct the trained generative model 210 to generate a recommendation request. For example, generate a first recommendation request.

[0061] In this regard, reference may also be made exemplarily to Figure 3B . Figure 3B FIG. shows a block diagram of a process 300B for generating a recommendation request according to some embodiments of the present disclosure. The process 300B is an exemplary embodiment of block 310 of the process 300A.

[0062] In Figure 3B , at block 311, the model training system 150 may select a first input template (for example, randomly select one from the first set of input templates 241) and the context information of a user from the first set of input templates 241. For example, select the context information of a user (i.e., the context information of the first user) from the historical context information 243 of the historical user.

[0063] At block 312, the model training system 150 may fill the context information of the user determined in block 311 above into the selected first input model to obtain a first model input. For example, insert movies B, C, and D that User A has watched into the above "K2" to obtain a first model input. For example, the first input template may be "Please act as User A of a movie recommendation system. You have watched movies B, C, and D. Now you come to the movie recommendation system again. Please interact reasonably and put forward your demands for this movie recommendation."

[0064] At block 313, the model training system 150 provides a first model input to the trained generative model 210. Specifically, the model training system 150 provides the first model input to the trained generative model 210, and the trained generative model 210 generates a recommendation request based on the first model input. For example, the trained generative model 210 can generate a first recommendation request according to the first model input constructed based on the context information of the first user. For example, the first recommendation request can be "I am User A. I have watched Movie B, Movie C, and Movie D. Please recommend a good movie for me."

[0065] At block 314, the model training system 150 obtains the recommendation request generated by the trained generative model 210. The obtained recommendation request is used to execute subsequent block 320 in process 300A.

[0066] At block 320, the model training system 150 generates a recommendation result using the trained first recommendation model 220, at least based on the recommendation request. For example, the model training system 150 generates a first recommendation result using the trained first recommendation model 220, at least based on the first recommendation request. In an embodiment of the present disclosure, the trained first recommendation model 220 can process the first recommendation request based on the performance obtained from its pre-training to obtain the corresponding first recommendation result. For example, based on the movies B, C, and D that User A has watched, Movie I is recommended for User A. In some embodiments, in order to improve the processing effect of the trained first recommendation model 220, the trained intent recognition model can also be used to identify the recommendation intent of the first recommendation request, so as to improve the processing ability of the first recommendation model 220 for the first recommendation request.

[0067] Exemplarily, reference can be made to Figure 3C . Figure 3C FIG. shows a block diagram of process 300C for generating a recommendation result according to some embodiments of the present disclosure. Process 300C is an exemplary embodiment of block 320 of process 300A. In process 300, the trained first recommendation model 220 is exemplarily used as the trained recommendation model.

[0068] At block 321, the model training system 150 utilizes the trained intent recognition model to determine the recommendation intent in the recommendation request. For example, the model training system 150 can utilize the trained intent recognition model to determine the recommendation intent in the first recommendation request. Specifically, the trained intent recognition model can recognize the intent for movie attributes in the expression of the LLM. In some embodiments, the movie attributes can be the theme, era, genre, actors, director, etc. of the movie. Thus, the recommendation intent generated using the intent recognition model can strengthen the inclination, recommendation intent, etc. for movie attributes in the first recommendation request output by the LLM, so that the trained first recommendation model 220 can make better recommendations.

[0069] At block 322, the model training system 150 can provide the recommendation intent output by block 322 and the user's context information to the trained first recommendation model 220. For example, provide the recommendation intent and the context of the first user to the trained first recommendation model 220, thereby generating the corresponding recommendation result.

[0070] At block 323, the model training system 150 can obtain the recommendation result output by the trained first recommendation model 220. For example, the model training system 150 can obtain the first recommendation result output by the trained first recommendation model 220 when processing the first recommendation request. The obtained recommendation result is used to execute the subsequent block 330 in process 300A.

[0071] At block 330, the model training system 150 utilizes the trained generative model 210 to generate an evaluation result for the recommendation result based on the recommendation result. For example, the first evaluation result for the first recommendation result can be generated based on the first recommendation result. In the embodiments of the present disclosure, the evaluation result at least indicates whether the recommendation result meets the expected goal. Taking the first evaluation result as an example, the first evaluation result can at least indicate whether the specific movie provided in the first recommendation result can meet the requirements of "User A". Generally, the trained generative model 210 can determine whether the content provided in the first evaluation result can meet the expected goal through, for example, feature comparison (such as comparison with the features of the expected movie, whether the feature similarity with the reference target movie exceeds the similarity threshold). The trained generative model 210 can generate the first evaluation result that at least indicates whether the first recommendation result meets the expected goal based on the judgment on the local machine, and return the first evaluation result to the model training system 150. The first evaluation result can include, for example, any appropriate evaluation result such as a score, rating, star rating, evaluation text, etc. for the first recommendation result. For example, the trained generative model 210 can return the text information "Satisfied" or a score of 8 points (for example, out of 10 points) as the first evaluation result to the model training system 150 to indicate that the first recommendation result can meet the expected goal.

[0072] In some embodiments, the model training system 150 may also maintain a second input template, or a second indicator word, for the scenario of instructing the trained generative model 210 to generate an evaluation result. The second input template is at least used to define the description format of the recommendation result and the generation instruction of the evaluation result. Thus, it can conveniently and accurately instruct the trained generative model 210 to generate a first evaluation result for the first recommendation result based on the first recommendation result. Correspondingly, in the case where the second input template is maintained, the model training system 150 may fill the first recommendation result into the second input template for the trained generative model to obtain a second model input. Similar to the first input template, the model training system 150 may provide the second model input to the trained generative model 210 to instruct it to process the first recommendation result to generate a first evaluation result, and after the processing is completed, obtain the first evaluation result output by the trained generative model 210.

[0073] In block 340, the model training system 150 determines whether the evaluation result indicates that the recommendation result meets the expected goal. If it does, the model training system 150 may choose to execute block 350. For example, it may be determined whether the first recommendation result meets the expected goal based on the first evaluation result.

[0074] In block 350, the model training system 150 may construct the recommendation request and the recommendation result into a training sample for training the second recommendation model 230. For example, the first recommendation request and the first recommendation structure may be constructed into a pair of training samples in the training sample set 110. The model training system 150 may use the first recommendation request as the input of the second recommendation model 230 and the first recommendation result as the output of the second recommendation model 230 to train the second recommendation model 230.

[0075] In block 370, the model training system 150 determines whether the number of training samples has reached a threshold. Or rather, the model training system 150 determines whether the number of statistically determined training samples has reached a threshold. For example, in the case where the current round is the first time to execute block 350, after block 350 is executed in this round, the number of training samples is 1. In subsequent rounds, when another pair of recommendation requests and recommendation results are used to complete a training, the number of training samples will be accumulated to 2. The model training system 150 may determine in block 370 whether the currently accumulated number of training samples has reached a threshold (such as the estimated number of training samples required to complete the training of the second model 230, a preset number threshold, etc.). If it has reached, the model training system 150 chooses to jump out to end the training of the second recommendation model 230.

[0076] In some embodiments, if the condition is not met, the model training system 150 may jump to the above box 310. The model training system 150 may reuse the trained generative model 210 to generate a new recommendation request (e.g., a second recommendation request relative to the first recommendation request) based on the context information input of the second user for iterative training. Specifically, the model training system 150 may start the next round of training at least by updating the context information of the user used in the current round (e.g., the context information of the first user). For example, the model training system 150 may use the context information of a second user different from the first user to generate a new recommendation request (described as the second recommendation request for convenience of description). Further, the model training system 150 may, in a similar manner as the processing for the "first recommendation request", at least execute box 320 based on the "second recommendation request" to use the trained first recommendation model to generate a third recommendation result based on at least the second recommendation request; execute box 330 again to use the trained generative model to generate a third evaluation result for the third recommendation result based on the third recommendation result; execute box 350 again to, in the case where the third evaluation result indicates that the third recommendation result meets the expected goal, construct the second recommendation request and the third recommendation result as a second training sample for implementing the next round of training of the second recommendation model. Thus, multiple training samples can be continuously determined in a cyclic manner to implement the iterative training of the second recommendation model 230 for a preset number of rounds.

[0077] In some embodiments, after the model training system 150 chooses to jump out, the model training system 150 may also store the training samples that have been used. For example, it stores the training sample set 110 for subsequent use, such as for other recommendation models, for example, a recommendation model that performs the same type of recommendation task as the second recommendation model 230.

[0078] In some embodiments, if the first evaluation result indicates that the first recommendation result does not meet the expected goal, that is, the first recommendation result obtained in this round may not be able to form a first training sample with the first recommendation request. In this case, the model training system 150 can use the trained first recommendation model 210 to regenerate a recommendation result (for convenience of description, it can be referred to as the second recommendation result) at least based on the first recommendation request. For example, in block 340, if the model training system 150 is in the case where the first evaluation result indicates that the first recommendation result does not meet the expected goal, it can jump to block 320 to regenerate the second recommendation result at least based on the above first recommendation request. Similarly, after generating the second recommendation result, the model training system 150 can execute block 330 again to use the trained generative model 210 to generate a second evaluation result for the second recommendation result based on the second recommendation result. Correspondingly, the model training system 150 can execute block 340 again to determine that the second recommendation result meets the expected goal based on the second evaluation result.

[0079] Similarly, if the second evaluation result indicates that the second recommendation result meets the expected goal, the model training system 150 can execute block 350 to construct the first recommendation request and the second recommendation result as a training sample for training the second recommendation model.

[0080] It should be understood that if the second evaluation result still indicates that the second recommendation result does not meet the expected goal, the model training system 150 can continue to generate multiple evaluation results (such as the third evaluation result, the fourth evaluation result, etc.) in a manner similar to generating the second evaluation result. Thus, it can be determined whether the corresponding third recommendation result, fourth recommendation result, etc. can be constructed as a training sample with the first recommendation request. In this way, the model training system 150 can cyclically (or continuously) regenerate recommendation results when the recommendation results do not meet the requirements to find a recommendation result that matches the first recommendation request as much as possible.

[0081] In some embodiments, to avoid problems such as jamming and resource waste caused by, for example, searching for recommendation results multiple times for the same first recommendation request. In some embodiments, block 360 may further be included in 300A. In block 360, the model training system 150 may determine whether to continue searching for a more suitable recommendation result for the first recommendation request based on whether the number of rounds of recommendation results already generated for the recommendation request (e.g., the first recommendation request) exceeds a threshold. For example, based on the performance of the trained first recommendation model 220, for example, after a specific number of rounds, the results generated by the trained first recommendation model 220 no longer change significantly, then this specific number of rounds may be determined as the threshold. The model training system 150 may choose to directly select the currently obtained recommendation result (e.g., when reaching the fourth round, the fourth recommendation result is obtained) to construct a sample pair with the first recommendation request in the case where the number of rounds of generated recommendation results exceeds the threshold round, so as to execute block 350 and implement the training of the second recommendation model 230. In some embodiments, if the trained first recommendation model 220 also refers to the recommendation intention when outputting the recommendation result, the model training system 150 may execute block 323 after executing block 360 to achieve the purpose of providing the recommendation result.

[0082] In some embodiments, the first evaluation result also indicates the reason why the first recommendation result does not meet the expected goal. Specifically, the model training system 150 may further instruct the trained generative model 210 to provide the corresponding reason in the case where the first evaluation result indicates that the first recommendation result does not meet the expected goal. For example, the trained generative model 210 believes that in the first recommendation result, the movie attributes that specifically cause the failure to meet the expected goal (e.g., genre mismatch, era does not meet the requirements, etc.). Thereby, the model training system 150 can prompt and provide references for the less valuable part of the first recommendation result, so as to optimize the first recommendation result and / or the trained first recommendation model 220.

[0083] In this case, in some embodiments, if the second input template set 242 is configured in the model training system 150, the second input template in the second input template set 242 may further include a reason generation instruction. The reason generation instruction is used to instruct the generative model to output the reason why the recommendation result does not meet the expected goal. Thereby, in, for example, the execution stage of block 330, the model training system 150 can also make a request to the trained generative model 210 to directly provide the corresponding reason in the case where the first evaluation result indicates that the first recommendation result does not meet the expected goal. Thereby improving the interaction efficiency. For example, the text form of the second input template may be: "If you are satisfied, please reply'satisfied'. If you are not satisfied, please explain the reason. When explaining the reason for dissatisfaction, please express it based on the attributes of the movie, such as not liking the genre, era, type, actor, director, etc."

[0084] Accordingly, embodiments of the present disclosure can collect conversational recommendation training data and generate any specified number of training data based on the user information of any user, which can effectively solve the problem of the scale of the data set. Embodiments of the present disclosure apply the understanding and reasoning capabilities of a generative model (such as an LLM), can use real user information as prompt data to guide the interaction of the generative model, introduce user diversity, and reduce the regularity of the data set. In addition, during the data collection process, there is no need to additionally cut into online traffic for gray-scale experiments, which has no impact on the online environment and reduces the data collection cost.

[0085] Subsequently, according to an embodiment of the present disclosure, a first recommendation request is generated based on the context information input of a first user by using a trained generative model; a first recommendation result is generated by using a trained first recommendation model based at least on the first recommendation request; a first evaluation result for the first recommendation result is generated by using the trained generative model based on the first recommendation result, and the first evaluation result at least indicates whether the first recommendation result meets the expected goal; and if the first evaluation result indicates that the first recommendation result meets the expected goal, the first recommendation request and the first recommendation result are constructed as a first training sample for training a second recommendation model, and the trained first recommendation model and the second recommendation model have at least the same type of model output. Thus, training samples in a recommendation task can be generated by using a generative model, while maintaining the quality of the training samples, reducing the difficulty of constructing the training samples.

[0086] Example Process

[0087] Figure 4 FIG. shows a flowchart of a process 400 for constructing a model training sample according to some embodiments of the present disclosure. The process 400 can be implemented at the model training system 150.

[0088] At block 410, the model training system 150 uses a trained generative model to generate a first recommendation request based on the context information input of a first user.

[0089] At block 420, the model training system 150 uses a trained first recommendation model to generate a first recommendation result based at least on the first recommendation request.

[0090] At block 430, the model training system 150 uses the generative model to generate a first evaluation result for the first recommendation content based on the first recommendation content. In an embodiment of the present disclosure, the first evaluation result at least indicates whether the first recommendation result meets the expected goal.

[0091] At block 440, the model training system 150 determines whether the first evaluation result at least indicates that the first recommendation result meets the expected goal. If so, block 450 is executed.

[0092] At block 450, the model training system 150 constructs a first recommendation request and first recommendation content into a training sample for training a second recommendation model. In an embodiment of the present disclosure, the trained first recommendation model and the second recommendation model are used to recommend the same type of content.

[0093] In some embodiments, generating the first recommendation request based on the user's context information input includes: populating the context information of a first user into a first input template for a trained generative model to obtain a first model input, where the first input template is used to define a description format of the user's context information and an indication for generating a recommendation request; providing the first model input to the trained generative model; and obtaining the first recommendation request output by the trained generative model.

[0094] In some embodiments, the context information of the first user is determined based on the historical context information of historical users associated with the trained first recommendation model.

[0095] In some embodiments, generating the first recommendation result based at least on the first recommendation request includes: using a trained intent recognition model to determine the recommendation intent in the first recommendation request; and providing the context information of the first user and the recommendation intent to the trained first recommendation model; obtaining the first recommendation result output by the trained first recommendation model.

[0096] In some embodiments, generating a first evaluation result for the first recommendation result based on the first recommendation result includes: populating the first recommendation result into a second input template for a trained generative model to obtain a second model input, where the second input template is at least used to define a description format of the recommendation result and an indication for generating an evaluation result; providing the second model input to the trained generative model; and obtaining the first evaluation result output by the trained generative model.

[0097] In some embodiments, the second input template further includes a reason generation indication, which is used to indicate the reason why the recommendation result output by the generative model does not meet the expected target.

[0098] In some embodiments, if it is determined at block 440 that the first evaluation result indicates that the first recommendation result does not meet the expected target, blocks 420 and 430 can be repeatedly executed. In this case, process 400 further includes: using the trained first recommendation model to regenerate a second recommendation result based at least on the first recommendation request; using the trained generative model to generate a second evaluation result for the second recommendation result based on the second recommendation result; and if the second evaluation result indicates that the second recommendation result meets the expected target, constructing the first recommendation request and the second recommendation result into a training sample for training the second recommendation model.

[0099] In some embodiments, the first evaluation result further indicates the reason why the first recommendation result fails to meet the expected goal.

[0100] In some embodiments, process 400 further includes: using a trained generative model to generate a second recommendation request based on the context information input of a second user; using a trained first recommendation model to generate a third recommendation result based at least on the second recommendation request; using a trained generative model to generate a third evaluation result for the third recommendation result based on the third recommendation result; and if the third evaluation result indicates that the third recommendation result meets the expected goal, constructing the second recommendation request and the third recommendation result as a second training sample for training a second recommendation model.

[0101] Example Device

[0102] Figure 5 The block diagram of an apparatus 500 for constructing a model training sample according to some embodiments of the present disclosure is shown. The apparatus 500 may be implemented as or included in a model training system 150.

[0103] The apparatus 500 includes a first generation module 510 configured to generate a first recommendation request based on the context information input of a first user using a trained generative model. The apparatus 500 further includes: a second generation module 520 configured to generate a first recommendation result based at least on the first recommendation request using a trained first recommendation model. The apparatus 500 further includes: a third generation module 530 configured to generate a first evaluation result for the first recommendation result based on the first recommendation result using a trained generative model, the first evaluation result at least indicating whether the first recommendation result meets the expected goal. The apparatus 500 further includes: a training module 540 configured to, if the first evaluation result indicates that the first recommendation result meets the expected goal, construct the first recommendation request and the first recommendation result as a first training sample for training a second recommendation model, and the trained first recommendation model and the second recommendation model are used to recommend the same type of content.

[0104] In some embodiments, generating the first recommendation request based on the context information input of a user includes: filling the context information of the first user into a first input template for the trained generative model to obtain a first model input, the first input template being used to define the description format of the context information of the user and the generation indication of the recommendation request; providing the first model input to the trained generative model; and obtaining the first recommendation request output by the trained generative model.

[0105] In some embodiments, the context information of the first user is determined based on the historical context information of historical users associated with the trained first recommendation model.

[0106] In some embodiments, generating a first recommendation result based at least on a first recommendation request includes: using a trained intent recognition model to determine a recommendation intent in the first recommendation request; and providing context information of the first user and the recommendation intent to a trained first recommendation model; obtaining a first recommendation result output by the trained first recommendation model.

[0107] In some embodiments, generating a first evaluation result for the first recommendation result based on the first recommendation result includes: filling the first recommendation result into a second input template for a trained generative model to obtain a second model input, where the second input template is at least used to define a description format of the recommendation result and a generation indication of the evaluation result; providing the second model input to the trained generative model; and obtaining a first evaluation result output by the trained generative model.

[0108] In some embodiments, the second input template further includes a reason generation indication, and the reason generation indication is used to indicate a reason why the recommendation result output by the generative model does not meet the expected goal.

[0109] In some embodiments, the apparatus 500 further includes: a repeated execution module configured to, if the first evaluation result indicates that the first recommendation result does not meet the expected goal, use the trained first recommendation model to regenerate a second recommendation result based at least on the first recommendation request; use the trained generative model to generate a second evaluation result for the second recommendation result based on the second recommendation result; and if the second evaluation result indicates that the second recommendation result meets the expected goal, construct the first recommendation request and the second recommendation result into a training sample for training a second recommendation model.

[0110] In some embodiments, the first evaluation result further indicates a reason why the first recommendation result does not meet the expected goal.

[0111] In some embodiments, the apparatus 500 further includes: a cyclic execution module configured to use the trained generative model to generate a second recommendation request based on context information input of a second user; use the trained first recommendation model to generate a third recommendation result based at least on the second recommendation request; use the trained generative model to generate a third evaluation result for the third recommendation result based on the third recommendation result; and if the third evaluation result indicates that the third recommendation result meets the expected goal, construct the second recommendation request and the third recommendation result into a second training sample for training a second recommendation model.

[0112] The modules included in apparatus 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units in apparatus 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0113] Figure 6 FIG. shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure can be implemented. It should be understood that Figure 6 the illustrated electronic device 600 is merely exemplary and should not impose any limitation on the functionality and scope of the embodiments described herein.

[0114] As Figure 6 shown, the electronic device 600 is in the form of a general-purpose electronic device. The components of the electronic device 600 can include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 can be an actual or virtual processor and is capable of performing various processes according to programs stored in the memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing ability of the electronic device 600.

[0115] The electronic device 600 generally includes multiple computer storage media. Such media can be any accessible media available to the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, magnetic disks, or any other media that can be used to store information and / or data (such as training data for training) and can be accessed within the electronic device 600.

[0116] The electronic device 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 6As shown, a disk drive for reading from and writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading from and writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data medium interfaces. The memory 620 may include a computer program product 625 having one or more program modules configured to perform the various methods or acts of the various embodiments of the present disclosure.

[0117] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented in a single computing cluster or multiple computer machines capable of communicating via a communication connection. Thus, the electronic device 600 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.

[0118] The input device 650 can be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 660 can be one or more output devices such as a display, speaker, printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown) as needed via the communication unit 640, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 600, or communicate with any device that enables the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0119] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which one or more computer instructions are stored, and the one or more computer instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.

[0120] Aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0121] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, result in an apparatus that implements the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0122] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0123] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, segment of a program, or part of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.

[0124] The implementations of the present disclosure have been described above. The description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art in the field to understand the implementations disclosed herein.

Claims

1. A method for constructing model training samples, comprising: Using a trained generative model to generate a first recommendation request based on the context information input of a first user; Using a trained first recommendation model to generate a first recommendation result based at least on the first recommendation request; Using the trained generative model to generate a first evaluation result for the first recommendation result, the first evaluation result at least indicating whether the first recommendation result meets the expected goal; And If the first evaluation result indicates that the first recommendation result meets the expected goal, constructing the first recommendation request and the first recommendation result into a first training sample for training a second recommendation model, where the trained first recommendation model and the second recommendation model are used to recommend the same type of content.

2. The method according to claim 1, wherein generating a first recommendation request based on the context information input of a user comprises: Filling the context information of the first user into a first input template for the trained generative model to obtain a first model input, where the first input template is used to define the description format of the user's context information and the generation indication of the recommendation request; Providing the first model input to the trained generative model; And Obtaining the first recommendation request output by the trained generative model.

3. The method according to claim 1, wherein the context information of the first user is determined based on the historical context information of historical users associated with the trained first recommendation model.

4. The method according to claim 1, wherein generating a first recommendation result based at least on the first recommendation request comprises: Using a trained intent recognition model to determine the recommendation intent in the first recommendation request; Providing the context information of the first user and the recommendation intent to the trained first recommendation model; And Obtaining the first recommendation result output by the trained first recommendation model.

5. The method according to claim 1, wherein generating a first evaluation result for the first recommendation result based on the first recommendation result comprises: Filling the first recommendation result into a second input template for the trained generative model to obtain a second model input, where the second input template is at least used to define the description format of the recommendation result and the generation indication of the evaluation result; Providing the second model input to the trained generative model; And Obtaining the first evaluation result output by the trained generative model.

6. The method according to claim 5, wherein the second input template further includes a reason generation indication, and the reason generation indication is used to indicate the reason why the generative model outputs that the recommendation result does not meet the expected goal.

7. The method according to claim 1, further comprising: If the first evaluation result indicates that the first recommendation result does not meet the expected goal, Using the trained first recommendation model to regenerate a second recommendation result based at least on the first recommendation request; Using the trained generative model, generate a second evaluation result for the second recommendation result based on the second recommendation result; and If the second evaluation result indicates that the second recommendation result meets the expected goal, construct the first recommendation request and the second recommendation result as the training sample for training the second recommendation model.

8. The method according to claim 1, wherein the first evaluation result further indicates the reason why the first recommendation result does not meet the expected goal.

9. The method according to claim 1, further comprising: Using the trained generative model, generate a second recommendation request based on the context information input of a second user; Using the trained first recommendation model, generate a third recommendation result based on at least the second recommendation request; Using the trained generative model, generate a third evaluation result for the third recommendation result based on the third recommendation result; and If the third evaluation result indicates that the third recommendation result meets the expected goal, construct the second recommendation request and the third recommendation result as a second training sample for training the second recommendation model.

10. An apparatus for constructing a model training sample, comprising: A first generation module configured to generate a first recommendation request based on the context information input of a first user using a trained generative model; A second generation module configured to generate a first recommendation result based on at least the first recommendation request using a trained first recommendation model; A third generation module configured to generate a first evaluation result for the first recommendation result based on the first recommendation result using the trained generative model, the first evaluation result at least indicating whether the first recommendation result meets the expected goal; and A training module configured to, if the first evaluation result indicates that the first recommendation result meets the expected goal, construct the first recommendation request and the first recommendation result as a first training sample for training a second recommendation model, the trained first recommendation model and the second recommendation model being used to recommend the same type of content.

11. An electronic device, comprising: At least one processing unit; and At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to execute the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 9.