Resource recommendation and model training method and device based on large model

By applying a reinforcement learning generation model based on large models in the resource recommendation system, the problem that the existing technology is difficult to quickly adapt to user behavior and video content changes is solved, and a more accurate and diversified resource recommendation effect is achieved.

CN120031129APending Publication Date: 2025-05-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510112249.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing resource recommendation technology is difficult to quickly adapt to the dynamic changes in user behavior and video content, resulting in insufficient accuracy and diversity of recommendation results.

Method used

A resource recommendation method based on a large model is adopted, and a reinforcement learning generation model is used to generate multiple estimated resource representation information based on historical resource representation information, and the corresponding recommended resources are determined from the candidate resources.

Benefits of technology

By strengthening the dynamic decision-making ability of the learning generative model and the inference ability of the generative model, we can quickly adapt to changes in user behavior and video content, and improve the accuracy and diversity of recommended results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031129A_ABST
    Figure CN120031129A_ABST
Patent Text Reader

Abstract

The invention provides a resource recommendation and model training method and device based on a large model, relates to the field of artificial intelligence such as deep learning, generative models, reinforcement learning, natural language processing and big data processing, and can be applied to scenes such as information flow recommendation. The resource recommendation method based on the large model can comprise the steps that historical resource representation information is determined according to a first resource recommended to a first recommendation object in history; according to the historical resource representation information, M pieces of first estimated resource representation information are generated through a reinforcement learning generation model, and M is a positive integer larger than 1; and a reinforcement learning generation model is utilized to determine third resources corresponding to the first estimated resource representation information from the candidate second resources, and the third resources are recommended to the first recommendation object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of deep learning, generative models, reinforcement learning, natural language processing and big data processing, and specifically to a resource recommendation and model training method and device based on a big model. Background Art

[0002] At present, resource recommendation technology has been widely used in different scenarios. For example, after a user enters a resource recommendation platform, each time a recommendation request is made, multiple resources can be recommended and displayed to the user, and the resources may be videos, etc. Summary of the invention

[0003] The present invention provides a method and device for resource recommendation and model training based on a large model.

[0004] A resource recommendation method based on a large model, comprising:

[0005] Determining historical resource representation information based on the first resource that has been historically recommended to the first recommendation object;

[0006] According to the historical resource representation information, using a reinforcement learning generation model to generate M first estimated resource representation information respectively, where M is a positive integer greater than 1;

[0007] The reinforcement learning generation model is used to determine the third resources corresponding to each of the first estimated resource representation information from the candidate second resources respectively, and the third resources are recommended to the first recommendation object.

[0008] A reinforcement learning generation model training method, comprising:

[0009] Get training samples;

[0010] The reinforcement learning generation model is trained according to the training samples, and the reinforcement learning generation model is used to generate M first estimated resource representation information according to the historical resource representation information, and determine the third resources corresponding to each first estimated resource representation information from the candidate second resources, and the third resources are used to recommend to the first recommendation object. The historical resource representation information is determined based on the first resources that have been recommended to the first recommendation object in history, and M is a positive integer greater than 1.

[0011] A resource recommendation device based on a large model, comprising: an information determination module, an information generation module and a resource recommendation module;

[0012] The information determination module is used to determine historical resource representation information based on the first resource that has been recommended to the first recommendation object in history;

[0013] The information generation module is used to generate M first estimated resource representation information respectively according to the historical resource representation information by using a reinforcement learning generation model, where M is a positive integer greater than 1;

[0014] The resource recommendation module is used to use the reinforcement learning generation model to determine the third resources corresponding to each of the first estimated resource representation information from the candidate second resources respectively, and recommend the third resources to the first recommendation object.

[0015] A reinforcement learning generation model training device, comprising: a sample acquisition module and a model training module;

[0016] The sample acquisition module is used to acquire training samples;

[0017] The model training module is used to train the reinforcement learning generation model according to the training samples, and the reinforcement learning generation model is used to generate M first estimated resource representation information according to the historical resource representation information, and to determine the third resources corresponding to each first estimated resource representation information from the candidate second resources, and the third resources are used to recommend to the first recommendation object. The historical resource representation information is determined based on the first resources that have been recommended to the first recommendation object in history, and M is a positive integer greater than 1.

[0018] An electronic device, comprising:

[0019] at least one processor; and

[0020] a memory communicatively connected to the at least one processor; wherein,

[0021] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0022] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0023] A computer program product comprises a computer program / instruction, wherein the computer program / instruction implements the method described above when executed by a processor.

[0024] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0026] Figure 1 A flowchart of an embodiment of a resource recommendation method based on a big model described in the present disclosure;

[0027] Figure 2 A schematic diagram of the overall implementation process of the resource recommendation method based on a large model described in the present disclosure;

[0028] Figure 3 A flowchart of an embodiment of the reinforcement learning generation model training method described in the present disclosure;

[0029] Figure 4 A schematic diagram of the composition structure of the reinforcement learning generation model described in the present disclosure;

[0030] Figure 5 It is a schematic diagram of the composition structure of an embodiment 500 of a resource recommendation device based on a large model described in the present disclosure;

[0031] Figure 6 Schematic diagram of the composition structure of the reinforcement learning generation model training device embodiment 600 described in the present disclosure;

[0032] Figure 7 A schematic block diagram of an electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0033] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0034] In addition, it should be understood that the term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0035] Figure 1 Flow chart of an embodiment of the resource recommendation method based on a large model described in the present disclosure. Figure 1 As shown, the following specific implementation methods are included.

[0036] In step 101, historical resource representation information is determined based on a first resource that has been historically recommended to a first recommendation object (such as a first user).

[0037] In step 102, based on the historical resource representation information, a reinforcement learning generation model is used to generate M first estimated resource representation information, where M is a positive integer greater than 1.

[0038] In step 103, a reinforcement learning generation model is used to determine the third resources corresponding to each first estimated resource representation information from each candidate second resource, and recommend the third resource to the first recommendation object.

[0039] Taking video recommendation platforms as an example, they usually have the characteristics of high user participation and fast changes in user behavior. Users often watch multiple videos in succession, providing rich and varied interactive data. In addition, video content is updated quickly and new videos are constantly emerging, which requires video recommendation platforms to be able to quickly adapt to changes in user behavior and video content.

[0040] By adopting the scheme described in the embodiment of the method disclosed herein, resource recommendations can be made with the help of a reinforcement learning generation model. Reinforcement learning is good at handling dynamic interactive environments and can learn and optimize in the continuous interaction between users and systems (platforms), so that it can quickly adapt to changes in user behavior and video content. Moreover, based on the determined historical resource representation information, a reinforcement learning generation model can be used to generate multiple first estimated resource representation information respectively, and then the reinforcement learning generation model can be used to determine the third resource corresponding to each first estimated resource representation information from each second resource, and the third resource can be recommended, thereby utilizing the powerful reasoning ability of the reinforcement learning generation model, etc., and thereby improving the accuracy of the recommendation results.

[0041] It should be noted that the resources and the like in the embodiments described in this disclosure are not for a specific user and do not reflect the personal information of a specific user. In addition, the executor of the method described in this disclosure can obtain the first resource through various public, legal and compliant methods, such as obtaining it after the user's authorization (knowledge and consent). In the technical solution of this disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0042] After a user enters a resource recommendation platform, each time a recommendation request is issued (such as sliding up the screen), a resource recommendation can be made for the user. In this process, historical resource representation information can be determined first.

[0043] In some embodiments of the present disclosure, N fourth resources closest to the current time can be determined from the first resource, where N is a positive integer greater than 1, and the fourth resources are satisfactory resources that meet satisfactory evaluation criteria. Fourth resource vector representations corresponding to each fourth resource are obtained respectively, and a first satisfactory resource vector representation is generated based on each fourth resource vector representation. In addition, M fifth resources most recently recommended to the first recommendation object can be determined from the first resource, and fifth resource vector representations corresponding to each fifth resource are obtained respectively. Then, the first satisfactory resource vector representation and each fifth resource vector representation can be determined as historical resource representation information.

[0044] The specific value of N can be determined according to actual needs, such as 30. In addition, the specific content included in the satisfaction evaluation standard can also be determined according to actual needs, such as including: the first recommendation object has interactive behaviors such as liking and following the resource, and / or the first recommendation object has watched the resource for a time greater than a predetermined threshold.

[0045] Assuming that the number of first resources is 200, of which the number of satisfactory resources that meet the satisfactory evaluation criteria is 50, then the 30 satisfactory resources closest to the current time among these 50 satisfactory resources can be determined as fourth resources, and the fourth resource vector representation corresponding to each fourth resource can be obtained respectively, and then the first satisfactory resource vector representation can be generated according to each fourth resource vector representation.

[0046] In addition, the specific value of M may also be determined according to actual needs, such as 6. Assuming that the number of first resources is 200, the six first resources most recently recommended to the first recommendation object may be determined as the fifth resources, and the fifth resource vector representation corresponding to each fifth resource may be obtained respectively, and then the first satisfactory resource vector representation and each fifth resource vector representation may be determined as the required historical resource representation information.

[0047] It can be seen that the historical resource representation information obtained in the above manner includes both the satisfactory resource vector representation information of the first recommended object and the resource vector representation information of the most recently recommended resource, thereby enriching the content of the historical resource representation information and further improving the accuracy of subsequent processing results.

[0048] According to the historical resource representation information, the reinforcement learning generation model can be used to generate M first estimated resource representation information respectively. Afterwards, the reinforcement learning generation model can also be used to determine the third resource corresponding to each first estimated resource representation information from each second resource.

[0049] In some embodiments of the present disclosure, the M first estimated resource representation information may include: M first estimated resource vector representations, in addition, the second resource vector representations corresponding to each second resource may be obtained respectively, and the following processing may be performed for each first estimated resource vector representation: obtaining the similarity between the first estimated resource vector representation and each second resource vector representation, and obtaining the first estimated long-term value of each second resource respectively, determining the comprehensive evaluation results of each second resource respectively according to the similarity and the first estimated long-term value, and determining L third resources corresponding to the first estimated resource vector representation from each second resource according to the comprehensive evaluation results, where L is a positive integer and is less than the number of second resources. The value of L is usually 1, but may also be greater than 1 in some cases.

[0050] For example, assuming that there are 6 first estimated resource vector representations, namely first estimated resource vector representation 1 to first estimated resource vector representation 6, and assuming that there are 150 candidate resources, taking first estimated resource vector representation 1 as an example, the similarity between the first estimated resource vector representation 1 and the 150 second resource vector representations can be first obtained, and the first estimated long-term values ​​of the 150 second resources can be obtained respectively. In this way, for each second resource, a similarity and a first estimated long-term value can be obtained respectively, and then a comprehensive evaluation result of the second resource can be determined in combination with the similarity and the first estimated long-term value. For example, the similarity and the first estimated long-term value can be weighted and summed to obtain a comprehensive evaluation result of the second resource. Then, the 150 second resources can be sorted in descending order according to the values ​​of the comprehensive evaluation results, and the second resource in the first place after sorting can be determined as the third resource corresponding to the first estimated resource vector representation 1. For the first estimated resource vector representation 2 to the first estimated resource vector representation 6, the corresponding third resources can be determined in the same way.

[0051] It can be seen that in the above processing method, the third resource corresponding to each first estimated resource vector representation can be independently determined, thereby avoiding the error generated when processing a first estimated resource vector representation to affect the processing of other first estimated resource vector representations. Moreover, the third resource corresponding to each first estimated resource vector representation can be determined by combining the similarity with each second resource and the first estimated long-term value of each second resource, thereby improving the accuracy of the determined third resource.

[0052] In some embodiments of the present disclosure, the reinforcement learning generation model may include: an actor module, a critic module and a vector generation module, the vector generation module is used to generate each fourth resource vector representation, each fifth resource vector representation and each second resource vector representation, the actor module is used to generate a first satisfactory resource vector representation based on each fourth resource vector representation, and to generate M first estimated resource vector representations in sequence based on historical resource representation information, the critic module is used to generate a first estimated long-term value corresponding to each vector pair, each vector pair corresponds to a second resource, and each estimated resource vector representation forms a vector pair with each second resource vector representation.

[0053] Combined with the above introduction, Figure 2 Schematic diagram of the overall implementation process of the resource recommendation method based on the big model described in this disclosure. Figure 2 As shown, after the first recommendation object enters a resource recommendation platform, a session can be opened, and a session can include multiple recommendation requests, wherein, after each recommendation request, the fourth resource (satisfactory resource) and the fifth resource (previous resource) can be determined respectively, that is, the N fourth resources closest to the current time can be determined from the first resource, and the fourth resource is a satisfactory resource that meets the satisfactory evaluation standard, and the M fifth resources that were most recently recommended to the first recommendation object can be determined from the first resource, and then the vector generation module (not shown in the figure to simplify the drawing) can be used to generate the vector representation of each fourth resource and the vector representation of each fifth resource respectively, and the vector generation module can be used to generate the second resource vector representation of each second resource as a candidate resource respectively.

[0054] like Figure 2 As shown, the actor module can then be used to generate a first satisfactory resource vector representation based on each fourth resource vector representation, and the actor module can be used to sequentially generate M first estimated resource vector representations based on the first satisfactory resource vector representation and each fifth resource vector representation.

[0055] like Figure 2As shown, further, assuming that there are 150 second resources in total, then for each first estimated resource vector representation, it can be processed in the following manner: respectively obtain the similarity between the first estimated resource vector representation and the 150 second resource vector representations, and use the critic module to generate the first estimated long-term value corresponding to each second resource. For example, the first estimated resource vector representation can be composed of a vector pair with the 150 second resource vector representations, and a total of 150 vector pairs can be obtained. Then, the 150 vector pairs can be input into the critic module respectively, so as to obtain the first estimated long-term value output for each vector pair. Then, according to the similarity corresponding to each second resource and the first estimated long-term value, the comprehensive evaluation result of each second resource can be determined respectively, and then the second resources can be sorted in descending order according to the value of the comprehensive evaluation result, and the second resource that is at the first place after sorting is determined as the third resource corresponding to the first estimated resource vector representation.

[0056] like Figure 2 As shown, after determining respectively that each first estimated resource vector represents a corresponding third resource, each third resource can be recommended to the first recommendation object, and the third resource information recommended this time and the various feedback information generated by the first recommendation object for each third resource can be recorded in a sample pool, so as to perform the next resource recommendation based on the information in the sample pool, etc. The various feedback information may include consumption time (i.e. viewing time), interactive behavior, etc.

[0057] The above scheme can be mapped to various components in reinforcement learning, as follows: Agent: resource recommendation platform; Environment: first recommended object; Action: Figure 2 The third resource determination and recommendation process shown in ; Reward: After each resource recommendation, obtain various feedback information of the first recommended object.

[0058] Traditional resource recommendation methods tend to use known user preferences to determine recommended resources based on the fusion results of multi-objective scoring (such as completion scoring, fast scrolling scoring, interaction scoring, etc.). This can easily lead to information cocoons, and focuses on optimizing users' short-term behavior, ignoring users' long-term satisfaction, and is not sensitive enough to the dynamic changes of users and resources.

[0059] Reinforcement learning is good at handling dynamic interactive environments. It can learn and optimize in the continuous interaction between users and systems and quickly adapt to system changes. Moreover, by designing a reasonable reward mechanism, reinforcement learning can optimize users' long-term satisfaction and improve user retention rate, rather than just optimizing users' short-term consumption behavior. In addition, reinforcement learning can break the information cocoon through the exploration mechanism, recommend fresh content, increase recommendation diversity, and achieve a better balance between exploring new content and leveraging users' known preferences.

[0060] The reinforcement learning generation model used in the resource recommendation method disclosed in the present invention is pre-trained, and the training method of the reinforcement learning generation model is described below.

[0061] Accordingly, Figure 3 Flow chart of an embodiment of the reinforcement learning generation model training method described in the present disclosure. Figure 3 As shown, the following specific implementation methods are included.

[0062] In step 301, a training sample is obtained.

[0063] In step 302, the reinforcement learning generation model is trained according to the training samples. The reinforcement learning generation model is used to generate M first estimated resource representation information according to the historical resource representation information, and to determine the third resource corresponding to each first estimated resource representation information from the candidate second resources. The third resource is used to recommend to the first recommendation object. The historical resource representation information is determined based on the first resource that has been recommended to the first recommendation object in history, and M is a positive integer greater than 1.

[0064] By adopting the scheme described in the above method embodiment, a reinforcement learning generation model can be trained based on the acquired training samples, so that resource recommendations can be made with the help of the reinforcement learning generation model, thereby being able to quickly adapt to changes in user behavior and video content, and improving the accuracy of recommendation results, etc.

[0065] In some embodiments of the present disclosure, the training sample may include: N eighth resources, M ninth resources and M seventh resources, N is a positive integer greater than 1, the eighth resource is the N satisfactory resources that meet the satisfactory evaluation criteria and are determined from the sixth resource and are closest to the recommendation time of the seventh resource, the sixth resource is the resource recommended to the second recommendation object before the recommendation time of the seventh resource, and the ninth resource is the resource recommended to the second recommendation object for the most recent time before the recommendation time of the seventh resource.

[0066] That is, the resources in the same training sample correspond to the same second recommendation object (such as the second user). Suppose that a certain resource recommendation recommends 6 resources to the second recommendation object, namely, resources 1 to 6, and suppose that resources 7 to 12 were recommended to the second recommendation object in the last resource recommendation. In addition, suppose that 30 satisfactory resources are determined from the resources recommended to the second recommendation object before the recommendation time of resources 1 to 6, then the 30 satisfactory resources, resources 7 to 12 (previous resources) and resources 1 to 6 (current resources) can be used to form a training sample. In the same way, multiple training samples can be constructed simply and quickly, thereby improving the training efficiency of the model.

[0067] Accordingly, the constructed training samples can be used to train the reinforcement learning generation model. In some embodiments of the present disclosure, the M second estimated resource representation information may include: M second estimated resource vector representations, and the method of training the reinforcement learning generation model according to the training samples may include: respectively obtaining the eighth resource vector representation corresponding to each eighth resource, and generating a second satisfactory resource vector representation according to each eighth resource vector representation, respectively obtaining the ninth resource vector representation corresponding to each ninth resource and the seventh resource vector representation corresponding to each seventh resource, respectively generating M second estimated resource vector representations according to the second satisfactory resource vector representation, each ninth resource vector representation and each seventh resource vector representation, respectively generating M second estimated long-term values ​​according to each second estimated resource vector representation and each seventh resource vector representation, respectively determining a comprehensive loss according to each second estimated resource vector representation, each seventh resource vector representation and each second estimated long-term value, and updating the parameters of the reinforcement learning generation model according to the comprehensive loss.

[0068] That is, the vector representation of each satisfactory resource can be obtained respectively, and the second satisfactory resource vector representation can be generated according to the vector representation of each satisfactory resource, and the vector representation of each previous resource and the vector representation of each current resource can be obtained respectively. Each vector representation can include vector representation information such as semantic quantization identification (ID), time consumption factor, interaction factor, etc., which are concatenated together as the required vector representation. Afterwards, M second estimated resource vector representations can be generated according to the second satisfactory resource vector representation, the vector representation of each previous resource and the vector representation of each current resource, and M second estimated long-term values ​​can be generated according to each second estimated resource vector representation and the vector representation of each current resource.

[0069] In some embodiments of the present disclosure, the reinforcement learning generation model may include: an actor module, a critic module and a vector generation module, the vector generation module is used to generate each eighth resource vector representation, each ninth resource vector representation and each seventh resource vector representation, the actor module is used to generate a second satisfactory resource vector representation based on each eighth resource vector representation, and sequentially generate M second estimated resource vector representations based on the second satisfactory resource vector representation, each ninth resource vector representation and each seventh resource vector representation, the critic module is used to generate a second estimated long-term value corresponding to each vector pair, each vector pair corresponds to a seventh resource, each second estimated resource vector representation forms a vector pair with the corresponding seventh resource vector representation, and the seventh resource vector representation corresponding to any second estimated resource vector representation is: the seventh resource vector representation whose ranking position in each seventh resource vector representation is the same as the ranking position of the second estimated resource vector representation in each second estimated resource vector representation.

[0070] Figure 4Schematic diagram of the composition structure of the reinforcement learning generation model described in this disclosure. Figure 4 As shown, to simplify the drawing, the vector generation module is not represented.

[0071] like Figure 4 As shown, the actor module can adopt an encoder-decoder structure. Among them, the encoder can generate a second satisfactory resource vector representation based on the input eighth resource vector representations, that is, the encoder can learn to obtain the user satisfaction sequence representation, such as learning the cross-features between the satisfactory resources, etc., and output a new vector representation to represent the user's historical interest information, etc. Assuming that the input eighth resource vector representations are 512 dimensions respectively, and the number of eighth resources is 30, then the dimension of the generated second satisfactory resource vector representation can be 512*30. The decoder can generate M second estimated resource vector representations in sequence according to the second satisfactory resource vector representation, each ninth resource vector representation and each seventh resource vector representation, such as generating a second estimated resource vector representation 1, a second estimated resource vector representation 2, ..., a second estimated resource vector representation 6 in sequence.

[0072] like Figure 4 As shown, the critic module can use an encoder-decoder-multilayer perceptron (MLP) structure to estimate the long-term value of resources. If necessary, the encoder in the critic module can be shared with the encoder in the actor module. The multilayer perceptron can generate corresponding second estimated long-term values ​​for the input vector pairs, each vector pair can correspond to a seventh resource, and each second estimated resource vector representation can form a vector pair with the corresponding seventh resource vector representation. For example, assuming that there are 6 second estimated resource vector representations, namely second estimated resource vector representation 1 to second estimated resource vector representation 6, and assuming that there are 6 seventh resources, namely resource 1 to resource 6, then the second estimated resource vector representation 1 and resource 1 can be used to form a vector pair, and the second estimated resource vector representation 2 and resource 2 can be used to form a vector pair, and so on. The second estimated long-term value is a specific numerical value.

[0073] The reinforcement learning generation model disclosed in the present invention can combine the advantages of reinforcement learning and generative models, and make use of the dynamic decision-making ability, long-term benefit optimization ability and exploration ability of reinforcement learning, and the understanding ability and reasoning ability of the generative model to effectively improve the recommendation efficiency and the accuracy and diversity of the recommendation results.

[0074] During the training process, the comprehensive loss can be determined based on the second estimated resource vector representation, the seventh resource vector representation and the second estimated long-term value, and the parameters of the reinforcement learning generation model can be updated based on the comprehensive loss.

[0075] In some embodiments of the present disclosure, the actor loss (Actor Loss) and the critic loss (Critic Loss) may be obtained respectively, and then the actor loss and the critic loss may be multiplied by the corresponding weights respectively, and the two products may be added to obtain the comprehensive loss.

[0076] Then we have:

[0077] Loss=0.9*Actor Loss+0.1*Critic Loss; (1)

[0078] Among them, Loss represents the comprehensive loss, 0.9 and 0.1 represent the weights corresponding to the actor loss and the critic loss respectively. The specific values ​​are only for illustration and can be determined according to actual needs.

[0079] Through the above processing, the obtained comprehensive loss can integrate the actor loss and the critic loss at the same time, thereby improving the comprehensiveness and accuracy of the obtained comprehensive loss, and further improving the accuracy of subsequent model parameter update results.

[0080] In some embodiments of the present disclosure, a method for obtaining an actor's loss may include: for each second estimated resource vector representation, performing the following processing respectively: obtaining the cross entropy loss between the second estimated resource vector representation and the corresponding seventh resource vector representation, and obtaining the difference between the actual long-term value of the seventh resource corresponding to the second estimated resource vector representation and the second estimated long-term value corresponding to the second estimated resource vector representation, determining the product of the cross entropy loss and the difference as the resource loss corresponding to the second estimated resource vector representation, and determining the actor's loss according to the resource loss corresponding to each second estimated resource vector representation.

[0081] For example, assuming that there are a total of 6 second estimated resource vector representations, namely second estimated resource vector representation 1 to second estimated resource vector representation 6, and assuming that there are a total of 6 seventh resources, namely resource 1 to resource 6, then taking second estimated resource vector representation 1 as an example, the cross entropy loss between the second estimated resource vector representation 1 and the vector representation of resource 1 can be obtained, and the difference between the actual long-term value of resource 1 and the second estimated long-term value 1 corresponding to the second estimated resource vector representation 1 can be obtained, and the difference can be called advantage. Then the cross entropy loss can be multiplied by the difference, and the product obtained is determined as the resource loss corresponding to the second estimated resource vector representation 1. In the same way, the resource losses corresponding to the second estimated resource vector representation 2 to the second estimated resource vector representation 6 can be obtained respectively, and then the actor loss can be determined in combination with the 6 resource losses, such as the average of the 6 resource losses can be determined as the actor loss, or the sum of the 6 resource losses can be determined as the actor loss, etc. The specific method is not limited.

[0082] There is no restriction on how to obtain the actual long-term value of resources. For example, for the above-mentioned resource 1, the actual long-term value of resource 1 can be calculated based on various feedback information from different users on resource 1 using pre-set value assessment rules. In other words, the actual long-term value of resource 1 can be generated based on real online feedback results.

[0083] In addition, in some embodiments of the present disclosure, the method of obtaining the critic's loss may include: for each second estimated resource vector representation, performing the following processing respectively: obtaining the cross-entropy loss between the actual long-term value of the seventh resource corresponding to the second estimated resource vector representation and the second estimated long-term value corresponding to the second estimated resource vector representation, and determining the critic's loss based on the cross-entropy loss corresponding to each second estimated resource vector representation.

[0084] For example, assuming that there are a total of 6 second estimated resource vector representations, namely second estimated resource vector representation 1 to second estimated resource vector representation 6, and assuming that there are a total of 6 seventh resources, namely resource 1 to resource 6, then taking second estimated resource vector representation 1 as an example, the cross entropy loss of the actual long-term value of resource 1 and the second estimated long-term value 1 corresponding to the second estimated resource vector representation 1 can be obtained, and then the critic loss can be determined by combining the cross entropy losses corresponding to the 6 second estimated resource vector representations. For example, the average of the 6 cross entropy losses can be determined as the critic loss, or the sum of the 6 cross entropy losses can be determined as the critic loss, and the specific method is not limited.

[0085] It can be seen that by adopting the above processing method, the actor loss and critic loss can be determined by combining the estimated resource vector representation, the actual vector representation of the resource, the actual long-term value of the resource and the estimated long-term value of the resource, thereby improving the accuracy of the obtained actor loss and critic loss, and correspondingly improving the accuracy of the obtained comprehensive loss.

[0086] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the described order of actions, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure. In addition, for parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0087] The above is an introduction to the method embodiment. The following is a further explanation of the scheme disclosed in the present invention through an apparatus embodiment.

[0088] Figure 5FIG. 5 is a schematic diagram of the composition structure of an embodiment of a resource recommendation device based on a large model described in the present disclosure. Figure 5 As shown, it includes: an information determination module 501, an information generation module 502 and a resource recommendation module 503.

[0089] The information determination module 501 is used to determine historical resource representation information according to the first resource that has been recommended to the first recommendation object in history.

[0090] The information generation module 502 is used to generate M first estimated resource representation information respectively according to the historical resource representation information by using the reinforcement learning generation model, where M is a positive integer greater than 1.

[0091] The resource recommendation module 503 is used to use the reinforcement learning generation model to determine the third resources corresponding to each first estimated resource representation information from each candidate second resource, and recommend the third resource to the first recommendation object.

[0092] By adopting the scheme described in the above-mentioned device embodiment, resource recommendations can be made with the help of a reinforcement learning generation model. Reinforcement learning is good at handling dynamic interactive environments and can learn and optimize in the continuous interaction between users and systems, so that it can quickly adapt to changes in user behavior and video content. Moreover, based on the determined historical resource representation information, a reinforcement learning generation model can be used to generate multiple first estimated resource representation information respectively, and then the reinforcement learning generation model can be used to determine the third resource corresponding to each first estimated resource representation information from each second resource, and the third resource can be recommended, thereby utilizing the powerful reasoning ability of the reinforcement learning generation model, etc., and thereby improving the accuracy of the recommendation results.

[0093] In some embodiments of the present disclosure, the information determination module 501 can determine the N fourth resources closest to the current time from the first resource, where N is a positive integer greater than 1, and the fourth resources are satisfactory resources that meet the satisfactory evaluation criteria, and obtain the fourth resource vector representation corresponding to each fourth resource, and generate a first satisfactory resource vector representation based on each fourth resource vector representation. In addition, the M fifth resources most recently recommended to the first recommendation object can be determined from the first resource, and the fifth resource vector representation corresponding to each fifth resource can be obtained, and then the first satisfactory resource vector representation and each fifth resource vector representation can be determined as historical resource representation information.

[0094] In some embodiments of the present disclosure, the M first estimated resource representation information may include: M first estimated resource vector representations. In addition, the resource recommendation module 503 may respectively obtain the second resource vector representation corresponding to each second resource, and may respectively perform the following processing for each first estimated resource vector representation: obtain the similarity between the first estimated resource vector representation and each second resource vector representation, and respectively obtain the first estimated long-term value of each second resource, determine the comprehensive evaluation results of each second resource according to the similarity and the first estimated long-term value, and determine L third resources corresponding to the first estimated resource vector representation from each second resource according to the comprehensive evaluation results, where L is a positive integer and is less than the number of second resources.

[0095] In addition, in some embodiments of the present disclosure, the reinforcement learning generation model may include: an actor module, a critic module and a vector generation module, the vector generation module is used to generate each fourth resource vector representation, each fifth resource vector representation and each second resource vector representation, the actor module is used to generate a first satisfactory resource vector representation based on each fourth resource vector representation, and to generate M first estimated resource vector representations in sequence based on historical resource representation information, the critic module is used to generate a first estimated long-term value corresponding to each vector pair, each vector pair corresponds to a second resource, and each estimated resource vector representation forms a vector pair with each second resource vector representation.

[0096] Figure 6 FIG. 6 is a schematic diagram of the structure of the reinforcement learning generation model training device embodiment 600 described in the present disclosure. Figure 6 As shown, it includes: a sample acquisition module 601 and a model training module 602.

[0097] The sample acquisition module 601 is used to acquire training samples.

[0098] The model training module 602 is used to train the reinforcement learning generation model according to the training samples. The reinforcement learning generation model is used to generate M first estimated resource representation information according to the historical resource representation information, and to determine the third resource corresponding to each first estimated resource representation information from the candidate second resources. The third resource is used to recommend to the first recommendation object. The historical resource representation information is determined based on the first resource that has been recommended to the first recommendation object in history. M is a positive integer greater than 1.

[0099] By adopting the scheme described in the above-mentioned device embodiment, a reinforcement learning generation model can be trained based on the acquired training samples, so that resource recommendations can be made with the help of the reinforcement learning generation model, thereby being able to quickly adapt to changes in user behavior and video content, and improving the accuracy of recommendation results, etc.

[0100] In some embodiments of the present disclosure, the training sample may include: N eighth resources, M ninth resources and M seventh resources, N is a positive integer greater than 1, the eighth resource is the N satisfactory resources that meet the satisfactory evaluation criteria and are determined from the sixth resource and are closest to the recommendation time of the seventh resource, the sixth resource is the resource recommended to the second recommendation object before the recommendation time of the seventh resource, and the ninth resource is the resource recommended to the second recommendation object for the most recent time before the recommendation time of the seventh resource.

[0101] Accordingly, the model training module 602 can use the constructed training samples to train the reinforcement learning generation model. In some embodiments of the present disclosure, the M second estimated resource representation information may include: M second estimated resource vector representations, and the model training module 602 may train the reinforcement learning generation model according to the training samples. The method of training the reinforcement learning generation model may include: respectively obtaining the eighth resource vector representation corresponding to each eighth resource, and generating a second satisfactory resource vector representation according to each eighth resource vector representation, respectively obtaining the ninth resource vector representation corresponding to each ninth resource and the seventh resource vector representation corresponding to each seventh resource, respectively generating M second estimated resource vector representations according to the second satisfactory resource vector representation, each ninth resource vector representation and each seventh resource vector representation, respectively generating M second estimated long-term values ​​according to each second estimated resource vector representation and each seventh resource vector representation, determining a comprehensive loss according to each second estimated resource vector representation, each seventh resource vector representation and each second estimated long-term value, and updating the parameters of the reinforcement learning generation model according to the comprehensive loss.

[0102] In some embodiments of the present disclosure, the reinforcement learning generation model may include: an actor module, a critic module and a vector generation module, the vector generation module is used to generate each eighth resource vector representation, each ninth resource vector representation and each seventh resource vector representation, the actor module is used to generate a second satisfactory resource vector representation based on each eighth resource vector representation, and sequentially generate M second estimated resource vector representations based on the second satisfactory resource vector representation, each ninth resource vector representation and each seventh resource vector representation, the critic module is used to generate a second estimated long-term value corresponding to each vector pair, each vector pair corresponds to a seventh resource, each second estimated resource vector representation forms a vector pair with the corresponding seventh resource vector representation, and the seventh resource vector representation corresponding to any second estimated resource vector representation is: the seventh resource vector representation whose ranking position in each seventh resource vector representation is the same as the ranking position of the second estimated resource vector representation in each second estimated resource vector representation.

[0103] During the training process, the model training module 602 can determine the comprehensive loss based on each second estimated resource vector representation, each seventh resource vector representation and each second estimated long-term value, and can update the parameters of the reinforcement learning generation model based on the comprehensive loss.

[0104] In some embodiments of the present disclosure, the way in which the model training module 602 obtains the actor loss may include: for each second estimated resource vector representation, performing the following processing respectively: obtaining the cross entropy loss between the second estimated resource vector representation and the corresponding seventh resource vector representation, and obtaining the difference between the actual long-term value of the seventh resource corresponding to the second estimated resource vector representation and the second estimated long-term value corresponding to the second estimated resource vector representation, determining the product of the cross entropy loss and the difference as the resource loss corresponding to the second estimated resource vector representation, and determining the actor loss according to the resource loss corresponding to each second estimated resource vector representation.

[0105] In addition, in some embodiments of the present disclosure, the way in which the model training module 602 obtains the critic loss may include: performing the following processing for each second estimated resource vector representation: obtaining the cross entropy loss between the actual long-term value of the seventh resource corresponding to the second estimated resource vector representation and the second estimated long-term value corresponding to the second estimated resource vector representation, and determining the critic loss based on the cross entropy loss corresponding to each second estimated resource vector representation.

[0106] Figure 5 and Figure 6 The specific working process of the illustrated device embodiment can refer to the relevant description in the aforementioned method embodiment and will not be described in detail.

[0107] The scheme disclosed in the present invention can be applied to the field of artificial intelligence, especially to the fields of deep learning, generative models, reinforcement learning, natural language processing, and big data processing. Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.

[0108] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0109] Figure 7A schematic block diagram of an electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0110] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 to a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0111] Multiple components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0112] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI, Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP, Digital Signal Processing), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the methods described in the present disclosure may be executed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the method described in the present disclosure in any other appropriate manner (for example, by means of firmware).

[0113] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general programmable processor, which can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0114] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0115] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, Electronically Programmable Read-Only Memory), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM, Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0116] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0117] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0118] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0119] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0120] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A resource recommendation method based on a large model, comprising: Determining historical resource representation information based on the first resource that has been historically recommended to the first recommendation object; According to the historical resource representation information, using a reinforcement learning generation model to generate M first estimated resource representation information respectively, where M is a positive integer greater than 1; The reinforcement learning generation model is used to determine the third resources corresponding to each of the first estimated resource representation information from the candidate second resources respectively, and the third resources are recommended to the first recommendation object.

2. The method according to claim 1, wherein: The determining of historical resource representation information based on the first resource that has been historically recommended to the first recommendation object includes: Determine N fourth resources closest to the current time from the first resources, where N is a positive integer greater than 1, and the fourth resources are satisfactory resources that meet the satisfactory evaluation criteria, respectively obtain fourth resource vector representations corresponding to the fourth resources, and generate a first satisfactory resource vector representation according to the fourth resource vector representations; Determine, from the first resources, M fifth resources that were most recently recommended to the first recommendation object, and respectively obtain fifth resource vector representations corresponding to the fifth resources; The first satisfactory resource vector representation and each of the fifth resource vector representations are determined as the historical resource representation information.

3. The method according to claim 2, wherein: The M first estimated resource representation information includes: M first estimated resource vector representations; The determining of the third resources corresponding to the first estimated resource representation information from the candidate second resources respectively includes: Respectively obtain second resource vector representations corresponding to each of the second resources; And for each of the first estimated resource vector representations, the following processing is performed respectively: Obtaining the similarity between the first estimated resource vector representation and each of the second resource vector representations, and respectively obtaining the first estimated long-term value of each of the second resources; Determining comprehensive evaluation results of each of the second resources according to the similarity and the first estimated long-term value; According to the comprehensive evaluation result, L third resources corresponding to the first estimated resource vector are determined from each of the second resources, where L is a positive integer and is less than the number of the second resources.

4. The method according to claim 3, wherein: The reinforcement learning generation model includes: an actor module, a critic module and a vector generation module; The vector generation module is used to generate each of the fourth resource vector representations, each of the fifth resource vector representations and each of the second resource vector representations respectively; The actor module is used to generate the first satisfactory resource vector representation according to each of the fourth resource vector representations, and to sequentially generate the M first estimated resource vector representations according to the historical resource representation information; The critic module is used to generate the first estimated long-term value corresponding to each vector pair, each vector pair corresponds to one second resource, and the first estimated resource vector representation and each second resource vector representation form a vector pair.

5. A reinforcement learning generative model training method, comprising: Get training samples; The reinforcement learning generation model is trained according to the training samples, and the reinforcement learning generation model is used to generate M first estimated resource representation information according to the historical resource representation information, and determine the third resources corresponding to each first estimated resource representation information from the candidate second resources, and the third resources are used to recommend to the first recommendation object. The historical resource representation information is determined based on the first resources that have been recommended to the first recommendation object in history, and M is a positive integer greater than 1.

6. The method according to claim 5, wherein: The training sample includes: N eighth resources, M ninth resources and M seventh resources, N is a positive integer greater than 1, the eighth resource is the N satisfactory resources that meet the satisfactory evaluation criteria and are determined from the sixth resources and are closest to the recommendation time of the seventh resource, the sixth resource is the resource recommended to the second recommendation object before the recommendation time of the seventh resource, and the ninth resource is the resource recommended to the second recommendation object for the most recent time before the recommendation time of the seventh resource.

7. The method according to claim 6, wherein: The M second estimated resource representation information includes: M second estimated resource vector representations; The training of the reinforcement learning generation model according to the training samples includes: Respectively obtaining the eighth resource vector representation corresponding to each of the eighth resources, and generating a second satisfactory resource vector representation according to each of the eighth resource vector representations; Respectively obtaining a ninth resource vector representation corresponding to each of the ninth resources and a seventh resource vector representation corresponding to each of the seventh resources; Generate M second estimated resource vector representations respectively according to the second satisfactory resource vector representation, each of the ninth resource vector representations and each of the seventh resource vector representations; Generate M second estimated long-term values ​​respectively according to each of the second estimated resource vector representations and each of the seventh resource vector representations; A comprehensive loss is determined based on each of the second estimated resource vector representations, each of the seventh resource vector representations, and each of the second estimated long-term values, and parameters of the reinforcement learning generation model are updated based on the comprehensive loss.

8. The method according to claim 7, wherein: The reinforcement learning generation model includes: an actor module, a critic module and a vector generation module; The vector generation module is used to generate each of the eighth resource vector representations, each of the ninth resource vector representations, and each of the seventh resource vector representations respectively; The actor module is used to generate the second satisfactory resource vector representation according to each of the eighth resource vector representations, and to sequentially generate the M second estimated resource vector representations according to the second satisfactory resource vector representation, each of the ninth resource vector representations, and each of the seventh resource vector representations; The critic module is used to generate the second estimated long-term value corresponding to each vector pair respectively, each vector pair corresponds to one of the seventh resources respectively, each of the second estimated resource vector representations respectively forms a vector pair with the corresponding seventh resource vector representation, and the seventh resource vector representation corresponding to any second estimated resource vector representation is respectively: the seventh resource vector representation whose sorting position in each of the seventh resource vector representations is the same as the sorting position of the second estimated resource vector representation in each of the second estimated resource vector representations.

9. The method according to claim 8, wherein: The comprehensive loss is determined to include: Get the actor loss and critic loss respectively; The actor loss and the critic loss are respectively multiplied by the corresponding weights, and the two products are added to obtain the comprehensive loss.

10. The method according to claim 9, wherein: Obtaining the actor's loss includes: For each of the second estimated resource vector representations, the following processing is performed respectively: obtaining a cross entropy loss between the second estimated resource vector representation and the corresponding seventh resource vector representation, and obtaining a difference between the actual long-term value of the seventh resource corresponding to the second estimated resource vector representation and the second estimated long-term value corresponding to the second estimated resource vector representation, and determining the product of the cross entropy loss and the difference as the resource loss corresponding to the second estimated resource vector representation; The actor loss is determined according to the resource loss corresponding to each of the second estimated resource vector representations.

11. The method according to claim 9, wherein: Obtaining the critic loss involves: For each second estimated resource vector representation, the following processing is performed respectively: obtaining a cross entropy loss between the actual long-term value of the seventh resource corresponding to the second estimated resource vector representation and the second estimated long-term value corresponding to the second estimated resource vector representation; The critic loss is determined according to the cross entropy loss corresponding to each of the second estimated resource vector representations.

12. A resource recommendation device based on a large model, comprising: Information determination module, information generation module and resource recommendation module; The information determination module is used to determine historical resource representation information based on the first resource that has been recommended to the first recommendation object in history; The information generation module is used to generate M first estimated resource representation information respectively according to the historical resource representation information by using a reinforcement learning generation model, where M is a positive integer greater than 1; The resource recommendation module is used to use the reinforcement learning generation model to determine the third resources corresponding to each of the first estimated resource representation information from the candidate second resources respectively, and recommend the third resources to the first recommendation object.

13. The device according to claim 12, wherein: The information determination module determines N fourth resources closest to the current time from the first resource, where N is a positive integer greater than 1, and the fourth resources are satisfactory resources that meet the satisfactory evaluation criteria, obtains fourth resource vector representations corresponding to each of the fourth resources, and generates a first satisfactory resource vector representation based on each of the fourth resource vector representations, determines M fifth resources most recently recommended to the first recommendation object from the first resource, obtains fifth resource vector representations corresponding to each of the fifth resources, and determines the first satisfactory resource vector representation and each of the fifth resource vector representations as the historical resource representation information.

14. The device according to claim 13, wherein: The M first estimated resource representation information includes: M first estimated resource vector representations; The resource recommendation module respectively obtains the second resource vector representation corresponding to each of the second resources, and performs the following processing for each of the first estimated resource vector representations: obtains the similarity between the first estimated resource vector representation and each of the second resource vector representations, and respectively obtains the first estimated long-term value of each of the second resources, and determines the comprehensive evaluation results of each of the second resources according to the similarity and the first estimated long-term value, and determines L third resources corresponding to the first estimated resource vector representation from each of the second resources according to the comprehensive evaluation results, where L is a positive integer and is less than the number of the second resources.

15. The device according to claim 14, wherein: The reinforcement learning generation model includes: an actor module, a critic module and a vector generation module; The vector generation module is used to generate each of the fourth resource vector representations, each of the fifth resource vector representations and each of the second resource vector representations respectively; The actor module is used to generate the first satisfactory resource vector representation according to each of the fourth resource vector representations, and to sequentially generate the M first estimated resource vector representations according to the historical resource representation information; The critic module is used to generate the first estimated long-term value corresponding to each vector pair, each vector pair corresponds to one second resource, and the first estimated resource vector representation and each second resource vector representation form a vector pair.

16. A reinforcement learning generation model training device, comprising: Sample acquisition module and model training module; The sample acquisition module is used to acquire training samples; The model training module is used to train the reinforcement learning generation model according to the training samples, and the reinforcement learning generation model is used to generate M first estimated resource representation information according to the historical resource representation information, and to determine the third resources corresponding to each first estimated resource representation information from the candidate second resources, and the third resources are used to recommend to the first recommendation object. The historical resource representation information is determined based on the first resources that have been recommended to the first recommendation object in history, and M is a positive integer greater than 1.

17. The device according to claim 16, wherein: The training sample includes: N eighth resources, M ninth resources and M seventh resources, N is a positive integer greater than 1, the eighth resource is the N satisfactory resources that meet the satisfactory evaluation criteria and are determined from the sixth resources and are closest to the recommendation time of the seventh resource, the sixth resource is the resource recommended to the second recommendation object before the recommendation time of the seventh resource, and the ninth resource is the resource recommended to the second recommendation object for the most recent time before the recommendation time of the seventh resource.

18. The device according to claim 17, wherein: The M second estimated resource representation information includes: M second estimated resource vector representations; The model training module respectively obtains the eighth resource vector representation corresponding to each of the eighth resources, and generates a second satisfactory resource vector representation based on each of the eighth resource vector representations, respectively obtains the ninth resource vector representation corresponding to each of the ninth resources and the seventh resource vector representation corresponding to each of the seventh resources, respectively generates M second estimated resource vector representations based on the second satisfactory resource vector representation, each of the ninth resource vector representations and each of the seventh resource vector representations, respectively generates M second estimated long-term values ​​based on each of the second estimated resource vector representations and each of the seventh resource vector representations, determines a comprehensive loss based on each of the second estimated resource vector representations, each of the seventh resource vector representations and each of the second estimated long-term values, and updates the parameters of the reinforcement learning generation model based on the comprehensive loss.

19. The device according to claim 18, wherein: The reinforcement learning generation model includes: an actor module, a critic module and a vector generation module; The vector generation module is used to generate each of the eighth resource vector representations, each of the ninth resource vector representations, and each of the seventh resource vector representations respectively; The actor module is used to generate the second satisfactory resource vector representation according to each of the eighth resource vector representations, and to sequentially generate the M second estimated resource vector representations according to the second satisfactory resource vector representation, each of the ninth resource vector representations, and each of the seventh resource vector representations; The critic module is used to generate the second estimated long-term value corresponding to each vector pair respectively, each vector pair corresponds to one of the seventh resources respectively, each of the second estimated resource vector representations respectively forms a vector pair with the corresponding seventh resource vector representation, and the seventh resource vector representation corresponding to any second estimated resource vector representation is respectively: the seventh resource vector representation whose sorting position in each of the seventh resource vector representations is the same as the sorting position of the second estimated resource vector representation in each of the second estimated resource vector representations.

20. The device according to claim 19, wherein The model training module obtains the actor loss and the critic loss respectively, multiplies the actor loss and the critic loss by corresponding weights respectively, and adds the two products to obtain the comprehensive loss.

21. The device according to claim 20, wherein: The model training module performs the following processing for each second estimated resource vector representation: obtains the cross entropy loss between the second estimated resource vector representation and the corresponding seventh resource vector representation, and obtains the difference between the actual long-term value of the seventh resource corresponding to the second estimated resource vector representation and the second estimated long-term value corresponding to the second estimated resource vector representation, determines the product of the cross entropy loss and the difference as the resource loss corresponding to the second estimated resource vector representation, and determines the actor loss according to the resource loss corresponding to each second estimated resource vector representation.

22. The device according to claim 20, wherein: The model training module performs the following processing for each second estimated resource vector representation: obtains the cross entropy loss between the actual long-term value of the seventh resource corresponding to the second estimated resource vector representation and the second estimated long-term value corresponding to the second estimated resource vector representation, and determines the critic loss based on the cross entropy loss corresponding to each second estimated resource vector representation.

23. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.

24. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 11.

25. A computer program product, comprising a computer program / instruction, which implements the method according to any one of claims 1 to 11 when executed by a processor.