Recommendation method, training method and related products

Through an end-to-end recommendation method, object interaction records and item descriptions are used to determine the target recommendation order, which solves the problem of low accuracy of existing recommendation systems and achieves the matching of recommendation results with the object's interests and requirements.

CN120611099APending Publication Date: 2025-09-09SHUXING TECH (BEIJING) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510723334.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing recommendation systems have low accuracy when determining the order of recommendations and cannot simultaneously meet the conditions that the subject is interested in the items and the recommendation results match the requirements.

Method used

An end-to-end recommendation method is adopted to obtain the interaction records between the target object and the item, aggregate and determine the object description, combine it with the item description, and input it into the recommendation model to determine the target recommendation order, taking into account the object's interests and recommendation requirements.

Benefits of technology

The accuracy of items recommended to the target is improved, ensuring that the recommendation results are in line with the target's interests and meet the recommendation requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611099A_ABST
    Figure CN120611099A_ABST
Patent Text Reader

Abstract

The invention discloses a recommendation method, a training method and a related product. The method comprises the steps of obtaining a recommendation model and a target text, wherein the target text comprises description of a target object and description of a to-be-recommended article; the target text is input into the recommendation model, a target recommendation sequence output by the recommendation model is obtained, the target recommendation sequence comprises at least one target article in the to-be-recommended articles, and the at least one target article is an article recommended to the target object. In the method, the recommendation model can determine at least one target article recommended to the target object in an end-to-end manner, so that the accuracy of the recommendation sequence can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of recommendation technology, and in particular to a recommendation method, a training method and related products. Background Art

[0002] Currently, amidst the internet's information overload, recommendation systems have become a key technology for improving user experience and driving business growth. To meet user personalized needs, recommendation systems typically build recommendation models based on historical behavior, characteristics, and other data. Using algorithms such as collaborative filtering and deep learning, they explore users' potential interests and determine the order of recommended items. However, current methods often provide low accuracy in determining the order. Summary of the Invention

[0003] This application provides a recommendation method, training method, and related products, which determine at least one target item to recommend to a target subject in an end-to-end manner, thereby improving the accuracy of the at least one target item. The related products include a recommendation device, a training device, an electronic device, a computer-readable storage medium, and a computer program product.

[0004] In a first aspect, a recommendation method is provided, the recommendation method comprising:

[0005] Obtaining a recommendation model and a target text, wherein the target text includes a description of the target object and a description of the recommended item;

[0006] The target text is input into the recommendation model to obtain a target recommendation order output by the recommendation model, wherein the target recommendation order includes at least one target item among the items to be recommended, and the at least one target item is an item recommended to the target object.

[0007] In conjunction with any embodiment of the present application, obtaining the target text includes:

[0008] Obtaining a first interaction record between the target object and the item;

[0009] Determining a description of the target object by aggregating the first interaction records;

[0010] The target text is obtained based on the description of the target object and the description of the item to be recommended.

[0011] In combination with any embodiment of the present application, the determining a description of the target object by aggregating the first interaction records includes:

[0012] Determining a second interaction record by aggregating the first interaction record, where the second interaction record includes a record of the target object interacting with the item within a first time period and / or a record of the target object interacting with the item within a second time period, where a time span of the first time period is shorter than a time span of the second time period;

[0013] Based on the second interaction record, a description of the target object is determined.

[0014] In combination with any embodiment of the present application, the record of the target object's interaction with the item in the first time period and the record of the target object's interaction with the item in the second time period both include statistical results of the target object's interaction with the item.

[0015] In combination with any embodiment of the present application, the statistical results of the interaction between the target object and the item include the statistical results of the interaction between the target object and items of each subject type, and / or the statistical results of the interaction between the target object and items of each item type.

[0016] In combination with any embodiment of the present application, determining a description of the target object based on the second interaction record includes:

[0017] Based on the second interaction record and the first preset template, a description of the target object is determined, wherein the first preset template includes multiple statistical result fields, and the multiple statistical result fields include at least one of the following: a time field, a subject type field to which the item belongs, an item type field, an interaction behavior field, and an interaction number field.

[0018] In combination with any embodiment of the present application, determining a description of the target object based on the second interaction record and the first preset template includes:

[0019] Based on the second interaction record, the attributes of the target object and the first preset template, a description of the target object is determined, and the multiple statistical result fields also include the attributes of the object.

[0020] In conjunction with any embodiment of the present application, before obtaining the target text based on the description of the target object and the description of the item to be recommended, the method further includes:

[0021] Based on the information of the item to be recommended, a description of the item to be recommended is determined, where the information of the item to be recommended includes one or more of the following: the content of the item to be recommended, the title of the item to be recommended, the theme type of the item to be recommended, and the interaction index of the item to be recommended.

[0022] In conjunction with any embodiment of the present application, determining a description of the item to be recommended based on the information of the item to be recommended includes:

[0023] Based on the information of the item to be recommended and a second preset template, a description of the item to be recommended is determined.

[0024] In conjunction with any embodiment of the present application, inputting the target text into the recommendation model to obtain the target recommendation order output by the recommendation model includes:

[0025] The target text is input into the recommendation model to obtain a target recommendation order and at least one target reason output by the recommendation model, wherein the at least one target reason is a reason for determining that the at least one target item is recommended to the target object.

[0026] In combination with any embodiment of the present application, the target recommendation order is determined based on the at least one target reason.

[0027] In a second aspect, a training method is provided, the training method comprising:

[0028] Obtaining a first model to be trained and a training text, wherein the training text includes a description of a training object and a description of a training item;

[0029] Inputting the training text into the first to-be-trained model to obtain a training recommendation order output by the first to-be-trained model, wherein the training recommendation order includes at least one first candidate item among the training items, and the at least one first candidate item is an item recommended to the training subject;

[0030] Based on the training recommendation order, the parameters of the first model to be trained are updated to obtain a recommended model.

[0031] In combination with any embodiment of the present application, the updating of the parameters of the first to-be-trained model based on the training recommendation order to obtain the recommended model includes:

[0032] Determining, based on the training recommendation order, a first indicator and a second indicator, wherein the first indicator represents the degree of interest of the training subject in the at least one first candidate item, and the second indicator represents the degree of matching of the at least one first candidate item with target requirements, wherein the target requirements include requirements for items recommended to the subject;

[0033] Based on the first indicator and the second indicator, the parameters of the first model to be trained are updated to obtain the recommendation model.

[0034] In combination with any embodiment of the present application, the updating of the parameters of the first to-be-trained model based on the first indicator and the second indicator to obtain the recommendation model includes:

[0035] Based on the first indicator, the second indicator and the third indicator, the parameters of the first model to be trained are updated to obtain the recommendation model, and the third indicator represents whether the first model to be trained outputs at least one training reason for recommending the at least one first candidate item.

[0036] In combination with any embodiment of the present application, the updating of the parameters of the first to-be-trained model based on the first indicator and the second indicator to obtain the recommendation model includes:

[0037] Determining a reward value of the first to-be-trained model based on the first indicator and the second indicator;

[0038] Based on the reward value, the parameters of the first model to be trained are updated to obtain the recommendation model.

[0039] In combination with any embodiment of the present application, the target requirements include: the diversity of items recommended to the object meets the preset requirements, the items recommended to the object include items of preset types, the proportion of new items in the items recommended to the object is greater than or equal to a first threshold, and the new items are items whose release time is no more than a preset value from the current time.

[0040] In combination with any embodiment of the present application, obtaining the first model to be trained includes:

[0041] Obtaining a training prompt word, wherein the training prompt word includes a description of the training subject and the training item, and the training prompt word is used to guide the second to-be-trained model to determine an item of interest to the training subject from the training items based on the description of the training subject;

[0042] Inputting the training prompt word into the second model to be trained to obtain at least one second candidate item determined by the second model to be trained from the training items;

[0043] Based on the at least one second candidate item, the second model to be trained is updated to obtain the first model to be trained.

[0044] In combination with any embodiment of the present application, updating the second model to be trained based on the at least one second candidate item to obtain the first model to be trained includes:

[0045] determining a loss of the second to-be-trained model based on a difference between the at least one second candidate item and a positive sample, the positive sample including the item of interest to the training subject;

[0046] Based on the loss of the second model to be trained, the parameters of the second model to be trained are updated to obtain the first model to be trained.

[0047] In a third aspect, a recommendation device is provided, comprising:

[0048] an acquisition unit, configured to acquire a recommendation model and a target text, wherein the target text includes a description of a target object and a description of an item to be recommended;

[0049] A processing unit is used to input the target text into the recommendation model to obtain a target recommendation order output by the recommendation model, wherein the target recommendation order includes at least one target item among the items to be recommended, and the at least one target item is an item recommended to the target object.

[0050] In combination with any embodiment of the present application, the acquiring unit is further configured to:

[0051] Obtaining a first interaction record between the target object and the item;

[0052] Determining a description of the target object by aggregating the first interaction records;

[0053] The target text is obtained based on the description of the target object and the description of the item to be recommended.

[0054] In combination with any embodiment of the present application, the acquiring unit is further configured to:

[0055] Determining a second interaction record by aggregating the first interaction record, where the second interaction record includes a record of the target object interacting with the item within a first time period and / or a record of the target object interacting with the item within a second time period, where a time span of the first time period is shorter than a time span of the second time period;

[0056] Based on the second interaction record, a description of the target object is determined.

[0057] In combination with any embodiment of the present application, the record of the target object's interaction with the item in the first time period and the record of the target object's interaction with the item in the second time period both include statistical results of the target object's interaction with the item.

[0058] In combination with any embodiment of the present application, the statistical results of the interaction between the target object and the item include the statistical results of the interaction between the target object and items of each subject type, and / or the statistical results of the interaction between the target object and items of each item type.

[0059] In combination with any embodiment of the present application, the acquiring unit is further configured to:

[0060] Based on the second interaction record and the first preset template, a description of the target object is determined, wherein the first preset template includes multiple statistical result fields, and the multiple statistical result fields include at least one of the following: a time field, a subject type field to which the item belongs, an item type field, an interaction behavior field, and an interaction number field.

[0061] In combination with any embodiment of the present application, the acquiring unit is further configured to:

[0062] Based on the second interaction record, the attributes of the target object and the first preset template, a description of the target object is determined, and the multiple statistical result fields also include the attributes of the object.

[0063] In combination with any embodiment of the present application, the processing unit is further configured to:

[0064] Based on the information of the item to be recommended, a description of the item to be recommended is determined, where the information of the item to be recommended includes one or more of the following: the content of the item to be recommended, the title of the item to be recommended, the theme type of the item to be recommended, and the interaction index of the item to be recommended.

[0065] In combination with any embodiment of the present application, the processing unit is further configured to:

[0066] Based on the information of the item to be recommended and a second preset template, a description of the item to be recommended is determined.

[0067] In combination with any embodiment of the present application, the processing unit is further configured to:

[0068] The target text is input into the recommendation model to obtain a target recommendation order and at least one target reason output by the recommendation model, wherein the at least one target reason is a reason for determining that the at least one target item is recommended to the target object.

[0069] In combination with any embodiment of the present application, the target recommendation order is determined based on the at least one target reason.

[0070] In a fourth aspect, a training device is provided, comprising:

[0071] an acquiring unit, configured to acquire a first model to be trained and a training text, wherein the training text includes a description of a training object and a description of a training item;

[0072] a processing unit, configured to input the training text into the first to-be-trained model, and obtain a training recommendation order output by the first to-be-trained model, wherein the training recommendation order includes at least one first candidate item among the training items, the at least one first candidate item being an item recommended to the training subject;

[0073] An updating unit is used to update the parameters of the first model to be trained based on the training recommendation order to obtain a recommended model.

[0074] In combination with any embodiment of the present application, the updating unit is further configured to:

[0075] Determining, based on the training recommendation order, a first indicator and a second indicator, wherein the first indicator represents the degree of interest of the training subject in the at least one first candidate item, and the second indicator represents the degree of matching of the at least one first candidate item with target requirements, wherein the target requirements include requirements for items recommended to the subject;

[0076] Based on the first indicator and the second indicator, the parameters of the first model to be trained are updated to obtain the recommendation model.

[0077] In combination with any embodiment of the present application, the updating unit is further configured to:

[0078] Based on the first indicator, the second indicator and the third indicator, the parameters of the first model to be trained are updated to obtain the recommendation model, and the third indicator represents whether the first model to be trained outputs at least one training reason for recommending the at least one first candidate item.

[0079] In combination with any embodiment of the present application, the updating unit is further configured to:

[0080] Determining a reward value of the first to-be-trained model based on the first indicator and the second indicator;

[0081] Based on the reward value, the parameters of the first model to be trained are updated to obtain the recommendation model.

[0082] In combination with any embodiment of the present application, the target requirements include: the diversity of items recommended to the object meets the preset requirements, the items recommended to the object include items of preset types, the proportion of new items in the items recommended to the object is greater than or equal to a first threshold, and the new items are items whose release time is no more than a preset value from the current time.

[0083] In combination with any embodiment of the present application, the acquiring unit is further configured to:

[0084] Obtaining a training prompt word, wherein the training prompt word includes a description of the training subject and the training item, and the training prompt word is used to guide the second to-be-trained model to determine an item of interest to the training subject from the training items based on the description of the training subject;

[0085] Inputting the training prompt word into the second model to be trained to obtain at least one second candidate item determined by the second model to be trained from the training items;

[0086] Based on the at least one second candidate item, the second model to be trained is updated to obtain the first model to be trained.

[0087] In combination with any embodiment of the present application, the acquiring unit is further configured to:

[0088] determining a loss of the second to-be-trained model based on a difference between the at least one second candidate item and a positive sample, the positive sample including the item of interest to the training subject;

[0089] Based on the loss of the second model to be trained, the parameters of the second model to be trained are updated to obtain the first model to be trained.

[0090] In a fifth aspect, an electronic device is provided, comprising: a processor and a memory, the memory being used to store computer program code, the computer program code comprising computer instructions, and when the processor executes the computer instructions, the electronic device executes the first aspect and any one of its embodiments; when the processor executes the computer instructions, the electronic device either executes the second aspect and any one of its embodiments.

[0091] In the sixth aspect, another electronic device is provided, comprising: a processor, a sending device, an input device, an output device and a memory, wherein the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the first aspect and any embodiment thereof; when the processor executes the computer instructions, the electronic device may execute the second aspect and any embodiment thereof.

[0092] In the seventh aspect, a computer-readable storage medium is provided, in which a computer program is stored. The computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the first aspect and any embodiment thereof; when the program instructions are executed by the processor, the processor is caused to execute the second aspect and any embodiment thereof.

[0093] In an eighth aspect, a computer program product is provided, which includes a computer program or instructions, which, when the computer program or instructions are run on a computer, enable the computer to execute the above-mentioned first aspect and any embodiment thereof; or, when the computer program or instructions are run on a computer, enable the computer to execute the above-mentioned second aspect and any embodiment thereof.

[0094] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application.

[0095] In the present application, the description of the target object includes information about the target object, and the description of the item to be recommended includes information about the item to be recommended. After obtaining the recommendation model and the target text including the description of the target object and the description of the item to be recommended, the recommendation device inputs the target text into the recommendation model, which enables the recommendation model to determine the information about the target object and the item to be recommended based on the description of the target object and the description of the item to be recommended. Furthermore, based on the information about the target object and the information about the item to be recommended, the recommendation model can determine at least one target item to be recommended to the target object from the items to be recommended, and determine a target recommendation order for the at least one target item. This allows the recommendation model to determine the target recommendation order for at least one target item to be recommended to the target object in an end-to-end manner, thereby improving the accuracy of the target recommendation order. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the drawings required for use in the embodiments of the present application or the background technology will be described below.

[0097] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0098] Figure 1 A flowchart of a recommended method provided in an embodiment of the present application;

[0099] Figure 2 A flowchart of another recommended method provided in an embodiment of the present application;

[0100] Figure 3 A flowchart of a training method provided in an embodiment of the present application;

[0101] Figure 4 A flowchart of a training method and a recommendation method provided in an embodiment of the present application;

[0102] Figure 5 A schematic diagram of the structure of a recommended device provided in an embodiment of the present application;

[0103] Figure 6 A schematic diagram of the structure of a training device provided in an embodiment of the present application;

[0104] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0105] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0106] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0107] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0108] Before introducing the technical solutions of the embodiments of the present application, some technical terms in the embodiments of the present application are first introduced.

[0109] The object can be a user of the Internet platform.

[0110] Items can be any item on the Internet platform. Items can include notes, images, videos, text, audio, and products.

[0111] Interaction between an object and an item refers to an interaction between the object and the item. This interaction includes one or more of the following: clicking, browsing, liking, commenting, forwarding, adding to favorites, and purchasing. Interaction between an object and an item indicates the object's interest in the item. For example, if a subject browses a text related to sports, this indicates an interest in sports.

[0112] Prompts are used to guide the model through its tasks. Prompts include input data and instructions. Input data includes the data needed to perform the task. For example, if the task is translation, input data includes the data to be translated. Instructions include the task to be performed. For example, an instruction might include: "Please translate the input data into Chinese."

[0113] Optionally, the prompt word also includes one or more of the following: context and format requirements. Context is used to provide background information for performing the task or examples related to performing the task. Format requirements include requirements for the format of the output data.

[0114] Currently, in the context of information overload on the internet, recommendation systems have become a key technology for improving user experience and promoting business growth. To meet users' personalized needs, recommendation systems typically build recommendation models (such as the two-tower model). Based on data such as the object's historical behavior and characteristics, they use algorithms such as collaborative filtering and deep learning to explore the user's potential interests and thus identify items that the user may be interested in. However, the initial recommendation results generated by the model often have limitations, resulting in a low match between the initial recommendation results and the recommendation requirements. For example, if the recommendation requirement includes recommending videos, the initial recommendation results generated by the model are all text, and the initial recommendation results do not match the recommendation requirements. Therefore, after the model generates the initial recommendation results, it is necessary to further adjust the initial recommendation results based on the recommendation requirements to obtain the final recommendation results.

[0115] In other words, the current technology divides the process into two independent stages: obtaining the initial recommendation results based on the model to determine the items of interest to the object, and obtaining the final recommendation results by adjusting the initial recommendation results based on the recommendation requirements. That is, the current technology cannot determine the final recommendation results in an end-to-end manner.

[0116] During the initial recommendation stage, the model is guided by the subject's interest in the item, without considering whether the item matches the recommendation requirements. Furthermore, during the final recommendation stage, the model is guided by the same goal of whether the item matches the recommendation requirements, without considering the subject's interest in the item. Therefore, the final recommendation results, determined through these two stages, struggle to simultaneously satisfy both the following conditions: the subject's interest in the item in the final recommendation results and the final recommendation results matching the recommendation requirements. This can easily lead to low accuracy in the final recommendation results.

[0117] For example, suppose three items are recommended to a subject. The model's initial recommendation results include the subject's three most interesting items: text t1, text t2, and text t3. Since the recommendation requirement includes recommending a video to the subject, when adjusting the initial recommendation results based on the recommendation requirement, text t1 is replaced with video v1. Accordingly, the final recommendation results include video v1, text t2, and text t3. However, if the subject's interest in video v1 is low, recommending the final recommendation results to the subject may reduce their interest in the final recommendation results, thereby reducing the subject's user experience.

[0118] Based on this, an embodiment of the present application provides a recommendation method to determine items recommended to an object in an end-to-end manner, thereby improving the accuracy of items recommended to the object.

[0119] The recommendation method in the embodiments of the present application is performed by a recommendation device, which can be any electronic device capable of executing the technical solution disclosed in the embodiments of the present application. Optionally, the recommendation device can be one of the following: a computer or a server.

[0120] It should be understood that the method embodiment of the present application can also be implemented by a processor executing computer program code. The following describes the embodiment of the present application in conjunction with the drawings in the embodiment of the present application. Figure 1 , Figure 1 A flowchart of a recommended method provided in an embodiment of the present application.

[0121] 101. Obtain a recommendation model and a target text, where the target text includes a description of the target object and a description of the recommended item.

[0122] In the embodiment of the present application, the structure of the recommendation model can be any structure. The recommendation model has the ability to determine the items recommended to the object, and the ability to sort the items recommended to the object in order. Among them, the higher the order of the item, the higher the priority of the item in being recommended to the object. In one possible implementation method, the item is recommended to the object by displaying it on the page. The higher the order of the item, the closer the position of the item is to the top of the page. In another possible implementation method, the item is recommended to the object by displaying the items recommended to the object in sequence. The higher the order of the item, the earlier the item is displayed.

[0123] Optionally, the recommendation model's ability to determine items to recommend to a subject is acquired through training on a recommendation task. The recommendation task involves determining items of interest to the subject, i.e., the purpose of the recommendation task includes determining items of interest to the subject. Accordingly, when the recommendation model is trained on the recommendation task, the recommendation model is capable of determining items to recommend to the subject based on the items of interest to the subject.

[0124] In one possible implementation, a recommendation model can determine whether the subject is interested in an item based on a description of the subject, and thus determine which items to recommend to the subject. For example, if the subject's characteristics include that the subject likes digital products and the item is related to mobile phones, then the subject can be determined to have a high probability of being interested in the item, and thus mobile phone-related items can be recommended to the subject. For another example, if the subject's characteristics include that the items the subject has interacted with include items related to mobile phones, then based on the subject's characteristics, it can be determined that the subject is interested in mobile phones, and therefore, the subject can be determined to have a high probability of being interested in mobile phone-related items, and thus, mobile phone-related items can be recommended to the subject. For another example, if the subject's characteristics include that the subject browsed computer-related items and then keyboard-related items, then based on the subject's browsing history, it can be inferred that the subject is likely to browse computer peripherals other than keyboards. Therefore, it can be determined that the subject is likely to be interested in other peripherals (such as mice), and thus, items related to other peripherals can be recommended to the subject. These other peripheral items are items related to computer peripherals other than keyboards.

[0125] In another possible implementation, the recommendation model may determine whether the object is interested in the item based on the similarity between the object and the item, and further determine the item to recommend to the object.

[0126] In another possible implementation, items that a subject has interacted with are called interacted items. Based on the similarity between interacted items and items, the recommendation model can determine whether the subject is interested in the items and, therefore, recommend items to the subject. For example, if object d1 has interacted with item i1, then item i1 is considered an interacted item for object d1. Based on the similarity between item i1 and item i2, the recommendation model can determine whether object d1 is interested in item i2 and, therefore, whether to recommend item i2 to object d1.

[0127] In another possible implementation, objects that have interacted with an item are called interacting objects. Based on the similarity between interacting objects and objects, the recommendation model can determine whether the object is interested in the item and, therefore, recommend items to the object. For example, if object d1 has interacted with item i1, then object d1 is considered an interacting object for item i1. Based on the similarity between object d1 and object d2, the recommendation model can determine whether object d2 is interested in item i1 and, therefore, whether to recommend item i1 to object d2.

[0128] Optionally, the recommendation model determines the order of items recommended to the subject based on target requirements, wherein the target requirements include requirements for the items recommended to the subject.

[0129] Optionally, the target requirements include one or more of the following: the diversity of items recommended to the object meets preset requirements, the items recommended to the object include items of preset item types, the item with the highest recommendation order among the items recommended to the object is an item of preset item type, the similarity between any two items recommended to the object is less than or equal to a similarity threshold, the proportion of new items in the items recommended to the object is greater than or equal to a first threshold, and the new items are items whose release time does not exceed a second threshold from the current time.

[0130] The diversity of an item includes one or more of the following: the number of item types and the number of topic types. Item types include text, video, image, audio, and multimodal. Multimodal includes one or more of the following: text, video, image, and audio. The topic type of an item indicates the content of the item. For example, a document related to computers has a topic type of "computers," while a video related to food has a topic type of "food."

[0131] When the diversity of the items includes the number of item types of the items, the preset requirement includes that the number of item types of the items is greater than or equal to an item type threshold. If the number of item types of the items is greater than or equal to the item type threshold, it indicates that the number of item types of the items is greater, which means that the diversity of the items is greater.

[0132] When the diversity of items includes the number of thematic types of items, the preset requirement includes that the number of thematic types of items is greater than or equal to a thematic type threshold. If the number of thematic types of items is greater than or equal to the thematic type threshold, it indicates that the number of thematic types of items is large, which in turn indicates a greater diversity of items.

[0133] Therefore, in the case where the diversity of items recommended to the object meets the preset requirements, the diversity of items recommended to the object is richer.

[0134] The preset item type can be any item type. When the items recommended to the target include items of the preset item type, items of the preset item type can be better promoted. For example, if the target is a user of the target object and the recommended items are items published on the target platform, then if the target platform needs to promote videos, the preset item type can be set to videos, and the items recommended to users of the target platform can include videos.

[0135] New items are those whose release time is no more than the second threshold from the current time, indicating that the new items are recently released. If the proportion of new items in the recommended items is greater than or equal to the first threshold, it means that the proportion of new items in the recommended items is high.

[0136] The target recommendation order of an item is its recommendation priority. The higher the target recommendation order of an item, the higher it will be recommended to the recipient. If the item with the highest target recommendation order among the items recommended to the recipient belongs to a preset item type, the items of the preset item type can be promoted more effectively.

[0137] If the similarity between two items is less than or equal to the similarity threshold, it means that the degree of homogeneity between the two items is low. Therefore, if the similarity between any two items among the items recommended to the object is less than or equal to the similarity threshold, the diversity of the items recommended to the object can be enriched.

[0138] Optionally, training a recommendation model based on a recommendation task can enable the recommendation model to determine items recommended to a subject based on the subject's interests. Training a recommendation model based on target requirements can enable the recommendation model to determine the order of items recommended to a subject based on the target requirements. Therefore, when a recommendation model is trained based on a recommendation task and target requirements, the recommendation model can both determine items recommended to a subject based on the subject's interests and determine the order of items recommended to the subject based on the target requirements.

[0139] In the embodiment of the present application, the target text describes the target object and the recommended items, wherein the description of the target object and the description of the recommended items are both natural language descriptions.

[0140] Optionally, the description of the target object includes a description of the target object's attributes and / or a description of the first interaction record between the target object and an item. The target object interacts with at least one item, and accordingly, the first interaction record includes a record of the target object's interaction with the at least one item. For example, the description of the target object's attributes may include: the target object is 30 years old, male, and enjoys food and travel. The description of the first interaction record between the target object and an item may include: the target object viewed two videos related to food item A and saved one graphic note related to a travel guide to place B.

[0141] In an implementation of determining a description of a target object, after obtaining a first interaction record between the target object and an item, the recommendation device determines a description of the target object by aggregating the first interaction record.

[0142] In this implementation, by aggregating the first interaction record, the first interaction record can be summarized and concluded. Determining the description of the target object through aggregation can make the description of the target object more concise than that of the first interaction record, thereby reducing the amount of data in the description of the target object compared to the first interaction record.

[0143] Optionally, the recommendation device implements "determining a description of the target object by aggregating the first interaction records" by performing the following steps: aggregating the first interaction records to determine second interaction records, wherein the second interaction records include records of the target object interacting with the item within a first time period and / or records of the target object interacting with the item within a second time period, the first time period being shorter than the second time period. Based on the second interaction records, a description of the target object is determined.

[0144] The first time period is shorter than the second time period, indicating that the first time period is short-term compared to the second time period, while the second time period is long-term compared to the first time period. The second interaction record includes records of the target object's interactions with the item over a short period of time and / or records of the target object's interactions with the item over a long period of time. For example, the first time period may be the last seven days, while the second time period may be the last year.

[0145] Optionally, the recommendation device implements "determining a description of the target object based on the second interaction record" by performing the following steps: determining a description of the target object based on the second interaction record and a first preset template, wherein the first preset template includes multiple statistical result fields, and the multiple statistical result fields include at least one of the following: a time field, a subject type field to which the item belongs, an item type field, an interaction behavior field, and an interaction count field. The time field is used to describe the first time period and / or the second time period, the subject type field is used to describe the subject type of the item interacting with the target object, the item type field is used to describe the item type of the item interacting with the target object, the interaction behavior field is used to describe the interaction behavior between the target object and the item, and the interaction count field is used to describe the number of interactions between the target object and the item.

[0146] The fields in the first preset template can all be changed based on the second interaction record. Accordingly, based on the second interaction record and the first preset template, the description of the target object determined includes at least one of the following: the first time period and / or the second time period, the theme type of the item interacting with the target object, the item type of the item interacting with the target object, the interaction behavior between the target object and the item, and the number of interactions between the target object and the item. For example, the description of the target object includes: the target object has browsed videos related to food A twice in the past 7 days, and has collected 1 graphic note related to the travel guide to place B in the past year. Among them, the past 7 days are the first time period, the past year is the second time period, browsing and collecting are both interactive behaviors, 2 times and 1 are both the number of interactions, food A and the travel guide to place B are both theme types of items, and videos and graphic notes are both item types of items.

[0147] Optionally, the recommendation device implements "determining a description of the target object based on the second interaction record and the first preset template" by performing the following steps: determining a description of the target object based on the second interaction record, the attributes of the target object, and the first preset template, wherein the plurality of statistical result fields also include the attributes of the object. Specifically, the plurality of statistical result fields include the attributes of the object and at least one of the following: a time field, a field for the topic type to which the item belongs, an item type field, an interaction behavior field, and a field for the number of interactions.

[0148] Since the fields in the first preset template can all be changed based on the second interaction record, the description of the target object determined based on the second interaction record and the first preset template includes the attributes of the target object and at least one of the following: the first time period and / or the second time period, the subject type of the item interacting with the target object, the item type of the item interacting with the target object, the interaction behavior between the target object and the item, and the number of interactions between the target object and the item.

[0149] For example, the target object's description includes: the target object's age is 30, the target object's gender is male, and the target object likes food. The target object has viewed two videos related to food A in the last seven days and has collected one picture and text note related to a travel guide to place B in the last year. Among them, 30 years old, male, and a love of food are all attributes of the target object. The last seven days are the first time period, and the last year is the second time period. Viewing and collecting are both interactive behaviors, 2 and 1 are both the number of interactions. Food A and travel guide to place B are both the subject types of items, and videos and picture and text notes are both the item types of items.

[0150] After determining a description of the target object based on the second interaction record, the target text includes the second record. When determining items to recommend to the target object based on the target text, the items recommended to the target object can be determined based on the second interaction record. Furthermore, when determining items to recommend to the target object based on the second interaction record, the items recommended to the target object can be determined based on records of the target object's interactions with the items over a short period of time and / or records of the target object's interactions with the items over a long period of time.

[0151] Records of the target object's interactions with items in the short term are conducive to performing one or more of the following: determining the target object's immediate interests, determining the target object's temporary interests, verifying the target object's long-term interests, and correcting the target object's long-term interests. For example, if the target object's records of interactions with items in the short term include browsing items related to travel guides in the past 7 days, then it can be determined that the target object's immediate interests include travel. For another example, if the target object's records of interactions with items in the short term include browsing items related to medical health in the past 7 days, then it can be inferred that the target object or the target object's family may be unwell and therefore need to learn about medical health. In this case, it can be determined that the target object's temporary interests include medical health. For another example, if the target object's records of interactions with items in the short term include browsing items related to video editing in the past 7 days, and the target object's long-term interests include image processing, since video editing belongs to image processing, it can be verified that the target object's long-term interests are correct. For another example, the target object's short-term interaction record with objects includes browsing objects related to video clips in the past seven days, and the target object's long-term interests include image processing. Since image processing has a coarser granularity than video clips, correcting the target object's long-term interests to video clips can improve the accuracy of the target object's long-term interests.

[0152] Records of a target's interactions with items over a long period of time can be used to perform one or more of the following: determining the target's long-term interests; and predicting the target's future interests. For example, if a target's long-term interactions with items include purchasing ski equipment in December each year, it can be predicted that the target's interest in the next December will include skiing.

[0153] Optionally, the records of the target object's interactions with the item during the first time period and the records of the target object's interactions with the item during the second time period both include statistical results of the target object's interactions with the item. For example, the statistical results may include that the target object viewed items related to food A twice and collected an item related to a travel guide to place B.

[0154] Because the statistical results of items are smaller than the data related to the items, if both the records of the target object's interaction with the item in the first time period and the records of the target object's interaction with the item in the second time period include the statistical results of the target object's interaction with the item, the data processing workload for the second records can be reduced.

[0155] For example, in an embodiment of the present application, the statistical result of the interaction between the target object and the item is: the target object browsed 2 videos related to food A and collected 1 graphic note related to the travel guide to place B. In the current technology, the record of the interaction between the target object and the item includes: the embedded representation (or embedded vector) of video v1 related to food A, the embedded representation (or embedded vector) of video v2 related to food A, and the embedded representation (or embedded vector) of graphic note p1 related to the travel guide to place B. The embedded representation (or embedded vector) of video v1, the embedded representation (or embedded vector) of video v2, and the embedded representation (or embedded vector) of graphic note p1 are all data related to the item. In this example, the statistical result is smaller than the amount of embedded representation data of the item in the current technology.

[0156] Optionally, the statistical results of the target subject's interactions with items include statistical results of the target subject's interactions with items of each item type. Items with the same theme type are items of the same item type. For example, item i1 and item i2 are both related to food, i.e., the theme type of item i1 and the theme type of item i2 are both food. Therefore, item i1 and item i2 are items of the same item type. The statistical results of the target subject's interactions with items include that the target subject viewed food-related items twice.

[0157] In the embodiment of the present application, the recommended items can be any items. Optionally, the recommended items are items published on the target platform. For example, if the target platform is an internet platform, the subject can publish text, video, and other items on the target platform. The recommended items are texts and videos published on the target platform.

[0158] In an implementation method for determining a description of an item to be recommended, a recommendation device determines a description of the item to be recommended based on information about the item to be recommended, wherein the information about the item to be recommended includes one or more of the following: content of the item to be recommended, title of the item to be recommended, theme type of the item to be recommended, and interaction index of the item to be recommended.

[0159] The interaction indicator of a recommended item indicates the probability of interaction between the recommended item and the target audience. Optionally, the interaction indicator includes one or more of the following: number of exposures, number of favorites, number of reposts, number of likes, number of comments, number of purchases, number of clicks, number of views, and duration of views. The more times a recommended item is exposed, the higher the probability of interaction between the recommended item and the target audience. The more times a recommended item is favorited, the higher the probability of interaction between the recommended item and the target audience. The more times a recommended item is reposted, the higher the probability of interaction between the recommended item and the target audience. The more times a recommended item is liked, the higher the probability of interaction between the recommended item and the target audience. The more times a recommended item is commented on, the higher the probability of interaction between the recommended item and the target audience. The more times a recommended item is purchased, the higher the probability of interaction between the recommended item and the target audience. The more times a recommended item is clicked, the higher the probability of interaction between the recommended item and the target audience. The more times a recommended item is viewed, the higher the probability that the recommended item will interact with the object. The longer the recommended item is viewed, the higher the probability that the recommended item will interact with the object.

[0160] For example, the description of the recommended item includes: text t1 describes the food in place A, the tag of text t1 is food, the exposure number of text t1 is n1, and the number of times text t1 has been viewed is n2.

[0161] Optionally, the recommendation device implements “determining a description of the item to be recommended based on the information of the item to be recommended” by executing the following steps: determining a description of the item to be recommended based on the information of the item to be recommended and a second preset template.

[0162] Optionally, the second preset template includes at least one of the following: an item content field, an item title field, an item theme type field, and an item interaction index field. The content field is used to describe the content of the item to be recommended, the theme type field is used to describe the theme type of the item to be recommended, the item type field is used to describe the item type of the recommended item, and the interaction index field is used to describe the interaction index of the recommended item.

[0163] 102. Input the target text into the recommendation model to obtain a target recommendation order output by the recommendation model, wherein the target recommendation order includes at least one target item among the items to be recommended, and the at least one target item is an item recommended to the target object.

[0164] After the recommendation device inputs the target text into the recommendation model, the recommendation model can determine at least one target item to be recommended to the target object from the items to be recommended based on the description of the target object and the description of the item to be recommended, and determine the order of each target item in the at least one target item based on the recommendation priority to obtain the target recommendation order.

[0165] Optionally, the recommendation model is a large language model (LLM). Since the knowledge that can be acquired through training of an LLM is richer than that acquired through training of models other than the LLM (such as the twin-tower model), determining the target recommendation order based on the LLM can utilize richer knowledge to determine the target recommendation order, thereby effectively alleviating the information cocoon effect and improving the accuracy of the target recommendation order.

[0166] It should be understood that since the data modality that LLM can process is text, the modality of the data input to LLM should be text. Therefore, it is necessary to convert the information required to determine the target recommendation order (including information about the target object and information about the items to be recommended) into a natural language description and input this natural language description into the recommendation model so that the recommendation model can determine the target recommendation order based on the information about the target object and the information about the items to be recommended. Therefore, by inputting the target text including a description of the target object and a description of the item to be recommended into the recommendation model, the recommendation device can enable the recommendation model to determine the target recommendation order based on the description of the target object and the description of the item to be recommended.

[0167] Optionally, the recommendation model determines at least one item in which the target object is interested from the items to be recommended based on the description of the target object and the description of the item to be recommended, and then determines at least one target item to be recommended to the target object based on the at least one item in which the target object is interested.

[0168] Optionally, the recommendation model is trained based on the recommendation task and target requirements. Accordingly, the recommendation model can utilize the ability to determine the items recommended to the object based on the items in which the object is interested, and determine at least one target item from the items to be recommended. At the same time, the ability to determine the order of items recommended to the object based on the target requirements can be utilized to determine the order of each target item in at least one target item to obtain a target recommendation order. In this way, at least one target item in the target recommendation order can be an item of interest to the target object, and the target recommendation order can be matched with the target requirements. In this way, the recommendation model can determine the target recommendation order of at least one target item recommended to the target object in an end-to-end manner, thereby improving the accuracy of the target recommendation order.

[0169] In one possible implementation scenario, the items to be recommended are items published on a target platform, and the target objects are users of the target platform. Thus, the recommendation device can determine at least one item to recommend to the user from the items published on the target platform based on the recommendation model.

[0170] In an embodiment of the present application, the description of the target object includes information about the target object, and the description of the item to be recommended includes information about the item to be recommended. After obtaining the recommendation model and the target text including the description of the target object and the description of the item to be recommended, the recommendation device inputs the target text into the recommendation model, which enables the recommendation model to determine the information about the target object and the item to be recommended based on the description of the target object and the description of the item to be recommended. Furthermore, based on the information about the target object and the information about the item to be recommended, the recommendation model can determine at least one target item to be recommended to the target object from the items to be recommended, and determine a target recommendation order for the at least one target item. This allows the recommendation model to determine the target recommendation order for at least one target item to be recommended to the target object in an end-to-end manner, thereby improving the accuracy of the target recommendation order.

[0171] As an optional embodiment, the recommendation device performs the following steps during the process of executing the step of "inputting the target text into the recommendation model and obtaining the target recommendation order output by the recommendation model": inputting the target text into the recommendation model, obtaining the target recommendation order and at least one target reason output by the recommendation model, wherein the at least one target reason is the reason for determining at least one target item as an item recommended to the target object.

[0172] In this embodiment of the present application, the at least one target reason is the reason why the recommendation model recommends at least one target item to the target subject. Optionally, the target reason is determined based on one or more of the following: the target subject's interests, the target subject's interaction history with the item, the historical sequence of interactions between the target subject and the item, and the purpose of the target subject's interaction with the item.

[0173] For example, the at least one target item includes a mouse-related item, a display-related item, and a game-related item. The reason for recommending the mouse-related item to the target object and the reason for recommending the display-related item to the target object is that the target object is interested in computer peripherals. The reason for recommending the game-related item to the target object is that the target object's interaction history with items includes game-related items.

[0174] For another example, at least one target item includes a mouse-related item, a monitor-related item, and a game-related item. The reason for recommending mouse-related items to the target object and the reason for recommending monitor-related items to the target object are: the target object's historical interaction sequence with items includes: after browsing computer-related items, the target object also browsed keyboard-related items. Based on the object's historical interaction sequence, it can be inferred that the target object is more likely to browse computer peripherals other than keyboards next. The reason for recommending game-related items to the target object is: the target object's browsing of computer peripherals may be for the purpose of configuring a computer, and the purpose of configuring a computer may be to play games. Therefore, after recommending computer peripherals to the target object, game-related items may continue to be recommended to the target object.

[0175] Optionally, the target reasons correspond to the target items one-to-one, that is, the recommendation model outputs at least one target item and also outputs the reason for recommending each target item to the target object.

[0176] Optionally, one target reason corresponds to more than one target item. For example, the reason why the recommendation model recommends keyboard-related items to the target object and the reason why it recommends monitor-related items to the target object are both target reason r1, where target reason r1 is: the target object is interested in computer peripherals.

[0177] Optionally, more than one reason may be associated with a target item. For example, the reasons for a recommendation model to recommend game-related items to a target subject include target reason r2 and target reason r3. Target reason r2 is that the target subject's purpose for browsing computer peripherals may be to configure a computer, and the purpose of configuring a computer may be to play games. Therefore, after recommending computer peripherals to the target subject, game-related items may be recommended to the target subject. Target reason r3 is that the target subject's interests include games.

[0178] On one hand, the recommendation model outputs at least one target reason along with at least one target item, allowing the recommendation model to determine the at least one target item based on the at least one target reason, thereby improving the accuracy of the at least one target item. On the other hand, the recommendation model can be used to determine the accuracy of the at least one target item through the output of the at least one target reason. Optionally, a high accuracy of the target reason indicates that the recommendation model has a high accuracy in recommending the target item corresponding to the target reason to the target subject.

[0179] Optionally, the accuracy of the target reason is determined based on the attributes of the target object and / or the first interaction record between the target object and the item. For example, if the target reason is that the items that the target object is interested in include computer peripherals, if it is determined based on the description of the target object that the items that the target object is interested in do not include computer peripherals, then the accuracy of the target reason may be determined to be low.

[0180] In one possible implementation, the recommendation model first determines at least one target reason, and then determines at least one target item from among the items to be recommended based on the at least one target reason. Alternatively, the recommendation model is based on a chain of thought (COT), first determining at least one target reason, and then determining at least one target item from among the items to be recommended based on the at least one target reason.

[0181] The COT (Counter-Option-Typing) is a prompting strategy that instructs the recommendation model to identify at least one target item from a list of recommended items in a step-by-step manner. Based on the COT, the recommendation model divides the task of identifying at least one target item from a list of recommended items into x steps, where x is an integer greater than 1. Each step includes determining at least one target reason.

[0182] Optionally, by adjusting the temperature parameter (TP) of the recommendation model, the diversity of COT can be adjusted, thereby adjusting the diversity of the results (i.e., at least one target item) obtained by the recommendation model based on COT. Specifically, TP is positively correlated with the diversity of COT.

[0183] Optionally, after determining at least one target reason, the recommendation model determines a target recommendation order based on the at least one target reason.

[0184] See also Figure 2 , Figure 2 This is a flow chart of another recommended method provided in the embodiment of this application. Figure 2 As shown, the target object's attributes include cycling, fitness, beauty, fashion, travel, photography, food, career planning, ..., movies, and TV. The target object's first interaction records with items include those from the past 7 days and the past year. The past 7 days include 9 interactions with cycling-related items, 3 interactions with travel-related items, 5 interactions with fitness-related items, 2 interactions with technology-related items, 2 interactions with TV-related items, ... The past year includes 310 interactions with travel-related items, 429 interactions with fashion-related items, 272 interactions with cycling-related items, 2 interactions with career planning-related items, ...

[0185] After determining the description of the target object based on the attributes of the target object and the first interaction record, and determining the description of the item to be recommended based on the information of the item to be recommended, a target text can be generated based on the description of the target object and the description of the item to be recommended. After the target text is input into the recommendation model, the at least one target reason determined by the recommendation model includes: the target object has recently consumed 9 cycling-related content... As a cycling enthusiast who pursues freedom and health, his recent interests may focus on cycling equipment, route planning, and hiking, while he is also exploring photography and social activities. Based on at least one target reason, the recommendation model can determine at least one target item from the items to be recommended, as well as determine the target recommendation order of at least one target item. Figure 2 In the target recommendation order, the order of target items is: multimedia content related to cycling 1, multimedia content related to cycling equipment 2, multimedia content related to routes 3, multimedia content related to cycling 4, multimedia content related to photography 5, multimedia content related to hiking 6, multimedia content related to videography 7, multimedia content related to fitness 8, multimedia content related to photography 9, ...

[0186] The present application also provides a training method for training the recommendation model described above. The training method in the present application is performed by a training device, which can be any electronic device capable of executing the technical solutions disclosed in the present application. Optionally, the training device can be one of the following: a computer or a server.

[0187] It should be understood that the method embodiment of the present application can also be implemented by a processor executing computer program code. The following describes the embodiment of the present application in conjunction with the drawings in the embodiment of the present application. Figure 3 , Figure 3 A flowchart of a training method provided in an embodiment of the present application.

[0188] 301. Obtain a first model to be trained and a training text, wherein the training text includes a description of a training object and a description of a training item.

[0189] In the embodiment of the present application, the structure of the first model to be trained can be any structure.

[0190] Training items can be any items. They can be the same as or different from the items to be recommended. Optionally, training items can be items published on the target platform. For example, if the target platform is an internet platform, and the subject can publish text, videos, and other items on the target platform, the training items are text or videos published on the target platform.

[0191] It should be understood that the training method is the training process of the recommendation model, and the recommendation method is the inference process of the recommendation model. Therefore, there are multiple sets of corresponding content in the training method and the recommendation method. Among them, a set of corresponding content includes content in one training method and content in one recommendation method, and the content in the training method plays the same role in the training method as the content in the recommendation method plays in the recommendation method. Therefore, if there is corresponding content in the recommendation method for a certain content in the training method, the meaning and function of the content in the training method can be referred to the explanation of the corresponding content in the recommendation method, and the training method will not further explain the content.

[0192] In the embodiment of the present application, the training text includes a description of the training object and a description of the training item. The training text in the training method corresponds to the target text in the recommendation method. The meaning of the training text can be found in the description of the target text above.

[0193] The description of the training object is a natural language description. Optionally, the training device determines the description of the training object by aggregating the third interaction records. Based on the description of the training object and the description of the training item, a training text is obtained.

[0194] Optionally, the training device implements "determining a description of the training object by aggregating the third interaction records" by performing the following steps: aggregating the third interaction records to determine a fourth interaction record, wherein the fourth interaction record includes a record of the training object interacting with the object within a first time period and / or a record of the training object interacting with the object within a second time period, the first time period being shorter than the second time period. Based on the fourth interaction record, a description of the target object is determined.

[0195] Optionally, the record of the training subject's interaction with the object in the first time period and the record of the training subject's interaction with the object in the second time period both include statistical results of the training subject's interaction with the object.

[0196] Optionally, the statistical results of the interaction between the training object and the items include the statistical results of the interaction between the training object and items of each subject type, and / or the statistical results of the interaction between the training object and items of each item type.

[0197] Optionally, the training device implements “determining a description of the training object based on the fourth interaction record” by executing the following steps: determining a description of the training object based on the fourth interaction record and the first preset template.

[0198] Optionally, the training device implements "determining a description of the training object based on the fourth interaction record and the first preset template" by executing the following steps: determining a description of the training object based on the fourth interaction record, the attributes of the training object and the first preset template, and multiple statistical result fields also include the attributes of the object.

[0199] Optionally, before obtaining the target text based on the description of the training object and the description of the training item, the training device also determines the description of the training item based on the information of the training item. The information of the training item includes one or more of the following: the content of the training item, the title of the training item, the theme type of the training item, and the interaction indicator of the training item.

[0200] The description of the training item is also a natural language description. Optionally, the description of the training item includes one or more of the following: the content of the training item, the title of the training item, the tags of the training item, and the interaction index of the training item, where the interaction index of the training item indicates the probability of interaction between the training item and the subject.

[0201] Optionally, the training device implements “determining a description of the training item based on the information of the training item” by performing the following steps: determining a description of the training item based on the information of the training item and a second preset template.

[0202] 302. Input the training text into the first model to be trained, and obtain a training recommendation order output by the first model to be trained, wherein the training recommendation order includes at least one first candidate item among the training items, and the at least one first candidate item is an item recommended to the training subject.

[0203] After the training device inputs the training text into the first model to be trained, the first model to be trained can determine at least one first candidate item recommended to the training object from the training items based on the description of the training object and the description of the training items, and determine the order of each item in the at least one first candidate item based on the recommendation priority to obtain a training recommendation order.

[0204] Optionally, the first model to be trained is an LLM. Since the knowledge that can be acquired through training of an LLM is richer than that acquired through training of models other than the LLM (such as the twin-tower model), determining the target recommendation order based on the LLM can utilize the richer knowledge to determine the training recommendation order, thereby effectively alleviating the information cocoon effect and improving the accuracy of the training recommendation order.

[0205] 303. Based on the training recommendation order, update the parameters of the first model to be trained to obtain a recommended model.

[0206] Based on the training recommendation order, updating the parameters of the first model to be trained can make the training recommendation order output by the first model to be trained more accurate. Correspondingly, by updating the parameters of the first model to be trained to obtain a recommendation model, the accuracy of the target recommendation order determined by the recommendation model can also be improved.

[0207] In one possible implementation, the training device updates the parameters of the first model to be trained based on the difference between the training recommendation order and the ground truth (GT) of the recommendation order to obtain a recommendation model.

[0208] In this embodiment of the present application, the training text includes a description of the training subject and a description of the training items. After obtaining a first model to be trained and the training text, the training device inputs the training text into the first model to be trained, obtains a training recommendation order output by the first model to be trained, and the training recommendation order includes at least one first candidate item among the training items, wherein the at least one first candidate item is an item recommended to the training subject. Based on the training recommendation order, the parameters of the first model to be trained are updated to obtain a recommendation model.

[0209] As an optional implementation, the training device performs the following steps during the execution of the step of "updating the parameters of the first to-be-trained model based on the training recommendation order to obtain a recommended model":

[0210] 4001. Based on the training recommendation order, determine a first indicator and a second indicator, wherein the first indicator represents the degree of interest of the training subject in at least one first candidate item, and the second indicator represents the degree of matching between the at least one first candidate item and the target requirements, and the target requirements include requirements for items recommended to the subject.

[0211] In this embodiment of the present application, a first indicator can be used to determine the degree of interest of a training subject in at least one first candidate item in the training recommendation order, and a second indicator can be used to determine the degree of match between the at least one first candidate item and the target requirement. Optionally, both the first indicator and the second indicator are numerical values. A higher value for the first indicator indicates a higher degree of interest of the training subject in the at least one first candidate item. A higher value for the second indicator indicates a higher degree of match between the at least one first candidate item and the target requirement.

[0212] In one implementation for determining a training subject's level of interest in at least one first candidate item, a positive sample includes an item of interest to the training subject. After obtaining the positive sample, a training device determines the training subject's level of interest in the at least one first candidate item based on a difference between the positive sample and the at least one first candidate item in a training recommendation order, wherein a smaller the difference, the higher the training subject's level of interest in the at least one first candidate item.

[0213] Optionally, positive samples are determined based on records of the training subject's interactions with items. In one possible implementation, an interaction score for an item with which the training subject has interacted is determined based on one or more of the following: the training subject's interaction behavior with the item, the number of interactions between the training subject and the item, and the duration of the interaction between the training subject and the item. A higher interaction score for an item indicates a greater level of interest in the item by the training subject. Positive samples are then determined based on the interaction score.

[0214] For example, let's say the training subject interacted with item i1. The training subject's interactions with item i1 included browsing, adding to favorites, and commenting on it. Specifically, the training subject browsed item i1 for 10 seconds, added to favorites twice, and commented on it once. The interaction score for item i1 can be determined by taking a weighted sum of the browsing duration, number of favorites, and number of comments.

[0215] Optionally, after determining the interaction scores of all items that have interacted with the training object, the k items with the highest interaction scores may be determined as positive samples. Optionally, the items in the positive samples are sorted in descending order of interaction scores.

[0216] In another method of determining the degree of interest of a training subject in at least one first candidate item, a training device determines a first indicator based on a second similarity between the at least one first candidate item and the training subject, wherein a greater the second similarity, a higher degree of interest of the training subject in the at least one first candidate item.

[0217] The target requirements in the training method are the same as those in the recommendation method and will not be repeated here.

[0218] 4002. Based on the first indicator and the second indicator, update the parameters of the first model to be trained to obtain a recommended model.

[0219] The training device updates the parameters of the first model to be trained based on the first indicator, which can increase the training subject's interest in at least one first candidate item. The training device updates the parameters of the first model to be trained based on the second indicator, which can increase the degree of match between at least one first candidate item and the target requirement. Therefore, the training device updates the parameters of the first model to be trained based on the first indicator and the second indicator, which can enable the first model to be trained to both increase the training subject's interest in at least one first candidate item in the training recommendation order and improve the degree of match between the training recommendation order and the target requirement. Accordingly, the recommendation model obtained by updating the parameters of the first model to be trained has the ability to determine items recommended to the subject based on the items in which the subject is interested, and has the ability to determine an order that matches the target requirement.

[0220] In one possible implementation, a training device determines a reward value for a first model to be trained based on a first indicator and a second indicator. The greater the training subject's interest in the at least one first candidate item, the greater the reward value. The greater the degree of match between the at least one first candidate item and the target requirement, the greater the reward value. Based on the reward value, the parameters of the first model to be trained are updated to obtain a recommendation model.

[0221] Optionally, when both the first indicator and the second indicator are numerical values, the reward value is obtained by performing weighted summation on the first indicator and the second indicator.

[0222] Optionally, the reward value is determined based on the first indicator, the second indicator, and the reward function. For example, the independent variables of the reward function include the first indicator and the second indicator, and the dependent variable of the reward function includes the reward value.

[0223] In another possible implementation, the training device determines the loss of the first model to be trained based on the first indicator and the second indicator. The higher the training subject's interest in the at least one first candidate item, the smaller the loss of the first model to be trained. The higher the degree of match between the at least one first candidate item and the target requirement, the smaller the loss of the first model to be trained. Based on the loss of the first model to be trained, the parameters of the first model to be trained are updated to obtain a recommendation model.

[0224] In one possible implementation scenario, the target requirement is determined based on the business goal. For example, the business goal is to increase the diversity of videos recommended to the object, or the business goal may be to prioritize recommending the latest released items to the object, or the business goal may be to prioritize recommending commodities to the object. Since business goals may change over time, the target requirements may also change over time. In order to ensure that the order of items recommended to the object determined by the recommendation model matches the target requirement, when the target requirement changes, the second indicator is adjusted based on the changed target requirement. In this way, based on the first indicator and the second indicator, the parameters of the first model to be trained are updated to obtain a recommendation model, which can ensure that the order of items recommended to the object determined by the recommendation model matches the changed target requirement. In this way, by adjusting the second indicator, a recommendation model that matches the business goal can be obtained, thereby improving the efficiency of matching business goals.

[0225] Optionally, the training device determines a reward value for the first model to be trained based on the first indicator and the second indicator, wherein the reward value is obtained based on the first indicator, the second indicator, and the reward function. Based on the reward value, the parameters of the first model to be trained are updated to obtain a recommendation model. In this case, if the business objectives change, the reward value can be adjusted by adjusting the second indicator in the reward function, thereby obtaining a recommendation model that matches the business objectives. Furthermore, by adjusting the reward function, a better balance can be achieved between the business objectives and the items of interest to the identified subjects.

[0226] As an optional embodiment, the training device performs the following steps during the execution of the step of "updating the parameters of the first model to be trained based on the first indicator and the second indicator to obtain a recommendation model": updating the parameters of the first model to be trained based on the first indicator, the second indicator and the third indicator to obtain a recommendation model, wherein the third indicator represents whether the first model to be trained outputs at least one training reason for recommending at least one first candidate item.

[0227] The training reasons in the training method correspond to the target reasons in the recommendation method. The meaning of the training reasons can be found in the description of the target reasons above and will not be repeated here. Based on the third indicator, it can be determined whether the first model to be trained outputs at least one training reason for recommending at least one first candidate item. Optionally, the third indicator includes that the first model to be trained has output at least one training reason for recommending at least one first candidate item, or includes that the first model to be trained has not output at least one training reason for recommending at least one first candidate item. Optionally, in the case where the third indicator includes that the first model to be trained has output at least one training reason for recommending at least one first candidate item, the third indicator also includes at least one training reason.

[0228] Based on the description of target reasons in the recommendation method, it can be seen that outputting target reasons from the recommendation model can improve the accuracy of at least one target item and the accuracy of the target recommendation order, and also provide a basis for determining the accuracy of at least one target item. Therefore, updating the parameters of the first to-be-trained model based on the third metric enables the first to-be-trained model to determine at least one training reason for at least one first candidate item. Furthermore, the training recommendation order for the at least one first candidate item can be determined based on the at least one training reason, thereby improving the accuracy of the training recommendation order.

[0229] In one possible implementation, a training device determines a reward value for a first model to be trained based on a first indicator, a second indicator, and a third indicator. The reward value when the third indicator indicates that the first model to be trained has output at least one training reason is greater than the reward value when the third indicator indicates that the first model to be trained has not output at least one training reason. Based on the reward value, parameters of the first model to be trained are updated to obtain a recommended model.

[0230] Optionally, when the first indicator, the second indicator and the third indicator are all numerical values, the reward value is obtained by taking a weighted sum of the first indicator and the second indicator, wherein the value of the third indicator when indicating that the first model to be trained has output at least one training reason is greater than the value of the third indicator when indicating that the first model to be trained has not output at least one training reason.

[0231] Optionally, the reward value is determined based on the first indicator, the second indicator, the third indicator, and the reward function. For example, the independent variables of the reward function include the first indicator, the second indicator, and the third indicator, and the dependent variable of the reward function includes the reward value.

[0232] In another possible implementation, the training device determines the loss of the first model to be trained based on the first indicator, the second indicator, and the third indicator. The higher the degree of interest of the training subject in at least one first candidate item, the smaller the loss of the first model to be trained. The higher the degree of match between at least one first candidate item and the target requirement, the smaller the loss of the first model to be trained. The loss value when the third indicator indicates that the first model to be trained has output at least one training reason is smaller than the loss value when the third indicator indicates that the first model to be trained has not output at least one training reason. Based on the loss of the first model to be trained, the parameters of the first model to be trained are updated to obtain a recommendation model.

[0233] As an optional embodiment, the training device obtains a first model to be trained by performing the following steps: obtaining a training prompt word, wherein the training prompt word includes a description of the training subject and a training item, and the training prompt word is used to guide the second model to be trained to identify an item of interest to the training subject from the training items based on the description of the training subject; inputting the training prompt word into the second model to be trained to obtain at least one second candidate item identified by the second model to be trained from the training items; and updating the second model to be trained based on the at least one second candidate item to obtain the first model to be trained.

[0234] In this embodiment, the training device can obtain the first model to be trained by training the second model to be trained. Optionally, the second model to be trained is a pre-trained model.

[0235] The training prompt words are prompt words used to guide the first model to be trained to perform tasks. Specifically, the training prompt words are used to guide the second model to be trained to determine the items of interest to the training object from the training items based on the description of the training object. That is, the task that the first model to be trained needs to perform includes determining the items of interest to the training object from the training items based on the description of the training object.

[0236] After the training device inputs the training prompt word into the second model to be trained, the second model to be trained can, under the guidance of the training prompt word, perform the task of "determining the items of interest to the training object from the training items based on the description of the training object" to obtain at least one second candidate item.

[0237] The training device then updates the parameters of the second model to be trained based on the at least one second candidate item, thereby improving the second model's ability to perform the task of "identifying items of interest to the training subject from the training items based on the description of the training subject." Furthermore, when the first model to be trained is obtained by updating the parameters of the second model to be trained, the accuracy of the first model to be trained in identifying items of interest to the subject based on the description of the object can be improved.

[0238] In one possible implementation, the training device determines a loss for a second model to be trained based on a difference between at least one second candidate item and a positive sample, where the positive sample includes an item of interest to the training subject. Based on the loss of the second model to be trained, the parameters of the second model to be trained are updated to obtain the first model to be trained.

[0239] In this implementation, the training device supervises the second model to be trained based on the positive samples and obtains the first model to be trained by updating the parameters of the second model to be trained under this supervision. This achieves the first model to be trained based on supervised fine-tuning (SFT) of the second model to be trained.

[0240] In this embodiment, after obtaining a training prompt word, the training device inputs the training prompt word into the second to-be-trained model, thereby obtaining at least one second candidate item identified by the second to-be-trained model from the training items. Based on the at least one second candidate item, the parameters of the second to-be-trained model are then updated to obtain the first to-be-trained model. This allows the first to-be-trained model to identify items of interest to the training subject from the second candidate items based on a description of the training subject.

[0241] In this way, during the process of training the first model to be trained to obtain a recommendation model, the first model to be trained can utilize the following capabilities to determine at least one first candidate item: based on the description of the training subject, it can identify items of interest to the training subject from the training items. Furthermore, based on the at least one first candidate item, a first indicator and a second indicator can be determined. Based on the first and second indicators, the parameters of the first model to be trained can be updated to obtain a recommendation model. This improves the efficiency of training the recommendation model.

[0242] Optionally, the training prompt includes format requirements, wherein the format requirements include arranging the items that the subject is interested in in a recommended order. In this way, the second to-be-trained model is trained based on the training prompt to obtain the first to-be-trained model, so that the first to-be-trained model can output according to the format requirements.

[0243] Optionally, after the first model to be trained is obtained by performing SFT on the second model to be trained, reinforcement learning (RL) can be performed on the first model to be trained based on a group relative policy optimization (GRPO) algorithm to obtain a recommendation model.

[0244] As an optional embodiment, the first model to be trained is an LLM. Since the modality of data that can be processed by the LLM is text, the modality of the data input to the first model to be trained should be text. Therefore, it is necessary to convert the information required to determine the recommended training order (including information about the training object and information about the training item) into a natural language description, and input the natural language description into the first model to be trained, so that the first model to be trained can determine the recommended training order based on the information about the training object and the information about the training item. Therefore, the training device can enable the first model to be trained to determine the recommended training order based on the description of the training object and the description of the training item by inputting the training text including the description of the training object and the description of the training item into the first model to be trained.

[0245] Moreover, since the knowledge that can be obtained through pre-training of LLM is richer than the knowledge that can be obtained through training of models other than LLM (such as the twin-tower model), when the second model to be trained is a pre-trained model obtained by pre-training LLM, the first model to be trained obtained by training the second model to be trained also has richer knowledge. As a result, the recommendation model obtained by training the first model to be trained also has richer knowledge. In this way, the target recommendation order is determined based on the recommendation model, and richer knowledge can be used to determine the target recommendation order, which can effectively alleviate the effect of the information cocoon, thereby improving the accuracy of the target recommendation order.

[0246] See also Figure 4 , Figure 4 This is a flow chart of a training method and a recommendation method provided in an embodiment of the present application. Figure 4As shown, the input data includes: a description of the attributes of the training object, a description of the third interaction record between the training object and the item, training instructions, and a description of the training item. Feature engineering is performed on the input data to generate training prompts. Based on the description of the attributes of the training object and the description of the third interaction record between the training object and the item, a description of the training object can be determined. Based on the descriptions of the training object and the training item, a training text can be determined. Based on the training text and the training instructions, a training prompt can be determined. The training instructions are used to instruct the second to-be-trained model to determine a training recommendation order based on the descriptions of the training object and the training item.

[0247] Then, in Phase 1: Supervised Fine-tuning (SFT), the training prompt words obtained through feature engineering are input into the second model to be trained, and SFT can be performed on the second model to be trained to obtain the first model to be trained. Then, in Phase 2: Reinforcement Learning based on the GRPO (Relative Policy Optimization) algorithm, multiple sets of output results can be obtained, where each set of output results includes a training recommendation order for at least one first candidate item. Figure 4 In the example, O1, O2, ..., O6 is a set of output results, that is, at least one first candidate item in the training recommendation order can be: O1, O2, ..., O6. r1, r2, ..., r6 is also a set of output results, that is, at least one first candidate item in the training recommendation order can be: r1, r2, ..., r6. The parameters of the first model to be trained corresponding to different sets of output results are different. For example, when the parameter of the first model to be trained is c1, the set of output results obtained by the first model to be trained includes O1, O2, ..., O6. When the parameter of the first model to be trained is c2, the set of output results obtained by the first model to be trained includes r1, r2, ..., r6. Among them, c1 is different from c2.

[0248] Based on each group of output results, a reward value of the first model to be trained can be determined. Then, by normalizing the reward value corresponding to each group of output results (group norm), the normalized reward value corresponding to each group of output results can be obtained. Then, based on the parameters of the first model to be trained corresponding to the maximum value of the normalized reward value, the recommended model is determined. For example, the normalized reward value z1 corresponding to the output result j1 is larger than the normalized reward value j2 corresponding to the output result j2, then the recommended model can be determined based on the parameters of the first model to be trained corresponding to the output result j1. Optionally, the parameters of the first model to be trained corresponding to the maximum value of the normalized reward value are called target parameters, and the training device adjusts the parameters of the first model to be trained to the target parameters to obtain the recommended model. In Figure 4 In [ ] , the target parameters include: A1, A2, ..., A6. After obtaining the recommendation model, the target recommendation order can be determined based on the recommendation model.

[0249] The present embodiment also provides a comparison of the effectiveness of the recommendation model trained based on the training method described above and the open source method. Please refer to Table 1 below for a comparison of the effectiveness of the two.

[0250] Table 1

[0251] Evaluation indicator 1 Evaluation indicator 2 Evaluation indicator 3 Evaluation indicator 4 Open Source Model 0.3381 0.3762 0.4053 0.4802 Recommended Model 0.4137 0.4692 0.5311 0.5653

[0252] In Table 1, evaluation indicators 1 and 3 are both recall rates. Evaluation indicator 1 evaluates the recall rate of the top 5 recommended items, and evaluation indicator 3 evaluates the recall rate of the top 10 recommended items. Evaluation indicators 2 and 4 are both normalized discounted cumulative gain (NDCG). Evaluation indicator 2 evaluates the recall rate NDCG of the top 5 recommended items, and evaluation indicator 4 evaluates the recall rate NDCG of the top 10 recommended items.

[0253] As shown in Table 1, the evaluation index 1, evaluation index 2, evaluation index 3, and evaluation index 4 of the recommendation model are all greater than the evaluation index 1, evaluation index 2, evaluation index 3, and evaluation index 4 of the open source model. This shows that in terms of the four evaluation dimensions of evaluation index 1, evaluation index 2, evaluation index 3, and evaluation index 4, the recommendation effect of the recommendation model is better than that of the open source model.

[0254] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, personal information processing may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0255] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0256] The above describes in detail the method of the embodiment of the present application, and the following provides an apparatus of the embodiment of the present application.

[0257] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of a recommendation device provided in an embodiment of the present application. The recommendation device 1 includes: an acquisition unit 11 and a processing unit 12, wherein:

[0258] An acquisition unit 11 is configured to acquire a recommendation model and a target text, wherein the target text includes a description of a target object and a description of an item to be recommended;

[0259] The processing unit 12 is used to input the target text into the recommendation model to obtain a target recommendation order output by the recommendation model, wherein the target recommendation order includes at least one target item among the items to be recommended, and the at least one target item is an item recommended to the target object.

[0260] In combination with any embodiment of the present application, the acquiring unit 11 is further configured to:

[0261] Obtaining a first interaction record between the target object and the item;

[0262] Determining a description of the target object by aggregating the first interaction records;

[0263] The target text is obtained based on the description of the target object and the description of the item to be recommended.

[0264] In combination with any embodiment of the present application, the acquiring unit 11 is further configured to:

[0265] Determining a second interaction record by aggregating the first interaction record, where the second interaction record includes a record of the target object interacting with the item within a first time period and / or a record of the target object interacting with the item within a second time period, where a time span of the first time period is shorter than a time span of the second time period;

[0266] Based on the second interaction record, a description of the target object is determined.

[0267] In combination with any embodiment of the present application, the record of the target object's interaction with the item in the first time period and the record of the target object's interaction with the item in the second time period both include statistical results of the target object's interaction with the item.

[0268] In combination with any embodiment of the present application, the statistical results of the interaction between the target object and the item include the statistical results of the interaction between the target object and items of each subject type, and / or the statistical results of the interaction between the target object and items of each item type.

[0269] In combination with any embodiment of the present application, the acquiring unit 11 is further configured to:

[0270] Based on the second interaction record and the first preset template, a description of the target object is determined, wherein the first preset template includes multiple statistical result fields, and the multiple statistical result fields include at least one of the following: a time field, a subject type field to which the item belongs, an item type field, an interaction behavior field, and an interaction number field.

[0271] In combination with any embodiment of the present application, the acquiring unit 11 is further configured to:

[0272] Based on the second interaction record, the attributes of the target object and the first preset template, a description of the target object is determined, and the multiple statistical result fields also include the attributes of the object.

[0273] In combination with any embodiment of the present application, the processing unit 12 is further configured to:

[0274] Based on the information of the item to be recommended, a description of the item to be recommended is determined, where the information of the item to be recommended includes one or more of the following: the content of the item to be recommended, the title of the item to be recommended, the theme type of the item to be recommended, and the interaction index of the item to be recommended.

[0275] In combination with any embodiment of the present application, the processing unit 12 is further configured to:

[0276] Based on the information of the item to be recommended and a second preset template, a description of the item to be recommended is determined.

[0277] In combination with any embodiment of the present application, the processing unit 12 is further configured to:

[0278] The target text is input into the recommendation model to obtain a target recommendation order and at least one target reason output by the recommendation model, wherein the at least one target reason is a reason for determining that the at least one target item is recommended to the target object.

[0279] In combination with any embodiment of the present application, the target recommendation order is determined based on the at least one target reason.

[0280] In an embodiment of the present application, the description of the target object includes information about the target object, and the description of the item to be recommended includes information about the item to be recommended. After obtaining the recommendation model and the target text including the description of the target object and the description of the item to be recommended, the recommendation device inputs the target text into the recommendation model, which enables the recommendation model to determine the information about the target object and the item to be recommended based on the description of the target object and the description of the item to be recommended. Furthermore, based on the information about the target object and the information about the item to be recommended, the recommendation model can determine at least one target item to be recommended to the target object from the items to be recommended, and determine a target recommendation order for the at least one target item. This allows the recommendation model to determine the target recommendation order for at least one target item to be recommended to the target object in an end-to-end manner, thereby improving the accuracy of the target recommendation order.

[0281] See also Figure 6 , Figure 6 This is a schematic diagram of the structure of a training device provided in an embodiment of the present application. The training device 2 includes: an acquisition unit 21, a processing unit 22, and an update unit 23, wherein:

[0282] An acquisition unit 21 is configured to acquire a first model to be trained and a training text, wherein the training text includes a description of a training object and a description of a training item;

[0283] a processing unit 22 configured to input the training text into the first to-be-trained model, and obtain a training recommendation order output by the first to-be-trained model, wherein the training recommendation order includes at least one first candidate item among the training items, the at least one first candidate item being an item recommended to the training subject;

[0284] The updating unit 23 is configured to update the parameters of the first model to be trained based on the training recommendation order to obtain a recommended model.

[0285] In combination with any embodiment of the present application, the updating unit 23 is further configured to:

[0286] Determining, based on the training recommendation order, a first indicator and a second indicator, wherein the first indicator represents the degree of interest of the training subject in the at least one first candidate item, and the second indicator represents the degree of matching of the at least one first candidate item with target requirements, wherein the target requirements include requirements for items recommended to the subject;

[0287] Based on the first indicator and the second indicator, the parameters of the first model to be trained are updated to obtain the recommendation model.

[0288] In combination with any embodiment of the present application, the updating unit 23 is further configured to:

[0289] Based on the first indicator, the second indicator and the third indicator, the parameters of the first model to be trained are updated to obtain the recommendation model, and the third indicator represents whether the first model to be trained outputs at least one training reason for recommending the at least one first candidate item.

[0290] In combination with any embodiment of the present application, the updating unit 23 is further configured to:

[0291] Determining a reward value of the first to-be-trained model based on the first indicator and the second indicator;

[0292] Based on the reward value, the parameters of the first model to be trained are updated to obtain the recommendation model.

[0293] In combination with any embodiment of the present application, the target requirements include: the diversity of items recommended to the object meets the preset requirements, the items recommended to the object include items of preset types, the proportion of new items in the items recommended to the object is greater than or equal to a first threshold, and the new items are items whose release time is no more than a preset value from the current time.

[0294] In combination with any embodiment of the present application, the acquiring unit 21 is further configured to:

[0295] Obtaining a training prompt word, wherein the training prompt word includes a description of the training subject and the training item, and the training prompt word is used to guide the second to-be-trained model to determine an item of interest to the training subject from the training items based on the description of the training subject;

[0296] Inputting the training prompt word into the second model to be trained to obtain at least one second candidate item determined by the second model to be trained from the training items;

[0297] Based on the at least one second candidate item, the second model to be trained is updated to obtain the first model to be trained.

[0298] In combination with any embodiment of the present application, the acquiring unit 21 is further configured to:

[0299] determining a loss of the second to-be-trained model based on a difference between the at least one second candidate item and a positive sample, the positive sample including the item of interest to the training subject;

[0300] Based on the loss of the second model to be trained, the parameters of the second model to be trained are updated to obtain the first model to be trained.

[0301] In this embodiment of the present application, the training text includes a description of the training subject and a description of the training items. After obtaining a first model to be trained and the training text, the training device inputs the training text into the first model to be trained, obtains a training recommendation order output by the first model to be trained, and the training recommendation order includes at least one first candidate item among the training items, wherein the at least one first candidate item is an item recommended to the training subject. Based on the training recommendation order, the parameters of the first model to be trained are updated to obtain a recommendation model.

[0302] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0303] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device 3 includes a processor 31 and a memory 32. Optionally, the electronic device 3 also includes an input device 33 and an output device 34. The processor 31, the memory 32, the input device 33 and the output device 34 are coupled via a connector, and the connector includes various interfaces, transmission lines or buses, etc., which are not limited in the embodiments of the present application. It should be understood that in each embodiment of the present application, coupling refers to mutual connection in a specific manner, including direct connection or indirect connection through other devices, for example, connection through various interfaces, transmission lines, buses, etc.

[0304] The processor 31 may include one or more processors, for example, one or more central processing units (CPUs). In the case where the processor is a CPU, the CPU may be a single-core CPU or a multi-core CPU. Alternatively, the processor 31 may be a processor group consisting of multiple CPUs, wherein the multiple processors are coupled to each other via one or more buses. Alternatively, the processor may also be other types of processors, etc., which are not limited in the embodiments of the present application.

[0305] The memory 32 can be used to store computer program instructions and various computer program codes, including the program code for executing the solution of the present application. Optionally, the memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or portable compact disc read-only memory (CD-ROM), which is used for related instructions and data.

[0306] The input device 33 is used to input data and / or signals, and the output device 34 is used to output data and / or signals. The input device 33 and the output device 34 can be independent devices or an integrated device.

[0307] It can be understood that in the embodiment of the present application, the memory 32 can be used not only to store relevant instructions, but also to store relevant data. The embodiment of the present application does not limit the specific data stored in the memory.

[0308] It is understandable that Figure 7 Only a simplified design of an electronic device is shown. In actual applications, the electronic device may further include other necessary components, including but not limited to any number of input / output devices, processors, memories, etc., and all electronic devices that can implement the embodiments of the present application are within the scope of protection of the present application.

[0309] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0310] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here. Those skilled in the art will also clearly understand that the descriptions of the various embodiments of this application have different focuses. For the convenience and brevity of description, the same or similar parts may not be repeated in different embodiments. Therefore, for parts not described or not described in detail in a certain embodiment, reference can be made to the descriptions of other embodiments.

[0311] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0312] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0313] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0314] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0315] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by a computer program instructing related hardware to perform the processes. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A recommendation method, characterized in that: The recommended methods include: Obtaining a recommendation model and a target text, wherein the target text includes a description of the target object and a description of the recommended item; The target text is input into the recommendation model to obtain a target recommendation order output by the recommendation model, wherein the target recommendation order includes at least one target item among the items to be recommended, and the at least one target item is an item recommended to the target object.

2. The method according to claim 1, characterized in that The obtaining of the target text includes: Obtaining a first interaction record between the target object and the item; Determining a description of the target object by aggregating the first interaction records; The target text is obtained based on the description of the target object and the description of the item to be recommended.

3. The method according to claim 2, characterized in that Determining a description of the target object by aggregating the first interaction records includes: Determining a second interaction record by aggregating the first interaction record, where the second interaction record includes a record of the target object interacting with the item within a first time period and / or a record of the target object interacting with the item within a second time period, where a time span of the first time period is shorter than a time span of the second time period; Based on the second interaction record, a description of the target object is determined.

4. The method according to claim 3, characterized in that The record of the target object's interaction with the item in the first time period and the record of the target object's interaction with the item in the second time period both include statistical results of the target object's interaction with the item.

5. The method according to claim 4, characterized in that The statistical results of the target object's interaction with items include statistical results of the target object's interaction with items of each subject type, and / or statistical results of the target object's interaction with items of each item type.

6. The method according to any one of claims 3 to 5, characterized in that The determining, based on the second interaction record, a description of the target object includes: Based on the second interaction record and the first preset template, a description of the target object is determined, wherein the first preset template includes multiple statistical result fields, and the multiple statistical result fields include at least one of the following: a time field, a subject type field to which the item belongs, an item type field, an interaction behavior field, and an interaction number field.

7. The method according to claim 6, characterized in that The determining a description of the target object based on the second interaction record and the first preset template includes: Based on the second interaction record, the attributes of the target object and the first preset template, a description of the target object is determined, and the multiple statistical result fields also include the attributes of the object.

8. The method according to any one of claims 2 to 5, characterized in that Before obtaining the target text based on the description of the target object and the description of the item to be recommended, the method further includes: Based on the information of the item to be recommended, a description of the item to be recommended is determined, where the information of the item to be recommended includes one or more of the following: the content of the item to be recommended, the title of the item to be recommended, the theme type of the item to be recommended, and the interaction index of the item to be recommended.

9. The method according to claim 8, characterized in that The step of determining a description of the item to be recommended based on the information of the item to be recommended includes: Based on the information of the item to be recommended and a second preset template, a description of the item to be recommended is determined.

10. The method according to any one of claims 1 to 5, characterized in that Inputting the target text into the recommendation model to obtain the target recommendation order output by the recommendation model includes: The target text is input into the recommendation model to obtain a target recommendation order and at least one target reason output by the recommendation model, wherein the at least one target reason is a reason for determining that the at least one target item is recommended to the target object.

11. The method according to claim 10, characterized in that The target recommendation order is determined based on the at least one target reason.

12. A training method, characterized in that: The training method comprises: Obtaining a first model to be trained and a training text, wherein the training text includes a description of a training object and a description of a training item; Inputting the training text into the first to-be-trained model to obtain a training recommendation order output by the first to-be-trained model, wherein the training recommendation order includes at least one first candidate item among the training items, and the at least one first candidate item is an item recommended to the training subject; Based on the training recommendation order, the parameters of the first model to be trained are updated to obtain a recommended model.

13. The training method according to claim 12, characterized in that: The updating of the parameters of the first to-be-trained model based on the training recommendation order to obtain a recommended model includes: Determining, based on the training recommendation order, a first indicator and a second indicator, wherein the first indicator represents the degree of interest of the training subject in the at least one first candidate item, and the second indicator represents the degree of matching between the at least one first candidate item and target requirements, wherein the target requirements include requirements for items recommended to the subject; Based on the first indicator and the second indicator, the parameters of the first model to be trained are updated to obtain the recommendation model.

14. The training method according to claim 13, characterized in that: The updating of the parameters of the first to-be-trained model based on the first indicator and the second indicator to obtain the recommendation model includes: Based on the first indicator, the second indicator and the third indicator, the parameters of the first model to be trained are updated to obtain the recommendation model, and the third indicator represents whether the first model to be trained outputs at least one training reason for recommending the at least one first candidate item.

15. The training method according to claim 13, characterized in that: The updating of the parameters of the first to-be-trained model based on the first indicator and the second indicator to obtain the recommendation model includes: Determining a reward value of the first to-be-trained model based on the first indicator and the second indicator; Based on the reward value, the parameters of the first model to be trained are updated to obtain the recommendation model.

16. The training method according to any one of claims 13 to 15, characterized in that The target requirements include: the diversity of items recommended to the object meets preset requirements, the items recommended to the object include items of preset types, the proportion of new items in the items recommended to the object is greater than or equal to a first threshold, and the new items are items whose release time is no more than a preset value from the current time.

17. The training method according to any one of claims 12 to 15, characterized in that: The obtaining of the first model to be trained includes: Obtaining a training prompt word, wherein the training prompt word includes a description of the training subject and the training item, and the training prompt word is used to guide the second to-be-trained model to determine an item of interest to the training subject from the training items based on the description of the training subject; Inputting the training prompt word into the second model to be trained to obtain at least one second candidate item determined by the second model to be trained from the training items; Based on the at least one second candidate item, the second model to be trained is updated to obtain the first model to be trained.

18. The training method according to claim 17, characterized in that: The updating of the second model to be trained based on the at least one second candidate item to obtain the first model to be trained includes: determining a loss of the second to-be-trained model based on a difference between the at least one second candidate item and a positive sample, the positive sample including the item of interest to the training subject; Based on the loss of the second model to be trained, the parameters of the second model to be trained are updated to obtain the first model to be trained.

19. A recommendation device, characterized in that: The recommended device includes: an acquisition unit, configured to acquire a recommendation model and a target text, wherein the target text includes a description of a target object and a description of an item to be recommended; A processing unit is used to input the target text into the recommendation model to obtain a target recommendation order output by the recommendation model, wherein the target recommendation order includes at least one target item among the items to be recommended, and the at least one target item is an item recommended to the target object.

20. A training device, characterized in that: The training device comprises: an acquiring unit, configured to acquire a first model to be trained and a training text, wherein the training text includes a description of a training object and a description of a training item; a processing unit, configured to input the training text into the first to-be-trained model, and obtain a training recommendation order output by the first to-be-trained model, wherein the training recommendation order includes at least one first candidate item among the training items, the at least one first candidate item being an item recommended to the training subject; An updating unit is used to update the parameters of the first model to be trained based on the training recommendation order to obtain a recommended model.

21. An electronic device, characterized in that: include: A processor and a memory, the memory being configured to store computer program code, the computer program code comprising computer instructions, wherein when the processor executes the computer instructions, the electronic device executes the method according to any one of claims 1 to 11; When the processor executes the computer instructions, the electronic device may perform the method according to any one of claims 12 to 18.

22. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 11; When the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 12 to 18.

23. A computer program product, characterized in that The computer program product comprises a computer program or instructions; when the computer program or instructions are run on a computer, the computer is caused to perform the method according to any one of claims 1 to 11; When the computer program or instruction runs on a computer, the computer is caused to execute the method according to any one of claims 12 to 18.

Citation Information

Patent Citations

  • Article information recommendation method and device and electronic equipment

    CN116245592A

  • Shop recommendation method and device, electronic equipment and storage medium

    CN117273868A

  • Article cold start recommendation method and article cold start recommendation model training method

    CN117493666A

  • Personalized recommendation method and device, electronic equipment and readable storage medium

    CN117992672A

  • Drug recommendation method and device, electronic apparatus, and storage medium

    US20210407642A1