Resource recommendation method, model training method, device, equipment and medium

By acquiring the recommended reference features of the target object and dynamically adjusting the resource recommendation strategy using a reinforcement learning model, the problem of a fixed number of resources in existing systems is solved, and the personalized and accurate resource recommendation results are achieved, meeting the different needs of users in different scenarios.

CN120723977BActive Publication Date: 2026-03-27BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing recommendation systems lack flexible mechanisms to adjust the number of relevant resources, and cannot dynamically adjust the proportion of resources related to the target resource in the resource recommendation results according to user needs and context, thus failing to meet the diverse needs of different users in different scenarios.

Method used

By acquiring the recommendation reference features of the target object, the relevant resource recommendation parameters are determined, and the recommendation strategy is dynamically adjusted based on these parameters. The number of relevant resources is determined using reinforcement learning, and the value function and advantage function are combined to accurately evaluate the value of different numbers of resources and dynamically adjust the proportion of relevant resources in the resource recommendation results.

Benefits of technology

It enables flexible fulfillment of the diverse needs of target objects in different contexts, improves the accuracy and effectiveness of resource recommendations, and meets users' personalized needs in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723977B_ABST
    Figure CN120723977B_ABST
Patent Text Reader

Abstract

The present disclosure provides a resource recommendation method, a model training method, an apparatus, a device and a medium, relates to the technical field of artificial intelligence, and in particular to the technical field of big data, deep learning, intelligent recommendation and the like. The method comprises: in response to a target object selecting a target resource, obtaining a recommendation reference feature for the target object; determining a relevant resource recommendation parameter for the target object according to the recommendation reference feature, wherein the relevant resource recommendation parameter represents a proportion of recommended resources in a resource recommendation result for the target object, which have a relevance to the target resource satisfying a preset condition; and performing a resource recommendation operation according to the recommendation reference feature and an updated recommendation strategy obtained by updating the preset recommendation strategy according to the relevant resource recommendation parameter, to obtain the resource recommendation result for the target object.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of big data, deep learning, intelligent recommendation, etc. Specifically, it relates to a resource recommendation method, a model training method, an apparatus, a device, and a medium. BACKGROUND

[0002] With the rapid development of artificial intelligence, recommendation systems can analyze users' historical behaviors, interest preferences, and user attribute characteristics to generate personalized recommendation content, and have been widely applied in various application scenarios, especially in the fields of online video, e-commerce platforms, social media, etc. SUMMARY

[0003] The present disclosure provides a resource recommendation method, a model training method, an apparatus, a device, and a medium.

[0004] According to an aspect of the present disclosure, a resource recommendation method is provided, comprising: in response to a target object selecting a target resource, obtaining a recommendation reference feature for the target object; determining a related resource recommendation parameter for the target object according to the recommendation reference feature, wherein the related resource recommendation parameter represents a proportion of recommendation resources in a resource recommendation result for the target object that meet a preset condition of relevance to the target resource; and performing a resource recommendation operation according to the recommendation reference feature and an updated recommendation strategy obtained by updating the preset recommendation strategy with the related resource recommendation parameter, to obtain the resource recommendation result for the target object.

[0005] According to another aspect of the present disclosure, a model training method is provided, wherein the model comprises a consumption time determination module and a plurality of advantage value determination modules; the method comprises: obtaining a training sample, wherein the training sample comprises a sample recommendation reference feature, a sample related resource number, and a sample label; inputting the training sample into the consumption time determination module and the plurality of advantage value determination modules to obtain a sample consumption time and a plurality of sample advantage values; determining a total sample consumption time according to the sample consumption time and a target sample advantage value in the plurality of sample advantage values, wherein the target sample advantage value is determined according to the sample related resource number; adjusting network parameters of a related resource number determination model according to a difference between the total sample consumption time and the sample label to obtain a trained related resource number determination model.

[0006] According to another aspect of the present disclosure, a resource recommendation apparatus is provided, comprising: an acquisition module configured to acquire a recommendation reference feature for a target object in response to the target object selecting a target resource; a determination module configured to determine a relevant resource recommendation parameter for the target object according to the recommendation reference feature, wherein the relevant resource recommendation parameter represents a proportion of recommendation resources in a resource recommendation result for the target object that meet a preset condition in terms of relevance to the target resource; and a resource recommendation module configured to perform a resource recommendation operation according to the recommendation reference feature and an updated recommendation strategy obtained by updating a preset recommendation strategy according to the relevant resource recommendation parameter, to obtain the resource recommendation result for the target object.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as above.

[0008] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to perform the method as above.

[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method as above.

[0010] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0011] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:

[0012] Figure 1 An exemplary system architecture to which the resource recommendation method, the training method of the model and the agent according to the embodiments of the present disclosure can be applied is schematically shown;

[0013] Figure 2 A flowchart of the resource recommendation method according to the embodiments of the present disclosure is schematically shown;

[0014] Figure 3 A model architecture diagram of the advantage value determination model according to the embodiments of the present disclosure is schematically shown;

[0015] Figure 4 A flowchart of determining the target recommendation resource according to the embodiments of the present disclosure is schematically shown;

[0016] Figure 5 A flowchart of a model training method is schematically shown according to an embodiment of the present disclosure.

[0017] Figure 6A A schematic diagram of training sample construction is schematically shown according to an embodiment of the present disclosure

[0018] Figure 6B A model architecture schematic diagram of model training is schematically shown according to an embodiment of the present disclosure.

[0019] Figure 7 A module diagram of a resource recommendation apparatus is schematically shown according to an embodiment of the present disclosure.

[0020] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0021] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, which should be considered in a descriptive sense only. Thus, it will be apparent to one of ordinary skill in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.

[0022] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved in the technical solutions comply with relevant laws and regulations, necessary security measures are taken, and do not violate public order and good customs.

[0023] In the technical solutions of the present disclosure, the authorization or consent of the user is obtained before the user's personal information is acquired or collected.

[0024] In related embodiments, when the recommendation system recommends resources to the user, it usually recommends a fixed number of recommended resources to the user, wherein the recommended resources are usually related resources related to the current resource browsed by the user. This way lacks a flexible mechanism to adjust the number of related resources, and cannot decide how many related resources to display according to user needs and context environment, without considering the difference in user demand for recommended resources. In addition, since some users may prefer more related resources related to the current resource in some scenarios, while other users may prefer recommended resources completely different from the current resource, the above-mentioned recommendation method cannot dynamically adjust the number of related resources to adapt to the needs of different users.

[0025] To address the technical problem above, the present disclosure provides a resource recommendation method, comprising: in response to a target object selecting a target resource, obtaining a recommendation reference feature for the target object; determining a relevant resource recommendation parameter for the target object according to the recommendation reference feature, wherein the relevant resource recommendation parameter represents a proportion of recommendation resources in a resource recommendation result for the target object that meet a preset condition in terms of relevance to the target resource; and performing a resource recommendation operation according to the recommendation reference feature and an updated recommendation strategy obtained by updating the preset recommendation strategy according to the relevant resource recommendation parameter, to obtain the resource recommendation result for the target object.

[0026] According to an embodiment of the present disclosure, by adopting the method of obtaining a recommendation reference feature for a target object in response to the target object selecting a target resource, and determining a relevant resource recommendation parameter according to the recommendation reference feature, and dynamically adjusting a recommendation strategy by using the relevant resource recommendation parameter, the proportion of relevant resources related to the target resource in a resource recommendation result can be dynamically adjusted, the difference in the demand for relevant resources by a target object in different context environments can be flexibly met, the difference in the demand of different target objects in different scenarios can be accurately identified, and the resource recommendation effect can be improved.

[0027] Figure 1 An exemplary system architecture of the resource recommendation method, the training method of the model, and the apparatus according to an embodiment of the present disclosure is schematically shown.

[0028] It should be noted that, Figure 1 The system architecture shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture of the resource recommendation method, the training method of the model, and the apparatus can include a terminal device, but the terminal device can not need to interact with a server to implement the resource recommendation method, the training method of the model, and the apparatus provided by the embodiments of the present disclosure.

[0029] As Figure 1 shown, the system architecture 100 according to this embodiment can include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 can include various connection types, such as wired and / or wireless communication links, etc.

[0030] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as an example).

[0031] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, etc.

[0032] The server 105 can be a server providing various services, such as a background management server providing support for the content browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only as an example). The background management server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data, etc. obtained or generated according to user requests) to the terminal device.

[0033] It should be noted that the resource recommendation method and the model training method provided by the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103. Correspondingly, the resource recommendation apparatus provided by the embodiments of the present disclosure can also be arranged in the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0034] Alternatively, the resource recommendation method and the model training method provided by the embodiments of the present disclosure can also be generally executed by the server 105. Correspondingly, the resource recommendation apparatus provided by the embodiments of the present disclosure can be generally arranged in the server 105. The resource recommendation method and the model training method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the resource recommendation apparatus provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0035] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above-mentioned system is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks, and servers.

[0036] Figure 2 A flowchart of a resource recommendation method according to an embodiment of the present disclosure is shown.

[0037] As shown in Figure 2 The resource recommendation method of this embodiment includes operations S210-S230.

[0038] In operation S210, in response to a target object selecting a target resource, a recommendation reference feature for the target object is obtained.

[0039] The target object refers to a subject who performs operations such as resource selection in various recommendation systems, platforms, or scenarios. For example, the target object can be a consumer shopping on an e-commerce platform, and for another example, the target object can be a user watching videos on a video platform.

[0040] The target resource can be a resource selected by the target object from a resource list page. The target resource can be various forms of information, data, services, products, etc. For example, in a video platform, the target resource can be video content; in an e-commerce platform, the target resource can be various products; in a social media, the target resource can be various text, image, audio, video, etc. content.

[0041] The recommendation reference feature includes features of the target object in multiple dimensions, and the recommendation reference feature can include object behavior features, object attribute features, context features, and entry resource features. The object behavior features may, for example, include the click history, browsing time, etc. of the target object. The object attribute features can include the age, gender, region, etc. of the target object. The context features can include the time, place, state, etc. when the target object selects the target resource. The entry resource features can include the type, theme, label, etc. of the target resource.

[0042] It should be noted that the acquisition and use of the recommendation reference feature are known and agreed by the user, and comply with the relevant legal regulations and do not violate public order and good customs.

[0043] In operation S220, according to the recommendation reference feature, a related resource recommendation parameter for the target object is determined, and the related resource recommendation parameter represents the proportion of recommendation resources that meet a preset condition in terms of relevance to the target resource in the resource recommendation result for the target object.

[0044] For example, the preset condition can be a relevance threshold. When the relevance between the recommendation resource and the target resource is greater than the relevance threshold, the recommendation resource is considered to be a related resource, and when the relevance between the recommendation resource and the target resource is less than the relevance threshold, the recommendation resource is considered to be an unrelated resource.

[0045] The resource recommendation result can include relevant resources related to the target resource, or irrelevant resources unrelated to the target resource. The relevant resource recommendation parameter is used to represent the proportion of relevant resources in the resource recommendation result.

[0046] For example, in a video platform, after the target object selects a target video, each time the resource recommendation result used for recommendation to the target object has 6 resources, of which 3 are relevant resources and 3 are irrelevant resources. Therefore, the relevant resource recommendation parameter is 0.5.

[0047] In operation S230, a resource recommendation operation is performed according to the recommendation reference feature and the updated recommendation strategy obtained by updating the preset recommendation strategy according to the relevant resource recommendation parameter, to obtain a resource recommendation result for the target object.

[0048] The preset recommendation strategy can include only the number of resources recommended to the target object, or can include the number of resources recommended to the target object and the composition strategy of relevant resources and irrelevant resources in the recommended resources. The preset recommendation strategy is updated by using the relevant resource recommendation parameter, so that the proportion of relevant resources and irrelevant resources can be adjusted.

[0049] It should be noted that the relevant resource recommendation parameter can be dynamically updated according to the selection of different target resources by the target object. For example, after the target object browses a target resource, the target object can select the next target resource from the resource recommendation result for continuous browsing. At this time, the recommendation reference feature of the target object can be updated according to the target resource newly selected by the target object, and the relevant resource recommendation parameter can be updated according to the updated recommendation reference feature, so as to further update the recommendation strategy and dynamically adjust the proportion of relevant resources in the resource recommendation result.

[0050] According to the embodiments of the present disclosure, by acquiring the recommendation reference feature for the target object in response to the selection of the target resource by the target object, determining the relevant resource recommendation parameter according to the recommendation reference feature, and dynamically adjusting the recommendation strategy by using the relevant resource recommendation parameter, the proportion of relevant resources related to the target resource in the resource recommendation result can be dynamically adjusted, the difference in the demand for relevant resources by the target object in different context environments can be flexibly met, the difference in the demand of different target objects in different scenarios can be accurately identified, and the resource recommendation effect can be improved.

[0051] According to an embodiment of the present disclosure, determining the relevant resource recommendation parameter for the target object according to the recommendation reference feature can include: determining a target relevant resource number for the target object according to a recommendation reference feature and a relevant resource number determination model; the target relevant resource number represents a number of recommended resources in a resource recommendation result for the target object that meet a preset condition of relevance to the target resource; and determining the relevant resource recommendation parameter according to the target relevant resource number and a preset recommendation parameter.

[0052] The target relevant resource number represents a number of relevant resources included in the resource recommendation result.

[0053] The preset recommendation parameter represents a number of resource recommendation results displayed for the target object in response to the target object selecting the target resource, and is a recommendation number preconfigured in the recommendation system. For example, in a video platform, the preset recommendation parameter can be 6, and the target object selects a target video for viewing. Accordingly, 6 additional recommended videos are recommended for the target object.

[0054] The relevant resource recommendation parameter can be obtained according to a ratio of the target relevant resource number and the preset recommendation parameter.

[0055] Exemplarily, the relevant resource number determination model can be a deep neural network model combined with reinforcement learning. Reinforcement learning refers to using a value function and an advantage function to estimate the total value in the current state and the advantage of each action, respectively, so as to improve the learning efficiency and the stability of the model.

[0056] In the present embodiment, the value function can be used to evaluate the expected value of the user consumption time length in the state of the current recommendation reference feature. Through the value function, the relevant resource number determination model can estimate the consumption time length of a similar object group based on the recommendation reference feature, thereby eliminating individual bias. The advantage function is used to calculate the additional consumption time length benefit brought by selecting different target relevant resource numbers in the state of the current recommendation reference feature. The advantage function can reflect the influence of different target relevant resource numbers on the long-term consumption time length of the target object, thereby helping the relevant resource number determination model to dynamically adjust the recommended target relevant resource number.

[0057] According to an embodiment of the present disclosure, by introducing the relevant resource number determination model of reinforcement learning, by separating the value function and the advantage function, the relevant resource number determination model can accurately evaluate the value of different relevant resource numbers, and meet the different needs of different users for recommended resources.

[0058] According to an embodiment of the present disclosure, the related resource quantity determination model comprises a plurality of advantage value determination modules, each of which is configured with a candidate related resource quantity. According to the recommendation reference feature and the related resource quantity determination model, determining the target related resource quantity for the target object can comprise: inputting the recommendation reference feature into the plurality of advantage value determination modules respectively, and outputting the advantage value of each of the plurality of advantage value determination modules, wherein the advantage value represents the marginal contribution of the candidate related resource quantity to the consumption time length of the target object; and determining the target related resource quantity according to the maximum advantage value in the plurality of advantage values.

[0059] Exemplarily, the related resource quantity determination model can comprise two branches, the first branch can comprise a plurality of advantage value determination modules, and the second branch is a consumption time length determination module.

[0060] Each advantage value determination module is configured with a different candidate related resource quantity. For example, in the case of a preset recommendation parameter of 6, four advantage value determination modules can be set, and the candidate related resource quantities of the plurality of advantage value determination modules can be 0, 1, 2, and 3 respectively.

[0061] The recommendation reference feature is input into the plurality of advantage value determination modules respectively, the advantage function is estimated by the advantage value determination module, and a plurality of advantage values corresponding to different related resource quantities are output respectively. The recommendation reference feature is input into the consumption time length determination module, the value function is estimated by the consumption time length determination module, and the consumption time length is output.

[0062] According to the recommendation reference feature and the related resource quantity determination model, determining the target related resource quantity for the target object can also comprise: taking the object attribute feature in the recommendation reference feature as a category identifier, and extracting it into a corresponding feature representation through an encoder, processing the feature representation together with the object behavior feature, the context feature, and the entry resource feature in the recommendation reference feature through a multi-layer perceptron to generate a splicing vector. The splicing vector is input into the plurality of advantage value determination modules and the consumption time length determination module of the related resource quantity determination model respectively, and a plurality of advantage values and a consumption time length are output.

[0063] The maximum advantage value represents the maximum improvement of the consumption time length of the target object under the corresponding related resource quantity. The advantage value can be used to measure the additional improvement effect of the candidate related resource quantity on the consumption time length of the target object.

[0064] According to an embodiment of the present disclosure, determining the target related resource quantity according to the maximum advantage value in the plurality of advantage values can comprise: determining the advantage value determination module corresponding to the maximum advantage value in the plurality of advantage values as a target advantage value determination module; and determining the candidate related resource quantity of the target advantage value determination module as the target related resource quantity.

[0065] After the multiple advantage values ​​are determined by the module, the target advantage module and the number of related resources corresponding to the maximum advantage value can be determined by reverse query.

[0066] Figure 3 This diagram illustrates a model architecture for determining the number of related resource entries according to an embodiment of the present disclosure.

[0067] like Figure 3 As shown, the relevant resource count determination model may include a multilayer perceptron 320, an advantage value determination module, and a consumption duration determination module 340. The advantage value determination module may include multiple modules, such as a first advantage value determination module 331, a second advantage value determination module 332, a third advantage value determination module 333, and a fourth advantage value determination module 334. Each advantage value determination module can be configured with a different number of candidate relevant resources; for example, the first advantage value determination module 331 may have 0 candidate relevant resources, the second advantage value determination module 332 may have 1 candidate relevant resource, the third advantage value determination module 333 may have 2 candidate relevant resources, and the fourth advantage value determination module 334 may have 3 candidate relevant resources.

[0068] The recommended reference feature 310 is input into the multilayer perceptron 320 to generate a concatenated vector. This concatenated vector is then input into the consumption duration determination module 340, which outputs a consumption duration of 360. The concatenated vector is then input into the first dominance value determination module 331, the second dominance value determination module 332, the third dominance value determination module 333, and the fourth dominance value determination module 334, respectively, which output multiple dominance values. The maximum dominance value 350 is then determined from these multiple dominance values. Based on the maximum dominance value 350, the corresponding dominance value determination module is deduced as the target dominance value determination module. The number of candidate relevant resources in the target dominance value determination module is taken as the target relevant resource count. For example, if the target dominance value determination module is the third dominance value determination module 333, and since the third dominance value determination module 333 has 2 candidate relevant resources, then the target relevant resource count can be determined to be 2.

[0069] According to embodiments of this disclosure, by setting multiple advantage value determination modules and using each of these modules to output its own advantage value, and determining the number of target-related resources based on the maximum advantage value, the optimal number of target-related resources can be dynamically determined based on the user's real-time needs and behavior. This can better adapt to complex and ever-changing user needs and scenarios, and improve the effectiveness of resource recommendations.

[0070] According to an embodiment of the present disclosure, the resource recommendation result comprises at least one target recommended resource; and performing the resource recommendation operation according to the recommendation reference feature and the updated recommendation strategy obtained by updating the preset recommendation strategy according to the related resource recommendation parameter can comprise: recalling a plurality of candidate recommended resources for the target object according to the recommendation reference feature; and screening the plurality of candidate recommended resources according to the updated recommendation strategy to obtain the at least one target recommended resource.

[0071] Recalling a plurality of candidate recommended resources for the target object according to the recommendation reference feature can comprise: performing recall based on display label matching of the target resource; performing recall based on collaborative filtering of the similarity between the target object and the target resource; performing recall based on popular resources; and performing recall based on similarity matching of implicit representations of the target object and the resources learned by a double-tower model.

[0072] Screening the plurality of candidate recommended resources according to the updated recommendation strategy to obtain the at least one target recommended resource can comprise: after recalling the plurality of candidate recommended resources for the target object according to the recommendation reference feature, performing multi-level screening of the plurality of candidate recommended resources according to the updated recommendation strategy to obtain the target recommended resource.

[0073] For example, performing screening of the plurality of candidate recommended resources according to the updated recommendation strategy to a first precision to obtain a first preset number of first screened resources; performing screening of the first screened resources according to the updated recommendation strategy to a second precision to obtain a second preset number of second screened resources target recommended resources; and performing screening of the second screened resources according to the updated recommendation strategy to a third precision to obtain the at least one target recommended resource.

[0074] Screening to the first precision represents a coarse sorting stage of the candidate recommended resources. In this stage, a simple model such as a lightweight neural network can be used, and the main focus is on the core features of the resources, so as to quickly screen out the first preset number of first screened resources from the plurality of candidate recommended resources. For example, there are 5000 resources in the plurality of candidate recommended resources, and 1500 first screened resources are obtained by screening to the first precision. It should be noted that in the screening process in this stage, it is necessary to ensure that the proportion of related resources in the screened resources meets the related resource recommendation parameter. For example, the related resource recommendation parameter is 0.5, and therefore 750 related resources are included in the 1500 first screened resources.

[0075] The second-precision screening represents a fine arrangement stage of the candidate recommended resources. In this stage, a complex model such as a deep neural network can be used to comprehensively consider multi-dimensional features such as diversity of resources, context information of resources, and the like, and to screen a second preset number of candidate resources from the first-screened resources. For example, there are 1500 first-screened resources, and through the second-precision screening, 200 second-screened resources are obtained. It should be noted that in the screening process in this stage, it is necessary to ensure that the proportion of related resources in the screened resources meets the related resource recommendation parameter. For example, the related resource recommendation parameter is 0.5, and therefore, among the 200 second-screened resources, 100 related resources need to be included.

[0076] The third-precision screening represents a further fine arrangement stage of the candidate recommended resources. In this stage, a complex model such as a deep neural network can be used to comprehensively consider multi-dimensional features such as diversity of resources, context information of resources, and the like, and to screen a preset number of target recommended resources from the second-screened resources. For example, there are 200 second-screened resources, and through the third-precision screening, 10 target recommended resources are obtained. It should be noted that in the screening process in this stage, it is necessary to ensure that the proportion of related resources in the screened resources meets the related resource recommendation parameter. For example, the related resource recommendation parameter is 0.5, and therefore, among the 10 target recommended resources, 5 related resources need to be included.

[0077] According to the embodiments of the present disclosure, by recalling a plurality of candidate recommended resources for a target object according to the recommendation reference features, the universality and richness of the candidate resources are ensured. Then, the candidate recommended resources are screened by using the updated recommendation strategy to obtain the target recommended resources, which can ensure that in the subsequent fine arrangement stage, the resources can be screened based on the updated recommendation strategy, and the target recommended resources that meet the interests and needs of the target object are further accurately screened, thereby improving the efficiency and quality of resource recommendation.

[0078] According to the embodiments of the present disclosure, the updated recommendation strategy includes a related resource recommendation parameter and a preset recommendation parameter. According to the updated recommendation strategy, screening a plurality of candidate recommended resources to obtain at least one target recommended resource can include: dividing the plurality of candidate recommended resources into a first candidate recommended resource set and a second candidate recommended resource set according to the relevance between the candidate recommended resources and the target resources, wherein the first candidate recommended resource set includes candidate recommended resources that meet a preset condition in terms of relevance to the target resources, and the second candidate recommended resource set includes candidate recommended resources that do not meet the preset condition in terms of relevance to the target resources; under the constraints of the related resource recommendation parameter and the preset recommendation parameter, candidate recommended resources are screened from the first candidate recommended resource set and the second candidate recommended resource set respectively to obtain at least one target recommended resource.

[0079] Exemplarily, the relevance between the candidate recommended resources and the target resource can be determined based on a feature matching manner of the content of the resources. For example, the relevance is measured by calculating the similarity between the feature vectors by extracting the content features of the resources, such as keywords, tags, and the like of the resources. The relevance between the candidate recommended resources and the target resource can also be determined based on a user behavior collaborative filtering manner. For example, the user-resource behavior matrix is constructed by collecting the behavior data of the user on the resources, such as clicking, collecting, sharing, watching time, and the like. The relevance is measured by calculating the similarity of the scores of different resources by the user or the degree to which the resources are liked by a similar object group.

[0080] Exemplarily, the resources with the relevance higher than the relevance threshold value can be classified into the first candidate recommended resource set. The resources with the relevance lower than the relevance threshold value can be classified into the second candidate recommended resource set.

[0081] Optionally, in the first candidate recommended resource set, several most relevant resources can be screened according to the relevance. In the second candidate recommended resource set, several resources with unique value can be screened according to the heat, novelty, diversity, and the like of the resources. The resources screened from the two parts are merged to form a recommended list, which is displayed to the user.

[0082] According to the embodiments of the present disclosure, under the constraints of the related resource recommendation parameter and the preset recommendation parameter, the candidate recommended resources are screened from the first candidate recommended resource set and the second candidate recommended resource set respectively to obtain at least one target recommended resource can include: screening the first recommended number of candidate recommended resources from the first candidate recommended resource set according to the first recommended number determined according to the related resource recommendation parameter and the preset recommendation parameter; screening the second recommended number of candidate recommended resources from the second candidate recommended resource set according to the second recommended number determined according to the first recommended number and the preset recommendation parameter; and determining the target recommended resource according to the first recommended number of candidate recommended resources and the second recommended number of candidate recommended resources.

[0083] The first recommended number determined according to the related resource recommendation parameter and the preset recommendation parameter can be a first recommended number obtained by multiplying the related resource recommendation parameter and the preset recommendation parameter. For example,

[0084] The related resource recommendation parameter is 0.5, and the preset recommendation parameter is 6. Therefore, the first recommended number is 3, and the three candidate recommended resources with the highest relevance are screened from the first candidate recommended resource set.

[0085] The second recommended number determined according to the first recommended number and the preset recommendation parameter can be a second recommended number obtained by subtracting the first recommended number from the preset recommendation parameter. For example, in the case where the first recommended number is 3 and the preset recommendation parameter is 6, the second recommended number is 3, and the three candidate recommended resources are screened from the second candidate recommended resource set.

[0086] According to an embodiment of the present disclosure, by dividing the candidate recommended resources into two sets with different relevance levels, and screening from the two sets respectively according to the relevance resource recommendation parameters, the resources with high relevance can ensure that the basic needs of the user are met, while the resources with low relevance can increase the richness and freshness of the recommended results, so that the quality and diversity of the recommended results can be more finely controlled.

[0087] Figure 4 A flowchart for determining a target recommended resource according to an embodiment of the present disclosure is schematically shown.

[0088] As shown in Figure 4 , a plurality of candidate recommended resources 402 for a target object are recalled according to a recommended reference feature 401. According to the relevance between the candidate recommended resources 402 and the target resource, the plurality of candidate recommended resources 402 are divided into a first candidate recommended resource set 403 and a second candidate recommended resource set 404. A first recommended number is determined according to a relevance resource recommendation parameter 405 and a preset recommendation parameter 406, and a candidate recommended resource 407 of the first recommended number is screened from the first candidate recommended resource set 403. A second recommended number is determined according to a first recommended number 408 and a preset recommendation parameter 406, and a candidate recommended resource 409 of the second recommended number is screened from the second candidate recommended resource set 404. According to the candidate recommended resource 407 of the first recommended number and the candidate recommended resource 409 of the second recommended number, the target recommended resource 410 is determined. Thus, the quality and diversity of the recommended results can be more finely controlled.

[0089] According to an embodiment of the present disclosure, each of the at least one target recommended resource in the resource recommendation result has a resource weight. After the plurality of candidate recommended resources are screened to obtain the at least one target recommended resource according to the updated recommendation strategy, the resource recommendation method can further include: in the case where the resource recommendation result includes a target recommended resource of a preset category, adjusting the resource weight of the target recommended resource of the preset category according to a weight adjustment coefficient determined by each of the plurality of advantage value determination modules, to obtain an updated resource recommendation result.

[0090] The target recommended resource of the preset category can be a resource category related to the target resource.

[0091] The resource weight can be used to measure the exposure priority of the target recommended resource. The higher the weight of the target recommended resource, the higher the ranking of the target recommended resource in the recommended result, and the more display opportunities the target recommended resource obtains, and vice versa.

[0092] For example, when the target object prefers the related recommended resources, the weight of the related resources in the target recommended resources can be increased, so that the related resources are ranked higher, and thus the user can be quickly displayed. When the target object prefers the unrelated resources, the weight of the related resources in the target recommended resources can be reduced, so that the unrelated resources are ranked higher, so that the target object can quickly contact the recommended resources of different categories.

[0093] According to an embodiment of the present disclosure, by adjusting the resource weight of the target recommended resources by using the weight adjustment coefficient, an updated resource recommendation result is obtained, and the multiple target recommended resources can be sequentially displayed to the target object according to the exposure priority, so that the resource with the most interest of the target object is ranked higher, thereby improving the effect of resource recommendation.

[0094] According to an embodiment of the present disclosure, the weight adjustment coefficient can be determined according to an average advantage value determined according to the multiple advantage values and a maximum advantage value in the multiple advantage values.

[0095] According to an embodiment of the present disclosure, the weight adjustment coefficient can be determined according to an average advantage value determined according to the multiple advantage values and a maximum advantage value in the multiple advantage values.

[0096] Illustratively, the weight adjustment coefficient is calculated according to formula (1).

[0097] (1)

[0098] Wherein, x is the number of target related resources, X is the maximum value of the number of related resources, A is the advantage value output by the advantage value model, and the number of related resources The corresponding advantage value is denoted as . That is, if is negative, the weight is reduced, otherwise the weight is increased; if the greater the deviation from the average value, the greater the weight increase or decrease, otherwise the smaller.

[0099] Figure 5 The flowchart of the model training method according to an embodiment of the present disclosure is schematically shown.

[0100] The model includes a consumption time length determination module and multiple advantage value determination modules. As Figure 5 shown, the method includes operations S510-S540.

[0101] At operation S510, a training sample is obtained, where the training sample includes a sample recommendation reference feature, a sample relevant resource number, and a sample label.

[0102] At operation S520, the training sample is input into a consumption time determination module and a plurality of advantage value determination modules to obtain a sample consumption time and a sample advantage value.

[0103] At operation S530, a total sample consumption time is determined according to the sample consumption time and a target sample advantage value in the plurality of sample advantage values, where the target sample advantage value is determined according to the sample relevant resource number.

[0104] At operation S540, a network parameter of a relevant resource number determination model is adjusted according to a difference between the total sample consumption time and the sample label to obtain a trained relevant resource number determination model.

[0105] The consumption time determination module is used to evaluate an expected value of the user consumption time under the state of the current recommendation reference feature. Through the consumption time determination module, the consumption time of a similar object group can be estimated based on the recommendation reference feature, so as to eliminate individual bias. The plurality of advantage value determination modules are used to calculate an additional consumption time benefit brought by selecting different target relevant resource numbers under the state of the current recommendation reference feature, which can reflect the influence of different target relevant resource numbers on the long-term consumption time of the target object.

[0106] The sample recommendation reference feature includes a sample object behavior feature, a sample object attribute feature, a sample context feature, and a sample entry resource feature.

[0107] Figure 6A A schematic diagram of training sample construction according to an embodiment of the present disclosure is schematically shown.

[0108] As shown in Figure 6A , the sample reference feature can select different sample resources 601 based on the sample object, and the follow-up resources of the sample resources and the consumption time of the sample object are sorted in time sequence to obtain a plurality of different sample features, such as sample feature 1, sample feature 2, and sample feature 3. Each sample feature includes a sample reference feature S i , a sample resource number T i , and a sample label R i , where the sample recommendation reference feature is a sample object behavior feature, a sample object attribute feature, a sample context feature, and a sample entry resource feature; the sample resource number is a follow-up resource number, and the sample label R i is a consumption time. When training the model, a four-tuple (S i , T i , R i , S i+1) As a training sample, the input consumption time length determination module and the plurality of advantage value determination modules are trained to obtain a trained related resource number determination model.

[0109] The target advantage value can be the maximum advantage value in the plurality of sample advantage values. The target advantage value represents the additional consumption time length brought by the sample related resource number, which can also be referred to as a sample decay time length.

[0110] According to an embodiment of the present disclosure, by inputting the training sample into the consumption time length determination module and the plurality of advantage value determination modules, the sample consumption time length and the sample advantage value are obtained, and then the total sample consumption time length is obtained through the sample consumption time length and the target sample advantage value. Further, according to the difference between the total sample consumption time length and the sample label, the network parameters of the related resource number determination model are adjusted to obtain a trained related resource number determination model. The consumption time length determination module can be used to learn the relationship between the sample feature and the sample consumption time length, and the plurality of advantage value determination modules can be used to learn the relationship between the related resource number and the additional consumption time length, so as to accurately evaluate the value of the sample related resource number, thereby helping the related resource number determination model to dynamically adjust the recommended related resource number, accurately identifying the difference in demand of different target objects in different scenarios, and improving the resource recommendation effect.

[0111] According to an embodiment of the present disclosure, the training method of the above model can further include: determining a first loss value according to the sample consumption time length and the sample label; and adjusting the network parameters of the related resource number determination model according to the difference between the total sample consumption time length and the sample label, including: determining a second loss value according to the total sample consumption time length and the sample label; and adjusting the network parameters of the related resource number determination model according to a total loss value determined by the first loss value and the second loss value.

[0112] Exemplarily, the first loss value and the second loss value can be calculated using the same loss function, which can be, for example, a mean square error loss function, and the present disclosure is not limited thereto. The loss function can also be any other loss function. The weights of the first loss value and the second loss value can be the same or different.

[0113] According to an embodiment of the present disclosure, by determining the first loss value based on the consumption time length and the sample label and determining the second loss based on the total sample consumption time length and the sample label, the network parameters of the related resource number determination model can be adjusted to ensure that the advantage determination module learns the increase or decrease of the subsequent consumption time length after performing the corresponding sample related resource number, separate the value function and the advantage function, eliminate individual bias, and improve the accuracy and stability of the model.

[0114] Figure 6B A model architecture schematic diagram of model training according to an embodiment of the present disclosure is schematically shown.

[0115] As shown in Figure 6B , the training sample 610, such as a quadruple (S i , T i , R i , S i+1 ), is input into the multi-layer perceptron 620 to generate a splicing vector, and the splicing vector is input into the first advantage value determination module 631, the second advantage value determination module 632, the third advantage value determination module 633, and the fourth advantage value determination module 634, respectively, so that a plurality of sample advantage values can be output, and the target sample advantage value 650 can be determined from the plurality of sample advantage values. The splicing vector is input into the consumption duration determination module 640, and the sample consumption duration 660 can be obtained. The total consumption duration can be obtained according to the sample consumption duration 660 and the target sample advantage value 650. The first loss value is determined according to the sample consumption duration 660 and the sample label. The second loss value is determined according to the total sample consumption duration 670 and the sample label; and the total loss value determined according to the first loss value and the second loss value is used to adjust the network parameters of the related resource number determination model.

[0116] It should be noted that during training, there are two networks with the same structure, i.e. Figure 6B , and the network structures are denoted as the original network Q and the mirror network . Only before training of each hour of data, the parameters of the original network Q are copied to the mirror network . In addition, the mirror network will not perform any parameter update.

[0117] A quadruple (S i , T i , R i , S i+1 ) constitutes a sample, and the original network Q has two losses, denoted as loss1 and loss2, each with a weight of 0.5. Both losses use a mean square error loss function, and the labels label and the estimated values of the two losses are as follows:

[0118] For loss1, the label is . Wherein is the consumption duration; γ is a decay coefficient, which is generally 0.9 by experiment, and the decay coefficient can ensure the stability of the training; is the estimated subsequent recommended consumption duration, i.e., the estimated subsequent end duration after exiting the current video page; if the user in the training sample does not have subsequent consumption behavior on the same day, then is directly replaced by 0. The estimated value is the total consumption duration obtained by adding the sample consumption duration and the target sample advantage value.

[0119] For loss2, the label is , and loss1 is the same, so as to ensure that the advantage value determination module learns the increase or decrease of the subsequent average end length after performing the corresponding action as much as possible, separate the value function and the advantage function, eliminate individual bias, and improve the accuracy and stability of the model. The estimated value is the sample consumption length.

[0120] Figure 7 A module diagram of a resource recommendation apparatus according to an embodiment of the present disclosure is schematically shown.

[0121] As shown in Figure 7 , the resource recommendation apparatus 700 includes an acquisition module 710, a determination module 720, and a resource recommendation module 730.

[0122] The acquisition module 710 is configured to, in response to a target object selecting a target resource, acquire a recommendation reference feature for the target object.

[0123] The determination module 720 is configured to determine, according to the recommendation reference feature, a related resource recommendation parameter for the target object, wherein the related resource recommendation parameter represents a proportion of recommendation resources that meet a preset condition in terms of relevance to the target resource in a resource recommendation result for the target object.

[0124] The resource recommendation module 730 is configured to perform a resource recommendation operation according to the recommendation reference feature and an updated recommendation strategy obtained by updating a preset recommendation strategy according to the related resource recommendation parameter, to obtain a resource recommendation result for the target object.

[0125] According to an embodiment of the present disclosure, the determination module 720 includes a first determination sub-module and a second determination sub-module.

[0126] The first determination sub-module is configured to determine, according to the recommendation reference feature and a related resource number determination model, a target related resource number for the target object, wherein the target related resource number represents a quantity of recommendation resources that meet a preset condition in terms of relevance to the target resource in a resource recommendation result for the target object.

[0127] The second determination sub-module is configured to determine, according to the target related resource number and a preset recommendation parameter, the related resource recommendation parameter.

[0128] According to an embodiment of the present disclosure, the related resource number determination model includes a plurality of advantage value determination modules, and each of the plurality of advantage value determination modules is configured with a candidate related resource number. The first determination sub-module includes an advantage value determination unit and a resource number determination unit.

[0129] The advantage value determination unit is configured to input the recommendation reference feature into each of the plurality of advantage value determination modules, and output an advantage value of each of the plurality of advantage value determination modules, wherein the advantage value represents a marginal contribution of the candidate related resource number to a consumption length of the target object.

[0130] The resource number determination unit is configured to determine the target relevant resource number according to the maximum advantage value in the plurality of advantage values.

[0131] According to an embodiment of the present disclosure, the resource number determination unit comprises a first determination sub-unit and a second determination sub-unit.

[0132] The first determination sub-unit is configured to determine that the advantage value determination module corresponding to the maximum advantage value in the plurality of advantage values is the target advantage value determination module.

[0133] The second determination sub-unit is configured to determine the candidate relevant resource number of the target advantage value determination module as the target relevant resource number.

[0134] According to an embodiment of the present disclosure, the resource recommendation module 730 comprises a recall sub-module and a screening sub-module.

[0135] The recall sub-module is configured to recall a plurality of candidate recommended resources for the target object according to the recommendation reference feature.

[0136] The screening sub-module is configured to screen the plurality of candidate recommended resources according to an updated recommendation strategy to obtain at least one target recommended resource.

[0137] According to an embodiment of the present disclosure, the updated recommendation strategy comprises a relevant resource recommendation parameter and a preset recommendation parameter. The screening sub-module comprises a division unit and a screening unit.

[0138] The division unit is configured to divide the plurality of candidate recommended resources into a first candidate recommended resource set and a second candidate recommended resource set according to the relevance between the candidate recommended resources and the target resource, wherein the first candidate recommended resource set comprises candidate recommended resources whose relevance to the target resource satisfies a preset condition, and the second candidate recommended resource set comprises candidate recommended resources whose relevance to the target resource does not satisfy the preset condition.

[0139] The screening unit is configured to screen candidate recommended resources from the first candidate recommended resource set and the second candidate recommended resource set respectively under the constraints of the relevant resource recommendation parameter and the preset recommendation parameter to obtain at least one target recommended resource.

[0140] According to an embodiment of the present disclosure, the screening unit comprises a first screening sub-unit and a second screening sub-unit.

[0141] The first screening sub-unit is configured to screen a first recommended number of candidate recommended resources from the first candidate recommended resource set according to a first recommended number determined according to the relevant resource recommendation parameter and the preset recommendation parameter.

[0142] The second screening sub-unit is configured to screen, from the second candidate recommendation resource set, a second recommended number of candidate recommendation resources according to the first recommended number and a second recommended number determined according to the preset recommendation parameter.

[0143] The resource determination sub-unit is configured to determine the target recommendation resource according to the first recommended number of candidate recommendation resources and the second recommended number of candidate recommendation resources.

[0144] According to an embodiment of the present disclosure, each of the at least one target recommendation resource in the resource recommendation result has a resource weight; and the resource recommendation device further comprises an updating module.

[0145] The updating module is configured to, in a case where the resource recommendation result comprises target recommendation resources of a preset category, adjust the resource weight of the target recommendation resources of the preset category according to a weight adjustment coefficient determined according to the respective advantage value of the plurality of advantage value determination modules, to obtain an updated resource recommendation result.

[0146] According to an embodiment of the present disclosure, the updating module further comprises a weight determination sub-module.

[0147] The weight adjustment coefficient is determined according to an average advantage value of the plurality of advantage values and a maximum advantage value in the plurality of advantage values.

[0148] According to an embodiment of the present disclosure, the weight determination sub-module comprises a deviation degree determination unit and an adjustment unit.

[0149] The deviation degree determination unit is configured to determine a deviation degree amplification factor according to a deviation degree between the average advantage value and the maximum advantage value.

[0150] The adjustment unit is configured to determine the weight adjustment coefficient according to a basic regulation factor obtained by normalizing the maximum advantage value and the deviation degree amplification factor.

[0151] It should be noted that the training device part of the image recognition model in the embodiments of the present disclosure corresponds to the training method part of the image recognition model in the embodiments of the present disclosure, and the description of the training device part of the image recognition model is specifically referred to the training method part of the image recognition model, which will not be repeated here.

[0152] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0153] According to an embodiment of the present disclosure, an electronic device comprises at least one processor and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as above.

[0154] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method as above.

[0155] According to an embodiment of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method as above.

[0156] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown schematically. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0157] As shown in Figure 8 The electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0158] Various components in the electronic device 800 are connected to the input / output (I / O) interface 805, including an input unit 806, such as a keyboard, a mouse, and the like; an output unit 807, such as various types of displays, speakers, and the like; a storage unit 808, such as a magnetic disk, an optical disk, and the like; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0159] The computing unit 801 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs various methods and processes described above, such as the resource recommendation method, the training method of the large model. For example, in some embodiments, the resource recommendation method, the training method of the large model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded onto the RAM 803 and executed by the computing unit 801, one or more steps of the resource recommendation method, the training method of the large model described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the resource recommendation method, the training method of the large model by any other suitable means, such as by means of firmware.

[0160] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0161] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0163] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0164] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0165] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0166] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, without limitation herein, so long as the desired results of the technology described in the present disclosure are achieved.

[0167] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A resource recommendation method, comprising: in response to a target object selecting a target resource, obtaining a recommendation reference feature for the target object; determining a relevant resource recommendation parameter for the target object according to the recommendation reference feature, wherein the relevant resource recommendation parameter represents a proportion of recommendation resources in a resource recommendation result for the target object that meet a preset condition in terms of relevance to the target resource; and performing a resource recommendation operation according to the recommendation reference feature and an updated recommendation strategy obtained by updating a preset recommendation strategy with the relevant resource recommendation parameter, to obtain the resource recommendation result for the target object; wherein determining the relevant resource recommendation parameter for the target object according to the recommendation reference feature comprises: inputting the recommendation reference feature into a plurality of advantage value determination modules in a relevant resource number determination model respectively, to output an advantage value of each of the plurality of advantage value determination modules, wherein each of the plurality of advantage value determination modules is configured with a candidate relevant resource number, and the advantage value represents a marginal contribution of the candidate relevant resource number to a consumption time length of the target object; determining a target relevant resource number according to a maximum advantage value in the plurality of advantage values, wherein the target relevant resource number represents a number of recommendation resources in the resource recommendation result for the target object that meet the preset condition in terms of relevance to the target resource; and determining the relevant resource recommendation parameter according to the target relevant resource number and a preset recommendation parameter.

2. The method of claim 1, wherein: determining the target relevant number according to the maximum advantage value in the plurality of advantage values comprises: determining an advantage value determination module corresponding to the maximum advantage value in the plurality of advantage values as a target advantage value determination module; and determining a candidate relevant resource number of the target advantage value determination module as the target relevant resource number.

3. The method of claim 1, wherein, the resource recommendation result comprises at least one target recommendation resource; performing the resource recommendation operation according to the recommendation reference feature and the updated recommendation strategy obtained by updating the preset recommendation strategy with the relevant resource recommendation parameter comprises: recalling a plurality of candidate recommendation resources for the target object according to the recommendation reference feature; and screening the plurality of candidate recommendation resources according to the updated recommendation strategy to obtain the at least one target recommendation resource.

4. The method of claim 3, wherein, the updated recommendation strategy comprises the relevant resource recommendation parameter and the preset recommendation parameter; screening the plurality of candidate recommendation resources according to the updated recommendation strategy to obtain the at least one target recommendation resource comprises: dividing the plurality of candidate recommendation resources into a first candidate recommendation resource set and a second candidate recommendation resource set according to relevance between the candidate recommendation resources and the target resource, wherein the first candidate recommendation resource set comprises candidate recommendation resources that meet the preset condition in terms of relevance to the target resource, and the second candidate recommendation resource set comprises candidate recommendation resources that do not meet the preset condition in terms of relevance to the target resource. The at least one target recommended resource is obtained by screening candidate recommended resources from the first candidate recommended resource set and the second candidate recommended resource set respectively under the constraint of the related resource recommendation parameter and the preset recommendation parameter.

5. The method of claim 4, wherein, The at least one target recommended resource is obtained by screening candidate recommended resources from the first candidate recommended resource set and the second candidate recommended resource set respectively under the constraint of the related resource recommendation parameter and the preset recommendation parameter. A first recommended number of candidate recommended resources is screened from the first candidate recommended resource set according to the first recommended number and the preset recommendation parameter; A second recommended number of candidate recommended resources is screened from the second candidate recommended resource set according to the second recommended number and the preset recommendation parameter; The target recommended resource is determined according to the first recommended number of candidate recommended resources and the second recommended number of candidate recommended resources.

6. The method of claim 1, wherein, The at least one target recommended resource in the resource recommendation result each has a resource weight; The method further comprises: In the case that the resource recommendation result includes target recommended resources of a preset category, a weight adjustment coefficient is adjusted according to the respective advantage value determined by the plurality of advantage value determination modules, the resource weight of the target recommended resources of the preset category is adjusted to obtain an updated resource recommendation result.

7. The method of claim 6, further comprising: The weight adjustment coefficient is determined according to an average advantage value determined by the plurality of advantage values and a maximum advantage value in the plurality of advantage values.

8. The method of claim 7, wherein, The weight adjustment coefficient is determined according to an average advantage value determined by the plurality of advantage values and a maximum advantage value in the plurality of advantage values. The weight adjustment coefficient is determined according to an average advantage value determined by the plurality of advantage values and a maximum advantage value in the plurality of advantage values. The model comprises a consumption time determination module and a plurality of advantage value determination modules; 9. A method of training a model, wherein, The method comprises: A training sample is obtained, wherein the training sample comprises a sample recommended reference feature, a sample related resource number, and a sample label, and the training sample comprises at least one of the following: text, image, audio, and video; The training sample is input into the consumption time determination module and the plurality of advantage value determination modules to obtain a sample consumption time and a sample advantage value; A total sample consumption time is determined according to the sample consumption time and a target sample advantage value in the plurality of sample advantage values, wherein the target sample advantage value is determined according to the sample related resource number; A network parameter of a related resource number determination model is adjusted according to a difference between the total sample consumption time and the sample label to obtain a trained related resource number determination model.

10. The method of claim 9, further comprising: A first loss value is determined according to the sample consumption time and the sample label; The network parameter of the related resource number determination model is adjusted according to a difference between the total sample consumption time and the sample label. ​ determine a second loss value according to the total sample consumption time length and the sample label; adjust network parameters of a relevant resource number determination model according to a total loss value determined according to the first loss value and the second loss value.

11. A resource recommendation apparatus, comprising: an acquisition module configured to acquire a recommendation reference feature for a target object in response to the target object selecting a target resource; a determination module configured to determine a relevant resource recommendation parameter for the target object according to the recommendation reference feature, wherein the relevant resource recommendation parameter represents a proportion of recommendation resources that meet a preset condition in terms of relevance to the target resource in a resource recommendation result for the target object; and a resource recommendation module configured to perform a resource recommendation operation according to the recommendation reference feature and an updated recommendation strategy obtained by updating a preset recommendation strategy according to the relevant resource recommendation parameter, to obtain the resource recommendation result for the target object. The determination module comprises a first determination submodule and a second determination submodule. The first determination submodule comprises: an advantage value determination unit configured to input the recommendation reference feature into a plurality of advantage value determination modules in a relevant resource number determination model respectively, and output an advantage value of each of the plurality of advantage value determination modules, wherein each of the plurality of advantage value determination modules is configured with a candidate relevant resource number, and the advantage value represents a marginal contribution of the candidate relevant resource number to a consumption time length of the target object; a resource number determination unit configured to determine a target relevant resource number according to a maximum advantage value in the plurality of advantage values, wherein the target relevant resource number represents a number of recommendation resources that meet the preset condition in terms of relevance to the target resource in the resource recommendation result for the target object; and the second determination submodule configured to determine the relevant resource recommendation parameter according to the target relevant resource number and a preset recommendation parameter.

12. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.

13. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-10.

14. A computer program product comprising a computer program stored on at least one of a readable storage medium and an electronic device, the computer program being executed by a processor to implement the method of any one of claims 1-10.

Citation Information

Patent Citations

  • Processing method, processing device and electronic equipment

    CN115344788A