Resource recommendation method and device, model training method and device, equipment and medium

By obtaining recommendation reference features and dynamically adjusting resource recommendation strategies through reinforcement learning models, the problem of inflexible resource number adjustment in the existing system is solved, and the personalized and accurate resource recommendation results are achieved.

CN120723977AActive Publication Date: 2025-09-30BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510897120.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing recommendation systems lack a flexible mechanism to adjust the number of related resources, and are unable to dynamically adjust the proportion of resources related to the target resource in the resource recommendation results according to user needs and contextual environment, and are unable to meet the different needs of different users in different scenarios.

Method used

By obtaining the recommendation reference features of the target object, determining the relevant resource recommendation parameters, and dynamically adjusting the recommendation strategy based on these parameters, the model is determined by the number of relevant resources using reinforcement learning. Combined with the value function and advantage function, the value of different numbers of resources is accurately evaluated, and the proportion of relevant resources in the resource recommendation results is dynamically adjusted.

Benefits of technology

It achieves the flexibility to meet the differentiated needs of target objects in different contextual environments, improves the accuracy and effectiveness of resource recommendations, and meets the personalized needs of users in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723977A_ABST
    Figure CN120723977A_ABST
Patent Text Reader

Abstract

The invention provides a resource recommendation method and device, a model training method and device, equipment and a medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of big data, deep learning, intelligent recommendation and the like. The method comprises the steps of selecting a target resource in response to a target object, and obtaining a recommendation reference feature for the target object; according to the recommendation reference features, related resource recommendation parameters for the target object are determined, and the related resource recommendation parameters represent the proportion of recommendation resources, of which the correlation with the target resources meets a preset condition, in a resource recommendation result for the target object; and according to the recommendation reference features and an updated recommendation strategy obtained after a preset recommendation strategy is updated by the related resource recommendation parameters, executing resource recommendation operation to obtain a resource recommendation result for the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to big data, deep learning, intelligent recommendation, and other technical fields. Specifically, it relates to a resource recommendation method, model training method, device, equipment, and medium. Background Art

[0002] With the rapid development of artificial intelligence, recommendation systems can analyze users' historical behaviors, interest preferences, and user attribute characteristics to generate personalized recommendations. They have been widely used in various application scenarios, especially in online videos, e-commerce platforms, social media and other fields. Summary of the Invention

[0003] The present disclosure provides a resource recommendation method, a model training method, an apparatus, a device, and a medium.

[0004] According to one aspect of the present disclosure, a resource recommendation method is provided, including: in response to a target object selecting a target resource, obtaining a recommendation reference feature for the target object; determining a relevant resource recommendation parameter for the target object based on the recommendation reference feature, wherein the relevant resource recommendation parameter represents the proportion of recommended resources whose relevance to the target resource meets a preset condition in the resource recommendation result for the target object; and executing a resource recommendation operation based on the recommendation reference feature and an updated recommendation strategy obtained after updating the preset recommendation strategy by the relevant resource recommendation parameter, to obtain a resource recommendation result for the target object.

[0005] According to another aspect of the present disclosure, a model training method is provided, wherein the model includes a consumption time determination module and multiple advantage value determination modules; the method includes: obtaining training samples, wherein the training samples include sample recommendation reference features, the number of sample-related resources, and sample labels; inputting the training samples into the consumption time determination module and the multiple advantage value determination modules to obtain sample consumption time and sample advantage values; determining the total sample consumption time based on the sample consumption time and the target sample advantage value among the multiple sample advantage values, wherein the target sample advantage value is determined based on the number of sample-related resources; adjusting the network parameters of the related resource number determination model based on the difference between the total sample consumption time and the sample label to obtain a trained related resource number determination model.

[0006] According to another aspect of the present disclosure, a resource recommendation device is provided, including: an acquisition module for acquiring a recommendation reference feature for the target object in response to a target object selecting a target resource; a determination module for determining relevant resource recommendation parameters for the target object based on the recommendation reference feature, wherein the relevant resource recommendation parameter represents the proportion of recommended resources whose relevance to the target resource meets a preset condition in the resource recommendation result for the target object; and a resource recommendation module for executing a resource recommendation operation based on the recommendation reference feature and an updated recommendation strategy obtained after updating the preset recommendation strategy by the relevant resource recommendation parameter, so as to obtain a resource recommendation result for the target object.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the above method.

[0009] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the above method when executed by a processor.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0012] Figure 1 Schematically illustrates an exemplary system architecture to which a resource recommendation method, a model training method, and an intelligent agent according to an embodiment of the present disclosure can be applied;

[0013] Figure 2 The following schematically shows a flow chart of a resource recommendation method according to an embodiment of the present disclosure;

[0014] Figure 3 A schematic diagram of a model architecture of a dominant value determination model according to an embodiment of the present disclosure is shown;

[0015] Figure 4 Schematically shows a flow chart of determining target recommended resources according to an embodiment of the present disclosure;

[0016] Figure 5 The following schematically shows a flow chart of a model training method according to an embodiment of the present disclosure;

[0017] Figure 6A Schematic diagram showing the construction of training samples according to an embodiment of the present disclosure

[0018] Figure 6B A schematic diagram of a model architecture for model training according to an embodiment of the present disclosure is shown schematically;

[0019] Figure 7 Schematically shows a module diagram of a resource recommendation device according to an embodiment of the present disclosure;

[0020] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0021] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0022] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0023] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0024] In a related embodiment, when recommending resources to a user, the recommendation system usually recommends a fixed number of recommended resources to the user, wherein the recommended resources are usually related resources related to the current resource being browsed by the user. This method lacks a flexible mechanism to adjust the number of related resources, cannot determine how many related resources to display based on user needs and context, and does not take into account the differences in user needs for recommended resources. In addition, because some users may prefer more related resources related to the current resource in certain scenarios, while other users may prefer recommended resources that are completely different from the current resource, the above recommendation method cannot dynamically adjust the number of related resources to adapt to the needs of different users.

[0025] Based on the above technical problems, the present disclosure provides a resource recommendation method, including: in response to a target object selecting a target resource, obtaining a recommendation reference feature for the target object; determining relevant resource recommendation parameters for the target object based on the recommendation reference feature, wherein the relevant resource recommendation parameter represents the proportion of recommended resources whose relevance to the target resource meets preset conditions in the resource recommendation results for the target object; and executing a resource recommendation operation based on the recommendation reference feature and an updated recommendation strategy obtained after updating the preset recommendation strategy by the relevant resource recommendation parameter to obtain a resource recommendation result for the target object.

[0026] According to the embodiments of the present disclosure, by adopting the response target object to select the target resource, obtaining the recommendation reference features for the target object, and determining the relevant resource recommendation parameters based on the recommendation reference features, and dynamically adjusting the recommendation strategy using the relevant resource recommendation parameters, the proportion of relevant resources related to the target resource in the resource recommendation results can be dynamically adjusted to flexibly meet the target object's demand differences for relevant resources in different contextual environments, accurately identify the different needs of different target objects in different scenarios, and improve the resource recommendation effect.

[0027] Figure 1 The exemplary system architecture of the resource recommendation method, model training method and device according to the embodiments of the present disclosure is schematically shown.

[0028] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the resource recommendation method, model training method, and apparatus may be applied may include a terminal device, but the terminal device may implement the resource recommendation method, model training method, and apparatus provided in the embodiments of the present disclosure without interacting with a server.

[0029] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0030] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).

[0031] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0032] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0033] It should be noted that the resource recommendation method and model training method provided in the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103. Accordingly, the resource recommendation apparatus provided in the embodiments of the present disclosure can also be set in the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0034] Alternatively, the resource recommendation method and model training method provided by the embodiments of the present disclosure may also generally be executed by the server 105. Accordingly, the resource recommendation apparatus provided by the embodiments of the present disclosure may generally be provided in the server 105. The resource recommendation method and model training method provided by the embodiments of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and that can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the resource recommendation apparatus provided by the embodiments of the present disclosure may also be provided in a server or server cluster that is different from the server 105 and that can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0035] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0036] Figure 2 The flowchart of the resource recommendation method according to an embodiment of the present disclosure is schematically shown.

[0037] like Figure 2 As shown, the resource recommendation method of this embodiment includes operations S210 to S230.

[0038] In operation S210 , in response to a target object selecting a target resource, a recommended reference feature for the target object is acquired.

[0039] The target object refers to the subject that performs operations such as resource selection in various recommendation systems, platforms, or scenarios. For example, the target object can be a consumer shopping on an e-commerce platform, or a user watching a video on a video platform.

[0040] A target resource is a resource selected by the target user from the resource list. This resource can be any form of information, data, services, or products. For example, on a video platform, the target resource might be video content; on an e-commerce platform, the target resource might be various products; and on social media, the target resource might be various text, images, audio, or video content.

[0041] Recommendation reference features include characteristics of the target object across multiple dimensions. These characteristics may include object behavior features, object attribute features, contextual features, and entry resource features. For example, object behavior features may include the target object's click history and browsing time. Object attribute features may include the target object's age, gender, and region. Contextual features may include the time, location, and status of the target object when selecting the target resource. Entry resource features may include the target resource's type, topic, and tags.

[0042] It should be noted that users are aware of and agree to the acquisition and use of recommended reference features, which complies with relevant laws and regulations and does not violate public order and good morals.

[0043] In operation S220 , a relevant resource recommendation parameter for the target object is determined based on the recommendation reference feature. The relevant resource recommendation parameter represents the proportion of recommended resources whose relevance to the target resource meets a preset condition in the resource recommendation result for the target object.

[0044] Exemplarily, the preset condition may be a correlation threshold. When the correlation between the recommended resource and the target resource is greater than the correlation threshold, the recommended resource is considered a relevant resource. When the correlation between the recommended resource and the target resource is less than the correlation threshold, the recommended resource is considered an irrelevant resource.

[0045] The resource recommendation results can include related resources that are related to the target resource, as well as unrelated resources that are unrelated to the target resource. The related resource recommendation parameter indicates the proportion of related resources in the resource recommendation results.

[0046] For example, on a video platform, after a target object selects a target video, there are 6 resource recommendation results for the target object each time, of which 3 are relevant resources and 3 are irrelevant resources. The relevant resource recommendation parameter is 0.5.

[0047] In operation S230 , a resource recommendation operation is performed according to the recommendation reference features and the updated recommendation strategy obtained by updating the preset recommendation strategy with the relevant resource recommendation parameters to obtain a resource recommendation result for the target object.

[0048] The preset recommendation strategy can include only the number of resources recommended to the target object; it can also include the number of resources recommended to the target object and the composition strategy of relevant and irrelevant resources in the recommended resources. The preset update recommendation strategy can be updated using the relevant resource recommendation parameters to adjust the ratio of relevant and irrelevant resources.

[0049] It should be noted that the relevant resource recommendation parameters can be dynamically updated based on the target resource selected by the target object. For example, after the target object has browsed a target resource, they can select the next target resource from the resource recommendation results to continue browsing. At this time, the target object's recommendation reference characteristics can be updated based on the target resource they newly selected. The relevant resource recommendation parameters can also be updated based on the updated recommendation reference characteristics, thereby further updating the recommendation strategy and dynamically adjusting the proportion of relevant resources in the resource recommendation results.

[0050] According to the embodiments of the present disclosure, by adopting the response target object to select the target resource, obtaining the recommendation reference features for the target object, and determining the relevant resource recommendation parameters based on the recommendation reference features, and dynamically adjusting the recommendation strategy using the relevant resource recommendation parameters, the proportion of relevant resources related to the target resource in the resource recommendation results can be dynamically adjusted to flexibly meet the target object's demand differences for relevant resources in different contextual environments, accurately identify the different needs of different target objects in different scenarios, and improve the resource recommendation effect.

[0051] According to an embodiment of the present disclosure, determining the relevant resource recommendation parameters for the target object based on the recommendation reference characteristics may include: determining a model based on the recommendation reference characteristics and the number of relevant resources to determine the number of target relevant resources for the target object; the number of target relevant resources represents the number of recommended resources whose relevance to the target resource meets preset conditions in the resource recommendation results for the target object; and determining the relevant resource recommendation parameters based on the number of target relevant resources and the preset recommendation parameters.

[0052] The number of target-related resources indicates the number of related resources included in the resource recommendation results.

[0053] The preset recommendation parameter represents the number of resource recommendations displayed to the target user in response to the target resource selection. This is the number of recommendations pre-configured in the recommendation system. For example, on a video platform, the preset recommendation parameter might be 6. If the target user selects a target video to watch, 6 additional recommended videos will be recommended to the target user accordingly.

[0054] The recommended parameters for related resources can be obtained based on the ratio of the number of target related resources and the preset recommended parameters.

[0055] For example, the model for determining the number of related resources can be a deep neural network model combined with reinforcement learning. Reinforcement learning uses a value function and an advantage function to estimate the total value and advantage of each action in the current state, respectively, thereby improving learning efficiency and model stability.

[0056] In this embodiment, the value function can be used to evaluate the expected value of the user's consumption time under the current recommended reference characteristics. Through the value function, the related resource number determination model can estimate the consumption time of a group of similar objects based on the recommended reference characteristics, thereby eliminating individual bias. The advantage function is used to calculate the benefits of additional consumption time brought about by selecting different target-related resource numbers under the current recommended reference characteristics. The advantage function can reflect the impact of different target-related resource numbers on the long-term consumption time of the target object, thereby helping the related resource number determination model to dynamically adjust the recommended target-related resource numbers.

[0057] According to the embodiments of the present disclosure, by introducing a reinforcement learning-based model for determining the number of related resources and separating the value function and the advantage function, the model for determining the number of related resources can accurately evaluate the value of different numbers of related resources and meet the different needs of different users for recommended resources.

[0058] According to an embodiment of the present disclosure, a model for determining the number of related resources includes multiple advantage value determination modules, each of which is configured with a candidate number of related resources. Determining a target number of related resources for a target object based on recommended reference features and the model for determining the number of related resources may include: inputting the recommended reference features into the multiple advantage value determination modules, outputting advantage values ​​for each of the multiple advantage value determination modules, wherein the advantage values ​​represent the marginal contribution of the candidate number of related resources to the consumption time of the target object; and determining the target number of related resources based on the maximum advantage value among the multiple advantage values.

[0059] Exemplarily, the model for determining the number of related resources may include two branches. The first branch may include multiple advantage value determination modules, and the second branch is a consumption duration determination module.

[0060] Each advantage value determination module is configured with a different number of candidate related resources. For example, when the preset recommendation parameter is 6, 4 advantage value determination modules can be set, and the number of candidate related resources of multiple advantage value determination modules can be 0, 1, 2, and 3 respectively.

[0061] The recommended reference features are input into multiple advantage value determination modules, which use them to estimate the advantage function and output multiple advantage values ​​corresponding to different numbers of related resources. The recommended reference features are input into the consumption duration determination module, which uses them to estimate the value function and output the consumption duration.

[0062] Determining the target number of related resources for the target object based on the recommended reference features and the related resource number determination model may further include: using object attribute features in the recommended reference features as category identifiers, extracting them into corresponding feature representations through an encoder, processing the feature representations together with object behavior features, context features, and entry resource features in the recommended reference features through a multi-layer perceptron to generate a concatenated vector. The concatenated vector is input into multiple advantage value determination modules and a consumption duration determination module of the related resource number determination model, respectively, to output multiple advantage values ​​and consumption durations.

[0063] The maximum advantage value indicates that the target object's consumption duration is maximized for the corresponding number of related resources. The advantage value can be used to measure the additional effect of the number of candidate related resources on the target object's consumption duration.

[0064] According to an embodiment of the present disclosure, determining the target number of related items based on the maximum advantage value among multiple advantage values ​​may include: determining the advantage value determination module corresponding to the maximum advantage value among multiple advantage values ​​as the target advantage value determination module; determining the number of candidate related resources of the target advantage value determination module as the target number of related resources.

[0065] After multiple advantage values ​​are determined by using multiple advantage value determination modules to output multiple advantage values, a target advantage module and the number of target-related resources corresponding to the maximum advantage value can be determined by reverse query.

[0066] Figure 3 The figure schematically shows a model architecture diagram of a model for determining the number of related resources according to an embodiment of the present disclosure.

[0067] like Figure 3 As shown, the model for determining the number of related resources may include a multi-layer perceptron 320, a dominance value determination module, and a consumption duration determination module 340. The dominance value determination module may include multiple modules, such as a first dominance value determination module 331, a second dominance value determination module 332, a third dominance value determination module 333, and a fourth dominance value determination module 334. Each dominance value determination module may be configured with a different number of candidate related resources. For example, the number of candidate related resources in the first dominance value determination module 331 is 0, the number of candidate related resources in the second dominance value determination module 332 is 1, the number of candidate related resources in the third dominance value determination module 333 is 2, and the number of candidate related resources in the fourth dominance value determination module 334 is 3.

[0068] The recommended reference feature 310 is input into the multi-layer sensor 320 to generate a splicing vector. The splicing vector is input into the consumption duration determination module 340 to output the consumption duration 360. The splicing vector is input into the first advantage value determination module 331, the second advantage value determination module 332, the third advantage value determination module 333, and the fourth advantage value determination module 334 respectively to output multiple advantage values. The maximum advantage value 350 is determined from the multiple advantage values. Based on the maximum advantage value 350, the corresponding advantage value determination module is deduced to be the target advantage value determination module, and the number of candidate related resources of the target advantage value determination module is used as the target number of related resources. For example, if the target advantage value determination module is the third advantage value determination module 333, and the number of candidate related resources of the third advantage value determination module 333 is 2, then the target number of related resources can be determined to be 2.

[0069] According to the embodiments of the present disclosure, by setting up multiple advantage value determination modules, and utilizing the multiple advantage value determination modules to output their respective advantage values, and determining the number of target-related resources based on the maximum advantage value among the advantage values, the optimal number of target-related resources can be determined based on the user's immediate needs and behavioral dynamics, which can better adapt to complex and changeable user needs and scenarios and improve the effectiveness of resource recommendations.

[0070] According to an embodiment of the present disclosure, the resource recommendation result includes at least one target recommended resource; based on the recommendation reference features and the updated recommendation strategy obtained after updating the preset recommendation strategy by the relevant resource recommendation parameters, executing the resource recommendation operation may include: recalling multiple candidate recommended resources for the target object based on the recommendation reference features; and screening multiple candidate recommended resources according to the updated recommendation strategy to obtain at least one target recommended resource.

[0071] Recalling multiple candidate recommendation resources for a target object based on recommendation reference features may include: recalling based on display tag matching of the target resource; recalling based on collaborative filtering based on the similarity between the target object and the target resource; recalling based on popular resources; and learning implicit representations of the target object and resources through a dual-tower model and performing similarity matching based on the implicit representation to achieve recall.

[0072] According to the updated recommendation strategy, screening multiple candidate recommendation resources to obtain at least one target recommendation resource may include: after recalling multiple candidate recommendation resources for the target object according to the recommendation reference features, performing multi-level screening on the multiple candidate recommendation resources according to the updated recommendation strategy to obtain the target recommendation resource.

[0073] For example, according to the updated recommendation strategy, multiple candidate recommended resources are screened with a first precision to obtain a first preset number of first screened resources; according to the updated recommendation strategy, the first screened resources are screened with a second precision to obtain a second preset number of second screened target recommended resources; according to the updated recommendation strategy, the second screened resources are screened with a third precision to obtain at least one target recommended resource.

[0074] The first precision screening represents the rough sorting stage of candidate recommended resources. In this stage, a simple model, such as a lightweight neural network, can be used, focusing on the core features of the resources to quickly screen out a first preset number of first screened resources from multiple candidate recommended resources. For example, there are 5,000 resources among the multiple candidate recommended resources, and 1,500 first screened resources are obtained through the first precision screening. It should be noted that during the screening process at this stage, it is necessary to ensure that the proportion of relevant resources in the screened resources meets the relevant resource recommendation parameters. For example, if the relevant resource recommendation parameter is 0.5, then 750 relevant resources must be included in the 1,500 first screened resources.

[0075] The second precision screening represents the stage of fine ranking of candidate recommended resources. In this stage, complex models, such as deep neural networks, can be used to comprehensively consider multi-dimensional features, such as resource diversity, resource context information, etc., to screen out a second preset number of multiple candidate resources from the first screening resources. For example, there are 1,500 first screening resources, and through the second precision screening, 200 second screening resources are obtained. It should be noted that during the screening process at this stage, it is necessary to ensure that the proportion of relevant resources in the screened resources meets the relevant resource recommendation parameters. For example, if the relevant resource recommendation parameter is 0.5, then the 200 second screening resources must include 100 relevant resources.

[0076] The third precision screening represents a further refinement stage of candidate recommended resources. In this stage, complex models, such as deep neural networks, can be used to comprehensively consider multi-dimensional features, such as resource diversity, resource context information, etc., to screen out a preset number of target recommended resources from the second screening resources. For example, there are 200 second-screened resources, and through the third precision screening, 10 target recommended resources are obtained. It should be noted that during the screening process at this stage, it is necessary to ensure that the proportion of relevant resources in the screened resources meets the relevant resource recommendation parameters. For example, if the relevant resource recommendation parameter is 0.5, then 5 relevant resources must be included in the 10 target recommended resources.

[0077] According to the embodiments of the present disclosure, by recalling multiple candidate recommendation resources for a target object based on recommendation reference features, the breadth and richness of the candidate resources are ensured. These candidate recommendation resources are then screened using an updated recommendation strategy to obtain the target recommended resource. This ensures that in the subsequent refined ranking stage, resources can be screened based on the updated recommendation strategy to further accurately select those that meet the interests and needs of the target object, thereby improving the efficiency and quality of resource recommendations.

[0078] According to an embodiment of the present disclosure, an update recommendation strategy includes relevant resource recommendation parameters and preset recommendation parameters. According to the update recommendation strategy, screening multiple candidate recommendation resources to obtain at least one target recommendation resource may include: dividing the multiple candidate recommendation resources into a first candidate recommendation resource set and a second candidate recommendation resource set according to the correlation between the candidate recommendation resources and the target resource, wherein the first candidate recommendation resource set includes candidate recommendation resources whose correlation with the target resource meets a preset condition, and the second candidate recommendation resource includes candidate recommendation resources whose correlation with the target resource does not meet the preset condition; under the constraints of the relevant resource recommendation parameters and the preset recommendation parameters, screening candidate recommendation resources from the first candidate recommendation resource set and the second candidate recommendation resource set respectively to obtain at least one target recommendation resource.

[0079] For example, the correlation between candidate recommended resources and target resources can be determined based on a feature matching method of resource content. For example, by extracting content features of resources, such as keywords and tags of resources, the correlation can be measured by calculating the similarity between feature vectors. The correlation between candidate recommended resources and target resources can also be determined based on collaborative filtering of user behavior. For example, user behavior data on resources, such as clicks, favorites, shares, and viewing time, can be collected to construct a user-resource behavior matrix. The correlation can be measured by calculating the similarity of user ratings for different resources or the degree to which resources are liked by similar object groups.

[0080] For example, resources with a correlation higher than a correlation threshold may be classified into a first candidate recommended resource set, and resources with a correlation lower than the correlation threshold may be classified into a second candidate recommended resource set.

[0081] Optionally, the first set of candidate recommended resources can be filtered out based on relevance to identify the most relevant resources. The second set of candidate recommended resources can be filtered out based on popularity, novelty, diversity, and other factors to identify resources with unique value. The resources selected from these two sets are combined to form a recommendation list, which is then presented to the user.

[0082] According to an embodiment of the present disclosure, under the constraints of relevant resource recommendation parameters and preset recommendation parameters, candidate recommendation resources are screened from a first candidate recommendation resource set and a second candidate recommendation resource set respectively to obtain at least one target recommended resource, which may include: screening a first recommended number of candidate recommendation resources from the first candidate recommendation resource set according to a first recommendation number determined by the relevant resource recommendation parameters and the preset recommendation parameters; screening a second recommended number of candidate recommendation resources from the second candidate recommendation resource set according to a second recommendation number determined by the first recommendation number and the preset recommendation parameters; and determining the target recommended resource based on the first recommended number of candidate recommendation resources and the second recommended number of candidate recommendation resources.

[0083] The first recommendation number determined based on the relevant resource recommendation parameter and the preset recommendation parameter can be obtained by multiplying the relevant resource recommendation parameter and the preset recommendation parameter. For example,

[0084] The relevant resource recommendation parameter is 0.5, and the preset recommendation parameter is 6, so the first recommendation number is 3, and the three candidate recommendation resources with the highest relevance are screened out from the first candidate recommendation resource set.

[0085] The second recommended number determined based on the first recommended number and the preset recommendation parameter can be obtained by subtracting the preset recommendation parameter from the first recommended number. For example, if the first recommended number is 3 and the preset recommendation parameter is 6, the second recommended number is 3, and 3 candidate recommendation resources are selected from the second candidate recommendation resource set.

[0086] According to the embodiments of the present disclosure, by dividing candidate recommendation resources into two sets with different relevance, and screening them from the two sets respectively according to the relevance resource recommendation parameters, highly relevant resources can ensure that the basic needs of users are met, while low-relevance resources can increase the richness and freshness of the recommendation results, thereby more finely controlling the quality and diversity of the recommendation results.

[0087] Figure 4 The flowchart of determining target recommended resources according to an embodiment of the present disclosure is schematically shown.

[0088] like Figure 4 As shown, multiple candidate recommendation resources 402 for the target object are recalled based on the recommendation reference feature 401. Based on the correlation between the candidate recommendation resources 402 and the target resource, the multiple candidate recommendation resources 402 are divided into a first candidate recommendation resource set 403 and a second candidate recommendation resource set 404. A first recommendation number is determined based on the relevant resource recommendation parameter 405 and the preset recommendation parameter 406, and the first recommendation number of candidate recommendation resources 407 are screened from the first candidate recommendation resource set 403. Based on the second recommendation number determined based on the first recommendation number 408 and the preset recommendation parameter 406, a second recommendation number of candidate recommendation resources 409 are screened from the second candidate recommendation resource set 404. Based on the first recommendation number of candidate recommendation resources 407 and the second recommendation number of candidate recommendation resources 409, the target recommendation resource 410 is determined. In this way, the quality and diversity of the recommendation results can be more finely controlled.

[0089] According to an embodiment of the present disclosure, at least one target recommended resource in the resource recommendation result each has a resource weight. After screening the multiple candidate recommended resources according to the updated recommendation strategy to obtain the at least one target recommended resource, the resource recommendation method may further include: when the resource recommendation result includes a target recommended resource of a preset category, adjusting the resource weight of the target recommended resource of the preset category according to the weight adjustment coefficient determined by the respective advantage values ​​of the multiple advantage value determination modules, to obtain an updated resource recommendation result.

[0090] The target recommended resource of the preset category may be a resource category related to the target resource.

[0091] Resource weights can be used to measure the exposure priority of target recommended resources. Target recommended resources with higher weights are ranked higher in the recommendation results and receive more exposure opportunities. Conversely, resources with lower weights are ranked lower in the recommendation results and receive less exposure.

[0092] For example, if the target user prefers relevant recommended resources, the weight of relevant resources in the target recommended resources can be increased, so that relevant resources are ranked higher and can be displayed to the user more quickly. If the target user prefers irrelevant resources, the weight of relevant resources in the target recommended resources can be reduced, so that irrelevant resources are ranked higher, allowing the target user to access recommended resources of different categories more quickly.

[0093] According to an embodiment of the present disclosure, by using a weight adjustment coefficient to adjust the resource weight of the target recommended resource, an updated resource recommendation result is obtained. Multiple target recommended resources can be displayed to the target object in sequence according to the exposure priority, so that the resources that the target object is most interested in are ranked higher, thereby improving the effect of resource recommendation.

[0094] According to an embodiment of the present disclosure, the weight adjustment coefficient may be determined in the following manner: the weight adjustment coefficient is determined according to an average advantage value determined from a plurality of advantage values ​​and a maximum advantage value among the plurality of advantage values.

[0095] According to an embodiment of the present disclosure, determining a weight adjustment coefficient based on an average advantage value determined from a plurality of advantage values ​​and a maximum advantage value among the plurality of advantage values ​​may include: determining a deviation amplification factor based on a degree of deviation between the average advantage value and the maximum advantage value; and determining a weight adjustment coefficient based on a basic control factor and a deviation amplification factor obtained by normalizing the maximum advantage value.

[0096] Schematically, the weight adjustment coefficient The calculation formula of is shown in formula (1).

[0097] (1)

[0098] Where: x is the number of target related resources, X is the maximum number of related resources, A is the advantage value output by the advantage value model, and the number of related resources The corresponding advantage value is recorded as That is, if If it is a negative number, it means the weight is reduced, otherwise it means the weight is increased; if The greater the deviation from the mean, the greater the force of weighting increase or decrease, otherwise the smaller the force of weighting decrease.

[0099] Figure 5 The flowchart of the model training method according to an embodiment of the present disclosure is schematically shown.

[0100] The model includes a consumption time determination module and multiple advantage value determination modules. Figure 5 The method shown includes operations S510 to S540.

[0101] In operation S510 , a training sample is obtained, where the training sample includes a sample recommendation reference feature, a number of sample-related resources, and a sample label.

[0102] In operation S520 , the training sample is input into a consumption duration determination module and a plurality of advantage value determination modules to obtain a sample consumption duration and a sample advantage value.

[0103] In operation S530 , a total sample consumption duration is determined according to the sample consumption duration and a target sample advantage value among the plurality of sample advantage values, wherein the target sample advantage value is determined according to the number of sample-related resources.

[0104] In operation S540 , network parameters of a model for determining the number of related resources are adjusted according to the difference between the total sample consumption time and the sample labels, thereby obtaining a trained model for determining the number of related resources.

[0105] The consumption duration determination module is used to evaluate the expected user consumption duration based on the current recommended reference features. This module can estimate the consumption duration of a group of similar users based on the recommended reference features, thereby eliminating individual bias. Multiple advantage value determination modules are used to calculate the additional consumption duration benefits of selecting different numbers of target-related resources under the current recommended reference features. This module can reflect the impact of different numbers of target-related resources on the target user's long-term consumption duration.

[0106] The sample recommendation reference features include sample object behavior features, sample object attribute features, sample context features and sample entry resource features.

[0107] Figure 6A The figure schematically shows a diagram of training sample construction according to an embodiment of the present disclosure.

[0108] like Figure 6A As shown, the sample reference feature can select different sample resources 601 based on the sample object, as well as the follow-up resources of the sample resource and the consumption time of the sample object, and obtain multiple different sample features according to the time series, such as sample feature 1, sample feature 2, and sample feature 3. Each sample feature includes a sample reference feature S i , number of sample resources T i and sample labels R i , where the sample recommendation reference features are sample object behavior features, sample object attribute features, sample context features and sample entry resource features; the number of sample resources is the number of follow-up resources, and the sample label R i is the consumption time. When training the model, the four-tuple (S i 、T i 、R i 、S i+1) as a training sample, the input consumption time determination module and multiple advantage value determination modules are trained to obtain a trained related resource number determination model.

[0109] The target advantage value can be the maximum advantage value among multiple sample advantage values. The target advantage value represents the additional consumption time benefit brought by the number of sample-related resources, which can also be called the sample decay time.

[0110] According to an embodiment of the present disclosure, by inputting training samples into a consumption duration determination module and multiple advantage value determination modules, the sample consumption duration and sample advantage value are obtained, and then the total sample consumption duration is obtained by the sample consumption duration and the target sample advantage value, and then according to the difference between the total sample consumption duration and the sample label, the network parameters of the related resource number determination model are adjusted to obtain a trained related resource number determination model. The consumption duration determination module can be used to learn the relationship between sample parameter features and sample consumption duration, and the multiple advantage value determination modules can be used to learn the relationship between the number of related resources and the additional consumption duration, so as to accurately evaluate the value of the sample related resource number, thereby helping the related resource number determination model to dynamically adjust the recommended related resource number, accurately identify the different needs of different target objects in different scenarios, and improve the resource recommendation effect.

[0111] According to an embodiment of the present disclosure, the training method of the above-mentioned model may further include: determining a first loss value based on the sample consumption time and the sample label; adjusting the number of relevant resources to determine the network parameters of the model based on the difference between the total sample consumption time and the sample label, including: determining a second loss value based on the total sample consumption time and the sample label; adjusting the network parameters of the model based on the number of relevant resources to determine the total loss value determined by the first loss value and the second loss value.

[0112] For example, the first loss value and the second loss value can be calculated using the same loss function, which can be, for example, a mean square error loss function. The present disclosure is not limited thereto, and the loss function can also be any other loss function. The weights of the first loss value and the second loss value can be the same or different.

[0113] According to the embodiments of the present disclosure, by determining the first loss value based on the consumption time and sample labels and the second loss based on the total sample consumption time and sample labels, the network parameters of the model for determining the number of related resources are adjusted. This can ensure as much as possible that the advantage determination module learns the increase or decrease in the subsequent consumption time after executing the corresponding number of sample-related resources, separates the value function and the advantage function, eliminates individual biases, and improves the accuracy and stability of the model.

[0114] Figure 6B A schematic diagram of a model architecture for model training according to an embodiment of the present disclosure is shown schematically.

[0115] like Figure 6B As shown, the training sample 610, such as the four-tuple (S i 、T i 、R i 、S i+1 ), is input into the multi-layer perceptron 620 to generate a splicing vector, which is respectively input into the first advantage value determination module 631, the second advantage value determination module 632, the third advantage value determination module 633, and the fourth advantage value determination module 634, and multiple sample advantage values ​​can be output. From the multiple sample advantage values, a target sample advantage value 650 can be determined. The splicing vector is input into the consumption time determination module 640 to obtain the sample consumption time 660. The total consumption time can be obtained based on the sample consumption time 660 and the target sample advantage value 650. The first loss value is determined based on the sample consumption time 660 and the sample label. The second loss value is determined based on the total sample consumption time 670 and the sample label; the network parameters of the model determined by the number of related resources are adjusted according to the total loss value determined by the first loss value and the second loss value.

[0116] It should be noted that during training, there are two networks with the same structure, namely Figure 6B The network structure is recorded as the original network Q and the mirror network , only before each hour of data training, the parameters of the original network Q are copied to the mirror network , in addition, the mirror network No parameter updates are performed.

[0117] A quadruple (S i 、T i 、R i 、S i+1 ) constitutes a sample. The original network Q has two loss losses, denoted as loss1 and loss2, with a weight of 0.5 each. Both losses use the mean square error loss function. The labels and estimated values ​​of the two losses are:

[0118] For loss1, label is .in is the consumption time; γ is the attenuation coefficient, which is generally taken as 0.9 through experiments. The attenuation coefficient can ensure the stability of training; is the estimated consumption duration of subsequent recommendations, that is, the estimated duration of subsequent consumption after exiting the current video page; if the user in the training sample no longer has subsequent consumption behavior on the same day, then Directly replace it with 0. The estimated value is the total consumption time obtained by adding the sample consumption time and the target sample advantage value.

[0119] For loss2, label is , similar to loss1, to ensure that the advantage value determination module learns the increase or decrease in the subsequent average terminal duration after executing the corresponding action. This separates the value function and the advantage function, eliminates individual bias, and improves the accuracy and stability of the model. The estimated value is the sample consumption duration.

[0120] Figure 7 The module diagram of the resource recommendation device according to an embodiment of the present disclosure is schematically shown.

[0121] like Figure 7 As shown, the resource recommendation device 700 includes an acquisition module 710 , a determination module 720 and a resource recommendation module 730 .

[0122] The acquisition module 710 is configured to acquire recommended reference features for the target object in response to the target object selecting the target resource.

[0123] The determination module 720 is configured to determine a relevant resource recommendation parameter for the target object based on the recommendation reference feature, wherein the relevant resource recommendation parameter represents the proportion of recommended resources whose relevance to the target resource meets a preset condition in the resource recommendation result for the target object.

[0124] The resource recommendation module 730 is configured to perform resource recommendation operations based on the recommendation reference features and the updated recommendation strategy obtained by updating the preset recommendation strategy with the relevant resource recommendation parameters, and obtain resource recommendation results for the target object.

[0125] According to an embodiment of the present disclosure, the determination module 720 includes a first determination submodule and a second determination submodule.

[0126] The first determination submodule determines the number of target-related resources for the target object based on the recommendation reference features and the number of related resources determination model, wherein the number of target-related resources represents the number of recommended resources in the resource recommendation results for the target object whose relevance to the target resource meets preset conditions.

[0127] The second determination submodule is used to determine the relevant resource recommendation parameters according to the number of target relevant resources and the preset recommendation parameters.

[0128] According to an embodiment of the present disclosure, the related resource number determination model includes multiple advantage value determination modules, each of which is configured with a candidate related resource number. The first determination submodule includes a advantage value determination unit and a resource number determination unit.

[0129] The advantage value determination unit is used to input the recommended reference features into multiple advantage value determination modules respectively, and output the advantage values ​​of the multiple advantage value determination modules respectively, wherein the advantage value represents the marginal contribution of the number of candidate related resources to the consumption time of the target object.

[0130] The resource number determination unit is used to determine the target-related resource number according to the maximum advantage value among the multiple advantage values.

[0131] According to an embodiment of the present disclosure, the resource number determining unit includes a first determining subunit and a second determining subunit.

[0132] The first determining subunit is configured to determine a dominant value determining module corresponding to a maximum dominant value among the plurality of dominant values ​​as a target dominant value determining module.

[0133] The second determining subunit is configured to determine the number of candidate related resources of the target advantage value determining module as the target related resource number.

[0134] According to an embodiment of the present disclosure, the resource recommendation module 730 includes a recall submodule and a screening submodule.

[0135] The recall submodule is used to recall multiple candidate recommendation resources for the target object based on the recommendation reference features.

[0136] The screening submodule is used to screen multiple candidate recommended resources according to the updated recommendation strategy to obtain at least one target recommended resource.

[0137] According to an embodiment of the present disclosure, the update recommendation strategy includes relevant resource recommendation parameters and preset recommendation parameters. The screening submodule includes a dividing unit and a screening unit.

[0138] a dividing unit, configured to divide the plurality of candidate recommended resources into a first candidate recommended resource set and a second candidate recommended resource set based on the correlation between the candidate recommended resources and the target resource, wherein the first candidate recommended resource set includes candidate recommended resources whose correlation with the target resource meets a preset condition, and the second candidate recommended resource set includes candidate recommended resources whose correlation with the target resource does not meet the preset condition;

[0139] The screening unit is used to screen candidate recommended resources from the first candidate recommended resource set and the second candidate recommended resource set respectively under the constraints of the relevant resource recommendation parameters and the preset recommendation parameters to obtain at least one target recommended resource.

[0140] According to an embodiment of the present disclosure, the screening unit includes a first screening sub-unit and a second screening sub-unit.

[0141] The first screening subunit is configured to screen a first number of candidate recommended resources from the first candidate recommended resource set according to the first recommended number determined by the relevant resource recommendation parameter and the preset recommendation parameter.

[0142] The second screening subunit is configured to screen a second number of candidate recommended resources from the second candidate recommended resource set according to the first number of recommendations and a second number of recommendations determined by preset recommendation parameters.

[0143] The resource determination subunit is configured to determine a target recommended resource based on the first recommended number of candidate recommended resources and the second recommended number of candidate recommended resources.

[0144] According to an embodiment of the present disclosure, at least one target recommended resource in the resource recommendation result each has a resource weight; the resource recommendation device further includes an updating module.

[0145] The updating module is used to adjust the resource weight of the target recommended resources of the preset category according to the weight adjustment coefficient determined by the respective advantage values ​​of multiple advantage value determination modules when the resource recommendation results include the target recommended resources of the preset category, so as to obtain the updated resource recommendation results.

[0146] According to an embodiment of the present disclosure, the updating module further includes a weight determination submodule.

[0147] A weight adjustment coefficient is determined according to an average advantage value determined from the plurality of advantage values ​​and a maximum advantage value among the plurality of advantage values.

[0148] According to an embodiment of the present disclosure, the weight determination submodule includes a deviation determination unit and an adjustment unit.

[0149] The deviation determination unit is used to determine the deviation amplification factor according to the deviation degree between the average advantage value and the maximum advantage value.

[0150] The adjustment unit is used to determine the weight adjustment coefficient based on the basic control factor and the deviation amplification factor obtained by normalizing the maximum advantage value.

[0151] It should be noted that the training device part of the image recognition model in the embodiment of the present disclosure corresponds to the training method part of the image recognition model in the embodiment of the present disclosure. The description of the training device part of the image recognition model specifically refers to the training method part of the image recognition model and will not be repeated here.

[0152] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0153] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0154] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to enable a computer to execute the above method.

[0155] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the above method when executed by a processor.

[0156] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0157] like Figure 8 As shown, electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of electronic device 800 may also be stored in RAM 803. Computing unit 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.

[0158] Multiple components in electronic device 800 are connected to input / output (I / O) interface 805, including: an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0159] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the resource recommendation method and the large model training method. For example, in some embodiments, the resource recommendation method and the large model training method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the resource recommendation method and the large model training method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the resource recommendation method and the large model training method in any other appropriate manner (for example, by means of firmware).

[0160] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0161] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0162] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0163] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0164] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0165] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0166] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0167] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A resource recommendation method, comprising: In response to a target object selecting a target resource, obtaining recommended reference features for the target object; Determining, based on the recommendation reference features, a relevant resource recommendation parameter for the target object, wherein the relevant resource recommendation parameter represents the proportion of recommended resources whose relevance to the target resource meets a preset condition in the resource recommendation results for the target object; as well as A resource recommendation operation is performed based on the recommendation reference features and an updated recommendation strategy obtained by updating the preset recommendation strategy with the relevant resource recommendation parameters to obtain a resource recommendation result for the target object.

2. The method according to claim 1, wherein Determining the relevant resource recommendation parameters for the target object according to the recommendation reference features includes: Determining the number of target-related resources for the target object based on the recommendation reference features and the related resource number determination model, wherein the target-related resource number represents the number of recommended resources in the resource recommendation results for the target object whose relevance to the target resource meets a preset condition; Determine the related resource recommendation parameters according to the number of target related resources and preset recommendation parameters.

3. The method according to claim 2, wherein: The related resource number determination model includes a plurality of advantage value determination modules, each of which is configured with a candidate number of related resources; Determining a model based on the recommended reference features and the number of related resources to determine the target number of related resources for the target object includes: Inputting the recommendation reference features into the plurality of advantage value determination modules respectively, and outputting advantage values ​​of the plurality of advantage value determination modules respectively, wherein the advantage values ​​represent the marginal contribution of the number of candidate related resources to the consumption time of the target object; The number of target-related resources is determined according to a maximum advantage value among the plurality of advantage values.

4. The method according to claim 3, wherein: The determining the target-related number according to the maximum advantage value among the multiple advantage values ​​includes: determining a dominant value determination module corresponding to a maximum dominant value among the plurality of dominant values ​​as a target dominant value determination module; The number of candidate related resources of the target advantage value determination module is determined as the target related resource number.

5. The method according to claim 1, wherein The resource recommendation result includes at least one target recommended resource; The performing of the resource recommendation operation according to the recommendation reference feature and the updated recommendation strategy obtained by updating the preset recommendation strategy with the relevant resource recommendation parameters includes: Recalling a plurality of candidate recommendation resources for the target object according to the recommendation reference features; According to the update recommendation strategy, the multiple candidate recommendation resources are screened to obtain the at least one target recommendation resource.

6. The method according to claim 5, wherein: The update recommendation strategy includes relevant resource recommendation parameters and preset recommendation parameters; The filtering of the plurality of candidate recommended resources according to the updated recommendation strategy to obtain the at least one target recommended resource includes: Dividing the plurality of candidate recommended resources into a first candidate recommended resource set and a second candidate recommended resource set according to the correlation between the candidate recommended resources and the target resource, wherein the first candidate recommended resource set includes candidate recommended resources whose correlation with the target resource meets a preset condition, and the second candidate recommended resource set includes candidate recommended resources whose correlation with the target resource does not meet the preset condition; Under the constraints of the relevant resource recommendation parameters and the preset recommendation parameters, candidate recommended resources are screened from the first candidate recommended resource set and the second candidate recommended resource set respectively to obtain the at least one target recommended resource.

7. The method according to claim 6, wherein: The step of filtering candidate recommended resources from the first candidate recommended resource set and the second candidate recommended resource set respectively under the constraints of the relevant resource recommendation parameters and the preset recommendation parameters to obtain the at least one target recommended resource includes: Screening a first number of candidate recommended resources from the first candidate recommended resource set according to the first number of recommendations determined by the relevant resource recommendation parameter and the preset recommendation parameter; screening a second number of candidate recommended resources from the second candidate recommended resource set according to the first number of recommended resources and a second number of recommended resources determined by preset recommendation parameters; The target recommended resource is determined according to the first recommended number of candidate recommended resources and the second recommended number of candidate recommended resources.

8. The method according to claim 3, wherein: At least one target recommended resource in the resource recommendation result each has a resource weight; The method further comprises: In the case where the resource recommendation result includes target recommended resources of a preset category, the resource weights of the target recommended resources of the preset category are adjusted according to the weight adjustment coefficients determined by the respective advantage values ​​of multiple advantage value determination modules to obtain updated resource recommendation results.

9. The method according to claim 8, further comprising: The weight adjustment coefficient is determined according to an average advantage value determined from the plurality of advantage values ​​and a maximum advantage value among the plurality of advantage values.

10. The method according to claim 9, wherein: Determining the weight adjustment coefficient according to the average advantage value determined by the plurality of advantage values ​​and the maximum advantage value among the plurality of advantage values ​​includes: Determining a deviation magnification factor according to a degree of deviation between the average advantage value and the maximum advantage value; The weight adjustment coefficient is determined according to the basic control factor obtained by normalizing the maximum advantage value and the deviation amplification factor.

11. A method for training a model, wherein: The model includes a consumption duration determination module and a plurality of advantage value determination modules; The method comprises: Obtaining a training sample, wherein the training sample includes a sample recommendation reference feature, a number of sample-related resources, and a sample label; Inputting the training sample into the consumption duration determination module and the plurality of advantage value determination modules to obtain sample consumption duration and sample advantage value; Determining a total sample consumption time according to the sample consumption time and a target sample advantage value among the multiple sample advantage values, wherein the target sample advantage value is determined according to the number of resources related to the sample; According to the difference between the total sample consumption time and the sample label, the network parameters of the relevant resource number determination model are adjusted to obtain a trained relevant resource number determination model.

12. The method according to claim 11, further comprising: Determining a first loss value according to the sample consumption time and the sample label; The adjusting of the network parameters of the model for determining the number of related resources according to the difference between the total sample consumption time and the sample label includes: Determining a second loss value based on the total sample consumption time and the sample label; According to the total loss value determined by the first loss value and the second loss value, network parameters of the model for determining the number of related resources are adjusted.

13. A resource recommendation device, comprising: an acquisition module, configured to acquire, in response to a target object selecting a target resource, a recommended reference feature for the target object; a determination module, configured to determine, based on the recommendation reference features, a relevant resource recommendation parameter for the target object, wherein the relevant resource recommendation parameter represents a proportion of recommended resources whose relevance to the target resource meets a preset condition in the resource recommendation results for the target object; as well as The resource recommendation module is used to perform resource recommendation operations based on the recommendation reference features and the updated recommendation strategy obtained after the preset recommendation strategy is updated by the relevant resource recommendation parameters, and obtain resource recommendation results for the target object.

14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12. 16 . A computer program product, comprising a computer program, wherein the computer program is stored on at least one of a readable storage medium and an electronic device, and when the computer program is executed by a processor, implements the method according to claim 1 .

Citation Information

Patent Citations

  • Resource recommendation method and device

    CN111652674A

  • Model training method and device, resource recommendation method and device, electronic equipment and storage medium

    CN114564644A

  • Processing method, processing device and electronic equipment

    CN115344788A

  • Associated word recommendation method and device, electronic equipment and storage medium

    CN116975428A

  • Recommended resource screening strategy optimization method, resource screening method and related device

    CN118395001A