Resource recommendation method, method and device for training deep learning model, and medium
By integrating the interactive behavior characteristics of the target objects and using deep learning and reinforcement learning to train the model, the problem of inaccurate recommendations in the resource recommendation system is solved, achieving higher matching degree and user experience.
Patent Information
- Application Number
- CN202510885376.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-23
AI Technical Summary
In the existing technology, resource recommendation systems cannot accurately match users' actual needs and interests, resulting in a low consumption probability of recommended resources and difficulty in meeting users' actual needs.
By receiving the target object's interactive behavior characteristics during a specified period after performing a resource update operation, the model is trained using a deep learning model and reinforcement learning mechanism to determine the target resources and make recommendations.
It improves the accuracy of resource recommendations, enhances the user interaction experience, and meets the actual needs and changes in users' interests.
Smart Images

Figure CN120687680A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to technical fields such as resource recommendation, intelligent search, and big data. Background Art
[0002] With the rapid development of artificial intelligence technology, users can easily browse news, videos and other resource information through terminal devices such as smartphones. Recommending resource information to users based on their needs or preferences can make it easier for users to obtain resources of interest. Summary of the Invention
[0003] The present disclosure provides a resource recommendation method, a method for training a deep learning model, an apparatus, an electronic device, a storage medium, and a program product.
[0004] According to one aspect of the present disclosure, a resource recommendation method is provided, comprising: receiving interactive behavior characteristics of a target object for a specified resource in a specified time period after performing a resource update operation, wherein the resource update operation is used to update a current page resource to obtain the specified resource; fusing the interactive behavior characteristics associated with each of a plurality of specified time periods to obtain a target fusion characteristic; and determining a target resource based on the target fusion characteristic, and recommending the target resource to the target object.
[0005] According to another aspect of the present disclosure, a method for training a deep learning model is provided, including: receiving sample status, sample action and sample feedback information, the sample status including sample interaction behavior characteristics of the sample object for the sample specified resource in the sample specified time period after performing the sample resource update operation, the sample resource update operation being used to update the current sample page resource to obtain the sample specified resource, the sample action being a sample interaction intention index related to the sample candidate resource obtained by processing the sample status information through the deep learning model, the sample interaction intention index being used to determine the sample target resource from the sample candidate resource, and the sample feedback information representing the satisfaction of the sample object with the sample target resource; based on the reinforcement learning mechanism, the deep learning model is trained using the sample status, sample action and sample feedback information to obtain a trained deep learning model.
[0006] According to another aspect of the present disclosure, a resource recommendation device is provided, including: a first receiving module, used to receive the interactive behavior characteristics of a target object for a specified resource in a specified time period after performing a resource update operation, wherein the resource update operation is used to update the current page resource to obtain the specified resource; a fusion module, used to fuse the interactive behavior characteristics related to each of a plurality of specified time periods to obtain a target fusion characteristic; and a recommendation module, used to determine a target resource based on the target fusion characteristic and recommend the target resource to the target object.
[0007] According to another aspect of the present disclosure, a device for training a deep learning model is provided, including: a second receiving module for receiving sample status, sample action and sample feedback information, the sample status including the sample interaction behavior characteristics of the sample object for the sample specified resource in the sample specified time period after performing the sample resource update operation, the sample resource update operation being used to update the current sample page resource to obtain the sample specified resource, the sample action being a sample interaction intention index related to the sample candidate resource obtained by processing the sample status information through the deep learning model, the sample interaction intention index being used to determine the sample target resource from the sample candidate resource, and the sample feedback information representing the satisfaction of the sample object with the sample target resource; and a training module for training the deep learning model based on the reinforcement learning mechanism using the sample status, sample action and sample feedback information to obtain a trained deep learning model.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to an embodiment of the present disclosure.
[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method provided according to an embodiment of the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method provided according to the embodiment of the present disclosure when executed by a processor.
[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.
[0013] Figure 1 Schematically illustrates an exemplary system architecture to which the resource recommendation method and apparatus according to an embodiment of the present disclosure can be applied;
[0014] Figure 2 The following schematically shows a flow chart of a resource recommendation method according to an embodiment of the present disclosure;
[0015] Figure 3 Schematically illustrates a schematic diagram of the principle of a deep learning model according to an embodiment of the present disclosure;
[0016] Figure 4 Schematically shows a flow chart of a method for training a deep learning model according to an embodiment of the present disclosure;
[0017] Figure 5 The following schematically illustrates a method for training a deep learning model according to an embodiment of the present disclosure;
[0018] Figure 6 A block diagram of a resource recommendation device according to an embodiment of the present disclosure is schematically shown;
[0019] Figure 7 A block diagram schematically illustrates an apparatus for training a deep learning model according to an embodiment of the present disclosure; and
[0020] Figure 8 A schematic block diagram of an example electronic device that can be used to implement the resource recommendation method and the method for training a deep learning model according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0021] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0023] The inventors found that when recommending resources such as videos and news to users, the recommended resources have a low degree of match with the users' actual needs or interests, resulting in a low probability of relevant users consuming the recommended resources, making it difficult to meet the users' actual needs.
[0024] Embodiments of the present disclosure provide a resource recommendation method, a method for training a deep learning model, an apparatus, an electronic device, a storage medium, and a program product. The source recommendation method includes: receiving interactive behavior features of a target object for a specified resource during a specified period after performing a resource update operation, wherein the resource update operation is used to update the current page resource to obtain the specified resource; fusing the interactive behavior features associated with each of the multiple specified periods to obtain a target fusion feature; and determining a target resource based on the target fusion feature and recommending the target resource to the target object.
[0025] According to an embodiment of the present disclosure, by obtaining the interactive behavior characteristics of the target object for the specified resource in a specified time period, the interactive behavior characteristics can more accurately represent the interactive behavior preferences of the target object for the specified resource obtained after performing a resource update operation on the current page resource. By fusing the interactive behavior characteristics that characterize the interactive behavior preferences of multiple specified time periods, the target fusion characteristics can represent the changes in the preferences of the target object for the specified resource obtained after multiple resource update operations, and then the target object can learn the resource preference attributes and interactive behavior changes after performing a resource update operation due to dissatisfaction with the resource content or looking for new resources of interest based on the target fusion characteristics, thereby realizing interactive intention mining of the target object's interactive preferences. Determining the target resource based on the target fusion characteristics and recommending the target resource to it can enable the target resource to more accurately meet the target object's changing pattern of resource preferences, improve the matching degree between the target resource and the actual needs of the target object, and thus improve the accuracy of resource recommendation and the user interaction experience.
[0026] Figure 1 An exemplary system architecture to which the resource recommendation method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.
[0027] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure. This does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the resource recommendation method and apparatus may be applied may include a terminal device, but the terminal device may implement the resource recommendation method and apparatus provided by the embodiments of the present disclosure without interacting with a server.
[0028] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0029] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0030] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0031] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0032] It should be noted that the resource recommendation method provided in the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the resource recommendation apparatus provided in the embodiments of the present disclosure can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0033] Alternatively, the resource recommendation method provided in the embodiment of the present disclosure may also be generally executed by the server 105. Accordingly, the resource recommendation apparatus provided in the embodiment of the present disclosure may generally be provided in the server 105. The resource recommendation method provided in the embodiment of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and that can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the resource recommendation apparatus provided in the embodiment of the present disclosure may also be provided in a server or server cluster that is different from the server 105 and that can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0034] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0035] Figure 2 The flowchart of the resource recommendation method according to an embodiment of the present disclosure is schematically shown.
[0036] like Figure 2 As shown, the resource recommendation method includes operations S210 to S230.
[0037] In operation S210 , interaction behavior characteristics of a target object with respect to a specified resource in a specified period of time after performing a resource update operation are received.
[0038] In operation S220 , the interaction behavior features associated with each of the plurality of designated time periods are fused to obtain a target fused feature.
[0039] In operation S230 , a target resource is determined based on the target fusion feature, and the target resource is recommended to the target object.
[0040] According to an embodiment of the present disclosure, a resource update operation is used to update the current page resources to obtain a specified resource. For example, after performing a resource update operation on the product resources in the current product recommendation page, the product recommendation page can display the updated product resources as the specified interactive resource. The specified time period after the resource update operation can be the time period when the target object browses the specified resource on the page. The target object can perform any type of interactive operation such as resource browsing, comment input, etc. on the specified resource obtained after the update. The interactive behavior characteristics can represent the interactive behavior attributes related to the interactive operation of the target object, for example, the interactive behavior characteristics can represent the browsing time, the number of like operations, etc. Alternatively, the interactive behavior characteristics can also represent the target object's preference for the specified resource displayed on the page, for example, the interactive behavior characteristics can also represent the resource theme type, resource content information, etc. of the specified resource on which the target object performs the browsing operation. The embodiment of the present disclosure does not limit the specific data attributes represented by the interactive behavior characteristics, as long as it is related to the specified resource or the target object's interactive behavior with respect to the specified resource.
[0041] In some embodiments, the interactive behavior feature includes at least one of the following: a browsing time feature, a comment content feature, and a resource content feature of a designated interactive resource.
[0042] In some embodiments, the browsing duration feature may represent the duration of the target object's browsing of one or more specified resources during a specified period. For example, the browsing duration feature may represent the target object's viewing duration of a video resource, the video completion rate, etc. during a specified period.
[0043] In some embodiments, a comment content feature may represent attribute information related to the comment information, such as the text content and sentiment type attributes of the comment information entered by the target subject for a specified resource during a specified period. For example, a comment content feature may be a text content feature encoding of the comment information "This TV series is really good" entered by the target subject, and a sentiment attribute encoding indicating the type "like".
[0044] In some embodiments, the resource content features of a specified interactive resource may include resource content features of the specified resource on which the target object has performed a target interactive operation, such as a browse operation or a like operation, within a specified time period. The resource content features may include resource theme semantic features, resource classification features, resource body semantic features, and the like of the specified resource.
[0045] In some embodiments, the specified time period may be a time period between multiple resource update operations performed by the target object, for example, a time period between two consecutive resource update operations performed by the target object on a page.
[0046] In some embodiments, the designated resource browsed by the target object during the designated period can be used as the current page resource. After the target object performs a resource update operation on the current page, the current designated resource can be updated to obtain an updated designated resource. The updated designated period can be the time period for displaying the updated designated resource.
[0047] For example, if the current page displays the first set of product resources, the target object performs a first resource update operation. Based on the first resource update operation, the current page displays the updated second set of product resources as designated resources. If the target object performs a second resource update operation on the page displaying the first set of product resources, the second set of product resources is updated as the page resources in response to the resource update operation, and the page can display a third set of product resources as designated resources.
[0048] In some embodiments, fusing the interaction behavior features related to multiple specified time periods may include fusing multiple interaction behavior features based on an attention mechanism to capture the interaction behavior preferences and resource content feature preferences for the updated specified resources in multiple specified time periods after the target object performs multiple resource update operations. This allows the target fusion features to more fully represent the changes in resource preference attributes and interaction behavior for the updated resources after the target object performs a resource update operation due to dissatisfaction with the resource content or looking for new resources of interest, so that the target resources determined based on the target fusion features can match the target object's interest changes and actual needs, thereby improving the accuracy of the target resources recommended to the target object.
[0049] In some embodiments, fusing the interaction behavior features related to multiple specified time periods to obtain a target fusion feature may include: fusing the multiple interaction behavior features related to the specified time period to obtain an initial fusion feature; and fusing the initial fusion features related to multiple execution time periods based on an attention mechanism to obtain a target fusion feature.
[0050] According to an embodiment of the present disclosure, the target object can perform multiple interactive operations in a specified time period, and multiple interactive behavior features related to the specified time period can correspond to the multiple interactive operations. For example, after the target object performs a resource update operation on the product recommendation page, it can perform multiple interactive operations of any operation type such as browsing, product consultation, collection, and purchase on the multiple product resources obtained after the update. The multiple interactive behavior features can be determined based on information related to the interactive operations, such as the interactive operation type, resource content, input information content, purchase behavior attributes, etc. corresponding to the multiple interactive operations.
[0051] In one embodiment, feature fusion can be performed on multiple interactive behavior features related to a specified time period based on a fusion function. The fusion function may include, for example, an accumulation function, a weighted average function, etc. The embodiment of the present disclosure does not limit the specific type of the fusion function.
[0052] In one embodiment, multiple interactive behavior features associated with a specified time period may be processed based on a neural network algorithm to obtain an initial fusion feature. Neural network algorithms may include, for example, long short-term memory algorithms, attention network algorithms, etc., which are not limited in the embodiments of the present disclosure.
[0053] By fusing multiple interactive behavior features in a specified time period, the initial fused features can be used to represent the target object's interaction preference for the specified resource obtained after performing a resource update operation. This can then be used to represent the target object's interaction preference change trend for the displayed page resource after performing multiple resource update operations based on the initial fused features corresponding to each of the multiple specified time periods.
[0054] In some embodiments, based on the fusion of initial fusion features related to multiple execution time periods based on an attention mechanism, it can include processing multiple initial fusion features based on an attention network algorithm such as Transformer, so as to capture the temporal relationship between multiple initial fusion features based on the attention mechanism, so that the target fusion feature can learn the target object's interaction behavior preference for the updated specified resource after performing resource update operations multiple times based on actual demand intentions such as resource exploration needs or dissatisfaction with the displayed resources, so that the target fusion feature can more accurately represent the target object's interest change trend and resource demand changes, so that the target resource determined based on the target fusion feature can better meet the actual needs of the target object.
[0055] In some embodiments, determining the target resource based on the target fusion feature may include: performing interaction intent detection on candidate resources based on the target fusion feature to obtain interaction intent indicators related to the candidate resources; and determining the target resource from multiple candidate resources based on the interaction intent indicators.
[0056] According to an embodiment of the present disclosure, detecting interaction intent for candidate resources based on target fusion features may include processing the target fusion features using any type of algorithm, such as a deep learning algorithm or a normalization algorithm, to obtain interaction intent indicators related to the candidate resources. For example, the target fusion features may be processed using a multi-layer perceptron algorithm to obtain interaction intent indicators corresponding to each of the multiple candidate resources.
[0057] In some embodiments, the interaction intent indicator can represent the target object's desire to perform an interactive operation on the candidate resource. For example, the interaction intent indicator can be a click-through rate indicator or a completion rate indicator for a video resource, a conversion rate indicator for a product resource, or any other type of weight or score that can represent the desire to perform an interactive operation on the candidate resource.
[0058] In some embodiments, determining a target resource from multiple candidate resources based on an interaction intention indicator may include sorting the multiple candidate resources according to the interaction intention indicator, and determining a preset number of target resources from the multiple sequentially arranged candidate resources based on the sorting position.
[0059] In some embodiments, determining a target resource from a plurality of candidate resources based on the interaction intent indicator may further include selecting, based on the interaction intent indicator, a candidate resource that satisfies an interaction intent indicator threshold as the target resource. For example, a candidate resource having an interaction intent score greater than or equal to a preset score threshold of 0.5 may be determined as the target resource.
[0060] It should be noted that the embodiments of the present disclosure do not limit the specific method of determining the target resource from the candidate resources based on the interaction intention indicator, as long as it can be based on the interaction intention indicator corresponding to the candidate resource.
[0061] In one embodiment, the resource identifier of the target resource determined from the candidate resources can be set in the resource recommendation list for the target object. The resource identifier of the target resource in the resource recommendation list can be determined based on the following method.
[0062] In response to the resource update operation of the target object being triggered, multiple candidate resources are recalled. Three screening stages are performed for the multiple candidate resources, namely, coarse ranking and scoring, fine ranking and scoring, and multi-model hybrid sorting. In the coarse ranking stage, multiple candidate resources are sorted based on the click-through rate of the candidate resources output by the general resource recommendation model to obtain the first stage sorting result. In the fine ranking and scoring stage, the preference features for the target object are input into another resource recommendation model with a larger parameter scale, and the click-through rate of the fine ranking and scoring stage is output. Based on the click-through rate of the fine ranking and scoring stage, the candidate resources in the first stage sorting result are re-sorted and screened to obtain the candidate resources screened in the second stage. In the multi-model hybrid sorting stage, the interactive behavior characteristics of the target object in multiple specified time periods are processed based on the method provided in the embodiment of the present disclosure to obtain the interactive intention index corresponding to each candidate resource screened in the second stage. The interactive intention index is used as a weight parameter to perform weighted summation on the click-through rate of each candidate resource screened in the second stage to obtain the third stage click-through rate corresponding to each candidate resource. The candidate resources screened in the second stage are screened based on the click-through rate in the third stage, and the top N candidate resources are obtained as target resources, where N is an arbitrary positive integer.
[0063] Figure 3 The schematic diagram schematically shows the principle of the deep learning model according to an embodiment of the present disclosure.
[0064] like Figure 3 As shown, the deep learning model may include a first fusion layer, a second fusion layer, and an output layer. The first fusion layer and the output layer may be constructed based on a multilayer perceptron algorithm, and the second fusion layer may be constructed based on an attention network algorithm.
[0065] The multiple interactive behavior features 310 corresponding to the multiple specified time periods in the sampling phase may include the first interactive behavior feature, the second interactive behavior feature... and the nth interactive behavior feature, where n is an integer greater than 2. The first fusion layer is used to perform feature fusion on the multiple interactive behavior features 310 to obtain initial fusion features corresponding to the multiple specified time periods. The second fusion layer is used to fuse the initial fusion features related to the multiple specified time periods based on the attention mechanism to obtain the target fusion feature. The output layer is used to process the target fusion feature to obtain the interactive intention index 320 related to the candidate resource. The interactive intention index 320 may include any type of recommendation index such as click-through rate index, browsing time index, etc. The target resource for recommendation to the target object is determined from the multiple candidate resources through the interactive intention index 320.
[0066] Based on the resource recommendation algorithm provided in the above embodiments, the embodiments of the present disclosure also provide a method for training a deep learning model. The method for training a deep learning model provided in the embodiments of the present disclosure will be described in detail below in combination with specific embodiments and drawings.
[0067] Figure 4 A flowchart of a method for training a deep learning model according to an embodiment of the present disclosure is schematically shown.
[0068] like Figure 4 As shown, the method for training a deep learning model includes operations S410 to S420.
[0069] In operation S410 , sample status, sample action, and sample feedback information are received.
[0070] In operation S420, based on a reinforcement learning mechanism, a deep learning model is trained using sample states, sample actions, and sample feedback information to obtain a trained deep learning model.
[0071] According to an embodiment of the present disclosure, the sample state includes the sample interaction behavior characteristics of the sample object with respect to the sample-specified resource in the sample-specified period after the sample resource update operation. The sample resource update operation is used to update the current sample page resource to obtain the sample-specified resource.
[0072] According to an embodiment of the present disclosure, the sample action is to process the sample status information through a deep learning model to obtain a sample interaction intention index related to the sample candidate resources. The sample interaction intention index is used to determine the sample target resource from the sample candidate resources, and the sample feedback information represents the satisfaction of the sample object with the sample target resource.
[0073] In some embodiments, the sample feedback information r is used as reward information (reward) for training the deep learning model through the reinforcement learning mechanism. The sample state S, sample action a and sample feedback information r can be processed based on the reinforcement learning mechanism to determine the policy gradient information for the deep learning model, and the model parameters of the deep learning model can be adjusted based on the policy gradient information to obtain a trained deep learning model.
[0074] In some embodiments, the sample state may include multiple sample states, which may be multiple sample interaction behavior features related to multiple sampling phases. For example, the multiple sample states may be represented as the i-th sample state S i and the i+1th sample state S i+1 , the i-th sample state S i and the i+1th sample state S i+1 They correspond to the i-th sampling phase i and the i+1-th sampling phase i+1, respectively. It should be noted that any sampling phase may include multiple sample designated time periods. i is any positive integer.
[0075] In one embodiment, the current state, sample actions, and sample feedback information used to train the deep learning model can be represented as a sample array (S i , a i, r i , S i+1 ), where a i is the sample action of the i-th sampling stage, r i is the sample feedback information of the i-th sampling stage, the sample array (S i , a i , r i , S i+1 ) can be determined based on the actual interactive operation behavior of the sample object.
[0076] In some embodiments, a model structure for a deep learning model can be constructed based on a reinforcement learning mechanism. For example, a policy model (Actor Model) can be constructed based on the deep learning model, and a criticism model (Critic Model) can be constructed to process the predicted action information and sample feedback information output by the policy model to obtain a Q value used to evaluate the resource recommendation ability of the deep learning model. By performing gradient calculation based on the Q value, policy gradient information for the deep learning model is obtained.
[0077] It should be noted that the technical terms involved in the method for training a deep learning model in the embodiment of the present disclosure, including but not limited to sample resource update operations, sample specified time periods, sample interactive behavior characteristics, etc., have the same or similar attributes as the technical terms involved in the resource recommendation method provided by the embodiment of the present disclosure, including but not limited to resource update operations, specified time period interactive behavior characteristics, etc., and the embodiments of the present disclosure will not be repeated here.
[0078] In some embodiments, the sample feedback information includes at least one of first sample feedback information and second sample feedback information.
[0079] In some embodiments, the first sample feedback information may be determined based on a sample target interaction behavior characteristic of a sample object with respect to a sample target resource. The sample target interaction behavior characteristic may indicate target interaction operations, such as browsing time and click operations, performed by the target object with respect to the target resource. For example, the first sample feedback information may be obtained by processing the sample object's interaction behavior characteristic based on a preset algorithm, such as a preset normalization algorithm or a neural network algorithm.
[0080] In one embodiment, the sample target interaction behavior feature can represent browsing time indicators for the sample target resource in multiple consecutive specified time periods during the sampling phase. The browsing time indicators for the multiple specified time periods can be processed based on a multi-layer perceptron algorithm to obtain first sample feedback information based on numerical representation. The first sample feedback information can thus represent the sample object's preference for the sample target resource after the sample resource operation is updated, thereby achieving timely feedback on the sample object's preference for the sample resource based on the first sample feedback information.
[0081] In some embodiments, the second sample feedback information may be determined based on a target interval duration, which represents a time period between different page view events, where a page view event represents a sample object opening a target page for displaying a sample target resource for viewing.
[0082] For example, the target page is a product display page on an e-commerce platform, and the sample target resources displayed on the product display page are product resources to be selected. During the sampling phase, the sample object opens the product display page for the first time to browse the recommended product resources, causing the first page browsing event to be triggered. When the sample object returns the product display page opened for the first time to run the next day, or closes the product display page opened for the first time, the first page browsing event ends. One hour later, when the sample object opens the product display page for the second time to browse product resources, it can be determined that the second page browsing event is triggered. The time period of one hour between the first page browsing event and the second page browsing event can be determined as the target interval duration.
[0083] In one embodiment, the interactive behavior characteristics of the sample objects can also be processed through a trained binary classification model to output an estimated target interval duration, and the second sample feedback information can be obtained by normalizing the target interval duration.
[0084] The second sample feedback information determined based on the target interval duration can indicate whether the sample target resources determined based on the deep learning model can enable the sample object to open the target page to browse the sample target resources in a longer period of time. In this way, the target interval duration can be used to indicate the degree of dependence and demand of the target object on the sample target resources recommended by the deep learning model in a longer period of time, so that the second sample feedback information can indicate the long-term dependence of the sample object on the sample target resources in the target page. Therefore, the deep learning model can be trained by processing the second sample feedback information based on the reinforcement learning mechanism, so that the trained deep learning model can improve the retention and dependence of the sample object by outputting target resources that are more consistent with the demand intention and interest change direction of the sample object, thereby improving the interactive experience.
[0085] In some embodiments, the second sample feedback information can be determined based on multiple target interval durations. For example, the second sample feedback information can be obtained by processing the average duration, variance duration, etc. of multiple target interval durations based on a multilayer perceptron algorithm.
[0086] In one embodiment, the first sample feedback information and the second sample feedback information may be based on numerical representation.
[0087] In some embodiments, the first sample feedback information and the second sample feedback information may be output based on different evaluation models, so as to adjust parameters of the deep learning model serving as the policy model through different evaluation models to improve training efficiency.
[0088] According to an embodiment of the present disclosure, based on the reinforcement learning mechanism, training a deep learning model using sample states, sample actions and sample feedback information may include: using a first evaluation model to process sample actions, first sample feedback information and sample states to obtain first target feedback information; using a second evaluation model to process sample actions, second sample feedback information and sample states to obtain second target feedback information; performing gradient calculation based on the first target feedback information and the second target feedback information to obtain policy gradient information for the deep learning model; and updating model parameters of the deep learning model based on the policy gradient information.
[0089] In some embodiments, the first evaluation model processes the sample action, the first sample feedback information, and the sample state to obtain the first target feedback information, which may include processing the first sample array (S i , a i , r i_1 , S i+1 ), where r i_1 Represents the first sample feedback information of the i-th sampling stage. The first target feedback information may include the first target value yi_1 output by the first target evaluation network and the first Q value Qi_1 output by the first online evaluation network. The first evaluation model includes the first online evaluation network and the first target evaluation network.
[0090] For example, the first target value yi_1 may be determined based on the following formula (1), and the first Q value Qi_1 may be determined based on the following formula (2).
[0091] y i_1=r i_1 +γQ′_1(s i+1 ,μ′(s i+1 ∣θ μ′ )|θ Q′_1 ) (1);
[0092] Qi_1=Q_1(s i ,a i ∣θ Q_1 ) (2);
[0093] Among them, Q′_1() is the first target evaluation network, Q_1() is the first online evaluation network, θ μ′ is the model parameter of the target policy model corresponding to the deep learning model, θ Q is the model parameter of the first online evaluation network, and γ is the discount coefficient.
[0094] In some embodiments, the second evaluation model processes the sample action, the second sample feedback information, and the sample state to obtain the second target feedback information, which may include processing the second sample array (S) using the second evaluation model. i , a i , r i_2 , S i+1 ), where r i_2 The second target feedback information may include the second target value yi_2 output by the second target evaluation network and the second Q value Qi_2 output by the second online evaluation network. The second evaluation model includes the second online evaluation network and the second target evaluation network.
[0095] For example, the second target value yi_2 can be determined based on the following formula (3), and the second Q value Qi_2 can be determined based on the following formula (4).
[0096] y i_2=r i_2 +γQ′_2(s i+1 ,μ′(s i+1 ∣θ μ′ )|θ Q′_2 ) (3);
[0097] Qi_2=Q_2(s i ,a i ∣θ Q_2 ) (4);
[0098] Among them, Q′_2() is the second target evaluation network, Q_2() is the second online evaluation network, θ μ′ is the model parameter of the target policy model corresponding to the deep learning model, θ Q is the model parameter of the second online evaluation network, and γ is the discount coefficient.
[0099] In some embodiments, the first target feedback information and the second target feedback information can be fused to obtain fused feedback information, the fused feedback information can be calculated based on the gradient algorithm to obtain the policy gradient information for the deep learning model, and the model parameters of the deep learning model can be updated based on the policy gradient information to obtain a trained deep learning model.
[0100] In some embodiments, performing gradient calculation based on the first target feedback information and the second target feedback information to obtain policy gradient information for the deep learning model may also include: using a loss function to process the first target feedback information and the second target feedback information respectively to obtain first loss data and second loss data; determining the policy gradient information based on the fused loss data determined by fusing the first loss data and the second loss data.
[0101] In some embodiments, the first target value yi_1 and the first Q value Qi_1 in the first target feedback information may be processed using a loss function to obtain a first loss value as the first loss data.
[0102] For example, the first loss value can be calculated based on the following formula (5). Q_1 ) represents the first loss value.
[0103] L(θ Q_1 ) = ∑ i (yi_1−Q_1 (si,ai|θ Q_1 )) 2 / N (5);
[0104] In some embodiments, the second target value yi_2 and the second Q value Qi_2 in the second target feedback information can be processed by using a loss function to obtain a second loss value as the second loss data.
[0105] For example, the second loss value can be calculated based on the following formula (6). Q_2 ) represents the second loss value.
[0106] L(θ Q_2 ) = ∑ i (yi_2−Q_2 (si,ai|θ Q_2 )) 2 / N (6);
[0107] In some embodiments, the first loss value and the second loss value can be fused based on a weighted fusion algorithm, and the fused loss value can be determined as fused loss data. Based on the fused loss data, the model parameters of the first and second online evaluation networks are adjusted to obtain trained first and second online evaluation networks. The trained first and second online evaluation networks are used to process the first and second sample arrays, respectively, to obtain first and second gradient information. Policy gradient information is determined by fusing the first and second gradient information.
[0108] Figure 5 The schematic diagram schematically shows the principle of the method for training a deep learning model according to an embodiment of the present disclosure.
[0109] like Figure 5As shown, the deep learning model in the strategy module M510 is used to interact with the sample objects in the interactive environment 501 to recommend sample target resources to the sample objects based on the sample interaction behavior characteristics in the sample specified time period. The strategy module M510 can obtain sample status and sample feedback information from the interactive environment 501, and store the sample status, sample feedback information, and sample action in the experience pool to facilitate training the deep learning model. The target strategy model in the strategy module M510 is used to transmit the processing result μ′(s ′) for the sample status to the first evaluation model M520 and the second evaluation model M530. i+1 ∣θ μ′ ).
[0110] The first evaluation model M520 and the second evaluation model M530 can each obtain a first sample array (S i , a i , r i_1 , S i+1 ) and the second sample array (S i , a i , r i_2 , S i+1 The first online evaluation network and the first target evaluation network of the first evaluation model process the first sample array (S i , a i , r i_1 , S i+1 ) determines the first target value yi_1 and the first Q value Qi_1 as the first target feedback information. The second online evaluation network and the second target evaluation network of the second evaluation model process the second sample array (S i , a i , r i_2 , S i+1 ) Determine the second target value yi_2 and the second Q value Qi_2 as the second target feedback information.
[0111] The first evaluation model M520 determines first loss data by processing the first target feedback information, and the second evaluation model M530 determines second loss data by processing the second target feedback information. A fused loss data is determined by fusing the first and second loss data, and policy gradient information for the deep learning model is determined based on the fused loss data. The policy gradient information is transmitted to the deep learning model, where it is used to update the model parameters of the deep learning model, resulting in a trained deep learning model.
[0112] The method for training a deep learning model provided by the embodiments of the present disclosure adopts the reinforcement learning mechanism principle of DDPG (Deep Deterministic Policy Gradient) and designs different evaluation models for first sample feedback information and second sample feedback information that respectively represent immediate feedback and long-term feedback of sample objects to respectively learn the immediate feedback requirements and long-term feedback requirements of sample objects. This method can simultaneously optimize multiple feedback information goals and improve the overall performance of the deep learning model in the accuracy of target resource detection.
[0113] Figure 6 The block diagram of the resource recommendation device according to an embodiment of the present disclosure is schematically shown.
[0114] like Figure 6 As shown, the resource recommendation device 600 includes: a first receiving module 610 , a fusion module 620 and a recommendation module 630 .
[0115] The first receiving module 610 is configured to receive interactive behavior characteristics of a target object with respect to a specified resource within a specified period of time after executing a resource update operation, wherein the resource update operation is used to update a current page resource to obtain the specified resource.
[0116] The fusion module 620 is configured to fuse the interaction behavior features associated with each of the multiple specified time periods to obtain a target fusion feature.
[0117] The recommendation module 630 is configured to determine a target resource based on the target fusion feature and recommend the target resource to the target object.
[0118] According to an embodiment of the present disclosure, the fusion module includes: a first fusion unit and a second fusion unit.
[0119] The first fusion unit is used to perform feature fusion on multiple interactive behavior features related to a specified time period to obtain an initial fusion feature.
[0120] The second fusion unit is used to fuse the initial fusion features related to the multiple execution periods based on the attention mechanism to obtain the target fusion features.
[0121] According to an embodiment of the present disclosure, the recommendation module includes: a detection unit and a determination unit.
[0122] The detection unit is used to detect the interaction intention of the candidate resources based on the target fusion characteristics and obtain the interaction intention index related to the candidate resources.
[0123] The determining unit is used to determine a target resource from multiple candidate resources based on the interaction intention indicator.
[0124] According to an embodiment of the present disclosure, the interactive behavior feature includes at least one of the following: a browsing time feature, a comment content feature, and a resource content feature of a designated interactive resource.
[0125] Figure 7 A block diagram of an apparatus for training a deep learning model according to an embodiment of the present disclosure is schematically shown.
[0126] like Figure 7 As shown, the device 700 for training a deep learning model includes: a second receiving module 710 and a training module 720.
[0127] The second receiving module 710 is used to receive sample status, sample action and sample feedback information. The sample status includes the sample interaction behavior characteristics of the sample object for the sample specified resource in the sample specified time period after performing the sample resource update operation. The sample resource update operation is used to update the current sample page resource to obtain the sample specified resource. The sample action is to process the sample status information through the deep learning model to obtain the sample interaction intention index related to the sample candidate resource. The sample interaction intention index is used to determine the sample target resource from the sample candidate resources. The sample feedback information represents the satisfaction of the sample object with the sample target resource.
[0128] The training module 720 is used to train the deep learning model based on the reinforcement learning mechanism using sample states, sample actions and sample feedback information to obtain a trained deep learning model.
[0129] According to an embodiment of the present disclosure, sample feedback information is determined based on at least one of the following: sample target interaction behavior characteristics of the sample object with respect to the sample target resource; target interval duration, where the target interval duration represents the time period between different page browsing events, and the page browsing event represents that the sample object opens the target page for displaying the sample target resource for browsing.
[0130] According to an embodiment of the present disclosure, the training module includes: a first obtaining unit, a second obtaining unit, a third obtaining unit and an updating unit.
[0131] The first obtaining unit is configured to process the sample action, the first sample feedback information and the sample state using the first evaluation model to obtain first target feedback information, wherein the first sample feedback information is determined based on the sample target interaction behavior characteristics of the sample object for the sample target resource.
[0132] The second obtaining unit is used to process the sample action, the second sample feedback information and the sample state using the second evaluation model to obtain second target feedback information, wherein the second sample feedback information is determined based on the target interval duration.
[0133] The third obtaining unit is used to perform gradient calculation based on the first target feedback information and the second target feedback information to obtain policy gradient information for the deep learning model.
[0134] The update unit is used to update the model parameters of the deep learning model based on the policy gradient information.
[0135] According to an embodiment of the present disclosure, the third obtaining unit includes: a first obtaining sub-unit and a second obtaining sub-unit.
[0136] The first obtaining subunit is used to use the loss function to process the first target feedback information and the second target feedback information respectively to obtain first loss data and second loss data.
[0137] The second obtaining subunit is used to determine the policy gradient information based on the fused loss data determined by fusing the first loss data and the second loss data.
[0138] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0139] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0140] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.
[0141] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.
[0142] Figure 8 A schematic block diagram of an example electronic device that can be used to implement the resource recommendation method and the method for training a deep learning model of an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0143] like Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.
[0144] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0145] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the resource recommendation method and the method for training a deep learning model. For example, in some embodiments, the resource recommendation method and the method for training a deep learning model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the resource recommendation method and the method for training a deep learning model described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the resource recommendation method and the method for training a deep learning model in any other appropriate manner (for example, by means of firmware).
[0146] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0147] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0148] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0150] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0151] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0152] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0153] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A resource recommendation method, comprising: Receiving interactive behavior characteristics of a target object with respect to a specified resource within a specified period of time after executing a resource update operation, wherein the resource update operation is used to update a current page resource to obtain the specified resource; Fusing the interaction behavior features associated with each of the plurality of designated time periods to obtain a target fusion feature; and A target resource is determined based on the target fusion feature, and the target resource is recommended to the target object.
2. The method according to claim 1, wherein The step of fusing the interaction behavior features associated with each of the plurality of designated time periods to obtain a target fusion feature includes: Performing feature fusion on the plurality of interactive behavior features associated with the specified time period to obtain an initial fused feature; and The target fusion feature is obtained by fusing the initial fusion features related to each of the plurality of execution time periods based on the attention mechanism.
3. The method according to claim 1 or 2, wherein: The determining the target resource based on the target fusion feature includes: Performing interaction intention detection on candidate resources based on the target fusion feature, obtaining an interaction intention index related to the candidate resources; and The target resource is determined from the plurality of candidate resources based on the interaction intention indicator.
4. The method according to claim 1, wherein The interactive behavior characteristics include at least one of the following: Browsing time characteristics, comment content characteristics, and resource content characteristics of specified interactive resources.
5. A method for training a deep learning model, comprising: Receive sample status, sample action, and sample feedback information, wherein the sample status includes sample interaction behavior characteristics of the sample object with respect to the sample-specified resource in a sample-specified period after executing a sample resource update operation, the sample resource update operation being used to update the current sample page resource to obtain the sample-specified resource, the sample action being a sample interaction intention indicator related to the sample candidate resource obtained by processing the sample status information through a deep learning model, the sample interaction intention indicator being used to determine a sample target resource from the sample candidate resources, and the sample feedback information representing the sample object's satisfaction with the sample target resource; Based on a reinforcement learning mechanism, the deep learning model is trained using the sample state, the sample action, and the sample feedback information to obtain a trained deep learning model.
6. The method according to claim 5, wherein: The sample feedback information is determined based on at least one of the following: Sample target interaction behavior characteristics of the sample object with respect to the sample target resource; The target interval duration indicates the time period between different page browsing events. The page browsing event indicates that the sample object opens a target page for displaying the sample target resource for browsing.
7. The method according to claim 6, wherein: The method of training the deep learning model based on the reinforcement learning mechanism using the sample state, the sample action, and the sample feedback information includes: Processing the sample action, the first sample feedback information, and the sample state using a first evaluation model to obtain first target feedback information, wherein the first sample feedback information is determined based on a sample target interaction behavior feature of the sample object with respect to the sample target resource; Processing the sample action, the second sample feedback information, and the sample state using a second evaluation model to obtain second target feedback information, wherein the second sample feedback information is determined based on the target interval duration; Performing gradient calculation based on the first target feedback information and the second target feedback information to obtain policy gradient information for the deep learning model; and Update model parameters of the deep learning model based on the policy gradient information.
8. The method according to claim 7, wherein: The performing gradient calculation based on the first target feedback information and the second target feedback information to obtain policy gradient information for the deep learning model includes: Using a loss function to process the first target feedback information and the second target feedback information respectively to obtain first loss data and second loss data; The policy gradient information is determined based on fused loss data determined by fusing the first loss data and the second loss data.
9. A resource recommendation device, comprising: A first receiving module is configured to receive interaction behavior characteristics of a target object with respect to a specified resource during a specified period of time after executing a resource update operation, wherein the resource update operation is configured to update a current page resource to obtain the specified resource; a fusion module, configured to fuse the interaction behavior features associated with each of the plurality of designated time periods to obtain a target fusion feature; and A recommendation module is used to determine a target resource based on the target fusion feature and recommend the target resource to the target object.
10. A device for training a deep learning model, comprising: A second receiving module is configured to receive sample status, sample action, and sample feedback information. The sample status includes sample interaction behavior characteristics of a sample object with respect to a sample-specified resource during a sample-specified period after executing a sample resource update operation. The sample resource update operation is used to update a current sample page resource to obtain the sample-specified resource. The sample action is a sample interaction intention index related to a sample candidate resource obtained by processing the sample status information through a deep learning model. The sample interaction intention index is used to determine a sample target resource from the sample candidate resources. The sample feedback information indicates the sample object's satisfaction with the sample target resource. A training module is used to train the deep learning model based on a reinforcement learning mechanism using the sample state, the sample action and the sample feedback information to obtain a trained deep learning model.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.
13. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
User intention prediction method and system based on multi-data fusion
CN114386688A
Information recommendation model training method, information recommendation method and equipment
CN116108282A
Resource recommendation method and device, equipment, medium and program product
CN116992123A
Resource recommendation method and device, electronic equipment and storage medium
CN117573973A
Resource display method and device, electronic equipment and storage medium
CN119848351A
Cited By
Information processing method and device, electronic equipment and storage medium
CN121462840A