Information recommendation method and device, electronic equipment and storage medium
By extracting features from user characteristics and scene information through a multi-task network, a personalized resource recommendation list is generated, which solves the problem of user overload on online platforms and improves recommendation effectiveness and user satisfaction.
Patent Information
- Application Number
- CN202511331974.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-01-16
AI Technical Summary
Users face a vast amount of product information on online platforms, making it difficult to quickly and accurately find options that match their preferences, leading to a problem of choice overload.
By extracting features from user characteristics, historical click sequences, and scene-related information through a multi-task network, and combining this with joint modeling of click-through rate and conversion rate, a personalized resource recommendation list is generated.
It improved the initial user experience and overall recommendation performance, addressed the cold start issue, and increased user satisfaction.
Smart Images

Figure CN121350342A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing, and more particularly to the field of artificial intelligence, specifically to an information recommendation method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the rapid development of the internet and information technology, people are increasingly accustomed to booking goods through online platforms, such as booking hotels. However, faced with a vast amount of product information on these platforms, users often find themselves overwhelmed by choices, struggling to quickly and accurately find options that truly suit their preferences. Summary of the Invention
[0003] This disclosure provides an information recommendation method, apparatus, electronic device, and storage medium.
[0004] According to one aspect of this disclosure, an information recommendation method is provided, comprising: extracting features from user characteristics of a target user, user historical click sequences, and resource features of candidate resources through a first processing network to obtain a first fusion feature; determining scene-related information corresponding to the recommended scenario of the target user, and extracting features from the scene-related information through a second processing network to obtain a scene-related vector, wherein the recommended scenario is a new user scenario or an old user scenario; extracting features from the first fusion feature and the scene-related vector through a scene processing network to obtain a second fusion feature; performing multi-task prediction based on the second fusion feature through a multi-task network to obtain a prediction result for each task, wherein the tasks include click tasks and conversion tasks; determining a comprehensive score for candidate resources based on the prediction results, and determining a resource recommendation list based on the comprehensive score.
[0005] The information recommendation method provided in this application, through a multi-task network, simultaneously focuses on click-through rate and conversion rate, and balances the influence between the two through joint modeling, thereby improving the overall recommendation effect. By combining the user characteristics of new users, the user's historical click sequence, and scene-related information, and through effective scene-related vectors and multi-task prediction, new users can immediately receive more personalized and accurate recommendations after their first interaction, improving the user's initial experience and thus alleviating the cold start problem. Through accurate personalized recommendations, user satisfaction is further improved.
[0006] According to another aspect of this disclosure, an information recommendation device is provided, comprising: a first processing module, configured to extract features from user characteristics of a target user, user historical click sequences, and resource features of candidate resources through a first processing network to obtain a first fusion feature; a second processing module, configured to determine scene-related information corresponding to the recommended scene of the target user, and extract scene-related vectors from the scene-related information through the second processing network, wherein the recommended scene is a new user scene or an old user scene; a scene processing module, configured to extract features from the first fusion feature and the scene-related vectors through a scene processing network to obtain a second fusion feature; a task prediction module, configured to perform multi-task prediction based on the second fusion feature through a multi-task network to obtain a prediction result for each task, wherein the tasks include click tasks and conversion tasks; and a comprehensive recommendation module, configured to determine the comprehensive score of the candidate resources based on the prediction results, and determine a resource recommendation list based on the comprehensive score.
[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the aforementioned information recommendation method.
[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the above-described information recommendation method.
[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described information recommendation method.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0012] Figure 1 This is a schematic diagram illustrating an exemplary implementation of an information recommendation method according to an exemplary embodiment of the present disclosure.
[0013] Figure 2 This is a schematic diagram illustrating an exemplary implementation of an information recommendation method according to an exemplary embodiment of the present disclosure.
[0014] Figure 3This is an exemplary schematic diagram of a first processing network according to an exemplary embodiment of the present disclosure.
[0015] Figure 4 This is an exemplary schematic diagram of a scene processing network according to an exemplary embodiment of the present disclosure.
[0016] Figure 5 This is an exemplary schematic diagram of a multitasking network according to an exemplary embodiment of the present disclosure.
[0017] Figure 6 This is a schematic diagram of an overall recommendation network according to an exemplary embodiment of the present disclosure.
[0018] Figure 7 This is an exemplary schematic diagram of an information recommendation device according to an exemplary embodiment of the present disclosure.
[0019] Figure 8 This is a schematic diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0021] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It involves both hardware and software technologies. AI hardware technologies generally include computer vision, speech recognition, natural language processing, as well as learning / deep learning, big data processing, and knowledge graph technologies.
[0022] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0023] Figure 1 This is a schematic diagram illustrating an exemplary implementation of an information recommendation method shown in this application, such as... Figure 1 As shown, this information recommendation method includes the following steps:
[0024] S101, the first processing network extracts features from the user characteristics of the target user, the user's historical click sequence, and the resource characteristics of the candidate resources to obtain the first fused features.
[0025] In this application, the target user refers to a user who needs to receive information push notifications on the recommendation platform.
[0026] The information recommendation method proposed in this application can be applied to various recommendation platforms. For example, it can be applied to platforms that need to recommend information and aim to achieve user conversion, such as hotel recommendation platforms, product recommendation platforms, and book recommendation platforms.
[0027] User characteristics are used to represent a user's specific profile. For example, user characteristics may include information such as the user's age, gender, education level, hobbies, and geographical location. This information helps to more accurately push resources that meet the user's needs.
[0028] Resource features are used to represent the specific attributes of candidate resources. For example, resource features may include the resource ID, category, price, and user reviews of the candidate resource. If it is a hotel resource, it may also include information such as the hotel's location, whether it offers airport pick-up service, and nearby tourist attractions.
[0029] The user's historical click sequence refers to the sequence of resources that the current user has clicked within a recently preset time period (e.g., within the last week) before the current information recommendation is made to the target user. It's easy to understand that in some cases, even if a new user hasn't booked or purchased resources, they may browse and click on the resources displayed to them before making a purchase or booking. After this user browses and clicks, the user's historical click sequence can be quickly obtained. If a user hasn't clicked anything, their historical click sequence can be set as the default click sequence during the initial recommendation.
[0030] In this application, the first fused feature obtained by the first processing network after extracting and fusing the user characteristics of the target user, the user's historical click sequence and the resource characteristics of the candidate resources can fully reflect the user's personalized needs and the multidimensional attributes of the resources.
[0031] S102, determine the scene-related information corresponding to the recommended scene of the target user, and extract the scene-related vector by the second processing network. The recommended scene is either a new user scene or an old user scene.
[0032] In this application, the current recommendation scenario (new user scenario or old user scenario) is mainly determined based on the user's order records for the resources to be recommended.
[0033] For example, in a hotel recommendation platform, if a user has never booked a hotel on the platform, the recommendation scenario for that user is defined as a new user scenario when making hotel recommendations to that user; if a user has previously booked a hotel on the platform, the recommendation scenario for that user is defined as a returning user scenario when making hotel recommendations to that user.
[0034] For example, in a product recommendation platform, if a user has never purchased a product on the platform, the recommendation scenario for that user is defined as a new user scenario when making product recommendations to that user; if a user has purchased a product on the platform before, the recommendation scenario for that user is defined as a returning user scenario when making product recommendations to that user.
[0035] In this application, the scene-related information corresponding to the recommended scene may be the scene features and scene identifiers corresponding to the recommended scene.
[0036] For example, when the recommended scenario is a new user scenario, the scenario identifier can be set to "0", and when the recommended scenario is an old user scenario, the scenario identifier can be set to "1".
[0037] It's easy to understand that new user scenarios correspond to new user scenario features and new user scenario identifiers; old user scenarios correspond to old user scenario features and old user scenario identifiers.
[0038] The second processing network can be an embedding layer, which can convert the corresponding scene features and scene identifiers into vectors through dimensional mapping, and use the vectors converted from scene features and scene identifiers through dimensional mapping as scene-related vectors.
[0039] S103, the second fusion feature is obtained by extracting features from the first fusion feature and the scene-related vector through the scene processing network.
[0040] In this application, a scene processing network is set up. The main function of the scene processing network is to further extract and fuse the first fused feature and scene-related vector obtained above through deep scene feature extraction to obtain deeper user and scene features, which are then represented by the second fused feature.
[0041] S104, Multi-task prediction is performed by a multi-task network based on the second fusion feature to obtain the prediction result for each task, which includes click tasks and conversion tasks.
[0042] In this application, the prediction tasks in the multi-task network include click tasks and conversion tasks.
[0043] Specifically, when performing multi-task prediction based on the second fusion feature through a multi-task network, the multi-task network can ultimately output the predicted click-through rate for the click task and the predicted conversion rate for the conversion task.
[0044] Predicted click-through rate can be understood as the predicted probability that a user will click on a resource after it has been shown to them.
[0045] The predicted conversion rate can be understood as the predicted probability that a user will book or purchase a resource after clicking on a resource recommended to them.
[0046] S105. Determine the comprehensive score of the candidate resources based on the prediction results, and determine the resource recommendation list based on the comprehensive score.
[0047] For any candidate resource, multiply the predicted click-through rate (CTR) and predicted conversion rate (PCC) of that resource to obtain the predicted CTR / PCC of that candidate resource.
[0048] In this application, the predicted click-through rate for each candidate resource is used as its corresponding comprehensive score.
[0049] After determining the comprehensive score for each candidate resource, the candidate resources are sorted in descending order of comprehensive score, and the top K candidate resources are selected to form a resource recommendation list.
[0050] For example, in a hotel recommendation platform, assuming there are 1,000 candidate hotels, after determining the comprehensive score for each candidate hotel, the candidate hotels are sorted in descending order of comprehensive score, and the top K candidate hotels are selected to form a hotel recommendation list. Hotel recommendations are then made to the target user based on this hotel recommendation list.
[0051] For example, in a product recommendation platform, assuming there are 1000 candidate products, after determining the comprehensive score corresponding to each candidate product, the candidate products are sorted in descending order of comprehensive score, and the top K candidate products are selected to form a product recommendation list. Product recommendations are then made to the target user based on this product recommendation list.
[0052] This application proposes an information recommendation method, comprising: extracting features from the user characteristics of the target user, the user's historical click sequence, and the resource characteristics of candidate resources through a first processing network to obtain a first fusion feature; determining the scene-related information corresponding to the target user's recommendation scene, and extracting features from the scene-related information through a second processing network to obtain a scene-related vector, wherein the recommendation scene is a new user scene or an old user scene; extracting features from the first fusion feature and the scene-related vector through a scene processing network to obtain a second fusion feature; performing multi-task prediction based on the second fusion feature through a multi-task network to obtain the prediction result for each task, including click tasks and conversion tasks; determining the comprehensive score of candidate resources based on the prediction results, and determining a resource recommendation list based on the comprehensive score. This application, through a multi-task network, simultaneously focuses on click-through rate and conversion rate, and balances the influence between these two through joint modeling, thereby improving the overall recommendation effect; by combining the user characteristics of new users, the user's historical click sequence, and scene-related information, and through effective scene-related vectors and multi-task prediction, new users can immediately receive more personalized and accurate recommendations after their first interaction, improving the user's initial experience and thus alleviating the cold start problem; and further improving user satisfaction through accurate personalized recommendations.
[0053] Figure 2 This is a schematic diagram illustrating an exemplary implementation of an information recommendation method shown in this application, such as... Figure 2 As shown, this information recommendation method includes the following steps:
[0054] S201, the first processing network extracts features from the user characteristics of the target user, the user's historical click sequence, and the resource characteristics of the candidate resources to obtain the first fused features.
[0055] To illustrate more clearly, Figure 3 This is an exemplary schematic diagram of a first processing network shown in this application, such as... Figure 3 As shown, the first processing network includes a graph neural network (GNN), a basic feature extraction network, a deep interest transformer (DIT), and a first fusion unit.
[0056] like Figure 3 As shown, the dense and sparse features of the basic feature extraction layer to be input are first determined based on user features and resource features.
[0057] Dense features are typically numerical and can be directly extracted from the data. Examples of dense features include user age, browsing time, resource price, and resource rating.
[0058] Sparse features are typically categorical variables or variables with a limited number of valid values. For example, sparse features may include user gender, user city, product category, user device type, etc.
[0059] Among them, Generative Neural Networks (GNNs) are deep learning models capable of processing graph-structured data. In recommender systems, users and resources can be viewed as nodes in a graph, while interactions between users and resources (such as clicks and purchases) constitute edges. By learning from the local neighborhoods of the graph structure, GNNs can effectively capture the relationships between users and resources, and even the potential associations between users and resources. Figure 3 As shown, the user's historical click sequence is input into the GNN to extract GNN features.
[0060] like Figure 3 As shown, the basic feature extraction network extracts features from dense features, GNN features, sparse features, user historical click sequences, and resource features respectively, and obtains the first to fifth feature vectors respectively.
[0061] Specifically, the dense features obtained above are input into the corresponding multilayer perceptron (MLP) in the basic feature extraction network to obtain the first feature vector output.
[0062] Specifically, the GNN features obtained above are input into the corresponding MLP in the basic feature extraction network to obtain the output second feature vector.
[0063] Specifically, the sparse features obtained above are input into the corresponding embedding layer in the basic feature extraction network to obtain the output third feature vector.
[0064] Specifically, the aforementioned user historical click sequence is input into its corresponding embedding layer in the basic feature extraction network to obtain the output fourth feature vector. It's easy to understand that in some cases, even if a new user hasn't booked or purchased resources, they may have browsed and clicked on the resources displayed to them before making a purchase or booking. After this user browses and clicks, the user's historical click sequence can be quickly obtained. If a user hasn't clicked anything, their historical click sequence can be set as the default click sequence during the initial recommendation.
[0065] Specifically, the aforementioned resource features are input into their corresponding embedding layers in the basic feature extraction network to obtain the output fifth feature vector.
[0066] After obtaining the fourth and fifth feature vectors, the fourth and fifth feature vectors are input into a deep interest network to extract user preferences, resulting in a user preference vector. This user preference vector represents the degree of preference of the target user for the current candidate resources.
[0067] In deep interest networks, a Transformer-based architecture is introduced to enhance the ability to model the evolution of user interests. In this framework, the Encoder first receives the fourth and fifth feature vectors. By employing a Self-Attention mechanism, the Encoder can effectively capture the potential relationships between recently viewed resources, thereby modeling the dynamic changes in user interests.
[0068] In the Decoder section, the current candidate resource is treated as a query. Combined with the user's historical behavior sequence, and matched with the user's preference vector, a preference vector representing the user's interest in the current candidate resource is generated. This preference vector reflects the user's potential preferences when considering the current candidate resource and provides a basis for subsequent recommendation decisions.
[0069] Finally, the first to third feature vectors and the user preference vector are input into the first fusion unit (in Figure 3 The first fusion unit (represented by a plus sign in a circle) is fused to obtain the first fusion feature.
[0070] User interests and preferences are multifaceted, and a single type of feature often cannot fully describe user behavior. The first processing network mentioned above, by fusing different features, can integrate different types of data, helping the model capture information from various aspects and improving the accuracy and effectiveness of recommendations. In the deep interest network, the encoder based on the Transformer architecture models the user's historical behavior through a self-attention mechanism, which can dynamically capture changes in the user's interests and reflect the user's latest interests in real time, rather than relying solely on static features. Finally, the first fused feature obtained after fusion allows the recommendation system to better adapt to the changing trends of user interests, recommending resources that best meet the user's current needs and improving user satisfaction.
[0071] S202, determine the scene-related information corresponding to the recommended scene of the target user, and extract the scene-related vector by the second processing network to obtain the scene-related vector. The recommended scene is either the new user scene or the old user scene.
[0072] In this application, the recommendation scenario for the target user is determined based on the target user's resource order records.
[0073] For example, in a hotel recommendation platform, if a user has never booked a hotel on the platform, the recommendation scenario for that user is defined as a new user scenario when making hotel recommendations to that user; if a user has previously booked a hotel on the platform, the recommendation scenario for that user is defined as a returning user scenario when making hotel recommendations to that user.
[0074] For example, in a product recommendation platform, if a user has never purchased a product on the platform, the recommendation scenario for that user is defined as a new user scenario when making product recommendations to that user; if a user has purchased a product on the platform before, the recommendation scenario for that user is defined as a returning user scenario when making product recommendations to that user.
[0075] In this application, the scene-related information corresponding to the recommended scene may be the scene features and scene identifiers corresponding to the recommended scene.
[0076] For example, when the recommended scenario is a new user scenario, the scenario identifier can be set to "0", and when the recommended scenario is an old user scenario, the scenario identifier can be set to "1".
[0077] In this application, scene-related vectors are obtained by extracting features from scene-related information through a second processing network, including: converting scene identifiers and scene features into scene identifier vectors and scene feature vectors respectively through the second processing network (the second processing network can be set as an embedding layer); and using the scene identifier vectors and scene feature vectors as scene-related vectors.
[0078] Recommendations to new users are more challenging due to the lack of historical data. By differentiating between new and returning users' scenarios as described above, recommendation strategies tailored to new users can be adopted. Embedding and extracting features from scenario identifiers (with "0" as the scenario identifier for new users) helps the recommendation system infer potential interests based on limited user characteristics (such as geographic location and interest tags), thereby providing more accurate recommendation results and compensating for the lack of historical data.
[0079] For existing users, the recommendation system already has a wealth of historical data to utilize. By transforming scenario identifiers (existing users are identified by "1"), the recommendation system can leverage users' historical behavioral data (such as booked hotels, purchased items, etc.) to gain a deeper understanding of user preferences, further improving the personalization and accuracy of recommendations.
[0080] S203, the second fusion feature is obtained by extracting features from the first fusion feature and the scene-related vector through the scene processing network.
[0081] To illustrate more clearly, Figure 4This is an exemplary schematic diagram of a scene processing network shown in this application, such as... Figure 4 As shown, the scene processing network includes a shared feature extraction subnetwork, a scene feature extraction subnetwork, and a second fusion unit.
[0082] like Figure 4 As shown, the first fused feature is input into the shared feature extraction subnetwork to obtain the extraction result of the shared feature subnetwork.
[0083] like Figure 4 As shown, the first fused feature and the scene-related vector are input into the scene feature extraction sub-network to obtain the extraction result of the scene feature extraction sub-network.
[0084] like Figure 4 As shown, based on the second fusion unit (in Figure 4 The second fusion unit (represented by a plus sign in a circle) fuses the extraction results of the shared feature subnetwork and the scene feature extraction subnetwork to obtain the second fused feature.
[0085] The following section will detail the processing steps of the shared feature extraction subnetwork and the scene feature extraction subnetwork.
[0086] like Figure 4 As shown, the shared feature extraction subnetwork includes a shared gating network and a first shared expert network. The first fused feature is input into the shared feature extraction subnetwork to obtain the shared feature extraction result. This includes: inputting the first fused feature into the shared gating network and the first shared expert network respectively to obtain the first weight output by the shared gating network and the output result of the first shared expert network respectively; and then performing a dot product on the first weight and the output result of the first shared expert network to obtain the shared feature extraction result.
[0087] The first shared expert network may include multiple different first shared expert sub-networks, which are used to process the first fusion feature from different perspectives. The output results of each first shared expert sub-network constitute the output result of the first shared expert network.
[0088] The first weight output by the shared gating network is a weight vector, which contains multiple weights that correspond one-to-one with the output results of each first shared expert subnetwork.
[0089] like Figure 4As shown, the scene feature extraction subnetwork includes a scene expert network, a first scene gating network, and a second scene gating network. The scene expert network includes a new user scene expert network and an old user scene expert network. The first fused feature and scene-related vector are input into the scene feature extraction subnetwork to obtain the extraction result of the scene feature extraction subnetwork, including: obtaining the scene expert network output result and the second weight based on the scene expert network and the first scene gating network, respectively. The inputs of the scene expert network and the first scene gating network are the first fused feature and the scene identifier vector; the scene identifier vector and the scene feature vector are input into the second scene gating network to obtain the third weight; the scene expert network output result is multiplied by the second weight and the third weight respectively to obtain the first sub-result and the second sub-result of the scene feature extraction subnetwork.
[0090] The scene expert network may include multiple different scene expert subnetworks, which are used to process the first fused feature and scene identifier vector from different perspectives. The output results of each scene expert subnetwork constitute the output result of the scene expert network.
[0091] Among them, the second weight output by the first scene gating network is a weight vector, which contains multiple weights that correspond one-to-one with the output results of each scene expert subnetwork.
[0092] The third weight output by the second scenario gating network is a weight vector, which contains multiple weights that correspond one-to-one with the output results of each scenario expert subnetwork.
[0093] In the first scene gating network, based on the scene identifier vector, the weights of the non-target scene expert networks in the scene expert network that are opposite to the target scene corresponding to the scene identifier vector are reset to 0, and only the feature contribution of the target scene expert network is retained.
[0094] The second scene gating network uses the scene identifier vector as a basis to reset the weight of the target scene expert network that is consistent with the target scene corresponding to the scene identifier vector to 0, and only retains the cross-scene feature supplement of the non-target scene expert network.
[0095] For example, suppose the scene expert network includes 10 scene expert subnetworks, divided into 5 new user scene expert subnetworks and 5 old user scene expert subnetworks. When the scene identifier vector represents a new user scene, the weights corresponding to the 5 old user scene expert subnetworks in the second weight output of the first scene gating network are reset to 0, while the weights corresponding to the 5 new user scene expert subnetworks in the second weight output of the first scene gating network are assigned normal values. Similarly, the weights corresponding to the 5 new user scene expert subnetworks in the third weight output of the second scene gating network are reset to 0, while the weights corresponding to the 5 old user scene expert subnetworks in the third weight output of the second scene gating network are assigned normal values. That is, the first scene gating network retains only the feature contributions from the new user scene expert subnetworks, and the second scene gating network retains only the cross-scene feature supplements from the old user scene expert subnetworks.
[0096] For example, suppose the scene expert network includes 10 scene expert subnetworks, divided into 5 new user scene expert subnetworks and 5 old user scene expert subnetworks. When the scene identifier vector represents an old user scene, the weights corresponding to the 5 new user scene expert subnetworks in the second weight output of the first scene gating network are reset to 0, while the weights corresponding to the 5 old user scene expert subnetworks in the second weight output of the first scene gating network are assigned normal values. Similarly, the weights corresponding to the 5 old user scene expert subnetworks in the third weight output of the second scene gating network are reset to 0, while the weights corresponding to the 5 new user scene expert subnetworks in the third weight output of the second scene gating network are assigned normal values. That is, the first scene gating network retains only the feature contributions of the old user scene expert subnetworks, and the second scene gating network retains only the cross-scene feature supplementation of the new user scene expert subnetworks.
[0097] The scene processing network described above includes a first shared expert network in the shared feature extraction subnetwork, which contains multiple subnetworks and can process the first fused features from different perspectives. These features are then weighted by the weights output by the shared gating network, capturing richer feature information. The first scene gating network in the scene feature extraction subnetwork retains only the feature contributions of the target scene expert subnetwork based on the scene identifier vector, allowing the model to focus on the key features of the current scene and better adapt to different scenarios. The second scene gating network, based on the scene identifier vector, retains only the cross-scene feature supplements from non-target scene expert networks. This allows the model to reference relevant information from other scenes when processing the current scene, preventing it from being overly limited to the current scene and improving its generalization ability.
[0098] S204, multi-task prediction is performed by a multi-task network based on the second fusion feature to obtain the prediction result for each task, which includes click tasks and conversion tasks.
[0099] To illustrate more clearly, Figure 5 This is an exemplary schematic diagram of a multi-task network shown in this application, such as... Figure 5 As shown, the multi-task network includes a click task sub-network, a conversion task sub-network, and a second shared expert network.
[0100] like Figure 5 As shown, a multi-task network performs multi-task prediction based on the second fusion feature to obtain the prediction result for each task. This includes: inputting the second fusion feature into a second shared expert network to obtain the output result of the second shared expert network; inputting the second fusion feature into the click gating network and click expert network included in the click task sub-network to obtain the fourth weight and the output result of the click expert network; processing the fourth weight, the output result of the click expert network, and the output result of the second shared expert network, and then inputting them into the click tower network included in the click task sub-network to obtain the predicted click-through rate corresponding to the click task; inputting the second fusion feature into the conversion gating network and conversion expert network included in the conversion task sub-network to obtain the fifth weight and the output result of the conversion expert network; processing the fifth weight, the output result of the conversion expert network, and the output result of the second shared expert network, and then inputting them into the conversion tower network included in the conversion task sub-network to obtain the predicted conversion rate corresponding to the conversion task.
[0101] The click expert network contains multiple click expert subnetworks.
[0102] The second shared expert network contains multiple second shared expert sub-networks.
[0103] The transformation expert network contains multiple transformation expert sub-networks.
[0104] The fourth weight output by the click-gated network is a weight vector, and the number of weights contained in this vector is the sum of the number of the click expert subnetwork and the number of the second shared expert subnetwork.
[0105] The fifth weight output by the transformation gating network is a weight vector, and the number of weights contained in this vector is the sum of the number of transformation expert subnetworks and the number of second shared expert subnetworks.
[0106] For example, suppose the click expert network contains 5 click expert sub-networks, the second shared expert network contains 5 second shared expert sub-networks, and the conversion expert network contains 5 conversion expert sub-networks. Then, the fourth weight output by the click gating network includes 10 weight values. Five of these weight values are multiplied one-to-one with the outputs of the five click expert sub-networks, and the other five weight values are multiplied one-to-one with the outputs of the five second shared expert sub-networks. The results of these multiplications are concatenated and input into the click tower network to obtain the predicted click-through rate (CTR) output by the click tower network. Similarly, the fifth weight output by the conversion gating network also includes 10 weight values. Five of these weight values are multiplied one-to-one with the outputs of the five conversion expert sub-networks, and the other five weight values are multiplied one-to-one with the outputs of the five second shared expert sub-networks. The results of these multiplications are concatenated and input into the conversion tower network to obtain the predicted conversion rate (CTR) output by the conversion tower network.
[0107] In the multi-task network design described above, a second shared expert network is used to extract common features among different tasks, enabling click and conversion tasks to share basic feature information and avoiding redundant feature extraction and calculation. Dedicated expert networks (click expert network, conversion expert network) and gating networks (click gating network, conversion gating network) are designed for click and conversion tasks respectively, which can specifically capture the unique features of each task. The multi-task joint prediction approach enables the model to learn more robust feature representations, reduces the risk of overfitting to single task data, improves the model's generalization ability in different scenarios, and improves the accuracy of predicting click-through rate and conversion rate.
[0108] S205. For any candidate resource, the product of the predicted click-through rate and the predicted conversion rate of the candidate resource is used as the comprehensive score of the candidate resource.
[0109] To illustrate more clearly, Figure 6 This is a schematic diagram of an overall recommendation network shown in this application. In addition to the first processing network, the second processing network, the scene processing network, and the multi-task network mentioned above, the overall recommendation network also includes a click-to-conversion rate calculation unit after the multi-task network. The click-to-conversion rate calculation unit is used to multiply the predicted click-through rate output by the click tower network and the predicted conversion rate output by the conversion tower network corresponding to the same candidate resource to obtain the predicted click-to-conversion rate corresponding to the candidate resource.
[0110] In this application, the predicted click-through rate for each candidate resource is used as its corresponding comprehensive score.
[0111] S206, Determine the resource recommendation list based on the overall score.
[0112] After determining the comprehensive score for each candidate resource, the candidate resources are sorted in descending order of comprehensive score, and the top K candidate resources are selected to form a resource recommendation list.
[0113] For example, in a hotel recommendation platform, assuming there are 1,000 candidate hotels, after determining the comprehensive score for each candidate hotel, the candidate hotels are sorted in descending order of comprehensive score, and the top K candidate hotels are selected to form a hotel recommendation list. Hotel recommendations are then made to the target user based on this hotel recommendation list.
[0114] For example, in a product recommendation platform, assuming there are 1000 candidate products, after determining the comprehensive score corresponding to each candidate product, the candidate products are sorted in descending order of comprehensive score, and the top K candidate products are selected to form a product recommendation list. Product recommendations are then made to the target user based on this product recommendation list.
[0115] This application embodiment details the internal processing logic of the first processing network, the second processing network, the scene processing network, and the multi-task network, accurately capturing feature information at each level. Finally, by multiplying the predicted click-through rate output by the click tower network and the predicted conversion rate output by the conversion tower network corresponding to the same candidate resource, the predicted click-through conversion rate corresponding to the candidate resource is obtained as the comprehensive score of the candidate resource, providing accurate personalized recommendations for users and further improving user satisfaction.
[0116] The online recommendation process has been described above; the training process of the model will be described in detail below.
[0117] In this application, the first processing network, the second processing network, the scene processing network, and the multi-task network together constitute a multi-task prediction model (that is, the multi-task prediction model can be regarded as the above). Figure 6 After removing the topmost unit that multiplies the predicted click-through rate and predicted conversion rate, the multi-task prediction model is obtained. This can also be understood as, after training the multi-task prediction model as described below, adding a unit after the click-tower and conversion-tower networks to multiply the predicted click-through rate and predicted conversion rate, resulting in the overall recommendation network used to calculate the comprehensive score of candidate resources during online recommendation. The training method for the multi-task prediction model includes the following steps:
[0118] First, obtain the first training sample set corresponding to the click task, and then obtain the second training sample set corresponding to the conversion task.
[0119] The first training sample set includes multiple click task training samples. Each click task training sample includes sample user features, sample user historical click sequence and sample resource features, as well as the recommendation scenario and scenario features when making recommendations to the sample user, and click sample tags.
[0120] In the first training sample set, each positive click sample resource corresponds to N associated negative click sample resources. Specifically, among the same batch of recommended resources recommended to sample users, the resources clicked by the user are selected as the positive click sample resources corresponding to the click task, and N resources not clicked by the user are selected as the negative click sample resources associated with that positive click sample resource.
[0121] For example, taking N as 2, suppose there are 10 resources recommended to user A, numbered 10 to 10. User A clicked on resource 1 but did not click on resources 2 to 10. Resource 1 is then designated as the positive sample resource (carrying a positive sample label) corresponding to user A's clicks. Two resources from resource 2 to 10 are randomly selected as negative sample resources (carrying negative sample labels) associated with resource 1's clicks. That is, at this point, three sample resources corresponding to user A are obtained.
[0122] The second training sample set includes multiple conversion task training samples. Each conversion task training sample includes sample user features, sample user historical click sequence and sample resource features, as well as the recommendation scenario and scenario features when making recommendations to the sample user, and conversion sample tags.
[0123] In the second training sample set, each positive conversion sample resource corresponds to M associated negative conversion sample resources. Specifically, among the same batch of recommended resources recommended to sample users, resources that successfully convert after user clicks (i.e., complete the purchase operation) are selected as positive conversion sample resources corresponding to the conversion task, and M resources that do not successfully convert after user clicks are selected as negative conversion sample resources associated with the positive conversion sample resource.
[0124] For example, taking M as a value of 3, suppose there are 20 resources recommended to user A, namely resource 1 to resource 20. User A clicked on resource 11 and completed the conversion, but clicked on resources 12 to 14 without successful conversion. In this case, resource 11 is regarded as a positive conversion sample resource (carrying a positive sample label), and resources 12 to 14 are regarded as negative conversion sample resources associated with resource 11 (carrying a negative sample label). That is, at this time, we have 4 sample resources corresponding to sample user A.
[0125] The values of M and N can be set according to the actual situation, and this application does not impose any restrictions on them.
[0126] Preferably, M is greater than N.
[0127] After obtaining the first training sample set and the second training sample set, the multi-task prediction model is iteratively trained based on the first training sample set and the second training sample set until the loss function converges, thus obtaining the trained multi-task prediction model.
[0128] In recommendation scenarios, both clicks and conversions are considered "low-probability events" (e.g., only 5 out of 100 recommended resources are clicked, and only 1 out of those 5 clicks results in a conversion). If the entire dataset is used for training, an imbalance problem will occur, with "far more negative samples than positive samples," causing the model to favor predicting "negative samples" and reducing prediction accuracy. The model training process described above avoids this bias caused by a fixed ratio of positive to negative sample associations.
[0129] Figure 7 This is an exemplary schematic diagram of an information recommendation device shown in this application, such as... Figure 7 As shown, the information recommendation device 700 includes a first processing module 701, a second processing module 702, a scene processing module 703, a task prediction module 704, and a comprehensive recommendation module 705, wherein:
[0130] The first processing module 701 is used to extract features from the user characteristics of the target user, the user's historical click sequence, and the resource characteristics of the candidate resources through the first processing network to obtain the first fused features;
[0131] The second processing module 702 is used to determine the scene-related information corresponding to the recommended scene of the target user, and to extract the scene-related information through the second processing network to obtain the scene-related vector. The recommended scene is either a new user scene or an old user scene.
[0132] The scene processing module 703 is used to extract features from the first fused feature and the scene-related vector through the scene processing network to obtain the second fused feature;
[0133] The task prediction module 704 is used to perform multi-task prediction based on the second fusion feature through a multi-task network to obtain the prediction result for each task, including click tasks and conversion tasks.
[0134] The comprehensive recommendation module 705 is used to determine the comprehensive score of candidate resources based on the prediction results, and to determine the resource recommendation list based on the comprehensive score.
[0135] This device utilizes a multi-task network to simultaneously monitor click-through rate and conversion rate, and balances their impact through joint modeling, thereby improving overall recommendation performance. By combining new user characteristics, historical click sequences, and scene-related information, and employing effective scene-related vectors and multi-task prediction, new users can receive more personalized and accurate recommendations immediately after their first interaction, enhancing their initial user experience and mitigating the cold start problem. Furthermore, precise personalized recommendations further improve user satisfaction.
[0136] Furthermore, the scene-related information includes scene identifiers and scene features. The second processing module 702 is also used to: convert the scene identifiers and scene features into scene identifier vectors and scene feature vectors respectively through the second processing network; and use the scene identifier vectors and scene feature vectors as scene-related vectors.
[0137] Furthermore, the first processing network includes a graph neural network (GNN), a basic feature extraction network, a deep interest network, and a first fusion unit. The first processing module 701 is also used for: determining dense features and sparse features based on user features and resource features; inputting the user's historical click sequence into the GNN to obtain GNN features; extracting features from the dense features, GNN features, sparse features, user historical click sequence, and resource features respectively through the basic feature extraction network to obtain first to fifth feature vectors respectively; inputting the fourth and fifth feature vectors into the deep interest network to extract user preferences to obtain user preference vectors; and inputting the first to third feature vectors and the user preference vectors into the first fusion unit for fusion to obtain the first fused feature.
[0138] Furthermore, the scene processing network includes a shared feature extraction subnetwork, a scene feature extraction subnetwork, and a second fusion unit. The scene processing module 703 is also used to: input the first fused feature into the shared feature extraction subnetwork to obtain the shared feature extraction result; input the first fused feature and the scene-related vector into the scene feature extraction subnetwork to obtain the scene feature extraction result; and fuse the shared feature extraction result and the scene feature extraction subnetwork based on the second fusion unit to obtain the second fused feature.
[0139] Furthermore, the shared feature extraction subnetwork includes a shared gating network and a first shared expert network. The scene processing module 703 is also used to: input the first fused features into the shared gating network and the first shared expert network respectively, and obtain the first weight and the output result of the first shared expert network respectively; perform a dot product on the first weight and the output result of the first shared expert network to obtain the shared feature extraction result of the subnetwork.
[0140] Furthermore, the scene feature extraction subnetwork includes a scene expert network, a first scene gating network, and a second scene gating network. The scene expert network includes a new user scene expert network and an old user scene expert network. The scene processing module 703 is also used to: obtain the scene expert network output result and the second weight based on the scene expert network and the first scene gating network, respectively, wherein the inputs of the scene expert network and the first scene gating network are the first fused feature and the scene identifier vector; input the scene identifier vector and the scene feature vector into the second scene gating network to obtain the third weight; and perform dot product between the scene expert network output result and the second weight and the third weight, respectively, to obtain the first sub-result and the second sub-result of the scene feature extraction subnetwork.
[0141] Furthermore, in the scene processing module 703, the first scene gating network uses the scene identifier vector as a basis to reset the weights of the non-target scene expert networks in the scene expert network that are opposite to the target scene corresponding to the scene identifier vector to 0; the second scene gating network uses the scene identifier vector as a basis to reset the weights of the target scene expert networks in the scene expert network that are consistent with the target scene corresponding to the scene identifier vector to 0.
[0142] Furthermore, in the second processing module 702, the recommended scenario for the target user is determined based on the target user's resource order records.
[0143] Furthermore, the multi-task network includes a click task sub-network, a conversion task sub-network, and a second shared expert network. The task prediction module 704 is also used for: inputting the second fusion feature into the second shared expert network to obtain the output result of the second shared expert network; inputting the second fusion feature into the click gating network and the click expert network included in the click task sub-network to obtain the fourth weight and the output result of the click expert network; processing the fourth weight, the output result of the click expert network, and the output result of the second shared expert network and then inputting them into the click tower network included in the click task sub-network to obtain the predicted click-through rate corresponding to the click task; inputting the second fusion feature into the conversion gating network and the conversion expert network included in the conversion task sub-network to obtain the fifth weight and the output result of the conversion expert network; processing the fifth weight, the output result of the conversion expert network, and the output result of the second shared expert network and then inputting them into the conversion tower network included in the conversion task sub-network to obtain the predicted conversion rate corresponding to the conversion task.
[0144] Furthermore, the comprehensive recommendation module 705 is also used to: for any candidate resource, multiply the predicted click-through rate and the predicted conversion rate of the candidate resource as the comprehensive score of the candidate resource.
[0145] Furthermore, the information recommendation device 700 also includes: a model training module, used to obtain a first training sample set corresponding to the click task, and a second training sample set corresponding to the conversion task; iteratively training the multi-task prediction model based on the first training sample set and the second training sample set until the loss function converges, thereby obtaining a trained multi-task prediction model, wherein the first processing network, the second processing network, the scene processing network, and the multi-task network together constitute the multi-task prediction model.
[0146] Furthermore, in the first training sample set, each positive click sample resource corresponds to N associated negative click sample resources; in the second training sample set, each positive conversion sample resource corresponds to M associated negative conversion sample resources.
[0147] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0148] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0149] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0150] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0151] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as information recommendation methods. For example, in some embodiments, the information recommendation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the information recommendation method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform information recommendation methods by any other suitable means (e.g., by means of firmware).
[0152] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0153] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0154] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0156] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0157] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0158] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0159] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An information recommendation method, comprising: extracting features of a target user, a user historical click sequence and resource features of a candidate resource through a first processing network to obtain first fusion features; determining scene-related information corresponding to a recommendation scene of the target user, extracting features of the scene-related information through a second processing network to obtain a scene-related vector, the recommendation scene being a new user scene or an old user scene; extracting features of the first fusion features and the scene-related vector through a scene processing network to obtain second fusion features; performing multi-task prediction according to the second fusion features through a multi-task network to obtain a prediction result of each task, the tasks including a click task and a conversion task; determining a comprehensive score of the candidate resource according to the prediction result, and determining a resource recommendation list according to the comprehensive score.
2. The method of claim 1, wherein, The scene-related information includes a scene identifier and a scene feature, and the extracting features of the scene-related information through the second processing network to obtain a scene-related vector comprises: transforming the scene identifier and the scene feature into a scene identifier vector and a scene feature vector through the second processing network respectively; taking the scene identifier vector and the scene feature vector as the scene-related vector.
3. The method of claim 2, wherein, The first processing network includes a graph neural network (GNN), a basic feature extraction network, a deep interest network and a first fusion unit, and the extracting features of the target user, the user historical click sequence and the resource features of the candidate resource through the first processing network to obtain the first fusion features comprises: determining dense features and sparse features according to the user features and the resource features; inputting the user historical click sequence into the GNN to obtain GNN features; extracting features of the dense features, the GNN features, the sparse features, the user historical click sequence and the resource features through the basic feature extraction network respectively to obtain first to fifth feature vectors respectively; inputting the fourth feature vector and the fifth feature vector into the deep interest network to extract user preferences to obtain a user preference vector; inputting the first to third feature vectors and the user preference vector into the first fusion unit for fusion to obtain the first fusion features.
4. The method of claim 3, wherein, The scene processing network includes a shared feature extraction subnetwork, a scene feature extraction subnetwork and a second fusion unit, and the extracting features of the first fusion features and the scene-related vector through the scene processing network to obtain the second fusion features comprises: inputting the first fusion features into the shared feature extraction subnetwork to obtain a shared feature subnetwork extraction result; inputting the first fusion features and the scene-related vector into the scene feature extraction subnetwork to obtain a scene feature extraction subnetwork extraction result; fusing the shared feature subnetwork extraction result and the scene feature extraction subnetwork extraction result based on the second fusion unit to obtain the second fusion features.
5. The method of claim 4, wherein, The shared feature extraction sub-network comprises a shared gating network and a first shared expert network, and the first fused feature is input into the shared feature extraction sub-network to obtain a shared feature sub-network extraction result, which comprises: The first fused feature is input into the shared gating network and the first shared expert network respectively to obtain a first weight and a first shared expert network output result respectively; The first weight and the first shared expert network output result are point multiplied to obtain the shared feature sub-network extraction result.
6. The method of claim 4, wherein, The scene feature extraction sub-network comprises a scene expert network, a first scene gating network and a second scene gating network, the scene expert network comprises a new user scene expert network and an old user scene expert network, and the first fused feature and the scene-related vector are input into the scene feature extraction sub-network to obtain a scene feature extraction sub-network extraction result, which comprises: The scene expert network and the first scene gating network are based on the first fused feature and the scene identification vector to obtain a scene expert network output result and a second weight respectively, wherein the inputs of the scene expert network and the first scene gating network are the first fused feature and the scene identification vector; The scene identification vector and the scene feature vector are input into the second scene gating network to obtain a third weight; The scene expert network output result is point multiplied with the second weight and the third weight respectively to obtain a first sub-result and a second sub-result of the scene feature extraction sub-network.
7. The method of claim 6, wherein, Wherein: The first scene gating network sets the weight of a non-target scene expert network opposite to a target scene corresponding to the scene identification vector in the scene expert network to 0 based on the scene identification vector; The second scene gating network sets the weight of a target scene expert network consistent with the target scene corresponding to the scene identification vector in the scene expert network to 0 based on the scene identification vector.
8. The method according to any one of claims 1-7, characterized in that, The recommended scene of the target user is determined based on the resource order record of the target user.
9. The method of claim 8, wherein, The multi-task network comprises a click task sub-network, a conversion task sub-network and a second shared expert network, and the multi-task network performs multi-task prediction based on the second fused feature to obtain a prediction result of each task, which comprises: The second fused feature is input into the second shared expert network to obtain a second shared expert network output result; The second fused feature is input into a click gating network and a click expert network included in the click task sub-network respectively to obtain a fourth weight and a click expert network output result; The fourth weight, the click expert network output result and the second shared expert network output result are processed and then input into a click tower network included in the click task sub-network to obtain a predicted click rate corresponding to the click task; The second fused feature is input into a conversion gating network and a conversion expert network included in the conversion task sub-network respectively to obtain a fifth weight and a conversion expert network output result; The fifth weight, the conversion expert network output result and the second shared expert network output result are input into a conversion tower network included in the conversion task sub-network after processing, to obtain a predicted conversion rate corresponding to the conversion task.
10. The method of claim 9, wherein, The determining of the comprehensive score of the candidate resource according to the prediction result comprises: For any candidate resource, the product of the predicted click rate and the predicted conversion rate corresponding to the candidate resource is taken as the comprehensive score corresponding to the candidate resource.
11. The method of claim 1, wherein, The first processing network, the second processing network, the scene processing network and the multi-task network jointly constitute a multi-task prediction model, and a training method of the multi-task prediction model comprises: obtaining a first training sample set corresponding to the click task, and obtaining a second training sample set corresponding to the conversion task; iteratively training the multi-task prediction model according to the first training sample set and the second training sample set until the loss function converges, to obtain a trained multi-task prediction model.
12. The method of claim 11, wherein, Wherein: In the first training sample set, each click positive sample resource corresponds to N associated click negative sample resources; In the second training sample set, each conversion positive sample resource corresponds to M associated conversion negative sample resources.
13. An information recommendation device, comprising: A first processing module is configured to extract features of a user feature of a target user, a user historical click sequence and a resource feature of a candidate resource through a first processing network to obtain a first fusion feature; A second processing module is configured to determine scene-related information corresponding to a recommendation scene of the target user, extract features of the scene-related information through a second processing network to obtain a scene-related vector, and the recommendation scene is a new user scene or an old user scene; A scene processing module is configured to extract features of the first fusion feature and the scene-related vector through a scene processing network to obtain a second fusion feature; A task prediction module is configured to perform multi-task prediction according to the second fusion feature through a multi-task network to obtain a prediction result of each task, and the task includes a click task and a conversion task; A comprehensive recommendation module is configured to determine a comprehensive score of the candidate resource according to the prediction result, and determine a resource recommendation list according to the comprehensive score.
14. An electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.
15. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-12.
16. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method of any one of claims 1-12.
Citation Information
Cited By
Content recommendation method and device, equipment, storage medium and program product
CN122286235A