Model training method and device, model recommendation method and device, medium, product and equipment

Through joint training of multimodal networks and multitask networks, the problem of low accuracy in recommendation of unpopular resources is solved, and more efficient resource recommendation model training and multi-scene adaptability are achieved.

CN120386919APending Publication Date: 2025-07-29HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411716629.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, it is difficult to accurately recommend unpopular resources based on collaborative filtering and behavioral sequence recommendation methods, resulting in low recommendation accuracy.

Method used

Through the scene network corresponding to the multimodal network, multi-task network, first network and multiple recommended scenarios, the user data and task tags of the full sample are obtained, the full sample features are determined, and the resource recommendation model is trained using multimodal features and public features to generate the target resource recommendation model.

Benefits of technology

It improves the accuracy and training efficiency of resource recommendations, avoids popular resources leading to unpopular resource learning, and enhances the generalization ability of the model and the adaptability of multi-scene multi-task recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386919A_ABST
    Figure CN120386919A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and provides a model training method and device, a model recommendation method and device, a medium, a product and equipment. The model training method comprises the following steps: acquiring full sample user data of a plurality of recommendation scenes and full task tags of the plurality of recommendation scenes, and determining full sample features of the full sample user data, and training a resource recommendation model according to the full sample features, the multi-modal network, the scene network and a first network for public feature extraction, thereby obtaining a target resource recommendation model. According to the scheme, the accuracy of resource recommendation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a resource recommendation model training method, a resource recommendation method, a resource recommendation model training device, a resource recommendation device, a computer-readable storage medium, a computer program product, and an electronic device. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims, and no statement herein is admitted to be prior art by inclusion in this section.

[0003] Recommending resources to users can improve their resource acquisition efficiency. Current mainstream recommendation methods include collaborative filtering and behavior sequence-based recommendations, which can discover new items that users may be interested in based on their historical behavior and make recommendations. Summary of the Invention

[0004] However, for unpopular resources with little user interaction data, due to insufficient interaction information, recommendation methods based on collaborative filtering and behavior sequence are difficult to accurately find users who are interested in unpopular resources, resulting in low accuracy in recommending unpopular resources.

[0005] To this end, a resource recommendation model training method and a resource recommendation method are urgently needed to improve the accuracy of resource recommendation.

[0006] In this context, embodiments of the present invention are intended to provide a resource recommendation model training method, a resource recommendation method, a resource recommendation model training device, a resource recommendation device, a computer-readable storage medium, a computer program product, and an electronic device.

[0007] According to the first aspect of the embodiments of the present disclosure, a method for training a resource recommendation model is provided. The resource recommendation model includes a multi-modal network, a multi-task network, a first network, and a plurality of scenario networks corresponding to each of a plurality of recommendation scenarios. The method includes: obtaining the full-scale sample user data of the plurality of recommendation scenarios and the full-scale task labels of the plurality of recommendation scenarios, and determining the full-scale sample features of the full-scale sample user data. The full-scale sample features include a first resource identification sequence corresponding to the resources in the resource behavior sequence of the sample user, a first resource multi-modal feature sequence corresponding to the first multi-modal features of the resources in the resource behavior sequence of the sample user, the attribute features of the sample user, and the private features of each recommendation scenario. According to the first multi-modal feature sequence, the second multi-modal feature of the resource to be recommended, and the multi-modal network, a first relationship vector is obtained; according to the resource identification, resource attribute features, resource statistical features of the resource to be recommended, the attribute features of the sample user, the first resource identification sequence, and the first network, a second relationship vector is obtained; according to the scenario network corresponding to each recommendation scenario, a third vector of the private features of each recommendation scenario is obtained; according to the first relationship vector, the second relationship vector, the third vector, and the multi-task network, a task prediction result for each recommendation scenario is obtained; based on the task prediction result and the full-scale task labels, the resource recommendation model is iteratively trained to obtain a target resource recommendation model for resource recommendation.

[0008] Optionally, the multi-task network includes a first multi-task network and a second multi-task network, and the structures of the first multi-task network and the second multi-task network are the same. The obtaining the task prediction result for each recommendation scenario according to the first relationship vector, the second relationship vector, the third vector, and the multi-task network includes: for each recommendation scenario, performing the following processing procedure to obtain the task prediction result for the recommendation scenario: according to the fusion result of the first relationship vector, the second relationship vector, and the third vector corresponding to the private features of the recommendation scenario, a first fusion vector of the recommendation scenario is obtained; inputting the first fusion vector into each task network related to the recommendation scenario in the first multi-task network to obtain a first task prediction result for the recommendation scenario; inputting the first relationship vector into each task network of the second multi-task network to obtain a second task prediction result for the recommendation scenario.

[0009] Optionally, iteratively training the resource recommendation model based on the task prediction result to obtain a target resource recommendation model for resource recommendation includes: for each recommendation scenario, determining the task labels corresponding to each recommendation scenario from the full set of task labels; obtaining the first task loss of the recommendation scenario according to the difference degree between the first task prediction result of each recommendation scenario and the task labels of the recommendation scenario; obtaining the second task loss of the recommendation scenario according to the difference degree between the second task prediction result of each recommendation scenario and the task labels of the recommendation scenario; obtaining the task loss of the recommendation scenario according to the fusion result of the first task loss and the second task loss of the recommendation scenario; obtaining the training loss according to the fusion result of the task losses of each recommendation scenario, and iteratively training each network in the resource recommendation model according to the training loss until a preset condition is met, to obtain the target resource recommendation model.

[0010] Optionally, the multimodal network includes a first self-attention model and a first cross-attention model. Obtaining the first relationship vector according to the first multimodal feature sequence, the second multimodal feature of the resource to be recommended, and the multimodal network includes: inputting the first multimodal feature sequence into the first self-attention model to process the first multimodal feature sequence, and obtaining a first output result; inputting the first output result and the second multimodal feature into the first cross-attention model, and obtaining the first relationship vector according to the output of the first cross-attention model.

[0011] Optionally, the first network includes a second self-attention model, a second cross-attention model, a shared embedding layer, and a first fully-connected layer. Obtaining the second relationship vector according to the resource identifier of the resource to be recommended, the resource attribute features, the resource statistical features, the attribute features of the sample user, the first resource identifier sequence, and the first network includes: mapping the resource identifier of the resource to be recommended, the resource attribute features, the resource statistical features, the attribute features of the sample user, and the first resource identifier sequence to the same feature space through the shared embedding layer; inputting the feature-mapped first resource identifier sequence features into the second self-attention model to process the first resource identifier sequence, and obtaining a second output result according to the output of the second self-attention model; inputting the second output result and the feature-mapped resource identifier of the resource to be recommended into the second cross-attention model, and obtaining a first candidate relationship vector according to the output of the second cross-attention model; inputting the first candidate relationship vector, the feature-mapped resource attribute features, the resource statistical features, and the attribute features of the sample user of the resource to be recommended into the first fully-connected layer, and obtaining the second relationship vector according to the output of the first fully-connected layer.

[0012] Optionally, the scenario network includes a feature mask layer, a second fully-connected layer, a batch normalization layer, a third fully-connected layer, and a dropout layer. The obtaining of the third vector of the private features of each recommended scenario according to the scenario network corresponding to each recommended scenario includes: for each recommended scenario, inputting the full sample features into the feature mask layer of the scenario network corresponding to the recommended scenario, so that the feature mask layer determines the private features of the recommended scenario from the full sample features; sequentially inputting the private features of the recommended scenario into the second fully-connected layer, the batch normalization layer, and the third fully-connected layer to obtain a third output result; adding the fourth output result of the second fully-connected layer and the third output result to obtain a fifth output result; inputting the fifth output result into the dropout layer, and obtaining the third vector according to the output of the dropout layer.

[0013] Optionally, the resource behavior sequence of the sample user is determined according to the historical operation behaviors of the sample user on resources, the first multi-modal feature is determined according to the fusion result of at least two of the image feature, text feature, audio feature, and video feature of the resource, and the second multi-modal feature is determined according to the fusion result of at least two of the image feature, text feature, audio feature, and video feature of the to-be-recommended resource.

[0014] According to a second aspect of the embodiments of the present disclosure, a resource recommendation method is provided. The method includes: determining a second multi-modal feature of a to-be-recommended resource; obtaining target user data of a target user, and determining from the target user data a first user attribute feature of the target user, a second resource identifier sequence corresponding to resources in the resource behavior sequence of the target user, and a second resource multi-modal feature sequence corresponding to the third multi-modal feature of the resources in the resource behavior sequence of the target user, where the resource behavior sequence is determined according to the historical operation behaviors of the target user on resources; obtaining a fourth relationship vector according to the second multi-modal feature sequence, the second multi-modal feature of the to-be-recommended resource, and a multi-modal network; obtaining a fifth relationship vector according to the resource identifier, resource attribute feature, resource statistical feature of the to-be-recommended resource, the first user attribute feature, the second resource identifier sequence, and a first network; obtaining a to-be-recommended scenario identifier, and determining a target user private feature of the target recommended scenario from the target user data according to the scenario network corresponding to the target recommended scenario indicated by the to-be-recommended scenario identifier, and obtaining a sixth vector of the target user private feature; obtaining a recommended resource of the target user in the to-be-recommended scenario according to the fourth relationship vector, the fifth relationship vector, the sixth vector, and a multi-task network; pushing the recommended resource to the to-be-recommended scenario of the client where the target user is located;

[0015] Among them, the multi-modal network, the multi-task network, the first network, and the scenario network corresponding to the target recommendation scenario respectively correspond to the multi-modal network, the multi-task network, the first network, and the scenario network in the target resource recommendation model described in the first aspect above.

[0016] Optionally, the multi-modal network includes a first self-attention model and a first cross-attention model. The obtaining of the fourth relationship vector between the resource to be recommended and the target user according to the second multi-modal feature sequence, the second multi-modal feature of the resource to be recommended, and the first attention model includes: inputting the second multi-modal feature sequence into the first self-attention model to process the second multi-modal feature sequence, and obtaining a sixth output result; inputting the sixth output result and the second multi-modal feature into the first cross-attention model, and obtaining the fourth relationship vector according to the output of the first cross-attention model.

[0017] Optionally, the first network includes a second self-attention model, a second cross-attention model, a shared embedding layer, and a first fully connected layer. The obtaining of the fifth relationship vector according to the resource identifier of the resource to be recommended, the resource attribute features, the resource statistical features, the first user attribute features, the second resource identifier sequence, and the first network includes: mapping the resource identifier of the resource to be recommended, the resource attribute features, the resource statistical features, the first user attribute features, and the second resource identifier sequence to the same feature space through the shared embedding layer; inputting the feature-mapped second resource identifier sequence features into the second self-attention model to process the second resource identifier sequence, and obtaining a seventh output result according to the output of the second self-attention model; inputting the seventh output result and the feature-mapped resource identifier of the resource to be recommended into the second cross-attention model, and obtaining a second candidate relationship vector according to the output of the second cross-attention model; inputting the second candidate relationship vector, the feature-mapped resource attribute features, the resource statistical features, and the first user attribute features of the resource to be recommended into the first fully connected layer, and obtaining the fifth relationship vector according to the output of the first fully connected layer.

[0018] Optionally, the scenario network includes a feature mask layer, a second fully connected layer, a batch normalization layer, a third fully connected layer, and a dropout layer. For the scenario network corresponding to the target recommended scenario indicated by the to-be-recommended scenario identifier, determining the target user private feature of the target recommended scenario from the target user data and obtaining the sixth vector of the target user private feature includes: inputting the full user feature corresponding to the target user data into the feature mask layer of the scenario network corresponding to the target recommended scenario, and determining the target user private feature of the target recommended scenario according to the output of the feature mask layer; sequentially inputting the target user private feature into the second fully connected layer, the batch normalization layer, and the third fully connected layer to obtain an eighth output result; adding the ninth output result of the second fully connected layer and the eighth output result to obtain a tenth output result; inputting the tenth output result into the dropout layer, and obtaining the sixth vector according to the output of the dropout layer.

[0019] Optionally, the multi-task network includes a first multi-task network. Obtaining the recommended resource of the target user in the to-be-recommended scenario according to the fourth relationship vector, the fifth relationship vector, the sixth vector, and the multi-task network includes: obtaining a second fusion vector of the to-be-recommended scenario according to the fusion result of the fourth relationship vector, the fifth relationship vector, and the sixth vector; inputting the second fusion vector into each task network related to the to-be-recommended scenario in the first multi-task network to obtain the output of each task network; and obtaining the recommended resource of the to-be-recommended scenario according to the fusion result of the outputs of each task network.

[0020] Optionally, determining the second multi-modal feature of the to-be-recommended resource includes: obtaining at least two of the image feature, text feature, audio feature, and video feature of the to-be-recommended resource; and determining the second multi-modal feature of the to-be-recommended resource according to the fusion result of the at least two features.

[0021] According to a third aspect of the embodiments of the present disclosure, there is provided a training apparatus for a resource recommendation model. The resource recommendation model includes a multi-modal network, a multi-task network, a first network, and a plurality of scenario networks corresponding to each of a plurality of recommendation scenarios. The apparatus includes: a full-sample feature determination module configured to obtain full-sample user data of a plurality of recommendation scenarios and full-task labels of the plurality of recommendation scenarios, and determine full-sample features of the full-sample user data, where the full-sample features include a first resource identifier sequence corresponding to resources in a resource behavior sequence of a sample user, a first resource multi-modal feature sequence corresponding to first multi-modal features of resources in the resource behavior sequence of the sample user, attribute features of the sample user, and private features of each recommendation scenario; a first relationship vector output module configured to obtain a first relationship vector according to the first multi-modal feature sequence, second multi-modal features of a resource to be recommended, and the multi-modal network; a second relationship vector output module configured to obtain a second relationship vector according to a resource identifier of the resource to be recommended, resource attribute features, resource statistical features, the attribute features of the sample user, the first resource identifier sequence, and the first network; a third vector output module configured to obtain a third vector of private features of each recommendation scenario according to the scenario network corresponding to each recommendation scenario; a scenario task prediction module configured to obtain a task prediction result of each recommendation scenario according to the first relationship vector, the second relationship vector, the third vector, and the multi-task network; and an iterative training module configured to iteratively train the resource recommendation model based on the task prediction result and the full-task labels to obtain a target resource recommendation model for resource recommendation.

[0022] According to a fourth aspect of the embodiments of the present disclosure, a resource recommendation device is provided. The device includes: a multi-modal feature determination module configured to determine a second multi-modal feature of a resource to be recommended; a target user data processing module configured to obtain target user data of a target user, and determine from the target user data a first user attribute feature of the target user, a second resource identifier sequence corresponding to resources in a resource behavior sequence of the target user, and a second resource multi-modal feature sequence corresponding to a third multi-modal feature of resources in the resource behavior sequence of the target user, where the resource behavior sequence is determined according to historical operation behaviors of the target user on resources; a fourth relationship vector output module configured to obtain a fourth relationship vector based on the second multi-modal feature sequence, the second multi-modal feature of the resource to be recommended, and a multi-modal network; a fifth relationship vector output module configured to obtain a fifth relationship vector based on the resource identifier, resource attribute features, resource statistical features of the resource to be recommended, the first user attribute feature, the second resource identifier sequence, and a first network; a sixth vector output module configured to obtain a target recommendation scenario identifier, and determine from the target user data a target user private feature of the target recommendation scenario according to a scenario network corresponding to the target recommendation scenario indicated by the target recommendation scenario identifier, and obtain a sixth vector of the target user private feature; a recommended resource determination module configured to obtain a recommended resource of the target user in the target recommendation scenario based on the fourth relationship vector, the fifth relationship vector, the sixth vector, and a multi-task network; a resource recommendation module configured to push the recommended resource to the target recommendation scenario of the client where the target user is located; where the multi-modal network, the multi-task network, the first network, and the scenario network corresponding to the target recommendation scenario respectively include the multi-modal network, the multi-task network, the first network, and the scenario network corresponding to the target resource recommendation model in the first aspect described above.

[0023] According to a fifth aspect of the present disclosure, a computer program product including instructions is provided. When it runs on a computer, it causes the computer to execute the steps of the training method of the resource recommendation model as described in the first aspect and / or the resource recommendation method as described in the second aspect.

[0024] According to a sixth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements the training method of the resource recommendation model as described in the first aspect and / or the resource recommendation method as described in the second aspect in the above embodiments.

[0025] According to a seventh aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the training method of the resource recommendation model described in the first aspect of the above embodiments and / or the resource recommendation method described in the second aspect.

[0026] According to the training method of the resource recommendation model, the resource recommendation method, the training device of the resource recommendation model, the resource recommendation device, the computer-readable storage medium, the computer program product, and the electronic device of the embodiments of the present disclosure, the full-sample user data is used to determine the full-sample features, and then, based on the full-sample features, the multi-modal network, the public feature extraction network, and the private feature extraction network are jointly used for multi-scenario and multi-task training to obtain a target resource recommendation model for resource recommendation. On the one hand, through the multi-modal network, the present disclosure can extract the multi-modal feature representations of resources, thereby uniformly expressing different resources through multi-modal information, improving the generalization ability of the resource recommendation model, avoiding the training process of the resource recommendation model being dominated by popular resources and ignoring the learning of unpopular resources, and thus improving the accuracy of resource recommendation. On the other hand, through the joint training of the multi-modal network, the public feature extraction network, and the private feature extraction network, the present disclosure can obtain a target resource recommendation model capable of multi-task and multi-scenario recommendation, without the need to separately train the resource recommendation model for each scenario, improving the training efficiency of the resource recommendation model in the case of multi-scenario and multi-task. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of example and not limitation, wherein:

[0028] Figure 1 A flowchart showing a method for training a resource recommendation model in an exemplary embodiment of the present disclosure;

[0029] Figure 2 A flowchart showing a method for obtaining multi-modal features through a large model in an exemplary embodiment of the present disclosure;

[0030] Figure 3 A flowchart showing a method for obtaining a first relationship vector in an exemplary embodiment of the present disclosure;

[0031] Figure 4 A flowchart showing a method for obtaining a second relationship vector in an exemplary embodiment of the present disclosure;

[0032] Figure 5Schematic flowchart of a method for obtaining a third vector in an exemplary embodiment of the present disclosure;

[0033] Figure 6 Schematic flowchart of a method for obtaining a task prediction result for each scenario in an exemplary embodiment of the present disclosure;

[0034] Figure 7 Schematic flowchart of a method for obtaining a target resource recommendation model in an exemplary embodiment of the present disclosure;

[0035] Figure 8 Schematic structural diagram of a resource recommendation model during training in an exemplary embodiment of the present disclosure;

[0036] Figure 9 Schematic flowchart of a resource recommendation method in an exemplary embodiment of the present disclosure;

[0037] Figure 10 Schematic composition diagram of a training device for a resource recommendation model in an exemplary embodiment of the present disclosure;

[0038] Figure 11 Schematic composition diagram of a resource recommendation device in an exemplary embodiment of the present disclosure;

[0039] Figure 12 Schematic structural diagram of an electronic device in an exemplary embodiment of the present disclosure.

[0040] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed embodiments

[0041] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to be able to fully convey the scope of the present invention to those skilled in the art.

[0042] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, medium, method, or computer program product. Therefore, the present invention can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0043] According to the embodiments of the present invention, there are provided a training method for a resource recommendation model, a resource recommendation method, a training device for a resource recommendation model, a resource recommendation device, a computer-readable storage medium, a computer program product, and an electronic device.

[0044] In this document, any number of elements in the drawings is for illustration and not for limitation, and any naming is for distinction only and does not have any limiting meaning.

[0045] The principles and spirit of the present invention are described in detail below with reference to several representative embodiments of the present invention. Summary of the Invention

[0047] The inventors of the present disclosure have discovered that resource recommendation models in related technologies have a problem of low recommendation accuracy for unpopular resources when recommending resources.

[0048] In view of the above content, the basic idea of the present disclosure is to provide a training method for a resource recommendation model, a resource recommendation method, a training device for a resource recommendation model, a resource recommendation device, a computer-readable storage medium, a computer program product, and an electronic device, and determine the full sample features through full sample user data, so as to jointly perform multi-scenario and multi-task training based on the full sample features and a multimodal network, a public feature extraction network, and a private feature extraction network to obtain a target resource recommendation model for resource recommendation. On the one hand, the present disclosure can extract multimodal feature representations of resources through a multimodal network, thereby uniformly expressing different resources through multimodal information, improving the generalization ability of the resource recommendation model, avoiding the training process of the resource recommendation model being dominated by popular resources and ignoring the learning of unpopular resources, and thus improving the accuracy of resource recommendation; on the other hand, the present disclosure can obtain a target resource recommendation model capable of multi-task and multi-scenario recommendation through joint training of a multimodal network, a public feature extraction network, and a private feature extraction network, without the need to train a resource recommendation model separately for each scenario, thereby improving the training efficiency of the resource recommendation model in multi-scenario and multi-task situations.

[0049] After introducing the basic principles of the present invention, various non-limiting embodiments of the present invention are described in detail below.

[0050] Overview of Application Scenarios

[0051] It should be noted that the following application scenarios are only provided to facilitate understanding of the spirit and principles of the present invention, and the embodiments of the present invention are not limited in this respect. On the contrary, the embodiments of the present invention can be applied to any applicable scenario.

[0052] In an exemplary application scenario, for different recommendation scenarios in the application client, the user's full features can be obtained, and the user's full sample features can be input into the target resource recommendation model trained according to the method of the present invention. The target resource recommendation model obtains the private scene features and public scene features of the scene to be recommended and the multimodal features of the resource by processing the full features, and recommends resources for each different recommendation scenario based on the private scene features, public scene features and multimodal features of the resource.

[0053] Exemplary Method

[0054] In an exemplary embodiment, the resource recommendation model in the present disclosure includes a multimodal network, a multitask network, a first network, and multiple scenario networks corresponding to each of the multiple recommendation scenarios. Based on this, exemplary, Figure 1 A flow chart showing a method for training a resource recommendation model in an exemplary embodiment of the present disclosure is shown. Figure 1 , the method comprising:

[0055] Step S110: Obtain full sample user data of multiple recommendation scenarios and full task labels of the multiple recommendation scenarios, and determine full sample features of the full sample user data, wherein the full sample features include a first resource identification sequence corresponding to resources in the resource behavior sequence of the sample user, a first resource multimodal feature sequence corresponding to the first multimodal feature of the resource in the resource behavior sequence of the sample user, attribute features of the sample user, and private features of each recommendation scenario;

[0056] Step S120: obtaining a first relationship vector based on the first multimodal feature sequence, the second multimodal feature of the resource to be recommended, and the multimodal network;

[0057] Step S130: obtaining a second relationship vector based on the resource identifier, resource attribute characteristics, resource statistical characteristics of the resource to be recommended, attribute characteristics of the sample user, the first resource identifier sequence, and the first network;

[0058] Step S140, obtaining a third vector of private features of each recommended scene according to the scene network corresponding to each recommended scene;

[0059] Step S150, obtaining a task prediction result for each recommendation scenario based on the first relationship vector, the second relationship vector, the third vector, and the multi-task network;

[0060] Step S160 , iteratively training the resource recommendation model based on the task prediction result and the full set of task labels to obtain a target resource recommendation model for resource recommendation.

[0061] Below, the specific implementation method of "step S110, obtaining full sample user data of multiple recommendation scenarios and full task labels of the multiple recommendation scenarios, and determining full sample features of the full sample user data" is first described in detail.

[0062] Resources refer to information that can be accessed through a network application, including text, audio, video, and images. For example, if the network application is a video application, resources can include various types of videos, such as comedy videos, fashion videos, and entertainment videos. If the network application is an audio application, resources can include various types of audio, such as songs, podcasts, and audiobooks.

[0063] In an optional embodiment, the full sample user data includes the first historical consumption behavior label data of multiple users on resources in multiple recommendation scenarios as the consumed resources. Taking the network application as a music application as an example, any user's consumption behavior on any resource in the music application (such as songs, audiobooks, podcasts, etc.) constitutes a data record, and the data record is a sample user data. For example, if user A purchased audiobook B at 22:22:22 on February 22, 2022, then the data record is a sample user data, and the sample label of the data record is audiobook B.

[0064] In an exemplary embodiment, the full sample features include private features and public features of each recommendation scenario. For example, the full sample features may include: a first resource identification sequence corresponding to resources in the sample user's resource behavior sequence, a first resource multimodal feature sequence corresponding to the first multimodal feature of resources in the sample user's resource behavior sequence, attribute features of the sample user, and private features of each recommendation scenario.

[0065] The resource behavior sequence of the sample user is determined based on the historical operation behavior of the sample user on the resource, and the first multimodal feature is determined based on the fusion result of at least two features of the resource's image feature, text feature, audio feature, and video feature.

[0066] In an exemplary embodiment, image features, text features, audio features, and video features can be combined to form a multimodal feature. Alternatively, the image features, text features, audio features, and video features can be weighted according to their weights, and the weighted sum can be used as the multimodal feature. The weight of each feature is determined based on the type of resource.

[0067] For example, for resources of the video type, the weight of video features is the highest, followed by the weight of audio features, then image features, and finally text features; for resources of the audio type, the weight of audio features is the largest, followed by the weight of text features, then image features, and finally video features; or for resources of the video type, the weight of video features is the largest, and the weights of audio features, image features, and text features are the same and less than the weight of video features. In other words, the weight of features of the same type as the resource type is the largest, so as to obtain the weights of different types of features, and then perform weighted summation on multiple different types of features according to the weights to obtain the fused multi-modal features.

[0068] For example, the resource behavior sequence of a sample user can be determined based on the resources clicked, browsed, and purchased by the sample user in the sample data before consuming the resources in the sample. For example, before the user in the sample consumes the resources in the sample, the operation behaviors on the resources are successively: click on Resource 1, browse Resource 2, purchase Resource 3, click on Resource 4, purchase Resource 5, then the resource behavior sequence of the sample user is (Resource 1, Resource 2, Resource 3, Resource 4, Resource 5).

[0069] In an exemplary implementation, the text features of a resource can be extracted from the text information related to the resource, such as the title and description information of the resource, the image features of the resource can be extracted from the displayed image of the resource, such as the cover of the resource, the audio features of the resource can be extracted from the audio information of the resource, and the video features of the resource can be extracted from the video information of the resource.

[0070] Taking a song as an example, the text features of the song can be extracted from the text description information related to the song, such as the lyrics, song name, and song style tags; the image features of the song can be extracted from the image information related to the song, such as the cover of the song; the audio features can be obtained according to the Mel-spectrogram of the song; the video features of the song can be obtained according to the MV (Music Video) of the song.

[0071] In an exemplary implementation, the multi-modal features of a resource can be extracted through a large model. Refer to Figure 2As shown, resource information can be pre-processed first. For example, resource information can be extracted and assembled into a format suitable for the large model, preparing for input into the large model to extract multimodal features. For example, if the resource is a song, information such as the playlist title, playlist cover image, song album cover image, and song audio can be extracted and converted into the input format required by the large model, thus pre-processing the song resource. After pre-processing, the pre-processed resource information can be input into the large model. For example, pre-processed text information such as title and keywords can be input into the large language model to obtain a text information representation vector; pre-processed image and video information can be input into the large visual model to obtain image and video information representation vectors; and a convolutional neural network can be used to extract the song's mel-spectrogram features to obtain audio features. The information representation vector output by the large model can then be post-processed to obtain the resource's multimodal features. For example, the multimodal information representation vector output by the large model can be compressed and reduced in dimension to obtain the final multimodal feature representation vector.

[0072] For example, after obtaining the resource behavior sequence of the sample user, the above Figure 2 The process shown processes each resource in the resource behavior sequence of the sample user to obtain the multimodal features of each resource, thereby obtaining a first multimodal feature sequence based on a sequence composed of the multimodal features of each resource in the resource behavior sequence of the sample user.

[0073] By extracting multimodal features through large models, we can enhance the feature representation of resources and improve the recommendation accuracy of resource recommendation models.

[0074] The specific implementation of “step S120, obtaining a first relationship vector according to the first multimodal feature sequence, the second multimodal feature of the resource to be recommended, and the multimodal network” is described in detail below.

[0075] In an exemplary embodiment, the resources to be recommended may include all resources in the resource library. The second multimodal feature is determined based on a fusion result of at least two features of the image feature, text feature, audio feature, and video feature of the resource to be recommended.

[0076] The method for determining the image features, text features, audio features, and video features of the resource to be recommended may refer to the description of the relevant contents in the above-mentioned step S110, which will not be repeated here.

[0077] In an exemplary embodiment, the multimodal network includes a first self-attention model and a first cross-attention model. Based on this, exemplary, Figure 3A schematic flowchart showing a method for obtaining a first relationship vector in an exemplary embodiment of the present disclosure. Refer to Figure 3 The method may include steps S310 to S320.

[0078] Wherein:

[0079] In step S310, the first multi-modal feature sequence is input into the first self-attention model to process the first multi-modal feature sequence, and a first output result is obtained.

[0080] For example, the first multi-modal feature sequence and the second multi-modal feature can be first transformed into the same feature vector space through a multi-modal feature transformation structure. For example, the first multi-modal feature sequence and the second multi-modal feature can be mapped into the same feature vector space through a feature embedding layer. Then, the mapped first multi-modal feature is input into the first self-attention model, and the first multi-modal feature sequence is processed through the Transformer (transformation) of the first self-attention model, and a first output result is obtained according to the processing result of the first self-attention model.

[0081] In step S320, the first output result and the second multi-modal feature are input into the first cross-attention model, and the first relationship vector is obtained according to the output of the first cross-attention model.

[0082] For example, the first output result and the second multi-modal feature can be input into the first cross-attention model, and cross-attention processing is performed on the first output result and the second multi-modal feature through the first cross-attention model to obtain the first relationship vector.

[0083] Through the above steps S310 to S320, the long-distance dependence relationship between the multi-modalities of the resource to be recommended and the user behavior resources can be captured through the multi-modal network, and rich context information can be generated, thereby improving the accuracy of resource recommendation.

[0084] Next, a detailed description of the specific implementation manner of "step S130, obtaining a second relationship vector according to the resource identifier, resource attribute features, resource statistical features of the resource to be recommended, the attribute features of the sample user, the first resource identifier sequence, and the first network" will be given.

[0085] In an exemplary implementation manner, the first network can be understood as a public network, that is, common features of multiple recommendation scenarios can be extracted through the first network, and the extracted common features can be used in each recommendation scenario. In this way, for the common features, they only need to be processed once in multiple recommendation scenarios, rather than being processed separately for each recommendation scenario, thereby improving the training efficiency and recommendation efficiency of the resource recommendation model for multiple recommendation scenarios.

[0086] In an exemplary embodiment, the first network includes a second self-attention model, a second cross-attention model, a shared embedding layer, and a first fully-connected layer.

[0087] Next, in combination with Figure 4 The specific implementation of step S130 will be described. Exemplarily, Figure 4 The flowchart shows a method for obtaining a second relationship vector in an exemplary embodiment of the present disclosure. Referring to Figure 4 , this method may include steps S410 to S440. Among them:

[0088] In step S410, the resource identifier, resource attribute features, resource statistical features of the to-be-recommended resource, the attribute features of the sample user, and the first resource identifier sequence are mapped to the same feature space through the shared embedding layer.

[0089] In an exemplary embodiment, the resource attribute features may include the publisher identifier of the resource, the publication time of the resource, the type of the resource, the tags of the resource, etc., the resource statistical features may include the click volume of the resource, the purchase volume, the purchase population distribution characteristics, the exposure volume, etc., and the attribute features of the sample user may include information such as the age, gender, occupation, and location of the sample user.

[0090] In an exemplary embodiment, the shared embedding layer can be understood as an Embedding layer, and through Embedding, different input features can be mapped to the same feature space.

[0091] In step S420, the first resource identifier sequence features after feature mapping are input into the second self-attention model to process the first resource identifier sequence, and a second output result is obtained according to the output of the second self-attention model.

[0092] For example, the first resource identifier sequence features after feature mapping can be input into the second self-attention model, and the second self-attention model performs self-attention processing on the first resource identifier sequence, and a second output result is obtained according to the processing result.

[0093] In step S430, the second output result and the resource identifier of the to-be-recommended resource after feature mapping are input into the second cross-attention model, and a first candidate relationship vector is obtained according to the output of the second cross-attention model.

[0094] For example, the first resource identifier sequence features after self-attention processing and the resource identifier of the to-be-recommended resource can be input into the second cross-attention model, and a first candidate relationship vector is obtained according to the output of the second cross-attention model.

[0095] In step S440, the first candidate relationship vector, the resource attribute features of the resource to be recommended after feature mapping, the resource statistical features, and the attribute features of the sample user are input into the first fully connected layer, and the second relationship vector is obtained according to the output of the first fully connected layer.

[0096] Exemplarily, after obtaining the first candidate relationship vector, the resource attribute features, the resource statistical features of the resource to be recommended obtained after mapping in step S510, the attribute features of the sample user, and the first candidate relationship vector obtained in step S430 can be input into the first fully connected layer. The first fully connected layer fuses these features to learn the non-linear mapping relationship between the features, thereby extracting more complex and higher-order feature representations.

[0097] For example, the bottom layer of the first network is a Shared Embedding Layer. The input features of this layer can include the resource features (Item Feature) corresponding to the resource to be recommended, the user profile (User Profile), the first resource identifier sequence (User Seq), and the resource identifier of the resource to be recommended (Target Item). Among them, the resource features include resource attribute features and resource statistical features. These features first pass through the shared embedding layer, and the shared embedding layer maps these features into a unified vector space, enabling different types of features to interact and learn in the same semantic space. After feature mapping, the mapped first resource identifier sequence feature and the resource identifier of the resource to be recommended will enter the Attention Block of the Transformer for correlation calculation. The Transformer can capture the long-range dependence relationship between the resource to be recommended and the user behavior through the self-attention mechanism and the cross-attention mechanism, and generate rich context information. Then, the output of the attention block and the other features after feature mapping are input into the fully connected layer for processing. The fully connected layer network learns the non-linear relationship between the features, and finally, the common feature representation vector, that is, the second relationship vector, is obtained according to the output of the fully connected layer.

[0098] The second relationship vector fully fuses information from multiple aspects such as the user portrait, resource features, and user behavior features. It can fully express the long-range dependence relationship between the user and the resource, providing strong support for improving the accuracy of resource recommendation.

[0099] Next, a detailed description will be given of the specific implementation manner of "step S140, obtaining the third vector of the private features of each recommendation scenario according to the scenario network corresponding to each recommendation scenario".

[0100] In an exemplary embodiment, the resource recommendation model of the present disclosure can be used to recommend resources for multiple recommendation scenarios. To achieve resource recommendation for multiple scenarios, a hierarchical system is adopted to separately process the common features and private features of multiple scenarios through the first network and the scenario network of each recommendation scenario, reducing the task coupling between multiple tasks and multiple scenarios. It can not only retain the private features of different recommendation scenarios but also integrate common features to the greatest extent, enabling a recommendation model to serve multiple recommendation scenarios simultaneously.

[0101] In an exemplary embodiment, the scenario network includes a feature masking layer, a second fully connected layer, a batch normalization layer, a third fully connected layer, and a dropout layer. The following further describes with reference to Figure 5 the specific implementation manner of step S140. Exemplarily, Figure 5 FIG. shows a flowchart of a method for obtaining a third vector in an exemplary embodiment of the present disclosure. Refer to Figure 5 , the method may include steps S510 to S540. Among them:

[0102] In step S510, for each recommendation scenario, input the full amount of sample features into the feature masking layer of the scenario network corresponding to the recommendation scenario, so that the feature masking layer determines the private features of the recommendation scenario from the full amount of sample features.

[0103] In an exemplary embodiment, the first layer of each scenario network is a feature masking layer. The feature masking layer can pre-configure the values of the input channels of the private features belonging to the recommendation scenario in the full amount of features to 1, and configure the values of the input channels of the full amount of features that do not belong to the private features of the recommendation scenario to 0. The private features belonging to the recommendation scenario are screened out from the full amount of features by means of masking.

[0104] Among them, the private features of each scenario can be custom-determined according to experience or requirements, and this exemplary embodiment does not make special limitations on this.

[0105] In step S520, input the private features of the recommendation scenario into the second fully connected layer, the batch normalization layer, and the third fully connected layer in sequence to obtain a third output result.

[0106] In an exemplary embodiment, the private features of the scenario extracted by the feature masking layer can be input into the second fully connected layer of the scenario network, output 1 is obtained according to the output of the second fully connected layer, then output 1 is input into the batch normalization layer, output 2 is obtained according to the output of the batch normalization layer, and then output 2 is input into the third connection layer, and output 3, that is, the third output result, is obtained according to the output of the third fully connected layer.

[0107] In step S530, the fourth output result of the second fully connected layer and the third output result are added to obtain a fifth output result.

[0108] For example, the output 1 of the second fully connected layer and the third output result, that is, the output 2 mentioned above, can be added to obtain the fifth output result.

[0109] In step S540, the fifth output result is input into the random deactivation layer, and the third vector is obtained according to the output of the random deactivation layer.

[0110] In an exemplary embodiment, the fifth output result is input into a random dropout layer, and the random dropout layer processes the fifth output result to prevent overfitting, thereby obtaining a third vector.

[0111] For example, each recommendation scenario corresponds to a scenario network. Each scenario network is capable of capturing more detailed and discriminative feature information within a specific scenario. By building a scenario network for each scenario, processing the scenario's private features, and then connecting the scenario networks of multiple scenarios in parallel in the resource recommendation model, we can simultaneously achieve better learning of scenario characteristics and avoid interference between scenario parameter learning.

[0112] The scene network structure of each scene is the same, which can include a feature mask layer (Mask Layer), a fully connected layer 1 (FCN), a batch normalization layer (BN), a fully connected layer 2 and a random dropout layer (Dropout), as well as a residual connection layer (Residual connection) connecting two fully connected layers. The feature mask layer can obtain the private features of each scene by filtering features. The fully connected layer and the residual connection layer between them can improve the expression ability of the deep network. The batch normalization layer can solve the problem of feature distribution drift and improve the stability of the model. The random dropout layer can effectively prevent overfitting and improve the robustness of the model. The scene network enables the model to provide rich private feature expressions for the cross-scene recommendation system, improving the accuracy of recommendations and the flexibility of the recommendation system.

[0113] The following describes in detail the specific implementation of "step S150, obtaining a task prediction result for each recommendation scenario based on the first relationship vector, the second relationship vector, the third vector, and the multi-task network."

[0114] In an exemplary embodiment, during the training process, the multi-task network in the present disclosure includes a first multi-task network and a second multi-task network. The first multi-task network and the second multi-task network have the same structure, but their inputs are different. The second multi-task network is used to assist in training the above-mentioned multi-modal network. It is only used in the training phase and will be removed during the actual recommendation phase.

[0115] Next, in conjunction with Figure 6 a further description of the specific implementation manner of step S150 will be given. Exemplarily, Figure 6 FIG. shows a schematic flowchart of a method for obtaining the task prediction result of each scenario in an exemplary embodiment of the present disclosure. Refer to Figure 6 and this method may include steps S610 to S630. Among them:

[0116] In step S610, according to the fusion result of the first relationship vector, the second relationship vector, and the third vector corresponding to the private feature of the recommendation scenario, the first fusion vector of the recommendation scenario is obtained.

[0117] For example, for each recommendation scenario, the first relationship vector, the second relationship vector, and the third vector corresponding to the private feature of this recommendation scenario can be concatenated to obtain the first fusion vector.

[0118] In step S620, the first fusion vector is input into each task network related to the recommendation scenario in the first multi-task network to obtain the first task prediction result of the recommendation scenario.

[0119] For example, the first multi-task network includes task networks corresponding to each scenario task of each recommendation scenario. Each recommendation scenario may include one or more scenario tasks. In other words, each recommendation scenario may include one or more task networks.

[0120] The first fusion vector can be input into each task network related to the recommendation scenario in the first multi-task network, and the first task prediction result of each recommendation scenario can be obtained according to the fusion result of the outputs of each task network.

[0121] In step S630, the first relationship vector is input into each task network of the second multi-task network to obtain the second task prediction result of the recommendation scenario.

[0122] For example, the first relationship vector determined according to the multi-modal features can be input into the fusion result of the outputs of each task network related to the recommendation scenario in the second multi-task network to obtain the second task prediction result of each recommendation scenario.

[0123] For example, the multi-task network may include a single-layer CGC (Cross-Gating Coordination) module, special expert networks corresponding to each task, and two shared expert networks. These components are integrated in the way of MMOE (Multi-gate Mixture-of-Experts) to handle the complexity of multi-task learning. The special expert networks are only used in the corresponding tasks, while the shared expert networks are shared among all tasks. The expert networks are integrated in the way of MMOE, and a gating network is used to determine the contribution degree of each expert network to each task, and the weights of different expert networks can be automatically adjusted according to the characteristics of each task. Suppose there are N tasks and K expert networks, where 2 are shared expert networks and K - 2 are task-related special expert networks, and each task corresponds to a task-related special expert network, that is, K - 2 = N. For the nth task, its output y_n can be determined by the following formula (1):

[0124]

[0125] where f n is the output function specific to the nth task, h k (x) is the output of the expert network, and g nk is the output of the gating network for the Kth expert network of the nth task. K ∈ [1, N + 2] because there are 2 shared expert networks, so there are N + 2 expert networks in total for N tasks.

[0126] Next, a detailed description will be given of the specific implementation manner of "Step S160: Iteratively train the resource recommendation model based on the task prediction result and the full set of task labels to obtain a target resource recommendation model for resource recommendation".

[0127] In an exemplary implementation manner, as described above, in the model training stage, the task prediction result may include the above-mentioned first task prediction result and second task prediction result.

[0128] Next, a detailed description will be given of a specific implementation manner of Step S160. Exemplarily, Figure 7 shows a schematic flowchart of a method for obtaining a target resource recommendation model in an exemplary embodiment of the present disclosure. Referring to Figure 7 , the method may include Step S710 to Step S750. Among them: Figure 7 In Step S710, for each recommendation scenario, determine the task labels corresponding to each recommendation scenario from the full set of task labels.

[0129]

[0130] For example, for each recommendation scenario, the task tags corresponding to each scenario task of the recommendation scenario can be determined from the full set of task tags.

[0131] In step S720, according to the degree of difference between the first task prediction result of each recommendation scenario and the task tag of the recommendation scenario, the first task loss of the recommendation scenario is obtained.

[0132] For example, for each recommendation scenario, the first task loss can be obtained according to the first cross-entropy loss between the first task prediction result of the recommendation scenario and the task tag of the recommendation scenario. Of course, other loss functions can also be used to determine the degree of difference in this step, and this exemplary embodiment does not make special limitations in this regard.

[0133] In step S730, according to the degree of difference between the second task prediction result of each recommendation scenario and the task tag of the recommendation scenario, the second task loss of the recommendation scenario is obtained.

[0134] For example, for each recommendation scenario, the second task loss can be obtained according to the second cross-entropy loss between the second task prediction result of the recommendation scenario and the task tag of the recommendation scenario. Of course, other loss functions can also be used to determine the degree of difference in this step, and this exemplary embodiment does not make special limitations in this regard.

[0135] In step S740, according to the fusion result of the first task loss and the second task loss of the recommendation scenario, the task loss of the recommendation scenario is obtained.

[0136] For example, for each recommendation scenario, the task loss of the recommendation scenario can be obtained according to the sum of the first task loss and the second task loss. Weights can also be configured for the first task loss and the second task loss, and the first task loss and the second task loss are weighted and averaged according to the weights to obtain the task loss of the recommendation scenario. This exemplary embodiment does not make special limitations in this regard. Among them, the magnitude relationship between the weights of the first task loss and the second task loss can be custom-determined according to requirements and experience, and this exemplary embodiment does not make special limitations in this regard. For example, if you want to highlight the effect of multi-modal features, the weight of the second task loss can be configured to be greater than the weight of the first task loss, so as to focus on training the above multi-modal network.

[0137] In step S750, the training loss is obtained according to the fusion result of the task loss of each recommendation scenario, and each network in the resource recommendation model is iteratively trained according to the training loss until a preset condition is met, and a target resource recommendation model is obtained.

[0138] For example, the scenario network and the task network of each recommendation scenario in the present disclosure are jointly trained, so that the final recommendation model can perform well in each recommendation task of each recommendation scenario. Based on this, during the training process, the loss function is determined according to the sum of the task losses of each recommendation scenario. That is, the training loss in step S750 is the sum value of the task losses of each scenario.

[0139] In an exemplary implementation manner, the preset condition may include that the degree of difference is less than a preset value or the number of training times reaches a preset number.

[0140] For example, the network parameters of each network in the resource recommendation model can be adjusted continuously according to the degree of difference in each training process, so that the model converges in the direction of decreasing the degree of difference for iterative training until the degree of difference is less than the preset value or the number of iterations reaches the preset number, and then the training ends. Thus, the target resource recommendation model that can be used for multi-scenario and multi-task recommendation is obtained according to the network parameters of each network at the end of the training.

[0141] Exemplarily, Figure 8 shows a schematic structural diagram of the resource recommendation model during the training process in an exemplary embodiment of the present disclosure. Refer to Figure 8 , the resource recommendation model may include a multi-modal network, a first network, each scenario network, a first multi-task network, and a second multi-task network. Among them, the resource multi-modal features obtained by the large model are mainly used for the multi-modal feature expression of the resource to be recommended and the multi-modal feature expression of the resources in the user resource behavior sequence. Figure 8 In, FCN is the above-mentioned fully connected layer, BN is the above-mentioned batch normalization layer, and Attention Block is the attention module. Target ltemMultimodal represents the multi-modal features of the resource to be recommended, User Seq Multimodal represents the user resource behavior multi-modal feature sequence composed of the multi-modal features of the resources in the user resource behavior sequence, User Profile represents the user attribute features, Item Feature represents the resource features (such as resource attribute features, resource statistical features), User Seq represents the sequence composed of the identifiers of the resources in the user resource behavior sequence, Target Item represents the resource identifier of the resource to be recommended, ScenesPrivate Feature identifies the private features of the recommendation scenario, and Scens Indictor represents the scenario identifier.

[0142] In the present disclosure, through a multi-modal network structure, long-distance dependence relationships between resources to be recommended and user behavior resources can be captured. By combining multi-modal feature information with other information and jointly inputting them into a multi-task network structure for training, the effects of multi-modal features can be fully exerted. Meanwhile, during the training phase, since the convergence speed of multi-modal features is lower than that of resource identification sequence features, by adding a second multi-task network to add a multi-modal auxiliary loss structure, the learning objective of the multi-modal network is made consistent with the modeling objective of the resource recommendation model, enabling the multi-modal network to fully learn the modeling objective of the recommendation model and accelerate convergence, fully exerting the out-of-domain information of multi-modal features, enriching the feature representation ability, and improving the recommendation effect.

[0143] In another exemplary embodiment, the number of training times of the multi-modal network can also be increased so that the multi-modal network can fully learn and converge.

[0144] After obtaining the target resource recommendation model through the above steps S110 to S160, the target resource recommendation model can be used for resource recommendation in multiple scenarios. As mentioned above, the second multi-task network is designed to assist in the training of the multi-modal network and is only used during the training phase. During the actual recommendation phase, only the second multi-task network and other networks are required to implement resource recommendation.

[0145] Next, in conjunction with Figure 9 a method for resource recommendation using the resource recommendation model obtained by the above training method will be described, that is, Figure 9 the multi-modal network, multi-task network, first network, and the scenario network corresponding to the target recommendation scenario mentioned in

[0146] Exemplarily, Figure 9 shows a schematic flowchart of a resource recommendation method in an exemplary embodiment of the present disclosure. Referring to Figure 9 , the method may include steps S910 to S970, where:

[0147] In step S910, the second multi-modal feature of the resource to be recommended is determined.

[0148] Exemplarily, a specific implementation manner of step S910 may include: obtaining at least two of the image feature, text feature, audio feature, and video feature of the resource to be recommended; and determining the second multi-modal feature of the resource to be recommended according to the fusion result of the at least two features.

[0149] Among them, the determination methods of the image features, text features, audio features, and video features and their fusion methods can refer to the relevant content in step S110 above, and will not be elaborated here.

[0150] In step S920, obtain the target user data of the target user, and determine the first user attribute feature of the target user, the second resource identifier sequence corresponding to the resources in the resource behavior sequence of the target user, and the second resource multi-modal feature sequence corresponding to the third multi-modal feature of the resources in the resource behavior sequence of the target user from the target user data.

[0151] In an exemplary implementation, the resource behavior sequence in step S920 is determined according to the historical operation behavior of the target user on the resources. Among them, the specific determination method of the resource behavior sequence can refer to the relevant content in step S110 above, and will not be elaborated here.

[0152] In an exemplary implementation, the target user can include any user for whom resource recommendation is to be performed. The target user data can include user portrait data such as the user's gender and age, and the historical operation data of the user on the resources.

[0153] Exemplarily, the specific implementation method in step S920 can refer to the relevant content in step S110 above, and will not be elaborated here.

[0154] In step S930, according to the second multi-modal feature sequence, the second multi-modal feature of the resource to be recommended, and the multi-modal network, obtain the fourth relationship vector.

[0155] Exemplarily, as described above, the multi-modal network includes a first self-attention model and a first cross-attention model. Based on this, an exemplary implementation of step S930 can include: inputting the second multi-modal feature sequence into the first self-attention model to process the second multi-modal feature sequence, and obtaining a sixth output result; inputting the sixth output result and the second multi-modal feature into the first cross-attention model, and obtaining the fourth relationship vector according to the output of the first cross-attention model.

[0156] Among them, the specific implementation method of step S930 can refer to the relevant implementation methods of step S120 above, and will not be elaborated here.

[0157] In step S940, according to the resource identifier, resource attribute features, resource statistical features of the resource to be recommended, the first user attribute feature, the second resource identifier sequence, and the first network, obtain the fifth relationship vector.

[0158] Exemplarily, as mentioned above, the first network includes a second self-attention model, a second cross-attention model, a shared embedding layer, and a first fully connected layer. Based on this, an exemplary implementation of step S840 may include: mapping the resource identifier, resource attribute characteristics, resource statistical characteristics, the first user attribute characteristics, and the second resource identifier sequence of the resource to be recommended to the same feature space through the shared embedding layer; inputting the second resource identifier sequence characteristics after feature mapping into the second self-attention model to process the second resource identifier sequence, and obtaining a seventh output result according to the output of the second self-attention model; inputting the seventh output result and the resource identifier of the resource to be recommended after feature mapping into the second cross-attention model, and obtaining a second candidate relationship vector according to the output of the second cross-attention model; inputting the second candidate relationship vector, the resource attribute characteristics, resource statistical characteristics, and the first user attribute characteristics of the resource to be recommended after feature mapping into the first fully connected layer, and obtaining the fifth relationship vector according to the output of the first fully connected layer.

[0159] For example, the specific implementation of step S940 may refer to the specific implementation of the above-mentioned step S130, and will not be repeated here.

[0160] In step S950, the identification of the scene to be recommended is obtained, and according to the scene network corresponding to the target recommendation scene indicated by the identification of the scene to be recommended, the target user private features of the target recommendation scene are determined from the target user data, and a sixth vector of the target user private features is obtained.

[0161] For example, the identifier of the scene to be recommended can be determined based on the recommended scene to which the target user belongs, which is not specifically limited in this exemplary embodiment. Each user can belong to one or more recommended scenes. Each recommended scene is configured with a corresponding recommended scene identifier, and each scene network of the target resource recommendation model is also configured with its corresponding recommended scene identifier. In this way, the scene network corresponding to the desired recommended scene can be determined from the target resource recommendation model using the recommended scene identifier.

[0162] Exemplarily, as described above, the scenario network includes a feature mask layer, a second fully connected layer, a batch normalization layer, a third fully connected layer, and a dropout layer. Based on this, an exemplary implementation of step S850 may include: inputting the full user features corresponding to the target user data into the feature mask layer of the scenario network corresponding to the target recommendation scenario, and determining the target user private features of the target recommendation scenario according to the output of the feature mask layer; sequentially inputting the target user private features into the second fully connected layer, the batch normalization layer, and the third fully connected layer to obtain an eighth output result; adding the ninth output result of the second fully connected layer and the eighth output result to obtain a tenth output result; inputting the tenth output result into the dropout layer, and obtaining the sixth vector according to the output of the dropout layer.

[0163] Specifically, the implementation of step S950 may refer to the specific implementation of step S140 above, which will not be elaborated here.

[0164] In step S960, according to the fourth relationship vector, the fifth relationship vector, the sixth vector, and the multi-task network, the recommended resources of the target user in the to-be-recommended scenario are obtained.

[0165] Exemplarily, as described above, in the model recommendation stage, the multi-task network only includes the first multi-task network, that is, the multi-task network in step S960 is the above-mentioned first multi-task network. Based on this, an exemplary implementation of step S960 may include: obtaining a second fusion vector of the to-be-recommended scenario according to the fusion result of the fourth relationship vector, the fifth relationship vector, and the sixth vector; inputting the second fusion vector into each task network related to the to-be-recommended scenario in the first multi-task network to obtain the output of each task network; and obtaining the recommended resources of the to-be-recommended scenario according to the fusion result of the outputs of each task network.

[0166] For example, the above-mentioned fourth relationship vector, fifth relationship vector, and sixth vector can be concatenated to obtain a second fusion vector. Then, the second fusion vector is input into each task network related to the to-be-recommended scenario in the first multi-task network to obtain the output results of each task network, and the output results of each task network are fused to obtain the recommended resources of the target user in the to-be-recommended scenario.

[0167] In step S970, the recommended resources are pushed to the to-be-recommended scenario of the client where the target user is located.

[0168] For example, after obtaining the recommended resources for the target user in the scenario to be recommended, the recommended resources can be pushed to the corresponding scenario to be recommended on the client where the target user is located. Taking the home page recommendation of a music client as an example of the scenario to be recommended, the recommended resources can be displayed on the home page of the music client.

[0169] In the present disclosure, by applying the capabilities of a large model, multi-modal features are extracted, and the world domain knowledge of the multi-modal features is applied to the recommendation system, enabling the model to better understand the user's intention and improve the recommendation accuracy. Through the multi-modal network, the long-distance dependence relationship between the multi-modalities of the resources to be recommended and the user behavior resources can be captured, enriching the resource representation information, effectively improving the feature information expression of cold-start resources and medium and long-tail resources (i.e., resources with little user interaction data), and enhancing the recommendation effect of cold-start resources and medium and long-tail resources.

[0170] In addition, it should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.

[0171] Exemplary Apparatus

[0172] The exemplary embodiment of the present disclosure further provides a training device for a resource recommendation model, where the resource recommendation model includes a multi-modal network, a multi-task network, a first network, and multiple scenario networks corresponding to each of the multiple scenarios to be recommended. Refer to Figure 10As shown in the figure, the training device 1000 of the resource recommendation model may include the following program modules: a full - sample feature determination module 1010, configured to obtain the full - sample user data of multiple recommendation scenarios and the full - task labels of the multiple recommendation scenarios, and determine the full - sample features of the full - sample user data. The full - sample features include the first resource identifier sequence corresponding to the resources in the resource behavior sequence of the sample user, the first resource multi - modal feature sequence corresponding to the first multi - modal features of the resources in the resource behavior sequence of the sample user, the attribute features of the sample user, and the private features of each recommendation scenario; a first relationship vector output module 1020, configured to obtain a first relationship vector according to the first multi - modal feature sequence, the second multi - modal feature of the resource to be recommended, and the multi - modal network; a second relationship vector output module 1030, configured to obtain a second relationship vector according to the resource identifier of the resource to be recommended, the resource attribute features, the resource statistical features, the attribute features of the sample user, the first resource identifier sequence, and the first network; a third vector output module 1040, configured to obtain a third vector of the private features of each recommendation scenario according to the scenario network corresponding to each recommendation scenario; a scenario task prediction module 1050, configured to obtain the task prediction result of each recommendation scenario according to the first relationship vector, the second relationship vector, the third vector, and the multi - task network; and an iterative training module 1060, configured to iteratively train the resource recommendation model based on the task prediction result and the full - task label to obtain a target resource recommendation model for resource recommendation.

[0173] In an exemplary implementation manner, based on the foregoing embodiments, the multi - task network includes a first multi - task network and a second multi - task network. The structures of the first multi - task network and the second multi - task network are the same. The scenario task prediction module 1050 may be specifically configured to: for each recommendation scenario, perform the following processing procedures to obtain the task prediction result of the recommendation scenario: obtain the first fusion vector of the recommendation scenario according to the fusion result of the first relationship vector, the second relationship vector, and the third vector corresponding to the private features of the recommendation scenario; input the first fusion vector into each task network related to the recommendation scenario in the first multi - task network to obtain the first task prediction result of the recommendation scenario; and input the first relationship vector into each task network of the second multi - task network to obtain the second task prediction result of the recommendation scenario.

[0174] In an exemplary implementation manner, based on the foregoing embodiments, iteratively training the resource recommendation model based on the task prediction result to obtain a target resource recommendation model for resource recommendation includes: for each recommendation scenario, determining task tags corresponding to each recommendation scenario from the full set of task tags; obtaining the first task loss of the recommendation scenario according to the degree of difference between the first task prediction result of each recommendation scenario and the task tags of the recommendation scenario; obtaining the second task loss of the recommendation scenario according to the degree of difference between the second task prediction result of each recommendation scenario and the task tags of the recommendation scenario; obtaining the task loss of the recommendation scenario according to the fusion result of the first task loss and the second task loss of the recommendation scenario; obtaining a training loss according to the fusion result of the task losses of each recommendation scenario, and iteratively training each network in the resource recommendation model according to the training loss until a preset condition is met to obtain a target resource recommendation model.

[0175] In an exemplary implementation manner, based on the foregoing embodiments, the multimodal network includes a first self-attention model and a first cross-attention model. The first relationship vector output module 1020 may be specifically configured to: input the first multimodal feature sequence into the first self-attention model to process the first multimodal feature sequence to obtain a first output result; input the first output result and the second multimodal feature into the first cross-attention model, and obtain the first relationship vector according to the output of the first cross-attention model.

[0176] In an exemplary implementation manner, based on the foregoing embodiments, the first network includes a second self-attention model, a second cross-attention model, a shared embedding layer, and a first fully connected layer. The second relationship vector output module 1030 may be specifically configured to: map the resource identifier of the resource to be recommended, the resource attribute features, the resource statistical features, the attribute features of the sample user, and the first resource identifier sequence to the same feature space through the shared embedding layer; input the feature-mapped first resource identifier sequence features into the second self-attention model to process the first resource identifier sequence, and obtain a second output result according to the output of the second self-attention model; input the second output result and the feature-mapped resource identifier of the resource to be recommended into the second cross-attention model, and obtain a first candidate relationship vector according to the output of the second cross-attention model; input the first candidate relationship vector, the feature-mapped resource attribute features, resource statistical features, and attribute features of the sample user of the resource to be recommended into the first fully connected layer, and obtain the second relationship vector according to the output of the first fully connected layer.

[0177] In an exemplary embodiment, based on the foregoing embodiments, the scenario network includes a feature mask layer, a second fully-connected layer, a batch normalization layer, a third fully-connected layer, and a dropout layer. The above-mentioned third vector output module 1040 may be specifically configured as follows: for each recommended scenario, input the full amount of sample features into the feature mask layer of the scenario network corresponding to the recommended scenario, so that the feature mask layer determines the private features of the recommended scenario from the full amount of sample features; sequentially input the private features of the recommended scenario into the second fully-connected layer, the batch normalization layer, and the third fully-connected layer to obtain a third output result; add the fourth output result of the second fully-connected layer and the third output result to obtain a fifth output result; input the fifth output result into the dropout layer, and obtain the third vector according to the output of the dropout layer.

[0178] In an exemplary embodiment, based on the foregoing embodiments, the resource behavior sequence of the sample user is determined according to the historical operation behavior of the sample user on resources, the first multi-modal feature is determined according to the fusion result of at least two of the image feature, text feature, audio feature, and video feature of the resource, and the second multi-modal feature is determined according to the fusion result of at least two of the image feature, text feature, audio feature, and video feature of the to-be-recommended resource.

[0179] An exemplary embodiment of the present disclosure further provides a resource recommendation device. Refer to Figure 11As shown in the figure, the resource recommendation device 1100 may include the following program modules: a multi-modal feature determination module 1110, configured to determine the second multi-modal features of the resource to be recommended; a target user data processing module 1120, configured to obtain the target user data of the target user, and determine the first user attribute features of the target user, the second resource identifier sequence corresponding to the resources in the resource behavior sequence of the target user, and the second resource multi-modal feature sequence corresponding to the third multi-modal features of the resources in the resource behavior sequence of the target user from the target user data, where the resource behavior sequence is determined according to the historical operation behavior of the target user on the resources; a fourth relationship vector output module 1130, configured to obtain a fourth relationship vector according to the second multi-modal feature sequence, the second multi-modal features of the resource to be recommended, and the multi-modal network; a fifth relationship vector output module 1140, configured to obtain a fifth relationship vector according to the resource identifier, resource attribute features, resource statistical features of the resource to be recommended, the first user attribute features, the second resource identifier sequence, and the first network; a sixth vector output module 1150, configured to obtain the to-be-recommended scenario identifier, determine the target user private features of the target recommended scenario from the target user data according to the scenario network corresponding to the target recommended scenario indicated by the to-be-recommended scenario identifier, and obtain a sixth vector of the target user private features; a recommended resource determination module 1160, configured to obtain the recommended resources of the target user in the to-be-recommended scenario according to the fourth relationship vector, the fifth relationship vector, the sixth vector, and the multi-task network; a resource recommendation module 1170, configured to push the recommended resources to the to-be-recommended scenario of the client where the target user is located; where the multi-modal network, the multi-task network, the first network, and the scenario network corresponding to the target recommended scenario respectively include the multi-modal network, the multi-task network, the first network, and the scenario network corresponding to the target recommended scenario in the above-mentioned target resource recommendation model.

[0180] In an exemplary implementation manner, based on the foregoing embodiment, the multi-modal network includes a first self-attention model and a first cross-attention model. The fourth relationship vector output module 1130 may be specifically configured to: input the second multi-modal feature sequence into the first self-attention model to process the second multi-modal feature sequence, and obtain a sixth output result; input the sixth output result and the second multi-modal features into the first cross-attention model, and obtain the fourth relationship vector according to the output of the first cross-attention model.

[0181] In an exemplary embodiment, based on the foregoing embodiments, the first network includes a second self-attention model, a second cross-attention model, a shared embedding layer, and a first fully-connected layer. The fifth relationship vector output module 1140 may be specifically configured to: map the resource identifier, resource attribute features, resource statistical features, the first user attribute features, and the second resource identifier sequence of the resource to be recommended to the same feature space through the shared embedding layer; input the feature-mapped second resource identifier sequence features into the second self-attention model to process the second resource identifier sequence, and obtain a seventh output result according to the output of the second self-attention model; input the seventh output result and the feature-mapped resource identifier of the resource to be recommended into the second cross-attention model, and obtain a second candidate relationship vector according to the output of the second cross-attention model; input the second candidate relationship vector, the feature-mapped resource attribute features, resource statistical features, and first user attribute features of the resource to be recommended into the first fully-connected layer, and obtain the fifth relationship vector according to the output of the first fully-connected layer.

[0182] In an exemplary embodiment, based on the foregoing embodiments, the scenario network includes a feature masking layer, a second fully-connected layer, a batch normalization layer, a third fully-connected layer, and a dropout layer. The sixth vector output module 1150 may be specifically configured to: input the full user features corresponding to the target user data into the feature masking layer of the scenario network corresponding to the target recommendation scenario, and determine the target user private features of the target recommendation scenario according to the output of the feature masking layer; sequentially input the target user private features into the second fully-connected layer, the batch normalization layer, and the third fully-connected layer to obtain an eighth output result; add the ninth output result of the second fully-connected layer and the eighth output result to obtain a tenth output result; input the tenth output result into the dropout layer, and obtain the sixth vector according to the output of the dropout layer.

[0183] In an exemplary embodiment, based on the foregoing embodiments, the multi-task network includes a first multi-task network. The recommended resource determination module 1160 may be specifically configured to: obtain a second fusion vector of the recommendation scenario to be recommended according to the fusion result of the fourth relationship vector, the fifth relationship vector, and the sixth vector; input the second fusion vector into each task network related to the recommendation scenario to be recommended in the first multi-task network to obtain the output of each task network; and obtain the recommended resource of the recommendation scenario to be recommended according to the fusion result of the outputs of each task network.

[0184] In an exemplary embodiment, based on the foregoing embodiments, the determination of the second multi-modal feature of the resource to be recommended includes: obtaining at least two of the image feature, text feature, audio feature, and video feature of the resource to be recommended; and determining the second multi-modal feature of the resource to be recommended according to the fusion result of the at least two features.

[0185] The specific details of each part in the above device have been described in detail in the corresponding method part of the above embodiments. For the details not disclosed, please refer to the embodiments in the method part above, and thus will not be elaborated herein.

[0186] Exemplary Storage Medium

[0187] The storage medium of the exemplary embodiment of the present invention will be described below.

[0188] In this exemplary embodiment, the above method can be implemented by a program product. For example, a portable compact disc read-only memory (CD-ROM) can be used and includes program code, and can be run on a device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0189] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0190] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0191] The program code contained on a readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.

[0192] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0193] Exemplary Computer Program Product

[0194] Exemplary embodiments of the present disclosure also provide a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the above-described method for training a resource recommendation model.

[0195] In one embodiment, the computer program product can be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium can be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, etc. signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory, hard disk drive (HDD), solid state drive (SSD), etc. Exemplarily, the computer program product can be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, Nand Flash, etc.

[0196] In one embodiment, the computer program product can be an intangible product containing a computer program. Exemplarily, the computer program product can be implemented as a virtual digital product, such as an executable file storing the computer program, an installation package, and other digital files.

[0197] The code of a computer program can be written in one or more programming languages. Examples of programming languages include C, Java, C++, Python, etc. The program code can be executed entirely on the user's computing device, or partially on the user's computing device, or as an independent software package, or partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or it can be connected to an external computing device (e.g., through an Internet connection provided by an operator).

[0198] A computer program can be carried or transmitted by signals such as electricity, magnetism, light, electromagnetic, infrared, etc. An electronic device can convert the signal carrying the computer program into a digital signal and then run the computer program. When the computer program runs on an electronic device, its code is used to cause the electronic device to execute (more specifically, to cause the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure. For example, it can execute the training method of the above-mentioned resource recommendation model, which includes the following steps: obtaining the full amount of sample user data of multiple recommendation scenarios and the full amount of task labels of the multiple recommendation scenarios, determining the full amount of sample features of the full amount of sample user data, where the full amount of sample features includes the first resource identifier sequence corresponding to the resources in the resource behavior sequence of the sample user, the first resource multimodal feature sequence corresponding to the first multimodal features of the resources in the resource behavior sequence of the sample user, the attribute features of the sample user, and the private features of each recommendation scenario; obtaining a first relationship vector according to the first multimodal feature sequence, the second multimodal feature of the resource to be recommended, and the multimodal network; obtaining a second relationship vector according to the resource identifier, resource attribute features, resource statistical features of the resource to be recommended, the attribute features of the sample user, the first resource identifier sequence, and the first network; obtaining a third vector of the private features of each recommendation scenario according to the scenario network corresponding to each recommendation scenario; obtaining the task prediction result of each recommendation scenario according to the first relationship vector, the second relationship vector, the third vector, and the multi-task network; and iteratively training the resource recommendation model based on the task prediction result and the full amount of task labels to obtain a target resource recommendation model for resource recommendation.

[0199] Executing the above method steps through a computer program, on the one hand, through a multimodal network, multimodal feature representations of resources can be extracted, so that different resources can be uniformly expressed through multimodal information, improving the generalization ability of the resource recommendation model, avoiding the training process of the resource recommendation model being dominated by popular resources and ignoring the learning of unpopular resources, and thus improving the accuracy of resource recommendation; on the other hand, through the joint training of the multimodal network, the public feature extraction network and the private feature extraction network, a target resource recommendation model capable of performing multi-task and multi-scenario recommendations can be obtained, without separately training a resource recommendation model for each scenario, improving the training efficiency of the resource recommendation model in the case of multi-scenario and multi-task.

[0200] Exemplary Electronic Device

[0201] Reference Figure 12 An electronic device according to an exemplary embodiment of the present disclosure will be described. The electronic device may include a processor and a memory. The memory stores executable instructions of the processor, such as a computer program. The processor executes the method steps of various exemplary embodiments of the present disclosure by executing the executable instructions. In addition, the electronic device may further include a display for displaying a graphical user interface.

[0202] Next, with reference to Figure 12 , an electronic device will be described by way of example in the form of a general-purpose computing device. It should be understood that Figure 12 The electronic device 1200 shown is merely an example and should not impose limitations on the functions and usage scope of the embodiments of the present disclosure.

[0203] As Figure 12 shown, the electronic device 1200 may include: a processor 1210, a memory 1220, a bus 1230, an I / O (input / output) interface 1240, a network adapter 1250, and a display 1260.

[0204] The memory 1220 may include volatile memory, such as RAM 1221 and a cache unit 1222, and may also include non-volatile memory, such as ROM 1223. The memory 1220 may further include one or more program modules 1224. Such program modules 1224 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. For example, the program module 1224 may include each module in the above device.

[0205] The processor 1210 may include one or more processing units, for example: the processor 1210 may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor and / or an NPU (Neural-Network Processing Unit), etc.

[0206] The processor 1210 can be used to execute the executable instructions stored in the memory 1220, such as the training method of the resource recommendation model described above, which includes the following steps: obtaining full sample user data of multiple recommendation scenarios and full task labels of the multiple recommendation scenarios, determining full sample features of the full sample user data, the full sample features including a first resource identification sequence corresponding to a resource in the resource behavior sequence of the sample user, a first resource multimodal feature sequence corresponding to a first multimodal feature of a resource in the resource behavior sequence of the sample user, attribute features of the sample user, and private features of each recommendation scenario; and determining a resource recommendation model based on the first multimodal feature sequence, the resource to be recommended, and the resource identification sequence. A first relationship vector is obtained based on the second multimodal features and multimodal network of the resource; a second relationship vector is obtained based on the resource identifier, resource attribute characteristics, resource statistical characteristics, attribute characteristics of the sample user, the first resource identifier sequence and the first network of the resource to be recommended; a third vector of private characteristics of each recommendation scenario is obtained based on the scenario network corresponding to each recommendation scenario; a task prediction result of each recommendation scenario is obtained based on the first relationship vector, the second relationship vector, the third vector and the multi-task network; the resource recommendation model is iteratively trained based on the task prediction result and the full set of task labels to obtain a target resource recommendation model for resource recommendation.

[0207] The above method is implemented through a computer program. On the one hand, the multimodal feature representation of resources can be extracted through the multimodal network, so that different resources can be uniformly expressed through multimodal information, thereby improving the generalization ability of the resource recommendation model, avoiding the training process of the resource recommendation model being dominated by popular resources and ignoring the learning of unpopular resources, and thus improving the accuracy of resource recommendation; on the other hand, through joint training of the multimodal network, the public feature extraction network and the private feature extraction network, a target resource recommendation model capable of multi-task and multi-scenario recommendation can be obtained, without the need to train the resource recommendation model separately for each scenario, thereby improving the training efficiency of the resource recommendation model in multi-scenario and multi-task situations.

[0208] The bus 1230 is used to implement connections between different components of the electronic device 1200, and may include a data bus, an address bus, and a control bus.

[0209] The electronic device 1200 can communicate with one or more external devices 1300 (such as a keyboard, a mouse, an external controller, etc.) through the I / O interface 1240.

[0210] The electronic device 1200 can communicate with one or more networks through the network adapter 1250. For example, the network adapter 1250 can provide mobile communication solutions such as 3G / 4G / 5G, or wireless communication solutions such as a wireless local area network, Bluetooth, near field communication, etc. The network adapter 1250 can communicate with other modules of the electronic device 1200 through the bus 1230.

[0211] The electronic device 1200 can display a graphical user interface through the display 1260, such as a graphical user interface for displaying the resources recommended to the user as described above.

[0212] Although Figure 12 not shown in the figure, other hardware and / or software modules can also be provided in the electronic device 1200, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0213] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0214] As can be seen from the above, the technical solution of the present disclosure can be implemented as a method, a device, a computer program product, a computer-readable storage medium, an electronic device, etc. Those skilled in the art can understand that various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, such as can be respectively referred to as "circuit", "module", or "device".

[0215] It should be understood that the present disclosure is not limited to the specific method steps or structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. Based on the specific embodiments provided by the present disclosure, those skilled in the art will readily think of other embodiments. Therefore, the specific embodiments provided by the present disclosure are only exemplary, and the scope and spirit of the present disclosure are pointed out by the claims, and should cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure.

Claims

1. A training method for a resource recommendation model, characterized in that, The resource recommendation model includes a multi-modal network, a multi-task network, a first network, and multiple scenario networks corresponding to each of multiple recommendation scenarios. The method includes: Obtaining the full-scale sample user data of multiple recommendation scenarios and the full-scale task labels of the multiple recommendation scenarios, and determining the full-scale sample features of the full-scale sample user data. The full-scale sample features include a first resource identifier sequence corresponding to the resources in the resource behavior sequence of the sample user, a first resource multi-modal feature sequence corresponding to the first multi-modal features of the resources in the resource behavior sequence of the sample user, the attribute features of the sample user, and the private features of each recommendation scenario; Obtaining a first relationship vector according to the first multi-modal feature sequence, the second multi-modal feature of the resource to be recommended, and the multi-modal network; Obtaining a second relationship vector according to the resource identifier of the resource to be recommended, the resource attribute features, the resource statistical features, the attribute features of the sample user, the first resource identifier sequence, and the first network; Obtaining a third vector of the private features of each recommendation scenario according to the scenario network corresponding to each recommendation scenario; Obtaining the task prediction results of each recommendation scenario according to the first relationship vector, the second relationship vector, the third vector, and the multi-task network; Performing iterative training on the resource recommendation model based on the task prediction results and the full-scale task labels to obtain a target resource recommendation model for resource recommendation.

2. The method according to claim 1, wherein The multi-task network includes a first multi-task network and a second multi-task network. The structures of the first multi-task network and the second multi-task network are the same. The obtaining the task prediction results of each recommendation scenario according to the first relationship vector, the second relationship vector, the third vector, and the multi-task network includes: For each recommendation scenario, performing the following processing procedure to obtain the task prediction result of the recommendation scenario: Obtaining a first fusion vector of the recommendation scenario according to the fusion result of the first relationship vector, the second relationship vector, and the third vector corresponding to the private features of the recommendation scenario; Inputting the first fusion vector into each task network related to the recommendation scenario in the first multi-task network to obtain a first task prediction result of the recommendation scenario; Inputting the first relationship vector into each task network of the second multi-task network to obtain a second task prediction result of the recommendation scenario.

3. The method according to claim 2, wherein The performing iterative training on the resource recommendation model based on the task prediction results to obtain a target resource recommendation model for resource recommendation includes: For each recommendation scenario, determining the task label corresponding to each recommendation scenario from the full-scale task labels; Obtaining a first task loss of the recommendation scenario according to the difference degree between the first task prediction result of each recommendation scenario and the task label of the recommendation scenario; Obtaining a second task loss of the recommendation scenario according to the difference degree between the second task prediction result of each recommendation scenario and the task label of the recommendation scenario; Obtain the task loss of the recommendation scenario according to the fusion result of the first task loss and the second task loss of the recommendation scenario; Obtain the training loss according to the fusion result of the task loss of each recommendation scenario, and iteratively train each network in the resource recommendation model according to the training loss until a preset condition is met to obtain the target resource recommendation model.

4. The method according to claim 1, wherein The multimodal network includes a first self-attention model and a first cross-attention model. The obtaining of the first relationship vector according to the first multimodal feature sequence, the second multimodal feature of the resource to be recommended, and the multimodal network includes: Input the first multimodal feature sequence into the first self-attention model to process the first multimodal feature sequence, and obtain a first output result; Input the first output result and the second multimodal feature into the first cross-attention model, and obtain the first relationship vector according to the output of the first cross-attention model.

5. A resource recommendation method, characterized in that, Include: Determine the second multimodal feature of the resource to be recommended; Obtain the target user data of the target user, and determine from the target user data the first user attribute feature of the target user, the second resource identifier sequence corresponding to the resources in the resource behavior sequence of the target user, and the second resource multimodal feature sequence corresponding to the third multimodal feature of the resources in the resource behavior sequence of the target user, where the resource behavior sequence is determined according to the historical operation behavior of the target user on the resources; Obtain a fourth relationship vector according to the second multimodal feature sequence, the second multimodal feature of the resource to be recommended, and the multimodal network; Obtain a fifth relationship vector according to the resource identifier, resource attribute feature, resource statistical feature of the resource to be recommended, the first user attribute feature, the second resource identifier sequence, and the first network; Obtain the to-be-recommended scenario identifier, and determine the target user private feature of the target recommendation scenario from the target user data according to the scenario network corresponding to the target recommendation scenario indicated by the to-be-recommended scenario identifier, and obtain a sixth vector of the target user private feature; Obtain the recommended resource of the target user in the to-be-recommended scenario according to the fourth relationship vector, the fifth relationship vector, the sixth vector, and the multi-task network; Push the recommended resource to the to-be-recommended scenario of the client where the target user is located; Wherein, the multimodal network, the multi-task network, the first network, and the scenario network corresponding to the target recommendation scenario respectively include the multimodal network, multi-task network, first network, and scenario network corresponding to the target recommendation scenario in the target resource recommendation model according to any one of claims 1 to 4.

6. A training device for a resource recommendation model, characterized in that The resource recommendation model includes a multimodal network, a multi-task network, a first network, and a plurality of scenario networks corresponding to each of a plurality of recommendation scenarios. The apparatus includes: The full-sample feature determination module is configured to obtain the full-sample user data of multiple recommendation scenarios and the full-task labels of the multiple recommendation scenarios, and determine the full-sample features of the full-sample user data. The full-sample features include the first resource identifier sequence corresponding to the resources in the resource behavior sequence of the sample user, the first resource multi-modal feature sequence corresponding to the first multi-modal features of the resources in the resource behavior sequence of the sample user, the attribute features of the sample user, and the private features of each recommendation scenario. The first relationship vector output module is configured to obtain a first relationship vector according to the first multi-modal feature sequence, the second multi-modal feature of the resource to be recommended, and the multi-modal network. The second relationship vector output module is configured to obtain a second relationship vector according to the resource identifier of the resource to be recommended, the resource attribute features, the resource statistical features, the attribute features of the sample user, the first resource identifier sequence, and the first network. The third vector output module is configured to obtain a third vector of the private features of each recommendation scenario according to the scenario network corresponding to each recommendation scenario. The scenario task prediction module is configured to obtain the task prediction results of each recommendation scenario according to the first relationship vector, the second relationship vector, the third vector, and the multi-task network. The iterative training module is configured to iteratively train the resource recommendation model based on the task prediction results and the full-task labels to obtain a target resource recommendation model for resource recommendation.

7. A resource recommendation device, characterized in that, It includes: The multi-modal feature determination module is configured to determine the second multi-modal feature of the resource to be recommended. The target user data processing module is configured to obtain the target user data of the target user, and determine the first user attribute features of the target user, the second resource identifier sequence corresponding to the resources in the resource behavior sequence of the target user, and the second resource multi-modal feature sequence corresponding to the third multi-modal features of the resources in the resource behavior sequence of the target user according to the target user data. The resource behavior sequence is determined according to the historical operation behavior of the target user on the resources. The fourth relationship vector output module is configured to obtain a fourth relationship vector according to the second multi-modal feature sequence, the second multi-modal feature of the resource to be recommended, and the multi-modal network. The fifth relationship vector output module is configured to obtain a fifth relationship vector according to the resource identifier of the resource to be recommended, the resource attribute features, the resource statistical features, the first user attribute features, the second resource identifier sequence, and the first network. The sixth vector output module is configured to obtain the target recommendation scenario identifier, and determine the target user private features of the target recommendation scenario from the target user data according to the scenario network corresponding to the target recommendation scenario indicated by the target recommendation scenario identifier, and obtain a sixth vector of the target user private features. The recommended resource determination module is configured to obtain the recommended resources of the target user in the target recommendation scenario according to the fourth relationship vector, the fifth relationship vector, the sixth vector, and the multi-task network. A resource recommendation module, configured to push the recommended resource to the to-be-recommended scenario of the client where the target user is located; Among them, the multi-modal network, the multi-task network, the first network, and the scenario network corresponding to the target recommendation scenario respectively include the multi-modal network, the multi-task network, the first network, and the scenario network in the target resource recommendation model according to any one of claims 1 to 4.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 5.

9. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 5.

10. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method according to any one of claims 1 to 5.