Cover picture style determination method and device, model training method and device and storage medium

By obtaining the characteristics of the target resources in multiple recommendation scenarios and using the style model to determine the target cover image style, the problem of lack of personalization of the cover image style is solved, personalized cover image style adaptation is achieved in different scenarios, and the user experience is improved.

CN120707668APending Publication Date: 2025-09-26BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510661743.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, the method of determining the cover image style of the resource lacks personalization, resulting in a single user experience and difficulty in adapting to the needs of different recommendation scenarios.

Method used

By obtaining the scene features, resource features and object features of the target resource in multiple recommended scenarios, the target cover image style is determined using the trained style model, and the cover image style is dynamically adjusted to adapt to the object requirements in different scenarios.

Benefits of technology

The cover image style is adapted to the object characteristics and scene characteristics, providing a personalized cover image style that meets the object requirements and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707668A_ABST
    Figure CN120707668A_ABST
Patent Text Reader

Abstract

The invention provides a cover picture style determination method and device, a model training method and device and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning and the like. According to the specific implementation scheme, to-be-recommended target objects corresponding to target resources in multiple recommendation scenes respectively are obtained, scene features corresponding to the multiple recommendation scenes respectively, first resource features of the target resources and first object features of the target objects are obtained, and a trained style model is utilized to recommend the target objects to the multiple recommendation scenes. And according to the scene feature, the first resource feature and the first object feature, determining target cover image styles corresponding to the target resource in the plurality of recommendation scenes, so that the provided cover image styles are matched with the object feature and the scene feature, dynamic adjustment of the cover image styles is realized, and the user experience is improved. Therefore, cover picture styles meeting object requirements can be provided in different recommendation scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as computer vision and deep learning, and in particular to a cover image style determination method, model training method, device and storage medium. Background Art

[0002] At present, the cover image style of a resource is usually determined based on the inherent attributes of the resource. Therefore, the cover image style of the resource provided to each user in different resource recommendation scenarios is the same. This method of determining the cover image style lacks personalization, the user experience is single, and it is difficult to adapt to the needs of different scenarios. Summary of the Invention

[0003] The present disclosure provides a cover image style determination method, model training method, device and storage medium.

[0004] According to one aspect of the present disclosure, a method for determining a cover image style is provided, the method comprising: obtaining target objects to be recommended corresponding to a target resource in multiple recommendation scenarios; obtaining scene features corresponding to each of the multiple recommendation scenarios, a first resource feature of the target resource, and a first object feature of the target object; and using a trained style model to determine the target cover image styles corresponding to the target resource in the multiple recommendation scenarios according to the scene features, the first resource features, and the first object features.

[0005] According to another aspect of the present disclosure, a method for training a style model is provided, the method comprising: obtaining training data, wherein the training data comprises: sample cover image styles respectively used by sample resources in multiple recommendation scenarios, second object features of sample objects that have interacted with the sample resources in the recommendation scenarios, scene features corresponding to each of the multiple recommendation scenarios, and second resource features of the sample resources; utilizing a style model to determine, based on the scene features, the second resource features, and the second object features, predicted cover image styles corresponding to the sample resources in the multiple recommendation scenarios; training the style model based on the predicted cover image styles and sample cover image styles corresponding to the sample resources in the multiple recommendation scenarios to obtain a trained style model.

[0006] According to another aspect of the present disclosure, a device for determining a cover image style is provided, the device comprising: a first acquisition module for acquiring target objects to be recommended corresponding to a target resource in multiple recommendation scenarios; a second acquisition module for acquiring scene features corresponding to each of the multiple recommendation scenarios, a first resource feature of the target resource, and a first object feature of the target object; a first determination module for determining, by using a trained style model, the target cover image styles corresponding to the target resource in the multiple recommendation scenarios according to the scene features, the first resource features, and the first object features.

[0007] According to another aspect of the present disclosure, a training device for a style model is provided, the device comprising: a third acquisition module for acquiring training data, wherein the training data comprises: sample cover image styles respectively used by sample resources in multiple recommendation scenarios, second object features of sample objects that have interacted with the sample resources in the recommendation scenarios, scene features corresponding to each of the multiple recommendation scenarios, and second resource features of the sample resources; a second determination module for determining, by using the style model, the predicted cover image styles respectively corresponding to the sample resources in the multiple recommendation scenarios according to the scene features, the second resource features, and the second object features; a training module for training the style model according to the predicted cover image styles and sample cover image styles respectively corresponding to the sample resources in the multiple recommendation scenarios to obtain a trained style model.

[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for determining the cover image style proposed in the present disclosure, or the method for training a style model.

[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method for determining the cover image style proposed in the present disclosure, or the method for training the style model.

[0010] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the method for determining the cover image style proposed in the present disclosure, or the steps of the method for training the style model proposed in the present disclosure.

[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0013] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;

[0014] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;

[0015] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;

[0016] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure;

[0017] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure;

[0018] Figure 6 This is an example diagram of the model structure of the style model;

[0019] Figure 7 is a schematic diagram according to a sixth embodiment of the present disclosure;

[0020] Figure 8 is a schematic diagram according to a seventh embodiment of the present disclosure;

[0021] Figure 9 is a schematic diagram according to an eighth embodiment of the present disclosure;

[0022] Figure 10 is a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0024] In related technologies, the cover image style of a resource is usually determined based on the inherent attributes of the resource. Therefore, the cover image style of the resource provided to each user in different resource recommendation scenarios is the same. This method of determining the cover image style lacks personalization, the user experience is single, and it is difficult to adapt to the needs of different scenarios.

[0025] In response to the above problems, the present disclosure proposes a cover image style determination method, model training method, device and storage medium. The scheme proposes the following technical concept: obtain the target objects to be recommended corresponding to the target resource in multiple recommendation scenarios, and obtain the scene features corresponding to each of the multiple recommendation scenarios, the first resource features of the target resource, and the first object features of the target object. Using the trained style model, according to the scene features, the first resource features and the first object features, the target cover image styles corresponding to the target resource in multiple recommendation scenarios are determined, so that the provided cover image style is adapted to the object features and scene features, and dynamic adjustment of the cover image style is realized, so that a cover image style that meets the object requirements can be provided in different recommendation scenarios.

[0026] Figure 1 It is a schematic diagram according to the first embodiment of the present disclosure, wherein it should be noted that the method for determining the cover image style of the embodiment of the present disclosure can be applied to a device for determining the cover image style, which can be an electronic device, or can be configured in an electronic device so that the electronic device can perform the function of determining the cover image style.

[0027] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, and other hardware devices with various operating systems, touch screens and / or display screens.

[0028] It should be noted that, in the following embodiments, the device for determining the cover image style is described as an electronic device as an example.

[0029] like Figure 1 As shown, the method for determining the cover image style may include the following steps:

[0030] Step 101: Obtain target objects to be recommended corresponding to a target resource in multiple recommendation scenarios.

[0031] The target resource in this embodiment may be various types of resources. For example, the target resource may be a video, an image, an article, a graphic, etc. This embodiment does not specifically limit the target resource.

[0032] The recommendation scenario may be one of multiple applications on the user client or one of multiple interactive interfaces of an application on the user client. The multiple applications may include multiple applications for recommending videos. The interactive interface refers to an interface that displays multiple resources to the user via a resource list.

[0033] For example, the number of recommendation scenarios is two, and the two recommendation scenarios can be a first interaction interface that uses a single-column resource list in the same application to display resources to the user, and a second interaction interface that uses a double-column resource list to display resources to the user. The first interaction interface can be the homepage interaction interface of the application, and the second interaction interface can be the discovery interaction interface of the application, or the first interaction interface can be the discovery interaction interface of the application, and the second interaction interface can be the homepage interaction interface of the application.

[0034] Step 102: Acquire scene features corresponding to each of a plurality of recommended scenes, a first resource feature of a target resource, and a first object feature of a target object.

[0035] Among them, the scene characteristics corresponding to the recommended scene refer to the characteristics of the recommended scene. For example, the scene characteristics may include the resource presentation form corresponding to the recommended scene, the user interaction method supported by the recommended scene, etc. This embodiment limits the scene characteristics corresponding to the recommended scene.

[0036] The target object refers to a virtual object created for an interactive object that interacts with a target resource in the real world.

[0037] The interaction object may be a user, or a robot, etc. This embodiment does not specifically limit the interaction object.

[0038] The target object may be identification information of an interactive object. For example, if the interactive object is a user, the target object may be user identification information, such as a user account. This embodiment does not specifically limit the target object.

[0039] The first object feature refers to a feature related to the target object. For example, the first object feature of the target object may include but is not limited to any feature associated with the target object, such as the target object's behavior feature, browsing feature, and operation feature, which is not specifically limited in this embodiment.

[0040] Among them, the first resource characteristics of the target resource refer to the characteristics of the target resource itself. For example, the first resource characteristics of the target resource may include but are not limited to the resource type, release time, interaction method, resource quality and resource format of the target resource, etc. This embodiment does not make specific limitations on this.

[0041] Step 103: Using the trained style model, according to the scene feature, the first resource feature and the first object feature, determine the target cover image styles corresponding to the target resource in multiple recommended scenes.

[0042] Among them, the trained style model is trained based on the sample cover image styles used by the sample resources in multiple recommendation scenarios, the scene features of each recommendation scenario, and the object features of the sample objects that have interacted with the sample resources in the recommendation scenarios.

[0043] It should be noted that the sample cover image styles used by the sample resources in multiple recommendation scenarios may be different, or may be partially the same, and this embodiment does not specifically limit this.

[0044] Among them, the cover image style is the visual presentation style and specifications of the cover image designed for the resource.

[0045] In this embodiment, the scene features, the first resource features, and the first object features may be input into a trained style model to obtain target cover image styles corresponding to the target resource in multiple recommended scenes through the trained style model.

[0046] The method for determining the cover image style of the embodiment of the present disclosure obtains the target objects to be recommended corresponding to the target resource in multiple recommendation scenarios, and obtains the scene features corresponding to each of the multiple recommendation scenarios, the first resource features of the target resource, and the first object features of the target object. By using a trained style model, the target cover image styles corresponding to the target resource in the multiple recommendation scenarios are determined according to the scene features, the first resource features and the first object features, so that the provided cover image style is adapted to the object features and the scene features, and dynamic adjustment of the cover image style is achieved, so that a cover image style that meets the object requirements can be provided in different recommendation scenarios.

[0047] Based on the above embodiments, after determining the target cover image style of the target resource in each recommended scenario, for each recommended scenario, multiple candidate cover images with the target cover image style can be generated based on the target resource and the target cover image style of the target resource in the recommended scenario, and the target cover image of the target resource in the resource recommendation scenario can be determined from the multiple candidate cover images, and the target resource with the target cover image can be sent to the target object.

[0048] In some embodiments, in order to improve the quality of the target cover image of the target resource sent to the target object, a possible implementation method for determining the target cover image of the target resource in the resource recommendation scenario from multiple candidate cover images can be: determining the scores of the candidate cover images in multiple preset evaluation dimensions, and performing weighted summation on the weights corresponding to each evaluation dimension and the scores of the candidate cover images in the corresponding evaluation dimensions to obtain a comprehensive score of the candidate cover image, and determining the cover image with the highest comprehensive score from multiple candidate cover images as the target cover image of the target resource in the resource recommendation scenario.

[0049] Among them, the above-mentioned multiple evaluation dimensions may include but are not limited to at least two of the following: image clarity dimension, image text and / or face truncation dimension, image border dimension, and image aesthetic dimension.

[0050] It should be noted that the weights corresponding to different evaluation dimensions are pre-set.

[0051] Among them, the weights corresponding to different evaluation dimensions may be different.

[0052] Among them, it can be understood that an example of the specific process of determining the score of the candidate cover image in the image border dimension is: performing border detection on the candidate cover image to obtain the border situation in the candidate cover image, and determining the score of the candidate cover image in the image border dimension based on the border situation in the candidate cover image.

[0053] The border situation is used to indicate whether the candidate cover image includes a border. As an example, if the border situation determines that the candidate cover image includes a border, the score of the candidate cover image in the image border dimension can be determined to be a first score value; if the border situation determines that the candidate cover image does not include a border, the score of the candidate cover image in the image border dimension can be determined to be a second score value. The first score value and the second score value are different. For example, the first score value can be 1 and the second score value can be 0.

[0054] For example, the candidate cover image is an image that includes a face. The aforementioned multiple preset evaluation dimensions include the image clarity dimension, the image face truncation dimension, the image border dimension, and the image aesthetics dimension. Correspondingly, the candidate cover image's score 1 in the image clarity dimension can be determined, and the candidate cover image's score 2 in the image face truncation dimension can be determined, the candidate cover image's score 3 in the image border dimension can be determined, and the candidate cover image's score 4 in the image aesthetics dimension can be determined. Assuming that the weight corresponding to the image clarity dimension is w1; the weight corresponding to the image face truncation dimension is w2; the weight corresponding to the image border dimension is w3; and the weight corresponding to the image aesthetics dimension is w4, correspondingly, a weighted sum can be taken based on the weights corresponding to each evaluation dimension and the candidate cover image's score in the corresponding evaluation dimension to obtain a comprehensive score A for the candidate cover image. The formula for obtaining the comprehensive score A can be:

[0055] A=w1*score1+w2*score2+w3*score3+w4*score4.

[0056] In some examples, in order to improve the quality of the cover image provided to the target object, when the above-mentioned preset multiple evaluation dimensions include: image clarity dimension, image text / face truncation dimension, image border dimension, and image aesthetics dimension, the relationship between the weights corresponding to the above-mentioned multiple evaluation dimensions can be: weight of image text / face truncation dimension > weight of image clarity dimension > image clarity dimension > image border dimension > image aesthetics dimension.

[0057] Among them, in order to enable the style model to accurately determine the target cover image style of the target resource in each recommended scene, in some embodiments, when the above-mentioned style model includes a feature extraction network and a tower network corresponding to each recommended scene, and the feature extraction network includes a first expert network corresponding to each recommended scene and a second expert network shared by each recommended scene, the above-mentioned use of the trained style model is a possible implementation method for determining the target cover image styles corresponding to the target resource in multiple recommended scenes according to the scene features, the first resource features and the first object features, such as Figure 2 shown.

[0058] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure.

[0059] like Figure 2 As shown, the method may include:

[0060] Step 201: Obtain target objects to be recommended corresponding to the target resource in multiple recommendation scenarios.

[0061] Step 202: Acquire scene features corresponding to each of the plurality of recommended scenes, a first resource feature of the target resource, and a first object feature of the target object.

[0062] It should be noted that, for the specific description of step 201 and step 202, reference can be made to the relevant description in other embodiments, which will not be repeated here.

[0063] Step 203 : For each recommendation scenario, the first object feature and the first resource feature are input into a first expert network corresponding to the recommendation scenario, and a first specific feature of the recommendation scenario is obtained through the first expert network corresponding to the recommendation scenario.

[0064] Step 204: Input the first object feature and the first resource feature into the second expert network, and obtain a first shared feature common to multiple recommendation scenarios through the second expert network.

[0065] In this embodiment, after the first object feature and the first resource feature are input into the second expert network, the second expert network may perform feature extraction on the first object feature and the first resource feature to obtain a first shared feature that is common to multiple recommendation scenarios.

[0066] Step 205: Input the first proprietary feature, the first shared feature, the scene feature corresponding to the recommended scene, and the first resource feature of the recommended scene into the tower network corresponding to the recommended scene to obtain a target cover image style of the target resource in the recommended scene.

[0067] In this embodiment, the tower network of each recommendation scenario can independently use shared features and proprietary features to ensure that the reasoning process of each recommendation scenario in determining the target cover image style of the target resource in the corresponding recommendation scenario is independent of each other and does not affect each other, and can efficiently and accurately determine the target cover image style of the target resource in each recommendation scenario.

[0068] In some embodiments, when the feature extraction network further includes: each recommendation scenario corresponds to a gating network, a possible implementation method of inputting the first unique feature, the first shared feature, the scene feature corresponding to the recommendation scenario, and the first resource feature of the recommendation scenario into the tower network corresponding to the recommendation scenario to obtain the target cover image style of the target resource in the recommendation scenario is: through the gating network corresponding to the recommendation scenario, the first unique feature and the first shared feature of the recommendation scenario, the scene feature corresponding to the recommendation scenario, and the first resource feature are fused to obtain a first fused feature; and the first fused feature is input into the tower network corresponding to the recommendation scenario to obtain the target cover image style of the target resource in the recommendation scenario. In this way, the gating network corresponding to the recommendation scenario fuses the input shared features and unique features, and inputs the fused features into the tower network corresponding to the recommendation scenario to determine the target cover image style of the target resource in the recommendation scenario. This improves the synergy between different recommendation scenarios while making the reasoning process of each recommendation scenario independent of each other to avoid interference, thereby efficiently and accurately determining the target cover image style corresponding to the target resource in each recommendation scenario.

[0069] In addition, in this embodiment, in order to enable the tower network corresponding to the recommended scene to accurately determine the target cover image style corresponding to the target resource in the corresponding recommended scene, this embodiment directly inputs the scene features of the recommended scene into the gating network to avoid unnecessary information interference, so that the tower network of the recommended scene can combine the scene features of the recommended scene, and process the reasoning process under the recommended scene in a targeted manner, and accurately determine the target cover image style corresponding to the target resource in the recommended scene.

[0070] In addition, in this embodiment, in order to enable the tower network corresponding to the recommended scene to conveniently capture the first resource feature of the target resource, the first resource feature of the target resource can also be input into the gating network corresponding to the recommended scene, so that the tower network of the recommended scene can conveniently capture the first resource feature of the target resource from the gating network of the recommended scene, and combine the first resource feature to accurately determine the target cover image style corresponding to the target resource in the recommended scene.

[0071] In this embodiment, a possible implementation method for fusing the first specific feature and the first shared feature of the recommendation scene, and the scene feature and the first resource feature of the recommendation scene, through the gating network corresponding to the recommendation scene, to obtain the first fused feature is: concatenating the embedding representations corresponding to the scene feature and the first resource feature of the recommendation scene to obtain a first concatenated embedding representation; and fusing the first specific feature, the first shared feature, and the first concatenated embedding representation of the recommendation scene through the gating network corresponding to the recommendation scene to obtain the first fused feature. Thus, the concatenation preserves the independence of the scene feature and the first resource feature, allowing the gating network to accurately fuse the received input features and accurately obtain the first fused feature.

[0072] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure. It should be noted that the style model in this embodiment includes a feature extraction network and a tower network corresponding to each recommendation scenario. The feature extraction network includes a first expert network corresponding to each recommendation scenario and a second expert network shared by all recommendation scenarios.

[0073] like Figure 3 As shown, the method may include the following steps:

[0074] Step 301: Obtain target objects to be recommended corresponding to the target resource in multiple recommendation scenarios.

[0075] Step 302: Acquire scene features corresponding to each of the plurality of recommended scenes, a first resource feature of the target resource, and a first object feature of the target object.

[0076] It should be noted that, for the specific implementation of step 301 and step 302, reference may be made to the relevant descriptions in other embodiments, which will not be repeated here.

[0077] Step 303 : For each recommendation scenario, the first object feature and the first resource feature are input into a first expert network corresponding to the recommendation scenario, and a first specific feature of the recommendation scenario is obtained through the first expert network corresponding to the recommendation scenario.

[0078] In some embodiments, when the style model may further include an embedding layer and a splicing layer, a possible implementation of step 303 may be: embedding representations of the first object feature and the first resource feature respectively through the embedding layer to obtain a first embedding representation of the first object feature and a second embedding representation of the first resource feature; splicing the first embedding representation and the second embedding representation through the splicing layer to obtain a third embedding representation; inputting the third embedding representation into the first expert network corresponding to the recommendation scenario, and obtaining the first specific feature of the first recommendation scenario through the first expert network corresponding to the recommendation scenario. Thus, the embedding layer accurately determines the embedding representations corresponding to the first object feature and the first resource feature, and splices the embedding representations corresponding to the first object feature and the first resource feature through the splicing layer, and inputs the spliced ​​embedding representations into the first expert network corresponding to the recommendation scenario. The splicing can preserve the independence of the embedding representations corresponding to the first object feature and the first resource feature, so that the first expert network corresponding to the recommendation scenario can accurately determine the specific features of the recommendation scenario.

[0079] In some embodiments, the embedding layer includes a first embedding layer and a second embedding layer, and the embedding layers are used to embed the first object feature and the first resource feature, respectively, to obtain a first embedded representation of the first object feature and a second embedded representation of the first resource feature, including: embedding the first object feature through the first embedding layer to obtain the first embedded representation of the first object feature; and embedding the first resource feature through the second embedding layer to obtain the second embedded representation of the first resource feature. Thus, by embedding the first object feature and the first resource feature through different embedding layers, the embedded representation of the first object feature and the embedded representation of the first resource feature can be accurately obtained.

[0080] In some embodiments, when the style model further includes a coding layer and the first object feature includes temporal behavior sequence data, a possible implementation method for embedding the first object feature through the first embedding layer to obtain a first embedded representation of the first object feature is: encoding the temporal behavior sequence data through the coding layer to obtain a behavior sequence encoding result corresponding to the temporal behavior sequence data; and embedding the behavior sequence encoding result through the first embedding layer to obtain a first embedded representation. Thus, the first embedded representation is accurately determined through the temporal behavior sequence data of the target object, facilitating subsequent processing based on the determined first embedded representation, thereby improving the accuracy of the subsequently determined target cover image style.

[0081] Step 304: Input the first object feature and the first resource feature into the second expert network, and obtain a first shared feature common to multiple recommendation scenarios through the second expert network.

[0082] Step 305: For each recommendation scenario, in the tower network corresponding to the recommendation scenario, based on the first proprietary feature of the recommendation scenario, the first shared feature, the scene feature corresponding to the recommendation scenario, and the first resource feature, determine the scores of the multiple candidate cover image styles corresponding to the target resource in the recommendation scenario. The scores are used to represent the probability of the target object interacting with the target resource having the candidate cover image style.

[0083] In some embodiments, the feature extraction network also includes: each recommended scene corresponds to a gating network, and in the tower network corresponding to the recommended scene, based on the first proprietary feature, the first shared feature, the scene feature corresponding to the recommended scene, and the first resource feature, a possible implementation method for determining the scores of multiple candidate cover image styles corresponding to the target resource in the recommended scene is: through the gating network corresponding to the recommended scene, the first proprietary feature and the first shared feature of the recommended scene, the scene feature corresponding to the recommended scene, and the first resource feature are fused to obtain a first fused feature, and the first fused feature is input into the tower network corresponding to the recommended scene to obtain the scores of multiple candidate cover image styles corresponding to the target resource in the recommended scene.

[0084] Step 306: Select the candidate cover image style with the highest score from the multiple candidate cover image styles as the target cover image style for the target resource in the recommendation scenario.

[0085] In this embodiment, the tower network corresponding to the recommended scene can accurately determine the target cover image style of the target resource in the recommended scene from multiple candidate cover image styles corresponding to the recommended scene.

[0086] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure. It should be noted that the style model training method of the embodiment of the present disclosure can be applied to a style model training device, which can be an electronic device, or can be configured in an electronic device so that the electronic device can perform the style model training function.

[0087] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, and other hardware devices with various operating systems, touch screens and / or display screens.

[0088] It should be noted that, in the following embodiments, the training device for the style model is described as an electronic device as an example.

[0089] like Figure 4 As shown, the method may include:

[0090] Step 401, obtaining training data, wherein the training data includes: sample cover image styles used by sample resources in multiple recommendation scenarios, second object features of sample objects that have interacted with sample resources in the recommendation scenarios, scene features corresponding to multiple recommendation scenarios, and second resource features of sample resources.

[0091] The sample resources refer to the resources used to train the style model.

[0092] The sample object refers to the object that has interacted with the sample resource.

[0093] Step 402: Using the style model, according to the scene feature, the second resource feature, and the second object feature, determine the predicted cover image styles corresponding to the sample resource in multiple recommendation scenes.

[0094] In this embodiment, the scene feature, the second resource feature, and the second object feature may be input into the style model to obtain the predicted cover image styles corresponding to the sample resource in multiple recommended scenes through the style model.

[0095] Step 403: Train the style model according to the predicted cover image styles and the sample cover image styles corresponding to the sample resources in multiple recommendation scenarios to obtain a trained style model.

[0096] In this embodiment, the loss value of the style model can be determined based on the predicted cover image styles and sample cover image styles corresponding to the sample resources in multiple recommendation scenarios, and the style model can be trained based on the loss value until the loss value meets the preset conditions.

[0097] The preset condition is the condition for model training to end. The preset condition can be configured based on actual needs. For example, the loss value meeting the preset condition can be a loss value less than a preset value, or it can be a loss value that approaches a stable state, meaning that the difference between the loss values ​​of two or more consecutive training runs is less than a set value, indicating that the loss value has basically stopped changing.

[0098] The style model training method provided by the embodiment of the present disclosure obtains the sample cover image styles used by the sample resources in multiple recommendation scenarios, the second object features of the sample objects that have interacted with the sample resources in the recommendation scenarios, the scene features corresponding to the multiple recommendation scenarios, and the second resource features of the sample resources, and uses the style model to determine the predicted cover image styles corresponding to the sample resources in the multiple recommendation scenarios according to the scene features, the second resource features, and the second object features, and trains the style model according to the predicted cover image styles and the sample cover image styles corresponding to the sample resources in the multiple recommendation scenarios, so as to accurately obtain the trained style model, which is conducive to improving the training efficiency and accuracy of the style model, and facilitates the subsequent determination of the cover image styles corresponding to the corresponding resources in different recommendation scenarios based on the trained style model.

[0099] In some embodiments, when the style model includes: a feature extraction network and a tower network corresponding to each recommendation scene, and the feature extraction network includes: a first expert network corresponding to each recommendation scene and a second expert network shared by each recommendation scene, a possible implementation method of using the style model to determine the predicted cover image styles corresponding to the sample resource in multiple recommendation scenes according to the scene features, the second resource features, and the second object features is as follows: Figure 5 shown.

[0100] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure.

[0101] like Figure 5 As shown, the method may include:

[0102] Step 501, obtaining training data, wherein the training data includes: sample cover image styles used by sample resources in multiple recommendation scenarios, second object features of sample objects that have interacted with sample resources in the recommendation scenarios, scene features corresponding to multiple recommendation scenarios, and second resource features of sample resources.

[0103] Step 502 : For each recommendation scenario, the second object feature and the second resource feature are input into a first expert network corresponding to the recommendation scenario, and a second specific feature of the recommendation scenario is obtained through the first expert network corresponding to the recommendation scenario.

[0104] In some embodiments, when the style model further includes an embedding layer and a splicing layer, a possible implementation method for inputting the second object feature and the second resource feature into the first expert network corresponding to the recommendation scenario and obtaining the second specific feature of the recommendation scenario through the first expert network corresponding to the recommendation scenario is as follows: embedding the second object feature and the second resource feature respectively through the embedding layer to obtain a fourth embedding representation of the second object feature and a fifth embedding representation of the second resource feature; splicing the fourth embedding representation and the fifth embedding representation through the splicing layer to obtain a sixth embedding representation; inputting the sixth embedding representation into the first expert network corresponding to the recommendation scenario, and obtaining the second specific feature of the recommendation scenario through the first expert network corresponding to the recommendation scenario. Thus, the embedding layer accurately determines the embedding representations corresponding to the second object feature and the second resource feature, splices the embedding representations corresponding to the second object feature and the second resource feature through the splicing layer, and inputs the spliced ​​embedding representations into the first expert network corresponding to the recommendation scenario. The splicing can preserve the independence of the embedding representations corresponding to the first object feature and the first resource feature, so that the first expert network corresponding to the recommendation scenario can accurately determine the specific feature of the recommendation scenario.

[0105] In some embodiments, to improve the accuracy of the embedded representations corresponding to the second object feature and the second resource feature, the embedding layer includes a first embedding layer and a second embedding layer, and the second object feature and the second resource feature are respectively embedded and represented by the embedding layers to obtain a fourth embedded representation of the second object feature and a fifth embedded representation of the second resource feature, including: embedding the second object feature by the first embedding layer to obtain the fourth embedded representation of the second object feature; embedding the second resource feature by the second embedding layer to obtain the fifth embedded representation of the second resource feature. Thus, by embedding the second object feature and the second resource feature respectively by different embedding layers, the embedded representation of the second object feature and the embedded representation of the second resource feature can be accurately obtained.

[0106] In some embodiments, to accurately train the style model, the style model further includes: an encoding layer, wherein the second object feature includes: temporal behavior sequence data; and a first embedding layer is used to embed and represent the second object feature. A possible implementation method for obtaining a fourth embedded representation of the second object feature is: encoding the temporal behavior sequence data through the encoding layer to obtain a behavior sequence encoding result corresponding to the temporal behavior sequence data; and embedding and representing the behavior sequence encoding result through the first embedding layer to obtain a fourth embedded representation. Thus, the encoding layer can accurately determine the behavior sequence encoding result of the temporal behavior sequence data, and the first embedding layer can embed and represent the behavior sequence encoding result, thereby improving the accuracy of the obtained fourth embedded representation.

[0107] Step 503: Input the second object feature and the second resource feature into the second expert network, and obtain a second shared feature common to multiple recommendation scenarios through the second expert network.

[0108] Step 504: input the second proprietary feature of the recommended scene, the second shared feature, the scene feature corresponding to the recommended scene, and the second resource feature into the tower network corresponding to the recommended scene to obtain a predicted cover image style of the sample resource in the recommended scene.

[0109] In some embodiments, when the feature extraction network further includes: each recommended scene corresponds to a gating network, the second proprietary feature, the second shared feature, the scene feature corresponding to the recommended scene, and the second resource feature of the recommended scene are input into the tower network corresponding to the recommended scene to obtain the predicted cover image style of the sample resource in the recommended scene. A possible implementation method is: through the gating network corresponding to the recommended scene, the second proprietary feature, the second shared feature, the scene feature corresponding to the recommended scene, and the second resource feature are fused to obtain the second fused feature; the second fused feature is input into the tower network corresponding to the recommended scene to obtain the predicted cover image style of the sample resource in the recommended scene.

[0110] In this embodiment, the input features are fused through the gating network corresponding to each recommendation scene, and the fused features are input into the tower network corresponding to the recommendation scene. This not only improves the representation ability of the style model and the synergy between recommendation scenes, but also effectively reduces computational overhead and interference, and improves the generalization ability of the style model.

[0111] In some embodiments, in order to enable the subsequently trained style model to accurately determine the cover image styles corresponding to the corresponding resources in multiple recommendation scenarios, the training data in this embodiment also includes: the activity of the interaction between the corresponding sample object and the sample resource in the recommendation scenario, and the second proprietary feature, the second shared feature, the scene feature corresponding to the recommendation scenario, and the second resource feature of the recommendation scenario are fused through the gating network corresponding to the recommendation scenario. A possible implementation method for obtaining the second fused feature can be: through the gating network corresponding to the recommendation scenario, the second proprietary feature, the second shared feature, the scene feature corresponding to the recommendation scenario, the second resource feature, and the activity are fused to obtain the second fused feature. In this way, the style model's ability to capture the intent of the sample object in a fine-grained manner can be significantly improved, and the accuracy of the trained model can be improved.

[0112] In some embodiments, a possible implementation method for fusing the second specific feature, second shared feature, scene feature, second resource feature, and activity level of the recommended scene through the gating network corresponding to the recommended scene to obtain the second fused feature is as follows: concatenating the embedding representations corresponding to the scene feature, second resource feature, and activity level of the recommended scene to obtain a second concatenated embedding representation; and fusing the second specific feature, second shared feature, and second concatenated embedding representation of the recommended scene through the gating network corresponding to the recommended scene to obtain the second fused feature. Thus, the concatenation preserves the independence of the scene feature, second resource feature, and activity level, allowing the gating network to accurately capture these features and avoid interference from irrelevant information, thereby helping the gating network to more accurately perform feature fusion and improve the performance of the model in the recommended scene.

[0113] Step 505 : Training the style model according to the predicted cover image styles and the sample cover image styles corresponding to the sample resources in multiple recommendation scenarios to obtain a trained style model.

[0114] In this embodiment, in the style model, multiple recommendation scenarios share the same set of expert networks, which provides unified shared features for each recommendation scenario, and each recommendation scenario can independently process the shared features common to multiple recommendation scenarios and the proprietary features of the corresponding recommendation scenario through its own corresponding tower network. This design can ensure that the reasoning process of each recommendation scenario will not interfere with each other, thereby avoiding the mutual influence between recommendation scenarios during the training process, and effectively improving the training efficiency and accuracy of the style model.

[0115] In order to clearly understand the present disclosure, Figure 6 The model structure of the illustrated style model is used to exemplify the process of training the style model. Figure 6 In the example, the number of recommended scenes is two, and Figure 6 The two recommended scenarios are represented by recommended scenario 1 and recommended scenario 2. Figure 6It can be seen that the style model shown in this example includes: an encoding layer, a first embedding layer connected to the encoding layer, a second embedding layer, a splicing layer connected to the first embedding layer and the second embedding layer respectively, a second expert network shared by the two recommendation scenes connected to the splicing layer, a first expert network 1 corresponding to the recommendation scene 1 connected to the splicing layer, a first expert network 2 corresponding to the recommendation scene 2 connected to the splicing layer, a gating network 1 corresponding to the recommendation scene 1, a gating network 2 corresponding to the recommendation scene 2, a tower network 1 corresponding to the recommendation scene 1, and a tower network 2 corresponding to the recommendation scene 2, wherein the gating network 1 is connected to the first expert network 1 and the second expert network respectively, the gating network 2 is connected to the first expert network 2 and the second expert network respectively, the tower network 1 is connected to the gating network 1, and the tower network 2 is connected to the gating network 2, and the input data of the gating network 1 also includes the second resource feature of the sample resource, the activity 1 of the interaction between the sample object and the sample resource in the recommendation scene 1, and the scene feature of the recommendation scene 1, and the input data of the gating network 2 also includes the second resource feature of the sample resource, the activity 2 of the interaction between the sample object and the sample resource in the recommendation scene 2, and the scene feature of the recommendation scene 2. Among them, it should be noted that Figure 6 The temporal behavior sequence data i in represents the temporal behavior sequence data of sample object i. The value of i ranges from 1 to n, where n is an integer greater than 1.

[0116] Among them, the example process of training style model is as follows Figure 7 shown.

[0117] Figure 7 is a schematic diagram according to a sixth embodiment of the present disclosure.

[0118] like Figure 7 As shown, this may include:

[0119] Step 701: Obtain training data, where the training data includes: sample cover image styles used by sample resources in two recommendation scenarios, temporal behavior sequence data of sample objects that have interacted with sample resources in the recommendation scenarios, scene features corresponding to the two recommendation scenarios and resource features of the sample resources, and the activity of the interaction between the sample objects and the sample resources in the recommendation scenarios.

[0120] Step 702: Input the temporal behavior sequence data into the encoding layer in the style model to obtain a behavior sequence encoding result, and input the behavior sequence encoding result into the first embedding layer to obtain a first target embedding representation.

[0121] It should be noted that the encoding layer in this embodiment may include an encoder, and the encoder may be used to encode the temporal behavior sequence data to obtain a temporal behavior sequence encoding result.

[0122] Step 703: Input the resource features of the sample resource into the second embedding layer to obtain a second target embedding representation of the resource features.

[0123] Step 704: input the first target embedding representation and the second target embedding representation into the concatenation layer to obtain a concatenation result.

[0124] Step 705: Input the splicing results into the first expert network and the second expert network respectively.

[0125] Step 706: For each recommendation scenario, the output of the first expert network corresponding to the recommendation scenario, the output of the second expert network, the resource characteristics and activity of the sample resources, and the scene characteristics corresponding to the recommendation scenario are input into the gating network corresponding to the recommendation scenario. The gating network fuses the input features to obtain fused features.

[0126] The gating network in this embodiment may include a fully connected layer and an output layer. Accordingly, the fully connected layer in the gating network may determine received input features and corresponding weights, fuse the input features based on the weights of the corresponding input features, and output the fused features through the output layer.

[0127] Step 707: Input the fusion features output by the gated network corresponding to the recommendation scene into the tower network corresponding to the recommendation scene, so as to determine the score of the sample object on multiple candidate cover image styles through the tower network corresponding to the recommendation scene, wherein the score is used to represent the probability of the sample object interacting with the sample resource having the candidate cover image style.

[0128] Among them, multiple candidate image styles may include square images, horizontal images, vertical images, large images, etc.

[0129] Step 708: Use the cover image style with the highest score among the multiple candidate cover image styles as the predicted cover image style of the sample resource in the recommendation scenario.

[0130] Step 709 : Training the style model according to the predicted cover image style and the style cover image style corresponding to the sample resources in the two recommendation scenarios to obtain a trained style model.

[0131] Among them, it should be noted that after obtaining the trained style model, after obtaining the target objects to be recommended corresponding to the target resource in the two recommendation scenarios, the scene features corresponding to the two recommendation scenarios, the object features of the target object and the resource features of the target resource can be input into the trained style model, so as to obtain the target cover image styles corresponding to the target resource in the two recommendation scenarios respectively through the trained style model. In this way, the provided cover image style is adapted to the object features and scene features, and the dynamic adjustment of the cover image style is realized, so that a cover image style that meets the object requirements can be provided in different recommendation scenarios, and a cover image with the target cover image style is generated for the target resource, and the target resource with the cover image is sent to the target object. While improving the quality of the cover image of the target resource sent to the target object, the recommendation effect can also be improved.

[0132] In order to implement the above embodiments, the present disclosure also provides a device for determining a cover image style.

[0133] Figure 8 is a schematic diagram according to the seventh embodiment of the present disclosure.

[0134] like Figure 8 As shown, the cover image style determination device 80 may include: a first acquisition module 801, a second acquisition module 802 and a first determination module 803, wherein:

[0135] The first acquisition module 801 is configured to acquire target objects to be recommended corresponding to the target resource in multiple recommendation scenarios.

[0136] The second acquisition module 802 is configured to acquire scene features corresponding to each of the plurality of recommended scenes, a first resource feature of the target resource, and a first object feature of the target object.

[0137] The first determination module 803 is used to use the trained style model to determine the target cover image styles corresponding to the target resource in multiple recommended scenes according to the scene characteristics, the first resource characteristics and the first object characteristics.

[0138] As a possible implementation of an embodiment of the present disclosure, the style model includes: a feature extraction network and a tower network corresponding to each recommendation scenario. The feature extraction network includes a first expert network corresponding to each recommendation scenario and a second expert network shared by all recommendation scenarios. The first determination module 803 includes:

[0139] A first determining unit is configured to input, for each recommendation scenario, the first object feature and the first resource feature into a first expert network corresponding to the recommendation scenario, and obtain a first specific feature of the recommendation scenario through the first expert network corresponding to the recommendation scenario;

[0140] A second determining unit is configured to input the first object feature and the first resource feature into a second expert network, and obtain a first shared feature common to multiple recommendation scenarios through the second expert network;

[0141] The third determination unit is used to input the first proprietary feature, the first shared feature, the scene feature corresponding to the recommended scene and the first resource feature of the recommended scene into the tower network corresponding to the recommended scene to obtain the target cover image style of the target resource in the recommended scene.

[0142] As a possible implementation of the embodiment of the present disclosure, the feature extraction network further includes: each recommendation scenario corresponds to a gating network, and a third determination unit is specifically configured to:

[0143] fusing, through a gating network corresponding to the recommendation scenario, a first proprietary feature and a first shared feature of the recommendation scenario, a scene feature corresponding to the recommendation scenario, and a first resource feature to obtain a first fused feature;

[0144] The first fusion feature is input into the tower network corresponding to the recommendation scene to obtain the target cover image style of the target resource in the recommendation scene.

[0145] As a possible implementation method of an embodiment of the present disclosure, the first proprietary feature and the first shared feature of the recommended scene, and the scene feature and the first resource feature corresponding to the recommended scene are fused through the gating network corresponding to the recommended scene to obtain the first fused feature. The implementation method can be: splicing the embedded representations corresponding to the scene feature and the first resource feature corresponding to the recommended scene to obtain a first spliced ​​embedded representation; and fusing the first proprietary feature, the first shared feature and the first spliced ​​embedded representation of the recommended scene through the gating network corresponding to the recommended scene to obtain the first fused feature.

[0146] As a possible implementation of the embodiment of the present disclosure, the third determining unit is specifically configured to:

[0147] In the tower network corresponding to the recommended scenario, based on the first proprietary feature of the recommended scenario, the first shared feature, the scenario feature corresponding to the recommended scenario, and the first resource feature, determining scores for multiple candidate cover image styles corresponding to the target resource in the recommended scenario, where the scores are used to represent the probability of the target object interacting with the target resource having the candidate cover image style;

[0148] From multiple candidate cover image styles, select the candidate cover image style with the highest score as the target cover image style for the target resource in the recommendation scenario.

[0149] As a possible implementation of the embodiment of the present disclosure, the style model further includes: an embedding layer and a splicing layer, and a first determination unit specifically configured to:

[0150] Embedding the first object feature and the first resource feature respectively through the embedding layer to obtain a first embedded representation of the first object feature and a second embedded representation of the first resource feature;

[0151] The first embedding representation and the second embedding representation are concatenated through a concatenation layer to obtain a third embedding representation;

[0152] The third embedding representation is input into a first expert network corresponding to the recommendation scenario, and a first proprietary feature of the first recommendation scenario is obtained through the first expert network corresponding to the recommendation scenario.

[0153] As a possible implementation method of an embodiment of the present disclosure, the embedding layer includes a first embedding layer and a second embedding layer. The first object feature and the first resource feature are respectively embedded and represented by the embedding layer to obtain a first embedded representation of the first object feature and a second embedded representation of the first resource feature. A possible implementation method can be: embedding the first object feature through the first embedding layer to obtain a first embedded representation of the first object feature; embedding the first resource feature through the second embedding layer to obtain a second embedded representation of the first resource feature.

[0154] As a possible implementation method of an embodiment of the present disclosure, the style model also includes: a coding layer, the first object feature includes: temporal behavior sequence data, the first object feature is embedded and represented by the first embedding layer, and a possible implementation method for obtaining a first embedded representation of the first object feature is: encoding the temporal behavior sequence data through the coding layer to obtain a behavior sequence encoding result corresponding to the temporal behavior sequence data; embedding the behavior sequence encoding result through the first embedding layer to obtain a first embedded representation.

[0155] It should be noted that the aforementioned explanation of the embodiment of the method for determining the cover image style is also applicable to the device for determining the cover image style of this embodiment, and will not be repeated here.

[0156] The device for determining the cover image style of the embodiment of the present disclosure obtains the target objects to be recommended corresponding to the target resource in multiple recommendation scenarios, and obtains the scene features corresponding to each of the multiple recommendation scenarios, the first resource features of the target resource, and the first object features of the target object. By using a trained style model, the device determines the target cover image styles corresponding to the target resource in the multiple recommendation scenarios according to the scene features, the first resource features and the first object features, so that the provided cover image style is adapted to the object features and the scene features, and dynamic adjustment of the cover image style is realized, so that a cover image style that meets the object requirements can be provided in different recommendation scenarios.

[0157] In order to implement the above embodiments, the present disclosure also provides a training device for a style model.

[0158] Figure 9 is a schematic diagram according to an eighth embodiment of the present disclosure.

[0159] like Figure 9 As shown, the style model training device 90 may include: a third acquisition module 901, a second determination module 902 and a training module 903, wherein:

[0160] The third acquisition module 901 is used to obtain training data, wherein the training data includes: sample cover image styles used by sample resources in multiple recommendation scenarios, second object features of sample objects that have interacted with sample resources in the recommendation scenarios, scene features corresponding to multiple recommendation scenarios, and second resource features of sample resources.

[0161] The second determination module 902 is used to determine the predicted cover image styles corresponding to the sample resource in multiple recommendation scenarios according to the scene features, the second resource features and the second object features using the style model.

[0162] The training module 903 is used to train the style model according to the predicted cover image styles and sample cover image styles corresponding to the sample resources in multiple recommendation scenarios to obtain a trained style model.

[0163] In one possible implementation of the embodiment of the present disclosure, the style model includes: a feature extraction network and a tower network corresponding to each recommendation scenario. The feature extraction network includes: a first expert network corresponding to each recommendation scenario and a second expert network shared by all recommendation scenarios. The second determination module 902 includes:

[0164] a fourth determining unit, configured to input the second object feature and the second resource feature into a first expert network corresponding to each recommendation scenario, and obtain a second specific feature of the recommendation scenario through the first expert network corresponding to the recommendation scenario;

[0165] a fifth determining unit, configured to input the second object feature and the second resource feature into a second expert network, and obtain a second shared feature common to multiple recommendation scenarios through the second expert network;

[0166] The sixth determination unit is used to input the second proprietary feature, the second shared feature, the scene feature corresponding to the recommended scene and the second resource feature of the recommended scene into the tower network corresponding to the recommended scene to obtain the predicted cover image style of the sample resource in the recommended scene.

[0167] In a possible implementation of the embodiment of the present disclosure, the sixth determining unit is specifically configured to:

[0168] fusing the second proprietary feature of the recommendation scene, the second shared feature, the scene feature corresponding to the recommendation scene, and the second resource feature through the gating network corresponding to the recommendation scene to obtain a second fused feature;

[0169] The second fusion feature is input into the tower network corresponding to the recommendation scene to obtain the predicted cover image style of the sample resource in the recommendation scene.

[0170] In a possible implementation of an embodiment of the present disclosure, the training data further includes: the activity of the interaction between the corresponding sample object and the sample resource in the recommendation scenario, and the second proprietary feature, the second shared feature, the scene feature corresponding to the recommendation scenario, and the second resource feature are fused through the gating network corresponding to the recommendation scenario to obtain the second fused feature. A possible implementation method is: through the gating network corresponding to the recommendation scenario, the second proprietary feature, the second shared feature, the scene feature corresponding to the recommendation scenario, the second resource feature, and the activity are fused to obtain the second fused feature.

[0171] A possible implementation of the embodiment of the present disclosure is to fuse the second proprietary feature, the second shared feature, the scene feature corresponding to the recommended scene, the second resource feature, and the activity level of the recommended scene through a gating network corresponding to the recommended scene to obtain a second fused feature. A possible implementation is: splicing the embedding representations corresponding to the scene feature, the second resource feature, and the activity level of the recommended scene to obtain a second spliced ​​embedded representation; and fusing the second proprietary feature, the second shared feature, and the second spliced ​​embedded representation of the recommended scene through a gating network corresponding to the recommended scene to obtain a second fused feature.

[0172] In one possible implementation of the embodiment of the present disclosure, the style model further includes: an embedding layer and a splicing layer, and a fourth determination unit, which is specifically used to: embed the second object feature and the second resource feature respectively through the embedding layer to obtain a fourth embedded representation of the second object feature and a fifth embedded representation of the second resource feature; splice the fourth embedded representation and the fifth embedded representation through the splicing layer to obtain a sixth embedded representation; input the sixth embedded representation into the first expert network corresponding to the recommendation scenario, and obtain the second proprietary feature of the recommendation scenario through the first expert network corresponding to the recommendation scenario.

[0173] A possible implementation of the embodiment of the present disclosure is that the embedding layer includes a first embedding layer and a second embedding layer, and the second object feature and the second resource feature are respectively embedded and represented by the embedding layer to obtain a fourth embedded representation of the second object feature and a fifth embedded representation of the second resource feature, including: embedding the second object feature by the first embedding layer to obtain the fourth embedded representation of the second object feature; embedding the second resource feature by the second embedding layer to obtain the fifth embedded representation of the second resource feature.

[0174] In a possible implementation of the embodiment of the present disclosure, the style model further includes: a coding layer, the second object feature includes: temporal behavior sequence data, and the second object feature is embedded and represented by the first embedding layer to obtain a fourth embedded representation of the second object feature. The implementation method is: encoding the temporal behavior sequence data through the coding layer to obtain a behavior sequence encoding result corresponding to the temporal behavior sequence data; embedding the behavior sequence encoding result through the first embedding layer to obtain a fourth embedded representation.

[0175] It should be noted that the aforementioned explanation of the embodiment of the training method of the style model is also applicable to the training device of the style model of this embodiment, and will not be repeated here.

[0176] The style model training device provided by the embodiment of the present disclosure obtains the sample cover image styles used by the sample resources in multiple recommendation scenarios, the second object features of the sample objects that have interacted with the sample resources in the recommendation scenarios, the scene features corresponding to the multiple recommendation scenarios, and the second resource features of the sample resources, and uses the style model to determine the predicted cover image styles corresponding to the sample resources in the multiple recommendation scenarios according to the scene features, the second resource features, and the second object features, and trains the style model according to the predicted cover image styles and the sample cover image styles corresponding to the sample resources in the multiple recommendation scenarios, so as to accurately obtain the trained style model, which is conducive to improving the training efficiency and accuracy of the style model, and facilitates the subsequent determination of the cover image styles corresponding to the corresponding resource features in different recommendation scenarios based on the trained style model.

[0177] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, comply with relevant laws and regulations, and do not violate public order and good morals.

[0178] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0179] Figure 10 is a block diagram of an electronic device according to one embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided for example only and are not intended to limit the implementation of the present disclosure as described and / or claimed herein.

[0180] like Figure 10 As shown, the electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the electronic device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0181] Multiple components in the electronic device 1000 are connected to the I / O interface 1005, including an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0182] The computing unit 1001 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above, such as the method for determining the cover image style, or the method for training the style model. For example, in some embodiments, the method for determining the cover image style, or the method for training the style model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the method for determining the cover image style or the method for training the style model described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute a method for determining a cover image style, or a method for training a style model, in any other appropriate manner (for example, by means of firmware).

[0183] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0184] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0185] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0186] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0187] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0188] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0189] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0190] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for determining a cover image style, comprising: Obtain target objects to be recommended corresponding to the target resource in multiple recommendation scenarios; Acquire scene features corresponding to each of the plurality of recommended scenes, a first resource feature of the target resource, and a first object feature of the target object; The trained style model is used to determine the target cover image styles corresponding to the target resource in the multiple recommended scenes according to the scene features, the first resource features, and the first object features.

2. The method according to claim 1, wherein The style model includes: a feature extraction network and a tower network corresponding to each recommendation scene, the feature extraction network includes a first expert network corresponding to each recommendation scene and a second expert network shared by each recommendation scene, and using the trained style model to determine the target cover image styles corresponding to the target resource in the multiple recommendation scenes according to the scene features, the first resource features, and the first object features, including: For each recommendation scenario, inputting the first object feature and the first resource feature into a first expert network corresponding to the recommendation scenario, and obtaining a first specific feature of the recommendation scenario through the first expert network corresponding to the recommendation scenario; Inputting the first object feature and the first resource feature into the second expert network, and obtaining a first shared feature common to the multiple recommendation scenarios through the second expert network; The first proprietary feature of the recommended scene, the first shared feature, the scene feature corresponding to the recommended scene, and the first resource feature are input into the tower network corresponding to the recommended scene to obtain the target cover image style of the target resource in the recommended scene.

3. The method according to claim 2, wherein: The feature extraction network further includes: each recommendation scene corresponds to a gating network, wherein the first specific feature of the recommendation scene, the first shared feature, the scene feature corresponding to the recommendation scene, and the first resource feature are input into the gating network corresponding to the recommendation scene to obtain a target cover image style of the target resource in the recommendation scene, including: fusing, through the gating network corresponding to the recommendation scenario, the first proprietary feature and the first shared feature of the recommendation scenario, the scene feature corresponding to the recommendation scenario, and the first resource feature to obtain a first fused feature; The first fusion feature is input into the tower network corresponding to the recommendation scene to obtain the target cover image style of the target resource in the recommendation scene.

4. The method according to claim 3, wherein: The step of fusing the first specific feature and the first shared feature of the recommendation scenario, the scene feature corresponding to the recommendation scenario, and the first resource feature through the gating network corresponding to the recommendation scenario to obtain a first fused feature includes: splicing the embedding representations corresponding to the scene feature corresponding to the recommended scene and the first resource feature to obtain a first spliced ​​embedding representation; The first specific feature of the recommendation scene, the first shared feature, and the first concatenated embedding representation are fused through the gating network corresponding to the recommendation scene to obtain the first fused feature.

5. The method according to claim 2, wherein: The step of inputting the first proprietary feature of the recommended scene, the first shared feature, the scene feature corresponding to the recommended scene, and the first resource feature into the tower network corresponding to the recommended scene to obtain a target cover image style of the target resource in the recommended scene includes: In the tower network corresponding to the recommended scenario, determining scores for a plurality of candidate cover image styles corresponding to the target resource in the recommended scenario based on the first proprietary feature of the recommended scenario, the first shared feature, the scenario feature corresponding to the recommended scenario, and the first resource feature, wherein the scores are used to represent the probability of the target object interacting with the target resource having the candidate cover image style; From the multiple candidate cover image styles, select the candidate cover image style with the highest score as the target cover image style of the target resource in the recommendation scenario.

6. The method according to any one of claims 2 to 5, wherein: The style model further includes an embedding layer and a concatenation layer. Inputting the first object feature and the first resource feature into a first expert network corresponding to the recommendation scenario, and obtaining a first specific feature of the recommendation scenario through the first expert network corresponding to the recommendation scenario, includes: Embedding the first object feature and the first resource feature respectively through the embedding layer to obtain a first embedded representation of the first object feature and a second embedded representation of the first resource feature; splicing the first embedding representation and the second embedding representation through the splicing layer to obtain a third embedding representation; The third embedded representation is input into a first expert network corresponding to the recommendation scenario, and a first specific feature of the first recommendation scenario is obtained through the first expert network corresponding to the recommendation scenario.

7. The method according to claim 6, wherein: The embedding layer includes a first embedding layer and a second embedding layer, and embedding the first object feature and the first resource feature respectively through the embedding layer to obtain a first embedded representation of the first object feature and a second embedded representation of the first resource feature, including: Embedding the first object feature through the first embedding layer to obtain a first embedded representation of the first object feature; The first resource feature is embedded and represented by the second embedding layer to obtain a second embedded representation of the first resource feature.

8. The method according to claim 7, wherein: The style model further includes: an encoding layer, the first object feature includes: temporal behavior sequence data, and the embedding representation of the first object feature by the first embedding layer to obtain a first embedded representation of the first object feature includes: Performing encoding processing on the temporal behavior sequence data through the encoding layer to obtain a behavior sequence encoding result corresponding to the temporal behavior sequence data; The behavior sequence encoding result is embedded and represented by the first embedding layer to obtain the first embedded representation.

9. A method for training a style model, comprising: Acquire training data, wherein the training data includes: sample cover image styles used by the sample resource in multiple recommendation scenarios, second object features of sample objects that interact with the sample resource in the recommendation scenarios, scene features corresponding to each of the multiple recommendation scenarios, and second resource features of the sample resource; Determining, using a style model, predicted cover image styles corresponding to the sample resource in the plurality of recommendation scenarios, respectively, based on the scenario feature, the second resource feature, and the second object feature; The style model is trained according to the predicted cover image styles and the sample cover image styles corresponding to the sample resources in the multiple recommendation scenarios to obtain a trained style model.

10. The method according to claim 9, wherein: The style model includes: a feature extraction network and a tower network corresponding to each recommendation scene, the feature extraction network includes: a first expert network corresponding to each recommendation scene, and a second expert network shared by each recommendation scene. The style model is used to determine the predicted cover image styles corresponding to the sample resource in the multiple recommendation scenes according to the scene features, the second resource features, and the second object features, including: For each recommendation scenario, inputting the second object feature and the second resource feature into a first expert network corresponding to the recommendation scenario, and obtaining a second specific feature of the recommendation scenario through the first expert network corresponding to the recommendation scenario; Inputting the second object feature and the second resource feature into the second expert network, and obtaining a second shared feature common to the multiple recommendation scenarios through the second expert network; The second proprietary feature of the recommended scene, the second shared feature, the scene feature corresponding to the recommended scene, and the second resource feature are input into the tower network corresponding to the recommended scene to obtain the predicted cover image style of the sample resource in the recommended scene.

11. The method according to claim 10, wherein: The feature extraction network further includes: each recommendation scene corresponds to a gating network, wherein the second specific feature of the recommendation scene, the second shared feature, the scene feature corresponding to the recommendation scene, and the second resource feature are input into the gating network corresponding to the recommendation scene to obtain a predicted cover image style of the sample resource in the recommendation scene, including: fusing, through the gating network corresponding to the recommendation scenario, the second proprietary feature of the recommendation scenario, the second shared feature, the scenario feature corresponding to the recommendation scenario, and the second resource feature to obtain a second fused feature; The second fusion feature is input into the tower network corresponding to the recommendation scene to obtain the predicted cover image style of the sample resource in the recommendation scene.

12. The method according to claim 11, wherein The training data further includes: activity of interaction between the sample object corresponding to the recommendation scenario and the sample resource; and the second fused feature is obtained by fusing the second specific feature of the recommendation scenario, the second shared feature, the scene feature corresponding to the recommendation scenario, and the second resource feature through the gating network corresponding to the recommendation scenario, including: The second proprietary feature of the recommendation scenario, the second shared feature, the scene feature corresponding to the recommendation scenario, the second resource feature, and the activity are fused through the gating network corresponding to the recommendation scenario to obtain the second fused feature.

13. The method according to claim 12, wherein: The fusing, through the gating network corresponding to the recommendation scenario, the second specific feature of the recommendation scenario, the second shared feature, the scenario feature corresponding to the recommendation scenario, the second resource feature, and the activity to obtain the second fused feature includes: Concatenating the scene feature corresponding to the recommendation scene, the second resource feature, and the embedding representation corresponding to the activity to obtain a second concatenated embedding representation; The second specific feature of the recommendation scene, the second shared feature, and the second concatenated embedding representation are fused through the gating network corresponding to the recommendation scene to obtain the second fused feature.

14. The method according to any one of claims 10 to 13, wherein: The style model further includes an embedding layer and a concatenation layer. Inputting the second object feature and the second resource feature into the first expert network corresponding to the recommendation scenario, and obtaining the second specific feature of the recommendation scenario through the first expert network corresponding to the recommendation scenario, includes: Embedding the second object feature and the second resource feature respectively through the embedding layer to obtain a fourth embedded representation of the second object feature and a fifth embedded representation of the second resource feature; splicing the fourth embedded representation and the fifth embedded representation through the splicing layer to obtain a sixth embedded representation; The sixth embedded representation is input into a first expert network corresponding to the recommendation scenario, and a second specific feature of the recommendation scenario is obtained through the first expert network corresponding to the recommendation scenario.

15. The method according to claim 14, wherein The embedding layer includes a first embedding layer and a second embedding layer, and embedding the second object feature and the second resource feature respectively through the embedding layer to obtain a fourth embedded representation of the second object feature and a fifth embedded representation of the second resource feature, including: Embedding the second object feature through the first embedding layer to obtain a fourth embedded representation of the second object feature; The second resource feature is embedded and represented by the second embedding layer to obtain a fifth embedded representation of the second resource feature.

16. The method according to claim 15, wherein The style model further includes: an encoding layer, the second object feature includes: temporal behavior sequence data, and the embedding representation of the second object feature by the first embedding layer to obtain a fourth embedded representation of the second object feature includes: Performing encoding processing on the temporal behavior sequence data through the encoding layer to obtain a behavior sequence encoding result corresponding to the temporal behavior sequence data; The behavior sequence encoding result is embedded and represented by the first embedding layer to obtain the fourth embedded representation.

17. A device for determining a cover image style, comprising: The first acquisition module is used to acquire target objects to be recommended corresponding to the target resource in multiple recommendation scenarios; A second acquisition module is configured to acquire scene features corresponding to each of the plurality of recommended scenes, a first resource feature of the target resource, and a first object feature of the target object; The first determination module is used to use a trained style model to determine the target cover image styles corresponding to the target resource in the multiple recommended scenes according to the scene characteristics, the first resource characteristics and the first object characteristics.

18. A training device for a style model, comprising: A third acquisition module is configured to acquire training data, wherein the training data includes: sample cover image styles used by the sample resource in multiple recommendation scenarios, second object features of sample objects that interacted with the sample resource in the recommendation scenarios, scene features corresponding to each of the multiple recommendation scenarios, and second resource features of the sample resource; A second determination module is configured to determine, using a style model, predicted cover image styles corresponding to the sample resource in the plurality of recommendation scenarios according to the scene feature, the second resource feature, and the second object feature; The training module is used to train the style model according to the predicted cover image styles and sample cover image styles corresponding to the sample resources in the multiple recommendation scenarios to obtain a trained style model.

19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 8, or the method of any one of claims 9 to 16.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 8, or the method according to any one of claims 9 to 16.

21. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 8, or the method according to any one of claims 9 to 16.