Cover selection model training method and apparatus, electronic device, and storage medium
By training the model in multiple stages and combining multi-label quality annotation and interactive behavior features, the problem of the singleness of cover image selection in the information flow recommendation system is solved, and the quality assessment and attractiveness prediction of cover images are realized, thereby improving click-through rate and content exposure efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2026-04-09
- Publication Date
- 2026-07-10
Smart Images

Figure CN122368673A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, particularly to fields such as artificial intelligence and computer vision, and can be used in application scenarios such as information flow recommendation. Specifically, it relates to cover selection model training methods, devices, electronic devices, and storage media. Background Technology
[0002] In existing information flow recommendation systems, each content item is typically configured with only one cover image, which may come from the original image uploaded by the user or be a unified cover selected from multiple candidate images by an offline basic optimization model. Summary of the Invention
[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for training a cover selection model.
[0004] According to a first aspect of this disclosure, a cover selection model training method is provided, comprising: acquiring a first training sample set, a second training sample set, and a third training sample set; the first training sample set includes multiple first sample images and multi-label quality annotations corresponding to each first sample image; the second training sample set includes multiple second sample images and effect annotations corresponding to each second sample image; the third training sample set includes multi-dimensional features and interactive behavior features, wherein the multi-dimensional features include at least target object features, resource features, and a cover image feature set corresponding to each resource cover image; based on historical images, determining a quantitative judgment standard corresponding to a preset quality type and an effect prediction standard corresponding to a preset attraction dimension; training an initial multimodal large model for quality assessment based on the quantitative judgment standard and the first training sample set and a first loss function to obtain a first model; training the first model for click ranking based on the effect prediction standard and the second training sample set and a second loss function to obtain a second model; and training the second model for interaction prediction based on the third training sample set and a third loss function to obtain a cover selection model.
[0005] According to a second aspect of this disclosure, a cover selection method is provided, comprising: determining the features of the object to be recommended based on the object to be recommended; determining the features of the target resource and the feature set of the multiple candidate cover images based on the target resource and the multiple candidate cover images corresponding to the target resource; obtaining the prediction effect score of each candidate cover image for the object to be recommended using a cover selection model based on the features of the object to be recommended, the features of the target resource, and the feature set of the multiple candidate cover images; the cover selection model is obtained by the method of the first aspect; traversing each prediction effect score, selecting the candidate cover image with the highest prediction effect score as the target cover image, and displaying the target cover image to the object to be recommended.
[0006] According to a third aspect of this disclosure, a cover selection model training apparatus is provided, comprising: a sample acquisition module for acquiring a first training sample set, a second training sample set, and a third training sample set; the first training sample set includes multiple first sample images and multi-label quality annotations corresponding to each first sample image; the second training sample set includes multiple second sample images and effect annotations corresponding to each second sample image; the third training sample set includes multi-dimensional features and interactive behavior features, wherein the multi-dimensional features include at least target object features, resource features, and a cover image feature set corresponding to each resource cover image; and a standard setting module for determining, based on historical images, respectively... The system includes: a quantitative judgment standard corresponding to each preset quality type and an effect prediction standard corresponding to each preset attraction dimension; a quality assessment training module, used to train the initial multimodal large model for multi-label low-quality identification quality assessment based on the quantitative judgment standard and the first training sample set, using the first loss function, to obtain the first model; a click ranking training module, used to train the first model for click ranking based on the effect prediction standard and the second training sample set, using the second loss function, to obtain the second model; and an interaction prediction training module, used to train the second model for recommendation interaction prediction based on the third loss function using the third training sample set, to obtain the cover selection model.
[0007] According to a fourth aspect of this disclosure, a cover selection device is provided, comprising: a first feature determination module, configured to determine the current user features of the target user based on the target user to be recommended; a second feature determination module, configured to determine the target resource features and a set of multiple candidate cover image features based on the target resource and multiple candidate cover images corresponding to the target resource; a model prediction module, configured to obtain a prediction performance score for each candidate cover image for the current user to be recommended based on the current user features of the target user, the target resource features, and the set of multiple candidate cover image features, using a cover selection model; the cover selection model is obtained by the method of the first aspect; and an image display module, configured to iterate through each prediction performance score, select the candidate cover image with the highest prediction performance score as the target cover image, and display the target cover image to the current user to be recommended.
[0008] According to a fifth aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the methods described in the embodiments of this disclosure.
[0009] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.
[0010] According to a seventh aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.
[0011] By adopting the scheme disclosed herein, the model is trained through multiple stages to enable it to have the ability to assess quality, predict attractiveness, and personalize matching. It dynamically selects the cover with the highest click potential online in combination with context, upgrading the cover selection from a static rule to a dynamic decision, effectively improving click-through rate and content distribution efficiency.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a flowchart illustrating the cover selection model training method according to an embodiment of the present disclosure; Figure 2 This is a schematic flowchart of a cover selection method according to an embodiment of the present disclosure; Figure 3 This is another flowchart illustrating the cover selection method according to an embodiment of this disclosure; Figure 4 This is a schematic diagram of the interactive prediction training process according to the embodiments of this disclosure; Figure 5 This is another flowchart illustrating the cover selection method according to an embodiment of this disclosure; Figure 6 This is a schematic diagram of the structure of a cover selection model training device according to an embodiment of the present disclosure; Figure 7 This is a schematic diagram of the cover selection device according to an embodiment of the present disclosure; Figure 8 This is a schematic diagram of a scenario for a cover selection model training method according to an embodiment of the present disclosure; Figure 9 This is a schematic diagram of a cover selection method according to an embodiment of the present disclosure; Figure 10 This is a structural diagram of an electronic device used to implement the cover selection model training method and / or cover selection method of the embodiments of this disclosure. Detailed Implementation
[0014] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0015] In this document, the term "and / or" merely describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The term "at least one" in this document indicates any combination of at least two of a plurality of elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this document refer to and distinguish between multiple similar technical terms, not to restrict the order or to limit there to only two. For example, "first feature" and "second feature" refer to two categories / two features; the first feature can be one or more, and the second feature can also be one or more.
[0016] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0017] Before introducing the technical solutions of the embodiments of this disclosure, the technical terms that may be used in this disclosure will be further explained: Cover image: refers to the representative image displayed along with the content in the information feed, used as a thumbnail or main image of the content in the list or recommendation feed.
[0018] Target audience: refers to the group of people for whom the best cover image needs to be customized or selected, which can be a specific user or user group.
[0019] Resources: These refer to content resources within the information flow system, namely various content items that can be recommended and displayed, including short videos, long videos, text and image content, posts, product cards, etc. Each resource typically corresponds to one content item and can be associated with multiple candidate cover images, text titles, tags, and other information.
[0020] In related technologies, the effectiveness of cover image recognition highly depends on the accuracy of multiple sub-models. If a sub-model fails to recognize the image accurately, a cover that should be judged as low-quality may be mistakenly selected, thus affecting the overall content distribution quality. Furthermore, the modeling for cover image quality recognition focuses on the aesthetics and objective quality attributes of the image itself, lacking a direct characterization of its attractiveness. This often results in the model deeming a cover high-quality but not necessarily possessing high click-through potential, or aligning with the interests and preferences of different users. In addition, since each piece of content has only a single cover image, if that image is judged as low-quality, it is easily filtered out or scattered in diversity strategies. Even if the content itself is high-quality, it may not receive sufficient exposure, affecting its ranking position and users' discovery of high-quality content. Moreover, existing offline optimization only makes a uniform selection based on content dimensions. In the online stage, it lacks the ability to make personalized cover image decisions in real time based on user profiles, interests, and real-time context. It cannot dynamically select the cover image that best stimulates the click-through intention of different users, resulting in significant deficiencies in the personalization and attractiveness of cover display.
[0021] To at least partially address one or more of the aforementioned problems and other potential issues, this disclosure proposes a cover selection model training method. Through multi-stage training, the model acquires the capabilities of quality assessment, attractiveness prediction, and personalized matching. It dynamically selects the cover with the highest click potential online in conjunction with the context, effectively improving click-through rate and content distribution efficiency.
[0022] This disclosure provides a method for training a cover selection model. Figure 1 This is a flowchart illustrating a cover selection model training method according to an embodiment of the present disclosure. This cover selection model training method can be applied to a cover selection model training device. The cover selection model training device is located in an electronic device. This electronic device includes, but is not limited to, fixed devices and / or mobile devices. For example, fixed devices include, but are not limited to, servers, which can be cloud servers or ordinary servers. For example, mobile devices include, but are not limited to, information flow recommendation devices, which can be mobile phones, tablets, etc. In some possible implementations, the cover selection model training method can also be implemented by a processor calling computer-readable instructions stored in memory. Figure 1 As shown, the training method for this cover selection model includes: S101. Obtain a first training sample set, a second training sample set, and a third training sample set; the first training sample set includes multiple first sample images and multi-label quality annotations corresponding to each first sample image; the second training sample set includes multiple second sample images and effect annotations corresponding to each second sample image; the third training sample set includes multi-dimensional features and interactive behavior features; the multi-dimensional features include at least target object features, resource features, and a cover image feature set corresponding to each resource cover image.
[0023] S102. Based on historical images, determine the quantitative judgment criteria corresponding to the preset quality types and the effect prediction criteria corresponding to the preset attraction dimensions.
[0024] S103. Based on the quantitative judgment criteria and the first training sample set, the initial multimodal large model is trained for quality assessment using the first loss function to obtain the first model.
[0025] S104. Based on the effect prediction criteria and the second training sample set, the first model is trained by click ranking based on the second loss function to obtain the second model.
[0026] S105. Using the third training sample set, the second model is interactively predicted and trained based on the third loss function to obtain the cover selection model.
[0027] Here, the first sample image refers to a type of sample image used to train the cover quality assessment capability, derived from a large number of historical cover images. Multi-label quality annotation refers to the labeling results of each first sample image across multiple quality dimensions; that is, the same image can be assigned multiple quality labels simultaneously, such as whether the clarity is acceptable, whether there is stretching or distortion, whether the image is too cluttered, whether there are black borders, and whether the composition is aesthetically pleasing, thus forming a multi-dimensional quality annotation. The second sample image refers to another type of sample image used to train the cover attractiveness and effect prediction capabilities, which can be selected from the first sample images. Effect annotation refers to the effect-related label or numerical value corresponding to each second sample image, which may include, but is not limited to, click-through rate, conversion rate, completion rate, user dwell time, or a calculated comprehensive score. Target object features refer to publicly available information feature vectors related to users or user groups receiving the recommendation service, which can be used to characterize user profiles. Resource features refer to features related to the recommended content resource itself, used to characterize the content's own attributes. The cover image feature set refers to a set of feature vectors extracted from each candidate cover image for each resource. Interactive behavior characteristics refer to the interactive behavior characteristics between the target object and the resource and its cover image during the historical recommendation and display process. These characteristics can be statistically analyzed and coded according to user dimension, resource dimension, or cover image dimension.
[0028] In this embodiment, three training sample sets can be constructed from business data. For example, cover images can be extracted from a historical content library, and each image can be multi-labeled for quality control across multiple dimensions such as clarity, stretching distortion, image density, and border conditions, using manual annotation or a rule-based model, thus constructing a first training sample set. Further, cover images with real interaction data can be selected from historical recommendation logs as second sample images, and effect annotations can be generated for each second sample image based on actual exposure, clicks, completion times, and dwell times, thus constructing a second training sample set. Even further, target object features can be extracted based on the user's publicly available information, resource features can be extracted from the content side, and cover image feature sets can be extracted for each cover image associated with each piece of content. Then, interaction behavior features can be constructed based on the historical interaction behavior between the user and the content and its cover images, thus constructing a third training sample set.
[0029] Here, "historical images" refers to a large amount of image data that has been collected and stored, including cover images that have been filtered and displayed on the platform. "Preset quality types" refers to several predefined image quality dimensions or categories, such as sharpness, whether it is stretched or distorted, whether there is obvious noise or compression, whether the composition is reasonable, and whether the image is too dense. "Preset attraction dimensions" refers to several predefined dimensions for measuring the appeal of cover images to users, such as click-through rate, completion rate, conversion rate, and dwell time, which can characterize the driving effect of images on user behavior from an effectiveness perspective.
[0030] In this embodiment, a historical image set containing a large number of cover images, their historical quality annotations, and user interaction data can be collected and organized first. These historical images are then statistically analyzed according to pre-defined quality types such as sharpness, stretching / distortion, image density, and borders. By observing the relationship between feature distribution and manual annotations across each quality dimension, a quantitative judgment standard corresponding to each quality type is gradually determined. For example, a threshold between sharpness and blurriness can be set for the sharpness feature, and a judgment range between reasonable and excessively dense image density can be set for image density. Multiple quality types can also be unified. Furthermore, based on the click, completion, conversion, and dwell time results of these historical images in actual exposure, effect data mining can be performed around pre-defined attraction dimensions. By constructing statistical indicators and simple prediction models, key feature patterns and scoring ranges of image performance under each attraction dimension can be extracted, thereby forming an effect prediction standard corresponding one-to-one with each pre-defined attraction dimension.
[0031] In this embodiment, a first sample image from the first training sample set can be input into an initial multimodal large model based on a quantification criterion. Subsequently, the model provides a corresponding score or probability prediction for each quality type and compares it with the pre-defined multi-label quality annotations for that image in the first training sample set. A pre-designed first loss function is used to calculate the error between the model's prediction and the true annotations. Specifically, the first loss function can be calculated separately for each quality dimension and used to form an overall optimization objective through weighted summation or multi-task joint learning. After multiple iterations of training, a first model with image quality assessment capabilities is obtained.
[0032] In this embodiment, second sample images from a second training sample set can be input into a first model based on an effect prediction criterion. Subsequently, the model outputs one or more scores representing click appeal or effect potential for each image, and compares these scores with the effect annotations recorded in the second training sample set. A pre-designed second loss function is used to calculate the error between the predicted result and the actual effect. Specifically, the second loss function can characterize the effect error of a single image using point regression or classification, or it can construct image pairs or image lists with the same content or exposure scene for sorting and optimization. After multiple iterations of training, a second model is obtained that simultaneously possesses image quality assessment and overall appeal prediction capabilities.
[0033] In this embodiment, a second model can be used as a foundation, and multidimensional features from a third training sample set can be input into the second model. Subsequently, the model outputs interaction prediction results for each set of multidimensional features, and compares these prediction results with the actual interaction behaviors recorded in the third training sample set, calculating the prediction error using a predefined third loss function. After multiple iterations of training, a cover selection model is obtained that simultaneously possesses image quality assessment, overall attractiveness prediction capabilities, and interaction behavior prediction capabilities.
[0034] The technical solution of this disclosure utilizes a first training sample set with multi-label quality annotations to enable a multimodal large model to possess image quality assessment capabilities. Combined with a second training sample set with effect annotations, the model is guided to simultaneously possess image quality assessment and overall attractiveness prediction capabilities. Then, through a third training sample set, the model is trained for personalized effect prediction, resulting in a cover selection model that simultaneously possesses image quality assessment, overall attractiveness prediction, and interaction behavior prediction capabilities. This allows for the real-time selection of cover images with the highest click-through potential for different users and content, taking into account user profiles and contextual features, thereby improving personalized matching capabilities while ensuring cover quality. This effectively alleviates the problem of insufficient exposure of high-quality content caused by a single cover image, comprehensively optimizing the overall effect of the information flow recommendation system while improving click-through rates and content distribution efficiency.
[0035] In some embodiments, the first training sample set is pre-constructed in the following manner: according to a preset quality type, multiple quality labels corresponding to each first sample image are obtained; the multiple quality labels correspond one-to-one with the preset quality type; according to the multiple quality labels corresponding to each first sample image, a quality conclusion for each first sample image is determined; the quality conclusion is qualified or low quality; according to the multiple quality labels and the quality conclusion, a multi-label quality label corresponding to each first sample image is generated; and the first training sample set is constructed according to each first sample image and its corresponding multi-label quality label.
[0036] Here, multiple quality labels refer to the labeling results given for the same first sample image in different preset quality type dimensions.
[0037] In this embodiment of the disclosure, each first sample image can be pre-identified and labeled in various quality types. For example, this process can be implemented through manual review, rule engines, or automatic identification by specialized sub-models. For instance, a sharpness detection model can output "sharp" or "blurry," a deformation detection model can output "normal" or "stretched," and a content security model can output "normal" or "risky." These specific labeling results obtained in each quality type dimension are then summarized to form multiple quality labels corresponding to the first sample image, ensuring that each preset quality type has a clear labeling record.
[0038] Here, the quality conclusion refers to the overall quality judgment result given for each first sample image based on multiple quality labels, which is used to classify the image in a global dimension. In this embodiment of the disclosure, the quality conclusion can be obtained by comprehensively evaluating the labels of each quality type, and finally determining whether the overall image meets the quality requirements for cover use, which can be classified as qualified or low quality.
[0039] In this embodiment of the disclosure, for each first sample image, multiple quality labels on all preset quality types can be read, and the overall quality conclusion of the image can be derived based on pre-configured quality judgment rules or strategies. For example, weights and thresholds can be set for different quality types. For instance, whether the content is risky, whether it is severely blurry, or whether it is severely distorted can be defined as hard non-compliance items. Once the relevant label is non-compliance, the image is directly judged as low quality. For other soft indicators such as image density and aesthetic quality, a weighted sum can be calculated based on the scores. When the comprehensive score is higher than the threshold, it is judged as qualified; when it is lower than the threshold, it is judged as low quality.
[0040] In this embodiment, the multiple quality labels obtained for each first sample image and their overall quality conclusions can be uniformly organized and encoded to construct a multi-label quality label vector suitable for model training. For example, one or more label bits can be assigned to each preset quality type to store the corresponding quality labeling results, and an additional label bit can be reserved to record the overall quality conclusion. These labels are then organized into a multi-dimensional label vector in a fixed order. Missing or inapplicable quality types can be filled with default values or special markers, thereby forming a multi-label structure that includes both fine-grained quality information and comprehensive quality conclusions.
[0041] In this embodiment of the disclosure, all first sample images can be paired with their respective generated multi-label quality annotations to form a standardized training sample set. For example, each first sample image can undergo image preprocessing and be encapsulated together with its corresponding multi-label quality annotation into a training sample record, ultimately integrating multiple processed samples into a first training sample set.
[0042] In this way, by first obtaining multiple quality labels for each image based on a preset quality type, then combining these fine-grained labels to derive an overall quality conclusion, and finally integrating the two into a multi-label quality label to construct the first training sample set, it is possible to simultaneously retain the local dimensional information of image quality and the global quality assessment at the training data level. This provides a supervisory signal for subsequent model learning, which is beneficial to improving the accuracy and robustness of the model and reducing the risk of misjudgment.
[0043] In some embodiments, the second training sample set is pre-constructed in the following manner: acquiring any second sample image and its corresponding historical interaction data; the quality conclusion of the second sample image is qualified, and the historical interaction data includes at least one of the statistically obtained historical click rate, historical playback completion rate, and historical dwell time; scoring the historical interaction data according to a preset scoring rule, and using the scoring result as the effect label corresponding to the second sample image; and constructing a second training sample set based on each second sample image and its corresponding effect label.
[0044] Here, historical interaction data refers to the user behavior statistics accumulated during the past real exposure of the second sample image, which is used to reflect the actual feedback of users when they see the image as a cover.
[0045] In this embodiment, images deemed qualified based on completed quality assessments can be selected first from the image set. These images are then used as second sample images to ensure that the samples used for effect modeling are not affected by significant low-quality factors. Subsequently, each second sample image can be associated with user behavior data accumulated during its online display to statistically obtain historical interaction data corresponding to that image. This includes at least one of the following: historical click-through rate when the image is used as a cover image, historical completion rate, and historical dwell time after a user clicks into the content. Specifically, during the statistical process, a minimum exposure threshold can be set to filter out second sample images with insufficient samples.
[0046] Here, the preset scoring rules refer to a predefined set of rules that convert historical interaction data into a unified numerical or graded score. This system maps behavioral indicators of different dimensions and scales into performance scores or labels that can be directly used for model training. In this embodiment, the preset scoring rules may include specific calculation methods such as normalization of each indicator, segmented scoring, weighted summation, and graded mapping.
[0047] In this embodiment, weights can be assigned to each indicator according to preset scoring rules, and the behavioral performance of multiple dimensions can be integrated into a single effect score according to a certain formula. Alternatively, each indicator can be mapped to multiple levels according to its percentile range, and then further combined into multi-level labels. Finally, the calculated numerical value or level result can be used as the effect label uniquely corresponding to the second sample image to characterize the overall attractiveness and content-carrying performance of the image in the historical scene.
[0048] In this embodiment, all second sample images can be paired with their respective generated effect labels to form a standardized training sample set. For example, each second sample image can undergo image preprocessing and be encapsulated together with its corresponding effect label into a training sample record. Finally, multiple processed samples are integrated into a second training sample set. Specifically, image samples with severely insufficient historical interaction data or abnormal scores can be removed. If necessary, sampling balance can be performed on different effect level ranges to avoid excessive bias towards high-frequency or low-frequency samples during training.
[0049] In this way, by selecting only images with qualified quality conclusions as second sample images and constructing effect labels based on their real historical interaction data, the quality factors and attractiveness factors are effectively decoupled at the training data level, avoiding interference caused by obviously low-quality images to effect modeling. At the same time, by introducing multi-dimensional behavioral indicators and using preset scoring rules to uniformly map them into effect scores, the supervision signals obtained by the model comprehensively consider the depth of content consumption and experience quality.
[0050] In some embodiments, the third training sample set is pre-constructed by: acquiring interaction data; the interaction data includes at least the interaction behaviors between the target object and the resource; for any interaction behavior, extracting the corresponding resource features, target object features, interaction behavior features, and cover image feature set; the cover image feature set includes at least prior features, posterior features, and cross features; and constructing the third training sample set based on each interaction behavior feature and its corresponding target object features, resource features, and cover image feature set.
[0051] Here, interaction data refers to the set of behavioral data generated and recorded during the operation of an information flow recommendation system when a target object and a resource actually come into contact, including during exposure, clicks, playback, dwell time, and interactions. Interactive behavior refers to a specific user behavior event or a type of user behavior event within the interaction data, and can be the basic unit for constructing behavioral samples between users and content / cover images.
[0052] In this embodiment, the behavior records of the target object and resources during recommendation, display, and playback can first be collected and organized from the logs generated by the online service. Exposure logs, click logs, playback logs, interaction logs, and other multi-source data are correlated and cleaned to remove obviously abnormal or duplicate records, resulting in a structured interaction data table. Specifically, the interaction data includes at least the interaction behaviors between the target object and the resource, such as whether a user clicked on a piece of content, the playback duration after clicking and whether the playback was completed, and whether likes, comments, or shares were generated.
[0053] Here, prior features refer to cover image-related features that can be obtained based on static information or offline models before the interaction occurs, independent of the real-time behavior of specific users. Posterior features refer to cover image effect features obtained by statistically analyzing or modeling these historical interaction behaviors after the cover image has been actually exposed to users and generated a certain scale of interaction. Cross features refer to features constructed by combining target object features or resource features with cover image features according to a certain combination relationship, used to characterize the interaction relationship between users or resources and the cover image.
[0054] In this embodiment, each specific interaction record can be treated as a processing unit. Resource features associated with the interaction are read, and corresponding target object features are obtained from the user profiling system. Furthermore, interaction behavior features can be extracted from the interaction data itself to characterize the outcome and context of the specific interaction. Further, based on the cover image used in the interaction, prior features of the cover image can be read from offline visual models and quality models. Posterior features of the cover image under overall traffic can be extracted from historical statistics. These features are then combined with those between the user and the cover image, and between resources and the cover image to generate cross-features, thereby obtaining a cover image feature set.
[0055] In some implementations, the process of constructing a cover image feature set can begin by extracting prior features for each cover image from the image content itself and offline quality and aesthetic evaluation models. These prior features reflect the objective quality and content attributes of the cover image under unexposed conditions. Subsequently, posterior features can be obtained based on the overall performance statistics of the cover image during the platform's historical recommendation and display processes. These posterior features reflect the cover image's performance among real user groups. Next, the target object features, resource features, and the prior and posterior features of the cover image can be combined and transformed to generate cross-features, which directly characterize the preference patterns of a certain type of user for a certain type of content under a certain type of cover image.
[0056] In this embodiment, target object features, resource features, cover image feature set, and interaction behavior features can be integrated according to a unified data format. Various features are encoded, normalized, and missing value processing is performed. Interaction behavior features are used as label signals, while other contextual behavior features are retained as part of the input features, thereby constructing a training sample. Finally, each training sample record can be organized to construct a third training sample set.
[0057] In some embodiments, resource features may include author, category, title, content tags, points of interest, etc., for in-depth characterization of inherent attributes of the content. In cover image features, prior features may include cover aesthetic score, cover element tags, clustering based on multiple multimodal vectors, and other prior cover information; posterior features may include posterior information of the cover image; combined cross features may include target object-cover image cross features and resource-cover image cross features.
[0058] In this way, the features of the target object, resource features, and prior, posterior, and cross features of the cover image can be fully integrated at the training data level. With real interactive behavior as the supervision signal, the cover selection model can not only understand the quality and content features of the cover image itself during training, but also capture the differentiated preferences of different audience groups for different content and cover combinations. This helps the model to more accurately predict the click probability and completion probability of a specific target under given resources and cover, realize personalized and dynamic selection of cover images in the online stage, and thus improve the overall click-through rate and content completion rate of the recommendation system.
[0059] In some embodiments, extracting interaction behavior features includes: obtaining the occurrence time of each original interaction behavior based on interaction data; and applying different time decay factors to the original interaction behaviors at different times based on a preset time decay function to obtain interaction behavior features.
[0060] Here, the original interaction behavior refers to a single behavior record that is directly read, which can be stored in the form of log lines. The occurrence time refers to the specific time information corresponding to each original interaction behavior, which can be recorded in the form of a timestamp or standard date and time.
[0061] In this embodiment, a structured interactive data table can be read from a log or data warehouse first, and the corresponding time field can be parsed from the records and converted into an occurrence time in a unified format. In particular, obviously abnormal time records can also be cleaned and corrected.
[0062] Here, the preset time decay function refers to a pre-defined mathematical function used to automatically calculate the degree to which the weight of a behavior decays over time based on the time interval between the behavior's occurrence time and the current time. This function can be an exponential decay function, a linear decay function, or a piecewise decay function. The time decay factor refers to the numerical weight calculated by the time decay function for a specific original interaction behavior. It corresponds one-to-one with the behavior's occurrence time, and its value is between 0 and 1. It is used to scale the contribution of the interaction behavior when constructing its features, thus amplifying the impact of recent behaviors and relatively weakening the impact of long-term behaviors.
[0063] In this embodiment, a time decay function suitable for the business scenario can be selected or configured first. Using the current training time or feature calculation time as a reference time point, the time difference between the occurrence time and the reference time is calculated for each original interaction behavior. This time difference is then used as input to the preset time decay function to calculate the corresponding time decay factor. Subsequently, this time decay factor is used to scale the contribution of the original interaction behavior in feature statistics. The count or duration of each behavior is multiplied by its corresponding time decay factor, and then all behaviors are summed or aggregated to obtain a set of time-weighted interaction behavior features.
[0064] Thus, introducing a time decay mechanism when extracting interactive behavior features effectively distinguishes the importance of recent and distant interactions in a user's historical behavior. This significantly increases the weight of recent clicks, views, and dwell times in the features, enabling the model to quickly capture and respond to the latest user interest trends. Simultaneously, the gradual reduction in weighting of distant behaviors avoids the excessive influence of outdated preferences on model predictions, improving the sensitivity of features to short-term interest changes. This overall enhances the cover selection model's ability to model the temporal evolution of interests, improving the accuracy of personalized predictions and its adaptability to real-time scenarios.
[0065] In some embodiments, the initial multimodal large model is trained for quality assessment based on a first loss function according to a quantification criterion and a first training sample set to obtain a first model, including: constructing a first prompt word according to a quantification criterion; and using the first prompt word and the first training sample set to train the initial multimodal large model for quality assessment based on a first loss function to obtain a first model; wherein the first loss function is a loss function based on multi-label binary cross-entropy.
[0066] In this embodiment, the meaning, judgment criteria, thresholds, or grade classifications of each quality dimension can first be summarized and described in text based on quantitative judgment standards. Subsequently, these descriptions can be transformed into natural language instructions or question templates adapted to multimodal large models. Specifically, task descriptions and necessary output format constraints can be added to the prompts, thereby forming structured and semantically clear first prompts.
[0067] In this embodiment, the constructed first prompt word can be combined with each first sample image in the first training sample set and input into an initial multimodal large model. The model, through the synergy of visual encoding and text instructions, performs multi-label prediction of the image's pass / fail status in each preset quality type, outputting probability values for each quality dimension. Subsequently, the multi-dimensional quality probability vector output by the model can be compared with the pre-given multi-label quality annotations of the image in the first training sample set. A first loss function is used to calculate the binary cross-entropy loss for each quality label, and the losses for all quality dimensions are weighted and summed or averaged to form a loss value for the multi-label task. The model parameters are iterated repeatedly on a large scale of samples, gradually adjusting them to make the prediction results in all quality dimensions as close as possible to the true annotations, ultimately obtaining a first model that converges and performs well on the image quality assessment task.
[0068] Thus, by using the first cue word, which is strictly aligned with the quantification criteria, the quality standards are explicitly injected into the multimodal large model in natural language. This allows the model to accurately understand the task objectives and the judgment criteria for each quality dimension during training, avoiding semantic biases that may arise from relying solely on implicit labels. By employing a first loss function based on multi-label binary cross-entropy, each quality dimension is modeled separately and jointly optimized within a unified framework. This enables the model to simultaneously learn and weigh the recognition capabilities of multiple quality dimensions, improving its ability to distinguish complex quality issues in detail and enhancing the reliability of overall judgments.
[0069] In some embodiments, determining a quantization criterion corresponding one-to-one with a preset quality type based on historical images includes: performing image quality analysis on any historical image to determine the quality index of the historical image under each preset quality type; determining the maximum and minimum quality indices corresponding to the preset quality type based on the quality indices of multiple historical images under the same preset quality type, and determining the quantization range of the preset quality type based on the maximum and minimum quality indices; and using the quantization range corresponding to each preset quality type as the quantization criterion.
[0070] Here, image quality analysis refers to the process of quantitatively evaluating and judging a given image across various preset quality categories using image processing algorithms, machine learning models, or rule systems. Quality metrics refer to the specific numerical values calculated for each preset quality category.
[0071] In this embodiment, multiple historical images can be selected first as the sample basis for constructing the quantitative judgment criteria. Then, for any historical image, the corresponding quality analysis module or model can be called sequentially for evaluation in each preset quality type dimension. Next, after obtaining the analysis results for each quality type, these results can be organized into a set of structured quality indicators, thereby determining the corresponding quality indicator value for each historical image under each preset quality type.
[0072] Here, the quantization interval refers to the range of values determined by summarizing the quality indicators of a large number of historical images and based on the distribution of quality indicators of all images under a certain preset quality type.
[0073] In this embodiment, for each preset quality type, all historical images with calculated quality indices under that quality type can be collected to form a sample set of indices for that quality type. Then, statistical analysis can be performed on this set. After necessary data cleaning and outlier handling, the minimum and maximum quality index values for that quality type can be found in the index set. The quantization interval for that preset quality type is then determined by using the minimum quality index value as the left endpoint of the quantization interval and the maximum quality index value as the right endpoint.
[0074] In this embodiment of the disclosure, the quantization range corresponding to each preset quality type can be uniformly organized and configured, and solidified into the quantization judgment standard of that quality type.
[0075] In some implementations, a strictly enumerated negative list strategy can be used to determine the quantitative judgment criteria. For example, 12 low-quality patterns can be defined, including unattractive borders, poor composition, blurry images, truncated images, disturbing images, poor layout, large portraits, dense text, obstruction, low-quality screenshots, and marketing content, with each pattern accompanied by quantifiable judgment criteria. Ultimately, the model outputs 13 labels (0-12), where 0 represents acceptable and the other numbers represent the corresponding low-quality type.
[0076] Thus, by utilizing the quality index distribution of a large number of historical images to determine the maximum and minimum values and overall quantization range for each quality type, the actual image quality level in the system can be objectively reflected, avoiding subjective bias and instability caused by manually setting thresholds. After defining the quantization range for each quality type, the numerical meaning of the quality indices has a unified and comparable scale, facilitating the maintenance of standard consistency and interpretability.
[0077] In some embodiments, a second model is obtained by training a first model for click ranking based on a second loss function according to an effect prediction criterion and a second training sample set, including: constructing a second prompt word according to the effect prediction criterion; and training the first model for click ranking based on the second loss function using the second prompt word and the second training sample set to obtain the second model.
[0078] In this embodiment, the effect dimension types, evaluation criteria, and output formats can first be organized and described in text based on the effect prediction standards. Subsequently, these descriptions can be transformed into natural language instructions or question templates adapted to a multimodal large model. Specifically, task descriptions and necessary output format constraints can be added to the prompt words, thereby forming a structured and semantically clear second prompt word.
[0079] In this embodiment, the constructed second prompt word is combined with each training sample in the second training sample set and input into a first model that already possesses quality assessment capabilities. This allows the model to output a corresponding click or effect prediction value for the sample, based on an understanding of the task requirements. Subsequently, the model's prediction result is compared with the pre-defined effect label for that sample in the second training sample set, and the error is calculated using a preset second loss function. For example, cross-entropy loss based on click probability, regression loss based on effect score, or ranking loss based on sample pairs can be used. The errors of each sample or sample pair are accumulated as the overall loss for the current training batch. After multiple training iterations, the model's ability to predict the effect and click ranking of the cover image when given the second prompt word is significantly improved, ultimately resulting in a second model that simultaneously possesses image quality assessment and overall attractiveness prediction capabilities.
[0080] Thus, by using a second cue word closely aligned with the performance prediction criteria, the business objectives of performance prediction and ranking are explicitly injected into the first model, which already possesses quality understanding capabilities, in natural language form. This enables the model to accurately understand its own training task based on multimodal representations. By utilizing a second training sample set labeled with real performance data and directly optimizing the deviation between the predicted results and the actual performance using a second loss function, quality cognition is transformed into the ability to rank business objectives.
[0081] In some embodiments, determining effect prediction standards corresponding one-to-one with preset attraction dimensions based on historical images includes: performing preference analysis on any historical image to determine the effect index of the historical image under each preset attraction dimension; determining the maximum and minimum effect indices corresponding to the preset attraction dimension based on the effect indices of multiple historical images under the same preset attraction dimension, and determining the quantization range of the preset attraction dimension based on the maximum and minimum effect indices; and using the quantization range corresponding to each preset attraction dimension as the effect prediction standard.
[0082] Here, preference analysis refers to the process of analyzing and mining users' preferences for certain attractive features based on their interactive behavior in real-world scenarios, combined with the attributes of the cover image itself across different attractiveness dimensions. Performance metrics refer to numerical measures calculated for a historical image for each preset attractiveness dimension, used to measure the actual performance of that image within that dimension.
[0083] In this embodiment, multiple historical images can be used as the sample base for constructing the effect prediction standard, and associated with their corresponding interaction logs. Then, for any historical image, the corresponding preference analysis module or model can be invoked sequentially for feature extraction on each preset attraction dimension. Next, after obtaining the features of each attraction dimension and the corresponding user behavior data, an effect metric can be constructed for each historical image on each attraction dimension. For example, the overall click-through rate, completion rate, and average dwell time of the image can be directly used as the basic indicators, or after filtering out a subset of traffic with certain attraction features, the increase in click-through rate and completion rate on that subset can be statistically analyzed as the effect metric for that attraction dimension, thereby determining the corresponding effect metric value for each historical image under each attraction dimension.
[0084] In this embodiment, for a given preset attraction dimension, all historical images with calculated effect index values under that dimension can be collected to form an effect index sample set. Then, statistical analysis can be performed on this set. After necessary data cleaning and outlier handling, the minimum and maximum effect index values for that attraction dimension are found in the index set. The quantization interval for the preset attraction dimension is then determined by using the minimum effect index value as the left endpoint of the quantization interval and the maximum effect index value as the right endpoint.
[0085] In this embodiment of the disclosure, the quantification interval corresponding to each preset attraction dimension can be uniformly organized and configured, and solidified into the quantification judgment standard of the attraction dimension.
[0086] In some implementations, a four-dimensional weighted evaluation system can be constructed when determining the criteria for predicting effectiveness, including 40% weight for aesthetics, 25% weight for attractiveness, 20% weight for hormones, and 15% weight for consistency between text and images. Each dimension uses a tiered scoring standard, and the final attractiveness score is calculated from 0 to 1 through weighted summation.
[0087] In this way, preference analysis uses real interaction data to evaluate the actual value of different attraction features, ensuring that each preset attraction dimension is no longer merely a subjective description but has clear performance metrics to support it. This allows for a more accurate reflection of the preference performance of different cover styles in real-world scenarios. By statistically analyzing the maximum and minimum performance metrics across multiple historical images and constructing quantification intervals, each attraction dimension has a unified numerical reference system. This facilitates horizontal comparisons of the performance of different images and styles on that dimension and also provides a standardized basis for labeling subsequent model training. Consequently, it enables more targeted optimization of click-through rates and content consumption experience during cover selection and sorting.
[0088] In some embodiments, using a second prompt word and a second training sample set, a second model is trained by clicking and ranking the first model based on a second loss function to obtain a second model. This includes: based on effect annotations, sorting multiple second sample images for the same resource according to the effect annotations to construct at least one set of sample pairs; each set of sample pairs includes two second sample images for the same resource with different effect annotations; inputting the sample pairs into the first model, and training the first model's ability to distinguish between sample images with different effect annotations using the second loss function to obtain the second model; the second loss function is a pairwise ranking loss function.
[0089] Here, a sample pair refers to a pair of data consisting of two training sample images selected from multiple candidate covers of the same resource according to their performance rating.
[0090] In this embodiment, for any given resource, multiple second sample images corresponding to it, along with each historical effect annotation, can be found from the second training sample set. Then, these candidate covers can be sorted from highest to lowest according to the effect annotations, resulting in an ordered list of effect annotations. Next, multiple sample pairs can be constructed using an adjacent pairing method or a bi-polar pairing method, ensuring that each sample pair consists of two images with different effect annotations from the same resource.
[0091] In this embodiment, for each pair of samples, the two images are combined with a unified second prompt word to form two input instances, which are then fed into the first model to obtain the model's prediction scores for the two images. Subsequently, the system calculates the loss value of the sample pair based on a preset pairwise ranking loss function. This loss function encourages the model to give higher prediction scores for covers with better historical performance and lower scores for covers with poor performance, thus widening the gap between the two. After multiple rounds of training iterations, the model's ability to predict the effectiveness of cover images and its click ranking ability are significantly improved when given the second prompt word, ultimately resulting in a second model that simultaneously possesses image quality assessment and overall attractiveness prediction capabilities.
[0092] In this way, the actual effects of different covers from the same resource are directly transformed into learning signals for the model. This allows the model to learn not just precise regression of absolute values, but rather relative ranking accuracy that better aligns with business needs. Because the sample pairs are limited to the same resource, the interference of resource content differences on the results is eliminated. The model focuses only on the differences in the attractive features of the covers themselves, thereby enhancing its ability to perceive and judge the fine-grained performance of the covers. Compared to pointwise loss, pairwise ranking loss is more robust to label noise and exposure imbalance. It can still learn stable relative preference patterns even when the effect label scale is inconsistent or the distribution is skewed, enabling the final second model to more accurately select covers that contribute more to click-through rate and overall performance.
[0093] In some embodiments, the second model is trained to predict interaction using a third training sample set and a third loss function to obtain a cover selection model, including: inputting the third training samples into the second model and training the prediction ability of the second model for interaction behavior using the third loss function to obtain a cover selection model.
[0094] In this embodiment, the constructed third training samples can be input into the second model, enabling the model to output one or more predicted values representing the recommendation effect for each sample based on its existing quality and attractiveness representation capabilities. Subsequently, the predicted values output by the model can be compared with the pre-given interaction behavior features of the image in the third training sample set, and the error of the model in the recommendation effect prediction task can be calculated using a third loss function. Specifically, the third loss function can be selected based on the type of recommendation effect annotation, such as regression loss, cross-entropy loss based on click probability, or list loss and ranking loss for ranking purposes. By iterating repeatedly on a large-scale sample, the model parameters are gradually adjusted to make its recommendation effect prediction as close as possible to the true annotation, ultimately obtaining the cover selection model.
[0095] Thus, building upon the second model's ability to distinguish between high and low cover quality and relative effectiveness within the same resource, this model further leverages performance annotations from real-world recommendation scenarios and a third loss function designed to align with recommendation goals. This directly aligns the model's learning objective with the final business metrics, upgrading it from relative ranking capabilities to absolute business goal prediction capabilities. Through continuous training on large-scale real-world recommendation data, the model automatically adapts to the statistical patterns of cover performance across different scenarios, user groups, and content types. This enhances its ability to characterize nonlinear relationships in complex environments, enabling the final cover selection model to more accurately select the cover with the best expected performance for each resource in the current recommendation scenario during actual deployment. This overall optimizes the content distribution efficiency, content stickiness, and behavioral conversion rate of the recommendation system.
[0096] This disclosure provides a cover selection method. Figure 2 This is a flowchart illustrating a cover selection method according to an embodiment of the present disclosure. This cover selection method can be applied to a cover selection device. The cover selection model training device and / or the cover selection device are located in an electronic device. The electronic device includes, but is not limited to, fixed devices and / or mobile devices. For example, fixed devices include, but are not limited to, servers, which can be cloud servers or ordinary servers. For example, mobile devices include, but are not limited to, information flow recommendation devices, which can be mobile phones, tablets, etc. In some possible implementations, the cover selection method can also be implemented by a processor calling computer-readable instructions stored in memory. Figure 2 As shown, the cover selection method includes: S201. Determine the characteristics of the objects to be recommended based on the objects to be recommended.
[0097] S202. Based on the target resource and multiple candidate cover images corresponding to the target resource, determine the features of the target resource and the feature set of multiple candidate cover images.
[0098] S203. Based on the features of the object to be recommended, the features of the target resource, and the feature sets of multiple candidate cover images, use the cover selection model to obtain the prediction score of each candidate cover image for the object to be recommended.
[0099] S204. Iterate through each prediction score, select the candidate cover image with the highest prediction score as the target cover image, and display the target cover image to the target audience.
[0100] Here, the object to be recommended refers to the entity that will receive the recommendation results. It can be a specific user, an anonymous user group, a terminal device, a channel, or a business scenario instance. The characteristics of the object to be recommended refer to the structured description of the state of the object to be recommended at the current moment or in the current scenario.
[0101] In this embodiment, the unique identifier or scene identifier of the object to be recommended can be parsed first based on the current recommendation request. Then, a pre-trained feature extraction model or embedding model can be called to convert high-dimensional sparse features such as text and category into low-dimensional dense vectors. The encoded features are then concatenated or grouped according to a predefined feature structure to construct a structured feature set of the object to be recommended.
[0102] Here, the target resource refers to the specific content entity that is currently prepared to be displayed to the recommended object. Candidate cover images refer to multiple cover images associated with the target resource that can be used for display. Target resource features refer to the structured feature description of the target resource itself. The candidate cover image feature set refers to a set of feature vectors obtained after image understanding for each candidate cover image.
[0103] In this embodiment, the target resource to be displayed to the recommended object can be determined first, then the basic information and structured metadata of the target resource can be obtained, and a list of all candidate cover image resources associated with the target resource can be queried from the cover management system. Subsequently, feature extraction can be performed on the target resource itself. For example, resource profile information can be read directly, and a text encoding model can be called to embed and encode text such as title, introduction, and tags, and attributes such as duration, publication period, and sensitivity can be standardized and encoded to construct the target resource features. Further, an image feature extraction process can be called for each candidate cover image to obtain the candidate cover feature set corresponding to the cover image, which includes at least the candidate cover image prior features, candidate cover image posterior features, and candidate cover image cross features.
[0104] Here, the prediction performance score refers to a numerical value output by the cover selection model after simultaneously considering the features of the object to be recommended, the features of the target resource, and the features of a candidate cover image. This value is used to quantify the expected recommendation performance of the candidate cover image under the current combination of object and resource.
[0105] In this embodiment of the disclosure, for each candidate cover image, an input instance of the cover selection model can be constructed according to the same input format. The features of the object to be recommended, the features of the target resource, and the feature set of the candidate cover image of the current candidate cover image are input into the cover selection model together, so that the model outputs the prediction effect score of the candidate cover image for the object to be recommended, thereby obtaining the prediction effect score of each candidate cover image for the object to be recommended.
[0106] In this embodiment, the scores of all candidate covers corresponding to the target resource can be iterated and compared first, and the candidate cover image with the highest score can be selected and marked as the target cover image for the recommended object and the target resource. Subsequently, the target cover image can be sent along with the information of the target resource itself, and the display terminal can render the target cover image in the information flow card, details page entry, or recommendation position according to the specifications, and display this cover version that is considered to have the best effect by the model to the recommended object.
[0107] The technical solution of this disclosure achieves differentiated cover selection for different users by finely characterizing the features of the recommended objects. Based on this, deep feature extraction is performed on the target content and candidate cover images, enabling the model to fully learn the matching relationship between content semantics and cover visual presentation, avoiding misleading covers caused by simply pursuing click-through rates. The cover selection model outputs a prediction performance score as a quantifiable utility indicator; and through concise decision rules, the score is transformed into actual cover display, thereby achieving real-time and automated cover decision-making.
[0108] In some implementations, the initial multimodal large model may include: an initial quality assessment agent, an initial click ranking agent, and an initial interaction prediction agent. Further, based on a quantitative judgment criterion and a first training sample set, the initial quality assessment agent in the initial multimodal large model can be trained for quality assessment using a first loss function to obtain a trained quality assessment agent. At this point, the trained quality assessment agent, the initial click ranking agent, and the initial interaction prediction agent constitute the first model. Even further, based on an effect prediction criterion and a second training sample set, the initial click ranking agent in the first model can be trained for click ranking using a second loss function to obtain a trained click ranking agent. At this point, the trained quality assessment agent, the trained click ranking agent, and the initial interaction prediction agent constitute the second model. Finally, using a third training sample set and a third loss function, the initial interaction prediction agent in the second model can be trained for interaction prediction to obtain a trained interaction prediction agent. At this point, the trained quality assessment agent, the trained click ranking agent, and the trained interaction prediction agent constitute the cover selection model.
[0109] Figure 3 This illustration shows another flowchart of the cover selection method in an embodiment of this disclosure. For example... Figure 3 As shown, it includes: S301. Resource Acquisition. Acquire historical content resources and extract resource characteristics.
[0110] S302. Cover Image Selection. Utilize historical content resources to perform quality assessment training on the multimodal large model, obtaining candidate cover images.
[0111] S303. Model Training. Based on historical resource content and corresponding historical interaction data, the multimodal large model is trained for click ranking and interaction prediction to obtain the cover selection model.
[0112] S304. Online Trigger. When a recommendation request is triggered, based on the list of content to be recommended, candidate covers, user features, and content features are aggregated for each piece of content in the list, and the online-deployed cover selection model is invoked to obtain the prediction score for each candidate cover.
[0113] S305. Optimal Display. Based on the performance score output by the model, the online recommendation system selects the highest-scoring cover image from the candidate covers for each content as the target cover image, and displays the combined result of the resource and the target cover image to the target audience.
[0114] In some implementations, an initial multimodal large model can be trained for quality assessment based on a first loss function, using a quantitative judgment criterion and a first training sample set, to obtain a multimodal large model with image quality assessment capabilities, i.e., a first model. Further, an effect prediction criterion and a second training sample set can be input into the first model to obtain the quality assessment result corresponding to any second training sample. This quality assessment result is then used as a teacher signal. Based on this second training sample, a click-ranking distillation loss term is constructed. Then, based on the click-ranking distillation loss term and a third loss function, the first model is trained for click-ranking using the second training sample, resulting in a second model with both image quality assessment and overall attractiveness prediction capabilities. Finally, a third training sample set can be input into the second model to obtain the quality assessment result and click-ranking result corresponding to any third training sample. These results are then used as a teacher signal. Based on this third training sample, an effect prediction distillation loss term is constructed. Then, based on the effect prediction distillation loss term and the third loss function, the second model is trained for recommendation effect prediction using the third training sample set, resulting in a cover selection model with image quality assessment, overall attractiveness prediction, and interaction behavior prediction capabilities.
[0115] In some implementations, when determining multiple candidate cover images corresponding to the target resource, for video content, frame extraction can be performed first to obtain the original frame sequence. Then, inter-frame deduplication can be performed based on a structural similarity index algorithm to obtain a set of keyframes. Next, low-quality frames can be filtered through multi-dimensional feature detection, including those with insufficient clarity, stretching distortion, borders, and dense elements. Simultaneously, candidate frames can be weighted with quality scores based on image aesthetic models, element label recognition models, and image-text relevance models. Finally, a final candidate cover set {cover 1, cover 2, ..., cover n} is generated.
[0116] In some implementations, when determining multiple candidate cover images corresponding to the target resource, for text and image content, all images can first be extracted from the text and image content. Then, a structural similarity index algorithm can be used for inter-frame deduplication to obtain a key image set. Next, low-quality images can be filtered through multi-dimensional feature detection, including those with insufficient clarity, stretching distortion, borders, and dense elements. Simultaneously, image quality scores can be weighted based on image aesthetic models, element label recognition models, and image-text relevance models. Finally, a final candidate cover set {cover 1, cover 2, ..., cover n} is generated.
[0117] Figure 4 A schematic diagram of the interactive prediction training process in an embodiment of this disclosure is shown, such as... Figure 4 As shown, a third training sample set can be constructed based on real-time online business data. Then, deep learning architectures such as Multi-Layer Perceptron (MLP) or Transformer are used for modeling, with user click behavior serving as a binary classification supervision signal to train the model. The training objective is to minimize the logarithmic loss function. For example, the logarithmic loss function... It can be represented as follows: in, This represents the total number of training samples in the third training sample set; Indicates the first The true label of each sample; The model predicts the first The click probability of a sample; The base is The natural logarithm of .
[0118] Figure 5 Another flowchart illustrating the cover selection method in an embodiment of this disclosure is shown, such as... Figure 5 As shown, during the model inference phase, multiple candidate covers can be evaluated simultaneously. Click-through rate prediction is performed on the cover image {cover1, cover2, ..., covern}, generating a corresponding score set {s1, s2, ..., sn}, and then a maximum selection strategy (i.e., selecting the cover image with the highest score) is applied. Determine the final cover image to be displayed to the user. The optimal cover can be represented by the following formula: in, This represents the function that takes the maximum value. Represents the set of candidate covers The first in One cover candidate; The model represents the first One candidate cover Predicted click-through rate score.
[0119] It should be understood that Figures 3 to 5 The schematic diagrams shown are merely illustrative and not limiting, and are scalable; those skilled in the art can use them as a basis. Figures 3 to 5 Even with various obvious changes and / or substitutions to the examples, the resulting technical solutions still fall within the scope of this disclosure.
[0120] This disclosure provides a cover selection model training device, such as... Figure 6 As shown, the device may include: a sample acquisition module 601, used to acquire a first training sample set, a second training sample set, and a third training sample set; the first training sample set includes multiple first sample images and multi-label quality annotations corresponding to each first sample image; the second training sample set includes multiple second sample images and effect annotations corresponding to each second sample image; the third training sample set includes multi-dimensional features and interactive behavior features, the multi-dimensional features including at least target object features, resource features, and a cover image feature set corresponding to each resource cover image; and a standard setting module 602, used to determine, based on historical images, a one-to-one correspondence with preset quality types. The system includes a quantitative judgment standard and an effect prediction standard corresponding to each preset attraction dimension; a quality assessment training module 603, which is used to train the initial multimodal large model for multi-label low-quality recognition quality assessment based on the quantitative judgment standard and the first training sample set and the first loss function to obtain the first model; a click ranking training module 604, which is used to train the first model for click ranking based on the effect prediction standard and the second training sample set and the second loss function to obtain the second model; and an interaction prediction training module 605, which is used to train the second model for recommendation interaction prediction based on the third loss function using the third training sample set to obtain the cover selection model.
[0121] In some embodiments, the first training sample set is pre-constructed in the following manner: according to a preset quality type, multiple quality labels corresponding to each first sample image are obtained; the multiple quality labels correspond one-to-one with the preset quality type; according to the multiple quality labels corresponding to each first sample image, a quality conclusion for each first sample image is determined; the quality conclusion is qualified or low quality; according to the multiple quality labels and the quality conclusion, a multi-label quality label corresponding to each first sample image is generated; and the first training sample set is constructed according to each first sample image and its corresponding multi-label quality label.
[0122] In some embodiments, the second training sample set is pre-constructed in the following manner: acquiring any second sample image and its corresponding historical interaction data; the quality conclusion of the second sample image is qualified, and the historical interaction data includes at least one of the statistically obtained historical click rate, historical playback completion rate, and historical dwell time; scoring the historical interaction data according to a preset scoring rule, and using the scoring result as the effect label corresponding to the second sample image; and constructing a second training sample set based on each second sample image and its corresponding effect label.
[0123] In some embodiments, the third training sample set is pre-constructed by: acquiring interaction data; the interaction data includes at least the interaction behaviors between the target object and the resource; for any interaction behavior, extracting the corresponding resource features, target object features, interaction behavior features, and cover image feature set; the cover image feature set includes at least prior features, posterior features, and cross features; and constructing the third training sample set based on each interaction behavior feature and its corresponding target object features, resource features, and cover image feature set.
[0124] In some embodiments, extracting interaction behavior features includes: obtaining the occurrence time of each original interaction behavior based on interaction data; and applying different time decay factors to the original interaction behaviors at different times based on a preset time decay function to obtain interaction behavior features.
[0125] In some embodiments, the quality assessment training module 603 includes: a first prompt word submodule, used to construct a first prompt word according to a quantitative judgment standard; and a first training submodule, used to perform quality assessment training on an initial multimodal large model based on a first loss function using the first prompt word and a first training sample set to obtain a first model; the first loss function is a loss function based on multi-label binary cross-entropy.
[0126] In some embodiments, the standard setting module 602 includes: a quality index submodule, used to perform image quality analysis on any historical image and determine the quality index of the historical image under each preset quality type; a quality interval submodule, used to determine the maximum and minimum quality index corresponding to the preset quality type based on the quality index of multiple historical images under the same preset quality type, and to determine the quantization interval of the preset quality type based on the maximum and minimum quality index; and a quantization standard module, used to use the quantization interval corresponding to each preset quality type as a quantization judgment standard.
[0127] In some embodiments, the click ranking training module 604 includes: a second prompt word submodule, used to construct a second prompt word according to the effect prediction criteria; and a second training submodule, used to train the first model for click ranking based on a second loss function using the second prompt word and a second training sample set to obtain a second model.
[0128] In some embodiments, the standard setting module 602 includes: an effect index submodule, used to perform preference analysis on any historical image and determine the effect index of the historical image under each preset attraction dimension; an effect interval submodule, used to determine the maximum and minimum effect index corresponding to the preset attraction dimension based on the effect indices of multiple historical images under the same preset attraction dimension, and to determine the quantization interval of the preset attraction dimension based on the maximum and minimum effect indices; and an effect standard module, used to use the quantization interval corresponding to each preset attraction dimension as an effect prediction standard.
[0129] In some embodiments, the second training submodule is used to: sort multiple second sample images for the same resource according to the effect labels, and construct at least one set of sample pairs; each set of sample pairs includes two second sample images for the same resource with different effect labels; input the sample pairs into the first model, and train the first model to distinguish sample images with different effect labels through the second loss function to obtain the second model; the second loss function is the pairwise sorting loss function.
[0130] In some embodiments, the interaction prediction training module 605 includes: a third training submodule, used to input a third training sample into a second model, and train the prediction ability of the second model for interactive behavior through a third loss function to obtain a cover selection model.
[0131] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0132] The cover selection model training device in this embodiment employs a multi-stage training strategy, utilizing sample sets with different annotations to sequentially enable the model to assess image quality, predict overall attractiveness, and predict personalized effects. During online recommendation, it dynamically selects the cover with the highest click potential based on contextual features, effectively alleviating the problem of insufficient exposure for high-quality content, thereby improving click-through rate and content distribution efficiency.
[0133] This disclosure provides a cover selection device, such as... Figure 7 As shown, the device may include: a first feature determination module 701, used to determine the current user features of the target user based on the target user to be recommended; a second feature determination module 702, used to determine the target resource features and multiple candidate cover image feature sets based on the target resource and multiple candidate cover images corresponding to the target resource; a model prediction module 703, used to obtain the prediction effect score of each candidate cover image for the current user to be recommended based on the current user features of the target user, the target resource features, and the multiple candidate cover image feature sets, using a cover selection model; the cover selection model is obtained through the method in the embodiments of this disclosure; and an image display module 704, used to traverse each prediction effect score, select the candidate cover image with the highest prediction effect score as the target cover image, and display the target cover image to the current user to be recommended.
[0134] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0135] The cover selection device in this embodiment can achieve differentiated cover selection for different users by finely characterizing the features of the recommended objects. Based on this, deep feature extraction is performed on the target content and candidate cover images, enabling the model to fully learn the matching relationship between content semantics and cover visual presentation, avoiding misleading covers caused by simply pursuing click-through rates. The cover selection model outputs a prediction performance score as a quantifiable utility indicator; and through concise decision rules, the score is transformed into actual cover display, thereby achieving real-time and automated cover decision-making.
[0136] This disclosure provides a scenario illustration of a cover selection model training method, such as... Figure 8 As shown.
[0137] As previously described, the cover selection model training method provided in this disclosure is applied to electronic devices. These electronic devices are intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers.
[0138] Specifically, the electronic device may perform the following operations: Obtain a first training sample set, a second training sample set, and a third training sample set. The first training sample set includes multiple first sample images and multi-label quality annotations corresponding to each first sample image. The second training sample set includes multiple second sample images and effect annotations corresponding to each second sample image. The third training sample set includes multi-dimensional features and interactive behavior features. The multi-dimensional features include at least target object features, resource features, and a cover image feature set corresponding to each resource cover image. Based on historical images, determine the quantitative judgment criteria corresponding to each preset quality type and the effect prediction criteria corresponding to each preset attraction dimension. Based on the quantitative judgment criteria and the first training sample set, train the initial multimodal large model for quality assessment using a first loss function to obtain a first model. Based on the effect prediction criteria and the second training sample set, train the first model for click ranking using a second loss function to obtain a second model. Using the third training sample set, train the second model for interaction prediction using a third loss function to obtain a cover selection model.
[0139] This disclosure provides a scenario illustration of a cover selection method, such as... Figure 9 As shown.
[0140] As previously described, the cover selection method provided in this disclosure is applied to electronic devices. These electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.
[0141] Specifically, the electronic device may perform the following operations: Based on the object to be recommended, determine the characteristics of the object to be recommended; based on the target resource and multiple candidate cover images corresponding to the target resource, determine the characteristics of the target resource and the feature set of multiple candidate cover images; based on the characteristics of the object to be recommended, the characteristics of the target resource, and the feature set of multiple candidate cover images, use the cover selection model to obtain the prediction effect score of each candidate cover image for the object to be recommended; iterate through each prediction effect score, and select the candidate cover image with the highest prediction effect score as the target cover image, and display the target cover image to the object to be recommended.
[0142] It should be understood that Figure 8 and Figure 9 The scene diagrams shown are merely illustrative and not restrictive; those skilled in the art can interpret them based on... Figure 8 and Figure 9 Even with various obvious changes and / or substitutions to the examples, the resulting technical solutions still fall within the scope of this disclosure.
[0143] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0144] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0145] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0146] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.
[0147] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0148] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the cover selection model training method and / or cover selection method. For example, in some embodiments, the cover selection model training method and / or cover selection method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, one or more steps of the cover selection model training method and / or cover selection method described above can be performed. Alternatively, in other embodiments, computing unit 1001 can be configured to perform the cover selection model training method and / or cover selection method by any other suitable means (e.g., by means of firmware).
[0149] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0150] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0151] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0152] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0153] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0154] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0155] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0156] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for training a cover selection model, comprising: Obtain the first training sample set, the second training sample set, and the third training sample set; The first training sample set includes multiple first sample images and multi-label quality annotations corresponding to each first sample image; the second training sample set includes multiple second sample images and effect annotations corresponding to each second sample image; the third training sample set includes multi-dimensional features and interactive behavior features; the multi-dimensional features include at least target object features, resource features, and a cover image feature set corresponding to each resource cover image. Based on historical images, quantitative judgment criteria corresponding to preset quality types and effect prediction criteria corresponding to preset attraction dimensions are determined respectively. Based on the quantitative judgment criteria and the first training sample set, the initial multimodal large model is trained for quality assessment based on the first loss function to obtain the first model; Based on the effect prediction criteria and the second training sample set, the first model is trained by click ranking based on the second loss function to obtain the second model; Using the third training sample set, the second model is interactively predicted and trained based on the third loss function to obtain the cover selection model.
2. The method according to claim 1, wherein, The first training sample set is pre-constructed in the following manner: Based on the preset quality type, multiple quality labels are obtained corresponding to each of the first sample images; the multiple quality labels correspond one-to-one with the preset quality type. Based on the plurality of quality labels corresponding to each first sample image, a quality conclusion is determined for each first sample image; the quality conclusion is either acceptable or low quality. Based on the multiple quality labels and the quality conclusions, generate multi-label quality labels corresponding to each of the first sample images; The first training sample set is constructed based on each of the first sample images and its corresponding multi-label quality annotations.
3. The method according to claim 2, wherein, The second training sample set is pre-constructed in the following manner: Acquire any of the second sample images and their corresponding historical interaction data; the quality conclusion of the second sample image is qualified, and the historical interaction data includes at least one of the statistically obtained historical click-through rate, historical playback completion rate, and historical dwell time. According to the preset scoring rules, the historical interaction data is scored, and the scoring results are used as the effect annotation corresponding to the second sample image. The second training sample set is constructed based on each second sample image and its corresponding effect annotation.
4. The method according to claim 1, wherein, The third training sample set is pre-constructed in the following manner: Acquire interaction data; the interaction data includes at least the interaction behavior between the target object and the resource. For any of the aforementioned interactive behaviors, extract the corresponding resource features, target object features, interactive behavior features, and cover image feature set; the cover image feature set includes at least prior features, posterior features, and cross features. The third training sample set is constructed based on each of the interactive behavior features and its corresponding target object features, resource features, and cover image feature sets.
5. The method according to claim 4, wherein, Extracting interactive behavior features includes: Based on the interaction data, the occurrence time of each original interaction behavior is obtained; Based on a preset time decay function, different time decay factors are applied to the original interactive behavior at different times to obtain the interactive behavior characteristics.
6. The method according to claim 1, wherein, The step of training an initial multimodal large model based on a first loss function, according to the quantification criteria and the first training sample set, to obtain a first model includes: Based on the aforementioned quantitative judgment criteria, a first prompt word is constructed; Using the first prompt word and the first training sample set, the initial multimodal large model is trained for quality assessment based on the first loss function to obtain the first model; the first loss function is a loss function based on multi-label binary cross-entropy.
7. The method according to claim 1, wherein, A quantitative judgment standard based on historical images, corresponding one-to-one with preset quality types, includes: Perform image quality analysis on any historical image to determine the quality index of the historical image under each preset quality type; Based on the quality indices of multiple historical images under the same preset quality type, determine the maximum and minimum quality indices corresponding to the preset quality type, and determine the quantization range of the preset quality type based on the maximum and minimum quality indices. The quantization interval corresponding to each preset quality type is used as the quantization judgment criterion.
8. The method according to claim 1, wherein, The step of training the first model by click ranking based on the effect prediction criteria and the second training sample set, and obtaining the second model by using the second loss function, includes: Based on the aforementioned effect prediction criteria, a second prompt word is constructed; Using the second prompt word and the second training sample set, the first model is trained by click ranking based on the second loss function to obtain the second model.
9. The method according to claim 1, wherein, Based on historical images, effect prediction criteria are determined to correspond one-to-one with preset attraction dimensions, including: Perform preference analysis on any historical image to determine the effect index of the historical image under each preset attraction dimension; Based on the effect metrics of multiple historical images under the same preset attraction dimension, determine the maximum and minimum effect metrics corresponding to the preset attraction dimension, and determine the quantization range of the preset attraction dimension based on the maximum and minimum effect metrics. The quantization interval corresponding to each of the preset attraction dimensions is used as the effect prediction standard.
10. The method according to claim 8, wherein, The step of using the second prompt word and the second training sample set to train the first model based on the second loss function to obtain the second model includes: Based on the effect annotations, multiple second sample images for the same resource are sorted according to the effect annotations to construct at least one set of sample pairs; each set of sample pairs includes two second sample images for the same resource with different effect annotations; The sample pairs are input into the first model, and the ability of the first model to distinguish sample images labeled with different effects is trained by the second loss function to obtain the second model; the second loss function is the pairwise ranking loss function.
11. The method according to claim 1, wherein, The step of using the third training sample set and training the second model for recommendation interaction prediction based on the third loss function to obtain the cover selection model includes: The third training sample is input into the second model, and the prediction ability of the second model for interactive behavior is trained by the third loss function to obtain the cover selection model.
12. A cover selection method, comprising: Based on the objects to be recommended, determine their characteristics; Based on the target resource and multiple candidate cover images corresponding to the target resource, determine the features of the target resource and the feature sets of multiple candidate cover images; Based on the features of the object to be recommended, the features of the target resource, and the feature sets of the multiple candidate cover images, a cover selection model is used to obtain the prediction score of each candidate cover image for the object to be recommended; the cover selection model is obtained by the method of any one of claims 1-11. Iterate through each prediction effect score, select the candidate cover image with the highest prediction effect score as the target cover image, and display the target cover image to the object to be recommended.
13. A cover selection model training device, comprising: The sample acquisition module is used to acquire the first training sample set, the second training sample set, and the third training sample set. The first training sample set includes multiple first sample images and multi-label quality annotations corresponding to each first sample image; the second training sample set includes multiple second sample images and effect annotations corresponding to each second sample image; the third training sample set includes multi-dimensional features and interactive behavior features; the multi-dimensional features include at least target object features, resource features, and a cover image feature set corresponding to each resource cover image. The standard setting module is used to determine, based on historical images, quantitative judgment standards corresponding to preset quality types and effect prediction standards corresponding to preset attraction dimensions. The quality assessment training module is used to perform multi-label low-quality identification quality assessment training on the initial multimodal large model based on the quantitative judgment criteria and the first training sample set and the first loss function to obtain the first model; The click ranking training module is used to train the first model by click ranking based on the effect prediction criteria and the second training sample set, and based on the second loss function, to obtain the second model. The interactive prediction training module is used to train the second model for recommendation interactive prediction based on the third loss function using the third training sample set to obtain the cover selection model.
14. A cover selection device, comprising: The first feature determination module is used to determine the current user features of the target user based on the target user to be recommended. The second feature determination module is used to determine the target resource features and the multiple candidate cover image feature sets based on the target resource and multiple candidate cover images corresponding to the target resource. The model prediction module is used to obtain the prediction performance score of each candidate cover image for the current user of the object to be recommended, based on the current user characteristics of the object to be recommended, the target resource characteristics, and the multiple candidate cover image feature sets, using a cover selection model; the cover selection model is obtained by the method of any one of claims 1-11. The image display module is used to traverse each prediction effect score, select the candidate cover image with the highest prediction effect score as the target cover image, and display the target cover image to the current user of the object to be recommended.
15. An electronic device comprising: At least one processor; as well as A memory that is communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, Computer instructions are used to cause a computer to perform the method according to any one of claims 1-12.
17. A computer program product comprising a computer program stored on a storage medium, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-12.