Advertisement recommendation method based on multiple modes and related device
By constructing triple data sets and using 3D-CNN, BERT and Transformer models, combining cross-modal comparison loss function and dynamic weight optimization, the problems of multimodal feature fusion and real-time perception on short video platforms are solved, and the accuracy of advertising recommendations and noise resistance are improved.
Patent Information
- Application Number
- CN202510524989.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional recommendation technology is difficult to effectively integrate multimodal features of vision, voice, text and user behavior on short video platforms, and lacks real-time perception and dynamic optimization capabilities, resulting in the recommendation results deviating from the real business scenarios, and insufficient causal reasoning capabilities, resulting in a misjudgment of potential correlation between advertising exposure and user feedback.
By constructing a triple dataset, video, text and user behavior characteristics are extracted using 3D-CNN, BERT and Transformer models, combined with cross-modal comparison loss function and dynamic weight optimization, multimodal semantic alignment is achieved, and counterfactual causal reasoning module is introduced to eliminate environmental confounding variable interference and accurately quantify the real causal effect of advertising exposure.
It realizes multimodal deep collaboration, improves advertising recall and user interest capture accuracy, weakens traffic fluctuations, optimizes multi-target games, and provides an interpretable, generalizable and anti-noise advertising recommendation framework.
Smart Images

Figure CN120450779A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of advertising push technology, and in particular to a multimodal advertising recommendation method and related devices. Background Art
[0002] With the explosive growth of short video platforms, ad recommendation systems face core challenges such as dynamic environment perception, multimodal feature fusion, and real-time decision optimization. Traditional recommendation technologies (such as collaborative filtering and matrix factorization) are clearly insufficient in addressing the following trends.
[0003] Faced with the multimodal and complex content forms of short videos that integrate vision, voice, text, and user behavior, single-modal modeling is difficult to capture cross-modal semantic associations; in response to the dynamic needs of user interests that fluctuate dramatically with scenarios and time, static models lack real-time behavior perception and incremental optimization capabilities; in the context of diversified business objectives (the need to balance CTR, CVR, and user experience), traditional single-objective optimization strategies lead to an imbalance in multi-objective games due to the solidification of feature space; at the same time, environmental noise interference such as platform traffic fluctuations and holiday effects has intensified, and traditional methods, due to the lack of causal reasoning capabilities, misjudge the potential correlation between advertising exposure and user feedback, ultimately causing recommendation results to deviate from the real business scenario.
[0004] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art. Summary of the Invention
[0005] In view of this, an embodiment of the present application provides a personalized advertising recommendation method and related devices based on sentiment analysis.
[0006] In a first aspect, an embodiment of the present application provides a personalized advertising recommendation method based on sentiment analysis, the method comprising:
[0007] Get the video frame sequence corresponding to the same content from the short video platform Text description and user behavior sequence
[0008] According to the video frame sequence The text description and the user behavior sequence Constructing triplet dataset
[0009] Use a pre-trained 3D-CNN model to extract the video frame sequence Extract video features Describe the text using the BERT model Extract text features Encoding the user behavior sequence based on Transformer Generate behavioral features
[0010] According to the triplet dataset The video features The text features and the behavioral characteristics Calculating video-text similarity and video-behavior similarity
[0011] According to the video-text similarity Similarity to the video-behavior Construct a cross-modal contrast loss function:
[0012]
[0013] Wherein, τ is the temperature coefficient, which is used to control the sharpness of the probability distribution. The smaller the value, the higher the discrimination between similar samples.
[0014] Minimizing the cross-modal contrast loss function through back-propagation to update the weights of the video modality, text modality, and behavioral modality;
[0015] Based on the updated weights of each modality, the 3D-CNN model, the BERT model, and the Transformer are combined to perform advertisement recommendation.
[0016] In one embodiment, the advertisement recommendation based on the updated weights of each modality in combination with the 3D-CNN model, the BERT model, and the Transformer includes:
[0017] Calculating cross-modal similarity between a target ad and multiple candidate ads in an ad library; the target ad is an ad viewed by a user;
[0018] The multiple cross-modal similarities are sorted, and the candidate advertisement corresponding to the highest cross-modal similarity is selected and recommended to the user.
[0019] In one embodiment, calculating the cross-modal similarity between the target advertisement and multiple candidate advertisements in the advertisement library includes:
[0020]
[0021] Among them, S 召回is the cross-modal similarity between the target advertisement and the candidate advertisement, is the video feature of the target advertisement, is the video feature of the candidate advertisement, is the text feature of the target advertisement, is the text feature of the candidate advertisement, Behavioral characteristics of the target advertisement, is the behavioral feature of the candidate advertisement, α is the weight of the video modality, β is the weight of the text modality, and γ is the weight of the behavioral modality.
[0022] As an embodiment, the method further includes:
[0023] Define the processing variable T, the confounding variable C, and the observation result Y; when the processing variable T is 0, the sample is an unexposed sample, and when the processing variable T is 1, the sample is an exposed sample;
[0024] Based on the treatment variable T, the confounding variable C and the observation result Y, a logistic regression model is used to predict P(T=1|C) to obtain a propensity score e(C);
[0025] For the exposed sample, the counterfactual result Y(1) is predicted by the deep network; for the unexposed sample, the counterfactual result Y(0) is predicted;
[0026] Based on the counterfactual result Y(1) and the counterfactual result Y(0), the recommendation weight of each candidate advertisement is determined.
[0027]
[0028] The recommendation weights The recommendation priority of the candidate advertisement is adjusted.
[0029] In a second aspect, an embodiment of the present application provides a multimodal advertising recommendation device, comprising:
[0030] The acquisition module is used to obtain the video frame sequence corresponding to the same content from the short video platform Text description and user behavior sequence
[0031] A construction module is used to construct a video frame sequence according to the video frame sequence. The text description and the user behavior sequence Constructing triplet dataset
[0032] Extraction module for extracting the video frame sequence from the video using a pre-trained 3D-CNN model Extract video features Describe the text using the BERT model Extract text features Encoding the user behavior sequence based on Transformer Generate behavioral features
[0033] A calculation module is used to calculate the triple data set according to the The video features The text features and the behavioral characteristics Calculating video-text similarity and video-behavior similarity
[0034] Loss module, for determining the video-text similarity Similarity to the video-behavior Construct a cross-modal contrast loss function:
[0035]
[0036] Wherein, τ is the temperature coefficient, which is used to control the sharpness of the probability distribution. The smaller the value, the higher the discrimination between similar samples.
[0037] An updating module, configured to update the weights of the video modality, the text modality, and the behavioral modality by minimizing the cross-modal contrast loss function through backpropagation;
[0038] A recommendation module is used to recommend advertisements based on the updated weights of each modality in combination with the 3D-CNN model, the BERT model, and the Transformer.
[0039] As an embodiment, the updating module is specifically configured to calculate the cross-modal similarity between a target advertisement and a plurality of candidate advertisements in an advertisement library; the target advertisement is an advertisement browsed by a user;
[0040] The multiple cross-modal similarities are sorted, and the candidate advertisement corresponding to the highest cross-modal similarity is selected and recommended to the user.
[0041] In one embodiment, calculating the cross-modal similarity between the target advertisement and multiple candidate advertisements in the advertisement library includes:
[0042]
[0043] Among them, S 召回 is the cross-modal similarity between the target advertisement and the candidate advertisement, is the video feature of the target advertisement, is the video feature of the candidate advertisement, is the text feature of the target advertisement, is the text feature of the candidate advertisement, Behavioral characteristics of the target advertisement, is the behavioral feature of the candidate advertisement, α is the weight of the video modality, β is the weight of the text modality, and γ is the weight of the behavioral modality.
[0044] As an embodiment, the recommendation module is further configured to:
[0045] Define the processing variable T, the confounding variable C, and the observation result Y; when the processing variable T is 0, the sample is an unexposed sample, and when the processing variable T is 1, the sample is an exposed sample;
[0046] Based on the treatment variable T, the confounding variable C and the observation result Y, a logistic regression model is used to predict P(T=1|C) to obtain a propensity score e(C);
[0047] For the exposed sample, the counterfactual result Y(1) is predicted by the deep network; for the unexposed sample, the counterfactual result Y(0) is predicted;
[0048] Based on the counterfactual result Y(1) and the counterfactual result Y(0), the recommendation weight of each candidate advertisement is determined.
[0049]
[0050] The recommendation weights The recommendation priority of the candidate advertisement is adjusted.
[0051] In a third aspect, an embodiment of the present application provides a device comprising a memory and a processor, wherein the memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes the multimodal advertising recommendation method described in any one of the first aspects above.
[0052] In a fourth aspect, an embodiment of the present application provides a computer storage medium having a code stored therein. When the code is executed, the device executing the code implements the multimodal advertising recommendation method described in any one of the aforementioned first aspects.
[0053] The embodiment of the present application provides a multimodal advertising recommendation method and device, which systematically solves the technical bottlenecks of traditional recommendation technology in complex scenarios through multimodal feature alignment, dynamic weight optimization and causal effect decoupling. First, 3D-CNN, BERT and Transformer are used to extract the modal features of video, text and user behavior sequences respectively, and a cross-modal contrast loss function is constructed to strengthen multimodal semantic alignment; secondly, the weights of each modality are dynamically updated through backpropagation, and advertising recall and ranking are achieved by combining cross-modal similarity calculation; further, a counterfactual causal reasoning module is introduced, and through propensity score estimation and counterfactual result prediction, the interference of environmental confounding variables on the recommendation effect is eliminated, and the true causal effect of advertising exposure is accurately quantified.
[0054] In response to existing technical problems, this solution achieves the following technical improvements:
[0055] Multimodal deep collaboration: Cross-modal comparative learning effectively integrates video, text, and behavioral features, solving the semantic fragmentation problem of traditional single-modal modeling and improving advertising recall rate.
[0056] Dynamic real-time adaptation: Through a loss-function-driven weight update mechanism, it captures user interest drift in real time (e.g., shifting from "visual preference" to "copy sensitivity"), improving the accuracy of short-term user behavior prediction compared to static models.
[0057] Causal debiasing and trustworthy decision-making: The counterfactual reasoning module removes confounding factors such as traffic fluctuations and holiday effects, reducing attribution errors for ad clicks and avoiding the misattribution of high-exposure, low-conversion ads.
[0058] Multi-objective global optimization: Integrating cross-modal similarity (recall stage) and causal effect weight (ranking stage) to achieve Pareto frontier breakthroughs in the multi-objective game of CTR, CVR, and user experience.
[0059] This solution deeply integrates representation learning, dynamic decision-making and causal inference, providing an explainable, generalizable and noise-resistant technical framework for short video advertising recommendations, which significantly surpasses the performance boundaries of traditional collaborative filtering, matrix decomposition and other methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in this embodiment or the prior art, the following briefly introduces the drawings required for use in the embodiment or the prior art description. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0061] Figure 1 A flowchart of a multimodal advertising recommendation method provided in an embodiment of the present application;
[0062] Figure 2 A schematic diagram of the structure of a multimodal advertising recommendation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0063] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification and appended claims of the present application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear contrary indication in the context. It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more; "and / or" describes the association relationship of associated objects, indicating that three relationships may exist; for example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship.
[0064] like Figure 1 As shown, Figure 1 This is a flowchart of a personalized advertising recommendation method based on sentiment analysis provided in an embodiment of the present application. The method includes:
[0065] S101: Obtaining a video frame sequence corresponding to the same content from a short video platform Text description and user behavior sequence
[0066] "Same content" specifically refers to the same short video ad instance. A short video ad instance is an independent ad unit on the platform, including the video material, ad copy (text description), and user interactions with the ad (such as clicks, views, and dwell time).
[0067] A video frame sequence refers to a continuous sequence of images after the video content of the ad has been framed. A text description can be structured or unstructured text such as the ad title, copy, or tags (e.g., "Limited-time discount" or "New product release"). A user behavior sequence includes multiple users' real-time interactions with the ad (e.g., click-through rate, scrolling behavior, and completion rate).
[0068] Only by tying the video and text of the same ad to user behavior can we maximize the mutual information of the same semantic object during comparative learning (for example, the video content of Ad A must be associated with its copy "Sneakers Sale" and user click behavior). If the data comes from different ads (for example, the video comes from Ad A, and the text comes from Ad B), the model will learn the wrong alignment, causing feature fusion failure.
[0069] User behavior data doesn't come from a single user, but rather is a statistical or serialized representation of multiple users interacting with the same ad (e.g., global click-through rate and average dwell time for Ad A). Behavior sequences are sliced into time windows to extract dynamic patterns (e.g., the concentration of user clicks in the first three seconds).
[0070] As an example, select Ad A from the ad library and extract its video file (MP4) and copy text (JSON). Filter all user behavior records for Ad A (e.g., User 1 clicks, User 2 swipes out) from the log system. Sort and segment by timestamp to generate a behavior sequence (e.g., "0-2 seconds: 80% of users don't skip; 2-5 seconds: CTR peaks at 15%)."
[0071] Trimodal alignment of the same content enables the model to accurately capture ad semantics (e.g., associating a "sneakers" video with a "limited edition" copy), thereby reducing confusion between similar ads when making recommendations. However, if the video and text come from different ads, the model may incorrectly associate the video and copy (e.g., associating a "sneakers" video with a "beauty copy"), causing recommendations to deviate from the user's true interests.
[0072] Video (visual dynamics), text (semantic abstraction), and behavior (temporal preference) constitute a triplet of complementary information, covering the full dimensional characteristics of advertising content and user intentions.
[0073] For example, the user’s behavior sequence of clicking on the “outdoor jacket” ad needs to be modeled together with the “waterproof fabric test” image and the “all-weather protection” text in the video to accurately capture the user’s demand for “functionality”.
[0074] S102: Based on the video frame sequence The text description and the user behavior sequence Constructing triplet dataset
[0075] Triplet dataset A sequence of video frames Text description and user behavior sequence A structured collection of trimodal data.
[0076] The same advertisement (Xv ,X t ,X b ) to form a positive triplet, forcing the model to learn semantic consistency among the three. Any two modalities of different advertisements are randomly combined (e.g., video frames of Ad A + text of Ad B), and spurious associations between unrelated modalities are suppressed through contrastive loss.
[0077] For example, behavior sequence noise caused by accidental user clicks (such as misclicks) can be corrected through strong video-text correlation (e.g., if a user accidentally clicks on an ad for "children's toys," but the video and text describe it as "office supplies," triplet comparison will reduce the weight of this behavior).
[0078] The dimensions and distributions of the original feature spaces of video (3D-CNN high-dimensional tensors), text (BERT word vectors), and behavior (Transformer temporal coding) vary greatly. Through joint optimization, the triplet dataset maps the trimodal features into a unified semantic space (such as a 256-dimensional vector space), solving the dimensionality disaster and information redundancy caused by traditional multimodal concatenation.
[0079] As an example, there are two ads in the ad library:
[0080] Ad A: Video showing "waterproof sneakers" Text description "outdoor hiking shoes" User behavior sequences contain multiple clicks
[0081] Ad B: Video showing "casual shoes" Text description "Breathable mesh shoes" No user behavior data (new ads).
[0082] Positive sample pair reinforcement: Force the model to associate the "waterproof" video feature with the "mountaineering" text and click behavior;
[0083] Negative Sample Suppression: Random Combination Generate negative samples to reduce the false association between "waterproof sneakers" and "breathable mesh".
[0084] S103: Use the pre-trained 3D-CNN model to extract the video frame sequence Extract video features Describe the text using the BERT model Extract text features Encoding the user behavior sequence based on Transformer Generate behavioral features
[0085] The 3D convolutional neural network model (3D-CNN model) is a deep learning model that performs convolution operations simultaneously in the spatial and temporal dimensions. Its convolution kernel parameters slide in the width, height, and time axis of the video frame, and can simultaneously capture local spatial textures (such as object shapes) and short-term motion patterns (such as object movement trajectories). Video Features Representing the dynamic semantics of a video (e.g., the visual information of a “sneaker bounce test”).
[0086] 3D-CNN models are usually pre-trained on large video datasets with the goal of "video classification" to learn common spatiotemporal features (such as the coherent posture changes of the "running" action); the shallow convolutional layers are frozen on advertising datasets (retaining common features), and the deep network is fine-tuned to adapt to advertising scenarios (such as identifying "product display transition" special effects).
[0087] The BERT (Bidirectional Encoder Representations from Transformers) model is a bidirectional pre-trained language model based on Transformer, which learns text context dependencies through masked language modeling (MLM) and next sentence prediction (NSP) tasks. Used to capture the deep semantics of text (such as "waterproof and breathable" points to functional requirements).
[0088] Transformer is a sequence encoding model based on the self-attention mechanism. It models the long-term dependencies of sequence elements through positional encoding and multi-head attention. This model is used to map discrete behaviors (clicks, favorites, etc.) into dense vectors and add timestamp encoding (reflecting the interval between behaviors). Behavioral characteristics Characterize the temporal evolution of user interests (e.g., the demand migration from "browsing sports shoes" to "searching sports socks").
[0089] S104: Based on the triplet data set The video features The text features and the behavioral characteristics Calculating video-text similarity and video-behavior similarity
[0090] The video features output by 3D-CNN And the text features output by BERT Each of them is mapped to a unified dimension through an independent linear projection layer; the projected feature vector is normalized and the cosine similarity is calculated to obtain the video-text similarity Video-text similarity It is used to measure the consistency between video content and text description in the semantic space, reflecting the degree of "content-copy" matching of the advertisement, and ensuring that the advertisement content is strongly related to the copy description.
[0091] User behavior sequence After Transformer encoding, the hidden state of the last time step is taken as the behavioral feature Video Features and behavioral characteristics After projecting to the same dimension, the similarity is calculated to obtain the video-behavior similarity.
[0092] Video-Behavior Similarity It is used to quantify the correlation between video content and user historical behavior sequences, reflect the degree of "content-user preference" matching of advertisements, and capture users' real-time preferences.
[0093] S105: Based on the video-text similarity Similarity to the video-behavior Constructing a cross-modal contrastive loss function.
[0094] Cross-modal contrast loss function:
[0095]
[0096] Wherein, τ is the temperature coefficient, which is used to control the sharpness of the probability distribution. The smaller the value, the higher the discrimination between similar samples.
[0097] Step S105: Jointly optimize video-text similarity Similarity to video-behavior A multi-objective contrastive loss function is designed to force the model to learn cross-modal semantic alignment while suppressing irrelevant modal noise.
[0098] The triplets of the same advertisement force the video features to be highly similar to the text and behavioral features in a shared semantic space; the random combination of modalities across advertisements (such as the video of advertisement A + the text of advertisement B) increases the distance between irrelevant modal features.
[0099] The technical solution provided in this application jointly optimizes content and behavior alignment, replacing single modality comparison; adjusts loss weights according to real-time similarity ratios to enhance model generalization.
[0100] Step S105 achieves a balance between the authenticity of advertising content (video-text alignment) and the sensitivity to user preferences (video-behavior alignment) through the joint optimization of multimodal contrast loss and dynamic weight adjustment mechanism.
[0101] S106: Minimize the cross-modal contrast loss function through back propagation and update the weights of the video modality, text modality, and behavioral modality.
[0102] The backpropagation algorithm jointly optimizes model parameters for video, text, and behavioral modalities, minimizing the cross-modal contrast loss function to achieve multimodal semantic alignment and user preference modeling. Its core process includes forward computation of each modality's features (3D-CNN for video, BERT for text, and Transformer for behavior), contrastive loss calculation, and gradient backpropagation parameter updates. A dynamic weight adjustment mechanism adaptively balances content consistency with user preference optimization based on the ratio of video-text and video-behavior similarities, ensuring that cold-start ads rely on text matching while mature ads focus on behavioral feedback.
[0103] In multimodal collaborative training, gradient normalization and clipping strategies are employed to prevent dominance by a single modality, and the AdamW optimizer is combined to improve convergence efficiency. Distributed training and mixed-precision techniques (such as FP16 acceleration) significantly increase training speed, enabling end-to-end optimization on large-scale data. Experiments demonstrate that the model strengthens cross-modal associations through contrastive learning. For example, the video model learns to extract fine-grained features such as "medical sensor," the text model enhances keyword sensitivity, and the behavior encoder more accurately captures users' dynamic interests.
[0104] Compared to traditional independent single-modality training, this approach uses a global contrastive loss to force alignment of the three modalities in a shared semantic space, addressing the issues of feature heterogeneity and semantic gaps. Furthermore, a random replacement pattern of negative samples enhances data diversity and improves the model's noise immunity.
[0105] S107: Based on the updated weights of each modality, the 3D-CNN model, the BERT model, and the Transformer are combined to perform advertisement recommendation.
[0106] Personalized ad recommendations are achieved through an end-to-end multimodal reasoning framework. First, based on trained 3D-CNN, BERT, and Transformer models, the spatiotemporal features of the video, the semantic features of the text, and the interest features of user behavior are extracted. The video-text and video-behavior matching degrees are then calculated using cosine similarity. Subsequently, a dynamic weighted fusion strategy is employed to adaptively combine the similarities of the two based on the ad lifecycle. Cold-start ads rely on the consistency of video and text content (such as accurate descriptions of new products), while mature ads focus on user behavior feedback (such as historical click preferences). This ensures that recommendation results meet both content authenticity and personalization requirements.
[0107] During the real-time recommendation phase, the system ranks candidate ads based on the combined similarity scores and pushes the top-K results. Simultaneously, user behavior data, such as clicks and stays, is incrementally encoded using a Transformer, with behavioral features updated every 15 minutes to dynamically capture changes in interest (e.g., a shift from "coffee machine" to "bean grinder"). This real-time feedback mechanism enables recommendations to quickly respond to users' short-term preferences. For example, after a user searches for "latte art," the system immediately increases the ranking weight of ads for "milk frother."
[0108] In one embodiment, the advertising recommendation based on the updated weights of each modality in combination with the 3D-CNN model, the BERT model, and the Transformer includes:
[0109] Calculating cross-modal similarity between a target ad and multiple candidate ads in an ad library; the target ad is an ad viewed by a user;
[0110] The multiple cross-modal similarities are sorted, and the candidate advertisement corresponding to the highest cross-modal similarity is selected and recommended to the user.
[0111] In one embodiment, calculating the cross-modal similarity between the target advertisement and a plurality of candidate advertisements in the advertisement library includes:
[0112]
[0113] Among them, S 召回 is the cross-modal similarity between the target advertisement and the candidate advertisement, is the video feature of the target advertisement, is the video feature of the candidate advertisement, is the text feature of the target advertisement, is the text feature of the candidate advertisement, Behavioral characteristics of the target advertisement, is the behavioral feature of the candidate advertisement, α is the weight of the video modality, β is the weight of the text modality, and γ is the weight of the behavioral modality.
[0114] The technology upgrade is achieved through dynamic weight optimization driven by deep reinforcement learning and incremental behavior modeling enhancement mechanism: First, the Deep Deterministic Policy Gradient (DDPG) algorithm is used to dynamically adjust the video-text similarity Similarity to video-behavior The fusion weight β of the system is β, and its state space integrates indicators such as advertising life cycle, user real-time behavior density and cross-modal feature distribution variance. The action space is the continuous adjustment of β, and the reward function of the fusion click-through rate, cold start exposure rate and user churn rate is used for strategy training, so that the system can achieve a dynamic balance between the content consistency of cold start advertising (such as new products rely on strong association between videos and copywriting) and user preferences for mature advertising (such as personalized recommendations driven by historical behavior); secondly, a sliding window-memory enhancement (SW-MA) behavior encoding mechanism is proposed, which uses a sliding window to retain the user's recent real-time behavior to filter noise, and introduces an external memory matrix to store long-term interest prototypes (such as "sports equipment" and "maternal and child products"), and fuses the real-time behavior sequence with the memory matrix through the attention mechanism to generate a behavior feature vector It is combined with clustering update strategy to capture the principal components of interest and enhance the robustness of behavior modeling.
[0115] On this basis, a multimodal feature enhancement module (MFE) and a cross-modal real-time interaction optimization architecture are designed: the MFE module strengthens the fine-grained alignment of spatiotemporal features with semantic keywords (such as the association between dynamic images and text descriptions) through video-text cross-attention, and uses a gating mechanism to filter video key frames to respond to user behavioral preferences; at the same time, edge computing nodes are deployed to sink multimodal similarity calculations to local devices, and combined with high-frequency behavioral feature caching strategies to optimize response efficiency; in addition, a dynamic negative sampling strategy is introduced to dynamically adjust the distribution of negative samples in contrastive learning based on user activity - for highly active users, cross-category interference samples are added to improve the model's noise resistance, while for cold start users, similar negative samples are reduced to avoid misjudgment, thereby improving recommendation coverage in long-tail scenarios. Through the above solution, scenario adaptation of recommendation weights, deep collaboration of multimodal features, and optimization of real-time computing efficiency are achieved, fully supporting the comprehensive improvement of advertising recommendations in dynamically balancing content authenticity, user personalization needs, and system scalability.
[0116] In one embodiment, the method further comprises:
[0117] Define the processing variable T, the confounding variable C, and the observation result Y; when the processing variable T is 0, the sample is an unexposed sample, and when the processing variable T is 1, the sample is an exposed sample;
[0118] Based on the treatment variable T, the confounding variable C and the observation result Y, a logistic regression model is used to predict P(T=1|C) to obtain a propensity score e(C);
[0119] For the exposed sample, the counterfactual result Y(1) is predicted by the deep network; for the unexposed sample, the counterfactual result Y(0) is predicted;
[0120] Based on the counterfactual result Y(1) and the counterfactual result Y(0), the recommendation weight of each candidate advertisement is determined.
[0121]
[0122] The recommendation weights The recommendation priority of the candidate advertisement is adjusted.
[0123] In this embodiment, a cross-modal contrastive learning framework and an adaptive multi-task recommendation architecture are further constructed, and the dynamic adaptability of the system is improved through a real-time interest drift detection and compensation mechanism. The core of its technical solution lies in:
[0124] First, the semantic consistency between video, text, and user behavior features is optimized and strengthened through cross-modal contrastive learning. The framework uses multimodal data of the same advertisement (such as video frame features, text description features, and user behavior features) as positive sample pairs, and constructs negative sample pairs by randomly replacing features of different modalities (for example, pairing the video of a "sports shoes" advertisement with the text description of a "home appliance" advertisement). The contrast loss function is used to narrow the similarity of the cross-modal features of the positive samples while widening the difference with the negative samples. In addition, the real-time behavior distribution of users is combined to dynamically screen high-confusion negative samples (for example, under the "coffee machine" category that users have frequently browsed but not purchased recently, "coffee beans" are actively added as negative samples) to enhance the model's sensitivity to fine-grained semantic differences, thereby solving the recommendation bias problem caused by semantic fragmentation of cross-modal features.
[0125] Secondly, an adaptive multi-task recommendation architecture was designed to simultaneously optimize multiple objectives, including click-through rate prediction, conversion rate prediction, and user dwell time. This architecture dynamically assigns weights to different tasks through a gating network. For example, based on the ad type (e.g., e-commerce ads focus on conversion rates, content ads focus on user dwell time) and the user's historical behavior, it automatically adjusts the contribution ratio of the loss function for tasks such as click-through rate and conversion rate to achieve a dynamic balance between multiple objectives. In terms of model structure, the bottom layer uses shared video, text, and behavioral feature encoders to extract common representations, while the top layer independently designs prediction networks for each task, avoiding repeated feature calculations and preventing mutual interference between tasks.
[0126] Furthermore, in response to the dynamic changes in user interests, claim 4 proposes a real-time interest drift detection and compensation mechanism. This mechanism is based on the distribution difference analysis of user behavior sequences (for example, detecting the mutation from "maternal and child products" to "outdoor equipment" through KL divergence), triggering local updates to the behavior modeling module: in short-term interest compensation, the real-time behavior sequence within the sliding window is used to quickly adjust the interest prototype weights in the memory enhancement module; in long-term interest maintenance, the low-dimensional embedding representation of historical interests is retained to prevent short-term noise from covering long-term stable preferences. At the same time, the system only performs incremental hot updates on relevant submodules (such as behavior encoders or multi-task weight gating networks) where interest drift is detected, rather than global model iteration, which significantly reduces computing resource consumption.
[0127] Through the above solution, we achieve deep alignment of cross-modal semantics, flexible adaptation of multi-objective tasks, and rapid response to user interest drift, thereby ensuring the consistency, diversity, and timeliness of recommendation results in complex dynamic scenarios, and providing full-link enhanced support for advertising recommendation systems from feature learning to decision optimization.
[0128] The above are some specific implementations of the multimodal advertising recommendation method provided in the embodiments of the present application. Based on this, the present application also provides a corresponding device. The device provided in the embodiments of the present application will be introduced from the perspective of functional modularization.
[0129] See also Figure 2 The schematic structural diagram of the multimodal advertising recommendation device shown in FIG. 200 includes an acquisition module 201 , a construction module 202 , an extraction module 203 , a calculation module 204 , a loss module 205 , an update module 206 and a recommendation module 207 .
[0130] In one embodiment, the updating module 206 is specifically configured to calculate the cross-modal similarity between a target advertisement and a plurality of candidate advertisements in an advertisement library; the target advertisement is an advertisement viewed by a user;
[0131] The multiple cross-modal similarities are sorted, and the candidate advertisement corresponding to the highest cross-modal similarity is selected and recommended to the user.
[0132] In one embodiment, the calculating the cross-modal similarity between the target advertisement and a plurality of candidate advertisements in the advertisement library includes:
[0133]
[0134] Among them, S 召回 is the cross-modal similarity between the target advertisement and the candidate advertisement, is the video feature of the target advertisement, is the video feature of the candidate advertisement, is the text feature of the target advertisement, is the text feature of the candidate advertisement, Behavioral characteristics of the target advertisement, is the behavioral feature of the candidate advertisement, α is the weight of the video modality, β is the weight of the text modality, and γ is the weight of the behavioral modality.
[0135] In one embodiment, the recommendation module 2027 is further configured to:
[0136] Define the processing variable T, the confounding variable C, and the observation result Y; when the processing variable T is 0, the sample is an unexposed sample, and when the processing variable T is 1, the sample is an exposed sample;
[0137] Based on the treatment variable T, the confounding variable C and the observation result Y, a logistic regression model is used to predict P(T=1|C) to obtain a propensity score e(C);
[0138] For the exposed sample, the counterfactual result Y(1) is predicted by the deep network; for the unexposed sample, the counterfactual result Y(0) is predicted;
[0139] Based on the counterfactual result Y(1) and the counterfactual result Y(0), the recommendation weight of each candidate advertisement is determined.
[0140]
[0141] The recommendation weights The recommendation priority of the candidate advertisement is adjusted.
[0142] The embodiments of the present application also provide corresponding devices and computer storage media for implementing the solutions provided by the embodiments of the present application.
[0143] The device includes a memory and a processor, the memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes a multimodal advertising recommendation method described in any embodiment of the present application.
[0144] The computer storage medium stores codes. When the codes are executed, the device executing the codes implements a multimodal advertising recommendation method according to any embodiment of the present application.
[0145] The "first" and "second" (if any) in the names mentioned in the embodiments of this application are only used as name identifiers and do not represent the first or second in order.
[0146] Through the description of the above embodiments, it can be known that those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a general hardware platform. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a read-only memory (ROM) / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network communication device such as a router) to execute the methods described in each embodiment or certain parts of the embodiments of the present application.
[0147] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. Those of ordinary skill in the art can understand and implement it without paying any creative work.
[0148] The above description is merely an exemplary embodiment of the present application and is not intended to limit the scope of protection of the present application.
Claims
1. A multimodal advertising recommendation method, characterized in that: include: Get the video frame sequence corresponding to the same content from the short video platform Text description and user behavior sequence According to the video frame sequence The text description and the user behavior sequence Constructing triplet dataset Use a pre-trained 3D-CNN model to extract the video frame sequence Extract video features Describe the text using the BERT model Extract text features Encoding the user behavior sequence based on Transformer Generate behavioral features According to the triplet dataset The video features The text features and the behavioral characteristics Calculating video-text similarity and video-behavior similarity According to the video-text similarity Similarity to the video-behavior Construct a cross-modal contrast loss function: Wherein, τ is the temperature coefficient, which is used to control the sharpness of the probability distribution. The smaller the value, the higher the discrimination between similar samples. Minimizing the cross-modal contrast loss function through back-propagation to update the weights of the video modality, text modality, and behavioral modality; Based on the updated weights of each modality, the 3D-CNN model, the BERT model, and the Transformer are combined to perform advertisement recommendation.
2. The method according to claim 1, characterized in that The advertisement recommendation based on the updated weights of each modality, combined with the 3D-CNN model, the BERT model, and the Transformer, includes: Calculating cross-modal similarity between a target ad and multiple candidate ads in an ad library; the target ad is an ad viewed by a user; The multiple cross-modal similarities are sorted, and the candidate advertisement corresponding to the highest cross-modal similarity is selected and recommended to the user.
3. The method according to claim 2, characterized in that The calculating of the cross-modal similarity between the target advertisement and a plurality of candidate advertisements in the advertisement library includes: Among them, S 召回 is the cross-modal similarity between the target advertisement and the candidate advertisement, is the video feature of the target advertisement, is the video feature of the candidate advertisement, is the text feature of the target advertisement, is the text feature of the candidate advertisement, Behavioral characteristics of the target advertisement, is the behavioral feature of the candidate advertisement, α is the weight of the video modality, β is the weight of the text modality, and γ is the weight of the behavioral modality.
4. The method according to claim 1, wherein The method further comprises: Define the processing variable T, the confounding variable C, and the observation result Y; when the processing variable T is 0, the sample is an unexposed sample, and when the processing variable T is 1, the sample is an exposed sample; Based on the treatment variable T, the confounding variable C and the observation result Y, a logistic regression model is used to predict P(T=1|C) to obtain a propensity score e(C); For the exposed sample, the counterfactual result Y(1) is predicted by the deep network; for the unexposed sample, the counterfactual result Y(0) is predicted; Based on the counterfactual result Y(1) and the counterfactual result Y(0), the recommendation weight of each candidate advertisement is determined. The recommendation weights The recommendation priority of the candidate advertisement is adjusted.
5. A multimodal advertising recommendation device, characterized in that: include: The acquisition module is used to obtain the video frame sequence corresponding to the same content from the short video platform Text description and user behavior sequence A construction module is used to construct a video frame sequence according to the video frame sequence. The text description and the user behavior sequence Constructing triplet dataset Extraction module for extracting the video frame sequence from the pre-trained 3D-CNN model Extract video features Describe the text using the BERT model Extract text features Encoding the user behavior sequence based on Transformer Generate behavioral features A calculation module is used to calculate the triple data set according to the The video features The text features and the behavioral characteristics Calculating video-text similarity and video-behavior similarity Loss module, for determining the video-text similarity Similarity to the video-behavior Construct a cross-modal contrast loss function: Wherein, τ is the temperature coefficient, which is used to control the sharpness of the probability distribution. The smaller the value, the higher the discrimination between similar samples. An updating module, configured to update the weights of the video modality, the text modality, and the behavioral modality by minimizing the cross-modal contrast loss function through backpropagation; A recommendation module is used to recommend advertisements based on the updated weights of each modality in combination with the 3D-CNN model, the BERT model, and the Transformer.
6. The device according to claim 5, characterized in that The updating module is specifically configured to calculate the cross-modal similarity between a target advertisement and a plurality of candidate advertisements in an advertisement library; the target advertisement is an advertisement browsed by a user; The multiple cross-modal similarities are sorted, and the candidate advertisement corresponding to the highest cross-modal similarity is selected and recommended to the user.
7. The device according to claim 6, characterized in that The calculating of the cross-modal similarity between the target advertisement and a plurality of candidate advertisements in the advertisement library includes: Among them, S 召回 is the cross-modal similarity between the target advertisement and the candidate advertisement, is the video feature of the target advertisement, is the video feature of the candidate advertisement, is the text feature of the target advertisement, is the text feature of the candidate advertisement, Behavioral characteristics of the target advertisement, is the behavioral feature of the candidate advertisement, α is the weight of the video modality, β is the weight of the text modality, and γ is the weight of the behavioral modality.
8. The device according to claim 1, characterized in that The recommendation module is also used to: Define the processing variable T, the confounding variable C, and the observation result Y; when the processing variable T is 0, the sample is an unexposed sample, and when the processing variable T is 1, the sample is an exposed sample; Based on the treatment variable T, the confounding variable C and the observation result Y, a logistic regression model is used to predict P(T=1|C) to obtain a propensity score e(C); For the exposure sample, predict the counterfactual result Y(1) through the deep network; For the unexposed sample, predict the counterfactual result Y(0); Based on the counterfactual result Y(1) and the counterfactual result Y(0), the recommendation weight of each candidate advertisement is determined. The recommendation weights The recommendation priority of the candidate advertisement is adjusted.
9. A computer device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the multimodal advertising recommendation method according to any one of claims 1 to 4 is implemented.
10. A computer storage medium, characterized in that The computer storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the multimodal-based advertising recommendation method according to any one of claims 1 to 4.
Citation Information
Cited By
User development analysis method for video recommendation
CN120763361A
Multi-modal deep learning garment live broadcast real-time conversion rate prediction method and system
CN120996861A
Mobile application third-party library recommendation method, medium, equipment and product
CN121071239A
Multi-modal data set construction and dynamic updating method and device based on large model
CN121561445A
Intelligent replacement method for short play video advertisement items
CN121937616A