An advertisement creative automatic generation and recommendation method and system based on deep learning

By constructing a multimodal fusion model using deep learning technology, the limitations of a single modality in advertising creative generation are solved. This enables the collaborative generation and accurate recommendation of multimodal creatives, improves the logical coherence of advertising creatives and user experience, and adapts to rapid market iteration.

CN122367552APending Publication Date: 2026-07-10
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Filing Date
2026-04-10
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing advertising creative generation and recommendation technologies have limitations in single-modal generation, resulting in homogenized creative content, logical confusion, insufficient recommendation accuracy, inability to iterate quickly, and high costs due to reliance on manual intervention.

Method used

Employing a deep learning-based multimodal fusion model, features are extracted from multi-source heterogeneous data to construct a feature library. Using Transformer, an improved diffusion model, and an encoder-decoder architecture, advertising copy, images, and video scripts are generated. Compliance, logical consistency, and performance prediction metrics are used for verification, enabling the collaborative generation and accurate recommendation of multimodal creatives.

Benefits of technology

It achieves full-chain automated generation of advertising creatives, improves the logic and coherence of content, reduces the cost of manual correction, quickly adapts to market demands, generates non-homogeneous creatives, and improves user experience and brand recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122367552A_ABST
    Figure CN122367552A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on deep learning's advertisement creative automatic generation and recommendation method and system, it is related to advertisement production and push technical field, method includes based on the construction feature library of multiple source heterogeneous data;Based on the target feature vector of demand instruction from feature library acquisition, input pre-trained multi-modal fusion deep learning model and obtain advertisement script sequence;Collaborative generation semantic alignment advertisement image and video script sequence, determine initial advertisement creative scheme;Based on channel style mapping rule and brand tonality constraint parameter carries out feature space transformation, obtains style adaptation after candidate creative scheme;Based on compliance, logical consistency and effect prediction index check, not pass then feedback, second generation, until target advertisement creative scheme;Sort distribution using intelligent recommendation algorithm, and dynamically optimize recommendation strategy in combination with real-time delivery feedback data.The application realizes the collaborative automatic generation of multi-modal creativity, quality check and accurate recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method and system for automatic generation and recommendation of advertising creatives based on deep learning. Background Technology

[0002] With the rapid development of the digital marketing market, the demand for advertising creatives is growing exponentially, and their quality and accuracy directly determine the conversion rate and return on investment (ROI) of advertising campaigns. Currently, the generation and recommendation process of advertising creatives mainly relies on manual creation or basic monomodal AI-assisted tools.

[0003] In the traditional advertising creative production model, operations staff or design teams need to manually complete tasks such as copywriting, graphic design, and short video script writing based on product selling points and placement requirements, and then select suitable content for placement based on experience.

[0004] However, this model has significant limitations: on the one hand, manual creation is limited by individual experience and inspiration, resulting in long production cycles and high costs, making it difficult to cope with massive deployment demands and the rapid iteration of the market, and easily leading to serious homogenization of creative content; on the other hand, although some existing technical solutions have introduced AI-assisted generation tools, they are mostly limited to independent processing of a single modality (such as generating only text), lacking deep integration with product characteristics, target audience preferences, and deployment scenario characteristics, resulting in a lack of semantic connection and logical coherence between the generated copy, images, or scripts, making it impossible to form a complete creative solution with collaborative linkage, and often still requiring a lot of manual intervention and correction before it can be put into use.

[0005] In addition, existing creative recommendation strategies are mostly based on simple mapping of single statistical indicators such as historical click-through rates, without combining user profiles, advertising scenarios and product characteristics for multi-dimensional matching, resulting in insufficient recommendation accuracy and a poor user experience.

[0006] Therefore, how to overcome the limitations of single-modal generation and achieve collaborative automatic generation, quality verification, and accurate recommendation of multimodal creatives has become a pressing technical problem in the current advertising technology field. Summary of the Invention

[0007] To overcome the shortcomings of existing technologies, break through the limitations of single-modal generation, and achieve collaborative automatic generation, quality verification, and accurate recommendation of multimodal creatives, this application provides a method and system for automatic generation and recommendation of advertising creatives based on deep learning.

[0008] Firstly, the objective of this invention is achieved through the following technical solution: A method for automatic generation and recommendation of advertising creatives based on deep learning, comprising: Based on multi-source heterogeneous data, core semantic features of products, user interest and preference features, and channel style constraint features are extracted to construct a feature library; Based on the delivery demand instruction, the target feature vector is obtained from the feature library and input into the pre-trained multimodal fusion deep learning model to obtain the advertising copy sequence. Based on the advertising copy sequence and the target feature vector, a semantically aligned advertising image and video script sequence is generated to determine the initial advertising creative scheme. Based on preset channel style mapping rules and brand tone constraint parameters, feature space transformation is performed on the copywriting style, image color distribution and video script narrative rhythm of the initial advertising creative scheme to obtain a style-adapted candidate creative scheme. The candidate creative solutions are verified based on compliance, logical consistency and effect prediction indicators. If the verification fails, a feedback signal is generated and sent back to the multimodal fusion deep learning model to perform secondary generation until the target advertising creative solution is determined. Based on the target audience profile, the characteristics of the delivery channels, and the preset conversion goals, the target advertising creative schemes are prioritized and distributed using intelligent recommendation algorithms; and the recommendation strategy is dynamically optimized by combining real-time delivery feedback data.

[0009] By adopting the above technical solution, the multi-source heterogeneous data includes multi-dimensional product attribute data, user behavior data, channel environment data, and historical high-performance creative data; the target advertising creative solution is a verified advertising creative solution. To achieve collaborative generation of multimodal creatives and significantly improve content logic and coherence, this application constructs a multimodal fusion deep learning model that deeply integrates core product semantics, user preferences, and channel style features. This model can collaboratively generate copy, images, and video scripts based on the same set of target feature vectors. This method breaks through the limitations of independent generation by existing single-modal tools, ensuring a high degree of alignment and logical consistency between copy semantics, visual images, and script narrative. It effectively solves the problems of fragmented and logically chaotic generated content, significantly reduces the cost of manual post-editing, and achieves a leap from "single-point assistance" to "automatic generation of the entire campaign." To achieve dynamic style adaptation of advertising content and accurately match brand tone and placement channel scenarios, this application introduces a feature space transformation mechanism based on channel style mapping rules and brand tone constraint parameters. This mechanism can adaptively adjust the stylistics, color distribution, and narrative rhythm of the initial creative solution. This allows generated creative content to automatically align with the specific style requirements and brand consistency guidelines of different advertising channels (such as the vibrancy of short videos and the simplicity of search ads), significantly improving user experience and brand recognition. It also solves the technical challenges of disconnect between creative content and advertising scenarios, and inconsistent styles found in existing technologies. Furthermore, it incorporates a multi-dimensional quality verification process, including compliance, logical consistency, and performance prediction metrics, and establishes an iterative update mechanism for ad verification and feedback. This application, through automated generation, adaptation, verification, and recommendation, changes the traditional model relying on manual creation, shortening the ad creative production cycle from "days / hours" to "seconds / minutes." This not only greatly reduces labor and time costs but also enables the rapid production of diverse and non-homogeneous creative solutions to meet massive advertising demands, adapting to the fast-paced iteration of the digital marketing market.

[0010] In a preferred embodiment of this application: the step of collaboratively generating a semantically aligned sequence of advertising images and video scripts based on the advertising copy sequence and the target feature vector includes: Based on the Transformer architecture, an advertising copy token sequence is generated using the target feature vector. Based on the improved diffusion probability model, the semantic embedding of the advertising copy token sequence is fused with the target feature vector across modalities to generate an advertising image that satisfies the constraints in both semantic content and visual style. The sequence generation model based on the encoder-decoder architecture uses the advertising copy token sequence, product selling point features and user preference features to autoregressively generate video script sequences. If the confidence level of any generated result is lower than the preset confidence threshold, then based on other generated deterministic results as constraints, the generation space of the model is resampled, and a globally consistent initial advertising creative scheme is calculated and determined.

[0011] By adopting the above technical solutions, the improved diffusion probability model is obtained through pre-training based on a noise prediction mechanism guided by semantic embedding. A confidence-based global consistency resampling mechanism is introduced. When the confidence of any modality's generation result is insufficient, the deterministic results of other modalities are used as constraints for resampling, effectively overcoming the common technical defects of "inconsistency between text and image" or "disconnect between script and visuals" in traditional multimodal generation. This mutually constrained generation method not only ensures a high degree of consistency in semantic content and visual style for advertising creatives but also significantly improves the logical coherence and user comprehension of the generated content.

[0012] In a preferred embodiment, the verification of the candidate creative solutions based on compliance, logical consistency, and effect prediction metrics includes: The pre-trained violation content recognition model is used to scan the text and image content of candidate creative solutions, identify and mark abnormal areas containing false advertising, sensitive words or violation visual elements, and eliminate solutions with compliance risks. The cosine similarity between the key semantic vector of the advertising copy and the deep visual feature vector of the advertising image is calculated, and the correlation score between the plot logic structure of the video script and the core selling points of the product is quantified. If the cosine similarity or correlation score is lower than the preset logical consistency threshold, it is determined to be a multimodal logical mismatch, and a feedback label indicating the mismatch type is generated. The effect prediction model trained based on historical delivery data calculates the estimated conversion index of candidate creative solutions. If the estimated conversion index is lower than the preset effect baseline, it is determined to be an inefficient solution. For schemes that are judged to be multimodal logic mismatched or inefficient, the corresponding feedback label or effect gap value is converted into a negative reward signal in the reinforcement learning algorithm, which is used to update the loss function of the multimodal fusion deep learning model, update the model parameters, and recalculate new candidate creative schemes.

[0013] By employing the aforementioned technical solutions, a three-dimensional verification method is constructed, encompassing compliance scanning, multimodal logic quantification, and effect prediction. This method can accurately identify and eliminate low-quality solutions containing false advertising, sensitive content, and mismatches between text and graphics logic. By transforming feedback labels or effect gap values ​​that fail verification into negative reward signals in reinforcement learning, the loss function and parameters of the multimodal fusion model are directly updated. This enables the multimodal fusion deep learning model to "self-correct" and continuously evolve against error patterns. This self-correction method not only significantly reduces the cost of manual review and compliance risks but also forces the model to actively avoid inefficient features during training. Consequently, it significantly improves the predicted conversion rate of the generated solutions and the internal logical consistency of the multimodal content, ensuring that the output creative ideas possess both safety and high commercial value.

[0014] In a preferred embodiment, this application describes the construction of a feature library based on extracting core semantic features of the product, user interest and preference features, and channel style constraint features from multi-source heterogeneous data, including: Collect structured attribute data and unstructured graphic descriptions of products, use a pre-trained language model to extract the core selling points of the products, calculate the co-occurrence frequency of each entity in historical high-conversion samples, and generate a core semantic feature vector of the products with semantic density weights. A user-item interaction graph is constructed based on the user's historical behavior sequence. The graph attention network is used to aggregate neighbor node information, extract the user's explicit preference features, and use the product's core semantic feature vector as the query key to retrieve the implicit intent features that are semantically aligned with it in the user behavior space. The features are then fused to generate a user interest preference feature vector, so as to realize the semantic space mapping between user preferences and product selling points. Analyze historical top creative samples from different advertising channels, extract visual color histograms, copywriting emotional polarity distribution, and video rhythm and beat features, cluster to generate channel style prototype vectors, and calculate the style compatibility matrix between the channel style prototype vectors and the product core semantic feature vectors. Select style constraint parameters with compatibility higher than a preset threshold as channel style constraint features. The semantically aligned user interest preference feature vectors and the channel style constraint features that have been filtered for compatibility are associated and mapped to the same feature space index to construct a feature library that includes static attributes and dynamic relationships.

[0015] By employing the aforementioned technical solution, core selling points of products with semantic density weights are extracted, and implicit intentions with semantic alignment are retrieved in the user behavior space using graph attention networks. This achieves a deep semantic space mapping between user preferences and product selling points, solving the intention recognition bias problem caused by relying solely on explicit behavior in traditional recommendations. Simultaneously, by calculating the style compatibility matrix between channel style prototypes and product semantics, highly compatible style constraint parameters are selected, ensuring that the generated creative content naturally aligns with the ecological characteristics of the distribution channels.

[0016] In a preferred embodiment, this application also includes: Real-time capture of click-through rate, conversion rate, and user dwell time feedback from the ad delivery platform, and mapping them into performance reward signals; The implicit intent weights in the user interest preference feature vector are updated using the effect reward signal: if a creative solution corresponding to a certain type of implicit intent feature receives a positive reward, its activation priority in the feature library is increased, and the semantic density weights of the corresponding selling points in the associated product core semantic feature vector are enhanced in reverse. Simultaneously, monitor external hot event streams and calculate the semantic similarity between hot keywords and existing product core semantic features in the feature library; if the similarity exceeds the dynamic injection threshold, generate a temporary enhanced feature vector based on the hot keywords and perform weighted fusion with the user interest preference feature vector to obtain a timeliness enhanced feature vector; The timeliness enhancement features are written into the real-time cache of the feature library and a lifespan is set. After the lifespan expires, the features are automatically removed or downgraded, so as to realize the adaptive evolution of the feature library to market dynamics.

[0017] By adopting the above technical solution, based on the dynamic evolution mechanism of the feature library during real-time delivery feedback, the weight of high-value implicit intentions and product selling points is reinforced through performance reward signals, forming a coupled reinforcement loop between feedback and features, enabling the system to quickly converge to the optimal feature combination. Simultaneously, the system introduces a monitoring mechanism for external hot topic events and a temporary feature injection mechanism, and sets a time-to-live (TTL) lifecycle to automatically manage time-sensitive features, solving the problems of delayed response to market hot topics and rigid feature libraries in traditional advertising systems.

[0018] In a preferred example, this application describes a feature space transformation of the copywriting style, image color distribution, and video script narrative rhythm of the initial advertising creative scheme, based on preset channel style mapping rules and brand tone constraint parameters. This includes: Construct a style transfer loss function, which includes a stylistic difference term, a color distribution distance term, and a narrative rhythm deviation term; By using the discriminator in the generative adversarial network, the style authenticity of the text, images and video scripts of the initial advertising creative scheme is judged, and the style discrimination gradient is obtained. Based on the style discrimination gradient and the brand tone constraint parameters, the latent space encoding of the generator is adjusted through the gradient backpropagation algorithm to minimize the style transfer loss function while keeping the original semantic content unchanged. The above transformation process is executed iteratively until the style discrimination gradient converges, resulting in a style-adapted candidate creative solution.

[0019] By employing the aforementioned technical solutions, a multimodal style transfer loss function encompassing register, color, and narrative rhythm was constructed. Combined with gradient feedback from an adversarial generative network discriminator, this enabled refined style reshaping of the initial creative proposal. Embedding brand tone constraint parameters into the discriminator input constrained the decision boundary, and introducing a content consistency regularization term while minimizing style loss, effectively addressed the common challenges of "semantic drift" or "loss of brand characteristics" during style transfer. This ensured that the output candidate creative proposals strictly adhered to channel style and brand tone requirements in visual, auditory, and textual representation, while maintaining the coherence of the original selling points in the semantic space, significantly enhancing the brand recognition of the advertising creative.

[0020] Secondly, the objective of this invention is achieved through the following technical solution: A deep learning-based automatic advertising creative generation and recommendation system is provided for executing the deep learning-based automatic advertising creative generation and recommendation method described above. The system includes: The feature library construction module is used to extract core semantic features of products, user interest and preference features, and channel style constraint features based on multi-source heterogeneous data, and to build a feature library. The initial creative generation module is used to obtain target feature vectors from the feature library based on the delivery demand instructions, input the target feature vectors into a pre-trained multimodal fusion deep learning model to obtain an advertising copy sequence, and generate semantically aligned advertising images and video script sequences based on the advertising copy sequence and the target feature vectors, thereby determining the initial advertising creative scheme. The style adaptation and transformation module is used to perform feature space transformation on the copywriting style, image color distribution and video script narrative rhythm of the initial advertising creative scheme based on preset channel style mapping rules and brand tone constraint parameters, so as to obtain the candidate creative scheme after style adaptation. The verification and feedback optimization module is used to verify the candidate creative schemes based on compliance, logical consistency and effect prediction indicators; if the verification fails, a feedback signal is generated and sent back to the multimodal fusion deep learning model to perform secondary generation until the target advertising creative scheme is determined. The intelligent recommendation and distribution module is used to prioritize and distribute the target advertising creative schemes based on the target audience profile, the characteristics of the delivery channel, and the preset conversion target, and to dynamically optimize the recommendation strategy by combining real-time delivery feedback data.

[0021] Thirdly, the objective of this invention is achieved through the following technical solution: A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned method for automatically generating and recommending advertising creatives based on deep learning.

[0022] Fourthly, the objective of this invention is achieved through the following technical solution: A computer program product includes a computer program / instructions that, when executed by a processor, implement the steps of a deep learning-based automatic generation and recommendation method for advertising creatives as described above.

[0023] In summary, this application includes at least one of the following beneficial technical effects: 1. A multi-dimensional feature library integrating product semantics, user intent, and channel style is constructed, breaking the data silos limitation in traditional advertising generation and realizing full-link automated generation from demand instructions to multi-modal creatives (copy, images, and videos); by introducing compliance, logical consistency, and effect prediction indicators to perform multi-dimensional quality verification and filtering of candidate solutions, and by driving the model to iterate twice through feedback signals, the problems of high risk of creative content violations, disconnect between text and image logic, and uncontrollable conversion effects in existing technical solutions are effectively solved. 2. By employing the collaborative work of Transformer, improved diffusion model and encoder-decoder architecture, cross-modal semantic depth alignment of text, images and video scripts is achieved; in particular, a confidence-based global consistency resampling mechanism is introduced, which uses the deterministic results of other modalities as constraints to resample when the confidence of the generated result of any modality is insufficient, effectively overcoming the technical defects of "text and image mismatch" or "script and video disconnection" commonly found in traditional multimodal generation. Attached Figure Description

[0024] Figure 1 This is a flowchart of a deep learning-based method for automatically generating and recommending advertising creatives in one embodiment of this application; Figure 2 This is another flowchart in one embodiment of a deep learning-based method for automatically generating and recommending advertising creatives. Detailed Implementation

[0025] The present application will be further described in detail below with reference to the accompanying drawings.

[0026] In one embodiment, such as Figure 1 As shown, this application discloses a method for automatic generation and recommendation of advertising creatives based on deep learning, which specifically includes the following steps: S1: Extract core semantic features of products, user interest and preference features, and channel style constraint features from multi-source heterogeneous data to construct a feature library.

[0027] In this embodiment, multi-source heterogeneous data refers to data sets originating from different business systems with varying data structures. Specifically, it includes: structured data (such as specification parameter tables and user profile tag libraries in product databases), semi-structured data (such as server logs and JSON-formatted user behavior records), and unstructured data (such as image and text comments on social media and historical advertising video files). The core semantic features of a product refer to a dense vector representation extracted from the product description using natural language processing technology, capable of representing the essential attributes and core selling points of the product. This not only includes the literal meaning of keywords but also implicitly contains the semantic weight of selling points in historical high-conversion scenarios. User interest and preference features refer to a fusion vector of explicit preferences (such as clicked and saved categories) and implicit intentions (such as potential purchase motives and aesthetic preferences) mined from users' historical interaction behavior, used to represent the user's acceptance of specific content styles.

[0028] Channel style constraint features refer to the quantitative parameters of visual color distribution patterns, copywriting style and emotional polarity distribution, and video narrative rhythm characteristics unique to specific advertising platforms (such as short video platforms and news feed platforms). These parameters are used to constrain the form of generated content to conform to the platform's ecosystem. The feature library is a high-dimensional vector index space used to store the aforementioned feature vectors, which have undergone semantic alignment and weighting, and their dynamic relationships, supporting rapid retrieval and updates.

[0029] Specifically, step S1 includes: S11: Collect structured attribute data and unstructured graphic descriptions of the product, use a pre-trained language model to extract the core selling points of the product, calculate the co-occurrence frequency of each entity in historical high-conversion samples, and generate a core semantic feature vector of the product with semantic density weights.

[0030] In this embodiment, the core selling point entity refers to the key attribute nodes extracted from the complex product description information that can directly drive users' purchase decisions, such as ultra-long battery life, aerospace-grade aluminum alloy, and quiet noise reduction. These are not only keywords but also semantic units carrying specific marketing value. Semantic density weight is a quantitative indicator used to characterize the contribution and activity of a certain selling point entity in historical high-conversion advertising scenarios; the higher the weight, the higher the co-occurrence frequency of the selling point with user clicks and conversion behaviors in past successful cases, and the higher attention priority it should be given when generating new creatives. Structured attribute data refers to parameter tables with fixed fields in the database, such as size, weight, and material codes, while unstructured graphic and text descriptions refer to free text on e-commerce detail pages, user reviews, and OCR text in promotional posters.

[0031] Specifically, the system first connects to the product database and content management system, reading structured attribute tables and unstructured text / image data in parallel. For unstructured data, the system calls a pre-trained large language model (such as LLaMA or BERT variants) to perform named entity recognition (NER) and key phrase extraction, initially screening out a set of candidate selling points. Subsequently, the system reviews the advertising logs from the past six months to construct a "selling point-conversion" co-occurrence matrix: it counts the frequency of each candidate selling point in historical high-conversion samples (defined as creatives with conversion rates higher than the top 20%) and the frequency of its occurrence in low-conversion samples.

[0032] For example, for a "cordless vacuum cleaner," the system identifies three candidate entities: "high suction power," "lightweight," and "long battery life." Statistical analysis reveals that "long battery life" appears 85% of the top 20% of high-conversion creatives, but only 30% of low-conversion creatives; conversely, the frequency of "high suction power" is similar in both categories. Based on this, the system calculates a semantic density weight of 0.9 (high) for "long battery life" and 0.5 (medium) for "high suction power." Finally, the system maps each selling point entity to a high-dimensional vector (Embedding) and uses the calculated semantic density weight as an additional scalar attribute of the vector or a specific dimension scaling factor embedded in the vector, generating a product core semantic feature vector with value tags.

[0033] S12: Construct a user-item interaction graph based on the user's historical behavior sequence, use a graph attention network to aggregate neighbor node information, extract explicit user preference features, and use the product's core semantic feature vector as the query key to retrieve implicit intent features that are semantically aligned with the user's behavior space. These features are then fused to generate a user interest preference feature vector, thereby achieving a semantic space mapping between user preferences and product selling points.

[0034] In this embodiment, the user-item interaction graph is a dynamically constructed bipartite or multipartite graph structure. Nodes include user ID, item ID, viewed page ID, and search keyword ID, while edges represent interactive behaviors such as clicking, adding to favorites, adding to cart, and purchasing, along with their timestamps. "Explicit preference features" refer to clear preferences expressed by users through direct interactive behaviors; for example, frequently clicking on "red" items represents an explicit preference for the color red. Implicit intention features are potential psychological motivations that users do not directly express but are inferred from behavioral sequences, such as pursuing cost-effectiveness, urgently needing stock, or aesthetic upgrades. These features are usually hidden in the temporal patterns of behavioral paths. Semantic space mapping refers to projecting feature vectors originally belonging to different modalities (such as user behavior and product attributes) into the same high-dimensional vector space, making semantically similar user intentions and product selling points spatially closer to each other.

[0035] Specifically, the system collects users' behavior sequences over the past 30 days in real time and constructs a user-centric subgraph. This subgraph is then encoded using a Graph Attention Network (GAT): the GAT model automatically learns the contribution weights of neighboring nodes (such as other products clicked by the user or categories browsed) to the user's current node through a multi-head attention mechanism, thereby aggregating and generating a user explicit preference vector containing contextual information. To uncover implicit intent, the system employs a "cross-modal retrieval" strategy: the product core semantic feature vector generated in step S11 is used as the query key for attention retrieval within the user behavior vector space.

[0036] For example, suppose a user has recently been frequently browsing "outdoor camping gear" and spending a lot of time on it, but hasn't made a purchase. The explicit features extracted by GAT might only point to categories like "tents" and "sleeping bags." When the system uses "lightweight and portable" (product selling point vector) as the query to retrieve the user's behavioral space, it finds that the user has a very high implicit interest in products related to keywords such as "ultra-light carbon fiber" and "foldable storage" (repeatedly comparing products even without purchasing). Based on this, the system infers that the user's implicit intent is "extreme lightweight needs" and generates a corresponding implicit intent vector. Finally, the system weightedly merges the explicit preference vector and the implicit intent vector using a gating mechanism to generate the final user interest preference feature vector.

[0037] S13: Analyze historical top creative samples from different advertising channels, extract visual color histograms, copywriting emotional polarity distribution, and video rhythm and beat features, cluster to generate channel style prototype vectors, calculate the style compatibility matrix between the channel style prototype vectors and the product core semantic feature vectors, and select style constraint parameters with compatibility higher than a preset threshold as channel style constraint features.

[0038] In this embodiment, the channel style prototype vector is a highly abstract summary of the content ecosystem of a specific advertising platform. It is not a single material, but rather the clustering center points of the platform's top-performing content in terms of visual (color, composition), auditory (rhythm, BGM type), and textual (emotional polarity, sentence length) features. The style compatibility matrix is ​​a two-dimensional mapping table used to quantify the degree of fit between the "product's core selling points" and the "channel style." Some hardcore technology selling points may have extremely high compatibility in a geek forum style, but low compatibility in a channel style that emphasizes aesthetics and lifestyle. Forcing such a fusion would result in creative incongruity. The channel style constraint features are the set of style parameters determined after compatibility filtering to be suitable for the product's placement on the current channel, used to guide the subsequent generation of creative forms.

[0039] Specifically, the system periodically crawls samples of "hot topics" or "top creative content" from each target channel over the past month (such as videos / articles with the top 1% likes). It then uses a multimodal encoder to extract their visual color histograms, copywriting sentiment scores (divided into positive, negative, and neutral proportions), and video shot switching frequency (rhythm). Subsequently, K-Means or DBSCAN clustering algorithms are used to divide the samples within the same channel into several "style clusters," and the center vector of each cluster is calculated as the "style prototype" for that channel.

[0040] For example, when analyzing "short video platforms," ​​the system clustered two main style archetypes: Archetype A is characterized by "fast pace, high saturation, and strong plot twists," while Archetype B is characterized by "slow pace, low saturation, and immersive experience." Next, the system calculated the compatibility of the product's core selling point vector (e.g., "precision instruments") with these two archetypes. Historical data regression analysis revealed that the conversion rate for "precision instruments" products was significantly higher in archetype B than in archetype A. Therefore, the compatibility score between "precision instruments" and archetype B was calculated to be 0.9, while it was only 0.3 with archetype A. The system set a threshold of 0.6, automatically filtering out archetype A and retaining only the color range, emotional tone, and rhythm parameters related to archetype B as channel style constraints for the product.

[0041] S14: Associate and map the semantically aligned user interest preference feature vectors with the channel style constraint features that have been filtered for compatibility to the same feature space index, and construct a feature library that includes static attributes and dynamic relationships.

[0042] In this embodiment, the feature space index is a high-dimensional vector database that supports millisecond-level retrieval based on semantic similarity (ANN Search). It stores not only the vectors themselves but also the dynamic relationships between them. Static attributes refer to the inherent characteristics of a product that do not change drastically over time (such as material and function), while dynamic relationships refer to weighted links that are updated in real time according to market trends, real-time user behavior, and channel trends (e.g., if a selling point becomes popular today, its correlation with user intent automatically strengthens). Association mapping refers to anchoring heterogeneous features (users, products, channels) from different sources under the same coordinate system, forming an interactive knowledge network.

[0043] Specifically, the system fuses the user interest preference feature vector generated in step S12 with the channel style constraint features selected in step S13. First, product quantization is used to compress and encode the high-dimensional vector to meet the needs of large-scale real-time retrieval. Then, the system establishes a multi-layer index structure in the vector database: the bottom layer stores the core semantic vector of the product as the basic anchor point; the middle layer establishes an inverse index of user intent vector and product selling points (i.e., the corresponding selling points can be quickly found through intent); and the top layer attaches the channel style constraint parameters as a filter mask.

[0044] For example, when receiving a generation task, the system does not simply concatenate all the data, but rather performs "dynamic assembly" through indexing. For instance, for "young female users" (user vector), the system retrieves the nearest implicit intent of "appearance economy" from the index, and then associates it with the "retro design" selling point of the product (product vector). At the same time, based on the distribution channel, it automatically loads style constraints such as "warm color tone and promotional tone" (channel vector). The system interacts with these three elements in a temporary context window to generate a composite feature package containing a complete semantic chain, and writes it to the real-time cache of the feature library.

[0045] S2: Based on the delivery demand instruction, obtain the target feature vector from the feature library, input it into the pre-trained multimodal fusion deep learning model to obtain the advertising copy sequence; based on the advertising copy sequence and the target feature vector, generate semantically aligned advertising image and video script sequence to determine the initial advertising creative scheme.

[0046] In this embodiment, the delivery requirement instruction refers to the natural language description or structured parameters input by the user regarding the ad generation task, including key information such as product name, promotional theme, and target audience. The ad copy sequence refers to the text token stream with a complete semantic structure generated by the model, constituting the core narrative logic of the ad. Semantic alignment refers to the high consistency of content across different modalities (text, image, video) in the deep feature space, meaning that image content and video plots accurately reflect the selling points and emotions described in the copy.

[0047] A pre-trained multimodal fusion deep learning model refers to a composite neural network built upon a cascaded Transformer encoder-decoder architecture and a diffusion probabilistic model. Its core function is to transform textual semantic instructions into spatially aligned image pixel matrices and temporally coherent video script sequences. Internally, it includes a cross-modal attention mechanism module to enable information interaction between different modalities. The cross-modal attention mechanism module is a key component within the multimodal fusion deep learning model, consisting of a multi-head self-attention layer and a cross-attention layer. It is used to calculate the correlation weight matrix between text tokens and latent variables in images or video frame features.

[0048] Specifically, step S2 includes: S21: Generate advertising copy token sequences based on target feature vectors using the Transformer architecture.

[0049] In this embodiment, a text generation module is constructed based on the Transformer Decoder architecture. The text generation module includes an input embedding layer, a positional encoding layer, a decoding layer (i.e., a multi-layer self-attention encoding layer), and a feedforward neural network layer. Target feature vector. It is a fixed-length floating-point vector with dimension D, obtained by weighted summation of product selling point feature vectors, user preference feature vectors, and channel style feature vectors, such as D=768. The output data is a sequence of advertising copy tokens. ,in is an integer index from the vocabulary, and L is the length of the generated sequence. The target feature vector... Through a linear transformation layer Mapping to the hidden layer dimension H of the Transformer model yields the initial context vector. ,in, For linear transformation layer The bias term.

[0050] Specifically, the output sequence is initialized with the start symbol [SOS]. During the i-th step of generation, the already generated sequence... Input embedding layer, and with The input is the same as that of the Transformer Decoder. The mask calculates the correlation between tokens within the current sequence from the attention layer, ensuring that only tokens before the current time step are considered during generation. The cross-attention layer... Using the decoder's output as the query, the output serves as the key and value, calculating the attention weights of product selling points and user preferences for the currently generated words. After passing through a feedforward network and a softmax layer, the probability distribution of each word in the vocabulary is output. The word with the highest probability is selected using either a greedy search or a beam search (with a beam width of 5). Termination condition: When the end-of-line character [EOS] is generated or the maximum length is reached. Stop at the specified time and output the complete ad copy token sequence. .

[0051] S22: Based on the improved diffusion probability model, the semantic embedding of the advertising copy token sequence is fused with the target feature vector across modalities to generate an advertising image that satisfies the constraints in both semantic content and visual style.

[0052] In this embodiment, an improved architecture based on the Latent Diffusion Model (LDM) is employed to generate an image generation module for producing advertising images. The improved architecture includes an encoder E and a decoder D from a variational autoencoder (VAE), as well as a denoising U-Net network. The U-Net network comprises downsampling blocks, intermediate blocks, and upsampling blocks; its key improvement lies in embedding a cross-attention layer at each resolution level. The input data is: noise latent variables. and joint condition vector Noise latent variables Let be the latent spatial noise matrix at time step t. Joint condition vector. Composed of advertising copy token sequences Semantic embedding vectors extracted by a text encoder (such as CLIP Text Encoder or BERT) , and the target feature vector The vector after linear projection layer mapping It is pieced together. The output data is an advertisement image. The latent variables are finally denoised using the VAE decoder D. Restored to pixel-level RGB image.

[0053] Specifically, the text encoder generated in step S21 is used to... Convert to semantic embedding vector . Target feature vector Mapped to style control vectors through fully connected layers Construct the joint condition vector Next, the initial noise latent variable is sampled from the standard Gaussian distribution N(0, I). Let the total number of denoising steps be T = 50.

[0054] As time step t decreases from T to 1: Time step embedding t and joint condition vector Input the U-Net network. In the cross-attention layer of the U-Net network, compute... Feature maps (as queries) and Attention matrix (as key / value) This step ensures that the generated image content is constrained by the text semantics, and the image style is constrained by the target feature vector. The Net output predicts noise. Update the latent variables according to the sampling formula of the diffusion model (such as the DDIM sampler): , where η is the scheduling coefficient.

[0055] The final pure latent variables Inputting the VAE decoder D, the pixel-level advertisement image is reconstructed. .

[0056] S23: A sequence generation model based on an encoder-decoder architecture, which uses autoregressive methods to generate video script sequences using advertising copy token sequences, product selling point features, and user preference features.

[0057] In this embodiment, an Encoder-Decoder architecture is employed. The encoder uses a bidirectional LSTM (Bi-LSTM) or a Transformer Encoder; the decoder uses an LSTM with attention mechanism or a Transformer Decoder. The input data is a concatenated vector. Includes ad copy token sequence Embedded representation, product selling point feature vector and user preference feature vector The output data is a video script sequence. It describes the storyboard, camera movement, and narration.

[0058] Specifically, in the encoding stage: the advertising copy token sequence Product selling points and features User preference characteristics Each is represented by an embedding.

[0059] The above embedded vectors are concatenated and then input into the encoder. The encoder extracts the contextual semantic state of the input sequence through multiple layers of bidirectional recurrent units or self-attention layers. .

[0060] During the code reception phase: the decoder input is initialized to the start symbol [SOS]. At step k, the decoder receives the output from the previous time step. and encoder status The attention mechanism includes: the decoder calculating the current hidden state and... The attention weights for each position within the text are dynamically focused on key action descriptions or specific style elements reflecting user preferences. The output shows the probability distribution of video script words at the current moment, selecting the word with the highest probability as... .

[0061] Output: Repeat the above process until an end marker is generated, resulting in a complete video script sequence. Its content includes scene descriptions, camera instructions (such as "push-in" and "close-up"), and corresponding narration.

[0062] S24: If the confidence level of any generated result is lower than the preset confidence threshold, then based on other generated deterministic results as constraints, the generation space of the model to which it belongs is resampled, and a globally consistent initial advertising creative scheme is calculated and determined.

[0063] In this embodiment, the confidence level of the generated result includes image confidence level. and script confidence Image confidence The generated image is calculated using a pre-trained image-text matching model (such as CLIP Image-Text MatchingHead). With copywriting Cosine similarity. Script confidence. Script generation using language model Relative to input conditions The conditional probability logarithmic mean (the reciprocal of perplexity), or the semantic implication score of the script and copy.

[0064] Specifically, calculate the image confidence score: Calculate the script confidence score: (Using a pre-trained semantic entailment model). Set a confidence threshold. =0.75.

[0065] In practical applications, the specific judgment logic includes: Scenario A (All Pass): If and Then output directly. As a consistent initial advertising creative solution.

[0066] Scenario B (Image not up to standard): If ,but and The confidence level is sufficient. Locking is now initiated. and For deterministic constraints, re-execute the image generation process in step S22.

[0067] During the re-denoising process, the semantic embedding of the text in the cross-attention layer is increased. The weighting coefficients (e.g., multiplied by a magnification factor α>1), or the introduction of classifier guidance, utilize... The key visual elements described in the text (such as "running" and "red") construct gradient guidance terms, forcing The update direction converges towards these elements. Repeat the denoising process until a new image is generated. Or the maximum number of retries (e.g., 3 times) is reached.

[0068] Scenario C (Script not up to standard): If ,but and The confidence level is sufficient. Locking is now initiated. and For deterministic constraints, re-execute the script generation process in step S23 and perform constraint enhancement operations: [The constraints will be...] Visual feature vectors are extracted using an image encoder. This is then used as an additional condition input into each step of the attention calculation in the decoder. This forces the script-generated model to be compatible with the scene description. The visual content must remain consistent (e.g., if a person is on the left in an image, the script must not describe them as "on the right"). Beam search is used to expand the search space and select the script sequence with the highest probability that satisfies the visual constraints.

[0069] Case D (Multiple failures): If multiple conditions are below the threshold, the modality with the highest confidence is retained as the constraint, and the above resampling operation is performed on the other modalities in sequence, or the process is returned to step S21 to regenerate the text.

[0070] S3: Based on the preset channel style mapping rules and brand tone constraint parameters, the feature space transformation is performed on the copywriting style, image color distribution and video script narrative rhythm of the initial advertising creative scheme to obtain the style-adapted candidate creative scheme.

[0071] In this embodiment, channel style mapping rules refer to predefined operational logic or mathematical mapping relationships that convert general creative content into a specific platform style, such as mapping a general color space to a high-saturation color space preferred by a certain platform. Brand tone constraint parameters refer to quantitative indicators that characterize the brand image, such as the RGB range of the brand's main color, the emotional polarity threshold of the copywriting style, and the pacing coefficient of the video narrative, used to ensure that the generated content does not deviate from the brand's core values. Feature space transformation refers to the process of adjusting the position of the content in the potential feature space through mathematical operations, while keeping the semantics of the content unchanged, so that it falls into the target style area. Candidate creative solutions refer to the combination of multimodal advertising content that has undergone style adaptation processing and is awaiting final verification.

[0072] Specifically, step S3 includes: S31: Construct a style transfer loss function, which includes a stylistic difference term, a color distribution distance term, and a narrative rhythm deviation term.

[0073] In this embodiment, a copywriting style rule library is pre-built, storing standard style templates for different advertising channels. These standard style templates predefine copywriting style, image color schemes, and video rhythm. The copywriting style defines a keyword library, sentence structure preferences (e.g., short vs. long sentences), frequency of modal particles, and emoji usage guidelines. Image color schemes define a standard color palette, a hue distribution histogram, and saturation and brightness thresholds. Video rhythm defines the average shot length (ASL), transition frequency, and background music beat matching curve. A style transfer loss function calculates the multidimensional distance between the current generated scheme and the target channel's style, outputting a scalar loss value. .

[0074] Specifically, based on the target advertising channel (e.g., "Xiaohongshu") and brand image, a style transfer loss function is dynamically assembled. The expression is:

[0075] in, These are weighting coefficients that are dynamically adjusted based on channel characteristics. Stylistic differences item. The calculation expression is: By extracting the initial copy stylistic feature vectors This includes confusion level, emotional polarity, and rhetorical density. The standard stylistic feature vector of the target channel; where Penalties for violating brand image (such as using prohibited words).

[0076] Color distribution distance ( ): Calculate the initial image Color histogram Gray-Level Co-occurrence Matrix (GLCM) texture features. Obtain the standard color distribution of the target channel. For example, a certain channel might prefer high saturation and warm colors. Calculate Earth Mover's Distance (EMD): .

[0077] Narrative rhythm deviation ( ): Parsing the initial video script Extracting the shot switching frequency sequence from the timeline. And mood fluctuation curve An ideal timeline template for acquiring target channels. For example, a certain channel might prefer a fast-paced, high-controversy narrative in the first 3 seconds. Calculate the Dynamic Time Warping (DTW) distance: .

[0078] S32: Using the discriminator in the generative adversarial network, the style authenticity of the text, images and video scripts of the initial advertising creative plan is judged, and the style discrimination gradient is obtained.

[0079] In this embodiment, the discriminator includes a text discriminator ( Image discriminator ) and video script discriminator ( The text discriminator is based on a finely tuned RoBERTa binary classification network. The input is a text sequence, and the output is the probability of "matching the target channel's style." The image discriminator... A convolutional network based on the PatchGAN architecture takes an image as input and outputs a probability map showing that local regions conform to the target color distribution. (Video script discriminator) Based on Temporal CNN or Transformer networks, the input is a sequence of script scenes or a sequence of video frames, and the output is a score of the narrative rhythm conformity.

[0080] Specifically, the copy of the initial advertising creative plan ,image and video script Input the corresponding discriminator respectively .

[0081] The steps involved in authenticity scoring include: Output probability This indicates the degree to which the copy conforms to the target channel's language style. Output probability graph This indicates the degree to which each region of the image conforms to the target color style; the average value is taken to obtain the result. . Output probability This indicates the degree to which the script's rhythm conforms to the target narrative pattern. Define the discriminant loss. The system does not update the discriminator parameters, but instead calculates... The gradient relative to the generator's latent space encoding Z:

[0082] It indicates how to modify the underlying encoding Z to make the generated content appear more like the target channel style in the eyes of the discriminator.

[0083] S33: Based on style discrimination gradient and brand tone constraint parameters, the latent space encoding of the generator is adjusted through gradient backpropagation algorithm to minimize the style transfer loss function while keeping the original semantic content unchanged.

[0084] In this embodiment, the brand tone constraint parameter defines the boundaries that a brand cannot cross, such as prohibiting the use of vulgar words, requiring the main color to include the brand blue, and requiring the narrative tone to be positive and uplifting. The brand tone constraint parameter is transformed into a hard constraint mask or a high-weight penalty item.

[0085] Specifically, let the initial latent encoding be... An initial scheme is obtained through the generator. The style loss gradient calculated in step S31 is compared with the discriminant gradient calculated in step S32. By fusion, a comprehensive driving gradient is obtained. :

[0086] in A balancing coefficient is used for style discrimination gradients; at the same time, semantic preservation constraints are introduced for the gradients. The CLIP model is used to calculate the similarity between G(Z) and the original product selling point text. If the similarity decreases, a back gradient is generated to force Z back to the semantically consistent region. Finally, the direction vector is updated. . Used to control the strength of semantic constraints, with values ​​such as 0.5, 1.0, and 1.5.

[0087] Update the latent encoding using an optimizer (such as Adam):

[0088] This is the updated potential encoding. During this process, the brand tone constraint parameter acts as a hard boundary: if the updated... The generated content triggers brand red lines in layout or color schemes. Instead of directly truncating the gradient of a specific dimension of the potential space Z (due to the difficulty in decoupling entangled features), the system constructs a brand compliance penalty loss term. .

[0089] Specifically, if the generated intermediate solution violates the brand's red lines (such as color deviation from the main color scheme or inclusion of prohibited words), the system calculates the distance between the solution's features and the brand's standard features. The greater the distance, the more likely it is to violate the brand's red lines. The larger the value.

[0090] The comprehensive driving gradient update formula is revised as follows:

[0091] in, This results in a high penalty coefficient. The optimizer minimizes the number of... The total loss naturally guides the potential coding Z to move towards areas that align with the brand's tone, without requiring manual specification of specific dimensions.

[0092] S34: Iterate through the above transformation process until the style discrimination gradient converges, and obtain the candidate creative solutions after style adaptation.

[0093] In this embodiment, steps S31 to S33 are repeated. In each iteration t: an intermediate solution is generated. Calculate the current and discrimination probability .renew .

[0094] Specifically, the norm of the style discrimination gradient is monitored. .like For example, if ϵ=1e−4, it means the discriminator can no longer distinguish the current scheme from the true target style sample, and style transfer is complete. Alternatively, monitor the total loss. The rate of change is considered convergent if the decrease is less than a threshold for K consecutive iterations. A maximum number of iterations is set. (e.g., 50 times) to prevent infinite loops.

[0095] When the convergence condition is met or the convergence is achieved Stop iterating when the time is right. Take the final latent encoding. The final style-adapted candidate creative solutions are obtained by decoding the generator G.

[0096] S4: Verify candidate creative solutions based on compliance, logical consistency, and effect prediction metrics. If the verification fails, a feedback signal is generated and sent back to the multimodal fusion deep learning model to perform secondary generation until the target advertising creative solution is determined.

[0097] In this embodiment, compliance verification refers to scanning the advertising content using a pre-trained classification or detection model to identify whether there are elements that violate laws, regulations, or platform prohibitions. Logical consistency verification refers to determining whether there are semantic conflicts or logical breaks between text, images, and video scripts by calculating the similarity or correlation scores between feature vectors of different modalities. Performance prediction metrics refer to the estimated values ​​of future advertising performance (such as click-through rate and conversion rate) output by an evaluation model trained on historical big data. Feedback signals are quantitative information generated when verification fails, containing the error type and correction direction, used to guide the model in iterative optimization. Secondary generation refers to the behavior of the model adjusting its internal parameters or generation strategy after receiving feedback signals, and re-executing the generation process to produce a better solution.

[0098] Specifically, step S4 includes: S41: Use a pre-trained violation content recognition model to scan the text and image content of candidate creative solutions, identify and mark abnormal areas containing false advertising, sensitive words or violation visual elements, and eliminate solutions with compliance risks.

[0099] In this embodiment, the violation content identification model includes a text sub-model and an image sub-model. The text sub-model is based on a binary or multi-label classification network finely tuned to the BERT-base-chinese architecture. The input is a sequence of advertising copy tokens, and the output is the violation probability and violation type label for each token. Violation type labels include absolute terms, false promises, and sensitive words. The image sub-model is based on an object detection and classification network based on the ResNet-50 or YOLOv8 architecture. The input is the RGB tensor of the advertising image, and the output is the bounding box coordinates and category of the violation elements in the image, such as prohibited signs, inappropriate images, missing QR codes, etc. A preset violation probability threshold is used. For example, a value of 0.9 is considered a violation if it exceeds 0.9.

[0100] S42: Calculate the cosine similarity between the key semantic vector of the advertising copy and the deep visual feature vector of the advertising image, and quantify the correlation score between the plot logic structure of the video script and the core selling points of the product; if the cosine similarity or correlation score is lower than the preset logical consistency threshold, it is judged as a multimodal logical mismatch, and a feedback label indicating the mismatch type is generated.

[0101] In this embodiment, cosine similarity is calculated by using a pre-trained multimodal alignment model (such as CLIP-ViT-L / 14) to extract the text embedding vectors of the text. and the visual embedding vector of the image .

[0102] Specifically, the text encoder of the CLIP model is invoked for processing. This yields the normalized text feature vector. Call CLIP's image encoder for processing. This yields the normalized image feature vector. Calculate cosine similarity: The cosine similarity value ranges from [-1, 1], with values ​​closer to 1 indicating greater semantic consistency between the image and text.

[0103] Building a product selling point knowledge graph based on graph neural networks (GNN) The video script is parsed into a sequence of plot nodes and a product selling point knowledge graph. The nodes include selling points and features, while edges represent logical relationships. Natural language processing tools are used to process the video script. Analysis into a set of key plot points .

[0104] Set a logical consistency threshold For example, 0.75. If The result was determined to be a "mismatch between text and image logic," and category tags were generated. And record the difference values. .like The system was flagged as having a "mismatch between script and selling point logic" and category tags were generated. And record the difference values. If both are below the threshold, a composite label is generated. If both are above the threshold, the solution passes the logical consistency dimension, and the process proceeds to step S43.

[0105] S43: The effect prediction model trained based on historical delivery data calculates the estimated conversion index of candidate creative solutions. If the estimated conversion index is lower than the preset effect baseline, it is judged as an inefficient solution.

[0106] In this embodiment, a preset effect baseline is read. It can be dynamically updated, for example, taking the average pCVR of the top 20% of similar creatives over the past 7 days. Specifically, deep features of candidate solutions are extracted, including: BERT embedding, ResNet features The LSTM hidden state is used. Contextual features of the current delivery scenario are concatenated. The assembled feature vector is then input into the performance prediction model. The performance prediction model outputs the estimated conversion metric. Read the preset effect baseline .like <, the proposed solution is determined to be an "inefficient solution", and the performance gap value is calculated. .like If the solution passes all the checks, it is marked as a "high-quality candidate solution" and can be directly entered into the recommendation queue.

[0107] S44: For solutions judged as having multimodal logic mismatch or inefficiency, the corresponding feedback label or effect gap value is converted into a negative reward signal in the reinforcement learning algorithm, which is used to update the loss function of the multimodal fusion deep learning model, update the model parameters, and recalculate new candidate creative solutions.

[0108] In this embodiment, a reward function R is defined. For a scheme that passes the verification, R = +1.0.

[0109] For unqualified solutions, construct a negative reward signal: If rejected due to logical mismatch: Where α is the penalty coefficient, such as 2.0. If rejected due to inefficiency: , where β is the penalty coefficient, such as 1.5. If multiple problems exist simultaneously, the reward value is the sum of the negative rewards for each problem. Specific feedback labels (such as "Text-Image Mismatch") are converted into conditional vector negative reward signals.

[0110] Specifically, the generation process of the unqualified solution is regarded as a trajectory τ, which includes the state (input features), action (generated token or pixel noise prediction) and final reward R.

[0111] The parameters of the multimodal fusion deep learning model are updated using the policy gradient method. The updated loss function is defined as follows: .

[0112] in This is the original supervised learning loss (such as cross-entropy or diffusion loss). Let be the probability distribution of the generative model. 'a' represents the generated action, and 's' represents the state. This term is used to reinforce the learning rate weights. Since R is negative, this term effectively increases the loss value for generating the incorrect sample, forcing the model parameters to... Adjust along the gradient descent direction to reduce the probability of generating similar errors (such as text-image mismatch or inefficient features) in the future.

[0113] In practice, for solutions deemed unqualified, the system can convert their error type (such as 'image-text mismatch') into a negative conditional vector. This negative vector is then input into the model during subsequent retry generation in the current session to suppress the probability of generating similar errors. Alternatively, mini-batch offline updates can be used, adding recently rejected samples and their corresponding negative reward signals to the training batch and performing a backpropagation to fine-tune the model weights of the multimodal fusion deep learning model.

[0114] Utilizing the updated multimodal fusion deep learning model parameters Keeping the original input conditions (product selling points, user preferences, etc.) unchanged, re-execute the generation process in S21-S24. Since the multimodal fusion deep learning model has received negative feedback about the reasons for the previous failure (reflected through loss function updates), the newly generated candidate solutions will show significant improvements in semantic alignment, logical coherence, or predicted performance metrics. The newly generated solutions are then fed back into S41-S43 for validation.

[0115] S5: Based on the target audience profile, the characteristics of the delivery channels, and the preset conversion goals, use intelligent recommendation algorithms to prioritize and distribute target ad creative solutions; and dynamically optimize the recommendation strategy by combining real-time delivery feedback data.

[0116] In this embodiment, the intelligent recommendation algorithm is a multi-objective ranking model used to calculate the distribution priority score of each creative scheme based on the target audience profile and conversion goals, under multiple constraints. The target audience profile refers to a multi-dimensional description of the target audience for advertising, including demographic characteristics, interests, and a set of spending power tags. Distribution channel characteristics refer to the attribute information of the specific platform on which the current advertisement will be placed, including user activity time periods, screen size preferences, and content consumption habits. Preset conversion goals refer to the specific marketing objectives set by the advertiser, such as maximizing clicks, maximizing conversions, or increasing brand exposure. Priority ranking refers to arranging multiple target advertising creative schemes from highest to lowest score based on a comprehensive calculation of multiple factors to determine the distribution order. Real-time delivery feedback data refers to user behavior data generated immediately after the advertisement goes live, including impressions, click-through rate, dwell time, and conversion behavior. Dynamic optimization of the recommendation strategy refers to the process of automatically adjusting the weight parameters or ranking logic in the recommendation algorithm based on real-time feedback data to adapt to market changes and user preferences.

[0117] In one embodiment, such as Figure 2 As shown, a method for automatic generation and recommendation of advertising creatives based on deep learning also includes: S101: Real-time capture of click-through rate, conversion rate, and user dwell time feedback from the ad platform, and map them into performance reward signals.

[0118] In this embodiment, the creative solution ID is received in real time via SDK tracking or API interface. The data includes: exposure timestamp, click events, conversion events (such as placing an order or registering), and user dwell time. The number of clicks is obtained. Conversion count Total number of exposures and average user dwell time .

[0119] ,in, For the preset weighting coefficients, such as , This is the benchmark value for the maximum average dwell time in the channel's historical statistics.

[0120] The storage structure combines a key-value database with a vector database. The data includes user interest and preference feature vectors. Product core semantic feature vector And associated mapping tables. Among them... These are implicit intention tags, such as pursuing cost-effectiveness or preferring a minimalist style; For feature embedding vectors, To activate the priority weight, the initial value is 1.0. For selling points (such as "ultra-long battery life" or "waterproof"), For semantic vectors, These are semantic density weights. The association mapping table is used to record implicit intentions. With product selling points Correlation strength matrix between .

[0121] Specifically, the system listens to the message queue and consumes user behavior logs from the ad serving engine in real time. For each generated creative solution... Aggregate its behavioral data within a preset time window (such as the most recent hour or the cumulative first 1000 exposures): count the number of clicks. and conversion count Calculate the average user dwell time. .

[0122] Specifically, if If it is determined to be strong positive feedback; but =0, indicating weak positive feedback; if If the data is significantly below the industry average (e.g., below the 20th percentile), it is considered a negative feedback.

[0123] The effect reward signal is obtained by performing normalization calculation. : .

[0124] S102: Update the implicit intent weights in the user interest preference feature vector using the effect reward signal: If a creative solution corresponding to a certain type of implicit intent feature receives a positive reward, its activation priority in the feature library will be increased, and the semantic density weight of the corresponding selling point in the associated product core semantic feature vector will be enhanced in reverse.

[0125] In this embodiment, a lightweight online update mechanism is used, updating only the metadata weights in the feature library without involving real-time backpropagation of the deep learning model's backbone parameters. Feature enhancement is performed on creative solutions marked as positive rewards: based on... Tracing back to the set of implicit intentions in the user interest and preference feature vectors used when generating the solution For example, the solution generation primarily activated two implicit intents: "outdoor adventure" and "durability." This involves tracing back to the set of selling points within the corresponding core semantic feature vector of the product. .

[0126] Traversal Each implicit intention in Update its activation priority weight in the feature library. :

[0127] in, The learning rate is set to the intention, such as 0.05. This ensures the weights gradually approach 1 without overflowing; higher rewards lead to faster growth. In subsequent generation tasks, the sampling module will... Weighted random sampling is performed based on the size. The higher the implicit intent, the greater the probability of being selected.

[0128] Query the related mapping table Find with Strongly related product selling points Update selling points semantic density weights (Semantic density weights are used to control the scaling factor of Cross-Attention in the generative model): .in, This is the selling point enhancement factor, such as 0.1. The increase in this means that when generating ad copy or images next time, the model will assign higher attention weights to the semantic embedding vectors of selling points such as "durable" and "waterproof", thus making these selling points more prominent in the generated content.

[0129] S103: Simultaneously, monitor external hot event flows and calculate the semantic similarity between hot keywords and existing product core semantic features in the feature library; if the similarity exceeds the dynamic injection threshold, generate a temporary enhanced feature vector based on the hot keywords and perform weighted fusion with the user interest preference feature vector to obtain a timeliness-enhanced feature vector.

[0130] In this embodiment, the system calls an external API at fixed intervals (e.g., every 10 minutes) to retrieve the current list of popular keywords. And its popularity index Pop(h). Each hot keyword is encoded using a pre-trained language model. hotspot vector Then, iterate through the existing product core semantic feature vectors in the feature library. All selling point vectors Calculate the cosine similarity between the hotspot vector and the selling point vector:

[0132] Set dynamic injection threshold The dynamic injection threshold can be dynamically adjusted over time, for example... (Based on the mean of historical similarity distribution) and standard deviation (or automatically lower the matching criteria during major holidays).

[0133] If it exists Then determine the hotspot With product selling points Highly correlated. The system initiates a gated fusion network to generate temporary enhanced feature vectors. :

[0134] in The hotspot fusion coefficient can be determined based on... Dynamic adjustment: The higher the popularity, The larger, Then With the current user interest and preference feature vector To merge. If The existence of and Semantically similar intent vectors Then it directly enhances The vector representation; if not, then... It will be added to the fusion pool as a new temporary intent vector, soon. As a new dimension, it is spliced ​​together and then processed through a learnable fusion weight matrix. Mapping back to the original feature dimension D yields the timeliness-enhanced feature vector. : , The weight matrix is ​​fused and has a dimension shape of [2D, D].

[0135] Furthermore, in another embodiment, an external hotspot event stream parsing model is run to monitor external hotspot event streams in parallel. This model architecture is based on the BERT-base-chinese pre-trained model, with a fully connected classification layer and a sequence labeling (CRF) layer added to its top layer. The model's input layer receives real-time crawled internet news headlines and social media trending search list text sequences. After processing through 12 Transformer blocks in the BERT encoder, the context embedding vector for each token is output. The sequence labeling layer uses these vectors to identify hot keyword entities. The fully connected layer then outputs the sentiment polarity score for each keyword. and spread popularity score .

[0136] Correspondingly, the system calls the semantic similarity calculation engine to extract the hot keywords. Convert to vector (Take the [CLS] label vector from the last layer of the BERT model), and calculate its relationship with the existing product core semantic feature vectors in the feature library. cosine similarity If the calculated maximum similarity Simmax exceeds the dynamic injection threshold θ (θ=0.6 for fashion products, θ=0.8 for durable consumer goods), the system determines that the hot topic has marketing value. At this time, the system activates the timeliness enhancement feature generator, which is a gated fusion network whose input is the hot topic feature vector. and user interest and preference feature vector The generator internally uses a Sigmoid gate function. Calculate the fusion coefficient and output the timeliness enhancement feature. The symbol ⊙ represents element-wise multiplication. For example, when "Olympics" becomes a hot topic and has a high semantic similarity to "sports equipment" products, the system generates a vtemp containing semantic information of "competition" and "glory," which can guide the generation of creative ideas in real time. This allows the generated advertising creatives to incorporate the Olympic theme into their narratives in an instant.

[0137] S104: Write timeliness enhancement features into the real-time cache of the feature library and set a lifespan. After the lifespan expires, the features will be automatically removed or downgraded to achieve adaptive evolution of the feature library to market dynamics.

[0138] In this embodiment, the generated timeliness enhancement features are written into the real-time cache area of ​​the feature library. This cache area is independent of the main feature library and is specifically used to store frequently updated dynamic features. The system sets a lifecycle for each written timeliness enhancement feature. The length of the lifecycle is determined based on the expected duration and decay curve of the hot event (e.g., 24 hours for breaking news and 168 hours for seasonal holidays).

[0139] For example, the system will generate timeliness enhancement features. The feature is written to a real-time cache (Redis cluster), which is independent of the feature database (MySQL vector index) and is specifically used to store frequently updated dynamic features. The system calls the lifecycle manager, which manages each written feature. Allocate a lifecycle . The calculation formula is ,in The base duration is set to 24 hours for breaking news and 168 hours for seasonal holidays. This is the attenuation factor for hotspot characteristics, such as setting it to 0.8.

[0140] The manager starts a background daemon that scans the cache every minute for the current time. Exceeding creation time For features that are close to expiring, perform automatic removal; for features that are close to expiring... The weights are adjusted according to the exponential decay function. A weighting process is applied, where γ is the decay rate constant. Within the feature's lifecycle, features not nearing expiration have the highest priority during model inference, forcing the multimodal fusion model to prioritize retrieving these features when generating creative content. Once the lifecycle expires, the feature no longer dominates the creative generation direction, thus enabling the feature library to adaptively evolve in response to market dynamics. This ensures that advertising creatives can both capitalize on current events and avoid generating outdated content after the trend has faded, maintaining the timeliness of marketing content.

[0141] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0142] In one embodiment, a deep learning-based automatic advertising creative generation and recommendation system is provided, which corresponds to the deep learning-based automatic advertising creative generation and recommendation method in the above embodiments.

[0143] This deep learning-based automatic advertising creative generation and recommendation system includes a feature library construction module, an initial creative generation module, a style adaptation and transformation module, a verification and feedback optimization module, and an intelligent recommendation and distribution module. Detailed descriptions of each functional module are as follows: The feature library construction module is used to extract core semantic features of products, user interest and preference features, and channel style constraint features based on multi-source heterogeneous data, and to build a feature library. The initial creative generation module is used to obtain target feature vectors from the feature library based on the delivery requirements, input the target feature vectors into a pre-trained multimodal fusion deep learning model to obtain an advertising copy sequence, and collaboratively generate semantically aligned advertising images and video script sequences based on the advertising copy sequence and the target feature vectors, thereby determining the initial advertising creative scheme. The style adaptation and transformation module is used to perform feature space transformation on the copywriting style, image color distribution and video script narrative rhythm of the initial advertising creative scheme based on preset channel style mapping rules and brand tone constraint parameters, so as to obtain the candidate creative scheme after style adaptation. The verification and feedback optimization module is used to verify candidate creative solutions based on compliance, logical consistency, and effect prediction metrics. If the verification fails, a feedback signal is generated and sent back to the multimodal fusion deep learning model to perform secondary generation until the target advertising creative solution is determined. The intelligent recommendation and distribution module is used to prioritize and distribute target ad creatives based on target audience profiles, channel characteristics, and preset conversion goals using intelligent recommendation algorithms; and dynamically optimize recommendation strategies by combining real-time delivery feedback data.

[0144] For specific limitations regarding the deep learning-based automatic advertising creative generation and recommendation system, please refer to the limitations of the deep learning-based automatic advertising creative generation and recommendation method mentioned above, which will not be repeated here. Each module in the aforementioned deep learning-based automatic advertising creative generation and recommendation system can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0145] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the deep learning-based automatic generation and recommendation method for advertising creatives.

[0146] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0147] In one embodiment, particularly according to embodiments of the present invention, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, embodiments of the present invention include a computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of a deep learning-based method for automatic generation and recommendation of advertising creatives. In such embodiments, the computer program can be downloaded and installed from a network via a communication module, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs the various functions defined in the present invention.

[0148] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0149] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for automatic generation and recommendation of advertising creatives based on deep learning, characterized in that, include: Based on multi-source heterogeneous data, core semantic features of products, user interest and preference features, and channel style constraint features are extracted to construct a feature library; Based on the delivery demand instruction, the target feature vector is obtained from the feature library and input into the pre-trained multimodal fusion deep learning model to obtain the advertising copy sequence; based on the advertising copy sequence and the target feature vector, a semantically aligned advertising image and video script sequence is generated to determine the initial advertising creative scheme; Based on preset channel style mapping rules and brand tone constraint parameters, feature space transformation is performed on the copywriting style, image color distribution and video script narrative rhythm of the initial advertising creative scheme to obtain a style-adapted candidate creative scheme. The candidate creative solutions are verified based on compliance, logical consistency and effect prediction indicators. If the verification fails, a feedback signal is generated and sent back to the multimodal fusion deep learning model to perform secondary generation until the target advertising creative solution is determined. Based on the target audience profile, the characteristics of the delivery channels, and the preset conversion goals, the target advertising creative schemes are prioritized and distributed using intelligent recommendation algorithms; and the recommendation strategy is dynamically optimized by combining real-time delivery feedback data.

2. The method for automatic generation and recommendation of advertising creatives based on deep learning according to claim 1, characterized in that, The step of collaboratively generating semantically aligned advertising image and video script sequences based on the advertising copy sequence and the target feature vector includes: Based on the Transformer architecture, an advertising copy token sequence is generated using the target feature vector. Based on the improved diffusion probability model, the semantic embedding of the advertising copy token sequence is fused with the target feature vector across modalities to generate an advertising image that satisfies the constraints in both semantic content and visual style. The sequence generation model based on the encoder-decoder architecture uses the advertising copy token sequence, product selling point features and user preference features to autoregressively generate video script sequences. If the confidence level of any generated result is lower than the preset confidence threshold, then based on other generated deterministic results as constraints, the generation space of the model is resampled, and a globally consistent initial advertising creative scheme is calculated and determined.

3. The method for automatic generation and recommendation of advertising creatives based on deep learning according to claim 1, characterized in that, The verification of the candidate creative solutions based on compliance, logical consistency, and effect prediction indicators includes: The pre-trained violation content recognition model is used to scan the text and image content of candidate creative solutions, identify and mark abnormal areas containing false advertising, sensitive words or violation visual elements, and eliminate solutions with compliance risks. The cosine similarity between the key semantic vector of the advertising copy and the deep visual feature vector of the advertising image is calculated, and the correlation score between the plot logic structure of the video script and the core selling points of the product is quantified. If the cosine similarity or correlation score is lower than the preset logical consistency threshold, it is determined to be a multimodal logical mismatch, and a feedback label indicating the mismatch type is generated. The effect prediction model trained based on historical delivery data calculates the estimated conversion index of candidate creative solutions. If the estimated conversion index is lower than the preset effect baseline, it is determined to be an inefficient solution. For schemes that are judged to be multimodal logic mismatched or inefficient, the corresponding feedback label or effect gap value is converted into a negative reward signal in the reinforcement learning algorithm, which is used to update the loss function of the multimodal fusion deep learning model, update the model parameters, and recalculate new candidate creative schemes.

4. The method for automatic generation and recommendation of advertising creatives based on deep learning according to claim 1, characterized in that, The feature library is constructed based on extracting core semantic features of products, user interest and preference features, and channel style constraint features from multi-source heterogeneous data, including: Collect structured attribute data and unstructured graphic descriptions of products, use a pre-trained language model to extract the core selling points of the products, calculate the co-occurrence frequency of each entity in historical high-conversion samples, and generate a core semantic feature vector of the products with semantic density weights. A user-item interaction graph is constructed based on the user's historical behavior sequence. The graph attention network is used to aggregate neighbor node information, extract the user's explicit preference features, and use the product's core semantic feature vector as the query key to retrieve the implicit intent features that are semantically aligned with it in the user behavior space. The features are then fused to generate a user interest preference feature vector, so as to realize the semantic space mapping between user preferences and product selling points. Analyze historical top creative samples from different advertising channels, extract visual color histograms, copywriting emotional polarity distribution, and video rhythm and beat features, cluster to generate channel style prototype vectors, and calculate the style compatibility matrix between the channel style prototype vectors and the product core semantic feature vectors. Select style constraint parameters with compatibility higher than a preset threshold as channel style constraint features. The semantically aligned user interest preference feature vectors and the channel style constraint features that have been filtered for compatibility are associated and mapped to the same feature space index to construct a feature library that includes static attributes and dynamic relationships.

5. The method for automatic generation and recommendation of advertising creatives based on deep learning according to claim 4, characterized in that, Also includes: Real-time capture of click-through rate, conversion rate, and user dwell time feedback from the ad delivery platform, and mapping them into performance reward signals; The implicit intent weights in the user interest preference feature vector are updated using the effect reward signal: if a creative solution corresponding to a certain type of implicit intent feature receives a positive reward, its activation priority in the feature library is increased, and the semantic density weights of the corresponding selling points in the associated product core semantic feature vector are enhanced in reverse. Simultaneously, monitor external hot event streams and calculate the semantic similarity between hot keywords and existing product core semantic features in the feature library; if the similarity exceeds the dynamic injection threshold, generate a temporary enhanced feature vector based on the hot keywords and perform weighted fusion with the user interest preference feature vector to obtain a timeliness enhanced feature vector; The timeliness enhancement features are written into the real-time cache of the feature library and a lifespan is set. After the lifespan expires, the features are automatically removed or downgraded, so as to realize the adaptive evolution of the feature library to market dynamics.

6. The method for automatic generation and recommendation of advertising creatives based on deep learning according to claim 1, characterized in that, The method, based on preset channel style mapping rules and brand tone constraint parameters, performs feature space transformation on the copywriting style, image color distribution, and video script narrative rhythm of the initial advertising creative scheme, including: Construct a style transfer loss function, which includes a stylistic difference term, a color distribution distance term, and a narrative rhythm deviation term; By using the discriminator in the generative adversarial network, the style authenticity of the text, images and video scripts of the initial advertising creative scheme is judged, and the style discrimination gradient is obtained. Based on the style discrimination gradient and the brand tone constraint parameters, the latent space encoding of the generator is adjusted through the gradient backpropagation algorithm to minimize the style transfer loss function while keeping the original semantic content unchanged. The above transformation process is executed iteratively until the style discrimination gradient converges, resulting in a style-adapted candidate creative solution.

7. A deep learning-based automatic advertising creative generation and recommendation system, characterized in that, The system is used to perform a deep learning-based automatic generation and recommendation method for advertising creatives as described in any one of claims 1 to 6, the system comprising: The feature library construction module is used to extract core semantic features of products, user interest and preference features, and channel style constraint features based on multi-source heterogeneous data, and to build a feature library. The initial creative generation module is used to obtain target feature vectors from the feature library based on the delivery demand instructions, input the target feature vectors into a pre-trained multimodal fusion deep learning model to obtain an advertising copy sequence, and generate semantically aligned advertising images and video script sequences based on the advertising copy sequence and the target feature vectors, thereby determining the initial advertising creative scheme. The style adaptation and transformation module is used to perform feature space transformation on the copywriting style, image color distribution and video script narrative rhythm of the initial advertising creative scheme based on preset channel style mapping rules and brand tone constraint parameters, so as to obtain the candidate creative scheme after style adaptation. The verification and feedback optimization module is used to verify the candidate creative schemes based on compliance, logical consistency and effect prediction indicators; if the verification fails, a feedback signal is generated and sent back to the multimodal fusion deep learning model to perform secondary generation until the target advertising creative scheme is determined. The intelligent recommendation and distribution module is used to prioritize and distribute the target advertising creative schemes based on the target audience profile, the characteristics of the delivery channel, and the preset conversion target, and to dynamically optimize the recommendation strategy by combining real-time delivery feedback data.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the deep learning-based automatic generation and recommendation method for advertising creatives as described in any one of claims 1 to 6.

9. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the deep learning-based automatic generation and recommendation method for advertising creatives as described in any one of claims 1 to 6.