Marketing originality automatic generation and optimization system based on deep learning

By employing few-sample stylization and cross-modal semantic alignment technologies, combined with state-aware units and dynamic optimization modules, the problem of generating and optimizing marketing creatives during the cold start phase of emerging brands was solved. This achieved brand consistency in creative content and user state-driven optimization, thereby improving marketing effectiveness.

CN121638252APending Publication Date: 2026-03-10BEIJING HEJIN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511837143.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In the initial launch phase of emerging consumer brands, existing deep learning models struggle to generate differentiated and brand-aligned marketing creatives, and lack effective optimization methods. In particular, with limited historical data and a limited budget, it is difficult to achieve personalized control and optimization of creative styles.

Method used

By employing few-sample stylization units and cross-modal semantic alignment units, a style reference set is constructed using a small number of reference creative materials. Combined with state-aware units and dynamic optimization modules, style consistency of creative content and user state-driven dynamic optimization are achieved.

Benefits of technology

In the absence of historical data, quickly create a differentiated visual and discourse system, generate creative content that aligns with the target brand's tone, and optimize online through user interaction feedback to improve the brand fit and business performance of the creative content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121638252A_ABST
    Figure CN121638252A_ABST
Patent Text Reader

Abstract

The invention discloses a marketing originality automatic generation and optimization system based on deep learning, and relates to the technical field of computers. Through a small sample stylization unit, the system can construct a style reference set only by relying on a small number of representative originality materials; the generated content distribution and the reference distribution are aligned in a unified depth feature space, so that the generated image and copywriting are closer to the target brand tonality in the dimensions of color, composition, tone and the like, and a differentiated vision and utterance system is quickly molded in the absence of large-scale brand data; in combination with a cross-modal semantic alignment unit, text information such as marketing targets, audience descriptions and the like and brand visual instructions are jointly coded into a unified semantic-style condition, so that the double constraints of what and what can be satisfied in the same generation process, and the risks of disjunction between copywriting and pictures and style deviation are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer technology, and particularly relates to a marketing creative automatic generation and optimization system based on deep learning. BACKGROUND

[0002] With the rapid growth of new consumer brands, such brands mostly face problems such as limited budget, time pressure and lack of historical advertising data accumulation in the cold start stage; in the early stage of brand construction, the brand visual system and discourse system are still in the process of exploration and construction, and it is necessary to improve user cognition and occupy the submarket through differentiated marketing creativity in a short time; at the same time, there are a large amount of marketing information with similar content forms in network media and social platforms, and the user's attention resources are limited, so the differentiation and novelty of creative content are highly required.

[0003] In order to improve the production efficiency of marketing content, various marketing creative automatic generation schemes are proposed in the industry; the existing schemes mostly rely on deep generation models pre-trained on large-scale public data sets, such as image generation models based on diffusion models and large language models, etc., which automatically generate advertising images and script contents by inputting product descriptions, target group labels and other structured or unstructured information; in order to make the generation results more consistent with the business objectives of advertising, some front-line technologies further introduce reinforcement learning, reward models or other optimization mechanisms based on feedback signals, and use the click-through rate, conversion rate and other performance data of historical advertising to fine-tune the generation model, so that the output is more consistent with the existing advertising strategy and effect preference.

[0004] However, the above technical schemes still have certain limitations in the cold start scene of new consumer brands; on the one hand, the pre-trained generation model mainly learns the Internet data distribution covering a wide range of fields, and the generation result often presents a relatively mainstream or averaged style, lacking the ability of personalized control for specific brand tone and sub-group appeal, and it is difficult to reflect the differentiated image that the new brand wants to shape; on the other hand, the scheme relying on historical advertising click-through rate, conversion and other performance data for optimization needs to accumulate enough feedback samples under a long time scale and a large scale of advertising, so as to train a stable reward model or realize effective strategy update; for new brands that have not formed stable advertising data, it is difficult to directly apply such optimization methods based on large-scale historical feedback in the cold start stage.

[0005] Therefore, in the cold start scene lacking sufficient historical advertising data, how to generate marketing creativity with high differentiation and brand fit using deep learning model under the conditions of limited budget and time, and how to effectively optimize the generation result based on a small amount of interactive feedback, are still technical problems to be solved. SUMMARY

[0006] In view of the aforementioned existing problems, the present invention is proposed.

[0007] This invention provides a deep learning-based system for automatically generating and optimizing marketing creative ideas, which solves the problem that existing pre-trained and feedback-optimized creative systems are insufficient to support the cold start needs of emerging brands in the absence of historical data.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0009] This invention provides a deep learning-based system for automatically generating and optimizing marketing creative ideas, comprising:

[0010] The creative generation module is used to receive input instructions containing marketing objectives and generate marketing creative content;

[0011] And a dynamic optimization module, which is communicatively connected to the creative generation module;

[0012] The creative generation module includes a small-sample stylization unit, which is configured to receive a style reference set consisting of a small number of reference creative materials, and optimize the parameters of the creative generation model based on the style reference set, so that the creative content output by the creative generation module is closer to the style reference set in terms of style representation, thereby improving the consistency between the generated creative content and the target brand style.

[0013] The dynamic optimization module includes a state-aware unit, which is configured to infer the user's state based on the user interaction sequence, generate an optimization strategy based on the user's state, and output it to the creative generation module to adjust the subsequently generated creative content.

[0014] As a preferred embodiment of the deep learning-based marketing creative automatic generation and optimization system described in this invention, the creative generation module further includes a cross-modal semantic alignment unit. The cross-modal semantic alignment unit is configured to jointly encode text description instructions and visual style instructions to generate a unified semantic-style condition vector, and control the creative generation module to ensure that the output content simultaneously satisfies the semantic constraints of the text description instructions and the style constraints of the visual style instructions when generating creative content.

[0015] As a preferred embodiment of the deep learning-based marketing creative automatic generation and optimization system described in this invention, the few-sample stylization unit includes a style encoder and a style injection network;

[0016] The style encoder is used to extract deep style features from the creative materials in the style reference set;

[0017] The style injection network is integrated into the creative generation module during the generation process. It is used to receive the deep style features and inject the deep style features as conditional information into the creative content generation process through a cross-attention mechanism.

[0018] The few-sample stylization unit is also configured to update the parameters of the creative generation module based on a loss function that measures the difference between the deep feature distribution of the generated creative content and the deep feature distribution of the style reference set.

[0019] As a preferred embodiment of the marketing creative automatic generation and optimization system based on deep learning described in this invention, the state perception unit includes a deep sequence coding model, which encodes the behavior type, content features and time information in the user interaction sequence and outputs a user state vector representing the user's current cognitive stage.

[0020] As a preferred embodiment of the deep learning-based marketing creative automatic generation and optimization system described in this invention, the dynamic optimization module further includes a narrative planning unit, which takes the user state vector as input and predicts and outputs a narrative outline containing multiple ordered stage themes through a sequence generation model.

[0021] After receiving the narrative outline, the creative generation module generates a series of creative content that is thematically coherent and progressive, based on the stage theme corresponding to the content to be generated.

[0022] As a preferred embodiment of the deep learning-based marketing creative automatic generation and optimization system described in this invention, the dynamic optimization module further includes a deep context decision unit, which takes the user state vector and candidate creative feature vector as input and estimates the expected return value of each candidate creative through a deep reward prediction network.

[0023] As a preferred embodiment of the deep learning-based marketing creative automatic generation and optimization system described in this invention, the deep context decision-making unit selects creative strategies based on a combination of the expected return value output by the deep return prediction network and the intrinsic exploration incentive value based on the model's prediction uncertainty. Specifically, for combinations of user states and creative types with high prediction uncertainty, the system increases the probability that the corresponding candidate creative will be selected for deployment.

[0024] As a preferred embodiment of the deep learning-based marketing creative automatic generation and optimization system described in this invention, the dynamic optimization module is configured to update its internal model parameters through online learning based on user interaction feedback obtained after the creative content is launched.

[0025] The updated model parameters are used to adjust the user state inference logic of the state-aware unit and the creative strategy selection logic of the deep context decision-making unit, so as to influence the subsequent output of the creative generation module through the optimization strategy.

[0026] As a preferred embodiment of the deep learning-based marketing creative automatic generation and optimization system described in this invention, the user interaction sequence includes at least one of the following user behaviors:

[0027] Exposure, click, dwell time, conversion, closing, or skipping; the deep sequence coding model is configured to encode the user behavior into corresponding discrete and / or continuous features.

[0028] As a preferred embodiment of the deep learning-based marketing creative automatic generation and optimization system described in this invention, the input instructions include at least one or more of the following: target audience attributes, media channels for placement, and marketing target types. The creative generation module and / or the dynamic optimization module are configured to adjust the content elements and presentation format of the creative content based on the input instructions.

[0029] The beneficial effects of this invention are as follows: The marketing creative automatic generation and optimization system based on deep learning proposed in this invention addresses the problems of emerging consumer brands lacking historical campaign data, average creative styles, and difficulty in continuous optimization during the cold start phase. It comprehensively introduces small sample style alignment, cross-modal semantic control, and user state-driven dynamic decision-making mechanisms, realizing a complete technical closed loop from brand style migration to adaptive optimization of the campaign process.

[0030] This invention utilizes a small-sample stylization unit, allowing the system to construct a style reference set based on only a small number of representative creative materials. It aligns the generated content distribution with the reference distribution within a unified deep feature space, ensuring that the generated images and copy are closer to the target brand's tone in dimensions such as color, composition, and tone. This enables the rapid creation of a differentiated visual and discourse system even in the absence of large-scale brand data. Combined with a cross-modal semantic alignment unit, marketing objectives, audience descriptions, and other textual information are jointly encoded with brand visual instructions into unified semantic-style conditions. This allows for simultaneous fulfillment of both "what to say" and "what to look like" constraints within the same generation process, reducing the risk of disconnect between copy and visuals and style deviation. Furthermore, this invention employs a state-aware unit to perform deep sequence modeling of user interaction sequences. The system can identify the user's cognitive stage and output a multi-stage brand story outline based on a narrative planning unit. This ensures that continuously exposed creative content has a coherent and progressive thematic relationship, overcoming the shortcomings of previous fragmented creative approaches and the difficulty in accumulating long-term brand assets. Meanwhile, the deep contextual decision-making unit comprehensively considers the estimated return and the uncertainty of model prediction when selecting specific creative solutions. Through intrinsic exploration incentives, it adaptively adjusts the exploration intensity for different combinations of user states and creative types. In the cold start phase where data is scarce, it effectively expands the creative space and accelerates the discovery of high-potential solutions. As campaign feedback accumulates, the dynamic optimization module updates internal parameters through online learning, gradually improving the accuracy of user state inference and creative selection. This allows the system to actively explore in the early stages and converge to better strategies in the later stages, thereby significantly improving the style fit and business performance of emerging brand marketing creatives under limited budget and time constraints. Attached Figure Description

[0031] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation on the scope of this application.

[0032] Figure 1 This is a schematic diagram of the framework of the deep learning-based marketing creative automatic generation and optimization system in the embodiment. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0034] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0035] For example, the terms “first” and “second” used in this application are only used to distinguish and describe similar objects, to differentiate the first object from another object, and are not used to describe a specific order or sequence, nor should they be interpreted as indicating or implying relative importance.

[0036] This application proposes a deep learning-based system for automatically generating and optimizing marketing creative ideas, combining... Figure 1 As shown, the system includes:

[0037] The creative generation module is used to receive input instructions containing marketing objectives and generate marketing creative content;

[0038] And a dynamic optimization module, which communicates and connects with the creative generation module;

[0039] The creative generation module includes a small-sample stylization unit, which is configured to receive a style reference set consisting of a small number of reference creative materials and optimize the parameters of the creative generation model based on the style reference set, so that the creative content output by the creative generation module is closer to the style reference set in terms of style representation, thereby improving the consistency between the generated creative content and the target brand style.

[0040] The dynamic optimization module includes a state awareness unit, which is configured to infer the user's state based on the user interaction sequence, generate optimization strategies based on the user's state, and output them to the creative generation module to adjust the creative content generated subsequently.

[0041] In one embodiment, the creative generation module further includes a cross-modal semantic alignment unit, which is configured to jointly encode text description instructions and visual style instructions to generate a unified semantic-style condition vector, and control the creative generation module to ensure that the output content simultaneously satisfies the semantic constraints of the text description instructions and the style constraints of the visual style instructions when generating creative content.

[0042] In one embodiment, the few-sample stylization unit includes a style encoder and a style injection network;

[0043] The style encoder is used to extract deep style features from creative materials in a style reference set;

[0044] The style injection network is integrated into the creative generation module to receive deep style features and inject them as conditional information into the creative content generation process through a cross-attention mechanism.

[0045] The few-sample stylization unit is also configured to update the parameters of the creative generation module based on a loss function that measures the difference between the deep feature distribution of the generated creative content and the deep feature distribution of the style reference set.

[0046] The steps for calculating distributional difference loss include:

[0047] Step a, to fix the feature space in the existing style encoder, select the layer index. Extract vector features from both the generated samples and the reference samples:

[0048] ,

[0049] in, Indicates the first The depth feature vector of each generated sample, Indicates the first The depth feature vector of each reference sample, The style encoder is indicated by the layer subscript. Feature mapping, Indicates the first One generated sample, Indicates the first One reference sample, superscript Indicates the generating domain, superscript Indicates the reference domain. To generate a sample index, For reference sample index, For selected and fixed layer index constants;

[0050] Step b, to facilitate downstream parameterization, performs a Gaussian approximation with diagonal shrinkage on the features of the two domains:

[0051] ,

[0052] ,

[0053] in, To generate the mean vector of domain features, To generate the domain feature covariance matrix, The reference domain feature mean vector, For the reference domain characteristic covariance matrix, To generate the number of samples in the domain batch, For the reference domain batch sample size, The diagonal contraction coefficient, for identity matrix The feature dimension is constant;

[0054] Step c, calculate the distribution difference and form the style alignment loss, including:

[0055] c1, maximum mean difference:

[0056] ,

[0057] ,

[0058] in, This is the maximum average difference measure. To and Different generated sample indexes, To and Different reference sample indexes, For Gaussian kernel function, For kernel bandwidth scalar, For the L2 distance, 2 in The square of the measurement value is represented in the middle.

[0059] ,

[0060] in, For multi-bandwidth kernel functions, The number of bandwidth components. For the first Individual core component weights, For the first One core bandwidth, For kernel component index;

[0061] c2, Kullback–Leibler divergence (Gaussian approximation):

[0062] ,

[0063] in, The KL divergence from the approximate Gaussian of the generating domain to the approximate Gaussian of the reference domain. For trace operation, for The inverse matrix, For determinant, It is the natural logarithm;

[0064] c3, second-order Wasserstein distance (Gaussian approximation):

[0065] ,

[0066] in, This is the squared quantity of the second-order Wasserstein distance. The L2 norm of the difference between the means for The square root of the principal value matrix, The square root operator for matrices;

[0067] Step d: Without changing the original training framework, weight multiple metrics into a style alignment loss and perform a reverse update on the generator parameters:

[0068] ,

[0069] ,

[0070] in, To achieve style alignment, a scalar loss is applied. For MMD weights, For KL weights, For Wasserstein weights, This is the set of trainable parameters for the creative generation module. The learning rate scalar For about The gradient;

[0071] ,

[0072] in, For the first Normalized weights of the metric For the corresponding trainable logarithmic weights, For weighted summation index, Update the step size and index for weights. For the set of values ​​indexed by the measure ;

[0073] Specifically, the above implementation revolves around aligning the distribution of the style reference set and the generated content in the same deep feature space. First, a semantic-style vector representation is obtained at a fixed layer in the style encoder. Then, two strategies are used to measure the difference: one is based on the kernel's maximum average difference, which does not depend on distribution assumptions and is suitable for small-sample style transfer; the other is to model the feature distribution using Gaussian approximation, giving KL and second-order Wasserstein distances respectively. The latter is consistent with commonly used distribution matching metrics, is computationally stable, and can reflect the coupling difference between mean and covariance. To improve training controllability, weight coefficients are introduced to form the total style alignment loss, and the generation parameters are updated using the standard gradient method. Covariance contraction suppresses numerical ill-posedness in small samples, and the layer and kernel bandwidth can be implemented by tuning parameters through the development set.

[0074] Specifically, the style reference set consists of a small number of representative creative materials provided by the brand. These materials can be existing offline materials, design drafts, or historically well-performing online materials. Before being input into the style encoder, they undergo uniform size scaling, color space normalization, and cropping to ensure that the depth features are within a consistent numerical range. In this embodiment, the number of samples in each style reference set can be set to a range of 5-100, with a default value of 20. When the number of materials exceeds the upper limit, samples that are newer and more representative of the target style can be selected based on manual scoring or time weighting. When the number of materials is less than the lower limit, training is still allowed, but the estimation will be stabilized by increasing the covariance shrinkage coefficient. In terms of numerical caliber, the generation domain and reference domain are sampled in mini-batch form with each parameter update. The number of generated samples and reference samples in a single batch can be set to 4-32 to balance computational efficiency and statistical stability. The typical value of the covariance shrinkage coefficient is between one ten-thousandth and one hundredth, and the specific value can be determined by comparing numerical stability and convergence speed on the development set. The feature dimension constants involved in the formula can be automatically determined based on the number of channels in the selected layer of the style encoder, generally within the range of 256 to 1024. The kernel bandwidth involved in the maximum average difference can be obtained by calculating the median distance between samples in the current batch and multiplying it by an empirical coefficient. The number of components in the multi-bandwidth kernel can be 3-5. The weights can be set uniformly during initialization and then automatically adjusted during training through logarithmic weight updates. Optionally, an upper and lower bound constraint can be introduced when combining weight coefficients to prevent a single metric from having too large a weight in the early training stage, leading to numerical instability. When the covariance matrix of the Gaussian approximation is detected to be close to singular or the condition number is too large, this embodiment can retain only the diagonal elements as a simplified approximation and record the situation in the log for subsequent investigation.

[0075] In one embodiment, the state-aware unit includes a deep sequence coding model, which encodes the behavior type, content features and time information in the user interaction sequence and outputs a user state vector representing the user's current cognitive stage.

[0076] In this embodiment, the user interaction sequence can be directly derived from the original event stream recorded chronologically in the advertising platform or business logs. Each event contains at least one behavior type marker from exposure, click, conversion, close, and skip, along with a corresponding creative identifier, media location information, and timestamp. The timestamp is recorded in milliseconds or seconds within a unified time zone. Specifically, in engineering implementation, behavior types can be converted into fixed-dimensional discrete features using discrete numbering or one-hot encoding. Content features can be obtained by looking up the creative identifier to obtain a pre-calculated vector representation. Time information can be input into the deep sequence coding model according to the time interval between events or a decay weight relative to the current time. Numerically, the system defaults to extracting the interaction sequence from the most recent 7 days for each user and retaining a certain number of recent events in chronological order, for example, retaining 20-100 events by default. Any excess events are discarded from oldest to newest to ensure the sequence length remains within a calculable range. When the number of events a user experiences within a set time window exceeds the upper limit, only the most recent events are retained; if the number of events is insufficient, the actual number is input. Furthermore, the dimension of the user state vector can be consistent with the dimension of the hidden layer of the deep sequence coding model, typically ranging from 64 to 512 dimensions. The specific value can be determined by combining performance on the validation set and online stress testing. Optionally, when the user has no interaction records during a cold start, the state awareness unit can output a pre-set initial state vector. This vector can consist of an all-zero vector or a small number of trainable parameters, with a flag indicating missing history so that downstream decision logic can differentiate and process it. When some event fields are missing, this embodiment can use the method of falling back to the default value or using the most recent valid value to fill in the missing values, while adding a missing flag to the features to ensure that the deep sequence coding model can still stably output the user state vector.

[0077] In one embodiment, the dynamic optimization module further includes a narrative planning unit, which takes the user state vector as input, predicts and outputs a narrative outline containing multiple ordered stage themes through a sequence generation model;

[0078] After receiving the narrative outline, the creative generation module generates a series of creative content that is thematically coherent and progressive, based on the stage theme corresponding to the content to be generated.

[0079] In one embodiment, the dynamic optimization module further includes a deep context decision unit, which takes the user state vector and the candidate creative feature vector as input and estimates the expected reward value of each candidate creative through a deep reward prediction network.

[0080] In one embodiment, when selecting an creative strategy, the deep context decision-making unit adopts a decision-making basis that combines the expected return value output by the deep return prediction network and the intrinsic exploration incentive value based on the model prediction uncertainty. Specifically, for user states and creative types with high prediction uncertainty, the system increases the probability that the corresponding candidate creative will be selected for delivery.

[0081] The intrinsic exploratory incentive value based on model prediction uncertainty is defined as follows:

[0082] Step e, the deep reward prediction network predicts each candidate idea at time... Output Mean-variance pairs obtained from sub-random forward (MCDropout / ensemble):

[0083] ,

[0084] in, Indicates support for candidate ideas At any moment Expected return estimate Indicates the first The average output of the next sample. Indicates the number of samples. Indicates the index of the decision moment. Indicates the candidate creative index;

[0085] Step f, constructing an uncertainty measure, includes the following steps:

[0086] f1, variance-based uncertainty (regression head)

[0087] ,

[0088] ,

[0089] in, This represents the between-sample variance arising from parameter uncertainty. This represents the average heteroscedasticity estimate of the model output. Indicates the first The prediction variance of sub-samples, Represents a measure of total uncertainty;

[0090] f2, entropy-type uncertainty (classification head)

[0091] ,

[0092] in, Indicates category The average predicted probability, Indicates the first Class probability of the next sample Indicates the number of categories. This represents a measure of uncertainty based on Shannon entropy. Represents the natural logarithm;

[0093] f3, Confidence interval width (interval type)

[0094] ,

[0095] in, Indicates confidence level The width of the interval below, This represents the quantiles of a standard normal distribution and their corresponding coverage probabilities. ;

[0096] Step g: After robustly normalizing the intra-batch metrics, linearly fuse them into stimulus values.

[0097] ,

[0098] ,

[0099] in, Represents a robust standardized function. For measurement The median within the batch, For measurement The absolute median difference within the batch, To prevent small constants from being divided by zero, To explore incentive values, The combined weights of the three metrics;

[0100] Step h generates a comprehensive score that can be used for ranking and sampling, and provides a sampling strategy:

[0101] ,

[0102] ,

[0103] in, For comprehensive scoring, For time-varying exploration coefficients, The probability of being selected. For softening temperature, For a moment The set of candidate ideas, sum of denominators, index Traversal ;

[0104] ,

[0105] ,

[0106] in, To update the step size, This represents the average incentive value within the batch at the previous time step. For the target incentive level, For the coefficient boundary, For interval truncation operators, The cardinality of the set;

[0107] Specifically, the above implementation coupling uncertainty measurement and strategy decision-making into a continuous pipeline: first, the mean and variance of the returns are obtained on the candidate set through multiple random forward passes; then, the uncertainty of the model is characterized by three paths: variance type, entropy type, and interval width, taking into account both the regression head and the classification head outputs respectively; the measurement is linearly combined after robust normalization to serve as the exploration incentive, avoiding weight imbalance caused by scale differences, and the exploration intensity is adjusted by a time-varying coefficient, making it higher in the early stage of deployment and gradually converging after feedback accumulation; the comprehensive score is linked with an additive structure and softened probability, which is used directly for ranking on the one hand, and for temperature-controlled random selection on the other hand, improving the coverage of novel ideas and rare states;

[0108] For example, the candidate creative set can consist of multiple creative instances to be delivered under the current ad slot. Each decision corresponds to a process of requesting specific creatives from users or traffic. When the deep reward prediction network performs multiple random forward calculations on this set, sampling can be achieved by enabling random deactivation or switching between multiple parameter snapshots. Numerically, the number of random forward samplings can be set to ten by default, but can be adjusted within the range of 5-20. Entropy uncertainty is suitable for situations where rewards are output in the form of discrete levels or click probabilities, while variance uncertainty is suitable for situations where rewards are output in the form of real-valued revenue. The quantile parameter for the confidence interval width can be set within the range of 90%-95% confidence. For example, at a 90% confidence level, 1.96 can be used as the corresponding standard normal quantile. To avoid differences in the numerical scales of different uncertainty measures leading to a single dominant exploration incentive, this embodiment uses robust normalization of the measures using the intra-batch median and absolute median difference. This normalization strategy is only activated when there are at least three candidate ideas in the same batch. When the number of candidates is too small, it can degenerate into simplified normalization using the mean and standard deviation. Furthermore, the combined weights of the exploration incentives can be set in the range of 0-1, with the sum of the three weights remaining constant at one. Initially, a uniform distribution can be used, and during the online learning phase, the weights are slowly adjusted based on actual reward performance. The temperature parameter in the softening probability can be set in the range of 0.1-1; the higher the temperature, the stronger the randomness, and the lower the temperature, the closer the strategy is to a greedy selection. Optionally, the time-varying exploration coefficient can be initialized to a large value during the cold start phase and decayed in fixed steps over time or with a sample size. When the average value of the exploration incentive within a batch is continuously lower than the preset target level, the decay rate can be appropriately slowed down. When a candidate idea is detected to have a non-numerical or infinite value in its comprehensive score, its score can be reset to the median within the batch or the candidate can be directly removed in this round of decision-making to ensure the computability of the final probability distribution.

[0109] In one embodiment, the dynamic optimization module is configured to update its internal model parameters through online learning based on user interaction feedback obtained after the creative content is delivered.

[0110] The updated model parameters are used to adjust the user state inference logic of the state-aware unit and the creative strategy selection logic of the deep context decision-making unit, so as to influence the subsequent output of the creative generation module through optimization strategy.

[0111] Furthermore, the internal model parameters can be updated in small incremental steps during online learning. In engineering implementation, this is usually triggered by time or sample size. For example, an update cycle is triggered by accumulating a certain number of valid user interactions or after a fixed time interval. In each update cycle, incremental training is performed only on data within the most recent time window to avoid excessive disturbance to previously stable regions. In this embodiment, the learning rate for online updates can be set to one-tenth to 1% of the offline training learning rate to ensure slow parameter convergence and reduce the risk of overfitting to recent noise. The user state inference logic and creative strategy selection logic share the same set of latest parameter snapshots, but a dual-channel mechanism is used during deployment. One set of parameters is used for online inference, while the other set is updated in the background. The new parameters are replaced as a whole after passing a simple consistency check to avoid inconsistencies in model state during inference. Optionally, when the number of feedback samples in a certain update cycle is lower than a preset threshold, the parameter update can be skipped and the previous version of parameters can be used. When abnormal situations such as numerical overflow or abnormally increased loss occur during online training, the system can automatically roll back to the most recent stable parameter snapshot and record the index information of the abnormal batch in the log. To reduce the risk of system failures due to network jitter or storage delays, this embodiment can also cache unuploaded interactive feedback data locally for a short period of time. The caching duration can be set from several minutes to one hour. The data can then be uploaded in batches when the cache is full or the connection is restored, so as to ensure the operability and robustness of the online learning process in the engineering environment.

[0112] In one embodiment, the user interaction sequence includes at least one of the following user behaviors:

[0113] Impressions, clicks, dwell time, conversions, closes, or skips; deep sequence coding models are configured to encode user behavior into corresponding discrete and / or continuous features;

[0114] In one embodiment, the input instructions include one or more of at least target audience attributes, media channels for placement, and marketing objective types, and the creative generation module and / or dynamic optimization module are configured to adjust the content elements and presentation format of the creative content based on the input instructions.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0116] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments. For example, all the embodiments above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.

Claims

1. A deep learning-based marketing creative automatic generation and optimization system, characterized in that, The application comprises: a creative generation module for receiving input instructions containing marketing targets and generating marketing creative content; and a dynamic optimization module in communication connection with the creative generation module; wherein the creative generation module comprises a small sample stylization unit configured to receive a style reference set composed of a small amount of reference creative materials, and optimize the parameters of a creative generation model based on the style reference set, so that the creative content output by the creative generation module is close to the style reference set in style representation, for improving the consistency of the generated creative content with the target brand style; the dynamic optimization module comprises a state perception unit configured to infer the user state according to the user interaction sequence, and generate an optimization strategy according to the user state and output it to the creative generation module for adjusting the subsequently generated creative content.

2. The deep learning based marketing creative auto-generation and optimization system of claim 1, wherein, The creative generation module further comprises a cross-modal semantic alignment unit configured to jointly encode the text description instruction and the visual style instruction, generate a unified semantic-style conditional vector, and control the creative generation module to make the output content meet the semantic constraints of the text description instruction and the style constraints of the visual style instruction at the same time when generating creative content. 3.The deep learning-based marketing creative automatic generation and optimization system of claim 1 or 2, wherein, The small sample stylization unit comprises a style encoder and a style injection network; the style encoder is used to extract the deep style features of the creative materials in the style reference set; the style injection network is integrated into the generation process of the creative generation module, and is used to receive the deep style features and inject the deep style features into the generation process of the creative content through a cross-attention mechanism; The small sample stylization unit is also configured to update the parameters of the creative generation module based on a loss function for measuring the difference between the deep feature distribution of the generated creative content and the deep feature distribution of the style reference set.

4. The deep learning based marketing creative auto-generation and optimization system of claim 1, wherein, The state perception unit comprises a deep sequence encoding model that encodes the behavior type, content features and time information in the user interaction sequence, and outputs a user state vector representing the current cognitive stage of the user.

5. The deep learning based marketing creative auto-generation and optimization system of claim 4, wherein, The dynamic optimization module further comprises a narrative planning unit that takes the user state vector as input, predicts and outputs a narrative outline containing multiple ordered stage themes through a sequence generation model; The creative generation module generates a series of creative content with thematic coherence and progression after receiving the narrative outline according to the stage theme corresponding to the content to be generated. 6.The deep learning based marketing creative auto-generation and optimization system of claim 4 or 5, wherein, The dynamic optimization module further comprises a deep context decision unit that takes the user state vector and candidate creative feature vector as input, and estimates the expected return value of each candidate creative through a deep return prediction network.

7. The deep learning based marketing creative auto-generation and optimization system of claim 6, wherein, The deep context decision unit adopts a decision basis that integrates the expected reward value output by the deep reward prediction network and the intrinsic exploration incentive value based on model prediction uncertainty when selecting creative strategies, wherein for user state and creative type combinations with high prediction uncertainty, the system increases the probability of the corresponding candidate creative being selected for delivery.

8. The deep learning based marketing creative auto-generation and optimization system of claim 6, wherein, The dynamic optimization module is configured to update its internal model parameters in an online learning manner according to user interaction feedback obtained after the delivery of creative content; The updated model parameters are used to adjust the user state inference logic of the state awareness unit and the creative strategy selection logic of the deep context decision unit, for affecting the subsequent output of the creative generation module through the optimization strategy. 9.The deep learning based marketing creative auto-generation and optimization system of claim 4, wherein, The user interaction sequence includes at least one of the following user behaviors: exposure, click, dwell time, conversion, close or skip; the deep sequence encoding model is configured to encode the user behavior into corresponding discrete features and / or continuous features. 10.The deep learning based marketing creative auto-generation and optimization system of claim 1 or 2, wherein, The input instructions include one or more of at least target population attributes, delivery media channels and marketing target types, and the creative generation module and / or the dynamic optimization module are configured to adjust the content elements and presentation forms of the creative content based on the input instructions.