A brand constraint-based advertisement material generation method and system

CN122550233APending Publication Date: 2026-08-11GUANGZHOU SANSHI CULTURE COMMUNICATION CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

1. 品牌风格一致性难以维持:现有多模态生成方案(如申请人上海海湃领客文化科技有限公司的“面向产品包装策划的AI驱动全营销内容设计方法”,公开号CN120408755A)提出了构建品牌调性母向量以监控风格一致性的方法,但其应用场景被限定为产品包装策划,且缺乏在生成过程中进行像素级强约束的机制,导致生成的全类型广告素材(如信息流、短视频)风格容易跑偏,仍需大量人工审核

Benefits of technology

1. 品牌风格一致性突破:通过将品牌母向量转换为门控约束注入生成网络,实现了对素材风格像素级的强约束。在模拟测试条件下(使用15个品牌的VI规范及5000条测试需求),品牌风格偏离度评分平均降低42.3%,自动校验通过率达到98.7%,大幅减少人工审核成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550233A_ABST
    Figure CN122550233A_ABST
Patent Text Reader

Abstract

This invention discloses a brand-constrained advertising creative generation method and system, relating to the fields of artificial intelligence and internet advertising technology. The method constructs a brand parent vector as a strong style constraint anchor point, injects it into the dual-tower Transformer multimodal generation backbone, and ensures style consistency of the generated creatives through cosine similarity verification. Simultaneously, it constructs a style fatigue quantification index system, triggering constrained adaptive style migration when a threshold is reached. Furthermore, it acquires creative specifications and target terminal performance parameters from multiple advertising platforms, dynamically selects the optimal encoding parameters through a reinforcement learning-based adaptive rendering decision model, and performs perceptual lossless compression on the creatives. This invention solves the problems of maintaining brand style consistency, requiring extensive manual adaptation of cross-platform creatives, and rigid style management in existing multimodal generation methods. It achieves fully automated generation of advertising creatives with highly consistent brand tone, efficient cross-platform distribution, and intelligent style evolution, significantly improving creative click-through rates and advertising efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and internet advertising technology, and in particular to a method and system for generating advertising materials based on brand constraints.

[0002] The user feedback data (including but not limited to click-through rate, conversion rate, and completion rate) involved in the implementation of this invention are all collected from monitoring data of advertising platforms with the user's active authorization and consent. This data does not contain any personally identifiable information and does not involve facial recognition data or sensitive personal information. The training data for the generated model comes from legally authorized commercial material libraries and proprietary materials provided by the brand, and no unauthorized web scraping data is used. The implementation of this invention fully complies with the relevant provisions of the "Personal Information Protection Law of the People's Republic of China" and the "Data Security Law of the People's Republic of China". Background Technology

[0003] In the field of internet advertising, generating high-quality, multi-platform compatible ad creatives is a core component. Current technologies primarily suffer from the following pain points: 1. Brand style consistency is difficult to maintain: Existing multimodal generation solutions (such as the "AI-driven full marketing content design method for product packaging planning" by the applicant, Shanghai Haipai Linker Culture Technology Co., Ltd., public number CN120408755A) propose a method to build a brand tone vector to monitor style consistency, but its application scenario is limited to product packaging planning, and it lacks a mechanism for strong pixel-level constraints during the generation process, which makes it easy for the style of all types of advertising materials (such as information flow and short video) to deviate, and still requires a lot of manual review.

[0004] 2. Multi-platform creative adaptation heavily relies on manual intervention: The same ad creative needs to be distributed to more than ten platforms, including ByteDance Ads and Tencent Ads, each with different requirements for creative resolution, aspect ratio, file size, and encoding format. Existing technologies (such as the "Creative Generation Method Based on Multimodal Dynamic Feedback" from Jingmeng Century (Beijing) Technology Co., Ltd., publication number CN120525585A) involve cross-platform format conversion, but they are general format conversions and lack intelligent encoding decisions based on real-time terminal capabilities. They cannot minimize file size while ensuring visual quality, and manual parameter tuning is extremely inefficient.

[0005] 3. Rigid Style Management and Inability to Perceive Content Fatigue: Existing solutions cannot quantify and perceive "content fatigue" in the market. The evolution of brand style relies entirely on human experience and judgment, failing to achieve intelligent and controllable style evolution while maintaining the brand's core tone. Furthermore, the "Advertising Creative Matching Method Based on Multimodal Content Generation" (Publication No. CN121808075A) applied for by Dagen Holdings Co., Ltd. discloses semantic alignment of text and images through cross-modal time anchoring and cultural fingerprint databases. However, its solution focuses on constructing semantic barriers against cultural taboos and does not provide a pixel-level constraint mechanism for the overall visual style of the brand during the generation process, nor does it address the issue of adaptive rendering of materials across multiple advertising platforms. The "Advertising Content Generation Method Based on Multimodal Fusion" (Publication No. CN121304247A) applied for by an entity in Rongcheng County discloses a material optimization method based on market adaptation adjustment vectors. However, its solution focuses on predicting the matching degree between materials and the market, rather than the structured constraints of the brand's visual specifications, and cannot solve the problem of brand style consistency. The application filed by individual applicant Chen Ping, entitled "A Method for Generating Ad Push Screens Based on Style Transfer" (Publication No. CN121353454A), discloses a method for style transfer using style-oriented parameters to drive generative AI. However, this style transfer is completed in a one-time process, lacking a continuous perception and adaptive management mechanism for brand style fatigue, and thus failing to achieve the organic evolution of brand style over long-term advertising campaigns. In summary, existing technologies have not yet provided an ad creative generation solution that can strongly constrain brand style during the generation process, while simultaneously achieving adaptive transfer of style fatigue and intelligent rendering across multiple platforms. This is precisely the core problem that this invention aims to solve. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an integrated fully automatic solution that seamlessly connects brand constraints, adaptive migration of style fatigue and multi-platform intelligent rendering, in order to address the above-mentioned deficiencies of the prior art. This solution enables the entire process of advertising materials from generation to cross-platform distribution to be completed automatically and intelligently under strong brand constraints.

[0007] To solve the above-mentioned technical problems, the technical solution disclosed in this invention is as follows: A method for generating advertising creatives based on brand constraints, comprising: S1. Construct the brand parent vector and inject it as a gating constraint into the dual-tower Transformer to generate the backbone; S2. Ensure style consistency of generated materials through cosine similarity verification; S3. Construct a quantitative index system for style fatigue, and trigger constrained style migration when fatigue occurs; S4. Collect multi-platform specifications and terminal performance parameters, and use a reinforcement learning model to dynamically decide the optimal encoding parameters; S5. Based on visual saliency, perform lossless compression to generate and output the final adapted material.

[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Breakthrough in Brand Style Consistency: By converting the brand's parent vector into a gated constraint and injecting it into the generative network, strong pixel-level constraints on the style of the materials were achieved. Under simulated testing conditions (using the VI specifications of 15 brands and 5000 test requirements), the brand style deviation score decreased by an average of 42.3%, and the automatic verification pass rate reached 98.7%, significantly reducing the cost of manual review.

[0009] 2. Revolution in Multi-Platform Content Production: The time for adapting a single piece of content to multiple platforms has been reduced from several hours of manual operation to less than 3 seconds. Furthermore, through reinforcement learning-based intelligent coding decisions, the Visual Quality Assessment Score (LPIPS) reaches the optimal level for similar image quality while 100% meeting platform file size limits, resulting in an overall improvement of approximately 50 times in human efficiency for ad delivery.

[0010] 3. Intelligent Brand Vitality Management: This feature introduces an adaptive style migration based on content fatigue index, preventing brand style from becoming rigid. In a three-month A / B testing experiment, the click-through rate decay period of the materials using this solution was 2.3 times longer than that of the fixed-style group, and the brand image keyword consistency score remained above 92 points, achieving an organic evolution of the brand that "changes without losing its essence." Attached Figure Description

[0011] Figure 1 This is an overall flowchart of the advertising material generation method in an embodiment of the present invention.

[0012] Figure 2 This is a module structure diagram of the system in an embodiment of the present invention.

[0013] Figure 3 A detailed diagram illustrating the process of constructing and constraining the brand parent vector.

[0014] Figure 4 A flowchart for multi-platform adaptive rendering and intelligent encoding decisions.

[0015] Figure 5 This is a flowchart of style fatigue perception and adaptive transfer. Detailed Implementation Example

[0016] This invention relates to a method and system for generating advertising materials based on brand constraints. Figure 1 The overall flow of the method is shown.

[0017] Step S1: Brand parent vector construction and constraint injection.

[0018] First, the system obtains three types of input data: (1) Advertising demand description text, such as "Launching a new summer lemon tea, focusing on refreshing and palatable taste, targeting Generation Z"; (2) High-definition images of the product, namely the packaging and actual product images of the new lemon tea; (3) Brand visual identity system specification documents, which clearly define the brand's standard color space, logo safe zone, font family, visual density range and other rules.

[0019] Next, a ResNet-50 deep residual network pre-trained on ImageNet (with the final fully connected classification layer removed) is used to extract a 2048-dimensional feature vector from the product image, covering visual features such as color distribution, compositional patterns, and texture patterns. Simultaneously, a BERT-base-Chinese pre-trained language model is used to encode the brand VI specification text, and the 768-dimensional latent vector corresponding to the [CLS] tag is taken as the brand specification semantic feature. The 2048-dimensional image features and 768-dimensional text features are adaptively fused through a cross-attention fusion layer to generate a 512-dimensional brand parent vector v_brand, serving as the brand's unshakeable style anchor. This fusion method can dynamically adjust the contribution ratio of image features and text features according to the weights of each visual element defined in the brand specification document.

[0020] This embodiment employs a dual-tower Transformer architecture for its generative backbone. The text tower and image tower each consist of 12 stacked Transformer blocks, each layer containing a multi-head self-attention mechanism (12 heads) and a feedforward network, with a hidden state dimension of 768. The two towers perform semantic alignment within a shared cross-modal latent space. The key innovation of this invention lies in adding a style constraint adaptation layer after the Multi-Head Self-Attention sub-layer of each Transformer block. Figure 3 As shown, this layer fuses the current hidden state vector h_i with the brand parent vector v_brand through a gating mechanism to achieve fine-grained constraints on the generation process. The fusion formula is: h'_i = h_i ⊙ σ(W_g · [h_i, v_brand] + b_g) Where h_i ∈ R^768 is the hidden state vector at the current i-th position, v_brand ∈ R^512 is the brand mother vector, [·,·] represents the linear transformation projected to 768 dimensions after vector concatenation, W_g ∈ R^(768×(768+512)) is the learnable gating weight matrix (768 dimensions after projection), b_g is the bias term, σ is the sigmoid activation function, and ⊙ represents element-wise multiplication. Through this gating mechanism, the brand mother vector can finely control the style expression of each layer of features, as if there is a strict "brand steward" in each layer of the network, ensuring that every step of the output does not deviate from the brand track.

[0021] Step S2: Constraint generation and style verification.

[0022] The system generates a backbone constrained by the brand's parent vector, directly outputting preliminary materials. The system then uses a feature extractor with the same structure as the style verification network to extract the 512-dimensional feature vector of the material and calculates its cosine similarity to the brand's parent vector, v_brand. This value ranges from -1 to 1, with a value closer to 1 indicating greater style consistency. In this embodiment, a style consistency threshold of 0.92 is set. If the similarity is lower than this value, the system determines the material is "off track," automatically triggering a regeneration mechanism to adjust the random seed and regenerate until it passes verification. This process completely replaces the manual "glance" review process, achieving closed-loop quality control.

[0023] Step S3: Style fatigue perception and adaptive transfer.

[0024] Even excellent creative content will eventually lead to market fatigue. After the creative content is deployed, the system aggregates daily feedback data such as click-through rate (CTR), conversion rate (CVR), and completion rate via the advertising platform API. This embodiment creatively defines the "Content Fatigue Index" (CFI), and its calculation formula is as follows: CFI = α × (CTR decay slope / baseline decay slope) + β × (1 - current conversion rate / historical average conversion rate) + γ × brand style deviation Among them, the CTR decay slope is the absolute value of the slope of the linear regression of CTR over the past 7 days; the baseline decay slope is the average normal decay slope of similar creatives in history; the current conversion rate is the average conversion rate over the past 3 days; the historical average conversion rate is the average conversion rate since the campaign started, excluding the first 3 days; and the brand style deviation is the normalized value of the Euclidean distance between the current creative feature vector and the brand parent vector v_brand (value range [0,1]). The weight coefficients α, β, and γ can be adjusted according to the brand's strategy, with initial default values ​​of 0.5, 0.3, and 0.2, respectively.

[0025] like Figure 5 As shown, when the system detects that the CFI exceeds the preset fatigue trigger threshold (e.g., 0.75), it determines that the style has entered a fatigue period and automatically triggers an adaptive style migration strategy. Migration is not about overturning the brand image, but rather about directional exploration within a high-dimensional sphere (a constrained space with cosine similarity ≥ 0.85) limited by the brand's parent vector. Specifically, it involves: obtaining the top K best-performing creatives (highest CTR and CFI not exceeding the threshold) before the CFI trigger within the current campaign period; calculating the mean v_best of the feature vectors of these creatives; and constructing a performance gradient direction vector Δv = normalize(v_best - v_brand). The current brand parent vector v_brand is then weighted and combined with Δv to generate the offset conditional vector v_brand' = normalize(v_brand + λ · Δv), where λ is the step size limited by the maximum deviation threshold (initial value 0.01, maximum not exceeding 0.05). This conditional vector is then fed into the decoder of the multimodal generation backbone, replacing the original brand parent vector as a style condition, guiding the generation of materials with new style tweaks but maintaining the core brand tone. The offset range is strictly limited to ensure that the brand "changes without losing its essence," achieving organic evolution.

[0026] Step S4: Multi-platform adaptive rendering and intelligent encoding.

[0027] After the creative materials pass style verification and potential style migration, they face the challenge of multi-platform distribution. The system obtains the latest creative material format specifications (supported resolution, aspect ratio, encoding format, file size limit, etc.) from the target advertising platforms (such as ByteDance Ads, Tencent Ads, Google Ads, etc.) through API interfaces. At the same time, through a lightweight monitoring module integrated into the advertising SDK, it transmits back parameters such as the user's terminal's network bandwidth, GPU model and decoding capability, and screen resolution in real time.

[0028] like Figure 4As shown, this embodiment models the multi-platform encoding parameter decision as a Markov Decision Process (MDP) and uses a Deep Q-Network (DQN) for solution. The state space S consists of network quality (poor / medium / good), GPU decoding capability (low / medium / high), visual complexity of the source material (complexity score calculated based on edge density and texture information), and target platform specification constraints. The action space A is a predefined list of encoding parameter combinations, including resolution (e.g., 720p, 1080p, 2K), video encoding format (H.264, H.265, AV1), bitrate level (low / medium / high), frame rate (24 / 30 / 60fps), and color depth (8bit / 10bit). The reward function R is designed as: R = ω1 × VQ_score - ω2 × (FileSize / Limit), where VQ_score is the visual quality score evaluated using perceptual quality metrics such as LPIPS, FileSize is the compressed file size, Limit is the platform file size limit, and ω1 and ω2 are balancing parameters. The DQN model includes an experience replay pool, which stores the (state, action, reward, next state) quadruple for each decision in the pool and samples it periodically for training, enabling online incremental learning.

[0029] The decision-making process is completed instantly. The model selects the optimal combination of encoding parameters from the action space that maximizes the expected cumulative reward, such as "1080p resolution, H.265 encoding, bitrate 2Mbps, frame rate 30fps". This decision minimizes the file size while ensuring visual quality, perfectly meeting the platform requirements.

[0030] Step S5: Perceive lossless compression and output.

[0031] Finally, to maximize subjective image quality within a limited file size, visual saliency detection is introduced. This embodiment uses the U²-Net saliency detection network based on deep learning. The input is the advertising material frame to be compressed, and the output is a pixel-level saliency probability map. Pixel regions with a probability value higher than 0.7 are defined as high-saliency regions (typically corresponding to product display areas, brand logos, call-to-action buttons like "Buy Now"), while the remaining regions are defined as low-saliency background regions.

[0032] During encoding and compression, a region-based bitrate control strategy is employed: the macroblock quantization parameter QP for highly saliency regions is set to a relatively small value, allocating approximately 90% of the available bitrate budget to ensure visual detail of key commercial information; for background regions (such as the sky and blurred textures), a larger quantization step size is used to significantly compress file size. The final generated material is visually virtually lossless, the file size perfectly meets the target platform's limitations, and the brand style is highly consistent, allowing for direct output into the deployment process. Example

[0033] The present invention also provides a system for implementing the above method, such as... Figure 2 As shown, the system includes: a brand parent vector construction and injection module, a constraint generation and style verification module, a style fatigue perception and adaptive migration module, a multi-platform adaptive rendering and intelligent encoding module, and a perceptual compression and output module. The functions of each module have been detailed in the aforementioned method embodiments and will not be repeated here.

Claims

1. A method and system for generating advertising creatives based on brand constraints, characterized in that, Includes the following steps: S1. Brand Mother Vector Construction and Constraint Injection: Obtain advertising requirement description text, product images, and brand visual identity system specification documents; use a deep convolutional neural network to extract the color, composition, and texture features of the product images, and use a semantic encoder to extract the font, visual density, and semantic features of the brand specification documents; perform multimodal fusion on the above features to construct a brand mother vector to represent the brand's visual style; use the brand mother vector as a constraint condition and inject it into each layer of the forward propagation process of the multimodal generation backbone based on the dual-tower Transformer architecture. S2. Constraint Generation and Style Verification: The multimodal generation backbone generates preliminary advertising materials under the constraints of the brand parent vector; the cosine similarity between the feature vector of the preliminary advertising materials and the brand parent vector is calculated. If the cosine similarity is lower than the preset style consistency threshold, the regeneration mechanism is automatically triggered until the generated materials pass the style verification. S3. Style Fatigue Perception and Adaptive Migration: After ad campaigns, click-through rates, conversion rates, and completion rates of generated creative materials are collected as feedback data; a quantitative indicator system for style fatigue is constructed, and the current brand's Content Fatigue Index (CFI) is calculated. When the CFI exceeds the preset fatigue trigger threshold, the adaptive style transfer strategy is automatically triggered. The style transfer strategy, within the constraint space of the brand parent vector, achieves a gradual evolution of brand style by adjusting the decoder input conditions of the multimodal generation backbone, and the transfer range is limited by a preset maximum deviation threshold. S4. Multi-platform adaptive rendering and intelligent encoding: Obtain the material format specifications of at least one target advertising platform, as well as the real-time network status and GPU decoding capability parameters of the target playback terminal; establish an adaptive rendering decision model based on reinforcement learning, and dynamically select the optimal combination of encoding parameters, including resolution, bit rate, frame rate and color depth, according to the material format specifications, network status and GPU decoding capability parameters. S5. Perceptual Lossless Compression and Output: The visual saliency detection algorithm is used to identify the product display area and call-to-action button area in the style-verified advertising material. A relatively higher bitrate is allocated to areas with visual saliency higher than a preset threshold, and a relatively lower bitrate is allocated to background areas. Under the premise of meeting the file size limit of the target advertising platform, the final multi-platform adapted advertising material is generated and output.

2. The method according to claim 1, characterized in that, In step S1, the method of injecting the brand parent vector into the generation backbone is to set a style constraint adaptation layer after each layer of the self-attention mechanism of the dual-tower Transformer generation backbone; the style constraint adaptation layer fuses the hidden state vector of the current layer with the brand parent vector through a gating mechanism to achieve style constraint on the generated content.

3. The method according to claim 1, characterized in that, The multimodal fusion in step S1 employs an attention mechanism, which adaptively allocates the fusion ratio of image features and semantic features according to the weight definition of each visual element in the brand specification document, thereby constructing the brand parent vector.

4. The method according to claim 1, characterized in that, The Content Fatigue Index (CFI) in step S3 is a comprehensive evaluation value obtained by weighting multiple indicators based on the feedback data, used to quantify the degree of brand style fatigue.

5. The method according to claim 1, characterized in that, In step S4, the adaptive rendering decision model based on reinforcement learning is a reinforcement learning model based on a deep Q-network (DQN). Its state space includes the terminal network status, GPU decoding capability, screen resolution, and material complexity. The action space is a combination of optional encoding parameters. The reward function provides feedback based on the visual quality evaluation score and file size of the generated material. The method also includes using the optimal combination of encoding parameters from each decision as training samples and replaying it into the DQN model to achieve online incremental learning of the model.

6. A method and system for generating advertising creatives based on brand constraints, characterized in that, include: Brand parent vector construction and injection module: used to perform the operation described in step S1 of claim 1; Constraint generation and style verification module: used to perform the operation described in step S2 of claim 1; Style fatigue perception and adaptive migration module: used to perform the operation described in step S3 of claim 1; Multi-platform adaptive rendering and intelligent encoding module: used to perform the operation described in step S4 of claim 1; Perception compression and output module: used to perform the operation described in step S5 of claim 1.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Product packaging planning-oriented AI-driven full-marketing content design method

    CN120408755A

  • Material generation method and device based on multi-modal dynamic feedback, computer equipment and readable storage medium

    CN120525585A

  • Advertisement content generation method based on multi-modal fusion

    CN121304247A

  • Advertisement push picture generation method and system based on style migration

    CN121353454A

  • Advertisement creativity matching method based on multi-modal content generation

    CN121808075A