An advertisement copy and visual collaborative design system fusing a multi-modal large model

CN122841013APending Publication Date: 2026-09-29HANGZHOU QUNZHI INFORMATION CONSULTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611048336.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]但是其在实际使用时,仍旧存在一些缺点,如文案生成路径与图像生成路径仅在输入层共享条件特征,内部变换层之间缺乏实时的跨模态信息交换,导致生成的文本与图像在局部细节上难以精确对齐;联合判别器仅在最终输出端提供整体匹配度反馈,无法在生成过程中逐层修正模态间的特征偏差,优化效率较低且容易陷入局部最优;缺乏对图文风格一致性的量化校验与反馈调节机制,无法自动识别并修正如“欢快文案搭配冷色调图像”或“科技感文案匹配粗糙纹理”等风格冲突问题,最终输出的广告内容往往需要人工二次调整,难以实现端到端的自动化协同设计

Benefits of technology

[0021]1、本发明通过设置双路交叉注意力协调模块,在文本生成路径与图像生成路径的每一对应变换层之间构建双向交叉注意力模块,以一路的中间激活值为查询、另一路的中间激活值为键和值,实时交换跨模态特征并叠加回原路径,实现了文案与视觉特征在深层网络中的逐层协同演化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841013A_ABST
    Figure CN122841013A_ABST
Patent Text Reader

Abstract

The application discloses a kind of fusion multimodal big model's advertisement script and visual collaborative design system, specifically related to advertisement intelligent generation technical field, including multi-modal joint embedding coding module, the semantic and style features of product text and reference image are extracted and spliced into joint embedding tensor;Double-path cross attention coordination module, cross attention module is realized by symmetry text and image generation path's cross-modal feature bidirectional exchange;Collaborative loss closed-loop control module, based on semantic similarity matrix generates collaborative loss scalar, adjusts attention weight in reverse to optimize text semantic alignment;Cross-modal style check module, calculate Mahalanobis distance from color-emotion, texture-style double dimension, trigger re-optimization when over threshold value;Alignment output module, script features are decoded into time-stamped structured text stream, visual features are decoded into hierarchical data packet and realize space-time alignment output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent advertising generation technology, and more specifically, to an advertising copywriting and visual collaborative design system that integrates multimodal large models. Background Technology

[0002] In the field of advertising creative design, the collaborative generation of copy and visual content is a core element in improving the effectiveness of advertising communication. With the development of deep learning technology, some solutions have attempted to use generative models to generate advertising text and images separately, but the two are still in the stage of independent generation and post-processing splicing, lacking in-depth semantic and stylistic collaboration.

[0003] The closest current technical solution involves receiving product keywords or descriptive text, extracting semantic features through a text encoder, and then inputting these features into a text generator and an image generator, respectively. The text generator outputs advertising copy, and the image generator outputs an advertising image semantically related to the copy. Both generators share the same conditional feature vector but have no direct intermediate feature interaction. The generated copy and image are evaluated holistically by a joint discriminator, which outputs an image-text matching score. This score is then used to update the parameters of both generators.

[0004] However, in practical use, it still has some shortcomings. For example, the text generation path and the image generation path only share conditional features at the input layer, and there is a lack of real-time cross-modal information exchange between the internal transformation layers, which makes it difficult to accurately align the generated text and images in local details. The joint discriminator only provides overall matching feedback at the final output end and cannot correct the feature deviation between modes layer by layer during the generation process. The optimization efficiency is low and it is easy to get trapped in local optima. There is a lack of quantitative verification and feedback adjustment mechanism for the consistency of text and image styles. It cannot automatically identify and correct style conflict problems such as "cheerful text paired with cool-toned images" or "tech-savvy text paired with rough textures". The final output of the advertising content often requires manual adjustment, making it difficult to achieve end-to-end automated collaborative design. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide an advertising copywriting and visual collaborative design system that integrates multimodal large models, thereby solving the problems mentioned in the background art through the following solutions.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an advertising copywriting and visual collaborative design system integrating a multimodal large model, including a multimodal joint embedding encoding module: receiving product text description and reference image, extracting text semantic vector and visual style vector, and concatenating them along the feature dimension to generate a joint embedding tensor;

[0007] Dual-path cross-attention coordination module: The input end is connected to the joint embedding tensor, and the text generation path and image generation path are run simultaneously. A cross-attention module is set between the corresponding transformation layers of the two paths. The intermediate activation value of one path is used as the query and the intermediate activation value of the other path is used as the key and value. The cross-modal features are exchanged and superimposed back to the original path, and the text feature matrix and visual feature map are output.

[0008] Collaborative loss closed-loop control module: Projects the text feature matrix into a semantic vector sequence, projects the visual feature map into a semantic vector set in the same space, calculates the element-wise similarity matrix to generate a collaborative loss scalar, and adjusts the weight matrix of all cross-attention modules in the dual-path cross-attention coordination module in reverse to optimize the process;

[0009] Cross-modal style verification module: Receives optimized copy features and optimized visual feature map, extracts global color histogram and texture statistics from visual map, extracts sentiment polarity vector and style description vector from copy, calculates the first distance between color histogram and sentiment polarity, and the second distance between texture statistics and style description. If any distance exceeds the preset threshold, a re-optimization signal is triggered and sent back to the collaborative loss closed-loop control module.

[0010] Alignment Output Module: Decodes the final text features into a text stream, decodes the final visual feature map into a layered visual data package, and aligns the text paragraphs with the corresponding visual layers in terms of timestamps.

[0011] Preferably, the multimodal joint embedding coding module includes a text preprocessing unit, which uses a TF-IDF keyword filtering algorithm to remove words in the product text description whose information gain value is less than or equal to a set threshold, and retains valid keywords whose information gain value is greater than the set threshold.

[0012] Preferably, the multimodal joint embedding encoding module further includes a text encoder and an image encoder. The text encoder adopts the RoBERTa-Large model, which is fine-tuned using an e-commerce advertising text dataset to output a text semantic vector. The image encoder adopts the ViT-L / 14 model, which is fine-tuned using a commercial advertising image dataset to output a visual style vector. The fine-tuning process includes freezing the underlying parameters, adding a task-specific output head, multi-task joint training, and model selection steps.

[0013] Preferably, the multimodal joint embedding encoding module adopts an adaptive weighted concatenation algorithm based on advertising feature dimensions. The weight coefficients of the text semantic vector and the visual style vector are adjusted according to the advertising scenario type. After weighting, the two vectors are concatenated along the feature dimension to form a joint embedding tensor, and L2 normalization constraints are applied to the interval [-1,1].

[0014] Preferably, in the dual-path cross-attention coordination module, both the text generation path and the image generation path are equipped with progressive feature transformation layers. Each layer consists of a self-attention sub-layer, a feedforward neural network sub-layer, and a cross-attention sub-layer. In the cross-attention sub-layer, the intermediate activation value of one path is used as the query and the intermediate activation value of the other path is used as the key and value. After being mapped to a unified attention feature space through independent linear projection layers, the cross-modal attention output is calculated and superimposed back to the original path output with a set residual coefficient.

[0015] Preferably, the dual-path cross-attention coordination module introduces a sparse attention mechanism: the marketing selling point feature location in the text and the product main feature region in the image are identified by a pre-trained advertising core region detection model, and full attention calculation and sliding window attention calculation are performed respectively.

[0016] Preferably, in the collaborative loss closed-loop control module, the text feature matrix is ​​projected into a semantic vector sequence S, and the visual feature map is projected into a semantic vector set V; the semantic similarity matrix M is calculated.

[0017] Preferably, when the cross-modal style verification module extracts the global color histogram vector H of the visual feature map, it uniformly divides the gray values ​​of the RGB three channels into intervals and generates a histogram by counting the number of pixels in the color intervals; when extracting the texture statistics vector B, it adopts the gray-level co-occurrence matrix method, sets the gray level, step size, and angle, and calculates the average value of four indicators: contrast, correlation, energy, and entropy, to obtain the texture statistics vector.

[0018] Preferably, when the cross-modal style verification module extracts the sentiment polarity vector E of the copy, it uses the BERT sentiment analysis model, fine-tuned on the advertising copy dataset, and outputs a 3-dimensional vector corresponding to the probabilities of positive, neutral, and negative, respectively; and extracts the style description vector S. ty When using a combination of keyword matching and semantic analysis, a 5-dimensional vector is output, corresponding to the probabilities of five styles: minimalist, luxurious, cute, technological, and natural. When calculating the first distance between the color histogram and the sentiment polarity, and the second distance between the texture statistics and the style description, Mahalanobis distance is used, and the first distance threshold and the second distance threshold are preset.

[0019] Preferably, when the alignment output module decodes the final copy features into a text stream, it uses a GPT-2 decoder and fine-tunes it, adds a linear projection layer at the input end to map the features to the word embedding space, and adds three parallel output heads: a text generation head that outputs the probability distribution of the next token, a semantic tag head that outputs the module tags, and a timestamp head that outputs the start and end timestamps; the decoded text stream includes four parts: title, body, selling points, and call to action, each part with semantic tags and timestamp information.

[0020] The technical effects and advantages of this invention are as follows:

[0021] 1. This invention establishes a bidirectional cross-attention coordination module by setting up a dual-path cross-attention module. This module is constructed between each corresponding transformation layer of the text generation path and the image generation path. The intermediate activation value of one path is used as the query, and the intermediate activation value of the other path is used as the key and value. Cross-modal features are exchanged in real time and superimposed back to the original path, thereby realizing the layer-by-layer collaborative evolution of text and visual features in deep networks.

[0022] 2. This invention constructs a collaborative loss closed-loop control module, projects the text feature matrix and visual feature map onto the same semantic space, calculates the element-wise similarity matrix, generates a collaborative loss scalar consisting of semantic alignment loss and diversity loss, and adjusts the weight matrix of all cross-attention modules in the dual-path cross-attention coordination module in real time through backpropagation. This enables the feature deviation between modalities to be corrected layer by layer during the generation process. Each time a set of text and visual features is generated, a loss calculation and weight update are performed until the collaborative loss is lower than a preset threshold, resulting in higher optimization efficiency and convergence accuracy.

[0023] 3. This invention introduces a cross-modal style verification module to quantitatively verify the advertising copy and visual style from two dimensions: color-emotion and texture-style. It extracts the global color histogram and texture statistics of the visual image, extracts the emotional polarity vector and style description vector of the copy, calculates the Mahalanobis distance and compares it with a preset threshold. If any distance exceeds the threshold, a re-optimization signal is triggered and sent back to the collaborative loss closed-loop control module. In the next iteration, the loss weight of the corresponding dimension is increased, realizing end-to-end automated collaborative design and significantly improving the style consistency and usability of the generated advertisement. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the overall system structure of the present invention;

[0025] Figure 2 This is a schematic diagram of the multimodal joint embedding coding of the present invention;

[0026] Figure 3 This is a schematic diagram of the dual-path cross-attention coordination of the present invention;

[0027] Figure 4 This is a schematic diagram of the collaborative loss closed-loop control of the present invention;

[0028] Figure 5 This is a schematic diagram of the cross-modal style verification of the present invention. Detailed Implementation

[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] like Figure 1 The system shown is an advertising copywriting and visual collaborative design system that integrates a multimodal large model, including a multimodal joint embedding encoding module: receiving product text descriptions and reference images, extracting text semantic vectors and visual style vectors, and concatenating them along the feature dimensions to generate a joint embedding tensor.

[0031] It should be specifically noted that the multimodal joint embedding coding module includes a text preprocessing unit, which uses the TF-IDF keyword filtering algorithm to remove words in the product text description with an information gain value less than or equal to 0.1, and retains effective keywords with an information gain value greater than 0.1.

[0032] The multimodal joint embedding coding module also includes a text encoder and an image encoder. The text encoder uses the RoBERTa-Large model, which is fine-tuned using a dataset of 1.2 million e-commerce advertising texts to output a 1024-dimensional text semantic vector. The image encoder uses the ViT-L / 14 model, which is fine-tuned using a dataset of 800,000 commercial advertising images to output a 1024-dimensional visual style vector. The fine-tuning process includes freezing the underlying parameters, adding a task-specific output head, multi-task joint training, and model selection steps.

[0033] In the multimodal joint embedding encoding module, an adaptive weighted concatenation algorithm based on advertising feature dimensions is adopted. The weight coefficients of the text semantic vector and the visual style vector are adjusted according to the advertising scenario type. After weighting, the two 1024-dimensional vectors are concatenated along the feature dimensions to form a 2048-dimensional joint embedding tensor, and L2 normalization constraints are applied to the interval [-1,1].

[0034] like Figure 2 As shown, it should be further explained that this module is a dedicated front-end coding module for the system, adopting a multimodal fine-grained coding architecture specific to commercial advertising, and customizing data processing rules and feature extraction dimensions for three core advertising scenarios: e-commerce detail pages, short video feeds, and outdoor posters.

[0035] The product text description received by the multimodal joint embedding encoding module is structured advertising-specific text, mandating the inclusion of four mandatory fields: core product parameters, marketing selling points, target audience, and marketing scenario. The length of a single text entry is limited to 20-200 characters. The module uses a TF-IDF keyword filtering algorithm to remove redundant text lacking substantial information, such as "very good" or "worth buying," retaining only valid keywords with an information gain greater than 0.1. The reference image is the original advertising material, uniformly preprocessed to a 512×512 pixel RGB three-channel image, with size normalization achieved through bilinear interpolation. Simultaneously, grayscale images, images with abnormal transparency channels, and low-quality images with a signal-to-noise ratio below 30dB are removed.

[0036] To further explain, the TF-IDF keyword filtering algorithm is used to measure the importance of words in a single text. The core logic is that the importance of a word is directly proportional to its frequency in a single text and inversely proportional to its frequency in the entire corpus. The TF-IDF value is the product of the word's term frequency and its inverse document frequency, used to determine the word's substantive information value. The module pre-calculates the IDF values ​​of all commonly used words based on a corpus of 1.2 million e-commerce advertising texts. For the input single product text description, after word segmentation, the TF-IDF value of each word is calculated: generic praise words without substantive information, due to their high frequency of occurrence in all advertisements, have extremely low IDF values, and their final TF-IDF value is usually less than 0.1, and are automatically removed; keywords containing core product information, due to their domain specificity, usually have TF-IDF values ​​greater than 0.1 and are retained for subsequent semantic encoding. In this invention, the 0.1 threshold is determined by statistically analyzing the TF-IDF value distribution of 100,000 redundant texts in the corpus.

[0037] The text encoding side uses the RoBERTa-Large model, fine-tuned with a dataset of 1.2 million e-commerce advertising texts, featuring 24 encoding layers, 1024 hidden layers, and a 16-head attention mechanism. During fine-tuning, the RoBERTa-Large model added marketing point weight prediction and sentiment classification tasks, enabling it to extract three core features from the advertising text: marketing semantic weight, selling point priority, and sentiment intensity, outputting a 1024-dimensional text semantic vector. The visual encoding side uses the ViT-L / 14 model, fine-tuned with a dataset of 800,000 commercial advertising images, with a fixed image block size of 14×14. During fine-tuning, four auxiliary tasks were added: color tone classification, texture classification, layout classification, and scene atmosphere classification, enabling it to extract these four specific visual features from the advertising images, outputting a 1024-dimensional visual style vector.

[0038] To further explain, fine-tuning is a standardized process for domain adaptation based on a general pre-trained model. This invention adopts a unified 5-step fine-tuning paradigm for both text and visual models:

[0039] Basic model initialization and parameter freezing: Load the original weights of RoBERTa-Large / ViT-L / 14 pre-trained on a general dataset, freeze the parameters of the bottom 18 layers (text) / 20 layers (visual) of the model, and only open the parameters of the top 6 layers (text) / 4 layers (visual) and the newly added task head for training to prevent the model from overfitting and retain general semantic / visual understanding capabilities.

[0040] Domain-specific labeled dataset preprocessing: 1.2 million e-commerce advertisement texts were segmented, denoised, and labeled (the weight level of each marketing selling point and the overall sentiment polarity were labeled); 800,000 commercial advertisement images were normalized in size, data augmented (random cropping, color fine-tuning) and labeled (color tone, texture, layout, scene atmosphere category were labeled), and divided into training set, validation set and test set in an 8:1:1 ratio.

[0041] Task-specific output head setup: Add a domain task head to the output of the original model encoder: Add two fully connected layers to the text side, corresponding to the marketing selling point weight prediction task and sentiment classification task respectively; Add four fully connected layers to the visual side, corresponding to the color tone, texture, layout and scene atmosphere classification tasks respectively.

[0042] Multi-task joint training: Training is performed using a joint loss function of "main task + auxiliary task". The main task is the semantic / visual feature reconstruction loss, and the auxiliary task is the loss of the newly added classification / prediction task. The AdamW optimizer is used, and a relatively low learning rate (1e) is set. -5 Train for 3-5 epochs, and evaluate the model performance on the validation set after each training round.

[0043] Model selection and feature extractor export: When the accuracy of the auxiliary task on the validation set reaches the preset threshold (text ≥ 92%, visual ≥ 88%), training is stopped, all task-specific output heads are removed, and only the encoder part is retained as the system's text / visual feature extractor, outputting a 1024-dimensional semantic vector or style vector.

[0044] An adaptive weighted concatenation algorithm based on advertising feature dimensions is adopted. A scene classifier (a binary classification network based on a fusion of ResNet-18 and TextCNN) determines whether the input advertisement belongs to one of three scenarios: e-commerce detail page, short video feed, or outdoor poster. Then, according to a predefined scene-weight mapping table, the weight coefficients of text and visual features are automatically adjusted: text weight 0.6 and visual weight 0.4 for e-commerce detail page scenarios; text weight 0.4 and visual weight 0.6 for short video feed scenarios; and text weight 0.3 and visual weight 0.7 for outdoor poster scenarios. After weighting, the two 1024-dimensional vectors are concatenated vertically along the feature dimensions to generate a 2048-dimensional joint embedding tensor. Simultaneously, an L2 normalization constraint is added to normalize the feature values ​​within the tensor to the [-1, 1] interval.

[0045] Dual-path cross-attention coordination module: The input end is connected to the joint embedding tensor, and the text generation path and image generation path are run simultaneously. A cross-attention module is set between the corresponding transformation layers of the two paths. The intermediate activation value of one path is used as the query and the intermediate activation value of the other path is used as the key and value. The cross-modal features are exchanged and superimposed back to the original path, and the text feature matrix and visual feature map are output.

[0046] like Figure 3 As shown, it should be specifically noted that in the dual-path cross-attention coordination module, both the text generation path and the image generation path are equipped with 8 progressive feature transformation layers. Each layer consists of a self-attention sub-layer, a feedforward neural network sub-layer, and a cross-attention sub-layer. In the cross-attention sub-layer, the intermediate activation value of one path is used as the query and the intermediate activation value of the other path is used as the key and value. After being mapped to the unified attention feature space through independent linear projection layers, the cross-modal attention output is calculated and superimposed back to the original path output with a residual coefficient of 0.3.

[0047] In the dual-path cross-attention coordination module, the number of attention heads is set to 8, and a sparse attention mechanism is introduced: the marketing selling point feature location in the text and the product main feature region in the image are identified by the pre-trained advertising core region detection model, and full attention calculation and sliding window attention calculation are performed respectively.

[0048] It should be further explained that this module is the core collaborative interaction module of the system, which constructs a dual-path symmetrical cross-attention mechanism for advertising, and realizes bidirectional real-time interaction and complementary correction of textual and visual modal features.

[0049] The input interface connects to the 2048-dimensional joint embedding tensor output by the multimodal joint embedding coding module, and includes a fully symmetrical text generation path and image generation path. Both paths are configured with 8 progressive feature transformation layers. Each transformation layer consists of a self-attention sublayer, a feedforward neural network sublayer, and a cross-attention sublayer. The corresponding transformation layers of the two paths achieve bidirectional feature interaction through the cross-attention sublayer.

[0050] The intermediate activation values ​​of the i-th layer of the text generation path are used as the query matrix. The intermediate activation values ​​of the i-th layer corresponding to the image generation path are used as the key matrix. Sum matrix After each of the three elements is mapped to a unified attention feature space through its own independent linear projection layer, the cross-modal attention output is calculated. The residual coefficient is then superimposed back onto the output of the i-th layer of the text generation path, and used as a query matrix. Simultaneously, the intermediate activation values ​​of the i-th layer of the image generation path are used as the query matrix. The intermediate activation values ​​of the i-th layer corresponding to the text generation path are used as the key matrix. Sum matrix After mapping through their respective independent linear projection layers, the cross-modal attention output is calculated. Similarly, the residual coefficient of 0.3 is superimposed back onto the output of the i-th layer of the image generation path. The residual coefficient of 0.3 was determined through comparative experiments on 100,000 sets of advertising data, which can avoid modal contamination while ensuring the feature interaction effect.

[0051] The original number of attention heads was 16, which was compressed to 8 heads. A sparse attention mechanism was introduced: a dual-path cross-attention coordination module uses a pre-trained advertising core region detection model to identify the feature locations corresponding to marketing points in the text and the feature regions corresponding to the main product in the image. The detection model uses a Faster R-CNN network, taking the advertising image and text sequence as input, and outputting the bounding box coordinates of the main product in the image and the position indices of the marketing point keywords in the text sequence. The model was trained on 5000 labeled advertising images and 20,000 labeled advertising texts. Full attention calculation is performed only on the core region features, while a sliding window attention calculation is used for the background region features, with the window size set to 7×7.

[0052] The dual-path cross-attention coordination module ultimately outputs a text feature matrix with a dimension of 1024×128 (where 128 is the maximum length of the text) and a visual feature map with a dimension of 512×512×32.

[0053] Collaborative loss closed-loop control module: Projects the text feature matrix into a semantic vector sequence, projects the visual feature map into a semantic vector set in the same space, calculates the element-wise similarity matrix to generate a collaborative loss scalar, and adjusts the weight matrix of all cross-attention modules in the dual-path cross-attention coordination module in reverse to optimize the system.

[0054] like Figure 4 As shown, it should be specifically noted that in the collaborative loss closed-loop control module, the text feature matrix is ​​projected into a sequence of 128 semantic vectors with a dimension of 256, and the visual feature map is projected into a set of 32 semantic vectors with a dimension of 256, V; the semantic similarity matrix M is calculated.

[0055] It should be further explained that this module is the core of the system's adaptive optimization. It constructs a cross-modal semantic alignment collaborative loss function and adjusts the weights of the dual-path cross-attention module in real time through backpropagation to achieve deep alignment between text and visual semantics.

[0056] For a 1024×128 text feature matrix, a 1×1 convolutional layer and an average pooling layer are used to reduce the dimensionality of each word vector, resulting in a sequence of 128 semantic vectors with a dimension of 256. , arrive The core semantic vectors correspond to the first 128 words in the original advertising copy. 128 is the maximum token length set by this invention; if the original copy has fewer than 128 words, it will be padded with zero vectors at the end. For a 512×512×32 visual feature map, a global average pooling layer is used to convert each channel feature map into a single vector, resulting in a set of 32 semantic vectors with a dimension of 256. , arrive The global semantic vectors of channels 1-32 in the visual feature map are represented by 32, which is the number of visual semantic channels set by this invention. Each channel is responsible for capturing a specific type of visual semantic feature. Through the above projection operation, the text features and visual features are mapped to the same 256-dimensional semantic space.

[0057] For each vector in the semantic vector sequence S and each vector in the semantic vector set V Semantic similarity is calculated using cosine similarity, resulting in a 128×32 semantic similarity matrix M. The specific formula is as follows:

[0058]

[0059] in, Representative vector with vector semantic similarity, The index representing the semantic vector sequence S. The index representing the semantic vector set V. Representative vector with vector The dot product (inner product) of . Representative vector The L2 norm (Euclidean norm). Representative vector The L2 norm (Euclidean norm).

[0060] The collaborative loss function of the present invention It consists of two parts: semantic alignment loss and diversity loss.

[0061]

[0062] in, , In this embodiment, the weighting coefficient is... Set to 0.8, Set to 0.2, , The parameter combinations were optimized using a grid search method on 50,000 sets of labeled commercial advertising verification data. The semantic alignment loss is calculated using the following formula:

[0063]

[0064] The diversity loss, used to avoid homogenization of output results, is calculated using the following formula:

[0065]

[0066] The calculated collaborative loss scalar The weight matrices of all eight cross-attention modules in the dual-path cross-attention coordination module are adjusted in reverse using the backpropagation algorithm. The optimizer is AdamW, and the learning rate is set to 1e. -5 The weight decay coefficient is set to 1e -4 Each time this module generates a preliminary set of text and visual features, it performs a collaborative loss calculation and weight update, forming a closed-loop optimization mechanism, until the collaborative loss is lower than the preset threshold of 0.15 or the number of iterations reaches 5.

[0067] Cross-modal style verification module: Receives optimized text features and optimized visual feature map, extracts global color histogram and texture statistics from the visual map, extracts sentiment polarity vector and style description vector from the text, calculates the first distance between the color histogram and sentiment polarity, and the second distance between the texture statistics and style description. If any distance exceeds the preset threshold, a re-optimization signal is triggered and sent back to the collaborative loss closed-loop control module.

[0068] It should be specifically noted that when the cross-modal style verification module extracts the global color histogram vector H of the visual feature map, it evenly divides the 0-255 grayscale values ​​of the RGB three channels into 16 intervals and generates a histogram with a dimension of 4096 by counting the number of pixels in 4096 color intervals. When extracting the texture statistics vector B, the gray-level co-occurrence matrix method is used, with the grayscale level set to 256, step size 1, and angles of 0°, 45°, 90°, and 135°. The average values ​​of four indicators, namely contrast, correlation, energy, and entropy, are calculated to obtain a 4-dimensional texture statistics vector.

[0069] When the cross-modal style verification module extracts the sentiment polarity vector E of the copy, it uses the BERT sentiment analysis model. After fine-tuning on a dataset of 200,000 advertising copy samples, it outputs a 3-dimensional vector corresponding to the probabilities of positive, neutral, and negative, respectively; and extracts the style description vector S. ty When using a combination of keyword matching and semantic analysis, a 5-dimensional vector is output, corresponding to the probabilities of five styles: minimalist, luxurious, cute, technological, and natural. When calculating the first distance between the color histogram and the sentiment polarity, and the second distance between the texture statistics and the style description, Mahalanobis distance is used, with a preset threshold of 2.5 for the first distance and a preset threshold of 1.8 for the second distance.

[0070] It should be further explained that this module is the core of the system's quality control, and proposes a two-dimensional quantitative verification method for advertising style, which quantitatively evaluates the consistency of advertising copy and visual style from two dimensions: color-emotion and texture-style.

[0071] For the optimized 512×512×32 visual feature map, after decoding into an RGB image, the continuous gray values ​​of 0-255 for each of the three RGB channels are evenly divided into 16 equal-width intervals (each interval consists of 16 consecutive values, such as 0-15 being interval 0, 16-31 being interval 1, and so on). All 512×512=262144 pixels in the image are scanned one by one, and the color interval to which the RGB value of each pixel belongs is counted. The number of pixels in each interval is accumulated, and the pixel statistics of the 4096 color intervals are arranged in a fixed order to generate a global color histogram vector H with a dimension of 4096. The RGB image was converted to a 256-level grayscale image using a Gray-Level Co-occurrence Matrix (GLCM), with each pixel's grayscale value ranging from 0 to 255. This eliminated the interference of color information on texture analysis. The number of grayscale levels was set to 256, the step size to 1, and the angles to 0°, 45°, 90°, and 135°. For each angle, a 256×256 GLCM was generated. The value of the (i,j)th element in the matrix represents the number of times the condition "a pixel has a grayscale value of i, and its adjacent pixel in the specified direction has a grayscale value of j" occurs in the image. The average values ​​of four indicators—contrast, correlation, energy, and entropy—were calculated at the four angles to obtain a 4-dimensional texture statistics vector B.

[0072] Among them, contrast measures the sharpness and depth of image texture, and is calculated as the weighted sum of the squares of the pixel value differences; correlation measures the linear correlation of image texture, and is calculated as the ratio of the covariance to the standard deviation of pixel values; energy measures the uniformity and regularity of image texture, and is calculated as the sum of the squares of matrix elements; entropy measures the complexity and randomness of image texture, and is calculated as the negative of the sum of the sum of element values ​​multiplied by the logarithms of the elements.

[0073] The optimized 1024×128 copy feature matrix needs to be restored to the original text sequence. Specifically, the system calls the inverse process of the GPT-2 decoder in the alignment output module: each 1024-dimensional feature vector in the copy feature matrix is ​​input into an independent linear inverse mapping layer, which outputs the corresponding 768-dimensional word embedding vector. This vector is then approximately restored to a token index sequence by the original BERT word embedding decoder (i.e., the pseudo-inverse of the embedding matrix corresponding to the BERT pre-trained vocabulary). Finally, it is decoded into a readable text string by the tokenizer. The inverse mapping layer is initialized with the transpose of the same projection matrix as during GPT-2 fine-tuning and trained on 100,000 advertising copy-feature pairing data points until the reconstruction perplexity is below 20. Then, the restored text string is input into the BERT sentiment analysis model (which has been fine-tuned on a dataset of 200,000 advertising copy data points) to extract the sentiment polarity vector E, which has a dimension of 3 and corresponds to the probability values ​​of positive, neutral, and negative sentiments, respectively.

[0074] Meanwhile, the cross-modal style verification module outputs a 5-dimensional style description vector S. ty These correspond to five advertising styles: minimalist, luxurious, cute, technological, and natural. The specific implementation steps are as follows:

[0075] Build a style keyword library, and create a style-specific keyword dictionary for each of the five styles:

[0076] Keywords for minimalist style: {"minimalist", "simple", "clean", "white space", "refreshing", "extremely minimalist"};

[0077] Keywords for luxury style: {"luxury", "high-end", "prestige", "refined", "precious", "elegant"};

[0078] Keywords for cute style: {"cute", "adorable", "girlish", "fun", "cartoonish", "sweet"};

[0079] Tech-themed keywords: {"technology", "intelligence", "futuristic", "cutting-edge technology", "digital", "AI"};

[0080] Keywords for natural style: {"natural", "natural", "environmentally friendly", "organic", "pristine", "fresh"}.

[0081] Design a hybrid scoring function to calculate the raw score Rc for each style c of the input text (text string):

[0082]

[0083] in, The number of style-related keywords that are matched in the copy; Total word count of the copy (normalization factor); The semantic similarity score is calculated as the cosine similarity between the overall embedding vector of the text (obtained by the RoBERTa-Large encoder) and the centroid vector of style c (pre-calculated from the average embedding of all training texts of style c). , As the weighting coefficient, in this embodiment =0.6, =0.4.

[0084] Original score The style description vector S is converted to a 5-dimensional style description vector using the Softmax function. ty The sum of all components is 1.

[0085] Furthermore, the fine-tuning steps for the BERT sentiment analysis model to extract the sentiment polarity vector E are as follows:

[0086] Load the pre-trained BERT-base Chinese model, freeze the parameters of the bottom 8 encoder layers (to retain general Chinese semantic understanding capabilities), and only open the parameters of the top 4 encoder layers and the newly added classification head for training to prevent small sample overfitting.

[0087] Using a dataset of 200,000 labeled e-commerce advertising copy, each copy was labeled with three categories: "positive / neutral / negative" (e.g., "limited-time offer" is positive, "no returns or exchanges for non-quality issues" is neutral, and "easily damaged" is negative), and divided into training set, validation set and test set in a ratio of 8:1:1.

[0088] The original pre-training task head of the BERT model is removed, and a fully connected classification layer is added to map the 768-dimensional semantic vector output by BERT into a 3-dimensional output vector. Then, the vector is normalized to a probability value (the sum of the three values ​​is 1) through the Softmax function, corresponding to the probabilities of positive, neutral, and negative emotions.

[0089] We employ the cross-entropy loss function, use the AdamW optimizer, and set the learning rate to 2e. -5 The training process is repeated for 3 epochs. After each training epoch, the accuracy is evaluated on the validation set. Training is stopped early when the accuracy on the validation set no longer improves for two consecutive epochs.

[0090] like Figure 5 As shown, this invention constructs an advertising style mapping database, which is trained from 100,000 labeled commercial advertising data. This database includes mappings from color histograms to sentiment polarity and from texture statistics to style descriptions. Based on the advertising style mapping database, the first distance D1 between the color histogram vector H and the sentiment polarity vector E, and the distance between the texture statistics vector B and the style description vector S are calculated. ty The second distance D2 between them is calculated using Mahalanobis distance to eliminate the dimensional differences between different feature dimensions.

[0091] The specific formula for calculating the Mahalanobis distance from the histogram vector H of the color to be detected to the corresponding distribution is as follows:

[0092]

[0093] The specific formula for calculating the Mahalanobis distance from the texture statistics vector B to the corresponding distribution is as follows:

[0094]

[0095] in, , Both represent the mean vector of the target distribution. The transpose operator is used for vectors. In this embodiment, the preset threshold for the first distance is set to 2.5, and the preset threshold for the second distance is set to 1.8. If D1 > 2.5 or D2 > 1.8, a re-optimization signal is triggered. The specific dimension information of style inconsistency (such as "red tones do not match negative emotions" or "rough textures do not match luxurious styles") is fed back to the collaborative loss closed-loop control module. In the next iteration, the loss weight of the corresponding dimension is increased (the weight coefficient increases by 0.5), guiding the generation of more consistent advertising content. If both distances are below the preset threshold, the verification passes, and the alignment output stage begins.

[0096] Alignment Output Module: Decodes the final text features into a text stream, decodes the final visual feature map into a layered visual data package, and aligns the text paragraphs with the corresponding visual layers in terms of timestamps.

[0097] It should be specifically noted that when the alignment output module decodes the final copywriting features into a text stream, it uses a GPT-2 decoder with fine-tuning. A linear projection layer is added at the input end to map the 1024-dimensional features to a 768-dimensional word embedding space, and three parallel output heads are added: a text generation head that outputs the probability distribution of the next token, a semantic tag head that outputs the module tags, and a timestamp head that outputs the start and end timestamps. The decoded text stream includes four parts: title, body, selling points, and call to action, each with semantic tags and timestamp information.

[0098] It should be further noted that this module is the system output end, realizing the spatiotemporal alignment and output of advertising content, and supporting advertising needs on multiple platforms.

[0099] For the final copy feature matrix, a finely tuned GPT-2 decoder based on a dataset of 500,000 advertising copy samples was used for decoding to generate a structured text stream. The text stream includes four parts: title, body text, selling points, and call to action, each with corresponding semantic tags and timestamp information. For the final visual feature map, a layered decoder was used to decode it into a layered visual data package. The data package includes four independent layers: background layer, product body layer, text layer, and decorative element layer. Each layer has corresponding position coordinates, transparency, and display duration information.

[0100] Furthermore, the fine-tuning steps of the GPT-2 decoder are as follows: load the pre-trained Chinese GPT-2 model, freeze the parameters of the bottom 12 layers (to retain the general language generation capability), add a linear projection layer at the input of the model, and map the 1024×128 text feature matrix (1024-dimensional features at each position) output by the alignment output module to the native 768-dimensional word embedding space of GPT-2, so as to achieve a seamless connection between cross-modal features and language generation space.

[0101] Using a dataset of 500,000 well-annotated e-commerce advertising copy, each copy is divided into four independent modules according to advertising logic: title, body, selling points, and call to action. The standard display duration for each module is also labeled, and the dataset is divided into training set, validation set, and test set in an 8:1:1 ratio.

[0102] Remove the original single-task text generation header of GPT-2 and add three parallel output headers:

[0103] Text generation header: Outputs the probability distribution of the next token, used to generate advertising copy content;

[0104] Semantic tag header: Outputs the module tag (title / body / selling point / call to action) corresponding to the current text fragment;

[0105] Timestamp header: Outputs the start and end timestamps of the current text segment.

[0106] Training is performed using a weighted multi-task loss function, with the total loss calculated as: 0.7 × text generation cross-entropy loss + 0.2 × semantic label classification loss + 0.1 × timestamp regression loss. The AdamW optimizer is used, with a learning rate of 5e. -5 Train for 4 epochs. Stop training when the text generation perplexity on the validation set is below 15 and the semantic label accuracy is above 95%, and export the trained decoder model.

[0107] Based on the semantic tags and timestamp information of the text stream, the corresponding visual layers and text paragraphs are spatiotemporally aligned: the title corresponds to the display time of the product main layer, the body text corresponds to the switching time of the background layer, the selling points correspond to the pop-up time of the decorative element layer, and the call to action corresponds to the display time of the text layer. In this embodiment, the total duration of the short video advertisement is set to 15 seconds, with the title display time being 0-3 seconds, the body text display time being 3-10 seconds, the selling points display time being 5-12 seconds, and the call to action display time being 10-15 seconds.

[0108] The alignment output module supports three output formats, each adapted to different advertising platforms, as detailed below:

[0109] Static poster format: Output a composite image in PNG format, and simultaneously output the corresponding text file in TXT format;

[0110] Short video format: Outputs 15-second short videos in MP4 format, with video resolution options of 720P, 1080P, and 4K.

[0111] Interactive Ad Format: Outputs interactive ad data packages in JSON format, including all layered visual elements and text content, supporting different interactive effects triggered by users clicking on different elements.

[0112] In addition, the alignment output module will generate an advertising performance evaluation report, including quantitative indicators such as image and text semantic similarity score, style consistency score, and target audience suitability score, providing advertisers with decision-making references.

[0113] Secondly, the accompanying drawings of the embodiments disclosed in this invention only involve structures related to the embodiments disclosed in this invention. Other structures can refer to general designs. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other.

[0114] Finally, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A collaborative design system for advertising copywriting and visuals that integrates a multimodal large model, characterized in that, include: Multimodal joint embedding coding module: Receives product text description and reference image, extracts text semantic vector and visual style vector, and concatenates them along the feature dimension to generate joint embedding tensor; Dual-path cross-attention coordination module: The input end is connected to the joint embedding tensor, and the text generation path and image generation path are run simultaneously. A cross-attention module is set between the corresponding transformation layers of the two paths. The intermediate activation value of one path is used as the query and the intermediate activation value of the other path is used as the key and value. The cross-modal features are exchanged and superimposed back to the original path, and the text feature matrix and visual feature map are output. Collaborative loss closed-loop control module: Projects the text feature matrix into a semantic vector sequence, projects the visual feature map into a semantic vector set in the same space, calculates the element-wise similarity matrix to generate a collaborative loss scalar, and adjusts the weight matrix of all cross-attention modules in the dual-path cross-attention coordination module in reverse to optimize the process; Cross-modal style verification module: Receives optimized copy features and optimized visual feature map, extracts global color histogram and texture statistics from visual map, extracts sentiment polarity vector and style description vector from copy, calculates the first distance between color histogram and sentiment polarity, and the second distance between texture statistics and style description. If any distance exceeds the preset threshold, a re-optimization signal is triggered and sent back to the collaborative loss closed-loop control module. Alignment Output Module: Decodes the final text features into a text stream, decodes the final visual feature map into a layered visual data package, and aligns the text paragraphs with the corresponding visual layers in terms of timestamps.

2. The advertising copywriting and visual collaborative design system integrating a multimodal large model as described in claim 1, characterized in that: The multimodal joint embedding coding module includes a text preprocessing unit, which uses a TF-IDF keyword filtering algorithm to remove words in the product text description whose information gain value is less than or equal to a set threshold, and retains valid keywords whose information gain value is greater than the set threshold.

3. The advertising copywriting and visual collaborative design system integrating a multimodal large model as described in claim 1, characterized in that: The multimodal joint embedding encoding module also includes a text encoder and an image encoder. The text encoder uses the RoBERTa-Large model, which is fine-tuned using an e-commerce advertising text dataset to output a text semantic vector. The image encoder uses the ViT-L / 14 model, which is fine-tuned using a commercial advertising image dataset to output a visual style vector. The fine-tuning process includes freezing the underlying parameters, adding a task-specific output head, multi-task joint training, and model selection steps.

4. The advertising copywriting and visual collaborative design system integrating a multimodal large model as described in claim 1, characterized in that: The multimodal joint embedding encoding module adopts an adaptive weighted concatenation algorithm based on advertising feature dimensions. It adjusts the weight coefficients of the text semantic vector and the visual style vector according to the advertising scenario type. After weighting, the two vectors are concatenated along the feature dimension to form a joint embedding tensor, and L2 normalization constraint is applied to the interval [-1,1].

5. The advertising copywriting and visual collaborative design system integrating a multimodal large model as described in claim 1, characterized in that: In the dual-path cross-attention coordination module, both the text generation path and the image generation path are equipped with progressive feature transformation layers. Each layer consists of a self-attention sub-layer, a feedforward neural network sub-layer, and a cross-attention sub-layer. In the cross-attention sub-layer, the intermediate activation value of one path is used as the query and the intermediate activation value of the other path is used as the key and value. After being mapped to a unified attention feature space through independent linear projection layers, the cross-modal attention output is calculated and superimposed back to the original path output with a set residual coefficient.

6. The advertising copywriting and visual collaborative design system integrating a multimodal large model as described in claim 1, characterized in that: The dual-path cross-attention coordination module introduces a sparse attention mechanism: it identifies the marketing selling point feature locations in the text and the product main feature regions in the image through a pre-trained advertising core region detection model, and performs full attention calculation and sliding window attention calculation respectively.

7. The advertising copywriting and visual collaborative design system integrating a multimodal large model as described in claim 1, characterized in that: The collaborative loss closed-loop control module projects the text feature matrix into a semantic vector sequence S and the visual feature map into a semantic vector set V; it then calculates the semantic similarity matrix M.

8. The advertising copywriting and visual collaborative design system integrating a multimodal large model according to claim 1, characterized in that: When the cross-modal style verification module extracts the global color histogram vector H of the visual feature map, it evenly divides the gray values ​​of the RGB three channels into intervals and counts the number of pixels in each color interval to generate a histogram. When extracting the texture statistics vector B, the gray-level co-occurrence matrix method is used. The gray-level level, step size, and angle are set, and the average value of four indicators, namely contrast, correlation, energy, and entropy, is calculated to obtain the texture statistics vector.

9. The advertising copywriting and visual collaborative design system integrating a multimodal large model according to claim 1, characterized in that: When the cross-modal style verification module extracts the sentiment polarity vector E of the copy, it uses the BERT sentiment analysis model. After fine-tuning on the advertising copy dataset, it outputs a 3-dimensional vector corresponding to the probabilities of positive, neutral, and negative, respectively; and extracts the style description vector S. ty At that time, a method combining keyword matching and semantic analysis is used to output a 5-dimensional vector, which corresponds to the probability of five styles: minimalist, luxurious, cute, technological, and natural. When calculating the first distance between the color histogram and sentiment polarity, and the second distance between texture statistics and style description, Mahalanobis distance is used, and the first distance threshold and the second distance threshold are preset.

10. The advertising copywriting and visual collaborative design system integrating a multimodal large model according to claim 1, characterized in that: When the alignment output module decodes the final copy features into a text stream, it uses a GPT-2 decoder and fine-tunes it. A linear projection layer is added at the input end to map the features to the word embedding space, and three parallel output heads are added: a text generation head that outputs the probability distribution of the next token, a semantic tag head that outputs the module tags, and a timestamp head that outputs the start and end timestamps. The decoded text stream includes four parts: title, body, selling points, and call to action. Each part carries semantic tags and timestamp information.