An advertisement creative generation method and device based on artificial intelligence, and a storage medium

By determining the brand feature vector set and atmosphere descriptors, and combining image generation models and saliency detection, advertising materials that conform to the brand style and are visually prominent are generated. This solves the problem of complex background textures or bright colors in existing technologies and achieves high-quality automated advertising creative generation.

CN122244216APending Publication Date: 2026-06-19BEIJING 180CHINA ADVERTISING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING 180CHINA ADVERTISING CO LTD
Filing Date
2026-04-14
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing ad generation technologies cannot effectively control the visual hierarchy of the image, resulting in the main product being submerged in the background, failing to highlight the core target of the advertisement, and generating background textures that are too complex or colors that are too vibrant, failing to meet the standards for ad placement.

Method used

By acquiring raw material data input by users, we determine the brand feature vector set and atmosphere description words, generate advertising copy using a natural language model, and generate an adapted background image in the target graphic template using an image generation model. We then combine saliency detection to calculate the saliency ratio and select candidate creative advertising images that meet the visual prominence criteria.

Benefits of technology

Ensure that the style of the generated ad creatives is consistent with the brand tone, improve the richness and click appeal of the ad content, achieve high-quality automated delivery, eliminate low-quality creatives with distracting backgrounds, and meet the standards for ad placement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122244216A_ABST
    Figure CN122244216A_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology, and in particular to an AI-based method, device, and storage medium for generating advertising creatives. The method includes acquiring raw material data comprising an original product image, brand logo, specified color scheme, and original keywords; determining a brand feature vector set based on the brand logo and specified color scheme; semantically parsing the original keywords to generate multiple sets of candidate advertising copy; acquiring a target image template and determining the product's main body area and background area; generating a suitable background image using an image generation model; fusing the suitable background image, the product's main body area, the candidate advertising copy, and the brand logo to obtain several sets of candidate creative advertising images; calculating the saliency ratio between the product's main body area and the background area in the candidate creative advertising images; and selecting candidate creative advertising images that meet preset output conditions based on the saliency ratio. This application improves the quality of generated advertising images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, device and storage medium for generating advertising creative ideas based on artificial intelligence. Background Technology

[0002] Currently, with the rapid development of e-commerce and mobile internet advertising, the demand for creative advertising images for single products in various scenarios has increased dramatically. Advertising creative generation technology aims to use computer automation algorithms to replace manual design, quickly integrate the main product with the generated background material, and mass-produce advertising images suitable for different advertising channels.

[0003] Existing ad generation technologies typically use AI models to directly generate a complete image containing the product and background based on the input text description, or generate a background image first and then paste the product in. However, this generation method lacks effective control over the visual hierarchy of the image. In pursuit of richness and artistry, AI models often generate backgrounds with overly complex textures or overly vibrant colors, causing the product to be visually overwhelmed by the background after integration. This fails to highlight the core promoted object of the advertisement, does not meet advertising placement standards, and therefore has room for improvement. Summary of the Invention

[0004] This application provides an artificial intelligence-based advertising creative generation method, device, and storage medium, which can improve the quality of generated advertising images.

[0005] The above-mentioned objective of this application is achieved through the following technical solution:

[0006] An artificial intelligence-based method for generating advertising creatives, the method comprising:

[0007] Obtain the original material data input by the user, which includes at least the original product image, brand logo, specified color scheme, and original keywords;

[0008] Determine the brand feature vector set based on the brand logo and the specified color scheme;

[0009] Semantic parsing is performed on the original keywords to identify product attribute words and atmosphere description words. Based on the product attribute words and atmosphere description words, multiple sets of candidate advertising copy are generated using a natural language model.

[0010] In response to an ad generation request, a target image and text template is obtained, the original product image is mapped to the target image and text template, and the main product area and background area are determined.

[0011] Initialize the image generation model, and obtain an adapted background image based on the brand feature vector set, the atmosphere descriptive words, and the background region;

[0012] The adapted background image, the main product area, the candidate advertising copy, and the brand logo are integrated to obtain several sets of candidate creative advertising images;

[0013] The candidate creative advertisement image is subjected to saliency detection, and the saliency ratio of the main product area to the background area is calculated.

[0014] Based on the significance ratio, candidate creative advertising images that meet the preset output conditions are selected as available candidate creative advertising images for output.

[0015] By adopting the above technical solutions, and by determining the brand feature vector set based on the brand logo and the specified color tone, it is possible to ensure that the subsequently generated advertising materials maintain a high degree of consistency with the brand tone in terms of style and color, avoiding deviation from brand norms. By using natural language models to generate multiple sets of candidate advertising copy, the dull product attributes can be transformed into creative copy with marketing appeal, thereby improving the richness of advertising content and click appeal. By mapping the original product image to the target image and text template to determine the main product area and background area, the operational boundaries of the image generation algorithm can be clarified, thereby providing an accurate canvas space for model generation while ensuring the integrity of the main product. By performing saliency detection on the candidate creative advertising images and calculating the saliency ratio for screening, the visual prominence of the main product relative to the generated background can be quantitatively evaluated, thereby automatically eliminating low-quality materials with backgrounds that are distracting and interfere with user attention, achieving high-quality automated delivery of advertising creatives.

[0016] In a preferred embodiment, this application can be further configured such that: determining the brand feature vector set based on the brand identifier and the specified color tone specifically includes:

[0017] A lightweight residual network is invoked to extract features from the brand logo, resulting in a contour feature vector and a color proportion feature vector.

[0018] Obtain the RGB value of the specified hue, and map the RGB value into a hue control vector for the HSV color space;

[0019] The brand feature vector set is constructed by fusing the contour feature vector, the color proportion feature vector, and the hue control vector.

[0020] By employing the above technical solutions, a lightweight residual network is used to extract features from the brand logo to obtain contour feature vectors and color proportion feature vectors. This captures the geometric complexity and color distribution patterns of the brand logo, providing fine-grained visual fingerprint information for the generation model. By mapping RGB values ​​to hue control vectors in the HSV color space, the perceptual attributes of color, such as warmth and coolness, and saturation, can be expressed more intuitively, thus guiding the generation model to generate brand color backgrounds that conform to human visual perception. By fusing the above features to construct a brand feature vector set, a unified and comprehensive style constraint interface can be formed, enabling precise control over both background texture and color during the generation process.

[0021] In a preferred embodiment, this application can be further configured as follows: in response to an advertisement generation request, obtaining a target image and text template, mapping the original product image to the target image and text template, and determining the main product area and background area, specifically including:

[0022] Based on the ad generation request, the scene size parameters are obtained, and the corresponding target image and text template is called from the template library according to the scene size parameters;

[0023] By using an image segmentation algorithm, the boundaries of foreground objects in the original product image are identified to obtain the main product image;

[0024] Align the main product image to the product placeholder coordinates in the target graphic template, and determine the area covered by the product placeholder coordinates as the main product area;

[0025] The area in the target graphic template other than the main product area is defined as the background area.

[0026] By adopting the above technical solution, and by obtaining scene size parameters based on the ad generation request and calling the corresponding target image and text template, it is possible to quickly adapt to the size specifications of different advertising channels, thereby achieving efficient production with one-time input and multi-terminal adaptation; by using image segmentation algorithms to identify the boundaries of foreground objects to obtain the main product image, it is possible to achieve automated and accurate image cutout, thereby removing interference from the original shooting background and ensuring clean and smooth product edges; by aligning the main product image to the product placement coordinates and determining the placement area as the main product area, it is possible to prevent the generated background elements from obscuring or encroaching on the core promoted object from a spatial logic perspective.

[0027] In a preferred example, this application can be further configured as follows: obtaining an adapted background image based on the brand feature vector set, the atmosphere descriptors, and the background region through the image generation model specifically includes:

[0028] Style parameters and semantic parameters are generated based on the brand feature vector set and the atmosphere descriptive words, respectively.

[0029] The style parameters and semantic parameters are input into the image generation model through a cross-attention mechanism, so that the image generation model can redraw pixels in the background area of ​​the target image template to obtain the adapted background image.

[0030] By adopting the above technical solution, style parameters and semantic parameters are generated based on the brand feature vector set and atmosphere descriptive words, respectively, thereby controlling the generated artistic style and content semantics. By inputting these parameters into the image generation model for pixel redrawing through a cross-attention mechanism, the model can be dynamically guided to focus on key features and generate content only in the background area, thereby synthesizing an adapted background image that not only conforms to the brand aesthetics and enhances the copywriting atmosphere, but also seamlessly integrates with the product subject.

[0031] In a preferred example, this application can be further configured such that the construction of the image generation model specifically includes:

[0032] Obtain a generative adversarial network model trained on a generative adversarial network, and construct a training sample set, which includes several pairs of sample background images, sample brand feature vectors, and sample atmosphere descriptive words;

[0033] The sample brand feature vector is input into the mapping network of the generative adversarial network model to obtain sample style parameters, and the sample semantic parameters are obtained based on the sample atmosphere descriptive words.

[0034] The sample style parameters and sample semantic parameters are input into the generator of the generative adversarial network model, and the predicted background image is output.

[0035] The predicted background image and the sample background image are input into the discriminator of the generative adversarial network model to obtain the discrimination result;

[0036] Based on the difference between the discrimination result and the true label, an adversarial loss function is calculated, and based on the feature difference between the predicted background image and the sample background image, a style consistency loss function is calculated.

[0037] Based on the adversarial loss function and the style consistency loss function, the network parameters of the generator and the network parameters of the discriminator are alternately updated using the backpropagation algorithm until the adversarial loss function and the style consistency loss function meet the corresponding preset conditions, and the adversarial network model is output to obtain the image generation model.

[0038] By adopting the above technical solution, and constructing a paired training sample set containing sample background images, sample brand feature vectors, and sample atmosphere descriptive words, the model can be provided with rich supervision information, thereby teaching the model to understand the mapping relationship between brand features and visual style. By introducing style parameters and semantic parameters and using backpropagation algorithm to alternately update network parameters based on adversarial loss function and style consistency loss function, the generator can be forced to maintain style without deviation while deceiving the discriminator, thereby training an image generation model with high-quality texture generation capability and strong style controllability.

[0039] In a preferred embodiment, this application can be further configured such that: the saliency detection of the candidate creative advertisement image, and the calculation of the saliency ratio between the main product area and the background area, specifically includes:

[0040] The candidate creative advertisement images are visually saliency detected using a visual saliency analysis algorithm to obtain a saliency heatmap.

[0041] Calculate the mean saliency of pixels located within the main product area in the saliency heatmap, and denot it as the main focus score;

[0042] Obtain the edge contour data of the main body area of ​​the product;

[0043] For each background pixel in the background region defined by the salient heatmap, calculate the shortest spatial Euclidean distance between the corresponding coordinate position and the pixel in the edge contour data.

[0044] Obtain the preset maximum interference distance parameter, and calculate the difference between the maximum interference distance parameter and the shortest spatial distance to obtain the distance difference.

[0045] The ratio of the distance difference to the maximum interference distance parameter is determined as the distance normalization parameter. The distance normalization parameter is multiplied by a preset reference space weight value to calculate the weighted weight corresponding to the candidate saliency pixel.

[0046] The saliency value of each background pixel in the background region in the saliency heatmap is fused with the corresponding weighting weight to obtain a spatially weighted saliency heatmap distributed in the background region.

[0047] Calculate the mean saliency of pixels in the spatially weighted saliency heatmap, and denote it as the background basic interference degree;

[0048] Obtain the maximum pixel saliency value in the spatially weighted saliency heatmap, and denote it as the background peak interference degree;

[0049] Using a preset weighting formula, the background base interference degree and the background peak interference degree are weighted and summed to obtain the comprehensive background interference index;

[0050] The significance ratio is obtained by calculating the ratio of the subject's attention to the comprehensive background interference index.

[0051] By adopting the above technical solution, the normalized weighted weight is calculated by calculating the shortest spatial Euclidean distance from each background pixel to the edge contour of the product subject and combining it with the maximum interference distance parameter. This can accurately depict the physical topological critical relationship of background noise close to the visual focus. By fusing the weighted weight with the original heatmap values, a spatially weighted significance heatmap is obtained, and the background basic interference degree and background peak interference degree are calculated accordingly. This can amplify and lock the high-risk noise values ​​close to the product edge. By weighted summing the above basic indicators and the maximum peak indicators, a comprehensive background interference index is obtained and compared with the subject attention to obtain a significance ratio. This can construct a composite-level evaluation mathematical parameter that is extremely sensitive to local abrupt highlights and overall complex layout. Thus, without human intervention, the objectivity and accuracy of quantitatively judging the risk of visual overshadowing of the subject in the composition can be greatly improved.

[0052] In a preferred embodiment, this application can be further configured such that: the step of filtering candidate creative ad images whose significance ratios satisfy preset output conditions as available candidate creative ad images output based on the significance ratio specifically includes:

[0053] Obtain a preset significance threshold, and compare the significance ratio with the significance threshold;

[0054] If the significance ratio is greater than or equal to the significance threshold, then the significance ratio is determined to meet the output condition, and the candidate creative ad image corresponding to the significance ratio is output as an available candidate creative ad image;

[0055] If the significance ratio is less than the significance threshold, then the significance ratio is determined not to meet the output condition.

[0056] By adopting the above technical solution, an automated quality control defense line can be established by comparing the salience ratio with a preset salience threshold, thereby quickly distinguishing qualified materials from unqualified ones. By determining not to output when the ratio does not meet the conditions, it is possible to effectively intercept inferior advertising images with visual logic errors, thereby significantly reducing the cost of subsequent manual screening and ensuring that every material finally delivered to the user complies with the basic principles of visual marketing.

[0057] In a preferred embodiment, this application can be further configured such that the advertising creative generation method also includes:

[0058] Obtain the number of available candidate creative ad images;

[0059] If the number of available candidate creative advertising images is less than the preset number, then adjust the atmosphere description words to obtain the adapted background image;

[0060] If the calculated significance ratio satisfies the output condition and the number of available candidate creative ad images is greater than or equal to the preset number, the final available candidate creative ad images are output.

[0061] By adopting the above technical solution, when the number of available candidate creative advertising images is insufficient, the atmospheric descriptive words are adjusted and the generation steps are repeated until the quantity reaches the target. This establishes a closed-loop backtracking and replenishment mechanism, which automatically attempts to generate backgrounds with less visual interference by reducing the semantic intensity of the descriptive words or replacing synonyms. This ensures the quality of individual images while rigidly meeting the user's demand for the quantity of materials delivered, thereby improving the system's delivery capability and robustness.

[0062] The second objective of this invention is achieved through the following technical solution:

[0063] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the aforementioned artificial intelligence-based advertising creative generation method.

[0064] The above-mentioned objective three of this application is achieved through the following technical solution:

[0065] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned artificial intelligence-based advertising creative generation method.

[0066] In summary, this application includes at least one of the following beneficial technical effects:

[0067] 1. By determining the brand feature vector set based on the brand logo and the specified color tone, it is possible to ensure that the subsequently generated advertising materials maintain a high degree of consistency with the brand tone in terms of style and color tone, avoiding deviation from brand specifications; by using a natural language model to generate multiple sets of candidate advertising copy, the dull product attributes can be transformed into creative copy with marketing appeal, thereby improving the richness of advertising content and click appeal; by mapping the original product image to the target image and text template to determine the main product area and background area, the operational boundaries of the image generation algorithm can be clarified, thereby providing a precise canvas space for model generation while ensuring the integrity of the main product; by performing saliency detection on candidate creative advertising images and calculating the saliency ratio for screening, the visual prominence of the main product relative to the generated background can be quantitatively evaluated, thereby automatically eliminating low-quality materials with backgrounds that are distracting and interfere with user attention, achieving high-quality automated delivery of advertising creatives;

[0068] 2. By constructing a paired training sample set containing sample background images, sample brand feature vectors, and sample atmosphere descriptive words, rich supervision information can be provided to the model, thereby teaching the model to understand the mapping relationship between brand features and visual style. By introducing style parameters and semantic parameters and using backpropagation algorithm to alternately update network parameters based on adversarial loss function and style consistency loss function, the generator can be forced to maintain style without deviation while deceiving the discriminator, thereby training an image generation model with high-quality texture generation capability and strong style controllability.

[0069] 3. By adjusting the mood description and repeating the generation process when the number of available candidate creative ad images is insufficient, a closed-loop backtracking and replenishment mechanism can be established. It automatically attempts to generate backgrounds with less visual interference by reducing the semantic intensity of the description or replacing synonyms. This ensures the quality of individual images while rigidly meeting the user's demand for the quantity of materials delivered, thus improving the system's delivery capability and robustness. Attached Figure Description

[0070] Figure 1 This is a flowchart illustrating the implementation of an artificial intelligence-based advertising creative generation method in one embodiment of this application.

[0071] Figure 2 This is a flowchart illustrating the implementation of an image generation model in an AI-based advertising creative generation method according to one embodiment of this application.

[0072] Figure 3 This is another implementation flowchart of an artificial intelligence-based advertising creative generation method in one embodiment of this application;

[0073] Figure 4 This is a schematic diagram of the internal structure of a computer device according to an embodiment of this application. Detailed Implementation

[0074] The following embodiments will help those skilled in the art to further understand the function of this application, but do not limit this application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of this application. These all fall within the protection scope of this application.

[0075] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0076] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0077] The present application will be further described in detail below with reference to the accompanying drawings.

[0078] In one embodiment, such as Figure 1 As shown, this application discloses an artificial intelligence-based advertising creative generation method, which specifically includes the following steps:

[0079] S10. Obtain the original material data input by the user. The original material data shall include at least the original product image, brand logo, specified color scheme and original keywords.

[0080] Specifically, raw material data refers to the various basic material information required in the process of generating advertising creatives. Among them, the original product image can be a real photo or rendered image containing the main body of the promoted product, such as a picture of a beverage bottle with a pure white background removed, or a product photo with a simple background. The brand logo usually refers to the advertiser's logo image. This logo can strengthen the brand identity in the subsequently generated advertising image. The specified color is the main color that the advertiser wants to present in the advertising image, usually given in the form of RGB or CMYK values, such as the brand's VI color, to ensure that the style of the generated advertisement matches the brand tone. The original keywords are text information entered by the user to describe the product's selling points, promotional theme, or desired scene atmosphere, such as "summer coolness" or "0 sugar, 0 calories". By obtaining this data, the necessary data foundation is provided for subsequent feature extraction and content generation.

[0081] S20. Determine the brand feature vector set based on the brand logo and specified color scheme.

[0082] Specifically, in order to understand and quantify a brand's visual style, the brand logo in image form and the specified color tone in numerical form need to be transformed into machine-readable vector features. This process involves parsing the visual attributes of the image, extracting information such as shape complexity and line features, and mapping individual color values ​​to a higher-dimensional feature space. By fusing these two types of information, a brand feature vector set that can comprehensively represent the brand's visual tone is constructed. This vector set will serve as a key constraint in the subsequent image generation process, guiding the generation model to produce background textures or color schemes that conform to the specific brand style, thereby preventing the generated advertising images from deviating from the brand specifications.

[0083] S30. Perform semantic analysis on the original keywords to identify product attribute words and atmosphere description words. Based on the product attribute words and atmosphere description words, use a natural language model to generate multiple sets of candidate advertising copy.

[0084] Specifically, the original keywords often contain a mixture of various types of information, which need to be classified and expanded using natural language processing technology. This allows information that must be retained, such as the objective specifications and promotional efforts of the product, to be identified as product attribute words, such as "buy one get one free" or "500ml". Information that can be used to describe the feeling of the scene and the emotional tone, which allows for creative expression, is identified as atmosphere description words, such as "dreamy" or "vibrant". Then, using a pre-trained natural language generation model, with product attribute words as a hard constraint (meaning that the generated copy must contain these words) and atmosphere description words as a soft guide, multiple sets of candidate copy that are easy to read and have marketing appeal are generated. For example, "vibrant" can be expanded to "releasing unlimited vitality", thereby achieving automated creation and optimization of copy.

[0085] S40. In response to the ad generation request, obtain the target image and text template, map the original product image to the target image and text template, and determine the main product area and background area.

[0086] Specifically, the ad generation request usually carries the desired output image size information, such as a 1080x1920 pixel vertical ad. Based on this size information, the most suitable layout template is matched from a pre-designed template library. This template specifies the approximate placement of each element. Then, the original product image is placed in the product display position preset in the template. In order for the subsequent generation algorithm to distinguish where the foreground is immovable and where the background needs to be drawn by the model, the boundaries need to be accurately defined. The pixel range occupied by the original product image in the template is marked as the main product area, and the remaining blank pixel range in the template left for the model to work with is marked as the background area, thereby clarifying the boundaries of the generation.

[0087] S50. Initialize the image generation model. Based on the brand feature vector set, atmosphere description words, and background area, obtain the adapted background image through the image generation model.

[0088] Specifically, after determining the generation range, i.e. the background area, the image generation model is started to draw the content. At this time, the previously constructed brand feature vector set is used as a style parameter to ensure that the color tone and texture style of the generated image do not deviate. At the same time, the atmosphere description words are used as semantic parameters to ensure that the generated image content conforms to descriptions such as "summer" and "ice and snow". Guided by these parameters, the image generation model only generates pixel data in the background area, while keeping the pixels of the main product area unchanged or treating it as an occlusion area. Finally, it outputs an adaptive background image that conforms to the brand tone and the copywriting atmosphere, and perfectly avoids the product position.

[0089] S60. Integrate the adapted background image, the main product area, the candidate advertising copy, and the brand logo to obtain several sets of candidate creative advertising images.

[0090] Specifically, this step involves assembling the generated components into the final image. Following the layer stacking logic, the generated adaptive background image is placed at the bottom layer, and the image corresponding to the main product area is placed in the middle layer to ensure that the product is not obscured by the background. After rendering the candidate advertising copy, it is placed at the top layer along with the brand logo. By adjusting their coordinate positions in the template, multiple complete and structured candidate creative advertising images are synthesized. These images visually possess all the elements of the finished advertisement.

[0091] S70. Perform saliency testing on the candidate creative advertising images and calculate the saliency ratio between the main product area and the background area.

[0092] Specifically, to quantitatively evaluate whether the generated advertising images effectively highlight the product, a visual saliency detection mechanism is introduced. By simulating the attention distribution mechanism when the human eye observes an image, a distribution map reflecting the strength of attractiveness in different parts of the image is generated, namely a saliency heatmap. In this heatmap, the higher the value, the more eye-catching it is. By separately counting the saliency values ​​in the main product area and the background area, and calculating the ratio between the two, it is possible to objectively determine whether the product or the background is more eye-catching in the current image, thus providing a quantitative basis for subsequent quality screening.

[0093] S80. Based on the significance ratio, select candidate creative advertising images whose significance ratios meet the preset output conditions as available candidate creative advertising images for output.

[0094] Specifically, a threshold for the salience ratio is set. For example, the preset output condition is a ratio greater than 1.2, which means that the product's attractiveness must be at least 1.2 times that of the background. The actual ratio calculated for each candidate creative ad image is compared with this threshold. If the ratio is higher than the threshold, it means that the image has a clear visual focus and highlights the product, making it a qualified ad creative. It is marked as usable and output. If the ratio is lower than the threshold, it means that the background is too flashy, and it is considered a waste and discarded. This mechanism automatically ensures the quality of delivered ad creatives.

[0095] In one embodiment, step S20, namely determining the brand feature vector set based on the brand identifier and the specified color tone, specifically includes:

[0096] S21. Use a lightweight residual network to extract features from the brand logo, and obtain the contour feature vector and the color proportion feature vector.

[0097] Specifically, a deep learning model with low computational complexity but strong feature representation capability is selected as the feature extractor. For example, ResNet-18 with fully connected layers removed or MobileNet with pruning optimization is used as the backbone network. The brand logo image is preprocessed by standardization and scaling while maintaining the aspect ratio and then input into the network. The shallow convolutional filters in the network are used to extract feature maps that reflect geometric information such as logo edges, inflection points, and closed areas. These feature maps are flattened and used as contour feature vectors. At the same time, instead of relying on complex color recognition models, pixel-level color histogram statistics or main color clustering analysis are performed on the preprocessed brand logo image to calculate the coverage area ratio of each main color in the image. These ratio values ​​are serialized into color proportion feature vectors. Thus, the digital deconstruction of the brand logo's form and color composition is achieved with almost no increase in computational burden.

[0098] S22. Obtain the RGB values ​​of the specified hue and map the RGB values ​​to a hue control vector for the HSV color space.

[0099] Specifically, given the limitations of the RGB color space in expressing perceptual differences and stylistic continuity of colors, directly using RGB values ​​makes it difficult for the generative model to understand style instructions such as "warm tone" or "high saturation." Therefore, through a color space transformation matrix formula, the user-input RGB values ​​of a specified hue, such as R=255, G=200, B=0, are converted into HSV (hue, saturation, lightness) color space values. The hue (H) component and saturation (S) component are extracted and processed through sine / cosine encoding or normalization to construct a one-dimensional or multi-dimensional hue control vector. This vector can more intuitively represent the attributes of colors in the latent space, ensuring that the subsequently generated background not only has similar color values ​​but also maintains a consistent warm / cool tendency and vibrancy with the brand's main color in terms of visual perception.

[0100] S23. The contour feature vector, color proportion feature vector and hue control vector are fused to construct the brand feature vector set.

[0101] Specifically, in order to provide a unified and comprehensive style constraint interface for image generation models, it is necessary to integrate the features from different sources and with different physical meanings. Typically, feature concatenation is used to concatenate the contour feature vector, color proportion feature vector, and tone control vector in the channel dimension or feature dimension. For example, a 128-dimensional contour feature, a 64-dimensional color proportion feature, and a 32-dimensional tone control vector can be concatenated into a 224-dimensional comprehensive vector. Alternatively, a lightweight multilayer perceptron (MLP) can be introduced as a fusion layer to further perform feature interaction and dimensionality reduction mapping on the concatenated vector, ultimately outputting a compact brand feature vector set rich in unique brand visual information.

[0102] In one embodiment, in step S40, in response to the advertisement generation request, a target image and text template is obtained, the original product image is mapped to the target image and text template, and the main product area and background area are determined, specifically including:

[0103] S41. Based on the ad generation request, obtain the scene size parameters, and call the corresponding target image and text template from the template library according to the scene size parameters.

[0104] Specifically, the ad generation request will clearly specify the required image specifications, such as "banner ad with a width of 970 pixels and a height of 250 pixels". First, the size parameter is parsed, and then a search is performed in the pre-built standardized template library. The template library stores various JSON or XML format layout description files defined for mainstream e-commerce banners, social media splash screens and other scenarios. These files not only define the total size of the canvas, but also specify the hierarchical relationship, relative coordinates and scaling ratio of each layer. Based on the closest size matching principle or the exact matching principle, the corresponding target image and text template file is retrieved.

[0105] S42. By using an image segmentation algorithm, the boundaries of foreground objects in the original product image are identified to obtain the main product image;

[0106] Specifically, a graph-based image segmentation algorithm is invoked to traverse and calculate features of the pixel matrix of the original product image. The algorithm calculates the classification probability or segmentation boundary of foreground object pixels and background environment pixels by analyzing the numerical differences of adjacent pixels in color space, texture gradient and regional connectivity. Based on the calculation result, an Alpha channel binarized mask with the same resolution as the original image is generated. Pixel regions identified as foreground objects are assigned a retention flag such as 1, and background regions are assigned a removal flag such as 0. Then, a masking operation is performed, and the product subject image without background noise and with smooth edges is extracted from the original image based on the mask.

[0107] S43. Align the main product image with the product placeholder coordinates in the target graphic template, and define the area covered by the product placeholder coordinates as the main product area.

[0108] Specifically, the center point coordinates, maximum width, and maximum height limits of the product layer are read from the target image template configuration file. The optimal scaling ratio required to fully accommodate the main product image within the limit frame is calculated. The determined main product image is scaled and translated to the specified coordinate position according to this ratio. Then, based on the actual projection range of the scaled and translated image on the canvas, a white projection mask with a completely black background is generated. The area covered by white pixels in this mask is precisely defined as the main product area.

[0109] S44. Define the area in the target graphic template other than the main product area as the background area.

[0110] Specifically, based on the product main area mask generated above, a pixel-level reverse operation is performed, that is, the area originally marked as white is set to black, and the area originally marked as black is set to white, thus obtaining a reverse mask. The area covered by the white pixels in the reverse mask is the background area. This area defines the operable range of the image generation model, which is equivalent to delineating a canvas margin for the model, ensuring that the textures and patterns generated by the model only fill the blank space around the product, achieving a seamless connection between the background and the product without overlap or conflict.

[0111] In one embodiment, in step S50, based on the brand feature vector set, atmosphere descriptors, and background region, an adapted background image is obtained through an image generation model, specifically including:

[0112] S50 generates style parameters and semantic parameters based on the brand feature vector set and atmosphere description words, respectively.

[0113] Specifically, in order to adapt to the input interface of the image generation model, the features obtained in the previous steps need to be format-converted. Through fully connected layers or linear mapping layers, the high-dimensional brand feature vector set is projected onto the style control space of the generation model, such as the W space of StyleGAN or the parameter space of AdaIN, to obtain style parameters that can directly adjust the color and texture of the generated image. At the same time, using a pre-trained multimodal text encoder, such as the Text Encoder of the CLIP model, the atmosphere descriptors are encoded into fixed-dimensional text embedding vectors, i.e. semantic parameters.

[0114] S51. Input the style parameters and semantic parameters into the image generation model through a cross-attention mechanism, so that the image generation model can redraw pixels in the background area of ​​the target image template to obtain an adapted background image.

[0115] Specifically, in the generative architecture based on Generative Adversarial Networks (GAN), a cross-attention module is embedded. Style parameters and semantic parameters are used as keys and values, respectively. The intermediate layer feature maps of the image generation model during the reverse denoising or upsampling process are used as queries. The attention weight matrix between the feature maps and the conditional parameters is calculated to dynamically guide the update direction of the feature maps. In each generation iteration, a background region mask is applied for forced constraints, that is, only newly generated pixels within the mask range are retained, while the regions outside the mask range are forcibly reset to the known state or ignored. Finally, after a complete number of generation steps, an adapted background image is output that not only contains the brand's aesthetic characteristics and matches the copywriting's meaning, but also perfectly fits the template layout in terms of spatial structure.

[0116] In one embodiment, such as Figure 2 As shown, the steps in constructing an image generation model specifically include:

[0117] S501. Obtain the generative adversarial network model trained on the generative adversarial network, and construct a training sample set. The training sample set includes several pairs of sample background images, sample brand feature vectors, and sample atmosphere descriptive words.

[0118] Specifically, to train a model that understands both design and branding, a high-quality domain-specific dataset needs to be constructed. First, tens of thousands of visually excellent sample background images are selected from a historically selected advertising material library. For each image, the logo and standard color of the brand to which it belongs are traced and processed into a sample brand feature vector using the aforementioned method. At the same time, using image description algorithms or manual annotation, each background image is labeled with tags describing its visual style or content, i.e., sample atmosphere description words, thereby constructing a triplet data of <background image, brand features, atmosphere description>, providing rich sample support for the supervised learning of the model.

[0119] S502. Input the sample brand feature vector into the mapping network of the generative adversarial network model to obtain the sample style parameters, and obtain the sample semantic parameters based on the sample atmosphere description words.

[0120] Specifically, within the model's internal architecture, a multilayer perceptron (MLP) is designed as a mapping network. Its role is to map the input sample brand feature vector to a more easily decoupled latent space and output sample style parameters. These parameters can control the global attributes of the generated image, such as hue and texture density. At the same time, a BERT or CLIP text encoder with pre-frozen weights is used to process the sample atmosphere descriptors and extract sample semantic parameter vectors that can represent the semantics of the text, providing content guidance for the generator.

[0121] S503. Input the sample style parameters and sample semantic parameters into the generator of the generative adversarial network model, and output the predicted background image.

[0122] Specifically, the generator typically employs a convolutional or Transformer-based architecture. At each layer, it injects sample style parameters using adaptive instance normalization or spatial adaptive normalization techniques to adjust the mean and variance of the feature map, thereby controlling the style. Simultaneously, it injects sample semantic parameters through cross-attention to guide the feature map in generating specific object shapes or scene structures. Starting from random noise, the generator gradually refines itself under these layered constraints, ultimately synthesizing a predicted background image that is similar in distribution to the real advertising background.

[0123] S504. Input the predicted background image and the sample background image into the discriminator of the generative adversarial network model to obtain the discrimination result.

[0124] Specifically, the discriminator is a convolutional neural network with binary classification capabilities. It alternately receives real sample background images and generator-generated predicted background images as input. By extracting features layer by layer, it finally gives a confidence score or feature vector between 0 and 1 at the output layer, which is the discrimination result. This result represents the probability judgment of the discriminator in considering the current input image to be a real image or a fake image.

[0125] S505. Based on the difference between the discrimination result and the true label, calculate the adversarial loss function, and based on the feature difference between the predicted background image and the sample background image, calculate the style consistency loss function.

[0126] Specifically, on the one hand, adversarial loss is calculated using binary cross-entropy loss or WGAN-GP loss function to quantify the degree to which the generator deceives the discriminator and the accuracy of the discriminator in identifying authenticity. On the other hand, in order to ensure that the generated background strictly adheres to the brand style, the Gram matrix distance between the predicted background image and the corresponding sample background image in the intermediate layer of the feature map of the pre-trained model such as VGG network is calculated as the style consistency loss. This loss forces the generator not only to draw realistically, but also to maintain consistency with the sample in terms of texture strokes and color distribution.

[0127] S506. Based on the adversarial loss function and the style consistency loss function, the network parameters of the generator and the discriminator are alternately updated using the backpropagation algorithm until the adversarial loss function and the style consistency loss function meet the corresponding preset conditions, and the adversarial network model is output to obtain the image generation model.

[0128] Specifically, using the Adam or SGD optimizer and following the standard training strategy of generative adversarial networks, such as updating the discriminator parameters multiple times for each generator parameter update, calculating the gradient of the total loss with respect to each network weight parameter using the chain rule, and backpropagating to update the weights, this game process is repeated until the adversarial loss tends to stabilize and the style consistency loss drops to an extremely low level, indicating that the model has the ability to generate high-quality and style-controllable backgrounds. At this point, the network weights of the generator are saved, and the image generation model is obtained.

[0129] In one embodiment, step S70, which involves performing saliency detection on the candidate creative advertisement image and calculating the saliency ratio between the main product area and the background area, specifically includes:

[0130] S71. Visual saliency analysis algorithm is used to detect the visual saliency of candidate creative advertising images and obtain a saliency heatmap.

[0131] Specifically, based on a global contrast-priority visual saliency analysis algorithm, the candidate creative advertising images are scanned at the pixel level. The algorithm calculates the degree of difference between each pixel in the image and its surrounding neighboring pixels or the statistical features of the entire image in terms of brightness, color, and spatial distribution. It quantifies the weight value of the pixel that attracts human visual attention. These weight values ​​are normalized and mapped to a single-channel grayscale image, i.e., a saliency heatmap. In this image, the grayscale value of the pixel, such as 0-255, directly maps the strength of the visual saliency calculated by the algorithm. The bright areas represent the visual focus, and the dark areas represent the ignored background.

[0132] S72. Calculate the average saliency of pixels located within the main product area in the saliency heatmap, and denot it as the main focus.

[0133] Specifically, using the product subject area mask determined in step S43 as a selection tool, the gray values ​​of all pixels falling within the white area of ​​the mask in the saliency heatmap are extracted. These gray values ​​are summed and then divided by the total number of pixels in the area to obtain an average value. This value objectively reflects the average attention level of the product subject in the current complex image compositing environment compared to other areas. The higher the value, the more eye-catching the product is.

[0134] S73. Obtain the edge contour data of the main product area.

[0135] Specifically, edge extraction operators such as the Canny or Sobel operators are called to process the previously generated product main body area mask, extract the boundary pixel coordinates at the junction of the foreground and background in the mask, and form a set of edge coordinate points denoted as E={(xi,yi)}. This coordinate set E accurately marks the physical boundary between the promoted product and the surrounding environment on the two-dimensional space matrix, providing an absolute starting point numerical reference for the subsequent measurement and differentiation of the anti-interference safety distance.

[0136] S74. For each background pixel in the background region defined by the salient heatmap, calculate the shortest spatial Euclidean distance between the corresponding coordinate position and the pixel in the edge contour data.

[0137] Specifically, the traversal algebraic calculation logic is initiated to sequentially read each background pixel point covered by the background region mask, denoted as P(x,y). The Euclidean distance formula is used to calculate the straight-line physical distance between the background pixel point P(x,y) and each independent boundary pixel point (xi,yi) in the edge coordinate point set E. Using the Euclidean distance method, by comparing all the solved distance scalars and extracting the minimum value, it is saved as the shortest spatial Euclidean distance corresponding to the background pixel point P(x,y), denoted as d. This algebraic calculation accurately quantifies the absolute physical distance of a single noise point approximating the core product in spatial coordinates.

[0138] S75. Obtain the preset maximum interference distance parameter, calculate the difference between the maximum interference distance parameter and the shortest spatial distance, and obtain the distance difference.

[0139] Specifically, a pre-defined pixel constant parameter used to define the effective danger limit radius of visual interference is read as the maximum interference distance parameter, denoted as Dmax. For example, Dmax = 300 pixels. The subtraction algebra operation rule is executed to subtract the shortest spatial Euclidean distance d obtained in the previous step from the maximum interference distance parameter, i.e., the formula Dmax-d is executed. When the distance d is greater than Dmax, resulting in a negative result for direct subtraction, a non-linear truncation function is activated, such as Max(Dmax-d,0), to force it to zero as a calculation fault tolerance mechanism. Finally, an absolute attenuation length value with a constant non-negative value is obtained and defined as the distance difference, denoted as Ddiff.

[0140] S76. The ratio of the distance difference to the maximum interference distance parameter is determined as the distance normalization parameter. The distance normalization parameter is multiplied by the preset reference space weight value to calculate the weighted weight corresponding to the candidate saliency pixel.

[0141] Specifically, the division-fractional algebraic operation logic divides the distance difference obtained in the previous step by the maximum interference distance parameter to calculate a distance normalization parameter, denoted as Norm, whose value is precisely locked in the interval between 0 and 1. That is, Norm = Ddiff / Dmax is calculated. When the noise point is 0 pixels away from the edge of the product, i.e., d is 0, this parameter directly reaches the peak value of 1. Then, a constant with a value greater than 1 is read as the reference spatial weight value Wbase, for example, Wbase = 2.0. This reference spatial weight value is multiplied by the obtained distance normalization parameter. That is, the weighted weight W customized for the coordinates of the specific background pixel is calculated using the formula W = Wbase × Norm. Thus, the anti-interference logic of the closer to the edge, the more severe the weight allocation is achieved by relying on rigorous algebraic operations.

[0142] S77. The saliency value of each background pixel in the saliency heatmap within the background region is fused with the corresponding weighting weight to obtain the spatial weighted saliency heatmap distributed in the background region.

[0143] Specifically, for each background pixel node P(x,y) in the background area, the original saliency value of the pixel in the initial saliency heatmap quantity matrix is ​​read and denoted as V_origin. This saliency value is combined with the corresponding weighting weight for calculation. For example, a new saliency value V_weighted is calculated and output using the formula V_weighted=V_origin×(1+W) after being weighted by the distance decay factor. After iterating through all background pixel coordinate points, all these updated saliency values ​​V_weighted are rewritten into a two-dimensional matrix to generate a spatially weighted saliency heatmap array. In this matrix array, the grayscale of the bright points close to the main product will be increased exponentially according to the above formula with a strong algebraic logic.

[0144] S78. Calculate the mean saliency of pixels located in the background region in the spatially weighted saliency heatmap, and denote it as the background basic interference degree.

[0145] Specifically, using background region feature masks as logical query constraints, spatially weighted saliency heatmaps are extracted. Figure 2 The saliency values ​​V_weighted in the dimensional matrix array that are hit by the background mask and dynamically changed due to spatial multiplication are summed and merged using a cumulative algebraic formula, i.e., Sum(V_weighted). The algebraic sum is then divided by the total number of effective pixel granularities N contained under the logical mask, i.e., the division formula Sum(V_weighted) / N is performed. The resulting global static arithmetic mean is used as the background base interference level and is denoted as Ibase. This average quantitative index, after eliminating the influence of extreme values, smoothly evaluates the overall background layout's eye-catching degree from a macroscopic mathematical mean level.

[0146] S79. Obtain the maximum pixel significance value located in the background area in the spatially weighted significance heatmap, and denot it as the background peak interference degree.

[0147] Specifically, it iterates through all significant values ​​V_weighted within the aforementioned spatially weighted saliency heatmap, and uses the extreme value function Max, which selects the maximum value, to isolate and extract the isolated extreme scalar that maximizes the absolute value within the matrix grid. That is, it executes the formula Ipeak=Max(V_weighted) and outputs it separately as the background peak interference degree Ipeak. This quantitative index, which is evaluated as a separate item, is simultaneously superimposed with the dual inducement conditions of the original pixel color block brightness V_origin being huge and the spatial approximation distance d being extremely small, resulting in an extremely high product parameter W. This allows it to directly rely on the extreme peak value to automatically anchor the most incongruous, glaring, and compositionally disruptive local dangerous high-brightness source within the background composition framework of the advertisement image.

[0148] S710. Using a preset weighting formula, the background base interference degree and the background peak interference degree are weighted and summed to obtain the comprehensive background interference index.

[0149] Specifically, two weighting coefficients α and β are set, such as α=0.7 and β=0.3, and then formula I is used. total = α×I base +β×I peak Perform calculations, where I base For background basic interference, I peak Using this weighted method, a comprehensive background interference index is constructed that takes into account both the overall background complexity and local glare interference, representing the peak background interference.

[0150] S711. Calculate the ratio of the subject's attention to the comprehensive background interference index to obtain the significance ratio.

[0151] Specifically, the calculated subject attention value is divided by the comprehensive background interference index value to obtain a dimensionless ratio. This ratio intuitively quantifies the strong contrast between the foreground and the background. The larger the ratio, the more visually prominent the foreground product is and the weaker the background is, which is ideal. Conversely, if the ratio is close to or even less than 1, it indicates that the background is too eye-catching and there is a quality risk of the background overshadowing the subject.

[0152] In one embodiment, step S80, namely, selecting candidate creative advertising images whose significance ratios satisfy preset output conditions as usable candidate creative advertising images, specifically includes:

[0153] S81. Obtain the preset significance threshold and compare the significance ratio with the significance threshold.

[0154] Specifically, based on the correlation analysis of a large amount of historical high-quality ad click-through rate data and visual saliency data, an empirical threshold that can guarantee ad conversion rate is preset, for example, 1.2. This threshold stored in the configuration file is read, and the currently calculated saliency ratio is logically compared with it.

[0155] S82. If the significance ratio is greater than or equal to the significance threshold, then the significance ratio is determined to meet the output condition, and the candidate creative advertising image corresponding to the significance ratio is output as an available candidate creative advertising image.

[0156] Specifically, when the comparison result shows that the ratio is higher than the threshold, it means that the image is visually correct. The image file is tagged as qualified and moved to the final delivery queue or directly pushed to the user for display, confirming its value as a valid advertising material.

[0157] S83. If the significance ratio is less than the significance threshold, then the significance ratio is determined not to meet the output condition.

[0158] Specifically, when the comparison results show that the ratio is below the threshold, it means that the background element is too eye-catching and is very likely to cause users to ignore the core product. The image is marked as obsolete or to be optimized and removed from the candidate delivery list, and is not used as the final output, thus ensuring delivery quality.

[0159] In one embodiment, such as Figure 3 As shown, the method for generating advertising creatives also includes:

[0160] S1. Get the number of available candidate creative ad images.

[0161] Specifically, after completing a round of generation and screening, a counter is used to count the total number of available candidate creative ad images that have been marked as qualified and entered the delivery queue. This real-time statistic is then compared with the expected delivery quantity set by the user in the request to check whether the number meets the user's expected delivery quantity requirement. For example, the user requests to generate 10 images, but only 4 are qualified after screening.

[0162] S2. If the number of available candidate creative ad images is less than the preset number, adjust the atmosphere description to obtain a suitable background image.

[0163] Specifically, when the number of qualified images is insufficient, for example, only 4 qualified images, an automatic backtracking and replenishment mechanism is triggered. Based on the preset degradation strategy, the parameters of the original atmosphere descriptors are adjusted. For example, the artistic weight parameter of the descriptors is reduced, or the original overly intense adjectives such as "explosive visual" are replaced with words with weaker semantic color and more peaceful visual effects, such as "dynamic visual," using a thesaurus. Based on this, the image generation model is re-driven to generate a batch of new adapted background images. These new background images with fine-tuned parameters are often more visually concise and easier to pass subsequent saliency detection.

[0164] S3. If the calculated significance ratio meets the output conditions and the number of available candidate creative ad images is greater than or equal to the preset number, the final available candidate creative ad images will be output.

[0165] Specifically, using the adjusted new descriptive terms, the process re-enters the "generation-synthesis-detection-screening" process. Each new image that passes the screening is added to the counter. This cycle continues until the total number of qualified images reaches or exceeds the user-set preset number, such as 10. At this point, a termination signal is triggered, the cycle stops, and the batch of sufficient and high-quality images is packaged and output, ensuring that visual quality is strictly guaranteed while rigidly meeting the user's delivery quantity requirements.

[0166] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0167] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores data such as raw material data, target image templates, image generation models, and candidate creative advertising images. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements an artificial intelligence-based advertising creative generation method.

[0168] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0169] Obtain the original material data input by the user. The original material data should include at least the original product image, brand logo, specified color scheme, and original keywords.

[0170] Determine the brand feature vector set based on the brand logo and specified color scheme;

[0171] Semantic parsing is performed on the original keywords to identify product attribute words and atmosphere description words. Based on the product attribute words and atmosphere description words, multiple sets of candidate advertising copy are generated using a natural language model.

[0172] In response to the ad generation request, obtain the target image and text template, map the original product image to the target image and text template, and determine the main product area and background area;

[0173] Initialize the image generation model. Based on the brand feature vector set, atmosphere description words, and background region, obtain the adapted background image through the image generation model.

[0174] By integrating the background image, the main product area, the candidate advertising copy, and the brand logo, several sets of candidate creative advertising images are obtained.

[0175] The saliency of candidate creative advertising images is tested, and the saliency ratio between the main product area and the background area is calculated.

[0176] Based on the significance ratio, candidate creative ad images that meet the preset output conditions are selected as available candidate creative ad images for output.

[0177] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0178] Obtain the original material data input by the user. The original material data should include at least the original product image, brand logo, specified color scheme, and original keywords.

[0179] Determine the brand feature vector set based on the brand logo and specified color scheme;

[0180] Semantic parsing is performed on the original keywords to identify product attribute words and atmosphere description words. Based on the product attribute words and atmosphere description words, multiple sets of candidate advertising copy are generated using a natural language model.

[0181] In response to the ad generation request, obtain the target image and text template, map the original product image to the target image and text template, and determine the main product area and background area;

[0182] Initialize the image generation model. Based on the brand feature vector set, atmosphere description words, and background region, obtain the adapted background image through the image generation model.

[0183] By integrating the background image, the main product area, the candidate advertising copy, and the brand logo, several sets of candidate creative advertising images are obtained.

[0184] The saliency of candidate creative advertising images is tested, and the saliency ratio between the main product area and the background area is calculated.

[0185] Based on the significance ratio, candidate creative ad images that meet the preset output conditions are selected as available candidate creative ad images for output.

[0186] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0187] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0188] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for generating advertising creatives based on artificial intelligence, characterized in that, The method for generating advertising creatives includes: Obtain the original material data input by the user, which includes at least the original product image, brand logo, specified color scheme, and original keywords; Determine the brand feature vector set based on the brand logo and the specified color scheme; Semantic parsing is performed on the original keywords to identify product attribute words and atmosphere description words. Based on the product attribute words and atmosphere description words, multiple sets of candidate advertising copy are generated using a natural language model. In response to an ad generation request, a target image and text template is obtained, the original product image is mapped to the target image and text template, and the main product area and background area are determined. Initialize the image generation model, and obtain an adapted background image based on the brand feature vector set, the atmosphere descriptive words, and the background region; The adapted background image, the main product area, the candidate advertising copy, and the brand logo are integrated to obtain several sets of candidate creative advertising images; The candidate creative advertisement image is subjected to saliency detection, and the saliency ratio of the main product area to the background area is calculated. Based on the significance ratio, candidate creative advertising images that meet the preset output conditions are selected as available candidate creative advertising images for output.

2. The advertising creative generation method according to claim 1, characterized in that, The step of determining the brand feature vector set based on the brand identifier and the specified color tone specifically includes: A lightweight residual network is invoked to extract features from the brand logo, resulting in a contour feature vector and a color proportion feature vector. Obtain the RGB value of the specified hue, and map the RGB value into a hue control vector for the HSV color space; The brand feature vector set is constructed by fusing the contour feature vector, the color proportion feature vector, and the hue control vector.

3. The advertising creative generation method according to claim 1, characterized in that, In response to the advertisement generation request, the process of obtaining a target image and text template, mapping the original product image to the target image and text template, and determining the main product area and background area specifically includes: Based on the ad generation request, the scene size parameters are obtained, and the corresponding target image and text template is called from the template library according to the scene size parameters; By using an image segmentation algorithm, the boundaries of foreground objects in the original product image are identified to obtain the main product image; Align the main product image to the product placeholder coordinates in the target graphic template, and determine the area covered by the product placeholder coordinates as the main product area; The area in the target graphic template other than the main product area is defined as the background area.

4. The advertising creative generation method according to claim 1, characterized in that, The process of obtaining an adapted background image based on the brand feature vector set, the atmosphere descriptive words, and the background region through the image generation model specifically includes: Style parameters and semantic parameters are generated based on the brand feature vector set and the atmosphere descriptive words, respectively. The style parameters and semantic parameters are input into the image generation model through a cross-attention mechanism, so that the image generation model can redraw pixels in the background area of ​​the target image template to obtain the adapted background image.

5. The advertising creative generation method according to claim 4, characterized in that, The construction of the image generation model specifically includes: Obtain a generative adversarial network model trained on a generative adversarial network, and construct a training sample set, which includes several pairs of sample background images, sample brand feature vectors, and sample atmosphere descriptive words; The sample brand feature vector is input into the mapping network of the generative adversarial network model to obtain sample style parameters, and the sample semantic parameters are obtained based on the sample atmosphere descriptive words. The sample style parameters and sample semantic parameters are input into the generator of the generative adversarial network model, and the predicted background image is output. The predicted background image and the sample background image are input into the discriminator of the generative adversarial network model to obtain the discrimination result; Based on the difference between the discrimination result and the true label, an adversarial loss function is calculated, and based on the feature difference between the predicted background image and the sample background image, a style consistency loss function is calculated. Based on the adversarial loss function and the style consistency loss function, the network parameters of the generator and the network parameters of the discriminator are alternately updated using the backpropagation algorithm until the adversarial loss function and the style consistency loss function meet the corresponding preset conditions, and the adversarial network model is output to obtain the image generation model.

6. The advertising creative generation method according to claim 1, characterized in that, The step of performing saliency detection on the candidate creative advertisement image and calculating the saliency ratio between the main product area and the background area specifically includes: The candidate creative advertisement images are visually saliency detected using a visual saliency analysis algorithm to obtain a saliency heatmap. Calculate the mean saliency of pixels located within the main product area in the saliency heatmap, and denot it as the main focus score; Obtain the edge contour data of the main body area of ​​the product; For each background pixel in the background region defined by the salient heatmap, calculate the shortest spatial Euclidean distance between the corresponding coordinate position and the pixel in the edge contour data. Obtain the preset maximum interference distance parameter, and calculate the difference between the maximum interference distance parameter and the shortest spatial distance to obtain the distance difference. The ratio of the distance difference to the maximum interference distance parameter is determined as the distance normalization parameter. The distance normalization parameter is multiplied by a preset reference space weight value to calculate the weighted weight corresponding to the candidate saliency pixel. The saliency value of each background pixel in the background region in the saliency heatmap is fused with the corresponding weighting weight to obtain a spatially weighted saliency heatmap distributed in the background region. Calculate the mean saliency of pixels in the spatially weighted saliency heatmap, and denote it as the background basic interference degree; Obtain the maximum pixel saliency value in the spatially weighted saliency heatmap, and denote it as the background peak interference degree; Using a preset weighting formula, the background base interference degree and the background peak interference degree are weighted and summed to obtain the comprehensive background interference index; The significance ratio is obtained by calculating the ratio of the subject's attention to the comprehensive background interference index.

7. The advertising creative generation method according to claim 1, characterized in that, The step of selecting candidate creative ad images whose significance ratios satisfy preset output conditions as usable candidate creative ad images based on the significance ratio specifically includes: Obtain a preset significance threshold, and compare the significance ratio with the significance threshold; If the significance ratio is greater than or equal to the significance threshold, then the significance ratio is determined to meet the output condition, and the candidate creative ad image corresponding to the significance ratio is output as an available candidate creative ad image; If the significance ratio is less than the significance threshold, then the significance ratio is determined not to meet the output condition.

8. The advertising creative generation method according to claim 1, characterized in that, The advertising creative generation method also includes: Obtain the number of available candidate creative ad images; If the number of available candidate creative advertising images is less than the preset number, then adjust the atmosphere description words to obtain the adapted background image; If the calculated significance ratio satisfies the output condition and the number of available candidate creative ad images is greater than or equal to the preset number, the final available candidate creative ad images are output.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the artificial intelligence-based advertising creative generation method as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the artificial intelligence-based advertising creative generation method as described in any one of claims 1 to 8.