Image-text advertisement content generation method and device based on marketing theory cascade deduction

By using a graphic advertising content generation method based on cascading deduction of marketing theory, the consistency and strategy constraints of marketing content generation in existing technologies are solved. This achieves unified and efficient generation of advertising copy and images, optimizes the quality and consistency of marketing content, and forms a closed-loop generation and evaluation process.

CN122472834APending Publication Date: 2026-07-28BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2026-06-29
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing marketing content generation solutions lack explicit marketing strategy constraints, have insufficient consistency between advertising copy and advertising images, and the generated results are prone to deviating from the preset marketing goals. They also lack an explainable strategy deduction link and quality evaluation closed loop, resulting in content that is complete in form but lacks resonance and conversion.

Method used

The graphic advertising content generation method based on marketing theory cascade deduction receives multimodal input data, analyzes product characteristics and generates standardized input data, uses marketing configuration items to constrain the generation of advertising copy and images, including product positioning, persuasion path, theme strategy and emotional strategy parameters, and combines visual control parameters to carry out a unified marketing logic development, and optimizes the generated results through a quality assessment model.

Benefits of technology

It enhances the correspondence between advertising copy and advertising images, improves the matching degree between generated results and marketing objectives, ensures the structure and consistency of content, and forms a closed-loop process of image and text advertising content generation, quality assessment and iterative optimization, thereby improving the efficiency and effectiveness of marketing content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122472834A_ABST
    Figure CN122472834A_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence and computer technology, and discloses a graphic advertisement content generation method and device based on marketing theory cascade deduction. The method comprises the following steps: receiving multi-modal input data of a target commodity; using a trained first target model to analyze the multi-modal input data, extracting commodity feature information, and controlling the commodity feature information to be fused with supplementary text instructions to generate standardized input data; based on the standardized input data, calling a trained second target model to sequentially execute according to a preset cascade reasoning template and a preset reasoning sequence; based on the generated marketing configuration item, using a trained third target model to generate an advertisement script; and based on the marketing configuration item and the advertisement script, generating visual control parameters to obtain a target advertisement picture. The method realizes collaborative control of advertisement script and picture generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and computer technology, and in particular to a method and apparatus for generating marketing graphic content based on marketing theory cascade deduction. Background Technology

[0002] Existing marketing content generation solutions typically rely on product images, descriptions, historical marketing materials, or user tags, directly calling large language models or image generation models to output ad copy and images. While these solutions can improve content production efficiency, they mostly remain at the level of surface-level semantic content completion or material rewriting, lacking unified modeling and intermediate control over product positioning, target audience, persuasion path, content theme, and target emotion. This leads to the following problems: First, the generated content lacks explicit marketing strategy constraints, making it difficult to consistently align with pre-set marketing goals; second, ad copy and images are usually generated separately, lacking a unified strategic intermediate layer, resulting in deviations in thematic focus, emotional tone, and persuasive logic; third, these solutions often struggle to simultaneously address theme focus, emotional expression, and persuasive logic, resulting in content that is formally complete but lacks resonance and conversion potential; fourth, the lack of an interpretable strategy derivation chain and quality assessment loop hinders subsequent reuse, optimization, and review. Summary of the Invention

[0003] This invention provides a method and apparatus for generating graphic and text advertising content based on cascaded deduction of marketing theory, in order to solve the problems of the generated graphic and text advertising content lacking explicit marketing strategy constraints, insufficient consistency between advertising copy and advertising images, and the generation results easily deviating from the preset marketing objectives.

[0004] In a first aspect, the present invention provides a method for generating graphic and text advertising content based on cascading deduction of marketing theory, the method comprising: Receive multimodal input data of the target product. The multimodal input data includes at least a product image of the target product and supplementary text instructions. The first target model is used to parse the multimodal input data, extract product feature information, and fuse the product feature information with supplementary text instructions to generate standardized input data; Based on standardized input data, the second target model is invoked to perform marketing theory deduction according to the preset cascade reasoning order, and marketing configuration items are generated. The marketing configuration items include at least product quadrant positioning parameters, persuasion path parameters, theme strategy parameters, emotional strategy parameters, and persuasion stage weight parameters. Based on the marketing configuration items, the third-party target model is invoked to generate advertising copy; Based on marketing configuration items and advertising copy, visual control parameters are generated, and the target advertising image is obtained based on the visual control parameters.

[0005] In this embodiment of the invention, by parsing the multimodal input data of the target product and generating standardized input data, product attribute information, user control conditions, and marketing needs can be uniformly expressed, providing a structured input foundation for subsequent marketing theory deduction. Furthermore, by using marketing configuration items obtained from cascaded marketing theory deduction to jointly constrain the generation of advertising copy and advertising images, the content of graphic and text advertisements can be made to revolve around a unified marketing logic, thereby enhancing the correspondence between advertising copy and advertising images and the degree of matching between the generated results and marketing objectives.

[0006] In one optional implementation, the preset cascading reasoning sequence includes: Product quadrant location based on Warn grid model; Selecting a persuasion path based on a refined processing probability model; Determine topic strategies based on topic hierarchy; Determine emotional strategies based on the target emotional set; Weights for the persuasion stage are assigned based on the IDA model.

[0007] In this embodiment of the invention, the stages in the cascaded reasoning template are not isolated but rather interdependent: product quadrant positioning determines the involvement characteristics and decision-making attributes of the product; communication path selection determines the persuasive methods for subsequent content; theme segmentation and target emotional tone determination limit the main expression axis and emotional direction of the advertising content; and persuasion path weight allocation controls the presentation emphasis of different information elements in the advertising copy and images. Through the above cascaded reasoning process, a unified marketing configuration can be formed to support the collaborative generation of advertising copy and images.

[0008] In one alternative implementation, the method further includes: Based on the Warn grid model, the product quadrant positioning parameters are obtained; Based on the elaboration possibility model and product quadrant positioning parameters, determine the persuasion path parameters; Based on the persuasion path parameters, topic hierarchy structure, and target sentiment set, determine the topic strategy parameters and sentiment strategy parameters; Based on the topic strategy parameters, emotional strategy parameters, and the Ida model, the weight parameters for the persuasion stage are determined.

[0009] In one optional implementation, the product feature information includes at least one of the following: product name, product description, appearance style, key selling points, potential target audience, applicable platforms, scenario cues, and visual style features. Supplementary text instructions include at least one of the following: marketing objectives, brand tone requirements, platform type constraints, audience characteristic descriptions, and original image retention controls.

[0010] In this embodiment of the invention, by introducing marketing objectives, brand tone requirements, platform type constraints, and audience characteristic descriptions into the supplementary text instructions, the input side can have clearer marketing control conditions.

[0011] In one alternative implementation, the advertising copy includes at least one of the following: headline text, a set of selling points descriptions, body text, a call to action statement, and visual guiding keywords; Based on the marketing configuration items, the third-party objective model is invoked to generate advertising copy, including: The third-objective model is invoked to generate advertising copy based on the theme strategy parameters, emotional strategy parameters, and persuasion stage weight parameters in the marketing configuration items.

[0012] In this embodiment of the invention, by organizing the advertising copy into headline text, a set of selling point descriptions, body text, call to action statements, and visual guiding keywords, the copy content can have a clearer structural division and provide a semantic basis for the subsequent generation of visual control parameters.

[0013] In one alternative implementation, the visual control parameters include at least one of the following: composition method, camera angle, primary and secondary visual objects, lighting model, color temperature range, background scene, color style, emotional atmosphere, and product display method.

[0014] In one alternative implementation, the method further includes: Based on preset marketing reference standards, the quality of advertising copy and target advertising images is evaluated using an evaluation model to generate a comprehensive score result; If the overall score is less than the preset score threshold, the ad copy and / or target ad image will be regenerated based on the overall score.

[0015] In this embodiment of the invention, the quality assessment process is used to check whether the generated results meet the preset marketing reference standards. The comprehensive score result can be used as the quality basis of the output results and as the judgment condition for whether to trigger regeneration, thereby forming a closed-loop process of graphic and text advertising content generation, quality assessment and iterative optimization.

[0016] In one alternative implementation, the quality assessment includes evaluating at least one of the following: semantic consistency of text and images, thematic relevance of content, visual perception quality, emotional resonance with users, consistency of persuasive logic, and integration of strategic elements.

[0017] In one alternative implementation, the method further includes: If the overall score is greater than or equal to the preset score threshold, the target marketing result will be output. The target marketing result includes at least the overall score, the advertising copy, the target advertising image, and the marketing configuration items.

[0018] In this embodiment of the invention, by outputting advertising copy, target advertising images, marketing configuration items, and comprehensive scoring results, a structured basis can be provided for subsequent manual review, marketing debriefing, and content re-editing.

[0019] Secondly, the present invention provides a graphic advertising content generation device based on marketing theory cascade deduction, the device comprising: The input receiving module is used to receive multimodal input data of the target product. The multimodal input data includes at least a product image of the target product and supplementary text instructions. The parsing module is used to parse the multimodal input data using the first target model, extract product feature information, and fuse the product feature information with supplementary text instructions to generate standardized input data; The strategy deduction module is used to call the second target model based on standardized input data to perform marketing theory deduction according to the preset cascading reasoning order, and generate marketing configuration items. The marketing configuration items include at least product quadrant positioning parameters, persuasion path parameters, theme strategy parameters, emotional strategy parameters, and persuasion stage weight parameters. The copy generation module is used to generate advertising copy based on marketing configuration items and by calling a third-party target model. The image generation module is used to generate visual control parameters based on marketing configuration items and advertising copy, and to obtain the target advertising image based on the visual control parameters.

[0020] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the graphic advertising content generation method of the first aspect or any corresponding embodiment described above. Attached Figure Description

[0021] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the first process of a graphic advertising content generation method based on marketing theory cascade deduction according to an embodiment of the present invention; Figure 2 This is a structural block diagram of a graphic advertising content generation device based on marketing theory cascade deduction according to an embodiment of the present invention; Figure 3This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0025] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0026] This embodiment provides a method for generating graphic and text advertising content, which can be used on electronic devices. Figure 1 This is a flowchart illustrating the first method for generating graphic and text advertising content based on the cascading deduction of marketing theory, such as... Figure 1 As shown, the process includes the following steps: Step S101: Receive multimodal input data of the target product. The multimodal input data includes at least a product image of the target product and supplementary text instructions.

[0027] Here, supplementary text instructions include at least one of the following: marketing objectives, brand tone requirements, platform type constraints, audience characteristic descriptions, and original image retention controls.

[0028] Multimodal input data can also include subjective preference parameters.

[0029] Subjective preference parameters reflect users' subjective preferences for target products, mainly including theme preference parameters and / or emotional preference parameters. Theme preference parameters can refer to the preference for themes related to the target product (such as retro style, sports style, etc.), while emotional preference parameters can refer to the emotional inclination that the target product is expected to convey (such as warmth, humor, etc.).

[0030] Supplementary text instructions can refer to instructions in text form that provide additional explanations and restrictions on operations or generated content related to the target product, mainly describing the marketing control conditions for the target product.

[0031] Marketing control conditions refer to the conditions that regulate and restrict various aspects of marketing activities in order to ensure that marketing activities proceed in the expected direction.

[0032] Marketing objectives can refer to specific goals that a marketing campaign aims to achieve, such as increasing sales of a target product by 20% within a month.

[0033] Brand tone requirements can refer to the style, temperament, and image characteristics of the target product. For example, the target product should have a young and fashionable style.

[0034] Platform type constraints can refer to the platforms used for marketing campaigns, such as promoting only on platform A.

[0035] Audience profile description can refer to the characteristic information of the target audience for a target product. For example, the target audience is college students aged 18 to 25 who like trendy culture and technology products.

[0036] In some implementations, if the user chooses to retain the features of the product image, the product image can be used as a control condition to perform partial redrawing and / or style transfer on the product image.

[0037] Here, product image features can include features such as subject structure, color distribution, and semantic information.

[0038] Partial redrawing refers to redrawing specific areas of an image based on a product image. By using the product image as a conditional input, it's possible to ensure that the redrawn portion maintains consistency with the rest of the original image in terms of style, color, and texture. For example, in a portrait photo, if you want to change the clothing, you can partially redraw the clothing while preserving the face and background.

[0039] Style transfer is the application of the style of one image to another. When using product images as conditional input, the target style is incorporated into the original image while preserving the content of the product image. For example, transferring the style of Van Gogh's oil paintings to an ordinary landscape photograph gives the photo the brushstrokes and color characteristics of Van Gogh's paintings.

[0040] It should be noted that product images can be directly imported into image editing tools (such as Photoshop) for editing, or they can be input into a pre-established image editing model to obtain a product image with partial redrawing and / or style transfer. This scene recognition model can be any suitable neural network model capable of performing this function. Image editing can also be performed in other ways, which are not limited in this application.

[0041] In some implementations, if the user chooses not to retain product image features, an image generation model can be used to directly generate a completely new image.

[0042] Step S102: The first target model is used to parse the multimodal input data, extract product feature information, and fuse the product feature information with supplementary text instructions to generate standardized input data.

[0043] Here, the visual content in the product image is converted into a text description using a first target model. The first target model can refer to a visual language model, or it can be a combination of an image classification model, an object detection model, a text recognition model, and a rule engine. This application does not limit the scope of the application.

[0044] In some implementations, product feature information includes at least one of the following: product name, product description, appearance style, key selling points, potential target audience, applicable platforms, scenario cues, and visual style features.

[0045] Among them, the product name can refer to the specific name of the target product; the product description can refer to the detailed description of the product, which can include information such as the function and specifications of the target product; the appearance style can refer to the unique style and characteristics presented by the product in appearance (such as sporty style, modern style, etc.); the key selling points can refer to the most attractive features and advantages of the product; the potential target audience refers to the people who may be interested in and purchase the product; the applicable platform can refer to the platform suitable for promoting the product; the scene clues can refer to the scene information in which the product is used or consumed (such as camping, hiking, etc.); and the visual style characteristics can refer to the visual characteristics presented by the product images, such as color, composition, and lighting.

[0046] In some implementations, the extracted product feature information and supplementary text instructions are organized according to certain formats and rules to form a unified data structure (i.e., standardized input data).

[0047] For example, for an image of a smartwatch, the first target model might output the following information: {"Product Name: Watch A", "Product Description: Equipped with the A1 system, featuring a high-definition circular screen, and functions such as heart rate monitoring and sleep monitoring", "Appearance Style: Stylish and minimalist", "Key Selling Points: Smart system, health monitoring functions", "Potential Target Audience: Sports enthusiasts, office workers", "Applicable Platform: Platform B", "Scenario Clues: Sports and fitness, daily wear", "Visual Style Characteristics: Metallic texture, blue color scheme"}.

[0048] At this point, the supplementary text instruction is set as follows: "The marketing objective is to boost product sales. The brand tone requirement is a strong sense of technology, suitable for promotion on platform C, and the target audience is young people aged 18-25." The standardized input data structure after merging is as follows: {"Product Name: Watch A", "Product Description: Equipped with the A1 system, featuring a high-definition circular screen, and functions such as heart rate monitoring and sleep monitoring", "Appearance Style: Stylish and minimalist", "Key Selling Points: Smart system, health monitoring function", "Potential Target Audience: Sports enthusiasts, office workers", "Applicable Platform: Platform B", "Scenario Clues: Sports and fitness, daily wear", "Visual Style Characteristics: Metallic texture, blue color scheme", "Supplementary Text Instruction": {"Marketing Objective: Boost product sales", "Brand Tone Requirement: Strong sense of technology", "Platform Type Constraint: Platform C", "Target Audience Description: Young people aged 18-25"}}.

[0049] Step S103: Based on standardized input data, the second target model is invoked to perform marketing theory deduction according to the preset cascading reasoning order, and marketing configuration items are generated. The marketing configuration items include at least product quadrant positioning parameters, persuasion path parameters, theme strategy parameters, emotional strategy parameters, and persuasion stage weight parameters.

[0050] Here, the second target model, which is executed sequentially according to the preset cascaded reasoning template and the preset reasoning order of the cascaded reasoning template, can be a single large language model, or it can be a collaborative execution of multiple role-based sub-agents, such as product positioning agent, strategy planning agent, copywriting agent and visual translation agent, to complete the reasoning.

[0051] In some implementations, the pre-defined cascading reasoning sequence includes: locating the product quadrant based on the Warn grid model; selecting the persuasion path based on the elaboration probability model; determining the topic strategy based on the topic hierarchy structure; determining the emotional strategy based on the target sentiment set; and allocating the weights of the persuasion stages based on the Ida model.

[0052] Here, the Warn Grid model refers to mapping target products to one of four quadrants based on consumer involvement (high or low) and product informativeness (rational or emotional). High involvement means that consumers invest more time and effort in information gathering and comparison during the purchase decision-making process; low involvement indicates that consumers make purchase decisions more casually. Rational information emphasizes the product's objective attributes such as function and performance, while emotional information focuses on the emotional experience and image the product evokes.

[0053] The propagation path selection based on the elaboration possibility model can refer to the following: when the target product is highly involved in rationality, the central path should be the main one, emphasizing parameters, evidence, and logical arguments; when the target product is low-involvement in emotionality, the peripheral path should be the main one, emphasizing atmosphere, story, symbols, and visual stimulation; for some highly involved emotional products, a dual-track strategy of peripheral path dominance and central path supplementation can be adopted.

[0054] Thematic division based on thematic hierarchy theory can refer to the output of product information based on the target product, which should have at least two layers: a core theme and auxiliary themes. The core theme is used to constrain the narrative axis of the entire copy and image, while the auxiliary themes are used to support parameter endorsements, price advantages, reviews, or supplementary information on scenarios.

[0055] Determining the target emotional tone based on the emotion transmission model can refer to selecting one or more target emotional tones from a pre-set set of emotions. The set of emotions may include trust, professionalism, humor, warmth, excitement, festive feeling, sense of identity, desire for exploration, etc.

[0056] The persuasion path weight allocation based on the IDA model can refer to calculating the weight distribution of the four stages of attention, interest, desire, and action according to the type of target product, platform form, and marketing objectives, and using this distribution to control the headline strength, body text length, selling point density, and call to action position.

[0057] In some implementations, product quadrant positioning parameters are obtained based on the Warn grid model; persuasion path parameters are determined based on the refinement possibility model and product quadrant positioning parameters; topic strategy parameters and emotional strategy parameters are determined based on the persuasion path parameters, topic hierarchy structure, and target sentiment set; and weight parameters for the persuasion stage are determined based on the topic strategy parameters, emotional strategy parameters, and the Ida model.

[0058] It should be noted that the marketing positioning parameters refer to the positioning results of the target product obtained based on the product quadrant positioning of the Warn grid model, the communication path parameters refer to the path results obtained based on the communication path selection of the refined processing possibility model and the weight allocation results of the persuasion path weight allocation based on the Ida model, and the content strategy parameters refer to the core theme and auxiliary theme obtained based on the theme division of the theme hierarchy theory and the target emotional tone obtained based on the target emotional tone of the emotional communication model.

[0059] Step S104: Based on the marketing configuration items, call the third target model to generate advertising copy.

[0060] Here, the third target model can refer to a large language model or other neural network model that can generate advertising copy; this application does not limit it.

[0061] In one optional implementation, the advertising copy includes at least one of the following: headline text, a set of selling points descriptions, body text, a call to action, and visual guiding keywords. The third-objective model is invoked to generate the advertising copy based on the theme strategy parameters, sentiment strategy parameters, and persuasion stage weight parameters in the marketing configuration items.

[0062] Here, the headline text can refer to the core appeal of the advertising copy, used to summarize the main content of the advertisement or highlight the key selling points of the target product.

[0063] A selling point description set can be a collection that describes the selling points of a target product in detail, clearly presenting the unique advantages and features of the target product to users and helping them understand the value of the target product.

[0064] The main body of the text can refer to the part that further elaborates on the target product information, brand story, usage scenarios, etc. It supplements and expands the title and selling points, allowing users to have a more comprehensive understanding of the target product.

[0065] A call to action is a statement that encourages users to take a certain action, such as purchasing a target product, learning more, or participating in an activity. It is a key part of guiding user conversion.

[0066] Visual guidance keywords can refer to keywords used to guide the visual design of advertisements and attract users' attention.

[0067] Step S105: Based on the marketing configuration items and advertising copy, generate visual control parameters, and obtain the target advertising image based on the visual control parameters.

[0068] In some implementations, visual control parameters can be generated using a trained fourth-objective model based on the advertising copy. This fourth-objective model can be the same large language model as the third-objective model, or other deep learning models capable of generating visual control parameters; this application does not impose any limitations on this.

[0069] In some implementations, visual control parameters can be input into a pre-established advertising image generation model to obtain the target advertising image. This advertising image generation model can be any suitable neural network model capable of performing this function.

[0070] In some alternative implementations, the visual control parameters include at least one of the following: composition method, camera angle, primary and secondary visual objects, lighting model, color temperature range, background scene, color style, emotional atmosphere, and product display method.

[0071] Composition refers to the way elements are organized and arranged in a picture, such as the rule of thirds and central composition.

[0072] Lens perspective refers to the position and angle of the camera when shooting. Different lens perspectives can show different visual effects and emotional atmospheres, such as overhead shots and low-angle shots.

[0073] The primary and secondary visual objects refer to the main visual object and the secondary visual object. The primary visual object is the most important expressive element in the picture and attracts the viewer's attention; the secondary visual object plays a supporting and supplementary role, helping to highlight the primary visual object and enrich the content of the picture.

[0074] A lighting model describes how light is distributed and propagates in a scene. Different lighting models produce different lighting effects, affecting the atmosphere and texture of the image, such as side lighting and backlighting.

[0075] Color temperature range refers to the color temperature of light. Different color temperatures will bring different visual feelings and emotional associations to people, such as warm colors and cool colors.

[0076] Background scene refers to the environment surrounding the main subject in the picture. Background scene can provide contextual information for the main subject and enhance the storytelling and appeal of the picture, such as natural background, shopping mall background, etc.

[0077] Color style refers to the overall combination and expression of colors in an image. Different color styles can convey different emotions and information, such as vibrant style and elegant style.

[0078] Emotional atmosphere refers to the feelings and atmosphere conveyed by a picture. It is created through a combination of factors such as composition, color, and lighting. Examples include a sad atmosphere and a joyful atmosphere.

[0079] Product display methods refer to the specific methods and forms of displaying target products, with the aim of highlighting the characteristics and advantages of the target products and attracting consumers' attention, such as close-up displays.

[0080] In this embodiment of the invention, a fourth target model is used to generate visual control parameters based on the advertising copy, and a target advertising image is obtained based on the visual parameters. This improves the correlation between the target advertising image and the preset marketing objective and reduces the subjective bias caused by manual design in related technologies.

[0081] In some alternative implementations, the method further includes: Step a1: Based on the preset marketing reference standards, use the evaluation model to evaluate the quality of the advertising copy and target advertising images, and generate a comprehensive score result; Here, the pre-set marketing reference standards refer to the criteria and standards established in advance to measure the effectiveness and quality of advertising, which need to comprehensively consider factors such as market demand, brand image, and target audience characteristics. No specific marketing reference standards are limited here.

[0082] In some implementations, the quality of advertising copy and target advertising images can be evaluated using an assessment model based on marketing reference standards, thereby generating multiple sets of scoring results. The visual language model is then used to automatically sort the scoring results of each set, and the scoring results within the preset sorting threshold (such as 25%, 30%, etc.) are selected and returned. Users can select one of the returned scoring results as the comprehensive scoring result, or directly use the top-ranked scoring result as the comprehensive scoring result.

[0083] In some implementations, the quality assessment includes evaluating at least one of the following: semantic consistency of text and images, thematic relevance of content, visual perception quality, emotional resonance with users, consistency of persuasive logic, and integration of strategic elements.

[0084] Image-text semantic similarity refers to the degree of matching between the semantic information conveyed by the advertising copy and the target advertising image. For example, if the advertising copy describes the target product as having "whitening and moisturizing" effects, and the image can clearly demonstrate the product's whitening and moisturizing effects, such as the model's skin becoming white and hydrated after using the product, then the image-text semantic similarity is relatively high.

[0085] Content theme relevance refers to the degree to which the content expressed in the advertising copy and images matches the marketing theme.

[0086] Visual perception quality refers to the intuitive feeling and quality level of an advertising image, including the clarity of the image, color matching, and composition.

[0087] User emotional resonance refers to the degree to which an advertisement can evoke emotional responses from its target audience. For example, advertising copy and images can create a warm and touching atmosphere, allowing consumers to relate to their own life scenarios and thus generating emotional resonance.

[0088] Consistent persuasive logic refers to the coherence and logical coherence of advertising copy in explaining product features, advantages, and guiding purchase decisions. For example, advertising copy might first introduce the product's unique functions, then explain the benefits these functions bring to consumers, and finally provide reasons and methods for purchasing, demonstrating clear and consistent logic.

[0089] The integration of strategic elements refers to whether the advertising copy effectively integrates strategic elements such as marketing objectives, target audience, and brand positioning into the copy and images.

[0090] Step a2: If the overall score is less than the preset score threshold, regenerate the advertising copy and / or target advertising image based on the overall score.

[0091] In some implementations, a comprehensive score can be obtained based on the scoring results of the advertising copy and the target advertising image in the above-mentioned dimensions.

[0092] In some implementations, the scores of one or more dimensions can be directly used as the comprehensive score result. For example, the score of content theme relevance can be used as the comprehensive score result. Alternatively, the total score of visual perception quality and user emotional resonance can be used as the comprehensive score result.

[0093] In some implementations, the scores can be directly added together to obtain a sum, which can then be used as the overall score. Alternatively, the scores can be weighted and averaged before being added together to obtain the overall score.

[0094] There are no restrictions on how the comprehensive score results are obtained.

[0095] In some implementations, a sports brand launches a new athletic shoe with the marketing goal of attracting young sports enthusiasts, highlighting the shoe's lightweight, comfortable, and stylish features.

[0096] The ad copy generated through the above steps is: "These sneakers are super lightweight, allowing you to exercise without any burden, and their stylish design will make you the center of attention." The target ad image shows a pair of sneakers placed on a plain white background, failing to convey the athletic scene or the shoe's lightweight feel. The ad copy and target ad image are then scored across various dimensions: Image-text semantic similarity: The copy emphasizes the shoes' lightness and style, but the images fail to effectively convey this lightness, only showcasing the shoes' appearance. Image-text semantic similarity score: 60. Content theme relevance: The copy revolves around the features of athletic shoes, but the images lack a sports scene and don't effectively align with the theme of "attracting young sports enthusiasts." Content theme relevance score: 55. Visual perception quality: The image clarity is acceptable, but the composition is simplistic and visually unappealing. Visual perception quality score: 65. User emotional resonance: Neither the copy nor the images evoke any emotional resonance with users regarding sports. User emotional resonance score: 50. Persuasive logic consistency: The copy is logically clear, introducing features before emphasizing advantages, but the images fail to effectively persuade the user. Persuasive logic consistency score: 60. Strategic element integration: The strategic elements for attracting young sports enthusiasts are not effectively integrated into the images. Strategic element integration score: 50.

[0097] The overall score, based on all scores, is 58 points.

[0098] Since the preset scoring threshold is 70 points, but the obtained comprehensive score of 58 points is less than the preset scoring threshold, it is necessary to regenerate the advertising copy and / or target advertising image.

[0099] In some implementations, the advertising copy can be directly input into the advertising image generation model to regenerate the target advertising image.

[0100] In some implementations, steps S101 to S105 can be repeated to regenerate the advertising copy and target advertising image.

[0101] In this embodiment of the invention, on the one hand, the quality of advertising copy and target advertising images can be comprehensively measured through multi-dimensional scoring; on the other hand, when the comprehensive score result is less than the threshold, the advertising copy and / or target advertising images are regenerated, which can continuously optimize the content and improve the competitiveness and marketing effect of the advertisement.

[0102] In some implementations, user opinions, evaluations, and suggestions regarding the generated ad copy and target ad images are collected. This reflects users' intuitive perception of whether the ad content meets their expectations and needs.

[0103] For example, an e-commerce platform generated advertising copy and images for a smartwatch. The advertising copy emphasized the watch's various functions, such as heart rate monitoring and activity tracking; the advertising images showcased the watch's performance in different scenarios. User feedback after viewing the ads might include comments such as, "The description of the watch's battery life in the copy is not prominent enough," or "The colors in the advertising images are too dark and not attractive enough."

[0104] This allows for the updating of supplementary text instructions based on user feedback, in order to regenerate ad copy and target ad images.

[0105] In this embodiment of the invention, user feedback reflects potential shortcomings in advertising copy and target advertising images. Updated and supplemented text instructions are regenerated to optimize advertising copy and target advertising images, thereby improving the attractiveness and persuasiveness of the advertisement and enhancing marketing effectiveness.

[0106] In this embodiment of the invention, on the one hand, by parsing the multimodal input data of the target product and generating standardized data, the correlation between marketing positioning and product characteristics is improved, and the possibility of deviation between the output content and the preset marketing objectives is reduced. On the other hand, advertising copy is generated based on marketing configuration items, and then advertising images are obtained based on the advertising copy, ensuring that the style of the advertising copy and the target advertising image are consistent, and improving the efficiency of marketing content generation.

[0107] In some implementations, the specific steps of the graphic advertising content generation method are as follows: Taking the marketing scenario of high-end flagship smartphones as an example, the system receives product images and supplementary text instructions such as "generate high-end marketing images and texts for social media platforms for photography enthusiasts and business users".

[0108] The visual semantic features of the target product, such as "circular multi-camera module, metal body, high-end appearance, and prominent photography attributes," were analyzed. The product was then identified as a high-involvement, emotionally driven decision-making product. A persuasion strategy was adopted, with the peripheral path as the main approach and the central path as a supplement. "Professional pride" and "identity recognition" were identified as the target emotional tone, "identity expression through image creation" was set as the core theme, and "parameter endorsement" was set as the auxiliary theme. At the same time, high weights were assigned to desire and interest.

[0109] Based on this configuration option, an advertising copy including a title, selling points, body text, and call to action is generated. The advertising copy is then further translated into visual control parameters, including visual parameters such as "low-light high-contrast background, exquisite and high-end texture, highlighting lens details, and dark luxury atmosphere". Finally, a set of consistent text and images with unified emotion and persuasive marketing appeal is output.

[0110] In this embodiment of the invention, the specific process of the graphic and text advertisement content generation method is as follows: S1. Multimodal input reception. Receives product images of the target product, supplementary text instructions, and optional theme preference parameters and sentiment preference parameters; among which, the supplementary text instructions are used to describe control conditions such as marketing objectives, brand tone, platform type, audience characteristics, or whether to retain original image features.

[0111] S2. Structural Analysis of Product Features. A visual language model is invoked to perform semantic analysis on product images, extracting product name, product description, appearance style, key selling points, potential target audience, applicable platforms, scene clues, and visual style features. These are then integrated with supplementary text instructions to form a standardized input data structure.

[0112] S3. Marketing Theory Parameter Derivation. Based on a standardized input data structure, the large language model is invoked and executed sequentially according to a preset cascading reasoning template: product quadrant positioning based on the Warn grid model, communication path selection based on the elaboration possibility model, topic division based on topic hierarchy theory, determination of target emotional tone based on the emotional communication model, and persuasion path weight allocation based on the Ida model.

[0113] S4. Copy Generation. Based on the structured marketing configuration items, the large language model is invoked to generate advertising copy containing a headline, a list of selling points, body text, a call to action, and visual keywords. The copy generation is constrained by the core theme, auxiliary theme, target emotional tone, product quadrant positioning, communication path selection, and weight allocation results.

[0114] S5. Visual Translation and Image Generation. Based on the structured marketing configuration items and advertising copy, the large language model is called to generate raw image prompts (equivalent to the visual control parameters mentioned above). The raw image prompts include at least the composition method, camera angle, main and secondary visual objects, lighting model, color temperature range, background scene, color style, emotional atmosphere, and product display method. The raw image prompts are then input into the image generation model to obtain the advertising image.

[0115] S6. Quality Assessment. Input the generated advertising copy and images into the assessment model, and combine them with preset marketing reference standards to score the semantic consistency of text and images, visual perception quality, content theme relevance, user emotional resonance, persuasive logic consistency, and integration of strategic elements, generating a quality assessment report (equivalent to the comprehensive score result mentioned above).

[0116] S7. Output Results. The output includes the generated results (equivalent to the target marketing results) including marketing configuration items, advertising copy, advertising images (equivalent to the target advertising images mentioned above), and a quality assessment report. In some implementations, steps S4 and S5 are iteratively regenerated based on the quality assessment report to obtain better marketing content.

[0117] This embodiment also provides a marketing content generation device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0118] This embodiment provides a marketing content generation device, such as... Figure 2 As shown, it includes: The input receiving module 201 is used to receive multimodal input data of the target product. The multimodal input data includes at least a product image of the target product and supplementary text instructions. The parsing module 202 is used to parse the multimodal input data using the first target model, extract product feature information, and fuse the product feature information with supplementary text instructions to generate standardized input data; The strategy deduction module 203 is used to call the second target model based on standardized input data to perform marketing theory deduction according to the preset cascading reasoning order, and generate marketing configuration items. The marketing configuration items include at least product quadrant positioning parameters, persuasion path parameters, theme strategy parameters, emotional strategy parameters, and persuasion stage weight parameters. The copy generation module 204 is used to generate advertising copy based on marketing configuration items and by calling a third-party target model. Image generation module 205 is used to generate visual control parameters based on marketing configuration items and advertising copy, and obtain target advertising images based on visual control parameters.

[0119] In some alternative embodiments, the apparatus further includes: The positioning module is used to locate the product quadrant based on the Warn grid model.

[0120] The selection module is used to choose a persuasion path based on a refined processing probability model.

[0121] The first determination module is used to determine the topic strategy based on the topic hierarchy structure.

[0122] The second determination module is used to determine the sentiment strategy based on the target sentiment set.

[0123] The allocation module is used to allocate weights for the persuasion stage based on the Aida model.

[0124] In some alternative embodiments, the apparatus further includes: The obtained unit is used to obtain the product quadrant positioning parameters based on the Warn grid model; The first determining unit is used to determine the persuasion path parameters based on the fine processing possibility model and product quadrant positioning parameters; The second determining unit is used to determine the topic strategy parameters and the sentiment strategy parameters based on the persuasion path parameters, the topic hierarchy structure, and the target sentiment set. The third determining unit is used to determine the weight parameters of the persuasion stage based on the topic strategy parameters, the emotion strategy parameters, and the Ida model.

[0125] In some alternative embodiments, the apparatus further includes: The first generation module is used to call the third target model to generate advertising copy based on the theme strategy parameters, emotional strategy parameters, and persuasion stage weight parameters in the marketing configuration items.

[0126] In some alternative embodiments, the apparatus further includes: The second generation module is used to evaluate the quality of advertising copy and target advertising images based on preset marketing reference standards and an evaluation model, and generate a comprehensive score result. The third generation module is used to regenerate the advertising copy and / or target advertising image based on the comprehensive score result if the comprehensive score result is less than the preset score threshold.

[0127] The marketing content generation apparatus provided in this embodiment of the invention can execute the graphic and text advertising content generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0128] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0129] The following is a detailed reference. Figure 3 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 301, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 302 or a program loaded from memory 308 into random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device. The processor 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0130] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0131] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 309, or installed from a memory 308, or installed from a ROM 302. When the computer program is executed by the processor 301, it performs the functions defined in the graphic advertising content generation method of the embodiments of the present invention.

[0132] Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0133] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0134] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for generating graphic and text advertising content based on cascading deduction of marketing theory, characterized in that, The method includes: Receive multimodal input data of the target product, wherein the multimodal input data includes at least a product image of the target product and supplementary text instructions; The multimodal input data is parsed using the first target model to extract product feature information, and the product feature information is fused with the supplementary text instructions to generate standardized input data; Based on the standardized input data, the second target model is invoked to perform marketing theory deduction according to the preset cascaded reasoning order, and marketing configuration items are generated. The marketing configuration items include at least product quadrant positioning parameters, persuasion path parameters, theme strategy parameters, emotional strategy parameters, and persuasion stage weight parameters. Based on the aforementioned marketing configuration items, the third-party target model is invoked to generate advertising copy; Based on the marketing configuration items and the advertising copy, visual control parameters are generated, and the target advertising image is obtained based on the visual control parameters.

2. The method according to claim 1, characterized in that, The preset cascaded reasoning sequence includes: Product quadrant location based on Warn grid model; Selecting a persuasion path based on a refined processing probability model; Determine topic strategies based on topic hierarchy; Determine emotional strategies based on the target emotional set; Weights for the persuasion stage are assigned based on the IDA model.

3. The method according to claim 2, characterized in that, The method further includes: Based on the Warn grid model, the product quadrant positioning parameters are obtained; Based on the aforementioned fine processing possibility model and the product quadrant positioning parameters, the persuasion path parameters are determined; Based on the persuasion path parameters, the topic hierarchy structure, and the target emotion set, the topic strategy parameters and the emotion strategy parameters are determined. Based on the topic strategy parameters, the emotional strategy parameters, and the Aida model, the weight parameters for the persuasion stage are determined.

4. The method according to claim 1, characterized in that, The product feature information includes at least one of the following: product name, product description, appearance style, key selling points, potential target audience, applicable platforms, scenario clues, and visual style features; The supplementary text instructions include at least one of the following: marketing objectives, brand tone requirements, platform type constraints, audience characteristic descriptions, and original image retention control conditions.

5. The method according to claim 1, characterized in that, The advertising copy shall include at least one of the following: headline text, a set of selling points descriptions, body text, call to action, and visual guiding keywords; The step of generating advertising copy by calling the third target model based on the marketing configuration items includes: The third target model is invoked to generate the advertising copy based on the theme strategy parameters, emotional strategy parameters, and persuasion stage weight parameters in the marketing configuration items.

6. The method according to claim 1, characterized in that, The visual control parameters include at least one of the following: composition method, camera angle, main and secondary visual objects, lighting model, color temperature range, background scene, color style, emotional atmosphere, and product display method.

7. The method according to claim 1, characterized in that, The method further includes: Based on preset marketing reference standards, an evaluation model is used to assess the quality of the advertising copy and the target advertising image, and a comprehensive score result is generated. If the overall score is less than a preset score threshold, the advertising copy and / or the target advertising image are regenerated based on the overall score.

8. The method according to claim 7, characterized in that, The quality assessment includes evaluating at least one of the following: semantic consistency of text and images, content thematic relevance, visual perception quality, user emotional resonance, consistency of persuasive logic, and integration of strategic elements.

9. A graphic advertising content generation device based on cascaded deduction of marketing theory, characterized in that, include: An input receiving module is used to receive multimodal input data of a target product, wherein the multimodal input data includes at least a product image of the target product and supplementary text instructions; The parsing module is used to parse the multimodal input data using the first target model, extract product feature information, and fuse the product feature information with the supplementary text instructions to generate standardized input data; The strategy deduction module is used to call the second target model to perform marketing theory deduction according to the preset cascaded reasoning order based on the standardized input data, and generate marketing configuration items. The marketing configuration items include at least product quadrant positioning parameters, persuasion path parameters, theme strategy parameters, emotional strategy parameters, and persuasion stage weight parameters. The copy generation module is used to generate advertising copy by calling a third target model based on the marketing configuration items; The image generation module is used to generate visual control parameters based on the marketing configuration items and the advertising copy, and to obtain the target advertising image based on the visual control parameters.

10. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method according to any one of claims 1 to 8.