AI intelligent image-text situation content accurate layout full-marketing generation method

By building a multi-semantic field group and a collaborative generation engine, the semantic misreading and situational mismatch problems in graphic and text generation are solved, the deep alignment between images and text is achieved, and the accuracy of marketing graphic and text and user perception consistency is improved.

CN120374783AActive Publication Date: 2025-07-25SHANGHAI HAIPAI LINGKE CULTURE TECH CO LTD

Patent Information

Application Number
CN202510864081.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In the prior art, the graphic and text generation model is prone to misunderstand the semantics of "negation" or "contrast" in the copy, and generates images that are opposite to the expected situation, resulting in weakening the sensual motivation and conversion effect of marketing graphic and text.

Method used

Build a multi-semantic field group, including idea items, reverse semantics, auxiliary emotions and tonal keywords, and perform regional attention intervention and structural-level semantic consistency judgment through the collaborative generation engine of the image generation sub-model and the semantic prediction sub-model, to ensure that the image is aligned with the text.

Benefits of technology

It significantly improves the marketing accuracy and user perception consistency of graphic and text content, and is especially suitable for high-demand scenarios such as brand content customization, social media marketing and smart advertising design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374783A_ABST
    Figure CN120374783A_ABST
Patent Text Reader

Abstract

The invention discloses an AI intelligent image-text situation content accurate layout full-marketing generation method, and particularly relates to the technical field of image-text recognition. The method comprises the following steps: performing semantic analysis on a text input by a user, extracting a deliberate graph item, reverse semantics, an auxiliary emotion item and a tonality keyword, and constructing a multi-semantic field group; the field groups are injected into a collaborative generation engine composed of an image generation sub-model and a semantic prediction sub-model, and the semantic prediction sub-model predicts graph attribute tags and guides the image generation sub-model to conduct regional attention intervention; after the image is generated, the matching degree of the image content and the field group structure is evaluated through a structure-level semantic consistency judgment mechanism, and regeneration is triggered if the image content and the field group structure are inconsistent; finally, the image content with the aligned structure is output and subjected to image-text combination typesetting with the original text, and image-text situation content meeting the multi-platform marketing application requirement is generated. According to the method, the semantic accuracy, the structural consistency and the visual expression quality are remarkably improved while the generation efficiency is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graphic and text recognition, and particularly to an AI intelligent graphic and text context content precise layout full-marketing generation method. Background Art

[0002] AI intelligent graphic and text context content precise layout full-marketing generation refers to using artificial intelligence technology to intelligently generate graphic and text content according to different marketing scenarios, and through precise content layout strategies, achieve omni-channel and integrated marketing communication. This method combines context understanding, content creation, and distribution optimization, aiming to improve marketing efficiency and user conversion effect.

[0003] The existing technologies have the following deficiencies: Contextually anti-logical image generation is a serious technical problem faced by current diffusion models in graphic and text generation. Specifically, it is manifested that the model misinterprets descriptions with "negative" or "contrastive" semantics in the copywriting, and then generates images contrary to the expected situation. For example, in the scenario of "emphasizing warmth in the subway on a cold winter night", the model may focus on the literal keyword "cold" and generate a picture full of ice and snow elements and people dressed coolly, completely deviating from the originally intended "warm" atmosphere. Such deviations stem from the ambiguity of Prompt semantic expression and the insufficient coverage of contrastive scenario samples in the training stage of the model, which may ultimately seriously mislead users' emotional cognition and weaken the emotional appeal and conversion effect of marketing graphics and texts. Summary of the Invention

[0004] The purpose of the present invention is to provide an AI intelligent graphic and text context content precise layout full-marketing generation method to solve the deficiencies in the background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: An AI intelligent graphic and text context content precise layout full-marketing generation method, including: Analyze the user input, extract the marketing intention items and their context semantics, and construct a multi-semantic field group including the main intention item, reverse semantics, auxiliary emotions, and tonality keywords; Inject the field group into a collaborative generation engine composed of an image generation sub-model and a semantic prediction sub-model. The semantic prediction sub-model is used to predict the graphic attribute labels to be generated; During the generation process, the image generation sub-model performs regional attention intervention according to the labels, making the main semantic features focus on the specified area of the image and dynamically suppressing the image expression of reverse word meanings; After generating the image, through a structural-level semantic consistency judgment mechanism, compare the image content structure based on the position weights of the semantic items in the field group. If they do not match, perform re-generation; Combine the image that meets the consistency requirements with the original text and output it as complete graphic and text context marketing content.

[0006] Preferably, parsing the user input and constructing a multi-semantic field group includes: Performing syntactic-level parsing on the input text to identify the core subject-predicate-object structure indicating the marketing intention, and extracting the main intention item therein; Based on the main intention item, invoking the reverse semantic dictionary and the domain ontology library to retrieve its logical opposite item and mark it as the reverse semantics; Identifying emotional words and intonation descriptions for modifying the main intention item through context analysis, and classifying them as auxiliary emotion items and tonality keywords respectively; Encapsulating the main intention item, reverse semantics, auxiliary emotion items, and tonality keywords into a multi-field structure, and combining them into a semantic field group vector according to the predefined field order.

[0007] Preferably, injecting the semantic field group into the collaborative generation engine composed of an image generation sub-model and a semantic prediction sub-model includes: Performing field identification and semantic labeling on each field item of the semantic field group; Constructing a multi-segment control vector according to the predefined field order, and allocating each field item to different model input channels; Injecting the main intention field and the auxiliary emotion field into the control channel of the image generation sub-model; Inputting the reverse semantic field and the tonality keyword field into the semantic prediction sub-model for generating graphic restriction labels and layout style prediction labels.

[0008] Preferably, the semantic prediction sub-model for predicting graphic attribute labels includes: Receiving the structured semantic field group vector as input; Based on the multi-label classification model established from the training semantic-visual contrast sample library, identifying the potential image elements corresponding to the semantic fields; Mapping each semantic field item to the corresponding graphic attribute label, including the type of visual element, position weight, and emotional expression characteristics; Outputting a set of structured graphic attribute labels as the control constraint basis for the image generation sub-model.

[0009] Preferably, the image generation sub-model performs regional attention intervention during the generation process, including: Receiving the graphic attribute labels output from the semantic prediction sub-model, and extracting the spatial position coordinates and weights corresponding to each semantic field; Constructing a multi-channel spatial position mask matrix, assigning an enhancement mark to the position mask of the main intention field and an inhibition mark to the position mask of the reverse semantic field; During the diffusion generation process, using the spatial position mask to weightedly adjust the normalized weights of the attention mapping layer; When generating a target area in an image, the channel activation threshold is adjusted in real time.

[0010] Preferably, the regional attention intervention includes: Establish a semantic field during the generation process, that is, the image channel mapping relationship, and bind the semantic weight to the feature channel through the channel attention mechanism; The activation intensity of the channels bound by the main intention field is dynamically up-regulated, and the channels bound by the reverse semantic field enter the cooling state; Use the SE mechanism to construct an interactive adjustment loop between channels, and periodically evaluate and strengthen the output of the main semantic visual flow channels; If the activation degree of the reverse area detected during generation exceeds the preset threshold, immediately trigger the pruning inhibition strategy to suppress the signal flow of the relevant feature map.

[0011] Preferably, the structural-level semantic consistency judgment mechanism includes: Analyze the preset position weights and sorting orders of each semantic item in the semantic field group in the structure template; Perform regional segmentation on the generated image and extract the spatial layout map of the graphic elements; Establish a spatial mapping relationship between the semantic items and the image elements, and evaluate whether the position expression of each semantic item in the image conforms to its preset weight area; Calculate the image structure alignment score. If the score is lower than the preset threshold, it is marked as a consistency failure and the image regeneration operation is triggered.

[0012] Preferably, use a multi-modal graphic-text adversarial consistency network to perform embedding comparison on the image and the semantic field group; Train a discriminator model to identify whether the significance ranking of the main semantic items, reverse semantic items, and tonal keywords in the image is consistent with the input structure; If the main semantic item does not occupy the significant area of the image, or the reverse semantic item is frequently displayed, the discriminant result is returned as inconsistent.

[0013] Preferably, the calculation steps of the image structure alignment score include: Perform a semantic region segmentation operation on the generated image, and extract the graphic elements and their position distributions in the image coordinate system; Map the image coordinate information to the structural expected position information of each field item in the semantic field group one by one to form a semantic-image space correspondence matrix; Calculate the distance error between the actual regional position and the expected weight position of each semantic item in the correspondence matrix, and the error is the normalized spatial offset value; Perform a weighted average on the normalized spatial offset values of all semantic items to obtain the structure alignment score result. If the score is lower than the set threshold, it is regarded as structurally inconsistent and the regeneration operation is triggered.

[0014] In the above technical solution, the technical effects and advantages provided by the present invention are as follows: 1. By constructing a structured semantic field group, a twin-model collaborative generation mechanism, and a regional attention intervention strategy, the present invention realizes the deep alignment of text semantics and image expression in terms of spatial structure, emotional style, and visual focus, effectively solving the common problems of semantic misreading, theme deviation, and context mismatch in traditional graphic-text generation, and significantly improving the marketing accuracy and user perception consistency of graphic-text content.

[0015] 2. The present invention introduces a structure-level consistency scoring, significance adversarial discrimination, and regeneration mechanism, establishing a closed-loop control path from input semantic recognition to output visual verification, ensuring a high degree of matching between image content and scene intention while guaranteeing automation efficiency, and is particularly suitable for application scenarios with extremely high requirements for context expression, such as brand content customization, social media marketing, intelligent advertising design, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a method mind map of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment, please refer to Figure 1 As shown, the AI intelligent graphic-text context content precise layout full-marketing generation method described in this embodiment includes: Analyze the user input, extract the marketing intention items and their context semantics, and construct a multi-semantic field group including the main intention item, reverse semantics, auxiliary emotions, and tonality keywords; Inject the field group into a collaborative generation engine composed of an image generation sub-model and a semantic prediction sub-model, and the semantic prediction sub-model is used to predict the graphic attribute labels to be generated; During the generation process, the image generation sub-model performs regional attention intervention based on the said tags, causing the main semantic features to focus on the specified area of the image and dynamically suppressing the image expressions of reverse semantic meanings; After the image is generated, through the structural-level semantic consistency judgment mechanism, the image content structure is compared based on the position weights of semantic items within the field group. If it does not match, regeneration is performed; The image that meets the consistency requirements is combined with the original text and output as complete graphic and text situational marketing content.

[0020] On the basis of extracting the main intention item, the present invention further introduces a "reverse semantic recognition mechanism". This mechanism performs semantic mapping on the main intention item by calling the domain semantic ontology library and the extended antonym dictionary to find its opposite, negative, contrasting or emotionally reversed meaning items. For example, when the main intention item is "warm", the system will identify words such as "cold" and "icy" as its reverse semantic items and classify them through the semantic label system. This reverse item is not used for image generation, but is used to construct a generation direction limit during the image generation process to suppress the expression deviation caused by model misjudgment or semantic ambiguity. This method has stronger ambiguity suppression ability compared with the traditional Prompt embedding method.

[0021] In addition to the main intention item and the reverse semantic item, the present invention also introduces two semantic dimensions of "auxiliary emotion item" and "tonality keyword", corresponding to the emotional words (such as "relaxed", "excited", "serene", etc.) that modify the intention item in the text and the overall style description of the text (such as "urban sense", "tech style", "fresh style", etc.). The auxiliary emotion item is extracted through the emotion dictionary and the emotion classifier, which can enhance the emotional touch of the image; while the tonality keyword is extracted and clustered based on the TF-IDF plus context encoder, mainly used for subsequent image style control and visual layout structure adaptation.

[0022] The above four types of semantic elements are uniformly encapsulated as a "semantic field group". During the construction of the field group, the present invention predefines the combination order of the four types of fields to ensure its structural consistency when used as the model control input. The order is generally defined as: the first field is the main intention item, the second field is the reverse semantic item, the third field is the auxiliary emotion item, and the fourth field is the tonality keyword; structural tags or vector paragraph boundary identifiers are used between each field to ensure that they can be accurately recognized and act independently when injected into the model.

[0023] The semantic field group will finally be vectorized and input into the generation control module. During the image generation process, this field group is not only used as supplementary information for the Prompt semantics, but also participates in operations such as controlling the model attention mechanism and adjusting the feature stacking weights, realizing precise limitation of the generation direction at the initial stage of generation, and significantly reducing risks such as misunderstanding the core intention of the Prompt, context inversion, and emotion mismatch.

[0024] For example: If the user input is "Generate a picture showing a warm atmosphere for a commuter relaxing in a cold subway", the system will identify in sequence: The main intention item is: "warm atmosphere"; The reverse semantic item is: "cold"; The auxiliary emotion item is: "relaxed"; The tonality keywords are: "subway", "commuter", "night", "indoor".

[0025] Then the above information is encapsulated into a structured field group and injected into the model control end in sequence.

[0026] Compared with the single Prompt parsing or CLIP similarity control methods in the prior art, this module has the following advantages: Multi-dimensional semantic cross-validation to enhance the accuracy of core intention positioning; Introduce a reverse semantic structure to form an anti-deviation mechanism for the direction of image generation; Emotion and tonality fields participate in the control to enhance the emotional expression and aesthetic alignment of the image; The standardization of the field group structure facilitates compatibility and integration with various AI generation model structures.

[0027] In summary, the multi-semantic field group construction module is one of the key innovative links of the present invention. By systematically organizing the semantic dimensions and standardizing the input structure, it provides a solid semantic foundation support for downstream image generation and graphic layout, and has high practicality and technological advancement in marketing scenarios.

[0028] In the present invention, the semantic field group, as an input control signal, first passes through the semantic classification and encoding module, which assigns semantic type labels and field sequence numbers to the field items respectively. This structure supports four types of core semantic fields, namely: main intention item, reverse semantic item, auxiliary emotion item, and tonality keyword. Among them, the main intention and auxiliary emotion are usually used to guide the main visual configuration of image generation, while the reverse semantic and tonality keyword are used to construct the negative restrictions and style trend control of image generation.

[0029] When this field group is injected into the collaborative generation engine, it is respectively routed to two sub-model channels. The image generation sub-model is generally a diffusion generation network or a Transformer image synthesis network, which receives the main intention item and the auxiliary emotion item as the attention weighted control source to adjust the visual focus distribution and local texture expression in the generation stage. At the same time, the semantic prediction sub-model receives the reverse semantics and the tonal keyword field, and is responsible for inferring the graphic attribute labels that the entire image should contain, including the specific visual elements that should appear (such as "coffee cup", "night light", "sofa", etc.), style labels (such as "urban night view", "warm color tone"), and layout logic information (such as "central subject", "diagonal arrangement", etc.).

[0030] The semantic prediction sub-model adopts a neural network model constructed based on the multi-label classification mechanism, and its training data comes from the pre-trained data set paired with semantic descriptions and image features. The model input is the embedded field group vector, and the output is a set of structured label sets. Each label item includes the visual category, the expression intensity (which can be converted into a control weight), and the image space position suggestion. For example, when the input semantics is "warm, night, relaxing", the predicted output label may be: "yellow warm light, high intensity, near the upper part of the image", "curtain background, medium intensity, centered and slightly to the right".

[0031] This prediction result is then transmitted to the image generation sub-model in real time to regulate the guiding trajectory of the image feature map in its diffusion step. The image generation model divides the hot areas of the feature map in the initial generation stage according to the label information, that is, by controlling the distribution of its attention weights, so that the areas associated with the "high-weight labels" obtain more image configuration resources, and the areas involved in the reverse semantics will trigger the spatial cooling mechanism to suppress the appearance of the corresponding image texture.

[0032] To enhance the collaborative ability between models, the present invention introduces a "control vector sharing mechanism", that is, the semantic prediction sub-model attaches the corresponding weight information when generating the graphic attribute labels, and this weight is directly used as one of the input dimensions of the image generation sub-model in the fusion stage to achieve the semantic-visual joint decision-making path.

[0033] In addition, to avoid semantic misjudgment or generation direction drift, the present invention also proposes a "field priority callback mechanism". When there is a deviation between the predicted label and the preliminary result of image generation (for example, the generated image lacks the main intention item, or over-expresses the reverse semantic elements), the system will dynamically adjust the control priority of each field in the semantic field group according to the callback determination logic, re-allocate the field routing path, and regenerate the image. In actual deployment, this mechanism significantly reduces the probability of anti-logical generation (such as "appearing cold visual elements in a warm scene"), and is especially suitable for graphic context generation tasks that need to express abstract concepts or emotional colors.

[0034] For example, if the input field group is: Main idea: "Warmth" Opposite semantic term: "cold" Auxiliary emotional items: "Relax, Peace of Mind" Tonality keywords: "city night scene, subway commuting" The semantic prediction sub-model will output the following graphic attribute labels: "Orange light, center area, high intensity" "Person leaning on sofa, medium intensity, lower right area of the image" "Window night scene background, low intensity, upper left corner of image" “Avoid snow, ice, blue and white tones” The image generation sub-model constructs the image content based on this, and finally outputs an image of a person surrounded by warm light and with a relaxed expression against the backdrop of the night scene outside the subway window. The overall style fits the tonality keywords, and the visual focus is on the main semantic item of "warmth".

[0035] In summary, the collaborative generation engine in the present invention innovatively implements the image generation control path driven by semantic prediction, structures high-dimensional semantic content into graphic control elements, and forms a controllable mapping closed loop of semantics → attribute → image. Compared with the existing single-channel Prompt-image generation scheme, it has higher expression consistency and lower probability of semantic deviation, and is particularly suitable for high-precision image and text synthesis tasks in scenarios such as advertising, brand content creation, and immersive interactive design.

[0036] The image generation sub-model described in the present invention can adopt a diffusion model structure, a Transformer-based image synthesis model, or a generation network combined with a U-Net backbone. By introducing a regional control module and a semantic channel adjustment logic, a direct path of "semantic field → image structure" is opened up, thereby realizing image attention space reconstruction under semantic dominance.

[0037] First, the image generation sub-model receives the graphic attribute labels output by the semantic prediction sub-model. The labels contain multiple structured fields, including: target graphic category, target region spatial coordinates (such as center point position, bounding box range), sentiment relevance score and suggested color style. Before entering the image generation stage, these labels are converted into two core control variables: spatial position mask matrix and channel weight assignment vector.

[0038] The model constructs a mask matrix consistent with the image size based on the spatial coordinate information in the semantic tags. Each mask contains a two-dimensional heatmap, representing the area where a specific semantic field (such as "warm light") is suggested to appear in the image. The mask corresponding to the main intention field is given an enhancement mark (such as a high-intensity area value), and the mask corresponding to the reverse semantic field is marked as a suppression area (low-intensity value or negative value). This mask is embedded in each attention layer during the image generation process. When performing attention calculations (such as Q-K-V weighting operations), the mask value is multiplied into the normalized attention distribution as a position-sensitive weight to achieve regional-level control.

[0039] For example, during each step of feature map update in the diffusion model, the attention output value will be multiplied by the corresponding mask value to enhance the activation degree of the main intention area and weaken or zero out the channel activity of the reverse semantic area, thus visually forming a "semantic hot zone".

[0040] To further strengthen the semantic mapping structure, the present invention introduces a channel attention loop mechanism in the image generation network. Each semantic field is bound to one or more groups of feature channels in the generation model, and this mapping is obtained from the semantic-channel response statistics during the training phase. The feature channels bound to the main semantic field are configured to be dynamically upregulated, and their channel weights are strengthened through the SE module (Squeeze-and-Excitation) during each forward propagation round, while the reverse semantic binding channels are set to a cooling state, and their weights will gradually decay to the background noise level.

[0041] The specific operation is as follows: In the output of each layer of feature map, global average pooling is used to extract the global response of each channel, and then a multi-layer perceptron (MLP) is used to calculate the channel importance score, and the activation ratio of each channel is dynamically adjusted according to the semantic field weight. For example, if the main intention item is "warm", the bound channels are channels 6, 12, and 21, and their corresponding weights can be adjusted from the initial 1.0 to more than 1.5; while for the channels 8, 19, etc. corresponding to the reverse field "cold", their weights are gradually reduced to less than 0.3 in each round.

[0042] To further optimize the accuracy of regional expression, the present invention designs an attention distillation mechanism based on a teacher model. The system uses a pre-trained visual saliency recognition model as the "teacher", and this model predicts the target region saliency map (heatmap) of each field according to the input semantics. The generation sub-model, as the "student", takes the teacher heatmap as a soft supervision target. During the generation process of each layer of attention weights, the consistency deviation loss with the teacher heatmap is calculated, and the optimization process will enable the student model to learn how to more accurately respond to the semantic focus area in space.

[0043] On this basis, the present invention further introduces a reverse regional suppression mechanism. For the spatial region identified by the reverse semantics, if its activation degree exceeds a threshold value (such as being equivalent to the main semantic region) in image generation, the system automatically triggers the pruning logic, sets the pixel gradient in the relevant region feature map to zero or reduces its signal-to-noise ratio, so as to eliminate the misleading detail expression in the final output image. For example, when there are obvious frost or blue light sensations in the "cold" region, the pruning mechanism will intervene in the generation of its color channels and texture details, making it appear as a "shadow area" or a low visual attention area.

[0044] In practical applications, this regional attention intervention mechanism significantly improves the consistency of graphic and text semantic expressions. For example, in a set of graphic and text tasks with the theme of "feeling warm in the subway on a city night": When the intervention mechanism is not used, the generated images often show people wearing thin clothes, a cold-colored background or the appearance of ice and snow elements; After using the control structure of the present invention, the main character in the image is located in the central light-heated area, the background is the city night view outside the window, and the warm-colored lights form a visual guide, and the emotional expression tends to be relaxed and warm, which highly matches the text semantics.

[0045] To sum up, the present invention realizes the high stability and expression accuracy of the image generation sub-model in processing complex context semantic tasks by constructing structures such as spatial position masks, channel weight allocation, attention distillation guidance, and regional suppression, providing a stronger semantic control basis for the intelligent marketing graphic generation system.

[0046] In the AI intelligent graphic and text context content generation method described in the present invention, the image structure-level semantic consistency judgment mechanism is a key link to ensure that the generated image is strictly aligned with the input text semantics. Different from the traditional graphic and text consistency method based on semantic similarity scoring, this mechanism emphasizes "structural consistency", that is, whether each semantic field (such as the main intention, auxiliary emotion, tonality keyword, etc.) in the text is accurately expressed in the image space according to the expected position, salience, and intensity.

[0047] The key technical paths of this mechanism include: image semantic structure parsing → spatial mapping matching → salience evaluation and adversarial judgment → structure scoring → re-generation control, and multiple algorithm modules operate in coordination.

[0048] First, the system receives the input of the semantic field group and constructs a "structural expectation template" based on the field order, field weight, and structural labels therein. This template specifies the ideal spatial position (such as the center, upper right) and relative salience weight (such as high, medium, weak) and other structural targets in the image for each semantic item.

[0049] Next, the system performs region segmentation on the generated image, using a semantic segmentation model (such as the DeepLab series or BLIP) to identify the spatial boundaries and category labels of each main object in the image, and extracts the central coordinates, bounding rectangles, and graphic categories of each visual region.

[0050] After that, the system constructs a mapping matrix between semantic items and image regions, and performs matching based on semantic label similarity and spatial distance. For example, if the semantic item is "warm light" and a "light source" region is identified in the image, the system compares the central point of this region with the expected position in the template for spatial distance.

[0051] Then, spatial offset error calculation is performed on each pair of matching items: that is, after normalizing the image coordinates, the Euclidean distance between the actual position and the target position is calculated, and this distance is normalized to form a "structural offset value" between 0 and 1.

[0052] The system performs weighted averaging on the structural offset values of all semantic items according to the weights in the field group, and the result is the image structure alignment score. If this score is lower than a preset threshold (such as 0.65), the system determines that there is a structural inconsistency and performs image regeneration.

[0053] To enhance the judgment ability of structural salience expression, the present invention constructs a text-image adversarial consistency network for testing whether the image generation result truly reflects the core expression of the text's main semantic items.

[0054] This mechanism uses a set of adversarial neural networks, which includes a discriminator and an embedder: The embedder encodes the semantic field group into a high-dimensional semantic tensor, including salience rankings such as the main intention, reverse semantics, and tonal keywords; The discriminator receives the image and the embedded tensor, evaluates whether the image accurately expresses the main semantic item in the main visual focus area, and determines whether there is an "opposite meaning" that is abnormally prominently expressed in the image; The discriminator outputs a consistency label (mapped to 0 to 1 through the Sigmoid function), and the higher the value, the better the degree of semantic structure alignment; If the consistency label score is lower than the set threshold, the system marks the image as "inconsistent" and executes the field group priority callback mechanism and image regeneration.

[0055] The advantage of this adversarial judgment network is that it not only evaluates "whether there is", but also judges "whether the key points are expressed strongly enough", achieving double alignment of image salience and semantic structure.

[0056] The present invention also introduces a visual salience heatmap analysis mechanism for supplementing the local accuracy judgment of the structure consistency score. This mechanism is executed through the following steps: Use a visual saliency detection model (such as SAM or SaliencyNet) to extract the top 5 most salient regions in the image and output them in the form of a heatmap (pixel value range from 0 to 1); Match the preset structural target positions of each semantic field with the center points of the image heatmap regions and calculate the normalized offset error; Calculate the mean of the structural offsets corresponding to all fields to form the "total saliency offset value"; If this offset value exceeds a threshold (such as 0.35), the regeneration process will be triggered, and the field group will be re - sorted, with the control weight of the main intention field increased.

[0057] This analysis mechanism enables the structural offset judgment to have local perception ability, avoiding misjudging images with locally successful expressions due to too low overall scores.

[0058] After any structural consistency judgment mechanism returns "failed", the system will call the "field group priority feedback mechanism": All field items are re - sorted according to the degree of mismatch generated in the previous time; The control weight of the main intention field is increased and it is preferentially injected into the model, and the reverse semantic field triggers the generation shielding mechanism; If the scores are not up to standard for two consecutive times, the system will mark the generated image as "irreparable" and prompt the user to modify the text or simplify the semantic structure.

[0059] In summary, through the integration of multiple mechanisms such as semantic - image structure mapping, spatial score calculation, adversarial judgment, and saliency offset analysis, the present invention has realized for the first time the quantifiable control of structural - level semantic consistency in the generation of graphic - text content, greatly improving the content quality, scene fit, and user's perceived trust, and is particularly suitable for AI - generation tasks in scenarios such as high - value brand content and precise marketing graphic placement.

[0060] After the system extracts semantic fields, generates images, evaluates and screens structural consistency, a set of highly - matching image content is obtained. When combining the image with the original text at this time, it is not simply "stacking" the text and the image, but based on the alignment logic of "semantic region → image region → layout structure", realizing structured graphic - text layout.

[0061] The specific operations include: Main intention item matching points: The main focus regions in the image (such as people, light sources) need to be structurally aligned with the main semantic items in the text. For example, "warm light" should be placed above the center of the picture to form a perceptual consistency with "warm atmosphere" in the copywriting; Auxiliary Emotion Modification Structure: Modal particles reflected in the text (such as "relaxed", "at ease") and the emotional scenes in the image (such as sitting postures, lighting tones) need to have a natural transition in perception, usually achieved with soft light, blurred edge processing or illustration areas for assistance; Tonal Keyword Style Coordination: Tonal descriptions such as "urban night view", "tech atmosphere", etc. that appear in the text will be mapped to the image background, filters or font styles to form overall style consistency.

[0062] Combining the content features and expression order of the image and text, the present invention uses a visual priority layout model (such as based on the Visual Guidance Model VIM) for graphic and text layout, enabling the visual focus to present the core intention first and maintaining the fluency of reading and the clarity of marketing logic.

[0063] The specific implementation process is as follows: Content Hierarchy Division: Divide the text into logical segments such as main titles, sub-titles, call-to-action (CTA) statements, etc., and extract the main body area, auxiliary background area, etc. from the image; Layout Area Definition: Define the relative positions of graphic and text blocks on the page, such as the image centered at the upper part, the main title placed at the lower edge, the CTA button located at the lower right corner, etc.; Format Fusion Output: Use HTML, SVG or AI export formats (such as InDesign, Figma API) to combine the final editable content for automatic output of marketing materials such as H5, social media posters, e-commerce front images, etc.

[0064] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any arbitrary combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server, data center, etc. that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0065] It should be understood that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally indicates an "or" relationship between the preceding and following associated objects, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this text can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0066] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application.

Claims

1. A precise layout and full-marketing generation method for AI intelligent graphic context content, characterized in that: including: Parse the user input, extract the marketing intention items and their contextual semantics, and construct a multi-semantic field group containing the main intention item, reverse semantics, auxiliary emotion, and tonality keywords; Inject the field group into a collaborative generation engine composed of an image generation sub-model and a semantic prediction sub-model. The semantic prediction sub-model is used to predict the graphic attribute labels to be generated; During the generation process, the image generation sub-model performs regional attention intervention according to the labels, making the main semantic features focus on the specified area of the image and dynamically suppressing the image expression of reverse semantic meanings; After generating the image, through a structural-level semantic consistency judgment mechanism, compare the image content structure based on the position weights of the semantic items in the field group. If they do not match, perform re-generation; Combine the image that meets the consistency requirements with the original text and output it as complete graphic context marketing content.

2. The AI intelligent graphic context content precise layout full-marketing generation method according to claim 1, wherein: The parsing of the user input and the construction of the multi-semantic field group include: Perform syntactic-level parsing on the input text, identify the subject-verb-object core structure representing the marketing intention, and extract the main intention item therein; Based on the main intention item, call the reverse thesaurus and the domain ontology library, retrieve its logical opposite item and mark it as reverse semantics; Identify the emotional words and intonation descriptions used to modify the main intention item through context analysis, and classify them as auxiliary emotion items and tonality keywords respectively; Encapsulate the main intention item, reverse semantics, auxiliary emotion item, and tonality keywords into a multi-field structure, and combine them into a semantic field group vector according to the predefined field order.

3. The AI intelligent graphic context content precise layout full-marketing generation method according to claim 1, characterized in that: Injecting the semantic field group into a collaborative generation engine composed of an image generation sub-model and a semantic prediction sub-model includes: Perform field identification and semantic labeling on each field item of the semantic field group; Construct a multi-segment control vector according to the predefined field order, and allocate each field item to different model input channels; Inject the main intention field and the auxiliary emotion field into the control channel of the image generation sub-model; Input the reverse semantic field and the tonality keyword field into the semantic prediction sub-model for generating graphic restriction labels and layout style prediction labels.

4. The AI intelligent graphic context content precise layout full marketing generation method according to claim 3, characterized in that: The semantic prediction sub-model for predicting graphic attribute labels includes: Receive the structured semantic field group vector as input; Based on the multi-label classification model established by the training semantic-visual contrast sample library, identify the potential image elements corresponding to the semantic fields; Map each semantic field item to the corresponding graphic attribute label, including the type of visual element, position weight, and emotional expression characteristics; Output a structured set of graphic attribute labels as the control constraint basis for the image generation sub-model.

5. The AI intelligent graphic context content precise layout full-marketing generation method according to claim 1, characterized in that: The image generation sub-model performs regional attention intervention during the generation process includes: Receive the graphic attribute labels output by the semantic prediction sub-model, and extract the spatial position coordinates and weights corresponding to each semantic field; Construct a multi-channel spatial position mask matrix, assign an enhancement mark to the position mask of the main intention field, and assign a suppression mark to the position mask of the reverse semantic field; During the diffusion generation process, use the spatial position mask to weightedly adjust the normalized weights of the attention mapping layer; When generating the target area in the image, adjust the channel activation threshold in real time.

6. The AI intelligent graphic context content precise layout full-marketing generation method according to claim 5, wherein: Regional attention intervention includes: Establish semantic fields during the generation process, that is, the image channel mapping relationship, and bind the semantic weights to the feature channels through the channel attention mechanism; The activation intensity of the channels bound by the main intention field is dynamically up-regulated, and the channels bound by the reverse semantic field enter the cooling state; Use the SE mechanism to construct an interactive adjustment loop between channels, and periodically evaluate and strengthen the output of the main semantic visual flow channels; If the activation degree of the reverse area is detected to exceed the preset threshold during generation, immediately trigger the pruning inhibition strategy to suppress the signal flow of the relevant feature maps.

7. The AI intelligent graphic context content precise layout full-marketing generation method according to claim 1, wherein: The structural-level semantic consistency judgment mechanism includes: Analyze the preset position weights and sorting orders of each semantic item in the semantic field group in the structure template; Perform regional segmentation on the generated image and extract the spatial layout map of the graphic elements; Establish the spatial mapping relationship between the semantic items and the image elements, and evaluate whether the position expression of each semantic item in the image conforms to its preset weight area; Calculate the image structure alignment score. If the score is lower than the preset threshold, it is marked as a consistency failure and the image regeneration operation is triggered.

8. The AI intelligent graphic context content precise layout full-marketing generation method according to claim 7, wherein: Among them, the judgment of the consistency between the image content structure and the semantic field group includes: Use the multi-modal graphic-text adversarial consistency network to perform embedding comparison on the image and the semantic field group; Train the discriminator model to identify whether the significance ranking of the main semantic items, reverse semantic items, and tonal keywords in the image is consistent with the input structure; If the main semantic item does not occupy the significant area of the image, or the reverse semantic item is frequently displayed, the discriminant result is returned as inconsistent.

9. The AI intelligent graphic context content precise layout full-marketing generation method according to claim 8, characterized in that: The calculation steps of the image structure alignment score include: Perform semantic region segmentation operations on the generated image, and extract the graphic elements and their position distributions in the image coordinate system; Map the image coordinate information one by one with the structural expected position information of each field item in the semantic field group to form a semantic-image space correspondence matrix; Perform distance error calculation on the actual regional position and the expected weight position of each semantic item in the correspondence matrix, and the error is the normalized spatial offset value; Perform weighted averaging on the normalized spatial offset values of all semantic items to obtain the structure alignment score result. If the score is lower than the set threshold, it is regarded as structurally inconsistent and the regeneration operation is triggered.

Citation Information

Patent Citations

  • Image semantic segmentation method and system based on differentiated contexts

    CN117253034A

  • Text association type short video multi-mode emotion recognition method and system

    CN117636196A

  • Image generation method and system, electronic equipment and readable storage medium

    CN117808923A

  • Cross-modal multi-text guided image generation method based on comparative learning

    CN118447132A

  • Vision-text collaborative abstract generation method and system based on multi-modal learning

    CN119862861A

Cited By

  • Marketing video auditing method based on AI

    CN120583273A

  • Image meaning analysis scene consistency evaluation system based on visual model

    CN120997650A

  • Scene consistency evaluation system for image meaning resolution based on visual model

    CN120997650B

  • Self-adaptive agent service platform based on unstructured data intelligent analysis

    CN121170345A

  • New media intelligent channel switching method, system and device and medium

    CN121561074A