AI intelligent graphic and text situational content precise layout full marketing generation method
By building a multi-semantic field group and a collaborative generation engine, the semantic misreading and situational mismatch problems in graphic and text generation are solved, the deep alignment between images and text is achieved, and the accuracy of marketing graphic and text and user perception consistency is improved.
Patent Information
- Application Number
- CN202510864081.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-26
AI Technical Summary
In the prior art, the graphic and text generation model is prone to misunderstand the semantics of "negation" or "contrast" in the copy, and generates images that are opposite to the expected situation, resulting in weakening the sensual motivation and conversion effect of marketing graphic and text.
Build a multi-semantic field group containing idea items, reverse semantics, auxiliary emotions and tonal keywords. Through the collaborative generation engine of the image generation sub-model and the semantic prediction sub-model, regional attention intervention and structural-level semantic consistency judgment are performed to ensure the depth alignment of the image and text semantics.
It significantly improves the marketing accuracy and user perception consistency of graphic and text content, and is especially suitable for application scenarios with high contextual expression requirements such as brand content customization, social media marketing and intelligent advertising design.
Smart Images

Figure CN120374783B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image and text recognition technology, and specifically to an AI intelligent image and text situational content precise layout full marketing generation method. Background Art
[0002] AI-powered, contextually accurate content generation for graphics and text across all marketing scenarios leverages artificial intelligence to intelligently generate graphics and text based on specific marketing scenarios. This approach, through precise content placement strategies, enables omnichannel, integrated marketing communications. This approach combines contextual understanding, content creation, and optimized distribution to improve marketing efficiency and user conversion.
[0003] The existing technology has the following shortcomings:
[0004] The generation of contextually counterintuitive images is a serious technical issue currently faced by diffusion models in image and text generation. Specifically, the model misinterprets descriptions with "negative" or "contrasting" semantics in the copy, resulting in images that contradict the intended context. For example, in a scenario titled "Emphasize warmth on a cold winter night in the subway," the model might focus on the literal keyword "cold," generating an image filled with icy elements and scantily clad characters, completely deviating from the intended "warm" atmosphere. This deviation stems from the ambiguity of the prompt's semantic expression and insufficient coverage of contrasting scene samples during the model's training phase. Ultimately, this can seriously mislead users' emotional cognition, weakening the emotional appeal and conversion effectiveness of marketing images and text. Summary of the Invention
[0005] The purpose of the present invention is to provide an AI intelligent graphic and text situational content precise layout full marketing generation method to address the shortcomings of the background technology.
[0006] To achieve the above objectives, the present invention provides the following technical solutions: an AI intelligent graphic and text contextual content precise layout full marketing generation method, comprising:
[0007] Parse user input, extract marketing intent terms and their contextual semantics, and construct a multi-semantic field group containing the main intent term, reverse semantics, auxiliary emotions, and tonality keywords;
[0008] Injecting the field group into a collaborative generation engine composed of an image generation sub-model and a semantic prediction sub-model, wherein the semantic prediction sub-model is used to predict the graphic attribute labels to be generated;
[0009] The image generation sub-model performs regional attention intervention according to the labels during the generation process, so that the main semantic features are focused on the specified area of the image and the image expression of the reverse meaning is dynamically suppressed;
[0010] After the image is generated, the structure-level semantic consistency judgment mechanism is used to compare the image content structure based on the position weights of the semantic items in the field group. If there is any discrepancy, regeneration is performed;
[0011] Combine images that meet consistency requirements with original text and output them as complete graphic and text contextual marketing content.
[0012] Preferably, parsing user input and constructing a multi-semantic field group includes:
[0013] Perform syntactic parsing on the input text to identify the subject-verb-object core structure that represents marketing intent and extract the main intent item.
[0014] Based on the main image item, calling the reverse semantic dictionary and the domain ontology library, retrieving its logical opposite item and marking it as reverse semantics;
[0015] Through contextual analysis, emotional words and intonation descriptions used to modify the main image items are identified and classified into auxiliary emotional items and tonal keywords respectively;
[0016] The main image item, reverse semantics, auxiliary emotional item and tonality keyword are encapsulated into a multi-field structure and combined into a semantic field group vector according to a predefined field order.
[0017] Preferably, injecting the semantic field group into the collaborative generation engine composed of the image generation sub-model and the semantic prediction sub-model includes:
[0018] Perform field identification and semantic labeling on each field item of the semantic field group;
[0019] Construct a multi-segment control vector according to the predefined field order and assign each field item to different model input channels;
[0020] Inject the main image field and auxiliary emotion field into the image generation sub-model control channel;
[0021] The reverse semantic field and the tonality keyword field are input into the semantic prediction sub-model to generate graphic restriction labels and layout style prediction labels.
[0022] Preferably, the semantic prediction sub-model for predicting graphic attribute labels includes:
[0023] Receives a structured semantic field group vector as input;
[0024] A multi-label classification model built based on a training semantic-visual comparison sample library identifies potential image elements corresponding to semantic fields;
[0025] Map each semantic field item to a corresponding graphic attribute label, including visual element type, position weight, and emotional expression characteristics;
[0026] Output a set of structured graphic attribute labels as the control constraint basis for the image generation sub-model.
[0027] Preferably, the image generation sub-model performs regional attention intervention during the generation process, including:
[0028] Receive the graphic attribute labels output by the semantic prediction sub-model and extract the spatial position coordinates and weights corresponding to each semantic field;
[0029] Construct a multi-channel spatial position mask matrix, assign the position mask of the main semantic field to the enhancement mark, and assign the position mask of the reverse semantic field to the suppression mark;
[0030] During the diffusion generation process, the spatial position mask is used to perform weighted adjustment on the normalized weights of the attention map layer;
[0031] As target regions are generated in the image, channel activation thresholds are adjusted in real time.
[0032] Preferably, regional attention interventions include:
[0033] During the generation process, semantic fields are established, i.e., image-channel mapping relationships, and semantic weights are bound to feature channels through the channel attention mechanism.
[0034] The activation intensity of the channel bound to the main semantic field is dynamically increased, and the channel bound to the reverse semantic field enters a cooling state;
[0035] The SE mechanism is used to construct an inter-channel interaction regulation loop to periodically evaluate and strengthen the output of the main semantic visual flow channel;
[0036] If the activation of the reverse region detected during generation exceeds the preset threshold, the pruning suppression strategy is immediately triggered to suppress the flow of related feature map signals.
[0037] Preferably, the structure-level semantic consistency judgment mechanism includes:
[0038] Analyze the preset position weight and sorting order of each semantic item in the semantic field group in the structure template;
[0039] Perform region segmentation on the generated image and extract the spatial layout of graphic elements;
[0040] Establish a spatial mapping relationship between semantic items and image elements, and evaluate whether the position expression of each semantic item in the image conforms to its preset weight range;
[0041] Calculate the image structure alignment score. If the score is lower than the preset threshold, it is marked as consistency failure and triggers the image regeneration operation.
[0042] Preferably, a multimodal image-text adversarial consistency network is used to perform embedding comparison between the image and the semantic field group;
[0043] The discriminator model is trained to identify whether the saliency ranking of the main semantic terms, reverse semantic terms, and tonal keywords in the image is consistent with the input structure;
[0044] If the main semantic item does not occupy a salient area of the image, or the reverse semantic item is frequently displayed, the discrimination result is returned as inconsistent.
[0045] Preferably, the steps of calculating the image structure alignment score include:
[0046] Perform semantic region segmentation on the generated image to extract graphic elements and their position distribution in the image coordinate system;
[0047] Mapping the image coordinate information to the expected structural position information of each field item in the semantic field group one by one to form a semantic-image space correspondence matrix;
[0048] Performing distance error calculation between the actual region position of each semantic item in the corresponding matrix and its expected weight position, wherein the error is a normalized spatial offset value;
[0049] The normalized spatial offset values of all semantic items are weighted averaged to obtain the structural alignment score. If the score is lower than the set threshold, it is considered structural inconsistency and the regeneration operation is triggered.
[0050] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0051] 1. By constructing a structured semantic field group, a twin-model collaborative generation mechanism, and a regional attention intervention strategy, the present invention achieves deep alignment between text semantics and image expression in terms of spatial structure, emotional style, and visual focus. This effectively solves the common problems of semantic misreading, theme deviation, and context mismatch in traditional image and text generation, and significantly improves the marketing accuracy and user perception consistency of image and text content.
[0052] 2. This invention introduces structural consistency scoring, saliency adversarial discrimination, and regeneration mechanisms, establishing a closed-loop control path from input semantic recognition to output visual verification. This ensures a high degree of match between image content and scene intent while guaranteeing automation efficiency. It is particularly suitable for application scenarios with extremely high requirements for contextual expression, such as brand content customization, social media marketing, and intelligent advertising design. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0054] Figure 1 This is a mind map of the method of the present invention. DETAILED DESCRIPTION
[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0056] For examples, see Figure 1 As shown, the AI intelligent graphic and text context content precise layout full marketing generation method described in this embodiment includes:
[0057] Parse user input, extract marketing intent terms and their contextual semantics, and construct a multi-semantic field group containing the main intent term, reverse semantics, auxiliary emotions, and tonality keywords;
[0058] Injecting the field group into a collaborative generation engine composed of an image generation sub-model and a semantic prediction sub-model, wherein the semantic prediction sub-model is used to predict the graphic attribute labels to be generated;
[0059] The image generation sub-model performs regional attention intervention according to the labels during the generation process, so that the main semantic features are focused on the specified area of the image and the image expression of the reverse meaning is dynamically suppressed;
[0060] After the image is generated, the structure-level semantic consistency judgment mechanism is used to compare the image content structure based on the position weights of the semantic items in the field group. If there is any discrepancy, regeneration is performed;
[0061] Combine images that meet consistency requirements with original text and output them as complete graphic and text contextual marketing content.
[0062] On the basis of extracting the main image item, the present invention further introduces a "reverse semantic recognition mechanism". This mechanism performs semantic mapping on the main image item by calling the domain semantic ontology library and the extended antonym dictionary, and searches for its opposite, negation, contrast or emotion reversal meaning items. For example, when the main image item is "warm", the system will identify words such as "cold" and "icy" as its reverse semantic items and classify them through the semantic labeling system. This reverse item is not used for image generation, but is used to construct generation directional constraints in the image generation process to suppress expression deviations caused by model misjudgment or semantic ambiguity. Compared with the traditional Prompt embedding method, this method has stronger ambiguity suppression capabilities.
[0063] In addition to the primary image term and the reverse semantic term, this paper also introduces two semantic dimensions: "auxiliary emotional terms" and "tonal keywords." These correspond to emotional terms that modify the primary image term (e.g., "relaxed," "excited," "serene," etc.) and overall textual descriptions (e.g., "urban," "technical," "fresh"), respectively. Auxiliary emotional terms are extracted using an emotional lexicon and sentiment classifier to enhance the emotional impact of the image. Tonal keywords, on the other hand, are extracted and clustered using TF-IDF and context encoders, primarily for subsequent image style control and visual layout structure adaptation.
[0064] The above four types of semantic elements are uniformly encapsulated as "semantic field groups." During the field group construction process, the present invention predefines the combination order of the four types of fields to ensure structural consistency when used as model control input. This order is generally defined as follows: the first field is the primary intention item, the second field is the reverse semantic item, the third field is the auxiliary emotional item, and the fourth field is the tonal keyword; structural tags or vector paragraph boundaries are used between each field to ensure that they can be accurately identified and act independently when injected into the model.
[0065] The semantic field group is ultimately vectorized and fed into the generation control module. During image generation, this field group not only serves as supplementary information for the prompt's semantics but also controls the model's attention mechanism and adjusts feature layer weights. This allows for precise control of the generation direction at the initial stage, significantly reducing the risks of misunderstanding the prompt's core intent, contextual inversion, and emotional mismatch.
[0066] For example, if the user input is "Please generate a warm image of an office worker relaxing in a cold subway," the system will identify the following in order:
[0067] The main image item is: "Warm atmosphere";
[0068] The reverse semantic item is: "cold";
[0069] The auxiliary emotion items are: “relaxation”;
[0070] The key words for the tone are: “subway”, “office workers”, “night”, and “indoor”.
[0071] Then encapsulate the above information into a structured field group and inject it into the model controller in sequence.
[0072] Compared with the existing single prompt parsing or CLIP similarity control methods, this module has the following advantages:
[0073] Multi-dimensional semantic cross-validation enhances the accuracy of core intent positioning;
[0074] Introducing reverse semantic structure to form a directional anti-bias mechanism for image generation;
[0075] The emotion and tonality fields participate in the control, enhancing the image’s perceptual expression and aesthetic alignment;
[0076] The field group structure is standardized to facilitate compatibility and integration with various AI generation model structures.
[0077] In summary, the multi-semantic field group construction module is one of the key innovative links of the present invention. By systematically organizing semantic dimensions and standardizing input structures, it provides a solid semantic foundation for downstream image generation and graphic layout, and is highly practical and technologically advanced in marketing scenarios.
[0078] In this paper, a semantic field group serves as an input control signal and first passes through a semantic classification encoding module, where each field item is assigned a semantic type label and a field sequence number. This structure supports four core semantic fields: primary image items, reverse semantic items, auxiliary emotion items, and tonal keywords. Primary image items and auxiliary emotions are typically used to guide the primary visual configuration of image generation, while reverse semantic items and tonal keywords are used to establish negative constraints and control stylistic tendencies in image generation.
[0079] When injected into the collaborative generation engine, this field group is routed to two sub-model channels. The image generation sub-model is typically a diffusion generation network or a Transformer image synthesis network. It receives the main image term and auxiliary emotion term as attention-weighted control sources, which are used to adjust the visual focus distribution and local texture expression during the generation phase. Meanwhile, the semantic prediction sub-model receives the reverse semantics and tonal keyword fields and is responsible for inferring the graphic attribute labels that the entire image should contain, including specific visual elements (such as "coffee cup," "night light," "sofa," etc.), style labels (such as "urban night scene," "warm tones"), and layout logic information (such as "central subject," "diagonal arrangement," etc.).
[0080] The semantic prediction sub-model uses a neural network model built based on a multi-label classification mechanism. Its training data comes from a pre-trained dataset that pairs semantic descriptions with image features. The model input is an embedded field group vector, and the output is a structured set of labels. Each label item includes a visual category, a strength (which can be converted into a control weight), and a suggested spatial location in the image. For example, when the input semantics are "warm, night, relaxing," the predicted output labels might be: "Yellow warm light, high intensity, near the top of the image" or "Curtain background, medium intensity, centered to the right."
[0081] The prediction results are then transmitted in real time to the image generation sub-model to control the guidance trajectory of the image feature map during its diffusion step. Based on the label information, the image generation model divides the feature map into hot zones during the initial generation phase. This involves controlling the distribution of attention weights, so that regions associated with "high-weight labels" receive more image configuration resources. Regions with negative semantics trigger a spatial cooling mechanism, suppressing the appearance of corresponding image texture.
[0082] To enhance the inter-model collaboration capability, the present invention introduces a "control vector sharing mechanism", that is, the semantic prediction sub-model also includes corresponding weight information when generating graphic attribute labels. This weight is directly used as one of the input dimensions of the image generation sub-model in the fusion stage, realizing a semantic-visual joint decision-making path.
[0083] Furthermore, to avoid semantic misjudgments or drift in generation direction, the present invention proposes a "field priority callback mechanism." When a discrepancy occurs between the predicted label and the initial image generation result (for example, the generated image lacks the primary semantic element or overrepresents the inverse semantic element), the system dynamically adjusts the control priority of each field in the semantic field group based on the callback judgment logic, reallocates the field routing path, and regenerates the image. In actual deployments, this mechanism significantly reduces the probability of counter-logical generation (e.g., "cold visual elements appearing in warm scenes") and is particularly suitable for image-text context generation tasks that require expressing abstract concepts or emotional overtones.
[0084] For example, if the input field group is:
[0085] Main item: "Warmth"
[0086] Opposite semantic term: "cold"
[0087] Auxiliary emotional items: "Relaxation, peace of mind"
[0088] Tonal keywords: "city night scene, subway commuting"
[0089] The semantic prediction sub-model will output the following graphic attribute labels:
[0090] "Orange light, center area, high intensity"
[0091] "Person leaning on sofa, medium intensity, lower right area of the image"
[0092] "Window night scene background, low intensity, upper left corner of image"
[0093] “Avoid snow, ice, blue and white tones”
[0094] The image generation sub-model constructs the image content based on this, and ultimately outputs an image of a person surrounded by warm light and with a relaxed expression against the backdrop of the night view outside the subway window. The overall style fits the tonal keywords, and the visual focus is concentrated on the main semantic item of "warmth".
[0095] In summary, the collaborative generation engine in the present invention innovatively implements an image generation control path driven by semantic prediction, structures high-dimensional semantic content into graphic control elements, and forms a controllable mapping closed loop of semantics → attributes → image. Compared with the existing single-channel Prompt-image generation scheme, it has higher expression consistency and lower probability of semantic deviation, and is particularly suitable for high-precision image and text synthesis tasks in scenarios such as advertising, brand content creation, and immersive interactive design.
[0096] The image generation sub-model described in the present invention can adopt a diffusion model structure, a Transformer-based image synthesis model, or a generation network combined with a U-Net backbone. By introducing a regional control module and semantic channel adjustment logic, it opens up a direct path from "semantic field → image structure" and realizes image attention space reconstruction under semantic dominance.
[0097] First, the image generation sub-model receives the image attribute labels output by the semantic prediction sub-model. These labels contain multiple structured fields, including the target image category, the spatial coordinates of the target region (such as the center point location and bounding box extent), the sentiment relevance score, and the suggested color style. Before entering the image generation stage, these labels are converted into two core control variables: a spatial position mask matrix and a channel weight assignment vector.
[0098] Based on the spatial coordinate information in the semantic labels, the model constructs a mask matrix of the same size as the image. Each mask contains a two-dimensional heat map representing the region in the image where a specific semantic field (such as "warm light") is suggested to appear. The mask corresponding to the main semantic field is assigned an enhancement label (such as a high-intensity area value), while the mask corresponding to the opposite semantic field is marked as a suppression area (low-intensity or negative value). This mask is embedded in each attention layer during the image generation process. During attention calculations (such as QKV weighting operations), the mask value is multiplied into the normalized attention distribution as a position-sensitive weight to achieve region-level control.
[0099] For example, when the feature map is updated at each step of the diffusion model, the attention output value will be multiplied by the corresponding mask value, enhancing the activation of the main image area and weakening or zeroing the channel activity of the reverse semantic area, thereby forming a "semantic hot zone" visually.
[0100] To further strengthen the semantic mapping structure, this paper introduces a channel attention loop mechanism into the image generation network. Each semantic field is bound to one or more sets of feature channels in the generation model. This mapping is derived from the semantic-channel response statistics during the training phase. The feature channels bound to the primary semantic field are configured to be dynamically up-regulated, with their channel weights strengthened in each forward propagation round via the SE module (Squeeze-and-Excitation). The reverse semantic binding channels are set to a cooling state, with their weights gradually decaying to the background noise level.
[0101] The specific operation is as follows: In each layer of feature map output, global average pooling is used to extract the global response of each channel. Then, a multi-layer perceptron (MLP) is used to calculate the channel importance score. The activation ratio of each channel is dynamically adjusted according to the semantic field weight. For example, if the main image item is "warm", the bound channels are channels 6, 12, and 21. Their corresponding weights can be adjusted from the initial 1.0 to above 1.5. Meanwhile, the weights of channels 8, 19, and so on corresponding to the reverse field "cold" are reduced to below 0.3 in each round.
[0102] To further optimize regional representation accuracy, this paper designs an attention distillation mechanism based on a teacher model. The system uses a pretrained visual saliency recognition model as a "teacher," which predicts a saliency map (heatmap) of the target region for each field based on the input semantics. A sub-model is generated as a "student," using the teacher's heatmap as a soft-supervisory target. During the generation of attention weights at each layer, the consistency deviation loss with the teacher's heatmap is calculated. This optimization process enables the student model to learn how to more accurately spatially respond to semantically focused regions.
[0103] Building on this foundation, the present invention further introduces a reverse region suppression mechanism. For spatial regions identified by reverse semantics, if their activation exceeds a threshold during image generation (e.g., comparable to that of the primary semantic region), the system automatically triggers pruning logic, setting pixel gradients in the feature map of the relevant region to zero or reducing their signal-to-noise ratio, thereby eliminating misleading detail in the final output image. For example, if a "cold" region exhibits noticeable frost or a blue light effect, the pruning mechanism will interfere with the generation of its color channels and texture details, rendering it as a "dark shadow" or low-attention area.
[0104] In practical applications, this regional attention intervention mechanism significantly improves the consistency of semantic expression between images and text. For example, in a set of image-text tasks with the theme of "feeling warmth in the subway at night in the city":
[0105] When no intervention mechanism is used, the generated images often show characters wearing thin clothes, cold backgrounds, or snowy elements;
[0106] After using the control structure of the present invention, the main body of the person in the image is located in the central light hot zone, with the background being the city night view outside the window. The warm light forms a visual guide, and the emotional expression tends to be relaxed and warm, which is highly consistent with the text semantics.
[0107] In summary, the present invention achieves high stability and expression accuracy of the image generation sub-model when processing complex situational semantic tasks by constructing structures such as spatial position masking, channel weight allocation, attention distillation guidance and regional suppression, providing a stronger semantic control foundation for the intelligent marketing image and text generation system.
[0108] In the AI-powered method for generating contextual content for images and text, a mechanism for determining semantic consistency at the image structure level is crucial for ensuring strict semantic alignment between the generated image and the input text. Unlike traditional image-text consistency methods based on semantic similarity scoring, this mechanism emphasizes "structural consistency," specifically, whether each semantic field in the text (such as the primary image, supporting emotions, and tonal keywords) is accurately expressed in the image space with the expected position, salience, and intensity.
[0109] The key technical paths of this mechanism include: image semantic structure analysis → spatial mapping matching → saliency assessment and adversarial judgment → structure scoring → regeneration control, combined with multiple algorithm modules to operate collaboratively.
[0110] First, the system receives a set of semantic fields as input and constructs a "structural expectation template" based on the field order, field weights, and structural labels. This template specifies structural targets for each semantic item, such as its ideal spatial position in the image (e.g., center, upper right), and relative salience weight (e.g., high, medium, weak).
[0111] Next, the system performs region segmentation on the generated image, using a semantic segmentation model (such as the DeepLab series or BLIP) to identify the spatial boundaries and category labels of the main objects in the image, and extract the center coordinates, bounding rectangle, and graphic category of each visual area.
[0112] The system then constructs a mapping matrix between semantic terms and image regions, matching them based on semantic label similarity and spatial distance. For example, if the semantic term is "warm light" and the "light source" region is identified in the image, the system compares the spatial distance between the center point of this region and the expected position in the template.
[0113] Then, a spatial offset error calculation is performed on each pair of matches: that is, after normalization with the image coordinates, the Euclidean distance between the actual position and the target position is calculated, and the distance is normalized to form a "structural offset value" between 0 and 1.
[0114] The system performs a weighted average of the structural offset values of all semantic items according to the weights in the field group. The result is the image structural alignment score. If the score falls below a preset threshold (such as 0.65), the system determines that there is a structural inconsistency and performs image regeneration.
[0115] In order to enhance the ability to judge the expression of structural significance, the present invention constructs an image-text adversarial consistency network to verify whether the image generation result truly reflects the core expression of the main semantic items of the text.
[0116] The mechanism uses a set of adversarial neural networks, which consists of a discriminator and an embedder:
[0117] The embedder encodes the semantic field group into a high-dimensional semantic tensor, which contains the main image, reverse semantics, tonal keywords and other saliency rankings;
[0118] The discriminator receives the image and the embedding tensor, evaluates whether the image accurately expresses the main semantic term in the main visual focus area, and determines whether there is any "reverse meaning" that is abnormally expressed in the image;
[0119] The discriminator outputs a consistency label (mapped from 0 to 1 by the Sigmoid function), where a higher value indicates a better degree of semantic structure alignment;
[0120] If the consistency label score is lower than the set threshold, the system marks the image as "inconsistent" and executes the field group priority callback mechanism and image regeneration.
[0121] The advantage of this adversarial judgment network is that it not only evaluates "whether it exists", but also judges "whether the key points are expressed strongly enough", achieving dual alignment of image saliency and semantic structure.
[0122] The present invention also introduces a visual saliency heatmap analysis mechanism to supplement the local accuracy judgment of the structural consistency score. This mechanism is implemented by the following steps:
[0123] Use a visual saliency detection model (such as SAM or SaliencyNet) to extract the top 5 most salient regions in the image and output them as a heat map (with values of 0 to 1 per pixel).
[0124] Match the preset structural target position of each semantic field with the center point of the image heat map area and calculate the normalized offset error;
[0125] Calculate the mean of the structural deviations corresponding to all fields to form the "total significant deviation value";
[0126] If the offset value exceeds a threshold (e.g., 0.35), the regeneration process is triggered, and the field groups are reordered, with the control weight of the main map field being increased.
[0127] This analysis mechanism makes the structure shift judgment locally aware and avoids misjudging images with successful local expression due to low overall scores.
[0128] When any structural consistency judgment mechanism returns "failed", the system will call the "field group priority feedback mechanism":
[0129] All field items are reordered according to the mismatch degree generated in the previous step;
[0130] The main semantic field increases the control weight and is injected into the model first, while the reverse semantic field triggers the generation shielding mechanism;
[0131] If the ratings fail to meet the standards for two consecutive times, the system will mark the generated image as “unrepairable” and prompt the user to modify the text or simplify the semantic structure.
[0132] In summary, the present invention integrates multiple mechanisms such as semantic-image structure mapping, spatial score calculation, adversarial judgment and significance offset analysis, and for the first time realizes quantifiable control of structural-level semantic consistency in graphic and text content generation, greatly improving content quality, scene fit and user perceived trust. It is particularly suitable for AI generation tasks in scenarios such as high-value brand content and precision marketing graphic and text delivery.
[0133] After semantic field extraction, image generation, structural consistency assessment, and screening, the system obtains a set of highly matched image content. When combining the image with the original text, it doesn't simply "stack" the text and image together. Instead, it uses the alignment logic of "semantic region → image region → typesetting structure" to achieve structured image and text arrangement.
[0134] Specific operations include:
[0135] Main image item matching point: The main focal area in the image (such as people, light source) must be structurally aligned with the main semantic item in the text. For example, "warm light" should be placed above the center of the image to create perceptual consistency with the "warm atmosphere" in the copy.
[0136] Auxiliary emotional modification structure: The transition between modal particles in the text (such as "relax" and "peace of mind") and the emotional scenes in the image (such as sitting posture and lighting tones) needs to be natural. This is usually achieved through soft lighting, blurred edge processing, or illustration areas.
[0137] Tonal keyword style coordination: Tonal descriptions such as "urban night scene" and "technological atmosphere" appearing in the text will be mapped to the image background, filter or font style to achieve overall style consistency.
[0138] Combining the content features and expression order of images and texts, the present invention adopts a visual priority layout model (for example, based on the visual guidance model VIM) for graphic and text typesetting, so that the visual focus gives priority to presenting the core intention and maintains reading fluency and clarity of marketing logic.
[0139] The specific implementation process is as follows:
[0140] Content hierarchy: Divide text into logical segments such as main title, subtitle, and call to action (CTA), and extract the main area and auxiliary background area from the image;
[0141] Layout area definition: defines the relative position of images and text blocks on the page, such as centering the image at the top, placing the main title at the bottom edge, and placing the CTA button in the lower right corner;
[0142] Format-fusion output: Combine final editable content using HTML, SVG, or AI export formats (such as InDesign and Figma API) for automatic output of marketing materials such as H5, social media posters, and e-commerce homepage images.
[0143] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0144] It should be understood that the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist at the same time, and B exists alone, where A and B may be singular or plural. In addition, the character " / " herein generally indicates that the objects associated with each other are in an "or" relationship, but it may also indicate an "and / or" relationship, which can be understood by referring to the context. A person of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0145] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. AI intelligent graphic and text situational content precise layout full marketing generation method, characterized by: include: Parse user input, extract marketing intent terms and their contextual semantics, and construct a multi-semantic field group containing the main intent term, reverse semantics, auxiliary emotions, and tonality keywords; Injecting the field group into a collaborative generation engine composed of an image generation sub-model and a semantic prediction sub-model, wherein the semantic prediction sub-model is used to predict the graphic attribute labels to be generated; The image generation sub-model performs regional attention intervention according to the labels during the generation process, so that the main semantic features are focused on the specified area of the image and the image expression of the reverse meaning is dynamically suppressed; After the image is generated, the structure-level semantic consistency judgment mechanism is used to compare the image content structure based on the position weights of the semantic items in the field group. If there is any discrepancy, regeneration is performed; Combine images that meet consistency requirements with original text and output them as complete graphic and text contextual marketing content.
2. The AI intelligent graphic and text contextual content precise layout full marketing generation method according to claim 1 is characterized by: Parsing user input and constructing multiple semantic field groups includes: Perform syntactic parsing on the input text to identify the subject-verb-object core structure that represents marketing intent and extract the main intent item. Based on the main image item, calling the reverse semantic dictionary and the domain ontology library, retrieving its logical opposite item and marking it as reverse semantics; Through contextual analysis, emotional words and intonation descriptions used to modify the main image items are identified and classified into auxiliary emotional items and tonal keywords respectively; The main image item, reverse semantics, auxiliary emotional item and tonality keyword are encapsulated into a multi-field structure and combined into a semantic field group vector according to a predefined field order.
3. The AI intelligent graphic and text contextual content precise layout full marketing generation method according to claim 1 is characterized by: Injecting semantic field groups into the collaborative generation engine consisting of the image generation sub-model and the semantic prediction sub-model includes: Perform field identification and semantic labeling on each field item of the semantic field group; Construct a multi-segment control vector according to the predefined field order and assign each field item to different model input channels; Inject the main image field and auxiliary emotion field into the image generation sub-model control channel; The reverse semantic field and the tonality keyword field are input into the semantic prediction sub-model to generate graphic restriction labels and layout style prediction labels.
4. The AI intelligent graphic and text contextual content precise layout full marketing generation method according to claim 3 is characterized by: The semantic prediction sub-model is used to predict graphic attribute labels including: Receives a structured semantic field group vector as input; A multi-label classification model built based on a training semantic-visual comparison sample library identifies potential image elements corresponding to semantic fields; Map each semantic field item to a corresponding graphic attribute label, including visual element type, position weight, and emotional expression characteristics; Output a set of structured graphic attribute labels as the control constraint basis for the image generation sub-model.
5. The AI intelligent graphic and text contextual content precise layout full marketing generation method according to claim 1 is characterized by: The image generation sub-model performs regional attention intervention during the generation process, including: Receive the graphic attribute labels output by the semantic prediction sub-model and extract the spatial position coordinates and weights corresponding to each semantic field; Construct a multi-channel spatial position mask matrix, assign the position mask of the main semantic field to the enhancement mark, and assign the position mask of the reverse semantic field to the suppression mark; During the diffusion generation process, the spatial position mask is used to perform weighted adjustment on the normalized weights of the attention map layer; As target regions are generated in the image, channel activation thresholds are adjusted in real time.
6. The AI intelligent graphic and text contextual content precise layout full marketing generation method according to claim 5 is characterized by: Regional attention interventions include: During the generation process, semantic fields are established, i.e., image-channel mapping relationships, and semantic weights are bound to feature channels through the channel attention mechanism. The activation intensity of the channel bound to the main semantic field is dynamically increased, and the channel bound to the reverse semantic field enters a cooling state; The SE mechanism is used to construct an inter-channel interaction regulation loop to periodically evaluate and strengthen the output of the main semantic visual stream channel. Specifically, the SE mechanism uses global average pooling to extract the global response of each channel in the feature map output of each layer, and then calculates the channel importance score through a multi-layer perceptron. The activation ratio of each channel is dynamically adjusted according to the semantic field weight. If the activation of the reverse region detected during generation exceeds the preset threshold, the pruning suppression strategy is immediately triggered to suppress the flow of related feature map signals.
7. The AI intelligent graphic and text contextual content precise layout full marketing generation method according to claim 1 is characterized by: The structure-level semantic consistency judgment mechanism includes: Analyze the preset position weight and sorting order of each semantic item in the semantic field group in the structure template; Perform region segmentation on the generated image and extract the spatial layout of graphic elements; Establish a spatial mapping relationship between semantic items and image elements, and evaluate whether the position expression of each semantic item in the image conforms to its preset weight range; Calculate the image structure alignment score. If the score is lower than the preset threshold, it is marked as consistency failure and triggers the image regeneration operation.
8. The AI intelligent graphic and text contextual content precise layout full marketing generation method according to claim 7 is characterized by: The judgment of the consistency between the image content structure and the semantic field group includes: Use a multimodal image-text adversarial consistency network to perform embedding comparison between images and semantic field groups; The discriminator model is trained to identify whether the saliency ranking of the main semantic terms, reverse semantic terms, and tonal keywords in the image is consistent with the input structure; If the main semantic item does not occupy a salient area of the image, or the reverse semantic item is frequently displayed, the discrimination result is returned as inconsistent.
9. The AI intelligent graphic and text contextual content precise layout full marketing generation method according to claim 8 is characterized by: The calculation steps of the image structure alignment score include: Perform semantic region segmentation on the generated image to extract graphic elements and their position distribution in the image coordinate system; Map the image coordinate information to the expected structural position information of each field item in the semantic field group one by one to form a semantic-image space correspondence matrix; Performing distance error calculation between the actual region position of each semantic item in the corresponding matrix and its expected weight position, wherein the error is a normalized spatial offset value; The normalized spatial offset values of all semantic items are weighted averaged to obtain the structural alignment score. If the score is lower than the set threshold, it is considered structural inconsistency and the regeneration operation is triggered.
Citation Information
Patent Citations
Image semantic segmentation method and system based on differentiated contexts
CN117253034A
Text association type short video multi-mode emotion recognition method and system
CN117636196A