Method and device for generating picture-text content, storage medium and electronic device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本申请实施例提供了一种图文内容的生成方法和装置、存储介质及电子设备,以至少解决相关技术中图文内容的生成效率比较低的技术问题
[0023]在本申请实施例中,采用以下步骤:获取目标对象的目标数据信息和待生成的目标图文内容的画布尺寸信息,其中,目标数据信息至少包括目标对象的图像信息和目标对象的描述信息;通过多模态语言模型对目标数据信息和画布尺寸信息进行处理,得到用于生成目标图文内容的目标提示词;通过图像生成模型对目标提示词进行处理,输出目标图文内容,解决了相关技术中图文内容的生成效率比较低的技术问题。
Smart Images

Figure CN122550751A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a method and apparatus for generating graphic content, a storage medium, and an electronic device. Background Technology
[0002] In interactive platforms, graphic posters serve as the primary visual expression of product information, widely used for product display, content distribution, and user interaction. Related technologies primarily rely on manual editing processes, demanding designers manually adjust layouts, color schemes, and text typesetting. This results in low production efficiency and difficulty in supporting high-frequency, multi-format content output. Alternatively, automated generation methods based on preset templates assemble content through fixed structures and element replacements. However, limited by the number of templates and rigid rules, this leads to repetitive styles and a lack of semantic understanding in the output content.
[0003] There is currently no effective solution to the technical problem of low efficiency in generating text and image content in the aforementioned related technologies. Summary of the Invention
[0004] This application provides a method and apparatus for generating graphic and textual content, a storage medium, and an electronic device to at least solve the technical problem of low efficiency in generating graphic and textual content in related technologies.
[0005] According to one aspect of the embodiments of this application, a method for generating graphic content is provided, comprising: acquiring target data information of a target object and canvas size information of the target graphic content to be generated, wherein the target data information includes at least image information of the target object and description information of the target object; processing the target data information and the canvas size information through a multimodal language model to obtain target prompt words for generating the target graphic content, wherein the multimodal language model is trained based on reward values calculated by multiple reward models; and processing the target prompt words through an image generation model to output the target graphic content.
[0006] Furthermore, the multimodal language model is trained using the following steps: processing sample data information and sample canvas size information through an initial multimodal language model to obtain multiple sets of predicted prompt words, wherein each set of predicted prompt words includes multiple predicted prompt words; calculating rewards for the multiple sets of predicted prompt words to obtain target reward values corresponding to each of the multiple predicted prompt words; updating the initial multimodal language model based on the target reward values to obtain the multimodal language model.
[0007] Further, reward calculation is performed on the multiple sets of predicted prompt words to obtain target reward values corresponding to each of the multiple predicted prompt words, including: processing the multiple predicted prompt words through the image generation model to output the initial image and text content corresponding to each of the multiple predicted prompt words; calculating the initial image and text content corresponding to each of the multiple predicted prompt words through multiple reward models to obtain initial reward values; and obtaining the target reward value based on the initial reward value.
[0008] Furthermore, the initial reward value is obtained by calculating the initial text and image content corresponding to the multiple predicted prompt words using multiple reward models, including: evaluating the click-through rate of the initial text and image content corresponding to the multiple predicted prompt words using a first reward model to obtain a first score; evaluating the text and image content quality of the initial text and image content corresponding to the multiple predicted prompt words using a second reward model to obtain a second score; and obtaining the initial reward value based on the first score and the second score.
[0009] Further, obtaining the initial reward value based on the first score and the second score includes: performing a weighted calculation based on the first score and the second score to obtain a target score; for any set of prediction prompts, calculating the average score corresponding to the set of prediction prompts based on the target score; and calculating the initial reward value based on the average score and the target score corresponding to the prediction prompts in the set of prediction prompts.
[0010] Furthermore, the click-through rate of the initial text and image content corresponding to the multiple predicted prompt words is evaluated using the first reward model to obtain a first score, including: performing block encoding on the initial text and image content to obtain a visual feature vector; performing lexical encoding on the text corresponding to the initial text and image content to obtain a sequence semantic feature vector; and obtaining the first score based on the visual feature vector and the sequence semantic feature vector.
[0011] Furthermore, the initial text and image content corresponding to the multiple predicted prompts is evaluated for quality using a second reward model to obtain a second score. This includes: obtaining a layout score based on the layout information of visual elements in the initial text and image content; obtaining a text layout score based on the text information in the initial text and image content; obtaining an appearance quality score based on the color information in the initial text and image content; obtaining a style consistency score based on the text style semantics and image style semantics in the initial text and image content; and obtaining the second score based on the layout score, the text layout score, the text layout score, and the style consistency score.
[0012] According to another aspect of the embodiments of this application, a method for generating graphic content is also provided, comprising: acquiring target data information of a target object uploaded by a client and canvas size information of the target graphic content to be generated, wherein the target data information includes at least image information and description information of the target object; processing the target data information and the canvas size information in a cloud server using a multimodal language model to obtain target prompt words for generating the target graphic content; processing the target prompt words using an image generation model to output the target graphic content; and returning the target graphic content to the client.
[0013] According to another aspect of the embodiments of this application, an apparatus for generating graphic content is also provided, comprising: an acquisition unit, configured to acquire target data information of a target object and canvas size information of the target graphic content to be generated, wherein the target data information includes at least image information of the target object and description information of the target object; a first processing unit, configured to process the target data information and the canvas size information through a multimodal language model to obtain target prompt words for generating the target graphic content; and a second processing unit, configured to process the target prompt words through an image generation model to output the target graphic content.
[0014] Furthermore, the multimodal language model is trained using the following apparatus: a third processing unit, used to process sample data information and sample canvas size information through the initial multimodal language model to obtain multiple sets of predicted prompt words, wherein each set of predicted prompt words includes multiple predicted prompt words; a calculation unit, used to perform reward calculation on the multiple sets of predicted prompt words to obtain target reward values corresponding to the multiple predicted prompt words respectively; and an update unit, used to update the initial multimodal language model based on the target reward values to obtain the multimodal language model.
[0015] Further, the calculation unit includes: a processing subunit, used to process the plurality of predicted prompt words through the image generation model and output the initial image and text content corresponding to the plurality of predicted prompt words respectively; a calculation subunit, used to calculate the initial image and text content corresponding to the plurality of predicted prompt words respectively through multiple reward models to obtain an initial reward value; and a determination subunit, used to obtain the target reward value based on the initial reward value.
[0016] Further, the calculation subunit includes: a first evaluation module, used to evaluate the click-through rate of the initial text and image content corresponding to the plurality of predicted prompt words using a first reward model, and obtain a first score; a second evaluation module, used to evaluate the text and image content quality of the initial text and image content corresponding to the plurality of predicted prompt words using a second reward model, and obtain a second score; and a determination module, used to obtain the initial reward value based on the first score and the second score.
[0017] Further, the determining module includes: a first calculation submodule, used to perform a weighted calculation based on the first score value and the second score value to obtain a target score value; a second calculation submodule, used to calculate the average score value corresponding to any set of predicted prompt words based on the target score value; and a third calculation submodule, used to calculate the initial reward value based on the average score value and the target score value corresponding to the predicted prompt words in the set of predicted prompt words.
[0018] Further, the first evaluation module includes: a first processing submodule, used to perform block encoding processing on the initial image and text content to obtain a visual feature vector; an encoding submodule, used to perform lexical encoding on the text corresponding to the initial image and text content to obtain a sequence semantic feature vector; and a first determining submodule, used to obtain the first score value based on the visual feature vector and the sequence semantic feature vector.
[0019] Further, the second evaluation module includes: a second determining submodule, used to obtain a layout score based on the layout information of visual elements in the initial graphic content; a third determining submodule, used to obtain a text layout score based on the text information in the initial graphic content; a fourth determining submodule, used to obtain an appearance quality score based on the color information in the initial graphic content; a fifth determining submodule, used to obtain a style consistency score based on the text style semantics and image style semantics in the initial graphic content; and a second processing submodule, used to obtain the second score based on the layout score, the text layout score, the text layout score, and the style consistency score.
[0020] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein, when the program is running, it controls the device where the storage medium is located to execute the above-described method for generating graphic content.
[0021] According to another aspect of the embodiments of this application, an electronic device is provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the above-described method for generating graphic content during runtime.
[0022] According to another aspect of the embodiments of this application, a computer program product is provided, including a computer program or instructions, which, when executed by a processor, implement the above-described method for generating graphic content.
[0023] In this embodiment, the following steps are adopted: obtaining target data information of the target object and canvas size information of the target graphic content to be generated, wherein the target data information includes at least image information and description information of the target object; processing the target data information and canvas size information through a multimodal language model to obtain target prompt words for generating target graphic content; processing the target prompt words through an image generation model to output the target graphic content, thereby solving the technical problem of low efficiency in generating graphic content in related technologies.
[0024] In this application, a multimodal language model is used to jointly model the received target object image information, descriptive information, and canvas size information to generate prompts that conform to visual structural specifications. The image generation model then outputs the target graphic content based on these prompts. This process integrates the previously manually designed multi-step operations (such as copywriting, composition planning, and style matching) into an end-to-end automated processing chain, reducing intermediate human intervention. Because the multimodal language model possesses the ability to jointly understand the semantics of images and text, it can adaptively generate structured prompts containing subject descriptions, composition guidance, and style instructions based on product characteristics and canvas proportions, thereby improving the consistency between the prompts and the input information. The image generation model renders based on the prompts, which helps reduce multiple redraws or manual corrections caused by ambiguous or incomplete prompts, improving the usability of a single generation. Without changing the underlying image generation capabilities, by optimizing the semantic quality and structural regularity of the input prompts, the technical effect of shortening the generation time of graphic content is achieved. Attached Figure Description
[0025] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0026] Figure 1 This is a hardware structure block diagram of a computer terminal provided according to Embodiment 1 of this application;
[0027] Figure 2 This is a flowchart of the method for generating graphic content according to Embodiment 1 of this application;
[0028] Figure 3 This is a schematic diagram of the method for generating graphic content according to Embodiment 1 of this application. Figure 1 ;
[0029] Figure 4 This is a schematic diagram of the method for generating graphic content according to Embodiment 1 of this application. Figure 2 ;
[0030] Figure 5 This is a schematic diagram of the method for generating graphic content according to Embodiment 1 of this application. Figure 3 ;
[0031] Figure 6 This is a flowchart of the method for generating graphic content according to Embodiment 2 of this application;
[0032] Figure 7 This is a schematic diagram of a dynamic special effects processing device according to Embodiment 3 of this application;
[0033] Figure 8 This is a structural block diagram of an electronic device provided according to Embodiment 5 of this application. Detailed Implementation
[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0036] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant regulations and standards of the relevant regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0037] Example 1
[0038] According to an embodiment of this application, a method for generating graphic content is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0039] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for generating graphic content is shown. Figure 1 As shown, the computer terminal (or mobile device) 10 may include a processor set 102 (the processor set 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA, and the processor set 102 may include a processor set, Figure 1 The data is illustrated using 102a, 102b, ..., 102n. A memory 104 is used for storing data, and a transmission module 106 is used for communication functions. In addition, it may include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0040] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0041] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image and text content generation method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-mentioned image and text content generation method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0042] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0043] The display may be, for example, a touchscreen LCD display that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0044] Under the aforementioned operating environment, this application provides the following: Figure 2 The method for generating the illustrated text content is shown. Figure 2 This is a flowchart of a method for generating graphic content according to Embodiment 1 of this application. The method for generating graphic content includes:
[0045] Step S201: Obtain the target data information of the target object and the canvas size information of the target graphic content to be generated. The target data information includes at least the image information of the target object and the description information of the target object.
[0046] Optionally, target data information of the target object can be obtained according to user needs. The target object can be a product of the current platform. Target data information includes, but is not limited to, image information and descriptive information.
[0047] Image information can be directly retrieved from the image library system. These are static images with a resolution of at least 480×480 pixels. Images can be the main image uploaded when the product is listed, or optional scene images, etc. Description information can be extracted from the attribute database and includes the name, core functionalities or usage scenario keywords, and target audience. After obtaining the description information, it can be cleaned to remove unsuitable words (e.g., subjective words) and repetitive expressions, retaining objective and concrete semantic content.
[0048] Canvas size information can be specified by the target use case, such as 1080×1920 pixels for a vertical display on a mobile device and 1200×628 pixels for a banner on the platform homepage. This parameter can be determined by front-end page configuration, API parameters, or preset templates, and is converted into integer width and height values after being received, serving as the basis for layout adaptation.
[0049] In an optional embodiment, during the input processing stage, image data can be normalized and scaled to a uniform resolution, text information can be segmented and encoded, and the canvas size can be converted into a floating-point aspect ratio and a joint embedding of pixel values. All three are simultaneously input into the subsequent multimodal oracle model as the contextual basis for generating prompt words, ensuring that the output content matches the target usage scenario in terms of visual structure.
[0050] Step S202: The target data information and canvas size information are processed by a multimodal language model to obtain target prompt words for generating target graphic content. The multimodal language model is trained based on reward values calculated by multiple reward models.
[0051] Optionally, semantic representations of image information, including the shape of the subject, color distribution, background environment, and key local regions, can be extracted through a visual encoder to obtain image features; the descriptive information can be transformed into a word vector sequence through word segmentation and semantic encoding to clarify attributes, usage scenarios, and value points, thus obtaining a text description; the canvas size information can be linearly mapped into a two-dimensional vector to reflect the aspect ratio and spatial constraints, helping the model understand the feasible range of content layout, thus obtaining the canvas size embedding.
[0052] The multimodal language model receives image features, text descriptions, and canvas size embeddings. After fusing the above modal information, the multimodal language model gradually generates prompts that conform to the visual narrative logic based on the structured expressive capabilities learned by the model through supervised fine-tuning of text and image content. The generation process follows a preset prompt template structure, which, for example, needs to include: a description of the main object (e.g., a transparent glass), style guidance words (e.g., minimalist style, low saturation), composition instructions (e.g., centered placement, occupying 60% of the screen), and canvas adaptation instructions (e.g., 15% white space at the top).
[0053] In an alternative embodiment, the model can combine visual focal points in the image with semantic focal points in the text during generation to align content priorities. For example, when the description mentions leak-proof design, the model tends to strengthen the visual presentation of the bottle opening or sealing structure in the prompt words.
[0054] In an optional embodiment, during the generation of prompts, the multimodal language model can dynamically weight the inputs of each modality through a self-attention mechanism. For example, when the canvas size is narrow and elongated (e.g., 1080×1920), the model tends to use composition instructions that conform to the reading habits of vertical screens, such as vertical arrangement, white space at the top, and bottom labels; if the background of the input image is cluttered, optimization suggestions such as blurred background and high-contrast foreground are automatically added to the prompts to improve the recognition of the subject.
[0055] In an optional embodiment, the multimodal language model may include: an image encoding module, a text encoding module, a size embedding module, and a multimodal fusion and prompt word generation module.
[0056] The input is a normalized image (e.g., a 480×480 RGB image) fed into the image encoding module. First, PatchEmbedding divides the image into 16×16 pixel blocks, then linear projection transforms it into a visual label sequence. Subsequently, a multi-layer encoder extracts semantic features, outputting a visual feature vector containing information such as global semantics (e.g., glass cup, metal bottle cap), local key regions (e.g., bottle opening structure, leak-proof sealing ring), color distribution (e.g., transparent subject, silver metal), and background complexity. The image encoding module does not perform image enhancement or cropping, maintaining the integrity of the original input and ensuring that the features are consistent with the real product (i.e., the target object).
[0057] The input is a descriptive text after cleaning (e.g., 500ml leak-proof glass water cup, baby-grade silicone seal) to the text encoding module. First, a word segmenter (BPE) divides the text into sub-word units, which are then mapped to word embedding sequences. After processing through multiple layers of self-attention mechanisms, the output is a semantically enhanced text vector, clearly defining attributes (material, capacity), functional characteristics (leak-proof, safe and non-toxic), and usage scenarios (infant, outdoor). This module filters out non-semantic words, retaining verifiable and visually appealing objective descriptions.
[0058] The size embedding module receives canvas size parameters (such as width W and height H) and converts them into a standardized two-dimensional vector of the form [W / H, min(W, H) / 1000], representing the aspect ratio and the logarithmic scaling of the size, respectively. This vector does not carry specific pixel values but expresses the relative relationships of spatial constraints (such as a vertical screen being narrow and long, a horizontal screen being wide and flat). As a learnable embedding vector, it is aligned dimensionally with visual and textual features and then concatenated for use by the subsequent fusion module. This design allows the model to generalize to any size, rather than relying on a fixed template.
[0059] The multimodal fusion and prompt word generation module can be composed of multiple layers of cross-modal cross-attention and autoregressive decoders. Image encoding vectors, text encoding vectors, and size embedding vectors are concatenated after being linearly projected to unify their dimensions and input into the cross-attention layer to achieve dynamic interaction between modalities. Attention weights are automatically used to determine: when the text emphasizes preventing leakage, the activation value of the bottle opening region in the visual module is enhanced; when the screen size is portrait, the model tends to activate generation paths related to the vertical layout.
[0060] The decoder generates prompts word by word in an autoregressive manner based on fused representations, strictly following the preset structure: [Main Object] + [Style Guidance] + [Composition Instructions] + [Canvas Adaptation].
[0061] For example, a transparent glass water cup with a low-saturation minimalist style is placed in the center and occupies 60% of the image height, with 15% of the top blank space left for the brand logo.
[0062] The generation process can also be controlled by temperature parameters (τ=0.7) and beam search (width=3) to ensure a balance between diversity and stability; before output, it is filtered by a syntax checker to ensure that there are no contradictory instructions (such as "center" and "left-align"), no illegal words, and a length of ≤200 characters.
[0063] The final output prompts are semantically complete, structurally sound, and visually executable natural language instructions, which are directly used as input to the downstream image generation model.
[0064] In an optional embodiment, the multimodal language model can dynamically extract style prototypes from a database of historically high-click-rate posters on the platform during the style guide word generation stage, and encode them into pluggable style embedding vectors. For example:
[0065] When the input product is detected to be a skin care product and the image contains water droplets and a soft background, the model automatically calls the light luxury style prototype and upgrades the style words from low saturation to matte texture, soft focus light and shadow, and pearly luster reflection.
[0066] When the input product is a children's toy and the description includes interactivity and safety, the model introduces a childlike illustration style prototype to generate descriptions with more emotional warmth, such as hand-painted watercolor borders, cartoon clouds, and bright macaron colors.
[0067] In an optional embodiment, user profile preference tags (such as a preference for warmth or a preference for Chinese style) can also be injected during the generation stage to dynamically adjust the generation style weight.
[0068] In an optional embodiment, the target prompt can be a structured, semantically complete, and visually executable natural language instruction, such as a transparent leak-proof glass baby bottle, with the main body placed in the center and occupying 65% of the screen height. A close-up of the bottle mouth shows the details of the silicone sealing ring. The background is a soft, gradient morning light, using low-saturation cream white and light wood tones. 18% of the top is left blank for the brand logo. The overall style is a warm and cozy mother and baby style, with no text interference, no people appearing, and the lighting simulates natural window light. The material is high-transparency glass + matte silicone.
[0069] In an optional embodiment, the multimodal language model is trained based on reward values calculated by multiple reward models. Different reward models can score the output of the multimodal language model from different evaluation dimensions, including but not limited to multimodal consistency (text-image matching degree) and logical coherence. The parameters of the multimodal language model are updated using rewards through reinforcement learning algorithms.
[0070] Step S203: Process the target prompt words using an image generation model to output the target image and text content.
[0071] Optionally, the target prompt words are input into the image generation model to perform semantic parsing on the target prompt words. For example, structural decomposition: automatically identify and label the four elements: main object, style guidance, composition instructions, and canvas adaptation.
[0072] Keyword standardization: Converting colloquial expressions into model terminology; Compliance verification: Calling the brand / platform's prohibited word library and automatically replacing inappropriate expressions; Implicit completion: Automatically supplementing default parameters based on category knowledge.
[0073] Then, based on the composition and canvas adaptation instructions in the prompt (such as center placement, occupying 60% of the height, and leaving 15% of the top blank space), a grayscale mask image with the same size as the target canvas is generated:
[0074] Main area (water glass): mask value = 1.0 (strong constraint); Logo area (top 15%): mask value = 0.3 (weak guidance); Background area: mask value = 0.0 (free generation); Other areas (such as edges): mask value = 0.1 (micro-constraint to prevent offset).
[0075] The mask image is scaled to the model input resolution (e.g., 1024×1024) using bilinear interpolation, and serves as a spatial condition map to guide the diffusion process in preserving or generating content in a specified area.
[0076] The image generation model is based on a multi-condition controlled latent diffusion architecture, which performs the following process in the latent space: the initial noise map is generated by sampling from a Gaussian distribution; in each denoising step, the model simultaneously receives three types of inputs: text semantic embedding (guiding content semantics); layout mask map (constraining spatial distribution); style / texture embedding (controlling visual representation).
[0077] Through a cross-attention mechanism, the model dynamically weighs the weights of various conditions: when glass and refraction appear simultaneously, the response of the lighting module is enhanced; when the top blank space coexists with the brand logo, the content generation in that area is suppressed; then, multi-step denoising is performed to output a high-resolution (e.g., 1080×1920) latent representation, which is then restored to an RGB image by the decoder.
[0078] In an optional embodiment, the generated original image can also undergo image enhancement processing. For example, an adaptive unsharpening mask can be used to enhance the sharpness of the subject's edges, improving recognizability; local super-resolution can be applied to the logo area to ensure that small text is clearly readable. Furthermore, a histogram matching algorithm can be used to align the color distribution of corresponding areas in the image with a reference color chart, reducing color cast.
[0079] In summary, by jointly modeling the received target object image information, descriptive information, and canvas size information using a multimodal language model, prompts conforming to visual structural specifications are generated. The image generation model then outputs the target text and image content based on these prompts. This process integrates the previously manually designed multi-step operations (such as copywriting, composition planning, and style matching) into an end-to-end automated processing chain, reducing intermediate human intervention. Because the multimodal language model possesses the ability to jointly understand the semantics of images and text, it can adaptively generate structured prompts containing subject descriptions, compositional guidance, and style instructions based on product characteristics and canvas proportions, thereby improving the consistency between the prompts and the input information. The image generation model renders based on the prompts, helping to reduce multiple redraws or manual corrections caused by ambiguous or incomplete prompts, improving the usability of a single generation. Without changing the underlying image generation capabilities, by optimizing the semantic quality and structural regularity of the input prompts, the technical effect of shortening the generation time of text and image content is achieved.
[0080] To improve the effectiveness of the multimodal language model, in the image and text content generation method provided in Embodiment 1 of this application, the multimodal language model is trained using the following steps: processing sample data information and sample canvas size information through the initial multimodal language model to obtain multiple sets of predicted prompt words, wherein each set of predicted prompt words includes multiple predicted prompt words; calculating rewards for the multiple sets of predicted prompt words to obtain target reward values corresponding to each of the multiple predicted prompt words; updating the initial multimodal language model based on the target reward values to obtain the multimodal language model.
[0081] Optionally, real-world scenario data can be collected as training samples. A sample may include: product images: main product image or scene image (such as skincare product bottle or sneaker appearance); product text information: structured or unstructured copy such as title, core selling points, ingredient description, and usage scenarios; canvas size information: the width and height of the target output poster (such as 1080×1920 or 800×600) to adapt to different distribution channels; and manually annotated high-quality prompts (optional): used only for initial alignment and not a training dependency.
[0082] The training samples are encoded into multimodal input vectors: images are processed by a visual encoder to extract structural and semantic features, text information is processed by a language encoder to convert it into semantic vectors, and canvas size is converted into numerical embeddings. The three are then fused and input into the initial multimodal language model as the context for generating prompts.
[0083] The initial multimodal language model does not output just one prompt word for each input sample. Instead, it uses a sampling strategy to generate multiple sets of diverse predicted prompt words at once. For example, for a glass water cup product, the model may simultaneously output five prompt words with different styles but all grammatically correct. Some emphasize "minimalist style," some highlight "morning light atmosphere," and some focus on "high-end quality," forming a candidate set containing multiple creative possibilities. This multi-path sampling mechanism allows the model to explore a richer expressive space in a single inference.
[0084] Then, rewards are calculated for multiple sets of predicted prompts. For example, the predicted prompts are input into a trained semantic effect association evaluator. The semantic effect association evaluator learns which language structures, word combinations, and expressions tend to produce results that are both highly convertible and aesthetically pleasing, based on the performance data of historical prompts and their corresponding generated images in real-world scenes. It identifies multi-dimensional features in the prompts, such as semantic density, style consistency, spatial instruction clarity, and sentiment intensity, and matches and compares them with patterns of historically high-performing prompts to output a comprehensive quality score.
[0085] In an optional embodiment, the target reward value is determined based on the comprehensive quality score output by the semantic effect association evaluator. For example, for all prompt words in each group, the group average of their comprehensive quality scores is calculated, and the relative advantage of each prompt word to the average value is calculated. This relative advantage is used as the target reward value. If the relative advantage is greater than 0, it means that the prompt word performs better than the average level in the current group and belongs to the relatively good solution. The model should increase the probability of it being sampled in the future. If the relative advantage is less than 0, it means that the prompt word does not reach the average level in the group and belongs to the "relatively bad solution". The model should appropriately reduce its generation tendency.
[0086] Building upon this, a dynamic balancing mechanism for the semantic space can be further introduced to prevent the model from getting stuck in local optima or style solidification during the optimization process. When a certain type of cue word consistently gains a high relative advantage across multiple groups, it is identified whether this type of expression is rapidly converging into a high-scoring template. Once the semantic diversity index is detected to be below a preset threshold, i.e., the cue words tend to be homogeneous in terms of word combination, structural layout, or emotional expression, an implicit exploration incentive will be applied to this type of high-frequency pattern, superimposing a small negative shift on its relative advantage. This guides the model to actively explore low-frequency but potentially high-value regions in the semantic space without directly interfering with the generation logic.
[0087] During the model parameter update phase, a policy gradient method based on relative advantage can be used, which assigns a gradient to increase the probability only to prompt words with a positive relative advantage, while suppressing prompt words with a negative advantage.
[0088] In an optional embodiment, all prompts within the same group can be relatively ranked to calculate the ranking distribution and relative advantage of the predicted prompts within that group. For example, in a group of five prompts, if a predicted prompt has a significantly higher score than the other four in the group, it is given a higher relative reward weight; conversely, if all prompts have similar scores, the update magnitude is reduced to avoid excessively perturbing the model parameters.
[0089] Through the aforementioned multi-round closed-loop optimization mechanism, the multimodal language model no longer relies on manual annotation or fixed templates. Instead, it learns autonomously under unsupervised conditions through a strategy gradient update driven by relative advantages within the group, a dynamic balance of semantic diversity, and a ranking-aware gradient weighting mechanism. This not only effectively suppresses the homogenization and style solidification of the generated results, avoiding getting stuck in local optima, but also continuously stimulates the model to explore diverse expressions in the semantic space. This allows the prompt words to have a high degree of creative diversity and scene adaptability while maintaining grammatical compliance and semantic integrity.
[0090] To improve the accuracy of reward value calculation, the image and text content generation method provided in Embodiment 1 of this application calculates rewards for multiple sets of predicted prompt words to obtain target reward values corresponding to each of the multiple predicted prompt words. This includes: processing the multiple predicted prompt words through an image generation model to output initial image and text content corresponding to each of the multiple predicted prompt words; calculating the initial image and text content corresponding to each of the multiple predicted prompt words through multiple reward models to obtain initial reward values; and obtaining target reward values based on the initial reward values.
[0091] Optionally, the predicted prompts are input into an image generation model, which accurately maps the complex semantics of the predicted prompts, such as style instructions, composition descriptions, material semantics, and color tendencies, into visual content. Each prompt triggers an image generation process independently, outputting the initial text and image content (i.e., the poster image) corresponding to its semantics. This ensures that the potential expressiveness of each prompt is realistically and completely visualized, rather than relying solely on semantic abstraction scoring at the linguistic level.
[0092] Multiple reward models are used to evaluate the visual quality of the generated initial text and image content from multiple dimensions to obtain its initial reward value. For example, a quality reward model evaluates the composition, color scheme, and lighting effects of the initial text and image content to obtain an aesthetic quality score. A visual realism reward model evaluates whether the initial text and image content conforms to common sense logic of the physical world, such as whether the shape, proportion, and texture of objects are reasonable, and whether there is distortion or artifacts, to obtain a realism score. Then, the corresponding initial reward value is determined based on the aesthetic quality score and the realism score. As another example, a pre-trained click-through rate prediction network is used to analyze image attributes strongly correlated with click behavior, such as visual saliency, color contrast, focus distribution, and text readability, to output a user preference score. A visual scoring model trained by designers can also be used to evaluate professional design indicators such as compositional balance, element hierarchy, white space rationality, style consistency, and texture realism, to output an aesthetic quality score. The aesthetic quality score and the user preference score are used as the initial reward value. For example, it can also detect whether there are violations of regulations in the image, such as missing brand logos, text obscuration, unbalanced proportions, or prohibited elements, output compliance penalty scores, and use the compliance penalty scores as initial reward values.
[0093] Based on the initial reward value, context normalization and relative advantage mapping mechanisms can be introduced to transform it into the final target reward value:
[0094] The initial reward values corresponding to all prompts within the same group are normalized within the group (e.g., Z-score standardization) to eliminate scoring bias caused by different product categories or canvas sizes; the difference between the initial reward value of the predicted prompt and the average value of the group is calculated as its relative advantage; further, the relative advantage is dynamically adjusted by combining semantic diversity constraints and historical pattern suppression factors to encourage exploratory expression and suppress pattern solidification, and finally output the target reward value of each prompt.
[0095] By realistically rendering predicted prompts as visual images and calculating initial reward values based on multi-dimensional visual feedback, the target reward value is ensured to accurately reflect the visual conversion potential and design compliance of the prompts, thereby improving the effectiveness of subsequent model predictions.
[0096] To improve the accuracy of calculating the initial reward value, the image and text content generation method provided in Embodiment 1 of this application calculates the initial image and text content corresponding to multiple predicted prompt words using multiple reward models to obtain the initial reward value. This includes: evaluating the click-through rate of the initial image and text content corresponding to multiple predicted prompt words using a first reward model to obtain a first score; evaluating the image and text content quality of the initial image and text content corresponding to multiple predicted prompt words using a second reward model to obtain a second score; and obtaining the initial reward value based on the first score and the second score.
[0097] Optionally, a first reward model and a second reward model are set up. The first reward model (URM, UserReward Model) is used to simulate and evaluate user click behavior on the initial text and image content. For example, its evaluation dimensions include: visual saliency (such as whether the subject is prominent and whether the color contrast is strong); focus distribution (visual attention area predicted by human eye heatmap); text readability and layout clarity (OCR recognition + font size and background contrast analysis); and emotional induction intensity (such as emotion tag classification), thereby obtaining the first score value.
[0098] The Designer Reward Model (DRM) is used to conduct a professional design aesthetic evaluation of the same initial graphic content. This evaluation considers factors such as: compositional balance (distribution of the center of gravity, rationality of symmetrical / asymmetrical structures); element hierarchy (clarity of primary and secondary information, smooth visual flow); white space and breathing room (appropriate use of negative space to avoid information overload); and stylistic consistency and texture (realistic material rendering, harmonious color scheme, and unified font and image styles). The model outputs a second score (range 0–1) representing the aesthetic quality of the initial graphic content under professional design standards.
[0099] Finally, the first and second scores are dynamically weighted and fused to obtain the initial reward value. A confidence calibration module can be further introduced: for each initial reward value, its prediction confidence derived from URM and DRM is calculated (e.g., by judging the variance, entropy, or consistency of the ensemble model output). If the confidence of either model's score for a certain image is lower than a threshold (e.g., entropy > 0.3), a manual review agent is activated. The review system samples human scores and feeds them back to the reward model for online incremental learning, avoiding systematic misjudgments due to data bias or noisy samples. Furthermore, when there is a significant discrepancy between the URM and DRM scores for the same image / text content (e.g., URM score ≥ 0.9 while DRM score ≤ 0.4, or vice versa), the context of the current delivery scenario is considered to determine which reward model's score is more likely to be affected by scenario bias or evaluation distortion. This identifies the reward model that needs priority correction, achieving accurate alignment and consistency of the reward signal.
[0100] In an optional embodiment, the evaluation diagram of the dual-reward model (i.e., the first reward model and the second reward model) is shown below. Figure 3 As shown, the poster rendering output evaluates the poster's user appeal using a user reward model (i.e., the first reward model) to obtain an appeal score. The designer reward model (i.e., the second reward model) scores the poster based on visual layout, text layout, appearance quality, and style consistency to obtain corresponding scores. These scores are then weighted and fused, and the weighted scores are used to optimize the cue word model (i.e., the multimodal language model mentioned above).
[0101] This dual-reward model collaborative evaluation mechanism achieves explicit decoupling and joint modeling of user behavior feedback and professional design standards, avoiding the problems of high click-through rates or aesthetically pleasing but non-converting results caused by a single indicator-driven approach.
[0102] To further improve the rationality of calculating reward values, in the method for generating text and image content provided in Embodiment 1 of this application, an initial reward value is obtained based on a first score value and a second score value, including: performing a weighted calculation based on the first score value and the second score value to obtain a target score value; for any set of predicted prompt words, calculating the average score value corresponding to the set of predicted prompt words based on the target score value; and calculating the initial reward value based on the average score value and the target score value corresponding to the predicted prompt words in the set of predicted prompt words.
[0103] Optionally, for each predicted prompt word, the corresponding URM score (first score) and DRM score (second score) of the generated image and text content are first obtained, and weights α and β are dynamically assigned according to the current delivery scenario to calculate the target score of the image and text content. α and β are dynamically determined by the scenario strategy library (e.g., α=0.4, β=0.6 during the brand period) to ensure that the reward orientation is consistent with the task objective.
[0104] All candidate prompts sampled in a single generation task (e.g., a group of 10) are considered as an evaluation batch. The mean of the target scores for all candidate samples within this batch is calculated and denoted as the batch average score. This average score represents the baseline performance level of the batch under the current model and scene settings, and is used to eliminate systematic biases caused by model sampling fluctuations, image generation noise, or scoring bias within the batch. For example, if a batch has generally high target scores due to the model's tendency to generate highly saturated styles, the average score will increase accordingly, providing a dynamic reference for subsequent relative scoring.
[0105] After obtaining the average score within the group, calculate the relative advantage of each prompt word relative to the average level within the group. You can also introduce the within-group dispersion correction factor to form the final initial reward value: Initial reward value = (target score – average score within the group) × (1 + γ × standardized dispersion).
[0106] Wherein, (target score – group average score): represents the relative improvement of the prompt word relative to the group average level. A positive number indicates that it is better than the group, and a negative number indicates that it is worse than the group, ensuring that the reward signal focuses on who is better, rather than how much better; Standardized dispersion: calculates the standard deviation of the target score for the group and normalizes it to the interval [0, 1], reflecting the level of diversity within the group. If the difference between samples within the group is small (low standard deviation), it indicates that the model is convergent and stable, and a smaller positive incentive (smaller γ) is given; if the difference between samples within the group is large (high standard deviation), it indicates active exploration and encourages diversity, and the reward difference is amplified (larger γ); γ is the dispersion adjustment coefficient, initially set to 0.3, which can be adaptively adjusted according to the training stage.
[0107] In an optional embodiment, the initial reward values of all historical batches can be collected periodically (e.g., every 100 batches) to construct their empirical distribution and calculate the global mean μ_global and standard deviation σ_global. If the overall deviation of the initial reward values of a batch exceeds |mean-μ_global|>2σ_global, a reward drift warning is triggered, and a resampling and model calibration process is initiated, thereby ensuring the long-term consistency of the reward system throughout the entire training cycle.
[0108] By weighting the URM and DRM scores to obtain the target score, and then calculating the relative advantage based on the batch mean, the volatility of single scores affected by noise or local preferences is effectively reduced. This makes the reward signal more focused on the relative merits between cue words, improving the stability and consistency of the reward signal and providing a more reliable and robust optimization basis for reinforcement learning.
[0109] To improve the accuracy of calculating the first score, in the image and text content generation method provided in Embodiment 1 of this application, the click-through rate of the initial image and text content corresponding to multiple predicted prompt words is evaluated by a first reward model to obtain the first score value. This includes: performing block encoding processing on the initial image and text content to obtain a visual feature vector; performing lexical encoding on the text corresponding to the initial image and text content to obtain a sequence semantic feature vector; and obtaining the first score value based on the visual feature vector and the sequence semantic feature vector.
[0110] Optionally, the generated initial image and text content is divided into several semantic region blocks (e.g., main product area, copywriting area, brand logo area, background decoration area, etc.), and local visual features of each region are extracted through convolution or attention mechanisms. All region features are aggregated (e.g., global average pooling or learnable pooling) to output a high-dimensional visual feature vector, which encodes visual cues in the image that are highly related to user attention, such as composition, color distribution, contrast, element layout, and visual center of gravity.
[0111] The original text associated with the image and text content is input into a text encoder, where it undergoes tokenization and is encoded into a serialized semantic vector. This vector not only contains keyword semantics (such as "limited-time discount" and "buy one get one free"), but also captures semantic structure (such as emotional tendency, intensity of urgency, and expression of trust) and pragmatic features (such as whether it contains numbers and whether it uses the second person), thereby quantifying the copy's potential to stimulate user clicks at a psychological level.
[0112] Visual feature vectors and sequential semantic feature vectors are input into a lightweight multimodal fusion module (such as a cross-attention mechanism or a combination of splicing and a fully connected network). This module learns the semantic relationship between the two, for example: does adding descriptive text to a red background enhance click expectation, and does adding abstract text to excessive white space reduce conversion motivation? The fused joint representation is fed into a regression head, which outputs a scalar value in the interval [0, 1], namely the first score (URM score), representing the predicted probability of the image and text content being clicked in a real user environment.
[0113] In an optional embodiment, during the visual feature extraction stage, not only static region features are extracted, but the user's visual scanning path on the poster can also be simulated using a spatiotemporal attention map. For example, a lightweight eye-tracking prediction subnetwork is trained based on real user eye-tracking data (from historical A / B tests or simulated click heatmaps) to output the expected gaze duration and saccade priority of image regions. Then, the expected gaze duration and saccade priority are incorporated into the region feature aggregation process, making the visual feature vectors closer to the cognitive priorities of real users.
[0114] In an optional embodiment, implicit embedding of user profiles can also be introduced: a lightweight user context vector is generated based on the current delivery channel, time period, and user historical behavior (such as click preferences and purchase categories), and then gating and fusing it with visual text joint representation.
[0115] By extracting visual features from the initial text and image content in blocks, performing semantic encoding on the text, and using a multimodal fusion mechanism to jointly model the synergistic relationship between visual layout and text semantics, this application shifts the calculation of the first score (URM score) from relying on manual rules or single-modal statistics to end-to-end prediction based on real user behavior data, thereby improving the accuracy of click-through rate assessment.
[0116] To improve the accuracy of calculating the second score, in the image and text content generation method provided in Embodiment 1 of this application, a second reward model is used to evaluate the image and text content quality of the initial image and text content corresponding to multiple predicted prompt words to obtain the second score. This includes: obtaining a layout score based on the layout information of visual elements in the initial image and text content; obtaining a text layout score based on the copywriting information in the initial image and text content; obtaining an appearance quality score based on the color information in the initial image and text content; obtaining a style consistency score based on the copywriting style semantics and image style semantics in the initial image and text content; and obtaining the second score based on the layout score, text layout score, text layout score, and style consistency score.
[0117] Optionally, object detection and instance segmentation technologies are used to automatically identify core visual elements (such as the product body, logo, promotional labels, buttons, borders, and white space) in the initial text and image content, and extract their position, size, proportion, and relative relationship on the canvas. Based on this, the degree of conformity with classic design principles (such as standard segmentation, symmetry, F-shaped visual flow, and visual center balance) is calculated to obtain a layout score. For example, if the product body deviates from the visual center by more than 30% and there is no guiding element to compensate, or if the overlap between the text area and the product area is too high, the layout score is lowered; if the element distribution presents a clear visual hierarchy (primary-secondary-secondary) and the white space ratio meets design specifications (such as ≥15%), a high score is obtained.
[0118] Then, text recognition and analysis are performed on the text area of the initial graphic content to extract layout parameters such as font type, font size level, line spacing, character spacing, alignment, and color contrast (against the background). The evaluation assesses whether it meets principles such as clarity, readability, hierarchy, and avoidance of crowding. For example: Is the main title font size 1.5–2 times that of the subtitle? Are more than 3 fonts used? Is there insufficient contrast between the text and the background color (e.g., light gray text on a white background)? Are there any improper line breaks or abrupt line breaks? These parameters are normalized and weighted to output a text layout score, quantifying the usability and professionalism of the copy in visual communication.
[0119] Secondly, the main color tone, secondary color distribution, color saturation, brightness contrast, and hue distribution entropy of the initial text and image content are extracted and compared with e-commerce industry design databases (such as the color distribution of conversion templates). An appearance quality scoring module is constructed by calculating indicators such as color harmony (e.g., using complementary / analogous color matching rules), visual complexity (whether the number of colors exceeds 5, resulting in clutter), and texture consistency (e.g., whether metallic materials are paired with high-gloss reflections, and whether gradients are natural). For example, while a combination of highly saturated red and fluorescent green is eye-catching, a lack of transition and neutral color buffering results in a perceived "cheap" feel and a lower score; conversely, a low-saturation Morandi color scheme with accent colors receives a high score. This score assesses whether the visual presentation possesses a "high-end" and "brand quality" feel.
[0120] Furthermore, style semantic encoding is performed on both the text and images: Text style tags are determined (e.g., lively promotion, high-end minimalism, professional authority, heartwarming narrative); image visual styles are identified (e.g., flat illustration, realistic photography, traditional Chinese ink painting, neon cyberpunk). The cosine similarity between the two in the style semantic space is calculated as a style consistency score. For example, high-end minimalist text paired with realistic product photography scores high, while heartwarming narrative text paired with a cyberpunk background scores low. This ensures that the text and images are consistent in emotion, tone, and aesthetics, avoiding cognitive dissonance caused by conflicting text and images.
[0121] Finally, a second score is obtained by weighting and combining the four scores of layout, typography, appearance quality and style consistency. For example, fixed weights are assigned based on the consensus of design experts or the contribution of historical excellent samples (e.g., layout 0.35, typography 0.25, appearance quality 0.20, style consistency 0.20), and the final second score (DRM score) is calculated by linear weighting.
[0122] By breaking down the quality assessment of text and image content into four quantifiable dimensions—layout rationality, text layout standardization, premium appearance, and consistency of text and image style—this approach not only effectively avoids the subjective bias and efficiency bottlenecks of traditional manual scoring, but also enables the model to accurately identify aesthetically pleasing but ineffective or effective but rough generated results. This guides the optimization of prompts towards high-quality posters that combine professional design standards with commercial expressiveness, thereby enhancing the visual credibility and brand suitability of automatically generated content.
[0123] In an alternative embodiment, such as Figure 4 The diagram shown illustrates that the automatic creation of graphic posters can be divided into the following steps: Step 1: Product Information Input: Obtain product images, selling point text, and canvas size. Real-time input and offline data parsing are supported to provide basic material input for poster generation.
[0124] Step 2: Multimodal feature fusion: The multimodal information (images, text, dimensions) is uniformly formatted into input features that the model can understand, thus completing feature fusion.
[0125] Step 3: Supervised fine-tuning of the prompt word model: Supervised fine-tuning is performed based on a multimodal large language model to learn the structure, grammar, and professional norms of e-commerce poster prompt words, so that the model has the ability to generate legal, standardized, and complete poster prompt words.
[0126] Step 4: Candidate prompt word sampling: After initializing the model with prompt words, sample and generate multiple candidate prompt words.
[0127] Step 5: Poster rendering: Input multiple candidate prompts into the image generation model to render multiple poster samples for evaluation and reward scoring.
[0128] Step 6: Dual Reward Model Evaluation: The user appeal of the poster is evaluated using the User Reward Model (URM, i.e., the first reward model), and the aesthetic quality of the poster is evaluated from four dimensions: visual layout, text layout, appearance quality, and style consistency using the Designer Reward Model (DRM, i.e., the second reward model). The two scores are weighted to obtain the weighted fusion reward value of the corresponding poster based on multiple candidate cue words, thereby evaluating the quality of the cue words.
[0129] Step 7: Optimize the prompt word model: Based on the comprehensive reward value, use Grouped Relative Policy Optimization (GRPO) to perform reinforcement learning iterations on the prompt word generation model, update the model weights, and make the model tend to generate prompt words with higher rewards.
[0130] Step 8: Use the optimized prompt word model to generate better prompt words.
[0131] Step 9: Generate better prompts: Output a final graphic poster that balances aesthetics and high conversion rate after image rendering.
[0132] In an alternative embodiment, the following can be employed: Figure 5 The architecture shown implements the production of graphic posters, including: a cue word initialization module, a cue word optimization module, and a generation and evaluation module. The cue word initialization module sets up a cue word generation model. The cue word optimization module includes a cue word sampler, a dual-reward scorer, and GRPO cue word model optimization. The cue word optimization module and the generation and evaluation module optimize the cue word generation model to generate high-quality cue words. The product title, text, and size are acquired, and multimodal feature fusion is performed on them. The cue word generation model processes the multimodal feature fusion features to obtain the cue words, and the final poster is obtained through the image renderer in the generation and evaluation module.
[0133] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0135] Example 2
[0136] According to embodiments of this application, a method for generating graphic and textual content is also provided, such as... Figure 6 As shown, it includes:
[0137] Step S601: Obtain the target data information of the target object uploaded by the client and the canvas size information of the target graphic content to be generated. The target data information includes at least the image information of the target object and the description information of the target object.
[0138] Step S602: In the cloud server, the target data information and canvas size information are processed by a multimodal language model to obtain target prompt words for generating target graphic content; the target prompt words are processed by an image generation model to output the target graphic content.
[0139] Step S603: Return the target image and text content to the client.
[0140] It should be noted that the specific steps for generating text and image content on the cloud server are the same as in Example 1, and will not be repeated here.
[0141] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0143] Example 3
[0144] According to an embodiment of this application, an apparatus for generating graphic content for implementing the above-described method for generating graphic content is also provided, such as... Figure 7 As shown, the device includes: an acquisition unit 701, a first processing unit 702, and a second processing unit 703.
[0145] The acquisition unit 701 is used to acquire target data information of the target object and canvas size information of the target graphic content to be generated, wherein the target data information includes at least image information of the target object and description information of the target object;
[0146] The first processing unit 702 is used to process the target data information and canvas size information through a multimodal language model to obtain target prompt words for generating target graphic content;
[0147] The second processing unit 703 is used to process the target prompt words through an image generation model and output the target image and text content.
[0148] In the graphic content generation apparatus provided in Embodiment 3 of this application, the acquisition unit 701 acquires target data information of the target object and canvas size information of the target graphic content to be generated. The target data information includes at least image information and description information of the target object. The first processing unit 702 processes the target data information and canvas size information through a multimodal language model to obtain target prompt words for generating the target graphic content. The second processing unit 703 processes the target prompt words through an image generation model and outputs the target graphic content, thus solving the technical problem of low efficiency in generating graphic content in related technologies.
[0149] In this application, a multimodal language model is used to jointly model the received target object image information, descriptive information, and canvas size information to generate prompts that conform to visual structural specifications. The image generation model then outputs the target graphic content based on these prompts. This process integrates the previously manually designed multi-step operations (such as copywriting, composition planning, and style matching) into an end-to-end automated processing chain, reducing intermediate human intervention. Because the multimodal language model possesses the ability to jointly understand the semantics of images and text, it can adaptively generate structured prompts containing subject descriptions, composition guidance, and style instructions based on product characteristics and canvas proportions, thereby improving the consistency between the prompts and the input information. The image generation model renders based on the prompts, which helps reduce multiple redraws or manual corrections caused by ambiguous or incomplete prompts, improving the usability of a single generation. Without changing the underlying image generation capabilities, by optimizing the semantic quality and structural regularity of the input prompts, the technical effect of shortening the generation time of graphic content is achieved.
[0150] Optionally, in the image and text content generation apparatus provided in Embodiment 3 of this application, the multimodal language model is trained using the following apparatus: a third processing unit, used to process sample data information and sample canvas size information through the initial multimodal language model to obtain multiple sets of predicted prompt words, wherein each set of predicted prompt words includes multiple predicted prompt words; a calculation unit, used to perform reward calculation on the multiple sets of predicted prompt words to obtain target reward values corresponding to the multiple predicted prompt words respectively; and an update unit, used to update the initial multimodal language model based on the target reward values to obtain the multimodal language model.
[0151] Optionally, in the image and text content generation apparatus provided in Embodiment 3 of this application, the calculation unit includes: a processing subunit, used to process the plurality of predicted prompt words through an image generation model and output the initial image and text content corresponding to the plurality of predicted prompt words respectively; a calculation subunit, used to calculate the initial image and text content corresponding to the plurality of predicted prompt words respectively through multiple reward models to obtain an initial reward value; and a determination subunit, used to obtain a target reward value based on the initial reward value.
[0152] Optionally, in the image and text content generation apparatus provided in Embodiment 3 of this application, the calculation subunit includes: a first evaluation module, used to evaluate the click-through rate of the initial image and text content corresponding to multiple predicted prompt words through a first reward model, and obtain a first score; a second evaluation module, used to evaluate the image and text content quality of the initial image and text content corresponding to multiple predicted prompt words through a second reward model, and obtain a second score; and a determination module, used to obtain an initial reward value based on the first score and the second score.
[0153] Optionally, in the image and text content generation apparatus provided in Embodiment 3 of this application, the determining module includes: a first calculation submodule, used to perform a weighted calculation based on a first score value and a second score value to obtain a target score value; a second calculation submodule, used to calculate the average score value corresponding to any set of predicted prompt words based on the target score value; and a third calculation submodule, used to calculate an initial reward value based on the average score value and the target score value corresponding to the predicted prompt words in the set of predicted prompt words.
[0154] Optionally, in the graphic content generation apparatus provided in Embodiment 3 of this application, the first evaluation module includes: a first processing submodule, used to perform block encoding processing on the initial graphic content to obtain a visual feature vector; an encoding submodule, used to perform lexical encoding on the text corresponding to the initial graphic content to obtain a sequence semantic feature vector; and a first determination submodule, used to obtain a first score based on the visual feature vector and the sequence semantic feature vector.
[0155] Optionally, in the graphic content generation apparatus provided in Embodiment 3 of this application, the second evaluation module includes: a second determining submodule, used to obtain a layout score based on the layout information of visual elements in the initial graphic content; a third determining submodule, used to obtain a text layout score based on the text information in the initial graphic content; a fourth determining submodule, used to obtain an appearance texture score based on the color information in the initial graphic content; a fifth determining submodule, used to obtain a style consistency score based on the text style semantics and image style semantics in the initial graphic content; and a second processing submodule, used to obtain a second score based on the layout score, text layout score, text layout score, and style consistency score.
[0156] It should be noted that the acquisition unit 701, the first processing unit 702, and the second processing unit 703 mentioned above correspond to steps S201 to S203 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the device, can run on the computer terminal 10 provided in Embodiment 1.
[0157] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0158] Example 4
[0159] Embodiments of this application may provide an electronic device, which may be any one of a group of electronic device terminals. Optionally, in this embodiment, the aforementioned electronic device may also be replaced by a terminal device such as a mobile terminal.
[0160] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0161] In this embodiment, the above-mentioned electronic device can execute the program code of the following steps in the method for generating graphic content: obtaining target data information of the target object and canvas size information of the target graphic content to be generated, wherein the target data information includes at least image information and description information of the target object; processing the target data information and canvas size information through a multimodal language model to obtain target prompt words for generating target graphic content; and processing the target prompt words through an image generation model to output the target graphic content.
[0162] The aforementioned electronic device can execute the program code for the following steps in the method for generating text and image content: The multimodal language model is trained using the following steps: the initial multimodal language model processes the sample data information and sample canvas size information to obtain multiple sets of predicted prompt words, wherein each set of predicted prompt words includes multiple predicted prompt words; rewards are calculated for the multiple sets of predicted prompt words to obtain the target reward values corresponding to the multiple predicted prompt words respectively; the initial multimodal language model is updated based on the target reward values to obtain the multimodal language model.
[0163] The aforementioned electronic device can execute the program code for the following steps in the method for generating graphic content: calculating rewards for multiple sets of predicted prompt words to obtain target reward values corresponding to each of the multiple predicted prompt words, including: processing the multiple predicted prompt words through an image generation model to output initial graphic content corresponding to each of the multiple predicted prompt words; calculating the initial graphic content corresponding to each of the multiple predicted prompt words through multiple reward models to obtain initial reward values; and obtaining target reward values based on the initial reward values.
[0164] The aforementioned electronic device can execute the program code for the following steps in the method for generating text and image content: calculating the initial text and image content corresponding to multiple predicted prompt words using multiple reward models to obtain initial reward values, including: evaluating the click-through rate of the initial text and image content corresponding to multiple predicted prompt words using a first reward model to obtain a first score; evaluating the quality of the text and image content corresponding to multiple predicted prompt words using a second reward model to obtain a second score; and obtaining the initial reward value based on the first score and the second score.
[0165] The aforementioned electronic device can execute the program code for the following steps in the method for generating text and image content: obtaining an initial reward value based on a first score and a second score, including: performing a weighted calculation based on the first score and the second score to obtain a target score; for any set of predicted prompt words, calculating the average score corresponding to that set of predicted prompt words based on the target score; and calculating the initial reward value based on the average score and the target score corresponding to the predicted prompt words in that set of predicted prompt words.
[0166] The aforementioned electronic device can execute the following steps in the method for generating graphic content: evaluate the click-through rate of the initial graphic content corresponding to multiple predicted prompt words through a first reward model to obtain a first score, including: performing block encoding processing on the initial graphic content to obtain a visual feature vector; performing lexical encoding on the text corresponding to the initial graphic content to obtain a sequence semantic feature vector; and obtaining the first score based on the visual feature vector and the sequence semantic feature vector.
[0167] The aforementioned electronic device can execute the following steps in the method for generating graphic content: It evaluates the quality of the initial graphic content corresponding to multiple predicted prompts using a second reward model to obtain a second score, including: obtaining a layout score based on the layout information of visual elements in the initial graphic content; obtaining a text layout score based on the text information in the initial graphic content; obtaining an appearance quality score based on the color information in the initial graphic content; obtaining a style consistency score based on the text style semantics and image style semantics in the initial graphic content; and obtaining the second score based on the layout score, text layout score, text layout score, and style consistency score.
[0168] Optionally, Figure 8 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 8 As shown, the electronic device 80 may include: one or more ( Figure 8(Only one is shown in the image) Processor 802 and memory 804. The electronic device 80 may also include a memory controller to control and manage the memory 804; the electronic device 80 may also include a peripheral interface to connect to a radio frequency module, an audio module, and a display screen, etc.
[0169] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image and text content generation method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the above-mentioned image and text content generation method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the electronic device 80 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0170] The processor can invoke the information and application program stored in the memory through the transmission device to perform the following steps: acquiring the target data information of the target object and the canvas size information of the target graphic content to be generated, wherein the target data information includes at least the image information and description information of the target object; processing the target data information and canvas size information through a multimodal language model to obtain target prompt words for generating the target graphic content; and processing the target prompt words through an image generation model to output the target graphic content.
[0171] Optionally, the processor may also execute program code for the following steps: The multimodal language model is trained using the following steps: the sample data information and sample canvas size information are processed through the initial multimodal language model to obtain multiple sets of prediction prompts, wherein each set of prediction prompts includes multiple prediction prompts; rewards are calculated for the multiple sets of prediction prompts to obtain target reward values corresponding to each prediction prompt; the initial multimodal language model is updated based on the target reward values to obtain the multimodal language model.
[0172] Optionally, the processor may also execute program code for the following steps: calculating rewards for multiple sets of predicted prompts to obtain target reward values corresponding to each of the predicted prompts, including: processing the multiple predicted prompts through an image generation model to output initial image and text content corresponding to each of the multiple predicted prompts; calculating the initial image and text content corresponding to each of the multiple predicted prompts through multiple reward models to obtain initial reward values; and obtaining target reward values based on the initial reward values.
[0173] Optionally, the processor may also execute program code that performs the following steps: calculates the initial reward value for the initial text and image content corresponding to the multiple predicted prompt words using multiple reward models, including: evaluating the click-through rate of the initial text and image content corresponding to the multiple predicted prompt words using a first reward model to obtain a first score; evaluating the text and image content quality of the initial text and image content corresponding to the multiple predicted prompt words using a second reward model to obtain a second score; and obtaining the initial reward value based on the first score and the second score.
[0174] Optionally, the processor may also execute program code that performs the following steps: obtaining an initial reward value based on a first score and a second score, including: performing a weighted calculation based on the first score and the second score to obtain a target score; for any set of prediction prompts, calculating the average score corresponding to the set of prediction prompts based on the target score; and calculating the initial reward value based on the average score and the target score corresponding to the prediction prompts in the set of prediction prompts.
[0175] Optionally, the processor may also execute program code for the following steps: evaluating the click-through rate of the initial text and image content corresponding to multiple predicted prompt words through a first reward model to obtain a first score, including: performing block encoding on the initial text and image content to obtain a visual feature vector; performing lexical encoding on the text corresponding to the initial text and image content to obtain a sequence semantic feature vector; and obtaining the first score based on the visual feature vector and the sequence semantic feature vector.
[0176] Optionally, the processor may also execute program code for the following steps: evaluating the quality of the initial text and image content corresponding to multiple predicted prompts using a second reward model to obtain a second score, including: obtaining a layout score based on the layout information of visual elements in the initial text and image content; obtaining a text layout score based on the text information in the initial text and image content; obtaining an appearance quality score based on the color information in the initial text and image content; obtaining a style consistency score based on the text style semantics and image style semantics in the initial text and image content; and obtaining a second score based on the layout score, text layout score, text layout score, and style consistency score.
[0177] Those skilled in the art will understand that Figure 8 The structure shown is for illustrative purposes only. Electronic device 80 can also be a smartphone, tablet computer, handheld computer, mobile internet device (MID), PAD and other terminal devices. Figure 8 This does not limit the structure of the aforementioned electronic device. For example, electronic device 80 may also include components that are more... Figure 8The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 8 The different configurations shown.
[0178] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0179] Example 5
[0180] Embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product can be used to store the program code executed by the method for generating text and image content provided in Embodiment 1.
[0181] Optionally, in this embodiment, the computer program product may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0182] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0183] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0184] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of units or modules may be electrical or other forms.
[0185] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0186] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0187] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0188] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for generating graphic and textual content, characterized in that, include: Obtain target data information of the target object and canvas size information of the target graphic content to be generated, wherein the target data information includes at least image information of the target object and description information of the target object; The target data information and the canvas size information are processed by a multimodal language model to obtain target prompt words for generating the target graphic content. The multimodal language model is trained based on reward values calculated by multiple reward models. The target prompt words are processed using an image generation model to output the target image and text content.
2. The method according to claim 1, characterized in that, The multimodal language model is trained using the following steps: The sample data and sample canvas size information are processed by the initial multimodal language model to obtain multiple sets of predicted prompt words, where each set of predicted prompt words includes multiple predicted prompt words. Rewards are calculated for the multiple sets of predicted prompts to obtain target reward values for each of the predicted prompts. The initial multimodal language model is updated based on the target reward value to obtain the multimodal language model.
3. The method according to claim 2, characterized in that, Rewards are calculated for the multiple sets of predicted prompts to obtain target reward values for each predicted prompt, including: The image generation model processes the multiple predicted prompt words and outputs the initial image and text content corresponding to each of the multiple predicted prompt words. The initial reward value is obtained by calculating the initial text and image content corresponding to the multiple predicted prompt words using multiple reward models; Based on the initial reward value, the target reward value is obtained.
4. The method according to claim 3, characterized in that, Initial reward values are obtained by calculating the initial text and image content corresponding to the multiple predicted prompts using multiple reward models, including: The first reward model is used to evaluate the click-through rate of the initial text and image content corresponding to the multiple predicted prompt words to obtain the first score. The second reward model is used to evaluate the quality of the initial text and image content corresponding to the multiple predicted prompt words, and a second score is obtained. The initial reward value is obtained based on the first score and the second score.
5. The method according to claim 4, characterized in that, The initial reward value is obtained based on the first score and the second score, including: The target score is obtained by weighting the first score and the second score. For any set of predicted prompt words, calculate the average score corresponding to that set of predicted prompt words based on the target score value; The initial reward value is calculated based on the average score and the target score corresponding to the predicted prompt words in the group of predicted prompt words.
6. The method according to claim 4, characterized in that, The first reward model is used to evaluate the click-through rate of the initial text and image content corresponding to the multiple predicted prompts, and a first score is obtained, including: The initial image and text content is segmented and encoded to obtain a visual feature vector; The text corresponding to the initial image and text content is lexicalized and encoded to obtain a sequence semantic feature vector; The first score is obtained based on the visual feature vector and the sequence semantic feature vector.
7. The method according to claim 4, characterized in that, The second reward model is used to evaluate the quality of the initial text and image content corresponding to the multiple predicted prompts, resulting in a second score, including: Based on the layout information of the visual elements in the initial text and image content, a layout score is obtained; Based on the text information in the initial text and image content, a text layout score is obtained; Based on the color information in the initial graphic content, the appearance texture score is obtained; Based on the text style semantics and image style semantics in the initial text and image content, a style consistency score is obtained; The second score is obtained based on the layout score, the text layout score, the text layout score, and the style consistency score.
8. A method for generating graphic and textual content, characterized in that, include: Obtain the target data information of the target object uploaded by the client and the canvas size information of the target graphic content to be generated, wherein the target data information includes at least the image information of the target object and the description information of the target object; In a cloud server, the target data information and the canvas size information are processed using a multimodal language model to obtain target prompt words for generating the target graphic content; the target prompt words are then processed using an image generation model to output the target graphic content. The target text and image content is returned to the client.
9. A device for generating graphic and textual content, characterized in that, include: The acquisition unit is used to acquire target data information of the target object and canvas size information of the target graphic content to be generated, wherein the target data information includes at least image information of the target object and description information of the target object; The first processing unit is used to process the target data information and the canvas size information through a multimodal language model to obtain target prompt words for generating the target graphic content; The second processing unit is used to process the target prompt words through an image generation model and output the target image and text content.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the storage medium is located to perform the method for generating graphic content according to any one of claims 1 to 8.
11. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, executes the method for generating graphic content according to any one of claims 1 to 8.
12. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the method for generating graphic content as described in any one of claims 1 to 8.