Picture material generation method and device, computer equipment and storage medium

By identifying and replacing element information in advertising image materials, image materials that meet new requirements are generated, solving the problem of unusable materials, improving material production efficiency and creative iteration speed, and enhancing the operational efficiency of the advertising system.

CN121505079APending Publication Date: 2026-02-10BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511489740.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In existing advertising image materials, elements such as brand logos and exclusive copy are deeply integrated with the main visual content of the image, lacking an effective element separation mechanism. This results in the inability to accurately identify, separate, and reuse the materials, leading to a waste of creative resources and reducing the efficiency of advertising material production and the speed of creative iteration.

Method used

By identifying element information in the original image material, generating prompt words, and using the target model to replace elements, new image material is generated by combining the requirements information, achieving accurate element identification and replacement, retaining most of the content of the original image while adapting to new requirements.

Benefits of technology

It improved the efficiency of material reuse, reduced the waste of creative resources, increased the production speed of advertising materials and the speed of creative iteration, and improved the overall operational efficiency of the advertising system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505079A_ABST
    Figure CN121505079A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a picture material generation method and device, computer equipment and a storage medium, and the method comprises the steps: recognizing included element information for each original picture material in an original picture material set, and obtaining an element information set; determining target element information from the element information set according to demand information of to-be-generated new picture content; and generating a cue word according to the demand information and the target element information, so that the target model generates a new picture material according to the cue word, the demand information and the target element information. Therefore, the element information in the original material can be accurately recognized, the target element can be determined according to needs, the targeted cue word is generated in combination with the needs to guide the target model to generate the new material, adaptation to the new needs is achieved while most of the content of the original picture is reserved, the material reuse efficiency and the production speed are remarkably improved, and the user experience is improved. The creative iteration of the advertisement materials is accelerated, and the overall operation efficiency of the advertisement system is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a method, apparatus, computer device and storage medium for generating image materials. Background Technology

[0002] In performance advertising, image creatives serve as the core conversion vehicle, and their creative quality directly impacts ad click-through rates and conversion rates. To improve campaign effectiveness, advertisers typically invest significant resources in creating image creatives that include brand elements such as logos, customized copy, and exclusive stickers. Some of these well-designed and creatively outstanding creatives not only perform exceptionally well in the current campaign context but also possess the potential for reuse across advertisers and industries, serving as high-quality creative templates to support the rapid production of new creatives.

[0003] However, the reuse of existing advertising image assets faces significant technical bottlenecks. Currently, brand logos, exclusive copy, and other elements are deeply integrated with the main visual content of the image, lacking an effective mechanism for separating these elements. This makes it impossible to accurately identify and separate personalized elements such as logos, copy, and icons, and also makes it difficult to preserve the composition, style, and visual logic of the core creative concept. This indivisible characteristic means that the core creative concept of the assets cannot be extracted independently, making it impossible to adapt to other advertisers' brand needs, derivative placement scenarios, or reuse. A large amount of high-quality creative resources are confined to a single advertiser or a single placement, resulting in serious creative waste.

[0004] The aforementioned technical deficiencies further restrict the efficiency of advertising material production and the overall effectiveness of the advertising system: on the one hand, advertisers need to repeatedly invest costs in developing new materials and cannot iterate quickly based on historical success cases; on the other hand, advertising platforms have difficulty in scalably utilizing existing high-quality materials and cannot provide efficient creative generation solutions for advertisers in different industries and brands, ultimately resulting in a low level of creative iteration speed for advertising materials and resource utilization efficiency of the delivery system. Summary of the Invention

[0005] In view of this, in order to solve the above-mentioned technical problems or some of the technical problems, the embodiments of the present invention provide a method, apparatus, computer equipment and storage medium for generating image materials.

[0006] In a first aspect, embodiments of the present invention provide a method for generating image materials, including: For each original image in the original image material set, identify the element information it contains to obtain the element information set; Based on the requirements for the content of the new image to be generated, the target element information is determined from the set of element information; Based on the requirement information and the target element information, prompt words are generated so that the target model can generate new image materials based on the prompt words, the requirement information, and the target element information.

[0007] In one possible implementation, identifying the element information contained in each original image in the original image material set includes: Standardized preprocessing is performed on each of the original image materials; The preprocessed original image materials are analyzed using a trained analytical model to extract element information contained in each original image material. The element information includes at least one of the following: text content and text position coordinates in the image, graphic content and graphic position coordinates and outline mask, position coordinates and outline mask of the brand logo, main color tone information, and image scene and style information.

[0008] In one possible implementation, determining the target element information from the element information set based on the requirement information of the new image content to be generated includes: Receive the requirement information, which includes at least one of the following: text requirements, graphic requirements, brand logo requirements, main color requirements, and image scene and style requirements for the new image content to be generated; Perform semantic association and fit analysis between the requirement information and the set of element information; Based on the analysis results, the element information that does not match the required information is determined as the target element information, which represents the element information that needs to be replaced for the original image material.

[0009] In one possible implementation, generating prompt words based on the demand information and the target element information includes: New text content is generated based on the text requirements in the requirement information, and the new text content is used to replace the text content in the target element information; And / or, generate new graphic content based on the graphic requirements in the requirement information, the new graphic content being used to replace the graphic content in the target element information; And / or, generate a new brand identity based on the brand identity requirements in the requirement information, the new brand identity being used to replace the brand identity in the target element information; And / or, generate new main color information based on the main color requirements in the requirement information, and use the new main color information to replace the main color information in the target element information; And / or, generate new image scene and style information based on the image scene and style requirements in the requirement information, and use the new image scene and style information to replace the image scene and style information in the target element information; The prompt word is generated based on the new text content, the new graphic content, the new brand logo, the new main color scheme, and the new image scene and style information.

[0010] In one possible implementation, the target model generates new image materials based on the prompt words, the demand information, and the target element information, including: After the target model replaces the target element information in the original image material according to the prompt words, it verifies and adjusts the image material after the replacement operation according to the requirement information to obtain the new image material.

[0011] In one possible implementation, the step of verifying and adjusting the image material after the replacement operation based on the requirement information includes: Based on the aforementioned requirements, verify at least one of the following aspects for the replaced image materials: completeness of element information, consistency of scene style, compliance with brand specifications, and visual harmony. If any verification fails, the failed content will be adjusted. And / or, calculate the semantic similarity score between the replaced image material and the requirement information through a semantic matching model; If the score is less than a preset threshold, the step of identifying the element information contained in each original image in the original image material set is re-executed.

[0012] In one possible implementation, the method further includes: The original image set is divided into multiple subsets with different image scenes and styles based on the image scene and style information of each original image material; For each subset, the original image material element information contained therein is identified to obtain the element information set corresponding to each subset; When a requirement is received, the target subset is determined based on the image scene and style requirements in the requirement information, and then the target element information is determined from the element information set corresponding to the target subset.

[0013] Secondly, embodiments of the present invention provide an image material generation apparatus, comprising: The recognition module is used to identify the element information contained in each original image in the original image material set, and obtain the element information set. The determination module is used to determine target element information from the element information set based on the requirement information of the new image content to be generated; The generation module is used to generate prompt words based on the requirement information and the target element information, so that the target model can generate new image materials based on the prompt words, the requirement information and the target element information.

[0014] Thirdly, embodiments of the present invention provide a computer device, including: a processor and a memory, wherein the processor is configured to execute an image material generation program stored in the memory to implement the image material generation method described in any one of the first aspects above.

[0015] Fourthly, embodiments of the present invention provide a storage medium storing one or more programs, which can be executed by one or more processors to implement the image material generation method described in any one of the first aspects.

[0016] The image material generation method provided in this invention identifies the element information contained in each original image material in an original image material set to obtain an element information set; determines target element information from the element information set based on the requirement information of the new image content to be generated; and generates prompt words based on the requirement information and the target element information, so that the target model can generate new image materials according to the prompt words, the requirement information, and the target element information. Thus, by accurately identifying element information in the original materials and determining target elements as needed, and generating targeted prompt words based on requirements to guide the target model in generating new materials, the method retains most of the content of the original images while adapting to new requirements. This avoids the waste of high-quality creative resources, significantly improves the efficiency of material reuse and production speed, accelerates the creative iteration of advertising materials, and ultimately improves the overall operational efficiency of the advertising system. Attached Figure Description

[0017] Figure 1 A schematic flowchart illustrating an image material generation method provided in an embodiment of the present invention; Figure 2 A flowchart illustrating another image material generation method provided in an embodiment of the present invention; Figure 3 A schematic flowchart illustrating another image material generation method provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an image material generation device provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] To facilitate understanding of the embodiments of the present invention, further explanations and descriptions will be provided below with reference to the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.

[0020] Figure 1 This is a flowchart illustrating a method for generating image materials according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method specifically includes: S11. For each original image in the original image material set, identify the element information it contains to obtain the element information set.

[0021] The image material generation method provided in this invention is applied to computer devices, including but not limited to servers, desktop computers, and tablets. It can be applied to the following scenarios: advertising platforms can batch replace elements such as images, text, and styles in historical image materials according to different advertiser needs, quickly generating new materials adapted to multiple industries. Specifically, it accurately identifies element information in the original materials and determines target elements as needed, generating targeted prompts based on requirements to guide the target model in generating new materials. This achieves adaptation to new requirements while retaining most of the original image content.

[0022] In this embodiment, the original image material set refers to multiple advertising images generated in history. First, each original image material is preprocessed (e.g., the image format is uniformly converted to a uniform general format, and scaled proportionally according to a preset resolution, etc.).

[0023] Element information may include, but is not limited to: text in the image material and its size and position in the image material, graphics (advertiser logo, stickers, etc.) and their size and position in the image material, color tone, style, etc.

[0024] Image segmentation models (e.g., U-Net or Mask R-CNN) are used to perform pixel-level analysis on the preprocessed images, accurately locating and segmenting graphic elements such as logos and custom stickers. The pixel-level contour mask for each graphic element is output, and its top-left corner coordinates (x, y), width, height, and alpha value are recorded as element information. Simultaneously, an OCR model (e.g., PaddleOCR or Tesseract) is used to scan the text regions in the image material. A text detection algorithm (e.g., EAST text detection algorithm) is combined to determine the location range of the text, extracting the text content while recording the font type, font size, color (RGB value), and layout direction (horizontal / vertical) as element information. Simultaneously, the overall scene and style of the image are analyzed through a text-image understanding model (e.g., CLIP or ViLT model), and descriptive text (e.g., "Summer beverage promotion scene, fresh watercolor style") is generated. Color extraction algorithms (e.g., K-Means clustering algorithm) are called to analyze the pixel distribution of the image and determine the main color tone with the highest proportion of a preset number of colors (output RGB values ​​and proportions).

[0025] The identified graphic element information (logo, sticker position, mask, size, etc.), text element information (text content, font attributes, position), style scene information, and main color tone information are encapsulated according to a preset data structure (e.g., JSON format) to form element information corresponding to a single original image, and a correspondence is established between each element information and its corresponding original image material. After all images in the original image material set have completed the above identification process, all element information is summarized to generate an element information set containing complete element information for each image, which is stored in the database to provide structured data support for subsequent target element selection and new material generation.

[0026] S12. Based on the requirements of the new image content to be generated, determine the target element information from the element information set.

[0027] In this embodiment, the requirement information represents the user input or external system (advertiser)'s requirement to replace or adapt elements in the new image content of the advertisement to be generated. This may include, but is not limited to: brand logo replacement (e.g., replacing the brand logo in the original advertisement with a new brand logo), copy replacement (e.g., requiring the generation of a new slogan or changing existing text content), graphic replacement (e.g., replacing a sticker or icon in the original image with a new graphic), color tone adjustment (e.g., adjusting the overall color tone to green), style adjustment (e.g., requiring the overall style to be adjusted to a technological style, etc.), and composition preservation requirement (e.g., requiring the original main character or scene to be retained, only replacing local decorations or color schemes).

[0028] Target element information refers to the element information that is selected from the set of element information based on the demand information and needs to be replaced or adjusted in the subsequent generation of new advertisements.

[0029] After receiving and parsing the requirement information, the system compares it one by one with the element information set to determine which elements should be used as target element information. Target element information can be determined in the following ways: 1. Keyword matching: When the requirement information includes "replace LOGO", the system will extract the element information corresponding to the brand LOGO from the element information set as the target element information; when the requirement information includes "modify text", the system will extract the element information corresponding to the text. 2. Rule reasoning: When the requirement information requests a change in the overall color scheme, the system will automatically lock the color scheme as the target element information. If the requirement information requests to retain the main composition, the system will automatically exclude the element information corresponding to people and background segmentation areas, and only use the LOGO, text, and other local element information as the target element information. 3. Semantic understanding: When the requirement information is input in natural language, such as "change this beach advertisement to an outdoor hiking style", the system will call the semantic parsing module to identify that "beach" corresponds to the original scene description information, while "outdoor hiking" corresponds to the new replacement requirement, thus extracting the style as the target element information.

[0030] As an example, in an original hotel advertisement image, the set of element information includes "text, logo, beach stickers, and a blue and gold color scheme". When the requirement is "replace with the Mountain Peak brand advertisement and change the color scheme to green and brown", the system will automatically identify the logo and color scheme as the target element information, while keeping the text and stickers unchanged.

[0031] S13. Generate prompt words based on the requirement information and target element information, so that the target model can generate new image materials based on the prompt words, requirement information and target element information.

[0032] In this embodiment, the prompts are control text descriptions input to the image generation model, typically composed of natural language, used to guide the model in generating image content that meets specified requirements. The target model is an artificial intelligence model used to generate or modify advertising images.

[0033] Specifically, the requirement information is converted into a natural language description and combined with the target element information to generate prompts. For example, if the requirement information is "replace the LOGO with the Mountain Peak brand," and the target element information includes the LOGO, its location, and size, the prompt will include "replace the Mountain Peak LOGO in the upper right corner of the original image." If the requirement information is "adjust the main color scheme to green and brown," and the target element information includes color scheme, the prompt will include "the overall color scheme is mainly green and brown." During the construction process, the system calls a language generation model (e.g., an LLM model, Large Language Model) to refine and expand the natural language, ensuring the prompts are semantically clear, stylistically consistent, and conform to the input habits of the generation model.

[0034] The generated prompts generally include, but are not limited to, the following: 1. Scene Description: Extracted from the requirements information, such as "outdoor mountain climbing scene, with blue sky and mountains as the background." 2. Element Replacement Instructions: Mapped from the target element information, such as "Draw a new logo at coordinates (x, y, w, h)." 3. Style and Color Requirements: Extracted from the requirements information, such as "The overall image style is natural and realistic, with green and brown as the main colors." 4. Composition Preservation Requirements: Ensure the original subject remains unchanged, such as "Preserve the figure's posture and background structure." This generates prompts. Prompts generated in this way not only contain semantic descriptions but also clear operational constraints, enabling the subsequent generation model to maintain consistency with the original image structure while adhering to user requirements.

[0035] As an example, let's consider the requirement to "replace a beach hotel advertisement with a mountain outdoor brand advertisement." The requirements are: brand replacement (replace beach hotel with mountain outdoor), and main color adjustment (change blue / gold to green / brown). Target element information includes: LOGO (coordinates x1, y1, w, h) and color scheme (blue, gold). The generated prompt is: "A realistic outdoor mountaineering advertisement image, retaining the original figures and composition, replace the top right corner (x1, y1) with the mountain logo, and adjust the overall color scheme to green and brown to present a natural and magnificent mountaineering atmosphere." After this prompt is input into the target model, the model can generate a new advertisement image that meets the requirements based on the prompt, requirement information, and target element information for display.

[0036] The image material generation method provided in this invention identifies the element information contained in each original image material in an original image material set to obtain an element information set; determines target element information from the element information set based on the requirement information of the new image content to be generated; and generates prompt words based on the requirement information and the target element information, so that the target model can generate new image materials according to the prompt words, the requirement information, and the target element information. Thus, by accurately identifying element information in the original materials and determining target elements as needed, and generating targeted prompt words based on requirements to guide the target model in generating new materials, the method retains most of the content of the original images while adapting to new requirements. This avoids the waste of high-quality creative resources, significantly improves the efficiency of material reuse and production speed, accelerates the creative iteration of advertising materials, and ultimately improves the overall operational efficiency of the advertising system.

[0037] Figure 2 This is a flowchart illustrating another image material generation method provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the method specifically includes: S21. Perform standardized preprocessing on each original image material; parse the preprocessed original image material using a trained parsing model, and extract the element information contained in each original image material. The element information shall include at least one of the following: text content and text position coordinates in the image, graphic content and graphic position coordinates and outline mask, position coordinates and outline mask of the brand logo, main color tone information, and image scene and style information.

[0038] In this embodiment, each original image material is standardized in terms of size, format, and color space to ensure the stability and consistency of the subsequent parsing process. For example, input images with different resolutions can be scaled to a uniform target size, the image format can be standardized to PNG or JPEG, and the color mode can be converted to RGB mode as needed to eliminate differences between the original image materials.

[0039] After preprocessing, the trained parsing model is invoked to parse the standardized original image materials. The parsing model can consist of a computer vision model and a multimodal understanding model, such as an image segmentation model (e.g., U-Net or Fast R-CNN), an optical character recognition (OCR) model (for recognizing text content), and a joint image-text understanding model (e.g., CLIP or ViLT). These models work together to extract element information from each original image material. Element information includes at least one or more of the following: 1. Text content in the image and its position coordinates: The OCR model recognizes the text information in the image and outputs the specific position coordinates of the text in the image (e.g., the x and y coordinates of the top-left corner, and the width and height of the text region). This information is used for subsequent text replacement or layout adjustments. 2. Graphic Content and its Position Coordinates and Contour Mask: Image segmentation models identify graphic objects such as stickers, icons, and decorative elements in the advertisement image, generating their corresponding position coordinates and contour masks (masks used to mark the precise shape region of the graphic). 3. Brand Logo Position Coordinates and Contour Mask: Object detection and segmentation models identify the advertiser's brand logo's position coordinates in the image and extract its corresponding contour mask for subsequent brand replacement or enhancement. 4. Main Color Information: Color clustering or histogram analysis extracts the main color information of the advertisement image, representing it in RGB (red, green, blue) numerical form to guide subsequent color replacement or style transfer. 5. Image Scene and Style Information: A textual understanding model automatically generates textual information describing the overall scene and style of the image, such as "beach resort scene, fresh and natural style" or "urban night scene, technological style." This information helps maintain semantic and visual consistency between the generated image and the original composition.

[0040] S22. Receive requirement information, which includes at least one of the following: textual requirements, graphic requirements, brand logo requirements, main color requirements, image scene and style requirements for the new image content to be generated; perform semantic association and adaptation analysis on the requirement information and the set of element information; determine the element information that does not match the requirement information as the target element information based on the analysis results, and the target element information represents the element information that needs to be replaced for the original image material.

[0041] In this embodiment, the requirement information is used to characterize the user's or external system's requirements for the content of the new image to be generated. The requirement information includes at least one of the following aspects: 1. Textual requirements: indicating the text content that needs to be added, modified, or deleted in the new image, for example, replacing the original text "Summer Beach Special Offer" with "Professional Mountaineering Equipment Limited-Time Offer". 2. Graphical requirements: indicating the graphic elements that need to be replaced or adjusted, for example, removing the "beach umbrella icon" from the original advertisement and adding a new "mountaineering flag icon". 3. Brand logo requirements: indicating the replacement or enhancement of the brand logo, for example, replacing the "beach hotel logo" in the original image with the "mountain peak outdoor logo". 4. Main color tone requirements: indicating the modification of the overall image color tone, for example, adjusting blue and gold to green and brown to match the style of the new brand. 5. Image scene and style requirements: indicating changes in the overall style or scene of the image, for example, replacing "vacation and leisure scene" with "outdoor mountaineering scene", or replacing "realistic style" with "illustration style".

[0042] After obtaining the requirement information, semantic association and suitability analysis are performed between the requirement information and the set of element information. Semantic association refers to analyzing the semantic matching relationship between the requirement information and the parsed element information using a Natural Language Processing Model (NLP) or a Multimodal Understanding Model (MUM). For example, if the requirement information specifies "replace the LOGO", the system will locate element information related to the brand identity in the element information set as the target element information and mark it as a potential replacement object.

[0043] Fit analysis refers to determining whether existing elements meet the requirements by calculating the degree of matching between the requirement information and the element information. For example, when the requirement information specifies "the main color tone is green and brown," if the main color tone extracted from the element information set is blue and gold, then this color tone information does not match the requirement information, has a low fit, and is identified as the target element information, which needs to be replaced.

[0044] Based on the results of semantic association and fit analysis, elements that do not match the requirements are identified and defined as target elements. Target elements represent the content that needs to be replaced or adjusted from the original image material. For example, if the requirement is "replace the logo and modify the color tone," and both the logo information and the main color tone information in the element information set do not match the requirements, then both the logo information and the main color tone information are simultaneously identified as target elements.

[0045] S23. Generate prompt words based on the demand information and target element information.

[0046] In this embodiment, new text content is generated based on the textual requirements in the requirement information, and this new text content is used to replace the textual content in the target element information; and / or, new graphic content is generated based on the graphic requirements in the requirement information, and this new graphic content is used to replace the graphic content in the target element information; and / or, a new brand logo is generated based on the brand logo requirements in the requirement information, and this new brand logo is used to replace the brand logo in the target element information; and / or, new main color information is generated based on the main color requirements in the requirement information, and this new main color information is used to replace the main color information in the target element information; and / or, new image scene and style information is generated based on the image scene and style requirements in the requirement information, and this new image scene and style information is used to replace the image scene and style information in the target element information; prompt words are generated based on the new text content, new graphic content, new brand logo, new main color information, and new image scene and style information.

[0047] Specifically, after determining the target element information, the next step is to generate replacement content for the target element information based on the requirement information, which will then serve as prompt words. The generation process includes at least one of the following aspects: 1. Text content replacement When the requirement information includes textual requirements, the natural language generation model is invoked to generate new text content based on the requirement information and the original text content. This new text content is then used to replace the original text content in the target element information. For example, if the requirement information is "Replace the text with 'Explore magnificent mountains and rivers, embark on an outdoor journey'", the system will generate the corresponding new text content and prepare it for subsequent rendering.

[0048] 2. Replacement of graphic content When the requirement information includes a graphic requirement, the graphics generation model is invoked to generate new graphic content. This new graphic content replaces the original graphic content in the target element information. For example, the "beach umbrella icon" in the original advertisement image is replaced with a "mountain climbing flag icon".

[0049] 3. Brand logo replacement When the requirement information includes a brand identity requirement, a new brand identity file is received or generated, resulting in a new brand identity. This new brand identity is used to replace the original brand identity in the target element information. For example, the "Beach Hotel LOGO" in the original image is replaced with the "Mountain Peak Outdoor LOGO".

[0050] 4. Replacement of main color tone information When the requirements include a primary color scheme, new primary color scheme information is generated using color mapping or color conversion algorithms. This new primary color scheme information replaces the original color scheme information in the target element information. For example, blue and gold are adjusted to green and brown to match the visual style of the new brand.

[0051] 5. Replacement of image scenes and styles When the requirement information includes image scene and style requirements, new image scene and style information is generated through an image understanding and generation model. This new image scene and style information is then used to replace the original scene and style information in the target element information. For example, "beach vacation scene, fresh and natural style" is replaced with "outdoor mountain climbing scene, naturalistic style".

[0052] After generating the replacement content, the system will construct prompts based on the new text content, new graphic content, new brand logo, new main color scheme, and new image scene and style information. These prompts will guide the model in generating new image content that meets the requirements. During the construction of the prompts, the system will semantically combine the replacement content with the retained content to ensure that the prompts accurately express the user's needs while preserving the main composition and visual features of the original image.

[0053] For example, in the application scenario of replacing "beach hotel advertisement" with "mountain peak outdoor advertisement", the prompt may be described as: "A realistic outdoor mountain climbing advertisement image, retaining the original figures and background structure, replacing the mountain peak logo in the upper right corner, adjusting the overall color tone to green and brown, and replacing the text content with 'Explore magnificent mountains and rivers, start an outdoor journey', presenting a natural and magnificent mountain climbing scene." S24. After the target model replaces the target element information in the original image material according to the prompt words, it verifies and adjusts the image material after the replacement operation according to the requirement information to obtain the new image material.

[0054] In this embodiment, the target model can be an image generation model, an image editing model, or a multimodal large model, etc. After receiving the prompt word, it performs a replacement operation on the target element information identified in the original image material. First, the target model locates the corresponding target element region in the original image material based on the semantic content of the prompt word. The location method can include: text position coordinates, graphic outline mask, brand logo position region, main color tone distribution range, and scene and style semantic tags, etc. Through this element information, the target model can accurately locate the element information that needs to be replaced in the original image.

[0055] The target model embeds new element information into the target area based on the generation requirements described in the prompt. For example, when the prompt includes "new text content," the model generates new text at the target text location coordinates, ensuring harmony between font style, color, and background environment. When the prompt includes "new graphic content," the model uses the outline mask of the original graphic to replace it with new graphic elements, while maintaining smooth edges and natural integration with the background. When the prompt includes "new brand logo," the model generates the brand logo in the specified coordinate area and fits it according to the perspective and lighting relationships of the image. When the prompt includes "new main color tone information," the model adjusts the color distribution of the overall image or a specific area to ensure the color tone matches the requirements while maintaining the overall visual harmony of the image. When the prompt includes "new image scene and style information," the model migrates and replaces the background environment, material texture, and style characteristics while maintaining the main content structure.

[0056] After the replacement operation is completed, the target model does not directly output the result, but instead performs a round of verification and adjustment, including: Based on the requirements information, verify at least one of the following for the replaced image materials: element information completeness, scene style consistency, brand specification compliance, and visual harmony; if any verification fails, adjust the failed content; and / or, calculate the semantic similarity score between the replaced image materials and the requirements information using a semantic matching model; if the score is less than a preset threshold, re-execute the step of identifying the element information contained in each original image material in the original image material set.

[0057] In this embodiment, element information integrity means checking whether the necessary element information, such as text, graphics, and brand logos, is retained in the image to avoid omissions or losses during the replacement process. For example, when the requirement information specifies that a particular brand logo should be retained, the system needs to confirm that the logo exists and is clearly visible in the replaced image.

[0058] Scene style consistency means verifying whether the scene and style of the newly generated image are consistent with the target description in the requirements information and maintain coordination with the overall composition of the original material. For example, if the requirements information specifies "minimalist style", the system needs to detect whether the image has complex textures or redundant backgrounds that contradict "minimalism".

[0059] Brand compliance means that when the requirements involve the brand logo, the generated logo must be verified to conform to the brand design guidelines, including color, proportion, font, and typography. For example, are the brand's standard colors applied correctly, and is the logo stretched or distorted?

[0060] Visual harmony refers to evaluating the light and shadow relationships, perspective effects, edge blending, and color matching of the replaced elements in the overall image to ensure that the final image looks natural.

[0061] If any of the above verification criteria fails, the system will automatically adjust the affected content. For example, it may regenerate the text, correct the color tone, adjust the marker position, or replace the style transfer model to improve image quality.

[0062] A semantic matching model can also be introduced to evaluate the semantic similarity between the generated image and the requirement information. This model uses natural language processing and visual semantic understanding technologies to map the image content and requirement information to the same semantic space and calculates a similarity score. When the similarity score is greater than or equal to a preset threshold (e.g., 0.85), it indicates that the generated image highly matches the requirement information and can be used as the final output. When the similarity score is less than the preset threshold, it means that the generated result does not meet the requirement information. In this case, the operation of "identifying the element information contained in each original image in the original image material set" is re-executed to ensure that subsequent generation processes can incorporate more suitable element information.

[0063] In one possible implementation, the original image set is divided according to the image scene and style information of each original image material to obtain multiple subsets with different image scenes and styles; for each subset, the element information of the original image materials contained therein is identified to obtain the element information set corresponding to each subset; when the requirement information is received, the target subset is determined according to the image scene and style requirements in the requirement information, and the target element information is determined from the element information set corresponding to the target subset.

[0064] In this embodiment, the image scene refers to the environment and semantic context presented in the image, such as "indoor office scene," "outdoor natural scene," and "city street scene." This type of information can usually be automatically identified using a trained image analysis model. Style information refers to the artistic style and expression of the image, such as realistic style or minimalist style. Style information can be extracted through a style recognition network, or style feature vectors can be extracted through a convolutional neural network and combined with an existing style classifier to complete the determination.

[0065] Based on the recognition results, the original image set is divided into multiple subsets according to scene and style characteristics. The image materials in each subset share similar scene and style features. After the subset division, each image material within a subset is analyzed individually to identify and extract element information. This results in a set of element information corresponding to each subset.

[0066] When a requirement is received, the image scene and style requirements contained in the requirement are parsed. For example, if the requirement is "to generate a minimalist outdoor advertising image", the target subset matching "outdoor + minimalist style" is determined from the pre-defined subsets.

[0067] After determining the target subset, the target element information is further identified from the element information set corresponding to that subset. This target element information is used to identify the new content that needs to be replaced or generated. For example, if the requirement is to replace text, the text element is located in the element information set of the target subset; if the requirement is to replace the brand logo, the corresponding brand logo element is located.

[0068] By pre-dividing scenes and styles, the selection range of element information can be significantly narrowed, avoiding the problem of inconsistent generation results caused by inconsistent styles or mismatched scenes, thereby improving the semantic matching degree and visual quality of new image materials.

[0069] This invention achieves high-precision and controllable updates to advertising images by classifying original image materials into scenes and styles, and combining element information extraction, requirement information parsing, target element identification, replacement content generation, and prompt-driven automated generation mechanisms. Its main technical effects include: Intelligent matching and efficient generation: By semantically associating requirement information with element information sets, it can accurately identify elements that need to be replaced, enabling targeted generation and improving efficiency. Visual and semantic consistency: When replacing text, graphics, brand logos, main color tones, and scene styles, it maintains the original composition and overall visual coordination, ensuring the new image is consistent in semantics and style. Controllable quality and feedback optimization: Through verification mechanisms and semantic matching scoring, it achieves automatic adjustment and iteration, ensuring the generated results meet user needs and allowing for retrospective optimization. Multi-scene and multi-style adaptability: By classifying the original material set into scenes and styles, it supports efficient processing and customized generation of various types of advertising materials. It enables intelligent, controllable, and rapid iterative updates of advertising image materials, significantly improving work efficiency and generation quality while reducing manual intervention costs.

[0070] As an example Figure 3 This is a flowchart illustrating another image material generation method provided in an embodiment of the present invention. The method specifically includes: 1. Input processing module: Input: Select a batch of original advertising image materials (image files) to be processed.

[0071] Operation: The system receives images and performs pre-standardization processing on each image (such as resizing and formatting).

[0072] 2. AI Analysis Module: Function: Uses pre-trained CV / NLP models (such as image segmentation models U-Net / Fast R-CNN + OCR models + image and text understanding models CLIP / ViLT) to parse input images.

[0073] Processing parameters / output: text_info: Recognizes all text content (copy) embedded in the image. sticker_info: Recognizes the position coordinates (x, y, width, height) and outline mask (mask) of stickers or graphic elements in the image. logo_info: Recognizes the position coordinates (x, y, width, height) and outline mask (mask) of the advertiser's logo. dominant_colors: Extracts the dominant color information (such as RGB values) of the image. content_descr (optional): Generates text describing the main scene and style of the image. For example, input an advertisement image for a "beach resort hotel". The module outputs: text_info = "Enjoy a five-star resort experience, limited-time special offer on seaside villas!"; logo_info = "Identified the logo located in the upper right corner, coordinates (x1, y1, w, h), mask M1"; identified several sticker_infos (such as beach umbrella icons); extracted dominant_colors = [blue(RGB), white(RGB), gold(RGB)].

[0074] 3. Content Adaptation and Command Module: Function: Based on parsed information and new requirements, adapt the copy and generate instructions for replacing elements.

[0075] Processing logic: Based on preset rules or external input (such as new advertiser information), determine the elements that need to be replaced (such as specifying to replace logo_info, or to replace a specific color in dominant_colors).

[0076] Using an NLP model (such as LLM), generate semantically coherent new text (new_text) based on the original text_info or content_descr and the new requirements. Prepare parameters for the elements to be replaced (new_logo_file, new_sticker_file, new_colors). Construct prompts (prompt_for_generation) to control AI image generation, emphasizing the preservation of the original composition and style, and specifying the location of the elements to be replaced (refer to the original coordinates x, y, w, h). For example (continued), the target is to replace it with the new brand "Shanfeng Outdoor". Instruction module: determines the logo and main color to be replaced. new_text = "Explore magnificent mountains and rivers, professional outdoor gear helps you reach the summit!"; Get new_logo_file(mountain logo); new_colors = [green(RGB), brown(RGB), gray(RGB)]; Construct prompt_for_generation: "An image with [original image style description], including [original main content description], with the top right corner (x1, y1) replaced by the mountain logo, and the main color tone adjusted to green, brown, and gray".

[0077] 4. AI Recombination Generation Module: Function: Uses image generation AI (usually combined with a conditional control model) to generate new images based on parsed information, prompt_for_generation, and replacement element materials.

[0078] Algorithm / Tools: Primarily relies on an image generation diffusion model (such as Stable Diffusion v2.1 or later), combined with ControlNet technology. ControlNet can take the edges, depth, segmentation map (generated based on logo_info / sticker_info mask, etc.) of the original image as input as structural guiding constraints.

[0079] Key parameters: input_image: Original image (for reference). control_map: Guide map (e.g., Scribble, Canny Edge, Segmentation Map) generated from the original element position coordinates and masks (especially the inverse mask of the area to be replaced). prompt: prompt_for_generation constructed in step 3. inpainting_mask (optional): Precisely specifies the area to be repainted (e.g., LOGO position). strength (repaint intensity): Affects the degree to which the content of the original image is preserved. denoising_steps: Affects the generation quality and time.

[0080] Output: Generate a new base image that retains the original main creative composition and style, but replaces the specified logo, text (through subsequent compositing), stickers and / or main color scheme (generated_base_image).

[0081] 5. Precise Element Synthesis Module: Function: This function performs pixel-level composite of the generated_base_image from step 4, the actual replacement elements (new_logo_file, new_sticker_file) provided in step 3, and the precise coordinates (logo_info / sticker_info.x,y,w,h) provided in step 2. Simultaneously, it renders the new text (new_text) onto the image after either retrieving the original parsed text or intelligently arranging it according to the composition.

[0082] Tools / Workflow: Use image processing libraries (such as OpenCV, Pillow) to perform image overlay compositing. Compositing parameters must include an alpha setting to achieve a smooth blend.

[0083] Output: Generates the final, usable new ad image (final_output_image).

[0084] 6. Batch processing and storage module: Function: Receives multiple input images and corresponding processing instructions, schedules the above modules for batch processing, and stores the generated new materials in a specified location (such as a file system or CDN).

[0085] For example, (overall process): The backend system receives a list of 50 historically popular images and the rule "Replace LOGO=A, Main Color Scheme=Blue-Green". The system automatically processes in parallel: for each image, steps 1-5 are executed. The LOGO position and mask information parsed in step 2 are used as control in step 4. After generating the base image, the system automatically composites the images using the new LOGO A at the original coordinates and replaces the main color scheme, generating 50 new materials with a unified style suitable for the new advertiser A, which are then stored in the material library for deployment.

[0086] Figure 4 This is a schematic diagram of the structure of an image material generation device provided in an embodiment of the present invention, as shown below. Figure 4 As shown, the device specifically includes: The recognition module 41 is used to identify the element information contained in each original image material in the original image material set, and obtain the element information set. The determining module 42 is used to determine target element information from the element information set based on the requirement information of the new image content to be generated; The generation module 43 is used to generate prompt words based on the requirement information and the target element information, so that the target model can generate new image materials based on the prompt words, the requirement information and the target element information.

[0087] In one possible implementation, the identification module is specifically used to perform standardized preprocessing for each of the original image materials; The preprocessed original image materials are analyzed using a trained analytical model to extract element information contained in each original image material. The element information includes at least one of the following: text content and text position coordinates in the image, graphic content and graphic position coordinates and outline mask, position coordinates and outline mask of the brand logo, main color tone information, and image scene and style information.

[0088] In one possible implementation, the determining module is specifically used to receive the requirement information, which includes at least one of the following: text requirements, graphic requirements, brand logo requirements, main color requirements, and image scene and style requirements for the new image content to be generated; Perform semantic association and fit analysis between the requirement information and the set of element information; Based on the analysis results, the element information that does not match the required information is determined as the target element information, which represents the element information that needs to be replaced for the original image material.

[0089] In one possible implementation, the generation module is specifically used to generate new text content based on the text requirements in the requirement information, and the new text content is used to replace the text content in the target element information. And / or, generate new graphic content based on the graphic requirements in the requirement information, the new graphic content being used to replace the graphic content in the target element information; And / or, generate a new brand identity based on the brand identity requirements in the requirement information, the new brand identity being used to replace the brand identity in the target element information; And / or, generate new main color information based on the main color requirements in the requirement information, and use the new main color information to replace the main color information in the target element information; And / or, generate new image scene and style information based on the image scene and style requirements in the requirement information, and use the new image scene and style information to replace the image scene and style information in the target element information; The prompt word is generated based on the new text content, the new graphic content, the new brand logo, the new main color scheme, and the new image scene and style information.

[0090] In one possible implementation, the generation module is specifically used to replace the target element information in the original image material according to the prompt word, and then verify and adjust the image material after the replacement operation according to the requirement information to obtain the new image material.

[0091] In one possible implementation, the processing module 44 is used to verify at least one of the following for the image material after the replacement operation based on the requirement information: element information integrity, scene style consistency, brand specification compliance, and visual coordination. If any verification fails, the failed content will be adjusted. And / or, calculate the semantic similarity score between the replaced image material and the requirement information through a semantic matching model; If the score is less than a preset threshold, the step of identifying the element information contained in each original image in the original image material set is re-executed.

[0092] In one possible implementation, the processing module is further configured to divide the original material set according to the image scene and style information of each original image material to obtain multiple subsets with different image scenes and styles; The identification module is also used to identify the original image material element information contained in each subset, and obtain the element information set corresponding to each subset; The determining module is further configured to, upon receiving demand information, determine a target subset based on the image scene and style requirements in the demand information, and then determine target element information from the element information set corresponding to the target subset.

[0093] The image material generation device provided in this embodiment can be as follows: Figure 4 The apparatus shown can perform, for example Figure 1-2 This involves all the steps in the method for generating image materials in Chinese, thereby achieving... Figure 1-2 For details on the technical effects of the image material generation method shown, please refer to [link / reference]. Figure 1-2 The relevant descriptions are presented concisely and will not be elaborated upon here.

[0094] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 5 The computer device 500 shown includes at least one processor 501, a memory 502, at least one network interface 504, and other user interfaces 503. The various components in the computer device 500 are coupled together via a bus system 505. It is understood that the bus system 505 is used to implement communication between these components. In addition to a data bus, the bus system 505 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 5 The general designated all buses as Bus System 505.

[0095] The user interface 503 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).

[0096] It is understood that the memory 502 in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate Synchronous DRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 502 described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0097] In some implementations, memory 502 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 5021 and application program 5022.

[0098] The operating system 5021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 5022 includes various applications, such as a media player and a browser, used to implement various application functions. The program implementing the method of this embodiment can be included in the application program 5022.

[0099] In this embodiment of the invention, by calling the program or instructions stored in memory 502, specifically the program or instructions stored in application program 5022, processor 501 executes the method steps provided in each method embodiment, including, for example: For each original image in the original image material set, identify the element information it contains to obtain the element information set; Based on the requirements for the content of the new image to be generated, the target element information is determined from the set of element information; Based on the requirement information and the target element information, prompt words are generated so that the target model can generate new image materials based on the prompt words, the requirement information, and the target element information.

[0100] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 501 or by instructions in the form of software. The processor 501 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 502. Processor 501 reads the information in memory 502 and, in conjunction with its hardware, completes the steps of the above method.

[0101] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.

[0102] For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0103] The computer device provided in this embodiment may be as follows: Figure 5 The device shown can perform, for example Figure 1-2 This involves all the steps in the method for generating image materials in Chinese, thereby achieving... Figure 1-2 For details on the technical effects of the image material generation method shown, please refer to [link / reference]. Figure 1-2 The relevant descriptions are presented concisely and will not be elaborated upon here.

[0104] This invention also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; it may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; and it may also include combinations of the above types of memory.

[0105] One or more programs in the storage medium can be executed by one or more processors to implement the image material generation method described above that is executed on the device side.

[0106] The processor is used to execute the image material generation program stored in the memory to implement the following steps of the image material generation method executed on the device side: For each original image in the original image material set, identify the element information it contains to obtain the element information set; Based on the requirements for the content of the new image to be generated, the target element information is determined from the set of element information; Based on the requirement information and the target element information, prompt words are generated so that the target model can generate new image materials based on the prompt words, the requirement information, and the target element information.

[0107] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0108] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0109] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for generating image materials, characterized in that, include: For each original image in the original image material set, identify the element information it contains to obtain the element information set; Based on the requirements for the content of the new image to be generated, the target element information is determined from the set of element information; Based on the requirement information and the target element information, prompt words are generated so that the target model can generate new image materials based on the prompt words, the requirement information, and the target element information.

2. The method according to claim 1, characterized in that, For each original image in the original image material set, the element information it contains is identified, including: Standardized preprocessing is performed on each of the original image materials; The preprocessed original image materials are analyzed using a trained analytical model to extract element information contained in each original image material. The element information includes at least one of the following: text content and text position coordinates in the image, graphic content and graphic position coordinates and outline mask, position coordinates and outline mask of the brand logo, main color tone information, and image scene and style information.

3. The method according to claim 1, characterized in that, The step of determining the target element information from the element information set based on the requirement information of the new image content to be generated includes: Receive the requirement information, which includes at least one of the following: text requirements, graphic requirements, brand logo requirements, main color requirements, and image scene and style requirements for the new image content to be generated; Perform semantic association and fit analysis between the requirement information and the set of element information; Based on the analysis results, the element information that does not match the required information is determined as the target element information, which represents the element information that needs to be replaced for the original image material.

4. The method according to claim 1, characterized in that, The step of generating prompt words based on the demand information and the target element information includes: New text content is generated based on the text requirements in the requirement information, and the new text content is used to replace the text content in the target element information; And / or, generate new graphic content based on the graphic requirements in the requirement information, the new graphic content being used to replace the graphic content in the target element information; And / or, generate a new brand identity based on the brand identity requirements in the requirement information, the new brand identity being used to replace the brand identity in the target element information; And / or, generate new main color information based on the main color requirements in the requirement information, and use the new main color information to replace the main color information in the target element information; And / or, generate new image scene and style information based on the image scene and style requirements in the requirement information, and use the new image scene and style information to replace the image scene and style information in the target element information; The prompt word is generated based on the new text content, the new graphic content, the new brand logo, the new main color scheme, and the new image scene and style information.

5. The method according to claim 1, characterized in that, The target model generates new image materials based on the prompt words, the requirement information, and the target element information, including: After the target model replaces the target element information in the original image material according to the prompt words, it verifies and adjusts the image material after the replacement operation according to the requirement information to obtain the new image material.

6. The method according to claim 5, characterized in that, The step of verifying and adjusting the image materials after the replacement operation based on the required information includes: Based on the aforementioned requirements, verify at least one of the following aspects for the replaced image materials: completeness of element information, consistency of scene style, compliance with brand specifications, and visual harmony. If any verification fails, the failed content will be adjusted. And / or, calculate the semantic similarity score between the replaced image material and the requirement information using a semantic matching model; If the score is less than a preset threshold, the step of identifying the element information contained in each original image in the original image material set is re-executed.

7. The method according to claim 1, characterized in that, The method further includes: The original image set is divided into multiple subsets with different image scenes and styles based on the image scene and style information of each original image material; For each subset, the original image material element information contained therein is identified to obtain the element information set corresponding to each subset; When a requirement is received, the target subset is determined based on the image scene and style requirements in the requirement information, and then the target element information is determined from the element information set corresponding to the target subset.

8. An image material generation device, characterized in that, include: The recognition module is used to identify the element information contained in each original image in the original image material set, and obtain the element information set. The determination module is used to determine target element information from the element information set based on the requirement information of the new image content to be generated; The generation module is used to generate prompt words based on the requirement information and the target element information, so that the target model can generate new image materials based on the prompt words, the requirement information and the target element information.

9. A computer device, characterized in that, include: A processor and a memory, the processor being configured to execute an image material generation program stored in the memory to implement the image material generation method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the image material generation method according to any one of claims 1 to 7.