Commercial art illustration similar style picture generation method and device, equipment and medium
By combining the Florence prompt word inference model, the Flux text image model, the SDXL model, and the IP-Adapter model, the problems of low efficiency and inaccurate style control in commercial art illustration generation are solved, and efficient and accurate commercial art illustration generation is achieved, meeting the needs of rapid iteration and creative diversity in commercial art creation.
Patent Information
- Application Number
- CN202510655262.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies make it difficult to efficiently generate commercial art illustrations that are similar in style to existing illustrations but with different elements and scenes, and it is also difficult to accurately control style migration, resulting in low output efficiency of commercial art works and limited creative expansion.
Using a combination of the Florence prompt word inference model, the Flux text image model, the SDXL model, the IP-Adapter model, and the reference image model, commercial art illustrations are generated through an automated process, including center cropping, element scene generation, and style transfer, to ensure the accurate inheritance of style features.
It achieves efficient generation of commercial art illustrations, shortens the processing time of a single image, achieves style similarity of over 90%, and provides strong flexibility in element scenes, meeting the needs of rapid iteration and visual image consistency of commercial projects.
Smart Images

Figure CN120707666A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and in particular to a method, device, equipment and medium for generating images in a style similar to commercial art illustrations. Background Art
[0002] In the field of commercial art illustration creation, it's often necessary to generate images that resemble existing illustrations in style, but with different elements and scenes. Traditionally, creators rely on experience to manually conceive elements and adjust stylistic details, which is extremely inefficient and makes it difficult to ensure accurate inheritance of the style.
[0003] Some existing AI-based image generation technologies either find it difficult to maintain the similarity of the original style when the elements and scenes change significantly, or lack precise control over style migration. They are unable to meet the demand of commercial art creation for high-quality, diverse images of similar styles, greatly limiting the output efficiency and creative expansion of commercial art works. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for generating pictures of similar styles to commercial art illustrations, thereby improving generation efficiency and accuracy.
[0005] In a first aspect, the present invention provides a method for generating pictures of similar style to commercial art illustrations, comprising the following steps:
[0006] Step 1: Input the set art illustration into the Florence prompt word reverse inference model to obtain the corresponding prompt word;
[0007] Step 2: Crop the art illustration to its center. First, determine the center of the art illustration and retain only the area within the set range to obtain a cropped image.
[0008] Step 3: input the prompt word into the flux text graph model, and the flux text graph model generates a primary fission picture containing the corresponding element scene;
[0009] Step 4: Input the cropped image into the reference image model and the IP-Adapter model respectively, input the primary fission image and prompt words into the SDXL model, and use the reference image model and the IP-Adapter model to guide the SDXL model to generate the required image.
[0010] In a second aspect, the present invention provides a device for generating pictures of similar styles to commercial art illustrations, comprising:
[0011] Generate prompt word module, input the set art illustration into the Florence prompt word reverse inference model to obtain the corresponding prompt word;
[0012] The cropping module performs center cropping on the art illustration. First, the center of the art illustration is determined, and only the area within the set range of the art illustration is retained to obtain a cropped image.
[0013] A text graph module inputs the prompt word into a flux text graph model, and the flux text graph model generates a primary fission picture containing the corresponding element scene;
[0014] Generate an image module, input the cropped image into the reference image model and IP-Adapter model respectively, input the primary fission image and prompt words into the SDXL model, and use the reference image model and IP-Adapter model to guide the SDXL model to generate the required image.
[0015] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0016] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in the first aspect when the program is executed by a processor.
[0017] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0018] 1. Efficiency: Compared to traditional manual creation methods, this invention significantly reduces the processing time for a single image through an automated model process, from inputting a reference image to generating a final image with a similar style. For example, a complex commercial illustration, which traditionally would take hours or even days to create manually, can be completed in minutes. This significantly improves the efficiency of commercial art production and meets the needs of rapid iteration for commercial projects.
[0019] 2. Highly Accurate Stylistic Similarity: Through the precise extraction of stylistic features by the reference image model, combined with the SDXL model and the IP-Adapter model, the generated image closely matches the original commercial art illustration in terms of color tonality, brushstroke texture, compositional rhythm, and other stylistic elements. Professional evaluation and objective style similarity algorithm testing have shown a style similarity rate exceeding 90%, effectively inheriting the artistic style of the original illustration and ensuring the consistency of the commercial brand's visual image.
[0020] 3. Highly flexible elements and scenes: The Flux text-based image model, combined with the rich cue words generated by the Florence cue word inference model, enables diverse transformations of elements and scenes. Whether switching from a fantasy magical world to a modern urban jungle, or transforming classical characters into sci-fi mecha, all can be flexibly achieved while maintaining a similar style. This provides broad creative space for commercial art creation and meets the diverse illustration needs of different commercial scenarios.
[0021] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0023] Figure 1 This is a flowchart of the method in Example 1 of the present invention;
[0024] Figure 2 This is a schematic diagram of the structure of the device in Example 2 of the present invention. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of this application have the following general ideas:
[0026] The Florence Cue Word Inference Model is a versatile visual model developed by the Microsoft Azure AI team. It is primarily used for various computer vision tasks, including image description, object detection, visual localization, and image segmentation. The model's core features include versatility, a unified representation method, sequence-to-sequence learning, large-scale dataset training, and multi-task learning. It is often used in image-to-text work to accurately describe the content of images.
[0027] Flux Text-to-Graph Model: A cutting-edge text-to-image model launched by BlackForest Labs. The FLUX model inherits multimodal processing capabilities and is expanded to 12 billion parameters. With its excellent complex scene generation capabilities and precise prompt word following features, it performs well in high-quality text-to-graph models.
[0028] The SDXL model is a text-to-image generation model based on the principles of diffusion models, boasting more robust model parameters and a complex architecture. It can generate high-quality images in a high-dimensional image space based on input information. In this paper, it is used to generate the image subject and initially establish the image style, capable of handling image generation tasks under complex semantic cues.
[0029] The IP-Adapter model assists the SDXL model in adapting and adjusting image styles. It utilizes a style transfer algorithm and an adaptive parameter adjustment mechanism. Based on specific style rules, it optimizes the color channel distribution, brushstroke and texture simulation parameters, and composition scale coefficients of the SDXL model-generated image, achieving precise style matching.
[0030] Reference Image Model: Based on a convolutional neural network (CNN) architecture, this model deeply mines the style characteristics of the reference image through multiple feature extraction layers. It extracts color hue distribution, brushstroke texture patterns (such as the probability distribution of brushstroke directions and the variation of line widths), and the visual center of gravity of the composition (such as the golden ratio of element distribution and symmetrical or asymmetrical structural features). It generates a style parameter vector, which is used to guide the SDXL and IP-Adapter models for accurate style transfer.
[0031] The main implementation is as follows:
[0032] The Florence prompt word inference model comprehensively analyzes the composition of elements in the image (such as character modeling, scene layout, and prop style), color matching, and brushstroke style (such as line thickness and density). Based on the large amount of image-text data it has been trained on, it generates detailed and accurate prompt words that cover both the image content description and style keywords, such as "retro steampunk style, a futuristic city surrounded by mechanical gears, and an explorer in a leather trench coat."
[0033] 2. Perform a center crop on the original image. First, determine the center of the image as (w / 2, h / 2). Then, retain 70% of the inner region of the image and remove 30%. This method preserves the style and texture characteristics of the main area of the image while avoiding interference from surrounding features, allowing the model to focus on preserving the style of the main area.
[0034] 3. Flux Text-Based Image Model Generates Reasonable Images: The cue words output by the large model are fed into the flux text-based image model. Based on the semantic information of the cue words, the model searches for matching image feature combinations within its learned image generation space and generates a primary fission image containing the corresponding elements and scenes. During this process, the model optimizes image details through adversarial learning, ensuring that the generated images meet certain standards for elemental completeness and scene plausibility. However, at this stage, style similarity still needs to be improved.
[0035] 4. The SDXL model works with the IP-Adapter model and the reference image model to achieve style transfer. The reference image model first extracts the style features of the cropped image, such as the color hue distribution range, the texture pattern of the brushstrokes, and the visual center of gravity of the composition, to generate a style parameter vector. The SDXL model uses the output image of the flux image model as a basis, and combines the style parameter vector with the adaptive adjustment of the IP-Adapter model to process the image's color mapping, brushstroke simulation, and texture overlay. For example, if the reference image is in the style of thick oil painting, the SDXL model overlays a texture that simulates the accumulation of oil paint on the generated image. At the same time, the IP-Adapter model adjusts the color saturation and contrast to levels similar to the reference image. The final output is a commercial art illustration with a large variation in element scenes and a highly similar style.
[0036] Example 1
[0037] like Figure 1 As shown, this embodiment provides a method for generating pictures of similar style to commercial art illustrations, including the following steps:
[0038] Step 1: Input the set art illustration into the Florence prompt word reverse inference model to obtain the corresponding prompt word;
[0039] Step 2: Crop the art illustration to its center. First, determine the center of the art illustration and retain only the area within the set range to obtain a cropped image.
[0040] Step 3: input the prompt word into the flux text graph model, and the flux text graph model generates a primary fission picture containing the corresponding element scene;
[0041] Step 4: Input the cropped image into the reference image model and the IP-Adapter model respectively, input the primary fission image and prompt words into the SDXL model, and use the reference image model and the IP-Adapter model to guide the SDXL model to generate the required image.
[0042] In this embodiment, preferably, step 2 is specifically as follows: center cropping the art illustration, first determining the center of the art illustration, that is, the position of half the length and width of the art illustration is the center of the art illustration, retaining 70% of the area inside the art illustration, and cropping and removing the outer 30% of the area to obtain a cropped image.
[0043] In this embodiment, preferably, the IP-Adapter model selects a weight value of 0.75, a weight type of style and composition, a merging type of norm_average, and an embedding layer parameter of k+mean(V)w / Cpenalty.
[0044] Based on the same inventive concept, this application also provides a device corresponding to the method in Example 1, see Example 2 for details.
[0045] Example 2
[0046] like Figure 2 As shown, in this embodiment, a device for generating pictures of similar styles to commercial art illustrations is provided, comprising:
[0047] Generate prompt word module, input the set art illustration into the Florence prompt word reverse inference model to obtain the corresponding prompt word;
[0048] The cropping module performs center cropping on the art illustration. First, the center of the art illustration is determined, and only the area within the set range of the art illustration is retained to obtain a cropped image.
[0049] A text graph module inputs the prompt word into a flux text graph model, and the flux text graph model generates a primary fission picture containing the corresponding element scene;
[0050] Generate an image module, input the cropped image into the reference image model and IP-Adapter model respectively, input the primary fission image and prompt words into the SDXL model, and use the reference image model and IP-Adapter model to guide the SDXL model to generate the required image.
[0051] In this embodiment, preferably, step 2 is specifically as follows: center cropping the art illustration, first determining the center of the art illustration, that is, the position of half the length and width of the art illustration is the center of the art illustration, retaining 70% of the area inside the art illustration, and cropping and removing the outer 30% of the area to obtain a cropped image.
[0052] In this embodiment, preferably, the IP-Adapter model selects a weight value of 0.75, a weight type of style and composition, a merging type of norm_average, and an embedding layer parameter of k+mean(V)w / Cpenalty.
[0053] Since the device described in the second embodiment of the present invention is used to implement the method of the first embodiment of the present invention, those skilled in the art will be able to understand the specific structure and variations of the device based on the method described in the first embodiment of the present invention, and therefore will not be described in detail here. All devices used in the method of the first embodiment of the present invention fall within the scope of protection of the present invention.
[0054] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment, see the third embodiment for details.
[0055] Example 3
[0056] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, any implementation method in the first embodiment can be implemented.
[0057] Since the electronic device described in this embodiment is the device used to implement the method in Example 1 of this application, based on the method described in Example 1 of this application, those skilled in the art will be able to understand the specific implementation of the electronic device of this embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of this application will not be described in detail here. As long as the device used by those skilled in the art to implement the method in the embodiment of this application falls within the scope of protection to be provided by this application.
[0058] Based on the same inventive concept, this application provides a storage medium corresponding to Example 1, see Example 4 for details.
[0059] Example 4
[0060] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, any implementation method in the first embodiment can be implemented.
[0061] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0062] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0063] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0064] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0065] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for generating pictures of similar style to commercial art illustrations, characterized by: The steps include: Step 1: Input the set art illustration into the Florence prompt word reverse inference model to obtain the corresponding prompt word; Step 2: Crop the art illustration to its center. First, determine the center of the art illustration and retain only the area within the set range to obtain a cropped image. Step 3: input the prompt word into the flux text graph model, and the flux text graph model generates a primary fission picture containing the corresponding element scene; Step 4: Input the cropped image into the reference image model and the IP-Adapter model respectively, input the primary fission image and prompt words into the SDXL model, and use the reference image model and the IP-Adapter model to guide the SDXL model to generate the required image.
2. The method for generating pictures of similar style to commercial art illustrations according to claim 1, characterized in that: The step 2 is specifically as follows: center cropping of the art illustration, first determining the center of the art illustration, that is, the position of half the length and width of the art illustration is the center of the art illustration, retaining 70% of the area inside the art illustration, and cropping and removing the outer 30% area to obtain a cropped image.
3. The method for generating pictures of similar style to commercial art illustrations according to claim 1, characterized in that: The IP-Adapter model selects a weight value of 0.75, a weight type of style and composition, a merging type of norm_average, and an embedding layer parameter of k+mean(V)w / C penalty.
4. A device for generating pictures of similar style to commercial art illustrations, characterized by: include: Generate prompt word module, input the set art illustration into the Florence prompt word reverse inference model to obtain the corresponding prompt word; The cropping module performs center cropping on the art illustration. First, the center of the art illustration is determined, and only the area within the set range of the art illustration is retained to obtain a cropped image. A text graph module inputs the prompt word into a flux text graph model, and the flux text graph model generates a primary fission picture containing the corresponding element scene; Generate an image module, input the cropped image into the reference image model and IP-Adapter model respectively, input the primary fission image and prompt words into the SDXL model, and use the reference image model and IP-Adapter model to guide the SDXL model to generate the required image.
5. The device for generating pictures of similar style to commercial art illustrations according to claim 4, characterized in that: The step 2 is specifically as follows: center cropping of the art illustration, first determining the center of the art illustration, that is, the position of half the length and width of the art illustration is the center of the art illustration, retaining 70% of the area inside the art illustration, and cropping and removing the outer 30% area to obtain a cropped image.
6. The device for generating pictures of similar style to commercial art illustrations according to claim 4, characterized in that: The IP-Adapter model selects a weight value of 0.75, a weight type of style and composition, a merging type of norm_average, and an embedding layer parameter of k+mean(V)w / C penalty.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 3 is implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.