Trademark image generation method and device based on large model, equipment and medium

By constructing prompt templates through feature extraction and clustering methods, and using large language models for feature annotation and model fine-tuning, the problem of generating images in special fields in existing technologies is solved, and efficient and personalized trademark image generation is achieved.

CN120672906APending Publication Date: 2025-09-19BEIJING AUGUST MELON TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510746553.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies have difficulty generating images that meet the image requirements of special fields, and existing large language models cannot effectively meet professional design needs when generating images.

Method used

The feature extraction model is used to extract the graphic layout, color combination and graphic elements of the trademark image, and the clustering method is used to cluster the images. A prompt template is constructed and input into the large language model for feature annotation. The generative model is fine-tuned to improve the generation capability. Finally, an image that meets the requirements is generated based on user input.

Benefits of technology

It achieves efficient generation of trademark images that meet users' personalized needs, reduces dependence on professional designers, and significantly shortens design cycles and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672906A_ABST
    Figure CN120672906A_ABST
Patent Text Reader

Abstract

The invention provides a trademark image generation method and device based on a large model, equipment and a medium. The method comprises the steps that image features are extracted through a feature extraction model based on trademark image data; performing image clustering by using a clustering method based on the image features to obtain a plurality of image subsets; constructing a prompt template, and setting feature category extraction logic; inputting the prompt template and the plurality of image subsets into a large language model, performing feature extraction on the images in the plurality of image subsets based on feature category extraction logic, and taking the extracted features as image tags to obtain an image set with tags; performing fine tuning training on the generative model by using the image with the label; and inputting the generation requirement information and / or the example pattern into the fine-tuned generative model to generate a preset number of first trademark generation images. According to the method, a user does not need to have a professional design background, and satisfactory image design can be obtained by providing simple input, so that the dependence on profession is reduced, and the design period and cost of the image are shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a method, device, equipment and medium for generating a trademark image based on a large model. Background Art

[0002] In some special fields, users need some images as signs for brand promotion, identification and differentiation, such as trademarks, appearance, works, etc. However, users themselves do not have the ability to design images, so they need to ask professional designers to come up with creative ideas and hand-draw according to customer needs. This method not only consumes a lot of time and manpower costs, but also the design results may not meet customer expectations. It is difficult to modify them later and it is difficult to meet the rapid changes in the market in a timely manner. Moreover, the image design in these special fields has creative requirements. It is meaningless to generate images randomly or under general conditions. For example, trademark design needs to be distinctive, so it needs to be significantly different from existing images or trademark images. For appearance design (such as packaging design), it also needs to be different from the appearance image of the market product to avoid losses caused by infringement.

[0003] With the rapid development of artificial intelligence and deep learning technologies, large-scale model-based trademark image generation technology has achieved remarkable results in the field of computer vision. Existing text-to-image and image-to-image models, such as Stable Diffusion, can generate corresponding images based on input text descriptions or reference images. However, the random or general conditional image generation of existing large language models cannot meet the image requirements of specific fields.

[0004] In view of this, there is an urgent need for an image generation method based on large models that can automatically generate images that meet the image generation requirements of special fields. It can efficiently generate images that meet design specifications and domain requirements based on user descriptions and reference examples. Summary of the Invention

[0005] In order to overcome the problems existing in the related art, the present disclosure provides a trademark image generation method, device, equipment and medium based on a large model to solve the technical problem that the generated images in the related art cannot meet the image requirements of special fields.

[0006] One or more embodiments of this specification provide a method for generating a trademark image based on a large model, comprising the steps of: Extracting image features based on the acquired trademark image data using a feature extraction model, the image features including graphic layout, color combination, and / or graphic elements; Using clustering methods based on image features to cluster images and obtain multiple image subsets; Constructing a prompt template, which includes prompt words for setting detailed task description information. The detailed task description information includes the execution process of the large language model feature extraction task, feature category extraction logic, and definitions of each feature label and sample images representing various graphic layout features. The feature categories include design style, graphic layout, color combination, graphic elements, and / or text elements. Input the prompt template and multiple image subsets into the large language model, extract features from the images in the multiple image subsets, and use the extracted features as labels to annotate the images; Fine-tune the image generation model using labeled images; The image generation requirement information and the sample pattern are input into the fine-tuned image generation model to generate a preset number of first trademark generation images.

[0007] One or more embodiments of this specification provide a large model-based trademark image generation device, including: A feature extraction module is used to extract image features based on the acquired trademark image data through a feature extraction model. The image features include graphic layout, color combination and graphic elements; A clustering module, used to cluster images using a clustering method based on image features to obtain multiple image subsets; A prompt template construction module is used to construct a prompt template. The prompt template includes prompt words for setting detailed task description information. The detailed task description information includes the execution process of the large language model feature extraction task, the feature category extraction logic, the definition of each feature label, and sample images of each representation class of graphic layout features, where the feature categories include design style, graphic layout, color combination, graphic elements and / or text elements; The image labeling module is used to input the prompt template and multiple image subsets into the large language model, extract features from the images in the multiple image subsets, and use the extracted features as labels to label the images; A model fine-tuning module, used to fine-tune the image generation model using labeled images; and The image generation module is used to input the image generation requirement information and the sample pattern into the fine-tuned image generation model to generate a preset number of first trademark generation images.

[0008] One or more embodiments of this specification provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for generating a trademark image based on a large model as described above is implemented.

[0009] One or more embodiments of this specification provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method for generating a trademark image based on a large model as described above is implemented.

[0010] The present disclosure provides a trademark image generation method, device, equipment and medium based on a large model. The advantage is that, considering that the generation of trademark images requires professional designers to creatively conceive and draw according to user needs, the present disclosure combines the large model to first construct various types of image data, and through prompt engineering, enables the model to complete the task of image feature labeling, and uses the feature-labeled images to fine-tune the instructions of the image generation model to improve the model's feature understanding ability and image generation ability; then, based on the generation requirement information and sample images input by the user, an image that better meets the user's personalized needs is generated. The user does not need to have a professional design or technical background, and only needs to provide simple input to obtain a satisfactory image design, which reduces dependence on professional designers and significantly shortens the image design cycle and cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 A flowchart of a method for generating a trademark image based on a large model provided in one or more embodiments of this specification; Figure 2 An example diagram of a trademark image input to a large language model provided in one or more embodiments of this specification; Figure 3 A block diagram of a trademark image generation device based on a large model provided for one or more embodiments of this specification; and Figure 4 A schematic diagram of the structure of a computer device provided in one or more embodiments of this specification. DETAILED DESCRIPTION

[0013] In order to help those skilled in the art better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0014] The present invention will be described in detail below with reference to specific implementation methods and the accompanying drawings.

[0015] Method Example According to an embodiment of the present invention, a method for generating a trademark image based on a large model is provided. Figure 1 FIG. 1 is a flow chart of a method for generating a trademark image based on a large model according to an embodiment of the present invention. The method for generating a trademark image based on a large model according to an embodiment of the present invention includes the following steps: Step S1: Based on the acquired trademark image data, image features are extracted using a feature extraction model. The image features include graphic layout, color combination and / or graphic elements.

[0016] Step S2: clustering the images using a clustering method based on the image features to obtain multiple image subsets.

[0017] Step S3: construct a prompt template and set feature category extraction logic, wherein the feature categories include design style, graphic layout, color combination, graphic elements and / or text elements.

[0018] In step S4, the prompt template and multiple image subsets are input into the large language model, features are extracted from the images in the multiple image subsets, and the extracted features are used as image labels to obtain a labeled image set.

[0019] In step S5, the generative model is fine-tuned using the labeled image set to improve the model's understanding and generation capabilities.

[0020] Step S6: inputting the generation requirement information and / or the example pattern into the fine-tuned generative model to generate a preset number of first trademark generation images.

[0021] The trademark image generation method based on a large model provided in this embodiment takes into account that the generation of trademark images requires professional designers to creatively conceive and draw according to user needs. The method of this embodiment combines the large model, first constructs various types of image data, and uses prompt engineering to enable the model to complete the task of image feature labeling, and uses feature-labeled images to fine-tune the instructions of the generative model to improve the model's feature understanding ability and image generation ability; then, based on the generation requirement information and sample images input by the user, an image that better meets the user's personalized needs is generated. The user does not need to have a professional design or technical background, and only needs to provide simple input to obtain a satisfactory image design, which reduces dependence on professional designers and significantly shortens the image design cycle and cost.

[0022] In this embodiment, in order to reduce the infringement risk of the generated images, the generated images are further deduplicated by combining the search-enhanced generation technology with existing images. The specific steps are as follows: Step S7: using the search enhancement generation technology, the first trademark generated image is compared with images in the image library for similarity, and the first trademark generated image with a similarity lower than a first preset value is retained.

[0023] In order to construct an image dataset that meets the generation needs of specific fields, it is necessary to conduct in-depth exploration of existing image data and explore the core elements contained in trademark images to accurately control the effect of generated images. Therefore, the image features that are more concerned in the feature extraction process of this embodiment are graphic layout, color combination and / or graphic elements.

[0024] In this embodiment, in order to improve the quality of image data, the extraction of image features by the feature extraction model in step S1 specifically includes the following steps: Step S11, rough image screening: extracting basic features of the image through a feature extraction model, and filtering out images with similarity higher than a second preset value based on the basic features through similarity calculation to obtain an initial image set.

[0025] In this embodiment, data screening is performed in a library of tens of millions of images, and the model for extracting image features is required to strike a balance between accuracy and efficiency. The feature extraction model in this embodiment selects the ResNet18 (Residual Network18) model with excellent accuracy and fast inference speed to extract image features. The cosine similarity between image features is calculated, and images with a similarity greater than 85%-95% are deduplicated.

[0026] Step S12, fine feature extraction: based on the initial image set, extract the image elements, text elements, graphic layout and / or color combination features of each image.

[0027] In this embodiment, the extraction of image elements, text elements and graphic layout features is specifically performed through the following steps: Step S121: Identify the category and position of the graphic elements in each image through the target detection model, and / or identify the content and position of the text elements in each image through the text detection model. When there are graphic elements and text elements, the relative position information is determined by the position coordinates of the graphic elements and the text elements, thereby determining the graphic and text layout.

[0028] The layout features of this embodiment are mainly used to identify the graphic elements, text elements and the layout of graphic elements and text elements in the image. For example, in trademark image design, specific graphic and text layout types include pure text, pure graphics, upper image and lower text, upper text and lower image, left image and right text, left text and right image, circular badge, graphics surrounding text, graphics semi-surrounding text, text surrounding graphics, text semi-surrounding graphics, etc. In this embodiment, the target retrieval model is used to identify and locate the category and position of graphic elements in the image by combining the target detection model and the OCR (Optical Character Recognition) model. The OCR model is used to identify the text content and position in the image. The relative position relationship between the graphic element position and the text position is determined by the position of the graphic element and the text to give the layout category to which they belong. Each image is allowed to have only one layout category.

[0029] In this embodiment, the purpose of extracting color information from an image is to identify the main color components in the image. This embodiment proposes an innovative and efficient color extraction method. Color extraction is not simply distinguishing, accumulating, and sorting the color values ​​of all pixels in the image. Two points should be noted: First, the importance of colors in different areas of an image varies. During feature recognition, attention should be paid to the colors of salient objects in the image, while the background color may be more intended to highlight salient objects. Therefore, attention is focused on closed irregular areas in the image. Second, the color values ​​of each pixel are discrete. Even if the colors appear to be the same, they may actually be different. Therefore, it is necessary to classify colors by setting a series of color value ranges. To implement the frequency of occurrence of color values ​​in the irregular graphic area in step S122, the RGB color values ​​of the graphic area are first mapped to predefined corresponding color values ​​of red, orange, yellow, green, cyan, blue, purple, white, and black. The Euclidean distance between the current pixel RGB value and the RGB value of each target color is calculated, and the color with the smallest distance is selected as the mapping result. The frequency of occurrence of the mapped colors is then counted to obtain the frequency of occurrence of each color category. The predefined RGB values ​​corresponding to red, orange, yellow, green, cyan, blue, purple, white, and black are: red (255, 0, 0), orange (255, 165, 0), yellow (255, 255, 0). , Green(0, 255, 0), Cyan(0, 255, 255), Blue(0, 0, 255), Purple(128, 0, 128), White(255, 255, 255), Black(0, 0,0).

[0030] In this embodiment, the image color combination features are extracted specifically through the following steps: In step S122, the image is edge detected and noise filtered using the OpenCV (Open Computer Vision Library) edge detection algorithm. Closed irregular shapes are then extracted, and the irregular shapes are treated as salient objects in the image. The frequency of occurrence of the color values ​​in each irregular shape region is calculated, and the frequency of occurrence of the color values ​​is statistically analyzed based on preset color value ranges to obtain the frequency of occurrence of each color category. The color categories with the highest frequency are selected as the color combination features of the image. For example, the color categories are sorted from high to low frequency, and the color categories with the highest frequency are selected as the color combination features of the image. For example, when generating a trademark image, 3-4 color categories can be selected as the color combination features of the image, as trademark designs are not complex and have relatively few elements. For a design, the top 7 color categories representing salient features are selected as the color combination features of the image. This also facilitates faster classification and computer processing in subsequent steps, saving computing resources.

[0031] Compared with using a model to extract color, this embodiment has a shorter processing time. This is because the number of images before clustering is still very large, and using a model to process color extraction requires a long inference time. This embodiment focuses on salient targets and extracts colors, greatly improving image processing time.

[0032] In step S13 , the acquired basic features and fine features of the image are encoded and combined to obtain combined features, and then clustered using a clustering method based on the graphic layout features to obtain multiple image subsets.

[0033] In another embodiment, image clustering is performed based on the graphic layout type. For example, the 11 graphic layout types listed in this embodiment can have 50-100 cluster centers for each cluster according to needs, and 50-100 images are extracted from each center. The sample data obtained in this way is highly relevant to the task, and the sample data covers various situations in which images appear, so as to ensure the versatility and applicability of the generative model.

[0034] To guide the large language model to the desired output, this embodiment requires setting a corresponding prompt template to guide the large language model in task execution. The prompt template includes prompt words that set detailed task description information. The detailed task description information includes the execution process of the large language model feature extraction task, feature category extraction logic, definition of each feature label, and sample images representing each type of graphic layout feature.

[0035] This embodiment uses a large language model to extract image features, mainly because the number of surrenders in the dataset is relatively large. Compared with manual labeling to determine and annotate the features of each image, it reduces the workload and difficulty of labeling.

[0036] In this embodiment, since the large model is not sensitive to the layout of the specific positional relationship between graphic elements and text elements, in addition to using the prompt template, sample images that characterize various types of graphic and text layout features are also given to the large model for learning. The sample images are images with labels such as graphic and text layout, design style, etc., so that the large model can better distinguish different layouts and give more reasonable reasoning results. In this embodiment, the sample images can be selected from images with graphic and text layout characteristics for annotation. Since the graphic and text layout result determined by step S121 may not be accurate, for example, the graphic and text structure in the original image may be that the image encloses the text, and the text encloses the image, it may be that the central image occupies a smaller area weight of the entire image, resulting in the graphic and text layout result determined in step S121 being the graphic surrounding the text. Therefore, the sample images with graphic and text layout characteristics are given to the large model for learning to make the large model have stronger understanding and generation capabilities. In one embodiment, the prompt template may include the following content: 1. Define the model's identity: You are an intelligent assistant that is given the following feature information of an image and can make inferences: 1. Design Elements Image elements and text elements: Which image elements and text elements are in the image? The image elements and text elements may be empty, which can be represented by "none".

[0037] 2. Graphic and text layout: The positional relationship between image elements and text elements in an image, such as pure text, pure graphics, image above and text below, text above and image below, image on the left and text on the right, text on the left and image on the right, circular badge, graphics surrounding text, graphics semi-surrounding text, text surrounding graphics, text semi-surrounding graphics, etc.

[0038] 3. Color combination: What colors does the image mainly consist of (such as: red, yellow, green, blue, purple, black).

[0039] 4. Design style: What is the overall design style of the image, such as traditional Chinese style, simple drawing, fresh and natural, cartoon, technological sense, etc.

[0040] --------- Example 1: Graphic Badge Logo Input: [image1.png] Output: Labeling results Design elements: (1) Graphic element: a circular gear pattern (2) Text element: Company name "XXX" Color: blue, black Design style: sense of technology Layout: Image above, badge logo below Note: The gear pattern and text of this trademark are surrounded by a badge style, which meets the definition of "graphic badge logo" --------- --------- Example 2: Logo with text and image Input: [image2.png] Output: Labeling results: Design elements: (1) Graphic element: an abstract green tree (2) Text element: Company name "XXX" 2. Color: green, white 3. Design style: natural and environmentally friendly 4. Layout: Text and images combined with logo Note: Graphics and text are separated, their positions do not overlap, and the layout of image on the left and text on the right is a "text-graphic combination logo" --------- II. Define the task output format: Output image elements - text elements - graphic and text layout - color combination - design style in the form of combined coding; III. Overview of the feature extraction task: Extract image elements + text elements + graphic and text layout + color combination + design style features from the input data. If there is an example image input, refer to the example image and its corresponding image description text information to extract the features of the image, and determine the codes of the extracted features according to the codes of each feature in Table 1.

[0041] Table 1. Codes of each feature

[0042] Reference Figure 2 As shown, it is an example diagram of a trademark image input provided in this embodiment to the large language model. Due to the strong learning and generalization ability of the large language model, the various feature information of the extracted picture is a general text description and is not suitable as an image label. For example, based on Figure 2 The extracted features are as follows: 1. Design elements (1) Graphic elements: A red square badge pattern with an abstract graphic inside that resembles an unfolded book or the seal script style character "gua", shaped like the letter "M" or an inverted "U". (2) Text elements: The Chinese character "八月瓜" and the pinyin "BA YUE GUA", where the Chinese character is in the seal script style and the pinyin is in uppercase sans-serif font. 2. Color: Red.

[0043] 3. Design style Combination of retro and modern. The seal script style Chinese characters carry the charm of traditional culture, while the pinyin font is relatively modern. The overall is simple and visually impactful.

[0044] 4. Typography layout Graphic and text combination logo: The graphic and text are separated and exist independently. The graphic is on the left, and the Chinese character and pinyin are on the right, forming a symmetrical and coordinated layout.

[0045] Therefore, it is also necessary to parse the extracted feature information to obtain keywords related to various features and format them to facilitate the learning of the large model; for example, the information extracted from the design style is "retro style, with classic Gothic fonts and decorative sun patterns", but only retro style information is needed as a label. In this embodiment, since the design style is not considered in the image feature extraction process obtained by steps 121 and 122, and the extracted graphic layout features may be inaccurate, and the large language model is not sensitive to distinguishing the layout between graphic elements and text elements when performing feature annotation on each image, the extracted graphic layout features may be inaccurate. Therefore, the image combination features obtained in step S13 are combined with the image features extracted in step S4, and the image features obtained after manual verification are obtained, and the features that are closer to the original image are selected as image features to ensure the accuracy of the feature labels, and the verified images are fed back to the large language model for model fine-tuning, thereby increasing the generation ability of the generative model. The verification steps are as follows: Step S21, displaying the image and the image combination features extracted based on step S13 and the features extracted by the large language model on the front-end interface; In step S22, based on the original image, manually verify and determine features such as image elements, text elements, graphic layout, color combination and design style, and use the verified features as labels for the image.

[0046] In this embodiment, large language model fine-tuning is a common method for enabling large language models to better follow task-specific instructions and produce the desired results. A reasonable instruction dataset is crucial for model effectiveness. Fine-tuning the large language model using multiple validated labeled image subsets achieves better performance with less data and computing resources. Fine-tuning methods include LoRA (Low-Rank Adaptation of Large Language Models) fine-tuning, QLoRA (Quantized LoRa) fine-tuning, and LoRA+MoE (Mixure of Experts) fine-tuning.

[0047] In this embodiment, image label information is used as input to a large language model to generate an image, and the corresponding image is used as a label to calculate the loss with the generated image, thereby improving the large language model's ability to understand and generate label information (graphic elements, text elements, graphic layout, color combination, design style).

[0048] In one embodiment, a model workflow of a generative model is further provided to implement step S6. The generative model workflow is specifically as follows: Step S61, input parameters: the input parameters include generation requirement information and sample images, and the generation requirement information is used to describe the specific requirements for image design; In one embodiment, the generation requirement information may include text information, voice information, or both. If the voice information is input, the voice information needs to be processed by a preset voice recognition module to obtain the image generation requirement text. The generation requirement information may include core image generation parameter requirements, such as the width and height of the generated image, batch size, denoising strength, number of denoising steps, prompt word relevance, sampler, and other parameters. Sample images are used to indicate the image tendency to be generated. The generation requirement information and sample images complement each other to more flexibly meet the user's personalized design needs. The generation requirement information can be positive prompt information or negative prompt information.

[0049] Step S62, sample image feature extraction: the sample image is input into the control generation model to obtain image features, which serve as control conditions for the generative model to generate images, and accurately control the results of the generated images.

[0050] In this embodiment, the control generation model uses the ControlNet model, which uses the Depth function to guide the processing of pictures, prompt words and random noise. ControlNet will extract a depth map based on the sample image, such as extracting features such as edges, colors, and layout from the sample image. Then, the model knows and imitates the overall style and design of the sample image to generate the corresponding image, making the generated image closer to user needs.

[0051] Step S63, generation requirement information completion: the input generation requirement information is converted into text, and extraction and completion operations are performed to obtain the completed image generation requirement text; the purpose is to modify the generation requirement information input by the user so that the generation requirement information is more in line with the style understood by the generative model, because the quality of the prompts given to the large language model has a great influence on the effect of model reasoning, and can more accurately and efficiently understand the user's specific preferences and emotional tendencies in a certain aspect.

[0052] Step S64, main model selection: select an open source generative model as the main model, such as the StableDiffusion model or the Qwen-VL (Qwen Large Vision Language Model) model as the main model for image generation, and input the image features obtained in the sample image feature extraction step and the completed image generation requirement text into the main model.

[0053] Step S65, parameter receiver: when it is determined that the image generation requirement text contains image generation core parameter requirements, the image output parameters of the generative model are modified according to the core parameter requirements. If there are no core parameter requirements in the image generation requirement text, the generative model generates the first trademark generation image according to the default parameters.

[0054] In this embodiment, the completion of the generated requirement information in step S63 can be implemented using the Stable Diffusion prompt manual combined with code to extract and complete the text keywords entered by the customer. The purpose is to make the input keywords more consistent with the style that can be understood by the large language model, so as to improve the quality of the prompt information of the large model and have a great impact on improving the effect of model reasoning; specifically, the following steps are included.

[0055] Step S632, keyword extraction; Step 1.1, Word Segmentation and Semantic Analysis: Use natural language processing technology to analyze the descriptive text entered by the user and extract key keywords. For example, from the sentence "I want a minimalist logo with a primary blue color and a geometric feel," we can extract the keywords "minimalist style," "blue gradient," and "geometric feel."

[0056] Step 1.2, keyword attribute classification: The extracted keywords are organized into categories such as style (such as "minimalist"), layout (such as "round badge"), color (such as "blue"), and elements (such as "line", "border") for input into the StableDiffusion model.

[0057] Step S632, keyword expansion and completion; Step 2.1, Keyword Semantic Expansion: Use pre-trained language models (such as GPT or T5) to semantically expand keywords, ensuring that generated prompts cover more details. For example, "geometric sense" can be expanded to "simple geometric pattern" or "regular symmetrical design"; and "blue" can be expanded to "dark blue gradient" or "light blue with gray."

[0058] Step 2.2: Combine image generation features to complete keyword details: Provide clear stylistic keywords for "Stable Diffusion" (e.g., "minimalist design, clean lines") to avoid vague expressions. Also, consider the trademark's industry background (precondition) and complete the relevant elements. For example, for a technology-related trademark, complete the keywords "technological sense," "futuristic style," and "high-tech texture."

[0059] Step 2.3, Add Negative Prompts: The descriptive text entered by the user will serve as a positive prompt. The system will automatically add default negative prompts to avoid common problems in generated images. For example, many image generation models often have problems generating human fingers and limbs. Adding negative prompts such as "mutated hands and fingers, deformed, disfigured, mutated, supernumerary limbs" can effectively alleviate this problem.

[0060] Step S633: Keyword formatting: In step 3.1, the completed keywords are converted to a prompt format that conforms to Stable Diffusion. The resulting keywords are then adjusted based on the preset part-of-speech weights to obtain an adjusted keyword sequence. For example, core keywords are adjusted to the front of the keyword sequence, while background or secondary requirements are placed at the back. For example, based on the preset part-of-speech weights, the keyword sequence is in the order of "<main description>, <detailed description - layout>, <style>, <color>". For example, "A minimalist logo with geometric shapes, featuring a blue gradient, modern design, clean lines".

[0061] 3.2 Add high-level modifiers that the model understands, for example, style modifiers: "hyper-realistic, vector art, flat design" and visual effect modifiers: "sharp edges, smooth gradients, vibrant colors".

[0062] In one embodiment, to improve the innovativeness of the image and reduce the risk of infringement, in step S7 of this embodiment, the generated image is compared with images in the image library using the search-enhanced generation technology for similarity, and the first trademark generated image with a similarity lower than a first preset value is retained, specifically including the following steps: In step S71, each first trademark generated image is compared with existing images in the image library using search-enhanced generation technology. If the similarity between the first trademark generated image and the existing image exceeds a preset similarity threshold of 50% to 80%, the corresponding first trademark generated image is discarded, and a qualified first trademark generated image is obtained. In this embodiment, the similarity threshold can be set based on the application domain. For example, for trademark image generation, the similarity threshold can be set to 60%, as trademark image review requirements require standardization and uniqueness, and relatively few design elements. For appearance image generation, the similarity threshold can be set to 75%, as the design elements are relatively rich.

[0063] Step S72: For each first trademark generated image that meets the conditions, a preset number of existing images with the highest similarity are obtained, and the first trademark generated image and a list of similar existing images are displayed on the front-end interface for user selection.

[0064] In order to make the generated image more in line with the user's requirements, this embodiment further provides an image regeneration step, the specific steps are as follows: Step S8, based on the first trademark generated image selected by the user, at least one image in the similar existing image list, and the change generation requirement information, input them into the generative model and generate a second trademark generated image and a similar existing image list with a similarity lower than a first preset value based on the generative model workflow.

[0065] In this embodiment, the generated image can be selected through the front-end display interface, or any one or two existing images in the list can be selected as sample images to input into the generative model for another image generation.

[0066] In this preferred embodiment, the first trademark generated image, the second trademark generated image, and the existing image can be circled, and the circled intention can be used as the modification generation requirement information as the input of the generative model. In this way, the user can circle the image to make the model notice the image areas and design elements that the user likes. Therefore, during the image regeneration process, the circled areas in the sample image will be given more weight, thereby retaining the design elements that the user is most concerned about. In addition to selecting sample images, users can also control the generated image by adjusting text keywords.

[0067] The method provided in this embodiment utilizes the powerful generation capability of a large language model. Through the user's descriptive text and sample images, the generated images are more in line with the user's personalized needs. The user does not need to have a professional design or technical background. Only simple input is required to realize the automation of image design, which reduces the dependence on professional designers, saves time and costs, and lowers the threshold for users to obtain customized images. It has broad application prospects.

[0068] Device embodiment According to an embodiment of the present invention, an image generation device based on a large language model is provided. Figure 3 FIG. 1 is a block diagram of an image generation apparatus based on a large language model provided in this embodiment. The image generation apparatus based on a large language model according to an embodiment of the present invention includes: The feature extraction module 10 is used to extract image features based on the acquired trademark image data through a feature extraction model. The image features include graphic layout, color combination and / or graphic elements.

[0069] The clustering module 20 is configured to perform image clustering based on image features using a clustering method to obtain multiple image subsets.

[0070] The prompt template construction module 30 is used to construct a prompt template and set feature category extraction logic, wherein the feature category includes design style, graphic layout, color combination, graphic elements and / or text elements.

[0071] The image labeling module 40 is used to input the prompt template and multiple image subsets into the large language model, extract features from the images in the multiple image subsets, and use the extracted features as image labels to obtain a labeled image set.

[0072] The model fine-tuning module 50 is used to fine-tune the generative model using labeled images to improve the model's understanding and generation capabilities.

[0073] The image generation module 60 is configured to input generation requirement information and / or example patterns into the fine-tuned generation model to generate a preset number of first trademark generation images.

[0074] The image generation method based on a large language model provided in this embodiment takes into account that generating trademark images requires professional designers to creatively conceive and draw according to user needs. This method combines the large language model to first construct various types of image data, and through prompt engineering, enables the model to complete the task of image feature labeling, and uses the feature-labeled images to fine-tune the instructions of the generative model to improve the model's feature understanding and image generation capabilities; then, based on the generation requirement information and sample images input by the user, an image that better meets the user's personalized needs is generated. The user does not need to have a professional design or technical background, and only needs to provide simple input to obtain a satisfactory image design, which reduces dependence on professional designers and significantly shortens the image design cycle and cost.

[0075] In this embodiment, in order to reduce the infringement risk of the generated images, an image screening module is also provided, which is used to compare the similarity of the first trademark generated image with the images in the image library through the retrieval enhancement generation technology, and retain the first trademark generated images whose similarity is lower than a first preset value.

[0076] In this embodiment, the feature extraction module 10 includes an image coarse screening submodule, a fine feature extraction submodule and a feature combination submodule; The image coarse screening module is used to extract basic features of the image through a feature extraction model, and screen out images with similarity higher than a second preset value based on the basic features through similarity calculation to obtain an initial image set.

[0077] In this embodiment, data screening is performed in a library of tens of millions of images, and the model for extracting image features is required to strike a balance between accuracy and efficiency. The feature extraction model in this embodiment selects the ResNet18 model with excellent accuracy and fast inference speed to extract image features. The cosine similarity between image features is calculated, and images with a similarity greater than 85%-95% are deduplicated.

[0078] The fine feature extraction submodule is used to extract the image elements, text elements, graphic and text layout and / or color combination features of each image based on the initial image set.

[0079] The feature combination submodule is used to encode the acquired basic features and fine features of the image and combine them to obtain image combination features, and then use a clustering method to cluster based on the image and text layout features to obtain multiple image subsets.

[0080] In this embodiment, the fine feature extraction submodule is provided with a layout feature extraction unit for extracting image elements and text elements of each image and determining the layout of the image and text, and a color extraction unit for extracting the color of each image, wherein: The layout feature extraction unit is used to identify the category and position of graphic elements in each image through the target detection model, and / or identify the content and position of text elements in each image through the text detection model. When there are graphic elements and text elements, the relative position information is determined by the position coordinates of the graphic elements and the text elements, thereby determining the graphic and text layout features.

[0081] The layout feature extraction unit of this embodiment is mainly used to identify the layout of graphic elements, text elements and graphic elements and text elements in the image. For example, in trademark image design, specific graphic and text layout types include pure text, pure graphics, upper image and lower text, upper text and lower image, left image and right text, left text and right image, circular badge, graphics surrounding text, graphics semi-surrounding text, text surrounding graphics, text semi-surrounding graphics, etc. In this embodiment, the target retrieval model is used to identify and locate the category and position of graphic elements in the image by combining the target detection model and the OCR model. The OCR model is used to identify the text content and position in the image. The relative position relationship between the graphic element position and the text position is determined by the position of the graphic element and the text to give the layout category to which they belong. Each image is allowed to have only one layout category.

[0082] The color extraction unit is used to extract image color information, with the goal of identifying the primary color components in the image. This embodiment proposes an innovative and efficient color extraction method. Color extraction is not simply distinguishing, accumulating, and sorting the color values ​​of all pixels in the image. Two points should be noted: First, the importance of colors in different areas of an image varies. During feature recognition, attention should be paid to the colors of salient targets in the image, while the background color may be more for highlighting salient targets. Therefore, attention should be paid to closed irregular areas in the image. Second, the color values ​​of each pixel are discrete. Even if the colors appear to be the same, the color values ​​themselves may be different. Therefore, it is necessary to classify colors by setting a series of color value ranges.

[0083] In this embodiment, the color extraction unit is used to extract closed irregular graphics after performing edge detection and noise filtering on the image using the OpenCV edge detection algorithm, and regard the irregular graphics as salient targets in the image. The frequency of occurrence of color values ​​in each irregular graphic area is calculated, and the frequency of occurrence of color values ​​is statistically analyzed based on preset color value ranges to obtain the frequency of occurrence of each color category, and the color category with the highest preset frequency is selected as the color combination feature of the image.

[0084] Compared with using models to extract color features, this method takes less time to process. Because the number of images before clustering is still very large, using models to process color extraction requires a long inference time. This method focuses on salient targets and extracts colors, greatly improving the image processing time.

[0085] In one embodiment, the prompt template may include the following content: 1. Define the model's identity: You are an intelligent assistant that is given the following feature information of an image and can make inferences: 1. Design Elements Image elements and text elements: Which image elements and text elements are in the image? The image elements and text elements may be empty, which can be represented by "none".

[0086] 2. Graphic and text layout: The positional relationship between image elements and text elements in an image, such as pure text, pure graphics, image above and text below, text above and image below, image on the left and text on the right, text on the left and image on the right, circular badge, graphics surrounding text, graphics semi-surrounding text, text surrounding graphics, text semi-surrounding graphics, etc.

[0087] 3. Color combination: What colors does the image mainly consist of (such as: red, yellow, green, blue, purple, black).

[0088] 4. Design style: What is the overall design style of the image, such as traditional Chinese style, simple drawing, fresh and natural, cartoon, technological sense, etc.

[0089] 2. Define the task output format: Output image elements, text elements, graphic layout, color combination, and design style in a combined coded format; 3. Feature extraction task overview: Extract image elements + text elements + graphic layout + color combination + design style features from the input data. If there is a sample image input, refer to the sample image and its corresponding image description text information to extract the image features, and determine the encoding of the extracted features based on the encoding of each feature in Table 1.

[0090] In this embodiment, the design style is not considered during the image feature extraction process. When the large language model is used to annotate each image, the layout between graphic elements and text elements is not sensitive. Therefore, the extracted graphic and text layout features may be inaccurate. Therefore, a manual annotation verification module is also provided, including: The display submodule is used to display the combined features of the image and the features extracted by the large language model on the front-end interface; The manual verification submodule is used to verify and determine the features of the image elements, text elements, graphic layout, color combination and design style based on the image, and use the verified features as the label of the image.

[0091] In this embodiment, the image generation module 60 includes the following submodules to generate the first trademark image, specifically including: The input parameter submodule is used to input parameters, including generation requirement information and sample images. The generation requirement information is used to describe the specific requirements for image design; In one embodiment, the generation requirement information can be text or voice. If the voice information is input, it is processed by a pre-configured voice recognition module to generate the image generation requirement text. The generation requirement information can include core image generation parameter requirements, such as the width and height of the generated images, the number of images, the redrawing amplitude, the number of denoising steps, the relevance of the prompt word, and other parameters. Sample images are used to indicate the desired image generation tendency. The generation requirement information and sample images complement each other to more flexibly meet the user's personalized design needs. The generation requirement information can be positive prompt information or negative prompt information.

[0092] The sample image feature extraction submodule is used to input the sample image into the control generation model to obtain image features, which serve as the control conditions for the generative model to generate images and accurately control the results of the generated images.

[0093] The generation requirement information completion submodule is used to convert the input generation requirement information into text, and perform extraction and completion operations to obtain the completed image generation requirement text. The main model selection submodule is used to select an open source generative model and input the image features obtained by the sample image feature extraction submodule and the image generation requirement text completed by the generation requirement information completion submodule into the main model.

[0094] The parameter receiving submodule is used to modify the image output parameters of the generative model according to the core parameter requirements when it is determined that the image generation requirement text contains image generation core parameter requirements. If there are no core parameter requirements in the image generation requirement text, the generative model generates the first trademark generation image according to the default parameters.

[0095] In order to improve the innovation of images and reduce the risk of infringement, the image screening module of this embodiment includes a first screening submodule and a second screening submodule; The first screening submodule is used to compare the similarity of each first trademark generated image with existing images in the image library through the retrieval enhancement generation technology. If the similarity between the first trademark generated image and the existing image is higher than the preset similarity threshold of 50% to 80%, the corresponding first trademark generated image is discarded and the first trademark generated image that meets the conditions is obtained.

[0096] The second screening submodule is used to generate images for each first trademark that meets the conditions, obtain a preset number of existing images with the highest similarity, and display the first trademark generated images and a list of similar existing images on the front-end interface for user selection.

[0097] In order to make the generated image more in line with the user's requirements, this embodiment also provides an image regeneration module for generating an image based on the first trademark selected by the user, at least one image in the similar existing image list, and change generation requirement information, inputting them into the generative model and generating a second trademark generated image and a similar existing image list with a similarity lower than a first preset value based on the generative model workflow.

[0098] In this embodiment, the generated image can be selected through the front-end display interface, or any one or two existing images in the list can be selected as sample images to input into the generative model for another image generation.

[0099] In this preferred embodiment, users can also circle the selected first trademark generated image, second trademark generated image, and existing image through the front-end display interface, and use the circled intention as the change generation requirement information as the input of the generative model. In this way, users can circle the image to make the model notice the image areas and design elements that the user is satisfied with. Therefore, during the image regeneration process, the circled areas in the sample image will be given more weight, thereby retaining the design elements that the user is most concerned about. In addition to selecting sample images, users can also control the generated image by adjusting text keywords.

[0100] The embodiment of the present invention is an apparatus embodiment corresponding to the above-mentioned method embodiment. The specific operations of the processing steps of each module can be understood by referring to the description of the method embodiment, and will not be repeated here.

[0101] like Figure 4 As shown, the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the trademark image generation method based on the large model in the above embodiment is implemented.

[0102] The present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for generating a trademark image based on a large model in the above-mentioned embodiment is implemented. Alternatively, when the computer program is executed by a processor, the method for generating a trademark image based on a large model in the above-mentioned embodiment is implemented. When the computer program is executed by the processor, the following method steps are implemented: Step S1: Based on the acquired trademark image data, image features are extracted using a feature extraction model. The image features include graphic layout, color combination and / or graphic elements.

[0103] Step S2: clustering the images using a clustering method based on the image features to obtain multiple image subsets.

[0104] Step S3: construct a prompt template and set feature category extraction logic, wherein the feature categories include design style, graphic layout, color combination, graphic elements and / or text elements.

[0105] In step S4, the prompt template and multiple image subsets are input into the large language model, features are extracted from the images in the multiple image subsets, and the extracted features are used as image labels to obtain a labeled image set.

[0106] In step S5, the generative model is fine-tuned using the labeled image set to improve the model's understanding and generation capabilities.

[0107] Step S6: inputting the generation requirement information and / or the example pattern into the fine-tuned generative model to generate a preset number of first trademark generation images.

[0108] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0109] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without making any creative efforts.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and the contents not described in detail in the specification of the present invention are common knowledge to those skilled in the art.

Claims

1. A trademark image generation method based on a large model, characterized in that: Including steps: Extracting image features based on the acquired trademark image data using a feature extraction model, the image features including graphic layout, color combination, and / or graphic elements; Using clustering methods based on image features to cluster images and obtain multiple image subsets; Constructing a prompt template and setting up feature category extraction logic, where the feature categories include design style, graphic layout, color combination, graphic elements, and / or text elements; Input the prompt template and multiple image subsets into the large language model, perform feature extraction on the images in the multiple image subsets based on feature category extraction logic, and use the extracted features as image labels to obtain a labeled image set; Fine-tune the generative model using a labeled image set; and The generation requirement information and / or the example pattern are input into the fine-tuned generative model to generate a preset number of first trademark generation images.

2. The method for generating a trademark image based on a large model according to claim 1, wherein: The following steps are also included: The first trademark generated images are compared with images in the image library for similarity by using the retrieval enhancement generation technology, and the first trademark generated images with similarity lower than a first preset value are retained.

3. The method for generating a trademark image based on a large model according to claim 1, wherein: Furthermore, based on the acquired trademark image data, image features are extracted using a feature extraction model, specifically including the following steps: Image coarse screening: extracting basic features of the image through a feature extraction model, and filtering out images with similarity higher than a second preset value based on the basic features through similarity calculation to obtain an initial image set; Fine feature extraction: Based on the initial image set, extract the image elements, text elements, graphic layout and / or color combination features of each image; and The acquired basic features and fine features of the image are encoded and combined to obtain the image composite features.

4. The method for generating a trademark image based on a large model according to claim 3, wherein: Furthermore, the color combination features of each image are extracted, which specifically includes the following steps: After edge detection and noise filtering of the image using an edge detection algorithm, closed irregular shapes are extracted and the irregular shapes are used as salient targets in the image. The frequency of occurrence of each color value in each irregular shape area is calculated, and the frequency of occurrence of the color values ​​is statistically analyzed based on the preset range of each color value to obtain the frequency of occurrence of each color category. The color category with the highest preset frequency is selected as the color combination feature of the image.

5. The method for generating a trademark image based on a large model according to claim 1, wherein: The prompt template is provided with prompt words of detailed task description information, and the detailed task description information includes the execution process of the large language model feature extraction task, feature category definition, feature category extraction logic, and various sample images representing graphic and text layout features.

6. The method for generating a trademark image based on a large model according to claim 1, wherein: It also includes setting up the model workflow for the generative model, which works as follows: Input parameters: Input parameters include generation requirement information and sample images. Generation requirement information is used to describe specific requirements for image design. Sample image feature extraction: The sample image is input into the control generation model to obtain image features, which are used as the control conditions for the generative model to generate images, thereby achieving the result of controlling the generated images; Generation requirement information completion: Convert the input generation requirement information into text, perform extraction and completion operations, and obtain the completed image generation requirement text; Main model selection: Select an open-source generative model as the main model, and input the image features obtained from the sample image feature extraction step and the completed image generation requirement text into the main model; and Parameter receiver: When it is determined that the image generation requirement text has set the image generation core parameter requirements, the image output parameters of the generative model are modified according to the core parameter requirements.

7. The method for generating a trademark image based on a large model according to claim 6, wherein: The generation requirement information includes image generation core parameter requirements; and / or text messages; and / or Voice message.

8. A trademark image generation device based on a large model, characterized in that: include: A feature extraction module, configured to extract image features based on the acquired trademark image data using a feature extraction model, wherein the image features include graphic layout, color combination, and / or graphic elements; A clustering module, used to cluster images using a clustering method based on image features to obtain multiple image subsets; A prompt template construction module is used to construct a prompt template. The prompt template includes prompt words for setting detailed task description information. The detailed task description information includes the execution process of the large language model feature extraction task, the feature category extraction logic, the definition of each feature label, and sample images of each representation class of graphic layout features, where the feature categories include design style, graphic layout, color combination, graphic elements and / or text elements; The image labeling module is used to input the prompt template and each image subset into the large language model, extract features from the images in each image subset, and use the extracted features as labels to label the images; A model fine-tuning module, used to fine-tune the generative model using labeled images; and The image generation module is used to input the generation requirement information and the sample pattern into the fine-tuned generation model to generate a preset number of first trademark generation images.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for generating a trademark image based on a large model according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating a trademark image based on a large model according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Traffic sign recognition method and system

    CN109389167A

  • Method and device for determining face color, medium and program product

    CN114445895A

  • Advertisement image generation method and system, terminal and medium

    CN118296176A

  • Extension device for orthodontics

    KR102688279B1

  • Image processing method and apparatus

    WO2021147670A1