Interactive icon automatic generation system based on artificial intelligence

By introducing semantic style fusion and adaptive noise scheduling, combined with an interactive optimization module, the problems of insufficient customization and interactive design in the icon generation system in the existing technology are solved, and efficient and personalized icon generation and brand consistency are achieved.

CN120672895APending Publication Date: 2025-09-19BEIJING YIYUANKU TECH CO LTD

Patent Information

Application Number
CN202511188354.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing icon generation systems lack customization capabilities, and the generation quality relies on manual parameter adjustment. They are difficult to adapt to icon generation tasks of different complexities. They also lack interactive design capabilities, consume high resources, lack user feedback and learning mechanisms, and are unable to meet the needs of instant feedback.

Method used

An AI-based interactive icon automatic generation system is adopted, which realizes the personalization and efficiency of icon generation through the semantic style fusion module, adaptive noise scheduling and interactive optimization module, combined with the adaptive noise scheduler and interactive optimization sub-module.

Benefits of technology

It improves the accuracy and flexibility of icon generation, optimizes resource utilization, achieves instant response and brand consistency, and reduces computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672895A_ABST
    Figure CN120672895A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence icon generation, and discloses an interactive automatic icon generation system based on artificial intelligence. The system comprises an input module used for obtaining semantic data and style description data; the preprocessing module is used for analyzing and coding the semantic data to obtain a semantic vector; obtaining a style vector corresponding to the style description data in a pre-constructed style library; the semantic style fusion module is used for matching according to the semantic vector and the style vector; the diffusion adjusting module is used for automatically generating diffusion parameters and updating the diffusion parameters according to user feedback; the diffusion generation module is used for gradually denoising according to diffusion parameters and semantic vectors on the basis of the initial noise image; the semantic style fusion module is used for introducing a semantic style fusion module in the denoising process of each step, restoring the guide image according to the corresponding key value, and outputting the denoised image of the current step; according to the method and the device, the icon conforming to the user intention can be more efficiently and accurately generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of icon generation, and more particularly to an interactive icon automatic generation system based on artificial intelligence. Background Art

[0002] In the field of automatic icon generation, existing icon generation systems primarily use the Stable Diffusion model as their underlying architecture. By interpreting user-entered keywords, style, and color preferences, they perform a "text-to-image" (T2I) conversion process, transforming text descriptions into visual representations. Furthermore, the CLIP (Contrastive Language-Image Pre-training) model is widely used to establish associations between text and images, laying the foundation for semantic understanding. Despite research on SVG (Scalable Vector Graphics), such as "path-level T2V diffusion" technology, existing technologies face numerous challenges in generating high-quality vector graphics, particularly in maintaining design consistency and brand style. Most systems still primarily output raster images and lack automatic vectorization capabilities.

[0003] There are many problems that need to be solved in current technologies. First, general models lack customization, and the quality of raw images relies on manual parameter adjustment, which cannot automatically adapt to icon generation tasks of different complexities; the style is controlled only by user prompts, lacks compliance with icon design specifications, has a high semantic mismatch rate, and is difficult to learn and adapt to the company's specific style. Second, the interactive design capabilities are insufficient. The interactive interface of existing AI design tools is limited to simple parameter adjustments, lacks real-time iteration and local editing functions, and has limited design flexibility. Third, resource consumption and operational efficiency are low. The existing system relies on large cloud models and lacks an end-cloud collaborative processing mechanism. The generation time is long and the resource consumption is high, making it difficult to meet the needs of instant feedback. Fourth, there is a lack of user feedback learning mechanism. After deployment, most models cannot be adaptively optimized according to user usage, making it difficult to improve generation quality from user feedback.

[0004] Therefore, how to efficiently and accurately realize the generation of personalized icons is a problem that those skilled in the art need to solve urgently. Summary of the Invention

[0005] In view of this, the present invention provides an interactive icon automatic generation system based on artificial intelligence, which can generate icons that meet user intentions more efficiently and accurately.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] An interactive icon automatic generation system based on artificial intelligence, comprising:

[0008] Input module, used to obtain semantic data and style description data.

[0009] A preprocessing module, configured to parse and encode the semantic data to obtain a semantic vector; and to obtain a style vector corresponding to the style description data in a pre-built style library through matching;

[0010] The semantic-style fusion module is configured to generate a key-value pair by pairing the semantic vector with the style vector.

[0011] A diffusion adjustment module is used to automatically generate diffusion parameters according to the semantic vector and update the diffusion parameters according to user feedback.

[0012] The diffusion generation module is used to initialize the noisy image and gradually denoise the initial noisy image according to the diffusion parameters and the semantic vector to obtain a final denoised image; and is used to introduce the semantic style fusion module in each denoising step, guide image restoration according to the corresponding key-value pairs, and output the denoised image of the current step.

[0013] Preferably, the diffusion regulation module includes an adaptive noise scheduling submodule.

[0014] The adaptive noise scheduling submodule is used to calculate information entropy according to the semantic vector; to determine complexity according to the information entropy, and to set a corresponding number of diffusion steps according to the complexity.

[0015] Preferably, the adaptive noise scheduling submodule increases the number of diffusion steps when the complexity increases, and conversely reduces the number of diffusion steps when the complexity decreases.

[0016] Preferably, the complexity is calculated as follows:

[0017]

[0018] Wherein, d is the feature dimension of the semantic vector, i and j are different summation indexes, and v is the semantic vector.

[0019] Preferably, the diffusion adjustment module further includes an interactive optimization submodule; the interactive optimization submodule is used to obtain and identify the user's fine-tuning actions on the icon, and adjust the current icon by updating the diffusion parameters.

[0020] Preferably, the interactive optimization submodule includes a human-computer interaction interface, a feedback recognition unit, a mapping unit and an immediate response unit.

[0021] The human-computer interaction interface is used to collect user fine-tuning actions.

[0022] The feedback recognition unit is used to generate corresponding action instructions according to the user's fine-tuning action.

[0023] The mapping unit is used to map the action instruction to the corresponding parameter to be adjusted to obtain a parameter weight adjustment signal.

[0024] The immediate response unit is used to apply the parameter weight adjustment signal to the diffusion generation module to update the icon.

[0025] Preferably, it further includes a consistency correction module, which is used to compare the current image with a preset template through local feature matching and correct the part that does not conform to the preset template.

[0026] A diffusion model for automatic icon generation, comprising:

[0027] The input layer is used to input the semantic vector, style vector, diffusion time step embedding vector and noise image.

[0028] An information fusion layer is used to perform cross-attention fusion on the semantic vector, the style vector, the diffusion time step embedding vector, and the noise image.

[0029] The encoder layer uses a multi-level encoding structure to perform multi-level feature extraction on the fused vector.

[0030] The CAFL layer is used to perform linear projection based on the semantic vector and the style vector to construct a key-value pair.

[0031] The decoder layer uses a multi-level decoding structure to restore the feature extraction results of the encoder layer layer by layer according to the key-value pairs.

[0032] Preferably, an adaptive noise scheduler is further included, which is used to dynamically adjust the number of diffusion steps according to the complexity of the semantic vector.

[0033] It can be seen from the above technical solutions that, compared with the prior art, the present invention provides an interactive icon automatic generation system based on artificial intelligence, which performs style fusion by introducing key-value pairs of semantic style fusion. On the basis of guiding image generation with basic prompt words, it can use style description to achieve further guidance and improve accuracy. Based on semantically guided style fusion, the present invention proposes interactive icon optimization, which can generate icons in a timely manner according to user operation feedback. The present invention proposes an adaptive icon generation method, which introduces adaptive noise adjustment, confirms the task difficulty according to the semantic description, and adjusts the generation resources required for generation according to the task difficulty, thereby improving system flexibility, optimizing the working mode, and saving computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0035] Figure 1 This is a schematic diagram of the structure of an interactive icon automatic generation system based on artificial intelligence provided in an embodiment of the present invention.

[0036] Figure 2 A schematic diagram of a diffusion model structure for automatic icon generation provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0038] Example 1

[0039] like Figure 1 The embodiment of the present invention discloses an interactive icon automatic generation system based on artificial intelligence, comprising:

[0040] Input module, used to obtain semantic data and style description data.

[0041] The preprocessing module is used to parse and encode the semantic data to obtain a semantic vector; and is used to obtain a style vector corresponding to the style description data in a pre-built style library through matching.

[0042] The semantic-style fusion module is used to generate key-value pairs based on pairing of semantic vectors and style vectors.

[0043] The diffusion adjustment module is used to automatically generate diffusion parameters based on semantic vectors and update the diffusion parameters based on user feedback.

[0044] The diffusion generation module is used to initialize the noisy image and gradually denoise it based on the diffusion parameters and semantic vectors to obtain the final denoised image. It is used to introduce the semantic style fusion module in each denoising step, guide image restoration according to the corresponding key-value pairs, and output the denoised image of the current step.

[0045] In this embodiment, the diffusion adjustment module can analyze the semantic complexity of the semantic data input by the user and control the diffusion process based on the complexity, thereby improving the flexibility of diffusion. At the same time, the diffusion adjustment module can also accept user feedback to achieve diffusion control.

[0046] To further implement the above technical solution, the input module is multimodal, supporting multiple input modes such as text, sketches, and voice, enabling comprehensive collection of user needs. Text input is extracted through a pre-trained language model (CLIP), sketch input is extracted through a convolutional neural network (CNN), and voice input is converted into text descriptions through a speech recognition model (Whisper), ultimately achieving unified encoding of multimodal data.

[0047] In addition, in the semantic-style fusion module, the semantic vector obtained by parsing and encoding the semantic data input by the multimodal input module is abstracted into a unique key to identify the core meaning of the content; the style vector retrieved from the style library is used as the value to form a pairing relationship with the semantic key to control the expression of the semantics and achieve precise single semantic guidance.

[0048] In this embodiment, a semantic style fusion module is constructed based on the style description. The key-value pairs obtained by the semantic style fusion module are injected into the diffusion generation process as intermediate conditions. In addition to obtaining the initial input information (the prompt word vector obtained from the semantic data), the diffusion generation process can further realize semantic-guided style fusion through key-value pairs, making the icon generation results more in line with user needs.

[0049] Further implementing the above technical solution, the diffusion adjustment module includes an adaptive noise scheduling submodule; the adaptive noise scheduling submodule is used to calculate information entropy based on the semantic vector; and is used to determine complexity based on the information entropy and set a corresponding number of diffusion steps based on the complexity.

[0050] This implementation plan proposes a method for determining complexity based on the information entropy of semantic vectors:

[0051]

[0052] Wherein, d is the feature dimension of the semantic vector, i and j are different summation indexes, and v is the semantic vector.

[0053] In this embodiment, semantic complexity is measured based on information entropy. When the semantic vectors are dispersed in the feature space, that is, when the difference in the values ​​of each dimension is small, the information entropy is high, indicating that the semantics contains multiple independent features, so multiple diffusion steps are required to achieve refined processing. When the vector distribution is concentrated and the information entropy is low, it means that the semantics focuses on a single feature, which means that it has a smaller processing requirement and can reduce diffusion iterations to improve efficiency. Therefore, as a feasible implementation method, the adaptive noise scheduling submodule increases the number of diffusion steps when the complexity increases, and conversely, reduces the number of diffusion steps when the complexity decreases. Exemplarily, when S<0.3, the number of steps N=20 is set, when 0.3≤S<0.7, N=50, and when S≥0.7, N=100.

[0054] Furthermore, in addition to the adaptive noise scheduler (ANS), one can also use a learning-based noise scheduler or dynamically adjust the noise level based on other metrics such as the clarity of the generated image.

[0055] The learning-based noise scheduler uses a neural network trained on a large amount of data to determine the noise level. It first prepares a dataset containing different semantic and stylistic information, as well as images that generate good results. It then uses semantic vectors and style vectors as input, and appropriate noise scheduling parameters as output labels, training the neural network to learn the mapping between the two. Once trained, when generating icons, the semantic and style information is input, and the network outputs a matching noise scheduling solution. For example, when encountering complex semantics and specific style combinations, it automatically outputs a noise variation strategy that produces a more detailed image.

[0056] The noise level is dynamically adjusted based on the clarity of the generated image. At each step of the diffusion generation process, algorithms such as the Structural Similarity Index (SSIM) and Peak Signal-to-Noise Ratio (PSNR) are used to evaluate the clarity of the denoised image. If the clarity is low and falls below a pre-set threshold, indicating insufficient image detail, the noise level is reduced, allowing the model to focus more on refining the image during the next denoising step, accelerating the generation of a clearer image. If the clarity is abnormal, such as signs of overfitting such as oversharpening or texture distortion, the noise level is increased to introduce more variation, increase the diversity of the generated image, and avoid image distortion. This allows the noise level to adjust in real time based on the image generation process.

[0057] To further implement the above technical solution, we propose a dual-loss constrained training mechanism. In addition to the standard LDM reconstruction loss, we add a semantic consistency contrast loss function: L_sem = 1-cos(CLIP(I_gen), CLIP(T)), where I_gen is the generated image and T is the original text description. Experiments show that this dual-loss mechanism reduces the semantic mismatch rate of the generated draft by 27%.

[0058] During training, the original text description and the corresponding initial noisy image are first input. The diffusion generation module gradually denoises the generated image I_gen based on the semantic and style vectors. At this point, the standard LDM reconstruction loss compares the pixel information of the generated image with the expected result, calculates the difference between the two, and adjusts the model parameters through backpropagation to ensure the accuracy of the image's visual effect. Simultaneously, the CLIP model is used to extract the feature vectors of the generated image I_gen and the original text description T. By calculating the cosine similarity between the two, the semantic consistency comparison loss L_sem is calculated. This loss value reflects the degree of deviation between the image semantics and the text semantics. The model adds these two loss values ​​as the total loss and continuously optimizes the model parameters using a gradient descent algorithm. This ensures that the generated image meets the visual reconstruction requirements while maximizing semantic consistency with the original text description, thereby reducing the semantic mismatch rate.

[0059] The above embodiment improves the flexibility of the model by adaptively adjusting the diffusion parameters. In addition, the present invention also introduces diffusion regulation based on interaction.

[0060] Furthermore, the diffusion adjustment module also includes an interactive optimization submodule. This module receives real-time user feedback and dynamically adjusts the denoising weights to achieve an immediate response. User feedback primarily refers to the fine-tuning performed by the user on the current icon, and the denoising weights refer to the attention weights.

[0061] In this embodiment, the interactive optimization submodule includes a human-computer interaction interface, a feedback recognition unit, a mapping unit, and an immediate response unit.

[0062] The human-computer interaction interface is used to collect the user's fine-tuning actions on the local area of ​​the icon.

[0063] The feedback recognition unit converts user actions into structured action instructions through event monitoring and semantic analysis; the mapping unit associates action instructions with the parameters to be adjusted in the diffusion model according to predefined key-value pair mapping rules, and generates a parameter weight adjustment signal with priority; the immediate response unit injects the adjustment signal into the noise scheduler and conditional embedding layer of the diffusion generation module, dynamically modifies the noise distribution in the denoising process, and realizes real-time update of the icon.

[0064] The interactive optimization engine's implementation principle is to convert local user adjustments into quantifiable semantic deviation signals, rather than directly comparing pixel differences. When a user adjusts an icon through the interface, the system first maps the user's actions into changes in predefined semantic key-value pairs in a mapping unit. This is then compared with the semantic feature vector extracted from the current generated image via the CLIP model for cross-modal similarity calculations, resulting in a semantic deviation matrix. This matrix is ​​normalized and then fed into the cross-modal fusion module, where weight parameters are generated through a dynamic routing mechanism. These weight parameters are applied to the noise scheduler and conditional embedding layer of the diffusion model. Specifically, the time step distribution of the noise scheduler is adjusted, so that regions related to the user's actions receive higher iteration weights in subsequent denoising. The conditional embedding layer also enhances the activation strength of feature channels corresponding to the new semantic values. Furthermore, the deviation signal is used to lock onto the corresponding regional features in the UNet network through an attention gating mechanism, prioritizing the correction of the noise prediction direction in that region during denoising, thereby achieving a precise response to the user's adjustment intentions.

[0065] Furthermore, based on the interactive optimization engine, each user interaction data is stored for model fine-tuning. The system stores each semantic deviation data, along with the original input and the intermediate feature maps from the generation process, as a triplet. Using a contrastive learning framework, the model's semantic mapping parameters are periodically updated, allowing the model to gradually memorize the user's semantic adjustment patterns. This ultimately forms a closed-loop optimization mechanism of "semantic manipulation-noise control-model evolution," ensuring that feedback adjustments meet visual modification requirements while maintaining semantic consistency.

[0066] In order to further implement the above technical solution, considering the importance of corporate brand image to icon design, this system introduces a brand consistency correction module before generating the result output. It compares the current image with the preset template through local feature matching and corrects the parts that do not conform to the preset template.

[0067] The specific steps include:

[0068] Brand preset parameters are loaded, and users pre-set unified design specifications such as brand colors, font styles, and graphic styles in the system. Before the brand consistency correction process starts, users can systematically preset brand design specifications through the visual interface provided by the system. In terms of color, not only can the main color be set, but also auxiliary colors and accent colors can be added, and the usage scenarios and proportions of each color in the icon can be clearly defined; in the font style setting link, parameters such as font type, font size, font weight, and character spacing can be accurately specified to ensure that the text part of the icon is consistent with the brand tone; in terms of graphic style, whether it is flat, realistic or minimalist, users can define details such as line thickness, corner curvature, light and shadow effects. These preset parameters will be stored in the system database as the core reference standard for subsequent correction work, laying the foundation for consistent control of icon generation.

[0069] Style matching and local adjustment. During the generation process, the current icon is compared with the brand preset template using a local feature matching method to detect style and color deviations, and local resampling technology is used to correct non-compliant parts. During the icon generation process, the system uses advanced local feature matching algorithms to perform a detailed pixel-by-pixel and element-by-element comparison of the currently generated icon with the preset brand template. By calculating multi-dimensional information such as the image's texture features, color histograms, and shape contours, the area where the icon deviates from the brand standard in style and color is accurately located. For detected problems, the system uses local resampling technology to make corrections. For example, when it is found that the color of an icon element deviates from the preset color value, the pixels in the area will be resampled to make its color value consistent with the brand standard; if the icon shape style does not match the template, the local lines and contours are adjusted to ensure that the graphic elements fit the brand's unified visual style, thereby achieving a high degree of unity between the icon details and brand specifications.

[0070] LoRA fine-tuning technology: The system uses low-rank adaptation (LoRA) technology to perform task-specific adjustments to the generative model, ensuring that the output icons meet multimodal semantic requirements while strictly adhering to brand standards. During the training process, the system targets the brand's preset parameters and multimodal semantic requirements, and inputs a large amount of icon data that meets brand standards into the model. The LoRA module automatically adjusts the newly added low-rank matrix parameters based on the difference between the input data and the target output, enabling the generative model to learn the brand's unique design style and semantic characteristics. After the LoRA fine-tuning model, in subsequent icon generation, it can directly output high-quality icons that both meet semantic requirements and strictly adhere to brand standards, effectively improving the accuracy of icon generation and brand consistency.

[0071] Example 2

[0072] like Figure 2 Based on the same inventive concept, an embodiment of the present invention discloses a diffusion model for automatic icon generation, including:

[0073] Input layer, used to input semantic vector, style vector, diffusion time step embedding vector and noise image;

[0074] Information fusion layer, used to cross-attentionally fuse the semantic vector, style vector, diffusion time step embedding vector, and noise image;

[0075] The encoder layer uses a multi-level encoding structure to perform multi-level feature extraction on the fused vector;

[0076] The CAFL layer is used to perform linear projection based on the semantic vector and style vector to construct key-value pairs;

[0077] The decoder layer uses a multi-level decoding structure to restore the feature extraction results of the encoder layer layer by layer according to the key-value pairs.

[0078] In order to further implement the above technical solution, an adaptive noise scheduler is also included, which is used to dynamically adjust the number of diffusion steps according to the complexity of the semantic vector.

[0079] In this embodiment, the diffusion model uses LDM reconstruction loss and semantic consistency contrast loss for training constraints.

[0080] In order to further implement the above technical solution, it also includes a first bottleneck layer and a second bottleneck layer; the CAFL layer is arranged between the first bottleneck layer and the second bottleneck layer; wherein, the first bottleneck layer is connected to the output of the encoder layer, and the second bottleneck layer is connected to the input of the decoder layer.

[0081] In this embodiment, the first bottleneck layer compresses the high-dimensional features output by the encoder into a low-dimensional semantic space, achieving abstract semantic refinement of the features (for example, converting specific features such as icon outline and color into abstract semantic representations such as "minimalist style" and "technical feel"). The second bottleneck layer restores the low-dimensional semantic features modulated by CAFL to the high-dimensional feature space required by the decoder, providing a semantically aligned feature foundation for subsequent image reconstruction. This "compression-modulation-restoration" structure avoids the computational redundancy of directly processing semantic information in the high-dimensional feature space, while forcing the model to capture the most critical semantic features through dimensionality compression. The CAFL layer, deployed between the two bottleneck layers, performs targeted modulation of the compressed low-dimensional features based on semantic and style vectors. Specifically, by constructing key-value pairs through linear projection, the CAFL layer converts input semantic conditions (such as "circular icon" and "flat style") into feature manipulation signals, specifically enhancing or suppressing semantically relevant feature channels (for example, strengthening circular outline features and suppressing redundant texture features). This design enables semantic control to influence the entire generation process with minimal computational cost, effectively establishing a "semantic conversion hub" between the encoder and decoder.

[0082] The compression mechanism of the dual bottleneck layers forces the model to refine core semantics in a low-dimensional space, preventing redundant information in high-dimensional features from interfering with semantic alignment. The CAFL layer further uses a key-value pair mechanism to filter features that match the input semantics, creating a cascade of "semantic filtering and reinforcement" that significantly improves the matching accuracy between the generated icon and the semantic vector. Because the bottleneck layer features are global, abstract representations of the encoder's output, CAFL's modulation of these features can globally influence the decoder's reconstruction process, ensuring that the overall shape and style of the generated icon align with the semantic conditions (for example, avoiding local semantic drift, such as generating a circular dial when semantics require a square one).

[0083] In addition, deploying the CAFL layer in the low-dimensional bottleneck space improves the model inference efficiency while ensuring the semantic control effect, compared to directly introducing semantic control in the high-dimensional feature layer of the encoder / decoder.

[0084] In this embodiment, the encoder layer in Unet includes multi-stage encoding and downsampling operations, that is, downsampling is performed once after each encoding to achieve step-by-step downsampling; similarly, the decoder layer includes multi-stage decoding and upsampling.

[0085] To further implement the above technical solution, a CAFL layer is added after each downsampling block on the encoder side. This layer linearly projects the feature maps using semantic and style vectors, enhancing the semantic expressiveness of low-level features. This embodiment uses CAFL to guide the generation process early, preventing subsequent generation from deviating from the semantic target and addressing the issue of insufficient semantic expression of low-level features.

[0086] A CAFL layer is introduced before each upsampling block on the decoder side, guiding the feature reconstruction process through a key-value pair mechanism to improve detail recovery accuracy. This embodiment uses the CAFL layer to provide explicit key-value pairs for the decoding process, making the detail generation process more directional.

[0087] In addition, a skip connection can be established between the encoder side and the decoder side, and a CAFL layer is introduced in the skip connection of each layer.

[0088] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0089] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An interactive icon automatic generation system based on artificial intelligence, characterized in that: include: Input module, used to obtain semantic data and style description data; A preprocessing module, configured to parse and encode the semantic data to obtain a semantic vector; and to obtain a style vector corresponding to the style description data in a pre-built style library through matching; A semantic-style fusion module, configured to generate a key-value pair by pairing the semantic vector with the style vector; a diffusion adjustment module, configured to automatically generate diffusion parameters according to the semantic vector and update the diffusion parameters according to user feedback; a diffusion generation module, configured to initialize a noise image and gradually perform denoising on the initial noise image according to the diffusion parameters and the semantic vector to obtain a final denoised image, i.e., a target icon; It is used to introduce the semantic style fusion module in each denoising process, guide image restoration according to the corresponding key-value pairs, and output the denoised image of the current step.

2. The interactive icon automatic generation system based on artificial intelligence according to claim 1, characterized in that: The diffusion regulation module includes an adaptive noise scheduling submodule; The adaptive noise scheduling submodule is used to calculate information entropy according to the semantic vector; to determine complexity according to the information entropy, and to set a corresponding number of diffusion steps according to the complexity.

3. The interactive icon automatic generation system based on artificial intelligence according to claim 2, characterized in that: The adaptive noise scheduling submodule increases the number of diffusion steps when the complexity increases, and conversely reduces the number of diffusion steps when the complexity decreases.

4. The interactive icon automatic generation system based on artificial intelligence according to claim 2 or 3, characterized in that: The complexity is calculated as follows: ; Wherein, d is the feature dimension of the semantic vector, i and j are different summation indexes, and v is the semantic vector.

5. The interactive icon automatic generation system based on artificial intelligence according to claim 1, characterized in that: The diffusion adjustment module further includes an interactive optimization submodule; the interactive optimization submodule is used to obtain and identify the user's fine-tuning action on the icon, and adjust the current icon by updating the diffusion parameters.

6. The interactive icon automatic generation system based on artificial intelligence according to claim 5, characterized in that: The interactive optimization submodule includes a human-computer interaction interface, a feedback recognition unit, a mapping unit and an immediate response unit; The human-computer interaction interface is used to collect user fine-tuning actions; The feedback recognition unit is used to generate corresponding action instructions according to the user's fine-tuning action; The mapping unit is used to map the action instruction to the corresponding parameter to be adjusted to obtain a parameter weight adjustment signal; The immediate response unit is used to apply the parameter weight adjustment signal to the diffusion generation module to update the icon.

7. The interactive icon automatic generation system based on artificial intelligence according to claim 1, characterized in that: It also includes a consistency correction module, which is used to compare the current image with a preset template through local feature matching and correct the part that does not conform to the preset template.

8. A diffusion model for automatic icon generation, characterized in that: include: Input layer, used to input semantic vector, style vector, diffusion time step embedding vector and noise image; an information fusion layer, configured to perform cross-attention fusion on the semantic vector, the style vector, the diffusion time step embedding vector, and the noise image; The encoder layer uses a multi-level encoding structure to perform multi-level feature extraction on the fused vector; A CAFL layer is used to perform linear projection based on the semantic vector and the style vector to construct a key-value pair; The decoder layer uses a multi-level decoding structure to restore the feature extraction results of the encoder layer layer by layer according to the key-value pairs.

9. The diffusion model for automatic icon generation according to claim 8, characterized in that: It also includes an adaptive noise scheduler for dynamically adjusting the number of diffusion steps according to the complexity of the semantic vector.

10. The diffusion model for automatic icon generation according to claim 8, characterized in that: The diffusion model uses LDM reconstruction loss and semantic consistency contrast loss for training constraints.

Citation Information

Patent Citations

  • Multi-modal man-machine interaction system and control method thereof

    CN106569613A

  • Intelligent LOGO generation method based on StyleGAN

    CN114219875A

  • Animation image style migration method and system based on Stable Diffusion

    CN117495662A

  • Figure graph model, model training method and device, image generation method and device and electronic equipment

    CN118657845A

  • Video generation method and device, computer program product and electronic equipment

    CN119583845A

Cited By

  • Artificial intelligence-based automatic calibration method and system for creative and cultural patterns

    CN121280220A

  • An AI-based automatic calibration method and system for cultural and creative patterns

    CN121280220B