Artificial intelligence-based automatic generation method and system for creative and creative graphs
By obtaining the target cultural and creative elements information for text encoding and visual feature extraction, the image is generated, and the unqualified images are adjusted in combination with the scoring mechanism, the problem of not being able to automatically generate and optimize cultural and creative images in the existing technology is solved, and high-quality and consistent automatic generation of cultural and creative images is achieved.
Patent Information
- Application Number
- CN202510594792.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
There is a lack of an intelligent cultural and creative image generation method that can generate images based on the target cultural and creative element information, and automatically extract aesthetic preferences based on the comparison results, and adaptively adjust and optimize unqualified images.
By obtaining the target cultural and creative elements information, text encoding and visual feature extraction are performed, the two are fused and visualized to generate an image that matches the input semantics. Then, by comparing the image with the target cultural and creative elements, combining the scoring mechanism to obtain the scoring results, and if it fails, the image will be modified according to its aesthetic preferences.
It realizes that high-quality images are automatically generated based on the target cultural and creative elements information, and the unqualified images are adjusted through aesthetic preferences, improving the quality and consistency of images and meeting the needs of automatic generation of high-quality cultural and creative images.
Smart Images

Figure CN120107419A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cultural and creative production, and in particular, relates to a method and system for automatically generating cultural and creative images based on artificial intelligence. Background Art
[0002] With the continuous development of artificial intelligence technology, the ability of computers to automatically generate images has been increasingly enhanced, and has received widespread attention in the fields of cultural and creative industries, digital collection design, and intangible cultural heritage pattern inheritance. In particular, the image generation system based on generative adversarial networks and diffusion models can generate pattern images with specific semantic and style characteristics based on input text or image prompt information, and has good application prospects in improving design efficiency and expanding visual expression space.
[0003] At present, there are many existing image generation methods based on text or multimodal input. Some methods use pre-trained language models to encode input text and input the encoded vector into the image generation model to achieve text-to-image conversion; there are also methods that extract the style feature vector of the image sample and combine it with text information to guide the model to generate an image that combines semantics and style. This type of method has made certain technical progress in general image generation, image style transfer, illustration synthesis, and other directions.
[0004] Currently, there is still a lack of an intelligent cultural and creative image generation method that can generate images based on the target cultural and creative element information, automatically extract aesthetic preferences based on the comparison results, and adaptively adjust and optimize unqualified images. Summary of the invention
[0005] The purpose of the embodiment of the present invention is to provide a method for automatically generating cultural and creative images based on artificial intelligence, aiming to solve the problems raised in the third part of the background technology.
[0006] The embodiment of the present invention is implemented as follows: a method for automatically generating cultural and creative images based on artificial intelligence, the method comprising: Acquire target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background and semantic tags; Perform text encoding based on the target cultural and creative element information, extract visual features based on the target cultural and creative element information, and fuse the text encoding and visual features; Visualize the fusion results and use the multimodal embedding vector as a conditional input to guide the model to generate images that match the input semantics, detect the image and the target cultural and creative element information, and judge the consistency of image generation; Compare the image with the target cultural and creative elements, obtain the scoring result based on the comparison result and the scoring mechanism, compare the scoring result with the scoring threshold, obtain the aesthetic preference based on the comparison result, and if it is judged as unqualified, modify the image based on the aesthetic preference.
[0007] Preferably, the steps of performing text encoding according to the target cultural and creative element information, performing visual feature extraction according to the target cultural and creative element information, and fusing the text encoding and the visual features specifically include: Performing text encoding according to the target cultural and creative element information, wherein the text encoding is to convert the target cultural and creative element information text information into a vector representation of a fixed length; Visual feature extraction is performed based on the target cultural and creative element information. The visual feature extraction extracts the image feature vector through a visual encoder; The text encoding and the visual features are fused to obtain a fusion result, wherein the fusion result is obtained by projecting the text and the visual features into the same vector space and then adding them together.
[0008] Preferably, the step of performing visualization processing according to the fusion result, taking the multimodal embedding vector as a conditional input, guiding the model to generate an image matching the input semantics, detecting the image and the target cultural and creative element information, and judging the consistency of the image generation specifically includes: Performing visualization processing according to the fusion result, wherein the visualization processing includes model selection and image generation, and the model selection includes a diffusion model and a generative adversarial network; The multimodal embedding vector is used as a conditional input to guide the model to generate images that match the input semantics and detect the information between the image and the target cultural and creative elements; The consistency of image generation is determined, including image fusion, style transfer, and semantic consistency.
[0009] Preferably, the steps of comparing the image with the target cultural and creative element, obtaining a scoring result based on the comparison result in combination with a scoring mechanism, comparing the scoring result with a scoring threshold, obtaining an aesthetic preference based on the comparison result, and modifying the image based on the aesthetic preference if the image is determined to be unqualified, specifically include: Comparing the image with the target cultural and creative element, obtaining a comparison result, and obtaining a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style; A scoring result is obtained according to the comparison result combined with the scoring mechanism, and a scoring threshold is obtained, where the scoring threshold is the qualified line of the image, and the scoring result is compared with the scoring threshold to obtain a comparison result; Aesthetic preferences are obtained based on the comparison results, including color tone preferences, graphic complexity preferences, and element composition ratio preferences. If the image is determined to be unqualified, the image is modified based on the aesthetic preferences.
[0010] Preferably, the element types include traditional painting, sculpture and calligraphy, the style features include ink painting style and meticulous painting style, the cultural background includes historical periods and regional characteristics, and the semantic tags are descriptive keywords for the elements.
[0011] Another object of an embodiment of the present invention is to provide an automatic generation system of cultural and creative images based on artificial intelligence, characterized in that the system comprises: A cultural and creative element information module is used to obtain target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background and semantic tags; The coding feature module performs text coding based on the target cultural and creative element information, extracts visual features based on the target cultural and creative element information, and integrates the text coding and visual features; The visualization processing module performs visualization processing based on the fusion results, takes the multimodal embedding vector as a conditional input, guides the model to generate images that match the input semantics, detects the image and the target cultural and creative element information, and determines the consistency of image generation; The aesthetic preference module compares the image with the target cultural and creative elements, obtains the scoring result based on the comparison result and the scoring mechanism, compares the scoring result with the scoring threshold, obtains the aesthetic preference based on the comparison result, and modifies the image according to the aesthetic preference if it is judged to be unqualified.
[0012] Preferably, the coding feature module includes: A text encoding unit, performing text encoding according to the target cultural and creative element information, wherein the text encoding is to convert the target cultural and creative element information into a vector representation of a fixed length; A visual feature extraction unit extracts visual features based on target cultural and creative element information. The visual feature extraction extracts image feature vectors through a visual encoder. The fusion unit fuses the text encoding and the visual features to obtain a fusion result, wherein the fusion result is obtained by projecting the text and the visual features into the same vector space and then adding them together.
[0013] Preferably, the visualization processing module includes: A visualization processing unit, performing visualization processing according to the fusion result, wherein the visualization processing includes model selection and image generation, and the model selection includes a diffusion model and a generative adversarial network; The image detection unit takes the multimodal embedding vector as a conditional input, guides the model to generate an image that matches the input semantics, and detects the image and the target cultural and creative element information; The consistency determination unit determines the consistency of image generation, wherein the consistency includes image fusion, style migration and semantic consistency.
[0014] Preferably, the aesthetic preference module comprises: A similarity determination unit compares the image with the target cultural and creative element, obtains the comparison result, and obtains a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style; A scoring unit, which obtains a scoring result according to the comparison result and the scoring mechanism, obtains a scoring threshold, wherein the scoring threshold is a qualified line of the image, and compares the scoring result with the scoring threshold to obtain a comparison result; The aesthetic preference unit obtains the aesthetic preference according to the comparison result, wherein the aesthetic preference includes the color tone preference, the graphic complexity preference and the element composition ratio preference. If the image is judged to be unqualified, the image is modified according to the aesthetic preference.
[0015] Preferably, the element types include traditional painting, sculpture and calligraphy, the style features include ink painting style and meticulous painting style, the cultural background includes historical periods and regional characteristics, and the semantic tags are descriptive keywords for the elements.
[0016] The embodiment of the present invention provides an artificial intelligence-based automatic generation method for cultural and creative images, which obtains target cultural and creative element information, performs text encoding according to the target cultural and creative element information, performs visual feature extraction according to the target cultural and creative element information, extracts image feature vectors through a visual encoder, fuses the text encoding and the visual features to obtain a fusion result, and performs visualization processing according to the fusion result. The model selection includes a diffusion model and a generative adversarial network, takes a multimodal embedding vector as a conditional input, guides the model to generate an image that matches the input semantics, detects the image and the target cultural and creative element information, judges the consistency of image generation, compares the image with the target cultural and creative element, obtains the comparison result, obtains a scoring mechanism, obtains a scoring result according to the comparison result combined with the scoring mechanism, obtains a scoring threshold, compares the scoring result with the scoring threshold, obtains the comparison result, obtains the aesthetic preference according to the comparison result, and modifies the image according to the aesthetic preference if it is judged to be unqualified, thereby solving the problem that the generated cultural and creative image cannot be scored in the existing automatic generation process of cultural and creative images, and cannot be modified according to the scoring result. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flowchart of a method for automatically generating cultural and creative images based on artificial intelligence provided by an embodiment of the present invention; Figure 2 A flowchart of the steps of performing text encoding according to target cultural and creative element information and performing visual feature extraction according to the target cultural and creative element information provided by an embodiment of the present invention; Figure 3 A flowchart of the steps of using a multimodal embedding vector as a conditional input, detecting information of an image and a target cultural and creative element, and determining consistency of image generation provided by an embodiment of the present invention; Figure 4 A flowchart of the steps of comparing the scoring result with the scoring threshold and modifying the image according to aesthetic preference if the image is judged to be unqualified provided in an embodiment of the present invention; Figure 5An architecture diagram of an automatic generation system of cultural and creative images based on artificial intelligence provided by an embodiment of the present invention; Figure 6 The architecture diagram of the coding feature module provided by the embodiment of the present invention; Figure 7 An architectural diagram of a visualization processing module provided by an embodiment of the present invention; Figure 8 An architectural diagram of an aesthetic preference module provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0019] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of this application, a first xx script may be referred to as a second xx script, and similarly, a second xx script may be referred to as a first xx script.
[0020] like Figure 1 As shown, an automatic generation method of cultural and creative images based on artificial intelligence is provided in an embodiment of the present invention, and the method includes: S100, obtaining target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background and semantic tags.
[0021] In this step, the target cultural and creative element information is obtained. Obtaining the target cultural and creative element information is a key pre-step in the automatic generation method of cultural and creative images. It is to comprehensively analyze and structure the cultural elements associated with user needs or target scenes, and provide semantic basis and style guidance for subsequent image generation. The target cultural and creative element information includes four dimensions: element type, style characteristics, cultural background and semantic labels, which are respectively expanded as follows: Element type refers to the core visual elements that constitute the cultural and creative image, which is the main content of the image generation process. Common element types include natural elements (such as landscapes, flowers, plants, and animals), artifact elements (such as ceramics, lacquerware, and clothing), symbol elements (such as totems, patterns, and characters), character elements (such as historical figures and mythological images), and architectural elements (such as ancient buildings and religious sites). Style features refer to the expression techniques and aesthetic styles of the cultural and creative elements in visual presentation, which directly affect the visual direction of image generation. Style features can be derived from user settings or automatically inferred by the system based on existing image samples or historical preferences. Cultural background refers to the historical origin, regional culture, folk beliefs or national characteristics of the cultural and creative elements. It provides semantic context and cultural symbol references, and is the basis for ensuring that the cultural and creative patterns have cultural depth and accuracy. Semantic tags are highly summarized and structured annotations of the above three dimensions of information, which facilitates the semantic vectorization processing of artificial intelligence models. Semantic tags are usually composed of structured tag words, which can be combined with natural language processing technology to automatically extract semantic keywords in texts, or standardized matching by manually annotated knowledge bases.
[0022] S200, performing text encoding according to target cultural and creative element information, performing visual feature extraction according to target cultural and creative element information, and fusing the text encoding and visual features.
[0023] In this step, text encoding is performed according to the target cultural and creative element information. Text encoding and visual feature extraction are performed based on the target cultural and creative element information, and the two are integrated. This is the core step to achieve high-quality generation of cultural and creative patterns. The purpose of text encoding is to convert unstructured or semi-structured text information such as "element type, style characteristics, cultural background and semantic labels" into semantic vectors that can be understood by artificial intelligence models. Visual feature extraction aims to provide reference style, line characteristics, color style and other art elements for cultural and creative pattern generation, ensuring that the output image visually meets cultural and style expectations; By converting structured semantic information and image style perception information into a unified deep representation, the image generation model's ability to understand content, style, and cultural context can be significantly improved, thereby generating cultural and creative patterns that meet expectations.
[0024] S300 performs visualization based on the fusion results, takes the multimodal embedding vector as a conditional input, guides the model to generate an image that matches the input semantics, detects the image and the target cultural and creative element information, and determines the consistency of image generation.
[0025] In this step, visualization processing is performed based on the fusion results. The fused multimodal embedding vector is used as a generation condition to guide the image generation model to generate a pattern image that is highly matched with the target cultural and creative element information. The generated image is then judged for consistency through a subsequent detection and matching mechanism to ensure the restoration and expression accuracy of the semantic content and style features. To ensure that the generated image is consistent with the original target cultural and creative element information, the image is automatically detected and analyzed in terms of content and style. An object detection algorithm (such as YOLOv8 and DETR) is used to identify whether the image contains key semantic elements. A consistency assessment is performed based on the detection results to determine whether the generated image meets the dual requirements of semantics and style of the target cultural and creative elements.
[0026] S400, comparing the image with the target cultural and creative element, obtaining a scoring result based on the comparison result combined with a scoring mechanism, comparing the scoring result with a scoring threshold, obtaining an aesthetic preference based on the comparison result, and modifying the image based on the aesthetic preference if it is determined to be unqualified.
[0027] In this step, the image is compared with the target cultural and creative elements. After the image is generated, it needs to be compared with the target cultural and creative elements, and the comparison score result is obtained according to the scoring mechanism. By comparing the score result with the preset score threshold, the system can determine whether the image meets the design requirements; If the score is lower than the threshold, it is judged as an unqualified image. Then the image is automatically modified and regenerated according to the aesthetic preference information modeled by the user or the system, so as to improve the image quality, fit and artistic expression; It not only has the ability to evaluate the quality of cultural and creative pattern generation, but can also automatically repair and regenerate images by adapting to aesthetic preferences, so that the generation results are highly consistent in terms of cultural expression, visual style and user aesthetics, meeting the actual application needs of automatic generation of high-quality cultural and creative images.
[0028] like Figure 2 As shown, as a preferred embodiment of the present invention, the steps of encoding text according to the target cultural and creative element information, extracting visual features according to the target cultural and creative element information, and fusing the text encoding and visual features specifically include: S201, performing text encoding according to target cultural and creative element information, wherein the text encoding is to convert the target cultural and creative element information text information into a vector representation of a fixed length.
[0029] In this step, text encoding is performed according to the target cultural and creative element information. Text encoding is the basic step in building a semantic-driven image generation model. By converting unstructured cultural and creative element text information into a fixed-length vector representation, the model can understand and process complex semantic content in a numerical form, thereby driving the image generation module to accurately express cultural imagery and visual style. Perform text preprocessing on the target cultural and creative element information and use a language model for embedding modeling. The preferred text encoding model can be a pre-trained language model based on the Transformer structure (such as BERT, RoBERTa, CLIP text encoder), which can capture the deep semantic associations and contextual dependencies between words in the text; The input text is feature extracted through a multi-layer Transformer structure. The encoder model obtains the global context semantics through the self-attention mechanism and finally outputs a fixed-length semantic vector to represent the overall semantic content of the cultural and creative element information. For example, if the CLIP text encoder is used, after encoding the above text, a vector representation with a dimension of 512 may be obtained, such as: [0.214, -0.132, 0.057, ..., 0.098] (a total of 512 dimensions); In this vector, each dimension represents a certain semantic feature extracted by the encoding model from the input text. For example, implicit features such as "flying posture", "sense of history", "female image", "religious meaning", and "Oriental charm" will find corresponding dimension weights in the vector space.
[0030] S202, performing visual feature extraction based on target cultural and creative element information, wherein the visual feature extraction extracts image feature vectors through a visual encoder.
[0031] In this step, visual feature extraction is performed based on the target cultural and creative element information, which is a key step to achieve style consistency and aesthetic control. By introducing a visual encoder to extract image feature vectors, the visual information such as style, composition, color and pattern contained in the reference image can be structured and converted into a computable feature expression, thereby guiding the image generation model to restore a specific visual style and artistic expression.
[0032] The core task of visual feature extraction is to use the trained visual encoder model to process the reference image that matches the target cultural and creative element information and output a fixed-dimensional image feature vector. This image feature vector can be fused with the text semantic vector and input into the generation model to achieve multi-modal information-driven pattern generation; The image is normalized and resized (such as scaling to 224×224), the image is divided into small blocks (patches) or convolution processing is performed, and a multi-layer visual Transformer or CNN extracts the image's style features, color distribution, composition information, local patterns, etc., and outputs a fixed-length vector (such as 512 dimensions or 1024 dimensions) to represent the overall visual expression of the image.
[0033] S203, fusing the text code and the visual features to obtain a fusion result, wherein the fusion result is obtained by projecting the text and the visual features into the same vector space and then adding them together.
[0034] In this step, the text encoding and visual features are fused to achieve a unified expression of semantic information and visual style information, so that the image generation model can understand both "what to generate" and "how to generate", that is, the fusion expression of content and style; The fusion method is: projecting the text encoding and visual features into the same vector space and then adding them to obtain the final fusion result, which is used as a priori condition or prompt input of the image generation model. The text encoding and visual features are the same in dimension, but they come from different modalities. Direct addition may cause semantic inconsistency and numerical scale imbalance. Before additive fusion, the two need to be mapped to the same feature space through a projection function.
[0035] like Figure 3 As shown, as a preferred embodiment of the present invention, the steps of performing visualization processing according to the fusion result, taking the multimodal embedding vector as a conditional input, guiding the model to generate an image matching the input semantics, detecting the image and the target cultural and creative element information, and judging the consistency of the image generation specifically include: S301, performing visualization processing according to the fusion result, wherein the visualization processing includes model selection and image generation, and the model selection includes a diffusion model and a generative adversarial network.
[0036] In this step, visualization is performed based on the fusion results. Visualization is the key stage for converting the fused multimodal vector (fusion of semantic and style information) into a visual image. Visualization processing includes model selection and image generation, among which model selection is crucial. It is necessary to select a suitable image generation architecture based on the characteristics, detail requirements and style complexity of the cultural and creative patterns, including diffusion models and generative adversarial networks (GANs). The visualization processing process selects a suitable architecture based on the different advantages of diffusion models or generative adversarial networks, with fusion vectors as the driving core, to achieve semantically accurate and stylish cultural and creative pattern generation, providing a high-quality image foundation for subsequent consistency detection and cultural and creative applications.
[0037] S302, taking the multimodal embedding vector as a conditional input, guiding the model to generate an image that matches the input semantics, and detecting the image and the target cultural and creative element information.
[0038] In this step, the multimodal embedding vector is used as a conditional input. This is a key step in building a "precise control of semantics and style" generation mechanism. The embedding vector combines the text semantic content and visual style features, and can be used as a conditional input to guide the image generation model to output an image that matches the target cultural and creative element information. In order to ensure the reliability and consistency of the generated results, the output image needs to be tested to determine whether it accurately restores the expected cultural and creative element features. Use image content recognition models (such as YOLOv8, DETR, CLIP visual embedding matching, etc.) to detect whether the image accurately contains the specified semantic elements. Use image style recognition networks or feature space comparison to detect the style of the generated image and determine whether it meets the target style.
[0039] S303, determining consistency of image generation, where the consistency includes image fusion, style migration, and semantic consistency.
[0040] In this step, the consistency of image generation is judged to ensure that the generated image matches the target cultural and creative element information, which is a key quality control link. The consistency not only focuses on whether the content contains the target semantic elements, but also includes the naturalness of image fusion, the accuracy of style transfer, and the integrity of semantic expression; Through a systematic analysis mechanism, the effectiveness of the generated image is comprehensively evaluated from three aspects: image fusion consistency, style transfer consistency, and semantic consistency, ensuring that the cultural and creative pattern has both cultural expression and visual beauty and style unity; Image fusion consistency refers to whether the visual transition between different content elements of the image is natural, whether the composition is coordinated, and whether the local details are unified. It is the basis for measuring whether the generated pattern has complete image expressiveness; style transfer consistency refers to whether the generated image successfully carries and reproduces the artistic style, expression techniques and visual texture specified in the target cultural and creative elements, such as whether it meets the style requirements of "heavy ink and color", "paper-cut style" and "mural texture"; semantic consistency refers to whether the image fully expresses the semantic elements specified in the text information, such as characters, animals, totems, cultural symbols, etc., and whether it has the specified actions, postures or composition meanings.
[0041] like Figure 4 As shown, as a preferred embodiment of the present invention, the steps of comparing the image with the target cultural and creative element, obtaining a scoring result based on the comparison result in combination with a scoring mechanism, comparing the scoring result with a scoring threshold, obtaining an aesthetic preference based on the comparison result, and modifying the image based on the aesthetic preference if the image is judged to be unqualified, specifically include: S401, comparing the image with the target cultural and creative element, obtaining the comparison result, and obtaining a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style.
[0042] In this step, the image is compared with the target cultural and creative elements. Comparing the image with the target cultural and creative elements is a key process to ensure that the generated image meets the user's preset requirements. The comparison result is based on the analysis and extraction of image content and style features, and the scoring mechanism is used to quantitatively evaluate the similarity between the preliminary generated image and the target style. By integrating the dual comparison mechanism of semantic layer and style layer, the scoring system can clearly indicate whether the image truly restores the specified cultural theme, pattern style and visual expression; the scoring mechanism is based on the above comparison results, and outputs a continuous score through preset weight rules and standardized evaluation models to quantify the similarity between the preliminary generated image and the target cultural and creative style. The comparison scoring mechanism not only supports systematic measurement of the content and style of the generated image, but also provides data basis for subsequent image optimization, so that the cultural and creative pattern generation process has quantifiable, traceable and iterative quality control capabilities.
[0043] S402, obtaining a scoring result based on the comparison result combined with a scoring mechanism, obtaining a scoring threshold, wherein the scoring threshold is a qualified line of the image, and comparing the scoring result with the scoring threshold to obtain a comparison result.
[0044] In this step, the scoring result is obtained based on the comparison result combined with the scoring mechanism, which is the core step for quantifying the degree of match between the image quality and the target cultural and creative element information. By comparing the scoring result with the preset scoring threshold, the system can clearly determine whether the image is qualified, and then decide whether to output, optimize or regenerate; The scoring threshold is the qualified line of image quality, which is used to automatically control the output threshold of the image generation process to ensure that the cultural expression and aesthetic style of the final image meet the actual usage requirements. The scoring threshold is set according to task requirements, user tolerance or target application scenarios. The threshold represents the quality standard line for the image to transform from "experimental" to "usability". If the scoring result is less than the scoring threshold, the image is judged to be qualified. If the scoring result is greater than the scoring threshold, the image is judged to be unqualified.
[0045] S403, obtaining aesthetic preferences according to the comparison results, wherein the aesthetic preferences include color tone preferences, graphic complexity preferences, and element composition ratio preferences. If the image is determined to be unqualified, the image is modified according to the aesthetic preferences.
[0046] In this step, obtaining aesthetic preferences based on the comparison results is a key step to improve the personalization and aesthetic consistency of image generation. By analyzing the deviations of unqualified images in terms of hue, complexity and composition, combined with the preset or learned aesthetic preference model, the image generation conditions can be automatically adjusted or the image can be directly modified, thereby outputting a pattern image that meets the visual preferences of the user or target audience; Aesthetic preferences can be obtained through explicit user settings and analysis of user historical behavior. Users can preset their favorite styles, or determine their aesthetic preferences based on the images they have selected, downloaded, and saved in the past by counting the color distribution, number of pixels, composition, etc. If it is judged as unqualified, the image is modified according to aesthetic preferences, such as adjusting the generation configuration, controlling the number of output pixels, and if necessary, only modifying the original image composition area or background part, retaining the qualified part.
[0047] like Figure 5 As shown, an automatic generation system of cultural and creative images based on artificial intelligence is provided in an embodiment of the present invention, and the system includes: The cultural and creative element information module 100 is used to obtain target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background and semantic tags.
[0048] In this system, the cultural and creative element information module 100 obtains the target cultural and creative element information. Obtaining the target cultural and creative element information is a key pre-step in the automatic generation method of cultural and creative images. It is to comprehensively analyze and structure the cultural elements associated with user needs or target scenes, and provide semantic basis and style guidance for subsequent image generation. The target cultural and creative element information includes four dimensions: element type, style characteristics, cultural background and semantic labels, which are respectively expanded as follows: Element type refers to the core visual elements that constitute the cultural and creative image, which is the main content of the image generation process. Common element types include natural elements (such as landscapes, flowers, plants, and animals), artifact elements (such as ceramics, lacquerware, and clothing), symbol elements (such as totems, patterns, and characters), character elements (such as historical figures and mythological images), and architectural elements (such as ancient buildings and religious sites). Style features refer to the expression techniques and aesthetic styles of the cultural and creative elements in visual presentation, which directly affect the visual direction of image generation. Style features can be derived from user settings or automatically inferred by the system based on existing image samples or historical preferences. Cultural background refers to the historical origin, regional culture, folk beliefs or national characteristics of the cultural and creative elements. It provides semantic context and cultural symbol references, and is the basis for ensuring that the cultural and creative patterns have cultural depth and accuracy. Semantic tags are highly summarized and structured annotations of the above three dimensions of information, which facilitates the semantic vectorization processing of artificial intelligence models. Semantic tags are usually composed of structured tag words, which can be combined with natural language processing technology to automatically extract semantic keywords in texts, or standardized matching by manually annotated knowledge bases.
[0049] The coding feature module 200 is used to perform text coding according to the target cultural and creative element information, extract visual features according to the target cultural and creative element information, and fuse the text coding and visual features.
[0050] In this system, the encoding feature module 200 performs text encoding according to the target cultural and creative element information. Based on the target cultural and creative element information, text encoding and visual feature extraction are performed and the two are integrated. This is the core step to achieve high-quality generation of cultural and creative patterns. The purpose of text encoding is to convert unstructured or semi-structured text information such as "element type, style characteristics, cultural background and semantic labels" into semantic vectors that can be understood by artificial intelligence models. Visual feature extraction aims to provide reference style, line characteristics, color style and other art elements for cultural and creative pattern generation, ensuring that the output image visually meets cultural and style expectations. By converting structured semantic information and image style perception information into a unified deep representation, the image generation model's ability to understand content, style, and cultural context can be significantly improved, thereby generating cultural and creative patterns that meet expectations.
[0051] The visualization processing module 300 is used to perform visualization processing based on the fusion results, take the multimodal embedding vector as a conditional input, guide the model to generate an image that matches the input semantics, detect the image and the target cultural and creative element information, and judge the consistency of image generation.
[0052] In this system, the visualization processing module 300 performs visualization processing according to the fusion result, and uses the fused multimodal embedding vector as a generation condition to guide the image generation model to generate a pattern image that is highly matched with the target cultural and creative element information, and through the subsequent detection and matching mechanism, the consistency of the generated image is judged to ensure the restoration degree and expression accuracy of the semantic content and style features; To ensure that the generated image is consistent with the original target cultural and creative element information, the image is automatically detected and analyzed in terms of content and style. An object detection algorithm (such as YOLOv8 and DETR) is used to identify whether the image contains key semantic elements. A consistency assessment is performed based on the detection results to determine whether the generated image meets the dual requirements of semantics and style of the target cultural and creative elements.
[0053] The aesthetic preference module 400 is used to compare the image with the target cultural and creative element, obtain a scoring result based on the comparison result combined with the scoring mechanism, compare the scoring result with the scoring threshold, and obtain the aesthetic preference based on the comparison result. If it is judged to be unqualified, the image is modified according to the aesthetic preference.
[0054] In this system, the aesthetic preference module 400 compares the image with the target cultural and creative elements. After the image is generated, it needs to be compared with the target cultural and creative elements, and the comparison score result is obtained according to the scoring mechanism. By comparing the score result with the preset score threshold, the system can determine whether the image meets the design requirements; If the score is lower than the threshold, it is judged as an unqualified image. Then the image is automatically modified and regenerated according to the aesthetic preference information modeled by the user or the system, so as to improve the image quality, fit and artistic expression; It not only has the ability to evaluate the quality of cultural and creative pattern generation, but can also automatically repair and regenerate images by adapting to aesthetic preferences, so that the generation results are highly consistent in terms of cultural expression, visual style and user aesthetics, meeting the actual application needs of automatic generation of high-quality cultural and creative images.
[0055] like Figure 6 As shown, as a preferred embodiment of the present invention, the coding feature module 200 includes: The text encoding unit 201 is used to perform text encoding according to the target cultural and creative element information, and the text encoding is to convert the target cultural and creative element information into a vector representation of a fixed length.
[0056] In this module, the text encoding unit 201 performs text encoding according to the target cultural and creative element information. Performing text encoding according to the target cultural and creative element information is a basic step in building a semantic-driven image generation model. By converting the unstructured cultural and creative element text information into a fixed-length vector representation, the model can understand and process complex semantic content in a numerical form, thereby driving the image generation module to accurately express cultural images and visual styles. Perform text preprocessing on the target cultural and creative element information and use a language model for embedding modeling. The preferred text encoding model can be a pre-trained language model based on the Transformer structure (such as BERT, RoBERTa, CLIP text encoder), which can capture the deep semantic associations and contextual dependencies between words in the text; The input text is feature extracted through a multi-layer Transformer structure. The encoder model obtains the global context semantics through the self-attention mechanism and finally outputs a fixed-length semantic vector to represent the overall semantic content of the cultural and creative element information. For example, if the CLIP text encoder is used, after encoding the above text, a vector representation with a dimension of 512 may be obtained, such as: [0.214, -0.132, 0.057, ..., 0.098] (a total of 512 dimensions); In this vector, each dimension represents a certain semantic feature extracted by the encoding model from the input text. For example, implicit features such as "flying posture", "sense of history", "female image", "religious meaning", and "Oriental charm" will find corresponding dimension weights in the vector space.
[0057] The visual feature extraction unit 202 is used to extract visual features according to the target cultural and creative element information. The visual feature extraction extracts image feature vectors through a visual encoder.
[0058] In this module, the visual feature extraction unit 202 extracts visual features based on the target cultural and creative element information. Extracting visual features based on the target cultural and creative element information is a key step in achieving style consistency and aesthetic control. By introducing a visual encoder to extract image feature vectors, the visual information such as style, composition, color and pattern contained in the reference image can be structured and converted into a computable feature expression, thereby guiding the image generation model to restore a specific visual style and artistic expression.
[0059] The core task of visual feature extraction is to use the trained visual encoder model to process the reference image that matches the target cultural and creative element information and output a fixed-dimensional image feature vector. This image feature vector can be fused with the text semantic vector and input into the generation model to achieve multi-modal information-driven pattern generation; The image is normalized and resized (such as scaling to 224×224), the image is divided into small blocks (patches) or convolution processing is performed, and a multi-layer visual Transformer or CNN extracts the image's style features, color distribution, composition information, local patterns, etc., and outputs a fixed-length vector (such as 512 dimensions or 1024 dimensions) to represent the overall visual expression of the image.
[0060] The fusion unit 203 is used to fuse the text code and the visual feature to obtain a fusion result, where the fusion result is the sum of the text and the visual feature after projecting them into the same vector space.
[0061] In this module, the fusion unit 203 fuses the text encoding and the visual features to achieve a unified expression of semantic information and visual style information, so that the image generation model can understand "what to generate" and "how to generate" at the same time, that is, the fusion expression of content and style; The fusion method is: projecting the text encoding and visual features into the same vector space and then adding them to obtain the final fusion result, which is used as a priori condition or prompt input of the image generation model. The text encoding and visual features are the same in dimension, but they come from different modalities. Direct addition may cause semantic inconsistency and numerical scale imbalance. Before additive fusion, the two need to be mapped to the same feature space through a projection function.
[0062] like Figure 7 As shown, as a preferred embodiment of the present invention, the visualization processing module 300 includes: The visualization processing unit 301 is used to perform visualization processing according to the fusion result, wherein the visualization processing includes model selection and image generation, and the model selection includes a diffusion model and a generative adversarial network.
[0063] In this module, the visualization processing unit 301 performs visualization processing according to the fusion result. Performing visualization processing according to the fusion result is a key stage for converting the fused multimodal vector (fusion semantics and style information) into a visual image; Visualization processing includes model selection and image generation, among which model selection is crucial. It is necessary to select a suitable image generation architecture based on the characteristics, detail requirements and style complexity of the cultural and creative patterns, including diffusion models and generative adversarial networks (GANs). The visualization processing process selects a suitable architecture based on the different advantages of diffusion models or generative adversarial networks, with fusion vectors as the driving core, to achieve semantically accurate and stylish cultural and creative pattern generation, providing a high-quality image foundation for subsequent consistency detection and cultural and creative applications.
[0064] The image detection unit 302 is used to take the multimodal embedding vector as a conditional input, guide the model to generate an image that matches the input semantics, and detect the image and the target cultural and creative element information.
[0065] In this module, the image detection unit 302 uses the multimodal embedding vector as a conditional input. Using the multimodal embedding vector as a conditional input is a key step in building a "precise control of semantics and style" generation mechanism. The embedding vector combines the text semantic content and visual style features, and can be used as a conditional input to guide the image generation model to output an image that matches the target cultural and creative element information. In order to ensure the reliability and consistency of the generated results, the output image needs to be detected to determine whether it accurately restores the expected cultural and creative element features; Use image content recognition models (such as YOLOv8, DETR, CLIP visual embedding matching, etc.) to detect whether the image accurately contains the specified semantic elements. Use image style recognition networks or feature space comparison to detect the style of the generated image and determine whether it meets the target style.
[0066] The consistency determination unit 303 is used to determine the consistency of image generation, where the consistency includes image fusion, style migration and semantic consistency.
[0067] In this module, the consistency determination unit 303 determines the consistency of image generation, which is a key quality control link to ensure that the generated image matches the target cultural and creative element information. The consistency not only focuses on whether the content contains the target semantic elements, but also includes the naturalness of image fusion, the accuracy of style transfer, and the integrity of semantic expression; Through a systematic analysis mechanism, the effectiveness of the generated image is comprehensively evaluated from three aspects: image fusion consistency, style transfer consistency, and semantic consistency, ensuring that the cultural and creative pattern has both cultural expression and visual beauty and style unity; Image fusion consistency refers to whether the visual transition between different content elements of the image is natural, whether the composition is coordinated, and whether the local details are unified. It is the basis for measuring whether the generated pattern has complete image expressiveness; style transfer consistency refers to whether the generated image successfully carries and reproduces the artistic style, expression techniques and visual texture specified in the target cultural and creative elements, such as whether it meets the style requirements of "heavy ink and color", "paper-cut style" and "mural texture"; semantic consistency refers to whether the image fully expresses the semantic elements specified in the text information, such as characters, animals, totems, cultural symbols, etc., and whether it has the specified actions, postures or composition meanings.
[0068] like Figure 8 As shown, as a preferred embodiment of the present invention, the aesthetic preference module 400 includes: The similarity determination unit 401 is used to compare the image with the target cultural and creative element, obtain the comparison result, and obtain a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style.
[0069] In this module, the similarity determination unit 401 compares the image with the target cultural and creative element. Comparing the image with the target cultural and creative element is a key process to ensure that the generated image meets the preset requirements of the user. The comparison result is based on the analysis and extraction of image content and style features, and the scoring mechanism is used to quantitatively evaluate the similarity between the preliminary generated image and the target style. By integrating the dual comparison mechanism of semantic layer and style layer, the scoring system can clearly indicate whether the image truly restores the specified cultural theme, pattern style and visual expression; the scoring mechanism is based on the above comparison results, and outputs a continuous score through preset weight rules and standardized evaluation models to quantify the similarity between the preliminary generated image and the target cultural and creative style. The comparison scoring mechanism not only supports systematic measurement of the content and style of the generated image, but also provides data basis for subsequent image optimization, so that the cultural and creative pattern generation process has quantifiable, traceable and iterative quality control capabilities.
[0070] The scoring unit 402 is used to obtain a scoring result based on the comparison result in combination with the scoring mechanism, obtain a scoring threshold, and the scoring threshold is the qualified line of the image. The scoring result is compared with the scoring threshold to obtain the comparison result.
[0071] In this module, the scoring unit 402 obtains the scoring result based on the comparison result combined with the scoring mechanism, which is the core step for quantifying the degree of match between the image quality and the target cultural and creative element information. By comparing the scoring result with the preset scoring threshold, the system can clearly determine whether the image is qualified, and then decide whether to output, optimize or regenerate; The scoring threshold is the qualified line of image quality, which is used to automatically control the output threshold of the image generation process to ensure that the cultural expression and aesthetic style of the final image meet the actual usage requirements. The scoring threshold is set according to task requirements, user tolerance or target application scenarios. The threshold represents the quality standard line for the image to transform from "experimental" to "usability". If the scoring result is less than the scoring threshold, the image is judged to be qualified. If the scoring result is greater than the scoring threshold, the image is judged to be unqualified.
[0072] The aesthetic preference unit 403 is used to obtain aesthetic preferences according to the comparison results. The aesthetic preferences include color tone preference, graphic complexity preference and element composition ratio preference. If the image is judged to be unqualified, the image is modified according to the aesthetic preference.
[0073] In this module, the aesthetic preference unit 403 obtains the aesthetic preference according to the comparison result, which is a key step to improve the personalization and aesthetic consistency of image generation. By analyzing the deviations of unqualified images in terms of hue, complexity and composition, combined with the preset or learned aesthetic preference model, the image generation conditions can be automatically adjusted or the image can be directly modified, so as to output a pattern image that meets the visual preferences of the user or the target audience; Aesthetic preferences can be obtained through explicit user settings and analysis of user historical behavior. Users can preset their favorite styles, or determine their aesthetic preferences based on the images they have selected, downloaded, and saved in the past by counting the color distribution, number of pixels, composition, etc. If it is judged as unqualified, the image is modified according to aesthetic preferences, such as adjusting the generation configuration, controlling the number of output pixels, and if necessary, only modifying the original image composition area or background part, retaining the qualified part.
[0074] In one embodiment, a computer device is provided, the computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented: Acquire target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background and semantic tags; Perform text encoding based on the target cultural and creative element information, extract visual features based on the target cultural and creative element information, and fuse the text encoding and visual features; Visualize the fusion results and use the multimodal embedding vector as a conditional input to guide the model to generate images that match the input semantics, detect the image and the target cultural and creative element information, and judge the consistency of image generation; Compare the image with the target cultural and creative elements, obtain the scoring result based on the comparison result and the scoring mechanism, compare the scoring result with the scoring threshold, obtain the aesthetic preference based on the comparison result, and if it is judged as unqualified, modify the image based on the aesthetic preference.
[0075] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor performs the following steps: Acquire target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background and semantic tags; Perform text encoding based on the target cultural and creative element information, extract visual features based on the target cultural and creative element information, and fuse the text encoding and visual features; Visualize the fusion results and use the multimodal embedding vector as a conditional input to guide the model to generate images that match the input semantics, detect the image and the target cultural and creative element information, and judge the consistency of image generation; Compare the image with the target cultural and creative elements, obtain the scoring result based on the comparison result and the scoring mechanism, compare the scoring result with the scoring threshold, obtain the aesthetic preference based on the comparison result, and if it is judged as unqualified, modify the image based on the aesthetic preference.
[0076] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps. Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0077] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0078] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
[0079] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for automatically generating cultural and creative images based on artificial intelligence, characterized in that: The method comprises: Acquire target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background and semantic tags; Perform text encoding based on the target cultural and creative element information, extract visual features based on the target cultural and creative element information, and fuse the text encoding and visual features; Visualize the fusion results and use the multimodal embedding vector as a conditional input to guide the model to generate images that match the input semantics, detect the image and the target cultural and creative element information, and judge the consistency of image generation; Compare the image with the target cultural and creative elements, obtain the scoring result based on the comparison result and the scoring mechanism, compare the scoring result with the scoring threshold, obtain the aesthetic preference based on the comparison result, and if it is judged as unqualified, modify the image based on the aesthetic preference.
2. The method for automatically generating cultural and creative images based on artificial intelligence according to claim 1, characterized in that: The steps of encoding text according to the target cultural and creative element information, extracting visual features according to the target cultural and creative element information, and fusing the text encoding and the visual features specifically include: Performing text encoding according to the target cultural and creative element information, wherein the text encoding is to convert the target cultural and creative element information text information into a vector representation of a fixed length; Visual feature extraction is performed based on the target cultural and creative element information. The visual feature extraction extracts the image feature vector through a visual encoder; The text encoding and the visual features are fused to obtain a fusion result, wherein the fusion result is obtained by projecting the text and the visual features into the same vector space and then adding them together.
3. The method for automatically generating cultural and creative images based on artificial intelligence according to claim 1, characterized in that: The steps of performing visualization processing according to the fusion result, taking the multimodal embedding vector as a conditional input, guiding the model to generate an image matching the input semantics, detecting the image and the target cultural and creative element information, and judging the consistency of the image generation specifically include: Performing visualization processing according to the fusion result, wherein the visualization processing includes model selection and image generation, and the model selection includes a diffusion model and a generative adversarial network; The multimodal embedding vector is used as a conditional input to guide the model to generate images that match the input semantics and detect the information between the image and the target cultural and creative elements; The consistency of image generation is determined, including image fusion, style transfer, and semantic consistency.
4. The method for automatically generating cultural and creative images based on artificial intelligence according to claim 1, characterized in that: The steps of comparing the image with the target cultural and creative element, obtaining a scoring result based on the comparison result in combination with a scoring mechanism, comparing the scoring result with a scoring threshold, obtaining an aesthetic preference based on the comparison result, and modifying the image based on the aesthetic preference if the image is determined to be unqualified, specifically include: Comparing the image with the target cultural and creative element, obtaining a comparison result, and obtaining a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style; A scoring result is obtained according to the comparison result combined with the scoring mechanism, and a scoring threshold is obtained, where the scoring threshold is the qualified line of the image, and the scoring result is compared with the scoring threshold to obtain a comparison result; Aesthetic preferences are obtained based on the comparison results, including color tone preferences, graphic complexity preferences, and element composition ratio preferences. If the image is determined to be unqualified, the image is modified based on the aesthetic preferences.
5. The method for automatically generating cultural and creative images based on artificial intelligence according to claim 1, characterized in that: The element types include traditional painting, sculpture and calligraphy, the style features include ink painting style and meticulous painting style, the cultural background includes historical periods and regional characteristics, and the semantic tags are descriptive keywords for the elements.
6. An artificial intelligence-based automatic generation system for cultural and creative images, characterized in that: The system comprises: A cultural and creative element information module is used to obtain target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background and semantic tags; The coding feature module performs text coding based on the target cultural and creative element information, extracts visual features based on the target cultural and creative element information, and integrates the text coding and visual features; The visualization processing module performs visualization processing based on the fusion results, takes the multimodal embedding vector as a conditional input, guides the model to generate images that match the input semantics, detects the image and the target cultural and creative element information, and determines the consistency of image generation; The aesthetic preference module compares the image with the target cultural and creative elements, obtains the scoring result based on the comparison result and the scoring mechanism, compares the scoring result with the scoring threshold, obtains the aesthetic preference based on the comparison result, and modifies the image according to the aesthetic preference if it is judged to be unqualified.
7. The system for automatically generating cultural and creative images based on artificial intelligence according to claim 6 is characterized in that: The coding feature module comprises: A text encoding unit, performing text encoding according to the target cultural and creative element information, wherein the text encoding is to convert the target cultural and creative element information into a vector representation of a fixed length; A visual feature extraction unit extracts visual features based on target cultural and creative element information. The visual feature extraction extracts image feature vectors through a visual encoder. The fusion unit fuses the text encoding and the visual features to obtain a fusion result, wherein the fusion result is obtained by projecting the text and the visual features into the same vector space and then adding them together.
8. The system for automatically generating cultural and creative images based on artificial intelligence according to claim 7 is characterized in that: The visualization processing module comprises: A visualization processing unit, performing visualization processing according to the fusion result, wherein the visualization processing includes model selection and image generation, and the model selection includes a diffusion model and a generative adversarial network; The image detection unit takes the multimodal embedding vector as a conditional input, guides the model to generate an image that matches the input semantics, and detects the image and the target cultural and creative element information; The consistency determination unit determines the consistency of image generation, wherein the consistency includes image fusion, style migration and semantic consistency.
9. The system for automatically generating cultural and creative images based on artificial intelligence according to claim 8 is characterized in that: The aesthetic preference module includes: A similarity determination unit compares the image with the target cultural and creative element, obtains the comparison result, and obtains a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style; A scoring unit, which obtains a scoring result according to the comparison result and the scoring mechanism, obtains a scoring threshold, wherein the scoring threshold is a qualified line of the image, and compares the scoring result with the scoring threshold to obtain a comparison result; The aesthetic preference unit obtains the aesthetic preference according to the comparison result, wherein the aesthetic preference includes the color tone preference, the graphic complexity preference and the element composition ratio preference. If the image is judged to be unqualified, the image is modified according to the aesthetic preference.
10. The system for automatically generating cultural and creative images based on artificial intelligence according to claim 9, characterized in that: The element types include traditional painting, sculpture and calligraphy, the style features include ink painting style and meticulous painting style, the cultural background includes historical periods and regional characteristics, and the semantic tags are descriptive keywords for the elements.
Citation Information
Patent Citations
Intelligent design system based on AI
CN117407556A
Multi-modal visual analysis processing system and method based on artificial intelligence
CN117876937A
Architecture graph rationality analysis and judgment method based on multi-modal feature fusion
CN119169643A
Commodity information processing and querying method and system
CN119377433A
Text-driven automatic image generation method and system
CN119444935A
Cited By
Image processing method and device, equipment and storage medium
CN120451514A
Method and device for generating costume design scheme embedded with cultural decoration symbols and medium
CN121188852A
Artificial intelligence-based automatic calibration method and system for creative and cultural patterns
CN121280220A
Emotion data graph generation method based on natural language understanding
CN121706786A