A method and system for automatically generating cultural and creative images based on artificial intelligence
By obtaining cultural and creative elements information for text encoding and visual feature extraction, integrating images to generate and scoring correction, the problem of not being able to automatically score and optimize in the generation of cultural and creative pictures is solved, and the generation of high-quality cultural and creative pictures is achieved.
Patent Information
- Application Number
- CN202510594792.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing automatic generation methods of cultural and creative pictures cannot effectively combine the target cultural and creative elements information, and the generated images cannot be automatically scored and adaptively adjusted and optimized according to the scoring results.
By obtaining the target cultural and creative elements information, text encoding and visual feature extraction, fusing them and inputting them as conditions to generate images, and using the scoring mechanism to judge the consistency and aesthetic preferences of the images, and automatically modify the unqualified images.
Automatic scoring and adaptive optimization of the generated images during the generation of cultural and creative images is realized, ensuring that the generated images are highly consistent with the target cultural and creative elements information and meeting the requirements of cultural depth and visual style.
Smart Images

Figure CN120107419B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cultural and creative production technology, and in particular relates to an artificial intelligence-based automatic generation method and system for cultural and creative images. Background Art
[0002] With the continuous development of artificial intelligence technology, computers are increasingly capable of automatically generating images, gaining widespread attention in fields such as cultural and creative industries, digital collection design, and the preservation of intangible cultural heritage patterns. In particular, image generation systems based on generative adversarial networks and diffusion models can generate pattern images with specific semantic and stylistic characteristics based on input text or image prompts, showing promising application prospects in improving design efficiency and expanding the scope of visual expression.
[0003] Currently, there are a variety of existing image generation methods based on text or multimodal input. Some methods use pre-trained language models to encode input text and input the encoded vector into an image generation model to achieve text-to-image conversion. Other methods extract style feature vectors from image samples and combine them with text information to guide the model to generate images that combine semantics and style. These methods have achieved certain technical progress in general image generation, image style transfer, and illustration synthesis.
[0004] Currently, there is still a lack of an intelligent cultural and creative image generation method that can generate images based on the target cultural and creative element information, automatically extract aesthetic preferences based on the comparison results, and adaptively adjust and optimize unqualified images. Summary of the Invention
[0005] The purpose of the embodiment of the present invention is to provide an automatic generation method of cultural and creative images based on artificial intelligence, aiming to solve the problems raised in the third part of the background technology.
[0006] The embodiment of the present invention is implemented as follows: a method for automatically generating cultural and creative images based on artificial intelligence, the method comprising:
[0007] Obtaining target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background, and semantic tags;
[0008] Perform text encoding based on target cultural and creative element information, extract visual features based on target cultural and creative element information, and fuse text encoding and visual features;
[0009] Visualize the fusion results and use the multimodal embedding vector as a conditional input to guide the model to generate an image that matches the input semantics, detect the image and the target cultural and creative elements, and determine the consistency of the generated image.
[0010] Compare the image with the target cultural and creative elements, obtain the scoring result based on the comparison result and the scoring mechanism, compare the scoring result with the scoring threshold, and obtain the aesthetic preference based on the comparison result. If it is judged to be unqualified, modify the image according to the aesthetic preference.
[0011] Preferably, the steps of encoding text according to the target cultural and creative element information, extracting visual features according to the target cultural and creative element information, and fusing the text encoding and visual features specifically include:
[0012] Performing text encoding according to the target cultural and creative element information, wherein the text encoding is to convert the target cultural and creative element information text information into a vector representation of a fixed length;
[0013] Visual feature extraction is performed based on the target cultural and creative element information. The visual feature extraction uses a visual encoder to extract image feature vectors.
[0014] The text encoding and the visual features are fused to obtain a fusion result, wherein the fusion result is obtained by projecting the text and the visual features into the same vector space and then adding them together.
[0015] Preferably, the steps of performing visualization processing based on the fusion results, using the multimodal embedding vector as a conditional input, guiding the model to generate an image that matches the input semantics, detecting the image and the target cultural and creative element information, and judging the consistency of the image generation specifically include:
[0016] Performing visualization processing based on the fusion results, wherein the visualization processing includes model selection and image generation, and the model selection includes a diffusion model and a generative adversarial network;
[0017] The multimodal embedding vector is used as a conditional input to guide the model to generate images that match the input semantics and detect information between the image and the target cultural and creative elements;
[0018] Determine the consistency of image generation, including image fusion, style transfer, and semantic consistency.
[0019] Preferably, the steps of comparing the image with the target cultural and creative element, obtaining a scoring result based on the comparison result in combination with a scoring mechanism, comparing the scoring result with a scoring threshold, obtaining an aesthetic preference based on the comparison result, and modifying the image based on the aesthetic preference if the image is determined to be unqualified, specifically include:
[0020] Comparing the image with the target cultural and creative element, obtaining a comparison result, and obtaining a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style;
[0021] A scoring result is obtained based on the comparison result combined with the scoring mechanism, and a scoring threshold is obtained. The scoring threshold is the passing line of the image, and the scoring result is compared with the scoring threshold to obtain a comparison result;
[0022] Aesthetic preferences are obtained based on the comparison results, including color tone preferences, graphic complexity preferences, and element composition ratio preferences. If the image is determined to be unqualified, the image is modified based on the aesthetic preferences.
[0023] Preferably, the element types include traditional painting, sculpture and calligraphy, the style features include ink painting style and fine brushwork style, the cultural background includes historical periods and regional characteristics, and the semantic tags are descriptive keywords for the elements.
[0024] Another object of an embodiment of the present invention is to provide an artificial intelligence-based automatic generation system for cultural and creative images, characterized in that the system includes:
[0025] A cultural and creative element information module obtains target cultural and creative element information, including element type, style characteristics, cultural background, and semantic tags;
[0026] The encoding feature module encodes text based on the target cultural and creative element information, extracts visual features based on the target cultural and creative element information, and integrates the text encoding and visual features;
[0027] The visualization processing module performs visualization processing based on the fusion results, uses the multimodal embedding vector as a conditional input, guides the model to generate images that match the input semantics, detects the image and the target cultural and creative elements, and determines the consistency of the image generation;
[0028] The aesthetic preference module compares the image with the target cultural and creative elements, obtains the scoring result based on the comparison result combined with the scoring mechanism, compares the scoring result with the scoring threshold, and obtains the aesthetic preference based on the comparison result. If it is judged to be unqualified, the image will be modified according to the aesthetic preference.
[0029] Preferably, the coding feature module includes:
[0030] A text encoding unit, which performs text encoding according to the target cultural and creative element information, wherein the text encoding is to convert the target cultural and creative element information into a vector representation of a fixed length;
[0031] The visual feature extraction unit extracts visual features based on the target cultural and creative element information. The visual feature extraction unit extracts image feature vectors through a visual encoder.
[0032] The fusion unit fuses the text encoding and the visual features to obtain a fusion result, wherein the fusion result is obtained by projecting the text and the visual features into the same vector space and then adding them together.
[0033] Preferably, the visualization processing module includes:
[0034] A visualization processing unit, performing visualization processing based on the fusion result, wherein the visualization processing includes model selection and image generation, wherein the model selection includes a diffusion model and a generative adversarial network;
[0035] The image detection unit takes the multimodal embedding vector as a conditional input, guides the model to generate an image that matches the input semantics, and detects the image and the target cultural and creative element information;
[0036] The consistency determination unit determines the consistency of image generation, wherein the consistency includes image fusion, style transfer and semantic consistency.
[0037] Preferably, the aesthetic preference module includes:
[0038] a similarity determination unit, comparing the image with the target cultural and creative element, obtaining a comparison result, and obtaining a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style;
[0039] A scoring unit, which obtains a scoring result based on the comparison result and the scoring mechanism, obtains a scoring threshold, wherein the scoring threshold is a passing line of the image, and compares the scoring result with the scoring threshold to obtain a comparison result;
[0040] The aesthetic preference unit obtains aesthetic preferences based on the comparison results, wherein the aesthetic preferences include color tone preferences, graphic complexity preferences, and element composition ratio preferences. If the image is determined to be unqualified, the image is modified according to the aesthetic preferences.
[0041] Preferably, the element types include traditional painting, sculpture and calligraphy, the style features include ink painting style and fine brushwork style, the cultural background includes historical periods and regional characteristics, and the semantic tags are descriptive keywords for the elements.
[0042] An embodiment of the present invention provides an artificial intelligence-based automatic generation method for cultural and creative images, which obtains target cultural and creative element information, performs text encoding based on the target cultural and creative element information, and extracts visual features based on the target cultural and creative element information. The visual feature extraction extracts image feature vectors through a visual encoder, fuses the text encoding and visual features to obtain a fusion result, and performs visualization processing based on the fusion result. Model selection includes a diffusion model and a generative adversarial network, uses a multimodal embedding vector as a conditional input, guides the model to generate an image that matches the input semantics, detects the image and the target cultural and creative element information, judges the consistency of image generation, compares the image and the target cultural and creative element, obtains a comparison result, obtains a scoring mechanism, obtains a scoring result based on the comparison result combined with the scoring mechanism, obtains a scoring threshold, compares the scoring result with the scoring threshold, obtains the comparison result, obtains aesthetic preferences based on the comparison result, and modifies the image according to the aesthetic preference if it is judged to be unqualified. This solves the problem that the generated cultural and creative images cannot be scored and modified according to the scoring results during the existing automatic generation of cultural and creative images. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A flowchart of a method for automatically generating cultural and creative images based on artificial intelligence provided by an embodiment of the present invention;
[0044] Figure 2 A flowchart of the steps of encoding text based on target cultural and creative element information and extracting visual features based on the target cultural and creative element information provided by an embodiment of the present invention;
[0045] Figure 3 A flowchart of the steps of using a multimodal embedding vector as a conditional input, detecting an image and target cultural and creative element information, and determining consistency in image generation, provided by an embodiment of the present invention;
[0046] Figure 4 A flowchart of the steps of comparing the scoring result with the scoring threshold and, if the image is judged to be unqualified, modifying the image based on aesthetic preferences according to an embodiment of the present invention;
[0047] Figure 5 An architectural diagram of an artificial intelligence-based automatic generation system for cultural and creative images provided by an embodiment of the present invention;
[0048] Figure 6 This is an architectural diagram of the coding feature module provided in an embodiment of the present invention;
[0049] Figure 7 An architectural diagram of a visualization processing module provided in an embodiment of the present invention;
[0050] Figure 8 This is an architectural diagram of the aesthetic preference module provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0052] It is understood that the terms "first," "second," etc., used herein may be used to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, a first xx script may be referred to as a second xx script, and similarly, a second xx script may be referred to as a first xx script without departing from the scope of this application.
[0053] like Figure 1 As shown in FIG, an embodiment of the present invention provides an automatic generation method of cultural and creative images based on artificial intelligence, the method comprising:
[0054] S100, obtaining target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background and semantic tags.
[0055] In this step, the target cultural and creative element information is obtained. Obtaining the target cultural and creative element information is a key prerequisite for the automatic generation method of cultural and creative images. It aims to comprehensively analyze and structure the cultural elements associated with user needs or target scenes, providing a semantic basis and style guidance for subsequent image generation. The target cultural and creative element information includes four dimensions: element type, style characteristics, cultural background, and semantic labels, which are expanded as follows:
[0056] Element type refers to the core visual elements that constitute cultural and creative images and is the main content of the image generation process. Common element types include natural elements (such as landscapes, flowers, plants, and animals), artifact elements (such as ceramics, lacquerware, and clothing), symbolic elements (such as totems, patterns, and characters), human elements (such as historical figures and mythological images), and architectural elements (such as ancient buildings and religious sites).
[0057] Stylistic features refer to the expression techniques and aesthetic style of the cultural and creative elements in visual presentation, which directly influence the visual direction of the generated image. Stylistic features can be derived from user settings or automatically inferred by the system based on existing image samples or historical preferences.
[0058] Cultural background refers to the historical origins, regional culture, folk beliefs, or ethnic characteristics of the cultural and creative elements. It provides semantic context and cultural symbolic references, and is the basis for ensuring that cultural and creative patterns have cultural depth and accuracy.
[0059] Semantic tags provide a highly summarized and structured annotation of the three dimensions mentioned above, facilitating semantic vectorization processing by AI models. Semantic tags typically consist of structured tag words and can be combined with natural language processing techniques to automatically extract semantic keywords from text, or standardized matching can be performed using manually annotated knowledge bases.
[0060] S200: perform text encoding based on target cultural and creative element information, extract visual features based on target cultural and creative element information, and fuse the text encoding and visual features.
[0061] In this step, text encoding is performed based on the target cultural and creative element information. Based on the target cultural and creative element information, text encoding and visual feature extraction are performed and the two are integrated. This is the core step to achieve high-quality generation of cultural and creative patterns. The purpose of text encoding is to convert unstructured or semi-structured text information such as "element type, style characteristics, cultural background and semantic labels" into semantic vectors that can be understood by artificial intelligence models. Visual feature extraction aims to provide artistic elements such as reference style, line characteristics, color style, etc. for cultural and creative pattern generation, ensuring that the output image is visually consistent with cultural and style expectations.
[0062] By converting structured semantic information and image style perception information into a unified deep representation, the image generation model's ability to understand content, style, and cultural context can be significantly improved, thereby generating cultural and creative patterns that meet expectations.
[0063] S300 performs visualization processing based on the fusion results, uses the multimodal embedding vector as a conditional input, guides the model to generate an image that matches the input semantics, detects the image and the target cultural and creative element information, and determines the consistency of the image generation.
[0064] In this step, visualization processing is performed based on the fusion results. The fused multimodal embedding vector is used as a generation condition to guide the image generation model to generate pattern images that highly match the target cultural and creative element information. The generated images are then judged for consistency through subsequent detection and matching mechanisms to ensure the restoration and accuracy of semantic content and style features.
[0065] To ensure that the generated image is consistent with the original target cultural and creative element information, the image is automatically detected and analyzed in terms of content and style. Object detection algorithms (such as YOLOv8 and DETR) are used to identify whether the image contains key semantic elements. Based on the detection results, a consistency assessment is performed to determine whether the generated image meets the dual requirements of semantics and style of the target cultural and creative elements.
[0066] S400, compare the image with the target cultural and creative element, obtain a scoring result based on the comparison result combined with the scoring mechanism, compare the scoring result with the scoring threshold, obtain aesthetic preferences based on the comparison result, and if it is judged to be unqualified, modify the image based on the aesthetic preferences.
[0067] In this step, the image is compared with the target cultural and creative elements. After the image is generated, it needs to be compared with the target cultural and creative elements, and the comparison score result is obtained according to the scoring mechanism. By comparing the score result with the preset score threshold, the system can determine whether the image meets the design requirements;
[0068] If the score is lower than the threshold, it is judged as an unqualified image. Then, the image will be automatically modified and regenerated based on the aesthetic preference information modeled by the user or the system, thereby improving the image quality, fit and artistic expression.
[0069] It not only has the ability to evaluate the quality of cultural and creative pattern generation, but can also automatically repair and regenerate images by adapting to aesthetic preferences, so that the generation results are highly consistent in terms of cultural expression, visual style and user aesthetics, meeting the actual application needs of automatic generation of high-quality cultural and creative images.
[0070] like Figure 2 As shown, as a preferred embodiment of the present invention, the steps of encoding text according to target cultural and creative element information, extracting visual features according to target cultural and creative element information, and fusing the text encoding and visual features specifically include:
[0071] S201, performing text encoding according to target cultural and creative element information, wherein the text encoding is to convert the target cultural and creative element information into a vector representation of a fixed length.
[0072] In this step, text encoding is performed based on the target cultural and creative element information. This is a fundamental step in building a semantically driven image generation model. By converting unstructured cultural and creative element text information into a fixed-length vector representation, the model can understand and process complex semantic content in a numerical form, thereby driving the image generation module to accurately express cultural imagery and visual style.
[0073] Perform text preprocessing on the target cultural and creative element information and use a language model for embedding modeling. The preferred text encoding model is a pre-trained language model based on the Transformer structure (such as BERT, RoBERTa, CLIP text encoder), which can capture the deep semantic associations and contextual dependencies between words in the text;
[0074] The input text is feature extracted through a multi-layer Transformer structure. The encoder model uses a self-attention mechanism to obtain global contextual semantics and ultimately outputs a fixed-length semantic vector representing the overall semantic content of the cultural and creative element information.
[0075] For example, if the CLIP text encoder is used, after encoding the above text, a vector representation with a dimension of 512 may be obtained, such as: [0.214, -0.132, 0.057, ..., 0.098] (a total of 512 dimensions);
[0076] In this vector, each dimension represents a certain semantic feature extracted by the encoding model from the input text. For example, implicit features such as "flying posture", "historical sense", "female image", "religious meaning", and "Oriental charm" will find corresponding dimension weights in the vector space.
[0077] S202, performing visual feature extraction based on target cultural and creative element information, wherein the visual feature extraction extracts image feature vectors through a visual encoder.
[0078] This step involves extracting visual features based on the target cultural and creative elements. This is a key step in achieving stylistic consistency and aesthetic control. By introducing a visual encoder to extract image feature vectors, the visual information contained in the reference image, such as style, composition, color, and pattern, can be structured and converted into computable feature representations, thereby guiding the image generation model to restore a specific visual style and artistic expression.
[0079] The core task of visual feature extraction is to use a trained visual encoder model to process a reference image that matches the target cultural element information and output a fixed-dimensional image feature vector. This image feature vector can be fused with the text semantic vector and input into the generative model to achieve multimodal information-driven pattern generation.
[0080] The image is normalized and resized (e.g., scaled to 224×224), divided into small patches or convolutional processing is performed, and a multi-layer visual Transformer or CNN extracts the image's style features, color distribution, composition information, local patterns, etc., and outputs a fixed-length vector (e.g., 512 or 1024 dimensions) representing the overall visual expression of the image.
[0081] S203: Fusing the text code and the visual features to obtain a fusion result, wherein the fusion result is obtained by projecting the text and the visual features into the same vector space and then adding them together.
[0082] In this step, the text encoding and visual features are fused to achieve a unified expression of semantic information and visual style information, so that the image generation model can understand both "what to generate" and "how to generate", that is, the fusion expression of content and style;
[0083] The fusion method is: projecting the text encoding and visual features into the same vector space and then adding them together to obtain the final fusion result. The fusion result serves as a priori conditions or prompt inputs for the image generation model. The text encoding and visual features are the same in dimension, but they come from different modalities. Direct addition may cause semantic inconsistency and numerical scale imbalance. Before additive fusion, the two need to be mapped to the same feature space through a projection function.
[0084] like Figure 3 As shown in FIG, as a preferred embodiment of the present invention, the steps of performing visualization processing based on the fusion result, using the multimodal embedding vector as a conditional input, guiding the model to generate an image that matches the input semantics, detecting the image and the target cultural and creative element information, and judging the consistency of the image generation specifically include:
[0085] S301, performing visualization processing based on the fusion result, wherein the visualization processing includes model selection and image generation, and the model selection includes a diffusion model and a generative adversarial network.
[0086] In this step, visualization is performed based on the fusion results. This is the key stage of converting the fused multimodal vector (fusion of semantic and style information) into a visual image.
[0087] Visualization processing includes model selection and image generation, among which model selection is crucial. It is necessary to select a suitable image generation architecture based on the characteristics, detail requirements and style complexity of the cultural and creative pattern, including diffusion models (Diffusion Models) and generative adversarial networks (GANs). The visualization processing process selects a suitable architecture based on the different advantages of diffusion models or generative adversarial networks, with fusion vectors as the driving core, to achieve semantically accurate and stylistically refined cultural and creative pattern generation, providing a high-quality image foundation for subsequent consistency detection and cultural and creative applications.
[0088] S302 uses the multimodal embedding vector as a conditional input to guide the model to generate an image that matches the input semantics and detect the image and target cultural and creative element information.
[0089] In this step, the multimodal embedding vector is used as a conditional input. This is a key step in building a "precise control of semantics and style" generation mechanism. The embedding vector combines the semantic content of the text with the visual style characteristics. It can be used as a conditional input to guide the image generation model to output an image that matches the target cultural and creative element information. In order to ensure the reliability and consistency of the generated results, the output image needs to be tested to determine whether it accurately restores the expected cultural and creative element characteristics.
[0090] Use image content recognition models (such as YOLOv8, DETR, CLIP visual embedding matching, etc.) to detect whether the image accurately contains specified semantic elements. Use image style recognition networks or feature space comparison to detect the style of the generated image and determine whether it matches the target style.
[0091] S303: Determine consistency of image generation, where consistency includes image fusion, style transfer, and semantic consistency.
[0092] In this step, the consistency of image generation is determined, ensuring that the generated image matches the target cultural and creative elements. This consistency not only focuses on whether the content contains the target semantic elements, but also includes the naturalness of image fusion, the accuracy of style transfer, and the completeness of semantic expression.
[0093] Through a systematic analysis mechanism, the effectiveness of generated images is comprehensively evaluated from three aspects: image fusion consistency, style transfer consistency, and semantic consistency, ensuring that cultural and creative patterns have both cultural expression and visual beauty and style unity.
[0094] Image fusion consistency refers to whether the visual transition between different content elements of the image is natural, whether the composition is coordinated, and whether the local details are unified. It is the basis for measuring whether the generated pattern has complete image expressiveness; style transfer consistency refers to whether the generated image successfully carries and reproduces the artistic style, expression techniques and visual texture specified in the target cultural and creative elements, such as whether it meets the style requirements of "heavy ink and color", "paper-cut style", "mural texture", etc.; semantic consistency refers to whether the image fully expresses the semantic elements specified in the text information, such as characters, animals, totems, cultural symbols, etc., and whether it has the specified actions, postures or composition meanings.
[0095] like Figure 4 As shown, as a preferred embodiment of the present invention, the steps of comparing the image with the target cultural and creative element, obtaining a scoring result based on the comparison result in combination with a scoring mechanism, comparing the scoring result with a scoring threshold, obtaining an aesthetic preference based on the comparison result, and modifying the image based on the aesthetic preference if the image is determined to be unqualified, specifically include:
[0096] S401, comparing the image with the target cultural and creative element, obtaining a comparison result, and obtaining a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style.
[0097] In this step, the image is compared with the target cultural and creative elements. This comparison is a key process to ensure that the generated image meets the user's preset requirements. The comparison result is based on the analysis and extraction of image content and style features, and the scoring mechanism is used to quantitatively evaluate the similarity between the initial generated image and the target style.
[0098] By integrating the dual comparison mechanism of semantic layer and style layer, the scoring system can clearly indicate whether the image truly restores the specified cultural theme, pattern style and visual expression; the scoring mechanism is based on the above comparison results, and outputs a continuous score through preset weight rules and standardized evaluation models to quantify the similarity between the preliminary generated image and the target cultural and creative style. The comparison scoring mechanism not only supports systematic measurement of the content and style of the generated image, but also provides data basis for subsequent image optimization, so that the cultural and creative pattern generation process has quantifiable, traceable and iterative quality control capabilities.
[0099] S402 , obtaining a scoring result based on the comparison result combined with a scoring mechanism, obtaining a scoring threshold, wherein the scoring threshold is a passing line for the image, and comparing the scoring result with the scoring threshold to obtain a comparison result.
[0100] In this step, the comparison results are combined with a scoring mechanism to generate a scoring result, which is the core step for quantifying the degree of match between the image quality and the target cultural and creative element information. By comparing the scoring result with the preset scoring threshold, the system can clearly determine whether the image is qualified and then decide whether to output, optimize, or regenerate it;
[0101] The scoring threshold is the passing line for image quality, which is used to automatically control the output threshold of the image generation process to ensure that the cultural expression and aesthetic style of the final image meet the actual usage requirements. The scoring threshold is set according to task requirements, user tolerance or target application scenarios. This threshold represents the quality standard line for the image to transform from "experimental" to "usable". If the scoring result is less than the scoring threshold, the image is judged to be a qualified image. If the scoring result is greater than the scoring threshold, the image is judged to be an unqualified image.
[0102] S403, obtaining aesthetic preferences based on the comparison results, wherein the aesthetic preferences include color tone preferences, graphic complexity preferences, and element composition ratio preferences. If the image is determined to be unqualified, the image is modified based on the aesthetic preferences.
[0103] In this step, obtaining aesthetic preferences based on the comparison results is a key step in improving the personalization and aesthetic consistency of image generation. By analyzing the deviations in hue, complexity, and composition of unqualified images, combined with a preset or learned aesthetic preference model, the image generation conditions can be automatically adjusted or the image can be directly modified to output a pattern image that meets the visual preferences of the user or target audience.
[0104] Aesthetic preferences can be obtained through user explicit settings and analysis of user historical behavior. Users can preset their favorite styles, or their aesthetic preferences can be determined by counting the tonal distribution, number of pixels, and composition of images that they have selected, downloaded, and saved in the past.
[0105] If it is judged to be unqualified, the image will be modified according to aesthetic preferences, such as adjusting the generation configuration, controlling the number of output pixels, and if necessary, only modifying the original image composition area or background part, retaining the qualified part.
[0106] like Figure 5 As shown in FIG, an embodiment of the present invention provides an artificial intelligence-based automatic generation system for cultural and creative images, the system comprising:
[0107] The cultural and creative element information module 100 is used to obtain target cultural and creative element information, where the target cultural and creative element information includes element type, style characteristics, cultural background and semantic tags.
[0108] In this system, the cultural and creative element information module 100 obtains target cultural and creative element information. Obtaining target cultural and creative element information is a key prerequisite for the automatic generation method of cultural and creative images. It aims to comprehensively analyze and structure the cultural elements associated with user needs or target scenes, providing a semantic basis and style guidance for subsequent image generation. The target cultural and creative element information includes four dimensions: element type, style characteristics, cultural background, and semantic label, which are expanded as follows:
[0109] Element type refers to the core visual elements that constitute cultural and creative images and is the main content of the image generation process. Common element types include natural elements (such as landscapes, flowers, plants, and animals), artifact elements (such as ceramics, lacquerware, and clothing), symbolic elements (such as totems, patterns, and characters), human elements (such as historical figures and mythological images), and architectural elements (such as ancient buildings and religious sites).
[0110] Stylistic features refer to the expression techniques and aesthetic style of the cultural and creative elements in visual presentation, which directly influence the visual direction of the generated image. Stylistic features can be derived from user settings or automatically inferred by the system based on existing image samples or historical preferences.
[0111] Cultural background refers to the historical origins, regional culture, folk beliefs, or ethnic characteristics of the cultural and creative elements. It provides semantic context and cultural symbolic references, and is the basis for ensuring that cultural and creative patterns have cultural depth and accuracy.
[0112] Semantic tags provide a highly summarized and structured annotation of the three dimensions mentioned above, facilitating semantic vectorization processing by AI models. Semantic tags typically consist of structured tag words and can be combined with natural language processing techniques to automatically extract semantic keywords from text, or standardized matching can be performed using manually annotated knowledge bases.
[0113] The coding feature module 200 is used to perform text coding according to the target cultural and creative element information, extract visual features according to the target cultural and creative element information, and fuse the text coding and visual features.
[0114] In this system, the encoding feature module 200 performs text encoding based on the target cultural and creative element information. Based on the target cultural and creative element information, text encoding and visual feature extraction are combined, and the two are integrated. This is the core step in achieving high-quality generation of cultural and creative patterns. The purpose of text encoding is to convert unstructured or semi-structured text information such as "element type, style characteristics, cultural background and semantic labels" into semantic vectors that can be understood by artificial intelligence models. Visual feature extraction aims to provide artistic elements such as reference style, line characteristics, color style, etc. for cultural and creative pattern generation, ensuring that the output image visually conforms to cultural and style expectations.
[0115] By converting structured semantic information and image style perception information into a unified deep representation, the image generation model's ability to understand content, style, and cultural context can be significantly improved, thereby generating cultural and creative patterns that meet expectations.
[0116] The visualization processing module 300 is used to perform visualization processing based on the fusion results, use the multimodal embedding vector as a conditional input, guide the model to generate an image that matches the input semantics, detect the image and the target cultural and creative element information, and judge the consistency of the image generation.
[0117] In this system, the visualization processing module 300 performs visualization processing based on the fusion results. Using the fused multimodal embedding vector as a generation condition, it guides the image generation model to generate a pattern image that is highly consistent with the target cultural and creative element information. The subsequent detection and matching mechanism then verifies the consistency of the generated image to ensure the restoration and accuracy of the semantic content and style features.
[0118] To ensure that the generated image is consistent with the original target cultural and creative element information, the image is automatically detected and analyzed in terms of content and style. Object detection algorithms (such as YOLOv8 and DETR) are used to identify whether the image contains key semantic elements. Based on the detection results, a consistency assessment is performed to determine whether the generated image meets the dual requirements of semantics and style of the target cultural and creative elements.
[0119] The aesthetic preference module 400 is used to compare the image with the target cultural and creative element, obtain a scoring result based on the comparison result combined with the scoring mechanism, compare the scoring result with the scoring threshold, and obtain the aesthetic preference based on the comparison result. If it is judged to be unqualified, the image is modified according to the aesthetic preference.
[0120] In this system, the aesthetic preference module 400 compares the image with the target cultural and creative elements. After the image is generated, it needs to be compared with the target cultural and creative elements and a comparison score is obtained based on the scoring mechanism. By comparing the score result with the preset score threshold, the system can determine whether the image meets the design requirements.
[0121] If the score is lower than the threshold, it is judged as an unqualified image. Then, the image will be automatically modified and regenerated based on the aesthetic preference information modeled by the user or the system, thereby improving the image quality, fit and artistic expression.
[0122] It not only has the ability to evaluate the quality of cultural and creative pattern generation, but can also automatically repair and regenerate images by adapting to aesthetic preferences, so that the generation results are highly consistent in terms of cultural expression, visual style and user aesthetics, meeting the actual application needs of automatic generation of high-quality cultural and creative images.
[0123] like Figure 6 As shown, as a preferred embodiment of the present invention, the coding feature module 200 includes:
[0124] The text encoding unit 201 is used to perform text encoding according to the target cultural and creative element information, and the text encoding is to convert the target cultural and creative element information into a vector representation of a fixed length.
[0125] In this module, the text encoding unit 201 performs text encoding based on the target cultural and creative element information. This is a fundamental step in constructing a semantic-driven image generation model. By converting unstructured cultural and creative element text information into a fixed-length vector representation, the model can understand and process complex semantic content in a numerical form, thereby driving the image generation module to accurately express cultural imagery and visual style.
[0126] Perform text preprocessing on the target cultural and creative element information and use a language model for embedding modeling. The preferred text encoding model is a pre-trained language model based on the Transformer structure (such as BERT, RoBERTa, CLIP text encoder), which can capture the deep semantic associations and contextual dependencies between words in the text;
[0127] The input text is feature extracted through a multi-layer Transformer structure. The encoder model uses a self-attention mechanism to obtain global contextual semantics and ultimately outputs a fixed-length semantic vector representing the overall semantic content of the cultural and creative element information.
[0128] For example, if the CLIP text encoder is used, after encoding the above text, a vector representation with a dimension of 512 may be obtained, such as: [0.214, -0.132, 0.057, ..., 0.098] (a total of 512 dimensions);
[0129] In this vector, each dimension represents a certain semantic feature extracted by the encoding model from the input text. For example, implicit features such as "flying posture", "historical sense", "female image", "religious meaning", and "Oriental charm" will find corresponding dimension weights in the vector space.
[0130] The visual feature extraction unit 202 is used to extract visual features based on the target cultural and creative element information. The visual feature extraction extracts image feature vectors through a visual encoder.
[0131] In this module, visual feature extraction unit 202 extracts visual features based on the target cultural and creative element information. This is a key step in achieving stylistic consistency and aesthetic control. By introducing a visual encoder to extract image feature vectors, the visual information contained in the reference image, such as style, composition, color, and pattern, can be structured and converted into computable feature representations, thereby guiding the image generation model to restore a specific visual style and artistic expression.
[0132] The core task of visual feature extraction is to use a trained visual encoder model to process a reference image that matches the target cultural element information and output a fixed-dimensional image feature vector. This image feature vector can be fused with the text semantic vector and input into the generative model to achieve multimodal information-driven pattern generation.
[0133] The image is normalized and resized (e.g., scaled to 224×224), divided into small patches or convolutional processing is performed, and a multi-layer visual Transformer or CNN extracts the image's style features, color distribution, composition information, local patterns, etc., and outputs a fixed-length vector (e.g., 512 or 1024 dimensions) representing the overall visual expression of the image.
[0134] The fusion unit 203 is used to fuse the text code and the visual features to obtain a fusion result. The fusion result is obtained by projecting the text and the visual features into the same vector space and then adding them together.
[0135] In this module, the fusion unit 203 fuses the text encoding and visual features to achieve a unified expression of semantic information and visual style information, so that the image generation model can understand both "what to generate" and "how to generate", that is, the fusion expression of content and style;
[0136] The fusion method is: projecting the text encoding and visual features into the same vector space and then adding them together to obtain the final fusion result. The fusion result serves as a priori conditions or prompt inputs for the image generation model. The text encoding and visual features are the same in dimension, but they come from different modalities. Direct addition may cause semantic inconsistency and numerical scale imbalance. Before additive fusion, the two need to be mapped to the same feature space through a projection function.
[0137] like Figure 7 As shown, as a preferred embodiment of the present invention, the visualization processing module 300 includes:
[0138] The visualization processing unit 301 is used to perform visualization processing based on the fusion result. The visualization processing includes model selection and image generation. The model selection includes a diffusion model and a generative adversarial network.
[0139] In this module, the visualization processing unit 301 performs visualization processing based on the fusion result. Visualization processing based on the fusion result is a key stage for converting the fused multimodal vector (fusion of semantic and style information) into a visual image;
[0140] Visualization processing includes model selection and image generation, among which model selection is crucial. It is necessary to select a suitable image generation architecture based on the characteristics, detail requirements and style complexity of the cultural and creative pattern, including diffusion models (Diffusion Models) and generative adversarial networks (GANs). The visualization processing process selects a suitable architecture based on the different advantages of diffusion models or generative adversarial networks, with fusion vectors as the driving core, to achieve semantically accurate and stylistically refined cultural and creative pattern generation, providing a high-quality image foundation for subsequent consistency detection and cultural and creative applications.
[0141] The image detection unit 302 is used to use the multimodal embedding vector as a conditional input to guide the model to generate an image that matches the input semantics and detect the image and target cultural and creative element information.
[0142] In this module, the image detection unit 302 uses the multimodal embedding vector as a conditional input. Using the multimodal embedding vector as a conditional input is a key step in building a "precise control of semantics and style" generation mechanism. The embedding vector combines the semantic content of the text with the visual style characteristics. It can be used as a conditional input to guide the image generation model to output an image that matches the target cultural and creative element information. To ensure the reliability and consistency of the generated results, the output image needs to be tested to determine whether it accurately restores the expected cultural and creative element characteristics.
[0143] Use image content recognition models (such as YOLOv8, DETR, CLIP visual embedding matching, etc.) to detect whether the image accurately contains specified semantic elements. Use image style recognition networks or feature space comparison to detect the style of the generated image and determine whether it matches the target style.
[0144] The consistency determination unit 303 is used to determine the consistency of image generation, where the consistency includes image fusion, style transfer, and semantic consistency.
[0145] In this module, the consistency determination unit 303 determines the consistency of image generation, a key quality control step to ensure that the generated image matches the target cultural and creative element information. This consistency not only focuses on whether the content contains the target semantic elements, but also includes the naturalness of image fusion, the accuracy of style transfer, and the completeness of semantic expression.
[0146] Through a systematic analysis mechanism, the effectiveness of generated images is comprehensively evaluated from three aspects: image fusion consistency, style transfer consistency, and semantic consistency, ensuring that cultural and creative patterns have both cultural expression and visual beauty and style unity.
[0147] Image fusion consistency refers to whether the visual transition between different content elements of the image is natural, whether the composition is coordinated, and whether the local details are unified. It is the basis for measuring whether the generated pattern has complete image expressiveness; style transfer consistency refers to whether the generated image successfully carries and reproduces the artistic style, expression techniques and visual texture specified in the target cultural and creative elements, such as whether it meets the style requirements of "heavy ink and color", "paper-cut style", "mural texture", etc.; semantic consistency refers to whether the image fully expresses the semantic elements specified in the text information, such as characters, animals, totems, cultural symbols, etc., and whether it has the specified actions, postures or composition meanings.
[0148] like Figure 8 As shown, as a preferred embodiment of the present invention, the aesthetic preference module 400 includes:
[0149] The similarity determination unit 401 is used to compare the image with the target cultural and creative element, obtain the comparison result, and obtain a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style.
[0150] In this module, the similarity determination unit 401 compares the image with the target cultural and creative elements. Comparing the image with the target cultural and creative elements is a key process to ensure that the generated image meets the user's preset requirements. The comparison result is based on the analysis and extraction of image content and style features, and the scoring mechanism is used to quantitatively evaluate the similarity between the preliminary generated image and the target style.
[0151] By integrating the dual comparison mechanism of semantic layer and style layer, the scoring system can clearly indicate whether the image truly restores the specified cultural theme, pattern style and visual expression; the scoring mechanism is based on the above comparison results, and outputs a continuous score through preset weight rules and standardized evaluation models to quantify the similarity between the preliminary generated image and the target cultural and creative style. The comparison scoring mechanism not only supports systematic measurement of the content and style of the generated image, but also provides data basis for subsequent image optimization, so that the cultural and creative pattern generation process has quantifiable, traceable and iterative quality control capabilities.
[0152] The scoring unit 402 is configured to obtain a scoring result based on the comparison result in combination with the scoring mechanism, obtain a scoring threshold, where the scoring threshold is the passing line of the image, and compare the scoring result with the scoring threshold to obtain a comparison result.
[0153] In this module, the scoring unit 402 generates a scoring result based on the comparison results and the scoring mechanism, which is the core step for quantifying the degree of match between the image quality and the target cultural and creative element information. By comparing the scoring result with the preset scoring threshold, the system can clearly determine whether the image is qualified and then decide whether to output, optimize, or regenerate it.
[0154] The scoring threshold is the passing line for image quality, which is used to automatically control the output threshold of the image generation process to ensure that the cultural expression and aesthetic style of the final image meet the actual usage requirements. The scoring threshold is set according to task requirements, user tolerance or target application scenarios. This threshold represents the quality standard line for the image to transform from "experimental" to "usable". If the scoring result is less than the scoring threshold, the image is judged to be a qualified image. If the scoring result is greater than the scoring threshold, the image is judged to be an unqualified image.
[0155] The aesthetic preference unit 403 is used to obtain aesthetic preferences based on the comparison results. The aesthetic preferences include color tone preference, graphic complexity preference, and element composition ratio preference. If the image is determined to be unqualified, the image is modified according to the aesthetic preferences.
[0156] In this module, the aesthetic preference unit 403 obtains aesthetic preferences based on the comparison results. This is a key step in improving the personalization and aesthetic consistency of image generation. By analyzing the deviations in hue, complexity, and composition of unqualified images and combining them with a preset or learned aesthetic preference model, it can automatically adjust the image generation conditions or directly modify the image to output a pattern image that meets the visual preferences of the user or target audience.
[0157] Aesthetic preferences can be obtained through user explicit settings and analysis of user historical behavior. Users can preset their favorite styles, or their aesthetic preferences can be determined by counting the tonal distribution, number of pixels, and composition of images that they have selected, downloaded, and saved in the past.
[0158] If it is judged to be unqualified, the image will be modified according to aesthetic preferences, such as adjusting the generation configuration, controlling the number of output pixels, and if necessary, only modifying the original image composition area or background part, retaining the qualified part.
[0159] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are performed:
[0160] Obtaining target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background, and semantic tags;
[0161] Perform text encoding based on target cultural and creative element information, extract visual features based on target cultural and creative element information, and fuse text encoding and visual features;
[0162] Visualize the fusion results and use the multimodal embedding vector as a conditional input to guide the model to generate an image that matches the input semantics, detect the image and the target cultural and creative elements, and determine the consistency of the generated image.
[0163] Compare the image with the target cultural and creative elements, obtain the scoring result based on the comparison result and the scoring mechanism, compare the scoring result with the scoring threshold, and obtain the aesthetic preference based on the comparison result. If it is judged to be unqualified, modify the image according to the aesthetic preference.
[0164] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor performs the following steps:
[0165] Obtaining target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background, and semantic tags;
[0166] Perform text encoding based on target cultural and creative element information, extract visual features based on target cultural and creative element information, and fuse text encoding and visual features;
[0167] Visualize the fusion results and use the multimodal embedding vector as a conditional input to guide the model to generate an image that matches the input semantics, detect the image and the target cultural and creative elements, and determine the consistency of the generated image.
[0168] Compare the image with the target cultural and creative elements, obtain the scoring result based on the comparison result and the scoring mechanism, compare the scoring result with the scoring threshold, and obtain the aesthetic preference based on the comparison result. If it is judged to be unqualified, modify the image according to the aesthetic preference.
[0169] It should be understood that, although the various steps in the flow chart of each embodiment of the present invention are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in order, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0170] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When executed, the program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).
[0171] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0172] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
[0173] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for automatically generating cultural and creative images based on artificial intelligence, characterized in that: The method comprises: Obtaining target cultural and creative element information, wherein the target cultural and creative element information includes element type, style characteristics, cultural background, and semantic tags; Perform text encoding based on target cultural and creative element information, extract visual features based on target cultural and creative element information, and fuse text encoding and visual features; Specifically, the method includes: performing text encoding based on target cultural and creative element information, wherein the text encoding is to convert the target cultural and creative element information into a vector representation of a fixed length; performing visual feature extraction based on the target cultural and creative element information, wherein the visual feature extraction extracts an image feature vector through a visual encoder; and fusing the text encoding and the visual features to obtain a fusion result, wherein the fusion result is obtained by projecting the text and visual features into the same vector space and then adding them together; Visualize the fusion results and use the multimodal embedding vector as a conditional input to guide the model to generate an image that matches the input semantics, detect the image and the target cultural and creative elements, and determine the consistency of the generated image. Compare the image with the target cultural and creative elements, obtain the scoring result based on the comparison result and the scoring mechanism, compare the scoring result with the scoring threshold, and obtain the aesthetic preference based on the comparison result. If it is judged to be unqualified, modify the image according to the aesthetic preference.
2. The method for automatically generating cultural and creative images based on artificial intelligence according to claim 1, characterized in that: The steps of performing visualization processing based on the fusion results, using the multimodal embedding vector as a conditional input, guiding the model to generate an image that matches the input semantics, detecting the image and the target cultural and creative element information, and judging the consistency of the image generation specifically include: Performing visualization processing based on the fusion results, wherein the visualization processing includes model selection and image generation, and the model selection includes a diffusion model and a generative adversarial network; The multimodal embedding vector is used as a conditional input to guide the model to generate images that match the input semantics and detect information between the image and the target cultural and creative elements; Determine the consistency of image generation, including image fusion, style transfer, and semantic consistency.
3. The method for automatically generating cultural and creative images based on artificial intelligence according to claim 1, characterized in that: The steps of comparing the image with the target cultural and creative element, obtaining a scoring result based on the comparison result in combination with a scoring mechanism, comparing the scoring result with a scoring threshold, obtaining an aesthetic preference based on the comparison result, and modifying the image based on the aesthetic preference if the image is determined to be unqualified, specifically include: Comparing the image with the target cultural and creative element, obtaining a comparison result, and obtaining a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style; A scoring result is obtained based on the comparison result combined with the scoring mechanism, and a scoring threshold is obtained. The scoring threshold is the passing line of the image, and the scoring result is compared with the scoring threshold to obtain a comparison result; Aesthetic preferences are obtained based on the comparison results, including color tone preferences, graphic complexity preferences, and element composition ratio preferences. If the image is determined to be unqualified, the image is modified based on the aesthetic preferences.
4. The method for automatically generating cultural and creative images based on artificial intelligence according to claim 1, characterized in that: The element types include traditional painting, sculpture and calligraphy, the style features include ink painting style and fine brushwork style, the cultural background includes historical periods and regional characteristics, and the semantic tags are descriptive keywords for the elements.
5. An artificial intelligence-based automatic generation system for cultural and creative images, characterized by: The system comprises: A cultural and creative element information module obtains target cultural and creative element information, including element type, style characteristics, cultural background, and semantic tags; The encoding feature module encodes text based on the target cultural and creative element information, extracts visual features based on the target cultural and creative element information, and integrates the text encoding and visual features; Specifically, it includes: a text encoding unit, which performs text encoding based on the target cultural and creative element information, and the text encoding is to convert the target cultural and creative element information text information into a vector representation of a fixed length; a visual feature extraction unit, which performs visual feature extraction based on the target cultural and creative element information, and the visual feature extraction extracts the image feature vector through a visual encoder; a fusion unit, which fuses the text encoding and visual features to obtain a fusion result, and the fusion result is the text and visual features projected into the same vector space and then added; The visualization processing module performs visualization processing based on the fusion results, uses the multimodal embedding vector as a conditional input, guides the model to generate images that match the input semantics, detects the image and the target cultural and creative elements, and determines the consistency of the image generation; The aesthetic preference module compares the image with the target cultural and creative elements, obtains the scoring result based on the comparison result combined with the scoring mechanism, compares the scoring result with the scoring threshold, and obtains the aesthetic preference based on the comparison result. If it is judged to be unqualified, the image will be modified according to the aesthetic preference.
6. The artificial intelligence-based automatic generation system for cultural and creative images according to claim 5 is characterized in that: The visualization processing module includes: A visualization processing unit, performing visualization processing based on the fusion result, wherein the visualization processing includes model selection and image generation, wherein the model selection includes a diffusion model and a generative adversarial network; The image detection unit takes the multimodal embedding vector as a conditional input, guides the model to generate an image that matches the input semantics, and detects the image and the target cultural and creative element information; The consistency determination unit determines the consistency of image generation, wherein the consistency includes image fusion, style transfer and semantic consistency.
7. The artificial intelligence-based automatic generation system for cultural and creative images according to claim 6 is characterized in that: The aesthetic preference module includes: a similarity determination unit, comparing the image with the target cultural and creative element, obtaining a comparison result, and obtaining a scoring mechanism, wherein the scoring mechanism is used to evaluate the similarity between the preliminary image and the target style; A scoring unit, which obtains a scoring result based on the comparison result and the scoring mechanism, obtains a scoring threshold, wherein the scoring threshold is a passing line of the image, and compares the scoring result with the scoring threshold to obtain a comparison result; The aesthetic preference unit obtains aesthetic preferences based on the comparison results, wherein the aesthetic preferences include color tone preferences, graphic complexity preferences, and element composition ratio preferences. If the image is determined to be unqualified, the image is modified according to the aesthetic preferences.
8. The artificial intelligence-based automatic generation system for cultural and creative images according to claim 7 is characterized in that: The element types include traditional painting, sculpture and calligraphy, the style features include ink painting style and fine brushwork style, the cultural background includes historical periods and regional characteristics, and the semantic tags are descriptive keywords for the elements.
Citation Information
Patent Citations
Intelligent service system and method for customizing cultural and creative products
CN119784471A
Cited By
Textile creative element generation method based on semantic topology and multi-dimensional style transfer
CN122636781A