An artistic image generation method for national style content creation
By constructing a database of Chinese-style art images with independent copyright data and a dual-stream feature extraction network, combined with physical rendering technology, the copyright risks and lack of physical topological logic in the existing models when generating Chinese-style content are solved, realizing the generation of high-fidelity and compliant Chinese-style art images, and meeting the needs of commercial and cultural protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIVERSITY OF MEDIA AND COMMUNICATIONS
- Filing Date
- 2026-02-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing generative AI models suffer from data copyright risks, lack of physical topological logic, and distortion of cultural aesthetics and material texture when generating Chinese-style (Guofeng) content, making them difficult to apply effectively in the fields of commerce and cultural preservation.
A multimodal graph database of traditional Chinese art based on proprietary copyright data is constructed. A dual-stream feature extraction network and physical rendering technology are used, combined with topological consistency loss and aesthetic rating loss. Through LoRA fine-tuning and ControlNet constrained generation model, high-fidelity generation of traditional Chinese art is achieved.
It achieves high-fidelity generation that complies with copyright, simulates physical processes, and integrates cultural aesthetics. The generated images have legal protection and practicality in the commercial and cultural protection fields, and the texture and realism are significantly improved.
Smart Images

Figure CN122134839A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of artificial intelligence generated content (AIGC), computer vision, deep learning, and digital cultural heritage protection, specifically to a method for generating traditional Chinese art images based on fine-tuning of proprietary copyright data, dual-stream feature extraction, and physical constraint optimization. Background Technology
[0002] In recent years, generative artificial intelligence (AIGC) technology based on deep learning has made groundbreaking progress. In particular, image generation algorithms (such as StableDiffusion and Midjourney) centered on the Denoising Diffusion Probability Model (DDPM) and the Latent Diffusion Model (LDM) have learned the latent distribution patterns of images through pre-training on large-scale image and text datasets, enabling the rapid generation of high-quality images from text descriptions. These technologies have been widely applied in fields such as commercial illustration, game asset creation, and film and television concept design, greatly improving content production efficiency.
[0003] Limitations of existing technology: While general-purpose large models perform excellently in generating realistic or anime-style content, existing technologies exhibit significant shortcomings when handling "Guofeng" content, which possesses unique physical properties, technological logic, and profound cultural connotations. These shortcomings are not merely aesthetic differences, but rather deeper deficiencies in data compliance, physical topology logic, and material simulation capabilities, primarily facing the following three major technical bottlenecks:
[0004] First, there are the "black box" risks of data copyright and the dilemma of commercial compliance. The training data for most current mainstream generative models comes from indiscriminate scraping of the entire internet, including a large amount of unauthorized contemporary Chinese-style illustrations, works by intangible cultural heritage inheritors, and digital museum collections. Due to the "black box" memory characteristics of deep neural networks, the models often implicitly learn and reproduce the unique style of specific artists (such as the knife techniques of a particular paper-cutting school). This generation method, lacking a "whitelist" rights confirmation mechanism, makes the generated Chinese-style images face extremely high copyright infringement risks in commercial scenarios such as cultural heritage IP development and brand collaborations, hindering the technology's practical application.
[0005] Second, the lack of physical topological logic leads to "similar in form but false in substance." Traditional Chinese art often follows strict physical craftsmanship logic, while general models based on statistical probability only focus on the visual correlation between pixels and lack an understanding of geometric topological structures.
[0006] Connectivity failure: Taking paper cutting as an example, its core technique requires that the pattern must be physically connected ("unbroken even after a thousand cuts"). Existing models often generate paper cutting images composed of a large number of discrete color blocks. Although the visual style is imitated, due to the broken lines (topological disconnection), it is impossible to use for laser engraving or physical cutting.
[0007] Structural fallacy: When generating shadow puppets, puppets, and other similar subjects, the models often ignore the mechanical principles of joint riveting, resulting in phenomena that violate the laws of physics, such as limb fusion or reverse bending, which lead to generated results that are only aesthetically pleasing but lose their practicality.
[0008] Third, there is a dual distortion in cultural aesthetic paradigms and material physical texture. The beauty of "Chinese style" lies in its unique cultural charm and the texture of materials accumulated over time, while general models often present a "Western-style bone structure" and a "digital plastic feel".
[0009] Aesthetic bias: The model often suffers from an imbalance of Western weights in the training data, leading to a biased understanding of Eastern aesthetics (such as the facial proportions of "three courts and five eyes" and scattered perspective composition), and even phenomena such as confusion of clothing from different dynasties and mixing of Chinese and Western architectural elements.
[0010] Lack of texture: The model struggles to reproduce the accumulated texture and oxidation / weathering marks of the mineral pigments in Dunhuang murals, and also fails to simulate the capillary diffusion effect of ink in Xuan paper fibers. The generated images are often too smooth and bright, lacking the historical weight and realistic tactile feel that traditional artworks should possess.
[0011] In conclusion, developing a dedicated image generation method based on proprietary copyrighted data, incorporating physical topological constraints, and capable of accurately reproducing Eastern aesthetics and material textures is key to breaking through current industry bottlenecks. Summary of the Invention
[0012] To address the numerous problems existing in the aforementioned background technologies, this invention proposes an art image generation method for creating content with a traditional Chinese style, aiming to overcome the limitations of general generative models that only achieve a superficial resemblance in the essence of traditional Chinese art. Addressing the difficulty of accurately processing the physical processes of paper-cutting (connectivity), shadow puppetry (joint structure), and unique textures such as mineral particles in murals and ink diffusion in existing AIGC methods, this invention constructs a comprehensive technical system encompassing bottom-level data ownership confirmation, mid-level feature decoupling, and top-level physical rendering. This method is particularly suitable for the digital creation of intangible cultural heritage (such as paper-cutting, shadow puppetry, and New Year paintings) and traditional arts (such as Dunhuang murals and ink paintings), IP derivative design, and immersive cultural experience scenarios.
[0013] Meanwhile, this invention addresses the copyright risks associated with the commercialization of traditional Chinese culture AIGC by constructing a vertical domain map with independent intellectual property rights and adopting a "whitelist" verification mechanism, thereby providing compliance guarantees for application scenarios such as intangible cultural heritage creative products, game art, and brand marketing.
[0014] Furthermore, this invention introduces a Chinese aesthetics scoring feedback and physical texture simulation technology to effectively coordinate the differences between Eastern and Western aesthetics, avoid the model misinterpreting Eastern cultural symbols, and generate professional-grade artistic images that combine traditional aesthetic norms with modern realism.
[0015] Specifically, this method integrates physical topology constraint mechanisms, dual-stream feature perception networks, and physically based rendering (PBR) post-processing techniques to achieve high-fidelity generation of specific Chinese style genres.
[0016] First, this invention constructs a multimodal atlas of traditional Chinese art based on independent copyright and semantic decoupling. Instead of relying on haphazardly crawled data from the entire internet, this invention establishes a dedicated database based on works by intangible cultural heritage inheritors with confirmed rights and high-resolution scans from museums. On this basis, an expert-guided preprocessing algorithm deeply decouples the original data along the "structure-material" dimension. For structure-dominated data such as paper-cutting and shadow puppetry, binary morphological processing and skeleton extraction algorithms are used to separate their geometric topology and connected component features. For material-dominated data such as murals and rock paintings, color space separation and frequency domain texture analysis techniques are used to extract the spectral distribution of mineral pigments and the weathering texture features of the carrier (such as wall plaster or Xuan paper).
[0017] Second, a dual-stream perceptual feature extraction network is constructed that integrates structure and material analysis. To accurately capture the complex features of traditional Chinese art, this invention designs a parallel dual-stream coding architecture:
[0018] TopologyStream focuses on capturing the geometric logic of images. This branch employs improved edge detection networks (such as the Holistically-NestedEdgeDetection network combined with dilated convolutions) to specifically extract the skeleton connectivity vector, edge closure index, and key node information of images. It aims to solve topological failure problems such as "broken lines" in paper-cutting and "structural disorder" in shadow puppetry generated by general models.
[0019] MaterialStream: This branch focuses on capturing the physical texture of an image. It uses a Gram matrix to calculate the correlation between feature channels and combines it with a statistical texture model to extract the graininess of pigments, the diffusion levels of ink, and historical weathering traces of the medium (such as cracks and peeling).
[0020] Third, a dynamic conditional adaptation and gating fusion mechanism based on physical perception is constructed. This invention introduces a Physics-AwarePromptAdapter, which maps the user's input natural language description into a multimodal conditional vector containing physical attribute constraints (such as "hollowed out," "mottled," and "semi-transparent"). Subsequently, an adaptive gated cross-attention mechanism is adopted to dynamically adjust the injection weights of dual-flow features according to the characteristics of the target Chinese style category. For example, when generating a paper-cutting style, the weight of the structural flow is automatically increased to ensure line connectivity; when generating a mural style, the weight of the material flow is increased to highlight color texture, thereby achieving accurate matching between style and content.
[0021] Fourth, a LoRA (Low-Rank Adaptation) fine-tuning strategy based on topological consistency and aesthetic scoring is implemented. This invention performs vertical domain low-rank adaptation (LoRA) fine-tuning on pre-trained diffusion models (such as StableDiffusion). During training, a novel composite loss function system is introduced to strongly constrain the generation process:
[0022] Topological Consistency Loss: By calculating the distance transformation difference between the generated image and the target structural skeleton, it penalizes broken connected components and forces the model to learn a topological structure that conforms to the physical process requirements.
[0023] AestheticScoreLoss: Based on a pre-trained Eastern aesthetics evaluation model, it guides the generated results to converge towards traditional compositional proportions (such as scattered perspective) and color systems (such as traditional Chinese colors), correcting Westernized aesthetic biases.
[0024] Fifth, a realistic post-processing module based on physically based rendering (PBR) is constructed. To eliminate the "digital smoothness" of the generated image, this invention introduces the physically based rendering process from computer graphics. Based on the depth information and semantic mask of the generated image, the normal map, roughness map, and height map are derived in reverse. Under virtual lighting conditions, diffuse reflection, specular reflection, and subsurface scattering effects are calculated to simulate the fibrous edges of Xuan paper, the layered projection of paper cutting, and the three-dimensional peeling effect of mural wall plaster, achieving museum-level visual restoration.
[0025] The objective of this invention is achieved through the following technical solution:
[0026] A method for generating artistic images for the creation of Chinese-style content includes the following steps:
[0027] Step 1: Construct a multimodal graph database of traditional Chinese art containing data on traditional Chinese art.
[0028] Step 2: Construct a dual-stream perceptual feature extraction network to decouple the traditional Chinese art data and generate structural topological features, texture and material features, and semantic imagery flow features;
[0029] Step 3: Construct a Physics-AwarePromptAdapter, which maps the user's natural language input into a multimodal condition vector containing physical attribute constraints, and uses a gated cross-attention mechanism to inject the multimodal condition vector into the Physics-AwarePromptAdapter to generate the adapter model.
[0030] Step 4: Using the multimodal condition vector generated in Step 3 and the structural topology features generated in Step 2 as structural conditions, we input the diffusion generation model that has been fine-tuned by LoRA strategy and ControlNet to generate an initial latent spatial representation with the physical characteristics of Chinese style.
[0031] Step 5: Based on the structural topology features and semantic imagery flow features generated in Step 2, a composite total loss function is constructed using the topology consistency loss function and the traditional Chinese aesthetics scoring loss function. The initial latent space representation generated in Step 4 is then subjected to adversarial constraints and iterative denoising optimization using the composite total loss function, and the denoised initial image is output.
[0032] Step 6: Using a physically based rendering post-processing enhancement module, the texture material features generated in Step 2 are used as material guides to perform texture mapping, light and shadow reconstruction and subsurface scattering simulation on the denoised initial image output in Step 5, and output the final artistic Chinese style image.
[0033] In step 1, a multimodal graph database based on traditional Chinese art is constructed, which includes data on traditional Chinese art. Specifically, this involves: high-precision scanning of four types of copyrighted images collected: paper-cutting, shadow puppetry, Dunhuang murals, and ink paintings; semantic segmentation using SAM (SegmentAnythingModel); and the creation of an image-text-structure-material quadruple index by combining manual annotations.
[0034] For paper cutting, extract the binary mask images of the negative and positive cuts;
[0035] A model of color erosion evolution based on time was established for the Dunhuang murals.
[0036] In step 2, the dual-stream sensing feature extraction network includes:
[0037] A shallow convolutional module with shared weights and two independent deep branches connected to the shallow convolutional module, namely a structure flow branch and a material flow branch;
[0038] The structured flow branch uses an HED network with dilated convolutions to expand the receptive field and specifically capture the hollow connectivity of paper-cutting or shadow puppetry.
[0039] The material flow branch employs a Gram matrix and a Perlin noise generator connected to the Gram matrix. The Gram matrix calculates the correlation of features in each channel, and the Perlin noise generator simulates irregular weathering textures.
[0040] In step 2, the structural topological features include: a skeleton connectivity vector based on distance transformation and an edge closure index.
[0041] In step 2, the texture material features include: the spectral reflectance distribution of mineral pigments, the fiber texture noise map of paper / fabric, and the color attenuation matrix based on a historical weathering model.
[0042] In step 2, the semantic imagery flow features include: high-dimensional semantic embeddings aligned with the CLIP model and the Chinese style-specific knowledge graph.
[0043] In step 3, the calculation method for the gating cross-attention mechanism is as follows:
[0044]
[0045] in, This is a gated cross-attention mechanism, where Q is the semantic embedding of the user's text input; For structural flow characteristic keys, V represents the material flow feature key; V represents the value vector. This is the scaling factor; An adaptive gating factor that automatically adjusts according to the target style (when generating a paper-cut style). When generating mural style ).
[0046] In step 4, the diffusion generation model is fine-tuned jointly by the LoRA strategy and ControlNet, i.e., the image generation model fine-tuned by the LoRA strategy. The fine-tuning strategy is as follows:
[0047] Based on StableDiffusionXL, the VAE encoder and the first half of the UNet backbone network are frozen. For each type of Chinese style, an independent LoRA rank matrix is trained. At the same time, an auxiliary ControlNet branch is trained to receive the topological flow features extracted in step 2 to enforce the line connectivity of the generated image.
[0048] In step 5, based on the structural topological features and semantic imagery flow features generated in step 2, a composite total loss function is constructed using the topological consistency loss function and the traditional Chinese aesthetics scoring loss function. Specifically, this includes: the composite total loss function. The calculation formula is:
[0049]
[0050] in, The standard noise prediction mean square error loss for the diffusion model;
[0051] The topological consistency loss function is used to penalize broken lines in paper-cutting and shadow puppetry, and calculates the difference in Euclidean distance transformation between the generated image and the target structure mask.
[0052] The style Gram matrix loss is used to constrain the color and brushstroke texture distribution of Dunhuang murals and ink paintings;
[0053] The loss function for scoring traditional Chinese aesthetics is the semantic alignment loss based on the pre-trained traditional Chinese aesthetics scoring model, which is used to enhance the artistic charm of the image.
[0054] , and The weights for each type of loss. The series of parameters represent the weights of each loss, and are dynamically adjusted according to the style and genre, such as in the paper-cutting mode. .
[0055] In step 6, a post-processing enhancement module based on physically based rendering is used. Using the texture material features generated in step 2 as material guides, texture mapping, lighting reconstruction, and subsurface scattering simulation are performed on the denoised initial image output in step 5 to output the final artistic Chinese-style image. Specifically, this includes:
[0056] The post-processing enhancement module based on physical rendering targets the paper-cutting in the denoised initial image. Using the texture material features generated in step 2 as material guides, it generates a normal map and a height map through texture mapping. Through light and shadow reconstruction, it calculates the edge projection in a virtual 3D lighting environment to simulate the thickness of the paper and outputs the final artistic Chinese style image.
[0057] The post-processing enhancement module based on physical rendering targets the Dunhuang murals in the denoised initial image: it overlays a crack mask based on fractal noise, performs nonlinear decay simulation oxidation on color saturation, completes subsurface scattering simulation, and outputs the final artistic Chinese style image.
[0058] For the shadow puppet and ink painting in the denoised initial image: no post-processing enhancement based on physical rendering is performed, and the denoised initial image output in step 5 is directly used as the final artistic Chinese style image.
[0059] Compared with the prior art, the present invention has the following advantages:
[0060] This invention meets the high requirements of professionalism, accuracy, and controllability for the creation of traditional Chinese style content in the art field, specifically:
[0061] 1. Building a Copyright-Compliant Commercial Security Barrier: Addressing the "copyright black box" problem caused by the indiscriminate data scraping of existing general-purpose models, this invention establishes a training system entirely based on "whitelist"-based rights confirmation data. By rigorously screening authorized works by intangible cultural heritage inheritors and museum-quality data, it ensures that every generated image in the style of paper-cutting, murals, or ink painting has a clearly traceable and legally compliant source. This mechanism completely eliminates potential copyright infringement risks, providing a solid legal guarantee for the development of traditional Chinese style IPs, brand collaborations, and the issuance of digital assets.
[0062] 2. Achieving physical-level process simulation and bridging the "virtual and physical manufacturing" link. Existing technologies often fail to materialize structural images such as paper-cutting and shadow puppetry due to broken lines or misaligned joints. This invention innovatively introduces a "topological consistency loss function" (…). The method employs the "ControlNet geometric constraint mechanism" to force the model to follow physical connectivity logic during the feature learning stage. Experimental results show that the physical connectivity rate of the paper-cutting patterns generated by this method has increased from about 60% in existing technologies to over 98%. The generated images can be directly imported into laser engraving machines or 3D printers for physical processing without manual restoration, achieving a seamless connection from AIGC to physical manufacturing.
[0063] 3. Recreating Museum-Grade Material Texture and Correcting Westernized Aesthetic Bias: To address the common "plastic" look and Westernized aesthetic issues in images generated by general models, this invention constructs a dual-stream perceptual feature extraction network and introduces Physically Based Rendering (PBR) technology into the post-processing workflow. By simulating the accumulation of mineral pigments, the capillary diffusion of Xuan paper, and the weathering marks of time, it successfully restores the unique microscopic texture of traditional Chinese art, improving texture realism by 40% compared to general models. Simultaneously, in conjunction with a traditional Chinese aesthetics scoring feedback mechanism, it effectively corrects erroneous Western perspective and composition, ensuring that the generated content possesses both historical depth and conforms to Eastern aesthetic norms. Attached Figure Description
[0064] Figure 1 This invention presents a schematic diagram of the overall process of an algorithm for generating artistic images for the creation of Chinese-style content.
[0065] Figure 2 Schematic diagram of the dual-stream sensing feature extraction network structure.
[0066] Figure 3 Schematic diagram of physical perception adaptation and gating cross-attention mechanism.
[0067] Figure 4 : Iterative optimization architecture diagram based on topology consistency loss and aesthetic rating loss.
[0068] Figure 5 : A schematic diagram of layered compositing in post-processing of physically based rendering (PBR).
[0069] Figure 6 : This is a diagram illustrating an embodiment of the present invention. Detailed Implementation
[0070] A method for generating artistic images for the creation of Chinese-style content includes the following steps:
[0071] Step 1: Construct a multimodal graph database of traditional Chinese art with independent intellectual property rights, and perform physical attribute decoupling preprocessing on the data;
[0072] Step 2: Construct a dual-stream perceptual feature extraction network to perform parallel feature encoding on the traditional Chinese art data, generating multi-dimensional traditional Chinese art features that include topological geometric information and physical material information;
[0073] Step 3: Construct a physical constraint-based dynamic prompt word adapter to map the input text into a multimodal conditional vector, and adopt a gated cross-attention mechanism to adaptively adjust the feature fusion weights according to the physical characteristics of the target style category to generate a unified guided conditional embedding.
[0074] Step 4: Construct a diffusion generation model based on LoRA strategy and ControlNet joint fine-tuning, embed the guiding conditions generated in step 3 into the injection model, and generate an initial representation with the physical characteristics of a specific Chinese style in the latent space.
[0075] Step 5: Introduce the topological consistency loss function and the traditional Chinese aesthetics scoring loss function to construct a composite total loss function to perform adversarial constraints and iterative optimization on the generation process, and output the denoised initial image;
[0076] Step 6: Construct a post-processing enhancement module based on physically based rendering (PBR). Through normal reconstruction, subsurface scattering simulation and procedural weathering processing, enhance the texture of the initial image and output the final artistic Chinese style image.
[0077] Furthermore, in step 1, a multimodal graph database of traditional Chinese art with independent intellectual property rights is constructed and preprocessing for physical attribute decoupling is performed, specifically including:
[0078] 1.1) High-precision data acquisition and cleaning: Establish a data acquisition standard based on a whitelist mechanism, only recording works by intangible cultural heritage inheritors with verified ownership and digital resources of museum collection level. The acquisition standard requires an image resolution of no less than 4K (3840×2160 pixels), a color depth of 16-bit, and a lossless TIFF or RAW format. Multi-scale quality assessment is performed on the acquired images, removing blurry, overexposed, or color-distorted samples. The Laplacian operator is used to calculate image sharpness scores, and a threshold is set to remove low-frequency blurry images; histogram equalization is used to detect brightness distribution and remove samples with uneven illumination.
[0079] 1.2) Decoupling of Structure and Material: To enable the model to independently learn "form" and "material," orthogonal feature decoupling processing is performed on the original data.
[0080] a) Topology extraction for structure-dominant data: For strongly structured data such as paper-cutting and shadow puppetry, the Otsu adaptive thresholding algorithm is used for binarization segmentation to obtain the foreground mask. Subsequently, morphological closing was applied to fill the tiny holes, and the Zhang-Suen thinning algorithm was used to extract the topological skeleton map with a single pixel width. This skeleton graph removes the interference of line thickness and color, retaining only pure geometric connectivity. Furthermore, an adjacency matrix from graph theory is used to construct the node relationship graph of the skeleton, marking endpoints, intersections, and loops, forming a metadata index of structural features.
[0081] b) Material Separation for Texture-Dominant Data: For strong texture data such as murals and ink paintings, frequency domain analysis is used for separation. The image is converted to the frequency domain using Fast Fourier Transform (FFT), and a high-pass filter is designed to truncate low-frequency structural components while retaining high-frequency texture components. Simultaneously, in the CIELAB color space, the K-Means clustering algorithm (K=5~8) is used to extract the dominant color palette, and the Euclidean distance of each pixel relative to the dominant color is calculated to generate a color distribution probability map. For the mural data, a U-Net-based semantic segmentation network is also needed to identify and extract the masks of "disease layers" such as cracks and peeling in the image. As an independent channel for weathering characteristics.
[0082] Furthermore, in step 2, a dual-stream perceptual feature extraction network is constructed. This network includes parallel structural topology flow branches and texture material flow branches, specifically including:
[0083] 2.1) TopologyStream Branch: This branch aims to extract geometric topological features that are invariant to translation, rotation, and scaling.
[0084] Network Architecture: An improved HED network is used as the backbone. To capture long-distance line connections in traditional Chinese art (such as continuous patterns in paper-cutting), the ordinary convolutional layers in the deep layers (SideOutput3,4,5) of the network are replaced with dilated convolutions, with dilation rates set to 2, 4, and 8 respectively. This design expands the receptive field by 3 times without increasing the number of parameters, enabling it to perceive global topological connectivity.
[0085] Feature Output: The output contains a three-dimensional structural feature vector. :
[0086] Connectivity features: Gradient fields generated by the DistanceTransform algorithm characterize the distance distribution from pixels to the skeleton center;
[0087] Closure feature: Calculate the Euler number in the edge graph to characterize the number and hierarchical relationship of holes in the pattern;
[0088] Key feature: Extract the coordinates of high curvature points and intersections in the skeleton to represent the geometric nodes of the pattern.
[0089] 2.2) MaterialStream Branch: This branch aims to extract statistical texture features that reflect physical properties.
[0090] Network architecture: A pre-trained VGG-19 network is used as the feature extractor. Fully connected layers are removed, and only convolutional layers are retained as texture encoders.
[0091] Feature Calculation: Input the original image and extract feature activation maps from four layers: conv1_2, conv2_2, conv3_4, and conv4_4. For each feature map layer, calculate the Gram matrix between its channels. The Gram matrix captures the correlation between features (e.g., the co-occurrence probability of "cinnabar red" and "graininess") by calculating the inner product of feature vectors, while ignoring specific spatial locations.
[0092] Weathering feature embedding: embedding the disease layer mask extracted in step 1 The material is encoded as a weathering feature vector and concatenated with the Gram matrix features to form the final material feature vector. This vector encodes the type of pigment, particle size, and degree of weathering over time.
[0093] Furthermore, in step 3, the construction of the dynamic condition vector adopts a physical perception adaptation and gated cross-attention mechanism, and its specific calculation method is as follows:
[0094] 3.1) Physics-AwarePromptAdapter: An adapter module based on a rule engine and CLIP text encoder is constructed. When the input text contains keywords related to a specific Chinese style, the adapter automatically retrieves a pre-defined physical attribute knowledge base and adds implicit constraint vectors. For example, when the intent to "paper cutting" is detected, semantic embeddings of "connected domain," "binary edge," and "hollow" are automatically injected; when the intent to "ink painting" is detected, semantic embeddings of "capillary diffusion" and "inkWash" are injected. These implicit vectors are weighted and fused with the explicit text vectors input by the user to generate enhanced semantic features. .
[0095] 3.2) Adaptive Gated Cross-Attention: To address the varying degrees of dependence on "form" and "material" among different Chinese style schools, an adaptive gating mechanism is introduced. A lightweight style classifier is constructed, taking the user's text prompts as input and outputting the probability distribution of whether the target style belongs to "structure-dominated" or "material-dominated" style, thereby calculating the gate control coefficient. The attention fusion formula is as follows:
[0096]
[0097] in:
[0098] The enhanced text semantic feature vector is used as a query.
[0099] The structural topological flow features extracted in step 2 are used as structural keys.
[0100] The texture material flow features extracted in step 2 are used as material keys.
[0101] This represents a weighted concatenation operation of feature channels to construct a hybrid key vector;
[0102] The value vector after linear mapping;
[0103] This is the scaling factor, with a value of 64.
[0104] An adaptive gating coefficient. When generating paper-cut style... Approaching 0.8~0.9 forces the model to pay close attention to the connectivity of lines; when generating a mural style, Approaching 0.2~0.3 allows the model to focus primarily on pigments and weathering textures.
[0105] Furthermore, in step 4, the image generation model, jointly fine-tuned by the LoRA strategy and ControlNet, specifically includes:
[0106] 4.1) Basic Generation Architecture: StableDiffusionXL (SDXL) is used as the basic model. SDXL has a larger number of parameters (2.6BUNet) and dual text encoders, which can more finely parse complex Chinese-style patterns.
[0107] 4.2) LoRA (Low-Rank Adaptation) fine-tuning strategy: Without compromising the generalization ability of the base model, LoRA technology is used to inject knowledge from the Chinese style vertical domain.
[0108] Injection location: Freeze most of the weights of the SDXL VAE encoder and the UNet backbone network, only projecting the matrix into the UNet's cross-attention layers. Injecting low-rank matrix and ,Right now .
[0109] Group training: Four independent LoRA weighted packages are trained for each of the four styles: "paper cutting," "shadow puppetry," "mural painting," and "ink painting." Rank is set. This ensures sufficient feature capacity to store complex artistic style information.
[0110] Training objective: To enable the model to learn the color distribution paradigms of specific art forms (such as the "red and white contrast" in paper-cutting and the "blue and green heavy colors" in murals) and compositional rules.
[0111] 4.3) ControlNet topology constraint branch: In order to make up for the lack of control over the geometric structure of the pure diffusion model, a ControlNet branch is connected in parallel.
[0112] Input conditions: Use the skeleton map or edge map extracted in step 2 as the input conditions for ControlNet. .
[0113] Feature injection: The structural features extracted by ControlNet are superimposed onto the corresponding layers of the main UNet through zero convolution layers.
[0114] Mechanism of action: During the generation process, ControlNet acts as a "geometric lock", which forces the generated image contour to be consistent with the input topology, thereby physically preventing the phenomenon of "line breakage" or "structural misalignment".
[0115] Furthermore, in step 5, fine-tuning employs a composite total loss function as a constraint. The calculation formula is:
[0116]
[0117] The meanings of each symbol are as follows:
[0118] (TotalLoss): Represents the weighted total loss of the model in a single training iteration, used to guide backpropagation gradient updates;
[0119] (DiffusionLoss): Represents the denoising and reconstruction loss of the diffusion model, specifically the predicted noise. Compared with real Gaussian noise The mean squared error (MSE) between the two is used to constrain the content fidelity of the generated image at the pixel level;
[0120] (Topological Consistency Loss): Represents the topological consistency loss, specifically the difference between the generated image skeleton and the target ground truth skeleton in the distance transformation space, used to constrain the connectivity of paper-cutting lines and the geometric correctness of the shadow puppet structure;
[0121] (GramMatrixStyleLoss): Represents style texture loss, specifically the Frobenius norm distance between Gram matrices calculated based on feature maps extracted from the VGG-19 network, used to constrain the consistency between the material texture (such as mineral pigment graininess, paper fiber texture) of the generated image and the target style.
[0122] (AestheticScoreLoss): This represents the loss in the Chinese style aesthetic score, which comes from the negative feedback of the pre-trained Chinese style aesthetic scorer (AestheticScorer) on the generated image. It is used to penalize the generated results that deviate from the Eastern aesthetic paradigm (such as scattered perspective and traditional color scheme).
[0123] : Represents the weighting coefficient of the denoising and reconstruction loss, usually set to 1.0 as the baseline;
[0124] : Represents the weighting coefficient of topological consistency loss. When generating structure-dominant styles (such as paper-cutting and shadow puppetry), this value is adaptively increased (suggested range 10.0~20.0) to force the model to prioritize physical connectivity.
[0125] : Represents the weighting coefficient for style texture loss. When generating material-dominated styles (such as murals or ink paintings), this value is adaptively increased (recommended range 15.0~25.0).
[0126] : Represents the weighting coefficient for aesthetic score loss, used to fine-tune the artistic quality of an image (suggested range 0.1~0.5).
[0127] 5.1) (Diffusion denoising loss):
[0128] This is the mean squared error loss of the standard diffusion model, used to calculate the predicted noise. With real noise Differences between them:
[0129]
[0130] This is used to ensure that the generated image content is basically consistent with the text description.
[0131] in:
[0132] : Represents the mathematical expectation, referring to the expected value at random sampling time steps over the entire training dataset. Latent variables of the original image and Gaussian noise Perform average calculations to ensure that the model learns the general distribution patterns of the data rather than individual cases;
[0133] : Represents the random time step in the diffusion process, and its value typically ranges from 1 to 2. (For example ), representing the degree to which noise has been added to the current image ( The larger the value, the stronger the noise and the blurrier the image.
[0134] : Represents the representation vector of the original, un-noiseed, authentic Chinese style image in the latent space;
[0135] : indicates that it follows the standard normal distribution The true noise tensor sampled in the middle, which is the actual noise target value added to the image in this training iteration;
[0136] : Represents the square of the L2 norm, i.e., the mean square error (MSE), used to quantify the pixel-level distance between predicted noise and true noise;
[0137] : Represents the neural network model to be trained (i.e., the fine-tuned UNet). The learnable parameters of the model (including LoRA weights);
[0138] : Indicates at time step The latent variables of the noisy image are derived from the original latent variables. and noise Through the diffusion forward formula = Calculated;
[0139] : This refers to the physical perception multimodal condition vector (ConditionVector) generated in step 3 through the gating cross-attention mechanism. It serves as a control signal to guide the model in generating specific structures and materials.
[0140] 5.2) (Topology consistency loss):
[0141] This is a loss term specifically designed for the structural characteristics of traditional Chinese style in this invention, used to penalize generated results that do not conform to physical connectivity. Its calculation formula is:
[0142]
[0143] in:
[0144] : Represents the average expected value within the current training batch;
[0145] : Represents the Euclidean distance, used to measure the difference between the topological feature field of the generated image and the topological feature field of the real image;
[0146] (DistanceTransform): Represents the Euclidean distance transformation operator. This operator calculates the distance from each foreground pixel to the nearest background pixel in the feature map, transforming the binarized skeleton line into a continuously changing gradient field (DistanceField).
[0147] (SoftSkeletonization): Represents a differentiable soft skeleton extraction operator. This operator extracts the central axis of an image through morphological erosion or a neural network-based thinning approximation.
[0148] : Represents the denoised image (or its corresponding binarization mask estimate) currently predicted by the model;
[0149] (GroundTruth): Represents the standard structural skeleton diagram corresponding to the real traditional Chinese style images in the training data;
[0150] : Represents the hyperparameter weight (WeightCoefficient) of the fracture penalty term. This is a large positive real number (e.g., This is used to adjust the proportion of the breakage penalty in the overall topology loss, forcing the model to prioritize solving the "connectivity" problem;
[0151] (BrokenComponentsCount): Represents the count of broken connected components. It discretizes physical structure errors by calculating the number of unexpected connected components in the generated image skeleton (e.g., a paper-cut pattern that should be connected becomes multiple fragments).
[0152] 5.3) (Style Gram matrix loss):
[0153] This loss is used to constrain the generated material texture. The Frobenius norm distance between the Gram matrix of each layer of the generated image's feature map and the Gram matrix of the target style is calculated. This loss ensures that the generated image possesses statistical features such as "mineral pigment graininess" and "rice paper fiber texture."
[0154] 5.4) (Loss in rating for traditional Chinese aesthetics):
[0155] We introduce a pre-trained Chinese-style aesthetic scoring network (AestheticScorer). This network is trained on a large-scale Chinese painting dataset and learns the unique scattered perspective composition and traditional color matching rules of the East.
[0156]
[0157] in:
[0158] The target aesthetic score threshold is used to guide the generated images to approach high-scoring samples (masterpieces) in the aesthetic dimension, correcting the Western aesthetic bias that the model may produce (such as focal perspective and high-saturation color clash).
[0159] : Maximum value function, takes the larger of 0 and the result calculated within the parentheses;
[0160] Aesthetic scoring function;
[0161] : Represents the denoised image (or its corresponding binarization mask estimate) currently predicted by the model.
[0162] 5.5) Weighting parameters :
[0163] The series of parameters are the weights of each loss, based on the gating coefficients in step 3. Dynamic adjustment. For example, when generating paper-cuts, the efficiency is greatly improved. The value (e.g., set to 15.0) prioritizes structural accuracy; when generating murals, this improves... The value (e.g., set to 20.0) is based on the material simulation accuracy.
[0164] Furthermore, in step 6, post-processing and optimization includes building a physically based rendering (PBR) module, specifically including:
[0165] 6.1) Reconstruction of normal and height maps:
[0166] To eliminate the "flat" look of AI-generated images, their three-dimensional geometric properties need to be reconstructed.
[0167] Using the Sobel operator or a pre-trained depth estimation network, pixel-level gradients are derived from the brightness channel of the generated image to generate a normal map. .
[0168] At the same time, based on the semantic segmentation results, different height values (HeightMap) are assigned to different regions. For example, in paper-cutting generation, the paper area is given a certain height, while the cutout area is given zero height, thereby simulating the physical thickness of the paper.
[0169] 6.2) Light and shadow simulation and subsurface scattering simulation: Construct a lighting model in a virtual three-dimensional space.
[0170] Edge projection: Based on the normal map and height map, ambient occlusion (AO) and drop shadow are calculated. Gaussian blur is applied to the shadow edges to simulate the soft shadow effect under diffuse light, giving the paper-cut lines a "floating" three-dimensional feel.
[0171] Subsurface Scattering (SSS): For materials like shadow puppets or jade, the dipole diffusion approximation model is used to simulate the scattering effect of light after it enters the medium, resulting in a translucent texture.
[0172] 6.3) Procedural weathering treatment: In order to restore the historical vicissitudes of traditional Chinese cultural relics, procedural texture generation technology is introduced.
[0173] Crack generation: A random crack network mask is generated using the Voronoi diagram (Thysen polygon) algorithm and then superimposed onto the mural image using a "multiply" mode.
[0174] Oxidation fading: A time-based color decay model was established. Based on the CIELAB color space, non-linear saturation decay was applied to specific hues (such as azurite), and a mottled texture based on Perlin noise was superimposed to simulate the oxidation and peeling effect of pigment layers over the years.
[0175] Media Simulation: Finally, a layer of high-definition scanned Xuan paper or silk fiber texture is superimposed, and the pixel position is finely adjusted using displacement mapping technology to make the image surface present a realistic paper texture, completely eliminating the smoothness of digital generation.
[0176] Through the deep collaboration of the above six steps, this invention achieves end-to-end Chinese style art generation from the data layer to the pixel layer, ensuring the high professionalism and authenticity of the generated results in the four dimensions of copyright, structure, material and aesthetics.
[0177] Specifically, the following section, with reference to the accompanying drawings, further explains a method for generating artistic images for the creation of content with a traditional Chinese style. Figure 1 As shown, the method of the present invention includes the following steps:
[0178] Step 1: Construct a multimodal graph database of traditional Chinese art based on independent copyright, and perform physical attribute decoupling preprocessing on the data. Specifically, this includes collecting and verifying intangible cultural heritage and museum collection data, establishing a whitelist mechanism, and performing structural binarization and material texture separation to ensure that the input meets physical constraints.
[0179] Step 2: Based on the decoupled data obtained in Step 1, a dual-stream perceptual feature extraction network is constructed. Specifically, a parallel processing architecture is adopted, using an improved HED network to extract structural topology flow features and a VGG network to extract texture material flow features, providing independent descriptions of "form" and "texture" for subsequent generation.
[0180] Step 3: Based on the text preprocessing results from Step 1 and the multidimensional features from Step 2, construct a dynamic conditional vector. Map implicit constraints through a physical awareness cue word adapter and employ an adaptive gating cross-attention mechanism to dynamically adjust the fusion weights of structure and material according to the target style category (such as paper-cutting or mural).
[0181] Step 4: Based on the fusion condition vector generated in Step 3, the diffusion model jointly fine-tuned by LoRA and ControlNet is invoked to generate images. This includes injecting condition vectors, applying geometric constraint control, and outputting an initial representation with specific physical properties in the latent space.
[0182] Step 5: Introduce topological consistency loss and aesthetic score loss to iteratively optimize the generation process. By calculating the distance transformation gradient and aesthetic score, adversarially correct broken lines and aesthetic biases, and output the denoised initial image.
[0183] Step 6: Based on the initial image generated in Step 5, perform Physically Based Rendering (PBR) post-processing. Through normal reconstruction, material overlay, and lighting simulation, the final output is a physically realistic Chinese-style art image.
[0184] To make the objectives, technical solutions, and effects of this invention clearer, the invention will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention. This embodiment uses "generating an artistic image of 'Tiger Descending the Mountain' with a 'non-heritage paper-cutting style'" as an example. Figure 6 Taking the example shown, this document details the specific execution process of an art image generation method for creating content with a traditional Chinese style, including six stages: graph construction, feature extraction, condition construction, model generation, iterative optimization, and physical rendering. Steps 1 and 2 are shown in the attached diagram. Figure 2 The process for step 3 is shown in the attached document. Figure 3 The loss calculation logic in step 5 is shown in the attached figure. Figure 4 As shown, the physical rendering process of the final embodiment is attached. Figure 5 As shown. Specifically:
[0185] Step 1: Decoupling of Graph Construction and Physical Properties. This step aims to obtain high-purity, copyright-free training data with clear physical properties, laying the foundation for accurate generation in the future.
[0186] 1.1) Data Collection: Access to the dedicated whitelist database for "Intangible Cultural Heritage Paper-cutting". This database contains 3,000 high-definition paper-cutting works created by collaborating national-level intangible cultural heritage inheritors. The image resolution is uniformly 4096×4096 pixels, and the format is lossless TIFF, ensuring the compliance of data sources.
[0187] 1.2) Decoupling preprocessing:
[0188] a) Binarization segmentation: Taking advantage of the high contrast of the "red paper and white background" paper-cutting, the image is converted to the LAB color space, and Otsu adaptive threshold segmentation is performed using the L channel (brightness) to obtain an accurate foreground mask.
[0189] b) Topological skeleton extraction: The Zhang-Suen thinning algorithm is applied to the foreground mask, and the edges are eroded through multiple iterations to extract the central skeleton line with a single pixel width. This skeleton line eliminates paper width interference and retains only the topological connectivity of the paper cutout.
[0190] c) Connectivity verification: The Union-Find algorithm is used to detect the number of connected components in the skeleton graph. If the number is greater than 1 (i.e., there is a break), the system automatically marks the sample as "requiring manual repair" and does not include it in the subsequent training set, ensuring the physical correctness of the training data.
[0191] Step 2: Dual-stream perceptual feature extraction. Based on the data processed in Step 1, the system starts a dual-stream network to extract features, such as... Figure 2 As shown.
[0192] 2.1) Topology Stream: The preprocessed binarized skeleton map is input into the improved HED edge detection network. To capture long-distance line connections in paper-cutting (such as the connection between a tiger's tail and body), layers 3 and 4 of the network employ dilated convolutions (dilation rates r=2,4) to expand the receptive field. The output structural feature vector... The closed loops (cutouts) and node connections of the pattern were encoded in detail.
[0193] 2.2) Material Stream: The original RGB image is input into the VGG-19 network, and shallow features from the conv1_2 layer are extracted. Its Gram matrix is calculated to capture the color value distribution of "Xuan paper red" and the graininess of the paper's microfibers, generating a material feature vector. .
[0194] Step 3: The dynamic conditional vector construction system receives the user's input text prompt: "A mighty tiger, descending the mountain, in the traditional paper-cutting style."
[0195] (Afiercetiger,descendingmountain,traditionalpaper-cutstyle), the fusion logic is as follows: Figure 3 As shown.
[0196] 3.1) Physical Awareness Adaptation: The adapter recognizes the keyword "Paper-cut," automatically retrieves the physical attribute library, and appends implicit constraint cues: "connectedlines," "binary silhouette," and "no disconnection." These cues are then converted into embedded vectors by the CLIP encoder. .
[0197] 3.2) Gating Fusion: The style classifier determines the target as "strong structure type" and outputs gating coefficients. .
[0198] 3.3) Vector Generation: According to the formula The system aggregates structural flow features with 85% weight and material flow features with 15% weight to generate the final guiding condition vector. This means that the generation process will prioritize ensuring the complete connectivity of the tiger's form, and then consider the texture of the red color.
[0199] Step 4: Model fine-tuning and generation. Based on the conditional vector generated in step 3, start the fine-tuned generation model to create images.
[0200] 4.1) Model Loading: Load the "Paper-cutting-specific LoRA" weights (Rank=32) fine-tuned based on StableDiffusionXL. These weights have been trained on a large amount of paper-cutting data in the early stage to learn the planar perspective method unique to paper-cutting.
[0201] 4.2) ControlNet guidance: Start the ControlNet branch, inject the structural features extracted in step 2 as geometric constraints into the main model, set the weight to 1.0, and force the generated contour to not deviate from the skeleton logic.
[0202] 4.3) Sampling Generation: A 50-step iterative denoising process is performed using the DPM++2MKarras sampler. An initial image latent variable with the physical properties of paper cutting is generated in the latent space.
[0203] Step 5: Iterative Optimization. In each iteration of the generated algorithm, calculate the total composite loss. Correct the results, such as Figure 4 As shown.
[0204] 5.1) Topological Consistency Constraint: Calculate the distance transform gradient field of the currently generated image. If a break occurs in the middle of the generated "tiger's tail," the distance gradient at the break point will undergo a drastic change, leading to topological loss. A surge occurs. Based on this feedback, the optimizer forces pixels to fill in the breakpoints and restore connectivity.
[0205] 5.2) Aesthetic Scoring Constraints: The traditional Chinese aesthetic scorer evaluates the image in real time. If the generated red is too fluorescent (does not conform to the traditional "Chinese red" color value), aesthetic loss occurs. It will guide the hue to shift towards true red (R=230, G=0, B=0).
[0206] Step 6: The initial image output from Physically Based Rendering (PBR) post-processing, while structurally correct, appears flat. This step gives it a realistic sense of material presence; the process is as follows... Figure 5 As shown.
[0207] 6.1) 3D reconstruction: The Sobel operator is used to calculate the gradient of the image edge and generate a normal map to simulate the small chamfers and undulations of the paper edge after being cut by scissors.
[0208] 6.2) Material overlay: Call the 4K level "Eternal Red" Xuan paper texture in the material library, and use displacement mapping technology to fine-tune the pixel position according to the texture grayscale to simulate the roughness of paper pulp fibers.
[0209] 6.3) Lighting and Shadow Simulation: Set up a 45-degree side light source in the virtual 3D environment. Calculate ambient occlusion (AO) and drop shadows based on the normal map and alpha channel. Apply Gaussian blur to the edges of the shadows to create a soft, diffused effect, producing the illusion of paper cutouts "floating" above the base.
[0210] 6.4) Final Output: The system ultimately outputs a high-definition artistic image of 4096×4096 pixels. In the image, the tiger's lines are smooth and completely connected (suitable for direct laser engraving), the paper surface shows fine plant fiber textures, and the edges are accompanied by realistic physical projections, perfectly replicating the artistic charm and physical texture of intangible cultural heritage paper-cutting.
Claims
1. A method for generating artistic images for the creation of content with a traditional Chinese style, characterized in that, Includes the following steps: Step 1: Construct a multimodal graph database of traditional Chinese art containing data on traditional Chinese art. Step 2: Construct a dual-stream perceptual feature extraction network to decouple the traditional Chinese art data and generate structural topological features, texture and material features, and semantic imagery flow features; Step 3: Construct a physical constraint-based dynamic prompt word adapter, which maps the user's input natural language description into a multimodal condition vector containing physical attribute constraints, and uses a gated cross-attention mechanism to inject the multimodal condition vector into the physical constraint-based dynamic prompt word adapter to generate the adapter model; Step 4: Using the multimodal condition vector generated in Step 3, and the structural topological features generated in Step 2 as structural conditions, as input to the diffusion generation model jointly fine-tuned by LoRA strategy and ControlNet, an initial latent space representation with the physical characteristics of Chinese style is generated. Step 5: Based on the structural topology features and semantic imagery flow features generated in Step 2, a composite total loss function is constructed using the topology consistency loss function and the traditional Chinese aesthetics scoring loss function. The initial latent space representation generated in Step 4 is then subjected to adversarial constraints and iterative denoising optimization using the composite total loss function, and the denoised initial image is output. Step 6: Using a physically based rendering post-processing enhancement module, the texture material features generated in Step 2 are used as material guides to perform texture mapping, light and shadow reconstruction and subsurface scattering simulation on the denoised initial image output in Step 5, and output the final artistic Chinese style image.
2. The method according to claim 1, characterized in that, In step 1, a multimodal graph database based on traditional Chinese art is constructed, which includes data on traditional Chinese art. Specifically, this involves scanning four types of images: paper-cutting, shadow puppetry, Dunhuang murals, and ink paintings. Semantic segmentation is performed using SAM, and an image-text-structure-material quadruple index is established by combining manual annotation. For paper cutting, extract the binary mask images of the negative and positive cuts; A model of color erosion evolution based on time was established for the Dunhuang murals.
3. The method according to claim 1, characterized in that, In step 2, the dual-stream sensing feature extraction network includes: A shallow convolutional module with shared weights and two independent deep branches connected to the shallow convolutional module, namely a structure flow branch and a material flow branch; The structured flow branch employs an HED network with introduced dilated convolutions; The material flow branch employs a Gram matrix and a Perlin noise generator connected to the Gram matrix. The Gram matrix calculates the correlation of features in each channel, and the Perlin noise generator simulates irregular weathering textures.
4. The method according to claim 1, characterized in that, In step 2, the structural topological features include: a skeleton connectivity vector based on distance transformation and an edge closure index.
5. The method according to claim 1, characterized in that, In step 2, the texture material features include: the spectral reflectance distribution of mineral pigments, the fiber texture noise map of paper / fabric, and the color attenuation matrix based on a historical weathering model.
6. The method according to claim 1, characterized in that, In step 2, the semantic imagery flow features include: high-dimensional semantic embeddings aligned with the CLIP model and the Chinese style-specific knowledge graph.
7. The method according to claim 1, characterized in that, In step 3, the calculation method for the gating cross-attention mechanism is as follows: in, This is a gated cross-attention mechanism, where Q is the semantic embedding of the user's text input; For structural flow characteristic keys, V represents the material flow feature key; V represents the value vector. This is the scaling factor; This is an adaptive gating coefficient that is automatically adjusted according to the target style.
8. The method according to claim 1, characterized in that, In step 4, the diffusion generation model is fine-tuned jointly by the LoRA strategy and ControlNet, specifically including: Based on StableDiffusionXL, the VAE encoder and the first half of the UNet backbone network are frozen. For each type of Chinese style, an independent LoRA rank matrix is trained. At the same time, an auxiliary ControlNet branch is trained to receive the topological flow features extracted in step 2 to enforce the line connectivity of the generated image.
9. The method according to claim 1, characterized in that, In step 5, based on the structural topological features and semantic imagery flow features generated in step 2, a composite total loss function is constructed using the topological consistency loss function and the traditional Chinese aesthetics scoring loss function. Specifically, this includes: the composite total loss function. The calculation formula is: in, The standard noise prediction mean square error loss for the diffusion model; The topology consistency loss function; For style Gram matrix loss; For scoring the aesthetics of traditional Chinese style; , and The weights for each type of loss.
10. The method according to claim 1, characterized in that, In step 6, a post-processing enhancement module based on physically based rendering is used. Using the texture material features generated in step 2 as material guides, texture mapping, lighting reconstruction, and subsurface scattering simulation are performed on the denoised initial image output in step 5 to output the final artistic Chinese-style image. Specifically, this includes: The post-processing enhancement module based on physically based rendering targets the paper-cutting in the denoised initial image. Using the texture material features generated in step 2 as material guides, it generates normal maps and height maps through texture mapping. Through light and shadow reconstruction, it calculates edge projection in a virtual 3D lighting environment to simulate the thickness of paper and outputs the final artistic Chinese style image. The post-processing enhancement module based on physical rendering targets the Dunhuang murals in the denoised initial image: it overlays a crack mask based on fractal noise, performs nonlinear decay simulation oxidation on color saturation, completes subsurface scattering simulation, and outputs the final artistic Chinese style image. For the shadow puppet and ink painting in the denoised initial image: no post-processing enhancement based on physical rendering is performed, and the denoised initial image output in step 5 is directly used as the final artistic Chinese style image.