Chu culture pattern diffusion generation method based on fusion of knowledge graph

CN122820904APending Publication Date: 2026-09-25WUHAN COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611156872.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-31
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供融合知识图谱的楚文化纹样扩散生成方法,可以有效解决现有技术中楚文化纹样静态呈现、知识关联碎片化、生成智能化不足的问题

Benefits of technology

本发明通过构建包含多维度语义关系的楚文化纹样知识图谱,将纹样的文化内涵、工艺特征、历史演变知识以结构化形式组织,实现了纹样知识与生成模型的有机融合,生成纹样不再仅是视觉图案的复制,而是承载深厚文化语义的知识驱动式创新生成;文化一致性保障:语义增强向量将知识图谱的实体嵌入与关系信息编码至扩散生成过程,有效增强了生成纹样与楚文化语义特征的一致性,避免了无知识约束生成导致的语义漂移问题,提升了生成纹样的文化准确性和历史真实性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820904A_ABST
    Figure CN122820904A_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence, and particularly discloses a Chu culture pattern diffusion generation method fusing a knowledge graph, which comprises the following steps: constructing a Chu culture pattern knowledge graph containing multi-dimensional semantic relationships; extracting and fusing multi-modal pattern feature representations of vision and semantics based on a convolutional neural network and a graph attention network; constructing a conditional generation model based on a denoising diffusion probability mechanism, embedding a semantic enhancement vector into a noise prediction network; and adopting an adversarial training strategy to jointly optimize generation quality. Through the technical scheme, the application can realize knowledge-driven intelligent generation of Chu culture patterns, guarantee the consistency of generated patterns and Chu culture semantics, output high-quality pattern image with cultural accuracy and artistic aesthetic feeling, and support users in customizing individualized patterns through natural language interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence, specifically involving a method for generating Chu culture patterns by integrating knowledge graphs. Background Technology

[0002] With the rapid development of information technology, digital technology is increasingly widely used in the protection and innovative inheritance of traditional culture, becoming an important means to promote the creative transformation and innovative development of outstanding Chinese traditional culture. As an important component of Chinese civilization, Chu culture, with its unique romantic style and exquisite craftsmanship, has left behind extremely rich decorative pattern resources in cultural relics such as bronzes, lacquerware, and silk fabrics. These patterns contain profound cultural connotations and unique aesthetic value, and are a precious cultural heritage of the Chinese nation. In the context of the digital age, the digital protection and innovative application of Chu cultural patterns has gradually become a research hotspot in the field of cultural inheritance. How to use artificial intelligence technology to achieve the intelligent generation and efficient dissemination of Chu cultural patterns has important theoretical significance and practical value for promoting national cultural spirit and enhancing cultural soft power.

[0003] For intelligent generation technology of cultural patterns, diffusion generation models have become a research frontier in the field of image generation due to their powerful generation capabilities and high-quality output. These models, through a progressive denoising inverse process, can generate image data conforming to a expected distribution from random noise, providing a new technical path for the intelligent creation of patterns. Meanwhile, knowledge graph technology, with its powerful knowledge representation and semantic association capabilities, demonstrates unique advantages in constructing structured knowledge systems and realizing deep semantic reasoning, providing an effective means for knowledge organization and intelligent services in the cultural field. Combining knowledge graphs with diffusion generation models can integrate rich cultural semantic knowledge while maintaining the artistic characteristics of patterns, thereby achieving innovative generation of patterns with deep cultural connotations. This technological direction is gradually attracting attention from the academic community and becoming a new research trend.

[0004] Existing technologies for processing Chu culture patterns primarily focus on static presentation and simple replication, exhibiting significant shortcomings in intelligent pattern generation and in-depth knowledge association. Specifically, while some existing solutions, such as the technical document with publication number CN306047494S, involve genealogical research on Chu culture decorative patterns, their design focus is solely on the static display and presentation of patterns, lacking dynamic generation and innovative diffusion capabilities, thus failing to meet the application needs of intelligent and innovative pattern design. Other existing solutions, such as the technical document with publication number CN303163920S, while achieving digital printing applications of Chu culture patterns, are essentially still static replication and reproduction of patterns, failing to address the intelligent generation and innovative diffusion mechanisms. More importantly, none of the aforementioned existing technologies involve the integrated application of knowledge graph technology, resulting in a fragmented organization of pattern knowledge, making it difficult to achieve structured expression and semantic association of pattern knowledge, and further unable to support intelligent reasoning and innovative generation of patterns based on deep semantic understanding. These technical deficiencies severely restrict the efficient dissemination and innovative application of Chu culture patterns in the digital age, making it difficult for existing pattern processing technologies to meet the demands of the times for intelligence, semanticization, and innovation. There is an urgent need for a new technical solution that can integrate the semantic association capabilities of knowledge graphs with the generation capabilities of diffusion models to achieve knowledge-driven intelligent generation and innovative diffusion of Chu culture patterns. Summary of the Invention

[0005] The purpose of this invention is to provide a method for generating Chu culture patterns by integrating knowledge graphs, which can effectively solve the problems of static presentation of Chu culture patterns, fragmented knowledge association, and insufficient intelligent generation in the existing technology.

[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: The method for generating Chu culture patterns by integrating knowledge graphs includes the following specific steps: Step 1: Construct a knowledge graph of Chu culture patterns; perform entity recognition and relation extraction on Chu culture decorative patterns to construct a structured knowledge graph containing multi-dimensional semantic relationships including pattern type, cultural connotation, craft characteristics, and historical evolution. Pattern entities include phoenix patterns, dragon patterns, cloud patterns, geometric patterns, and animal patterns. Relationship types include pattern association, craft technique association, historical background association, and application scenario association. Step 2: Extract pattern feature vectors; perform deep feature extraction on Chu culture pattern images based on convolutional neural networks to obtain visual feature vectors of the patterns. Simultaneously, combine these with semantic feature vectors from the knowledge graph, and generate a multi-modal pattern feature representation that integrates visual and semantic elements through a feature fusion module. Step 3: Construct a conditional diffusion generation model; construct a Chu culture pattern generation model based on a denoising diffusion probability mechanism. Step 1: Model building. The model encoder receives the multimodal pattern feature representation output from step 2 as conditional information. In the forward diffusion process, Gaussian noise is gradually added to the real pattern image, and in the reverse process, the target pattern image is gradually restored through the learned denoising network. Step 4: Generate conditional semantic enhancement. The entity embedding and relation information in the knowledge graph are encoded into semantic enhancement vectors and embedded into the noise prediction network of the diffusion model to enhance the consistency between the generated pattern and the cultural semantics, and realize the knowledge-driven innovative generation of the pattern. Step 5: Iterative optimization and pattern output. The adversarial training strategy is used to jointly optimize the generation quality of the diffusion model. The realism and cultural feature matching degree of the generated pattern are evaluated by the pre-trained discriminator. The model parameters are iteratively updated until convergence, and a high-quality pattern image that conforms to the semantic features and aesthetic style of Chu culture is output.

[0007] In some preferred embodiments, step 1: constructing a knowledge graph of Chu culture patterns; the construction of the knowledge graph adopts the form of triples to represent pattern knowledge, the subject is the pattern entity, the predicate is the relationship between entities, and the object is the associated entity or attribute value. The knowledge graph scale includes a preset number of pattern entity nodes and a preset number of semantic relationship edges, covering the complete knowledge system of patterns of Chu bronzes, lacquerware, and silk artifacts.

[0008] In some preferred embodiments, step 1: constructing a knowledge graph of Chu culture patterns; entity recognition adopts a conditional random field algorithm based on bidirectional long short-term memory network to perform sequence annotation on relevant literature and image annotations of Chu culture patterns, and the recognition accuracy reaches the preset accuracy requirements; relation extraction adopts a graph convolutional network with attention mechanism enhancement to realize the automatic extraction of multi-level semantic relations between pattern entities.

[0009] In some preferred embodiments, step 2: extracting pattern feature vectors; visual feature extraction uses a residual network as the backbone network to output a preset dimension depth feature vector; semantic feature extraction uses a graph attention network to perform message passing on the knowledge graph to output a preset dimension semantic embedding vector; the feature fusion module uses a cross-modal attention mechanism to fuse visual features and semantic features to generate a preset dimension multimodal pattern feature representation.

[0010] In some preferred embodiments, step 3: constructing a conditional diffusion generation model; the diffusion model adopts a diffusion time step of a preset number of steps, the noise prediction network adopts a U-shaped network structure, including a preset number of downsampling stages and a preset number of upsampling stages, the number of feature channels is gradually expanded from a preset initial value to a preset maximum value, and each stage uses residual connections and group normalization to stabilize the training process.

[0011] In some preferred embodiments, in step 3: constructing a conditional diffusion generation model, the conditional information is injected into each stage of the denoising network through an adaptive layer normalization module. The conditional embedding and the time step embedding are input into the network after being concatenated by dimension. The mean square error between the predicted noise and the real noise output by the decoder is used as the basic loss function of the diffusion model.

[0012] In some preferred embodiments, step 4: generating conditional semantic enhancement; the construction process of semantic enhancement vectors includes: mapping entity identifiers in the knowledge graph to a low-dimensional dense vector space, with the vector dimension being a preset dimension, the relation type being mapped to a learnable relation matrix, training the vector representations of entities and relations through a knowledge graph embedding algorithm, and splicing and fusing the semantic enhancement vectors with noisy image features in the channel dimension.

[0013] In some preferred embodiments, step 5: iterative optimization and pattern output; the adversarial training adopts a relativistic discriminator structure, the discriminator judges the realism probability of the generated pattern relative to the real pattern, the generator and discriminator are optimized alternately, the generator loss includes three components: diffusion reconstruction loss, semantic consistency loss and adversarial loss, the semantic consistency loss is determined by calculating the cosine similarity between the generated pattern features and the target semantic embedding.

[0014] In some preferred embodiments, a pattern post-processing step is also included; the pattern post-processing step performs super-resolution reconstruction on the output pattern image of step 5: iterative optimization and pattern output, improves the image resolution to a preset resolution, and at the same time uses a color correction algorithm to adjust the visual presentation of the pattern to make it more in line with the typical color system of Chu culture. The color correction adopts an end-to-end color mapping network based on convolutional neural network.

[0015] In some preferred embodiments, a generation quality assessment step is also included; the generation quality assessment step evaluates the pattern image output by the pattern post-processing step, and evaluates the quality of the generated pattern from two dimensions: cultural accuracy and artistic aesthetics. The cultural accuracy assessment indicators include pattern type recognition accuracy and semantic matching degree, and the artistic aesthetics assessment adopts a pre-trained pattern aesthetics evaluation network. The comprehensive quality score of the generated pattern is obtained by weighted fusion of the evaluation results of the two dimensions.

[0016] In some preferred embodiments, a user interaction step is also included; the user interaction step allows the user to specify the cultural theme, application scenario and style preference of the generated pattern through natural language description. The natural language input is converted into a semantic vector through a pre-trained language encoder. The semantic vector, as additional conditional information, is input into the diffusion model together with the multimodal pattern feature representation output in step 2, guiding the diffusion model to generate a pattern image that conforms to the user's intention.

[0017] Compared with the prior art, the present invention has the following beneficial effects: This invention constructs a knowledge graph of Chu culture patterns containing multi-dimensional semantic relationships, organizing the cultural connotations, craft characteristics, and historical evolution of the patterns in a structured form. This achieves the organic integration of pattern knowledge and the generation model, so that the generated patterns are no longer just a copy of visual patterns, but a knowledge-driven innovative generation carrying profound cultural semantics. Cultural consistency is guaranteed: the semantic enhancement vector encodes the entity embedding and relational information of the knowledge graph into the diffusion generation process, effectively enhancing the consistency between the generated patterns and the semantic features of Chu culture, avoiding the semantic drift problem caused by generation without knowledge constraints, and improving the cultural accuracy and historical authenticity of the generated patterns.

[0018] This invention utilizes a generative model based on a denoising diffusion probability mechanism. Through a progressive denoising reverse process, it recovers the target pattern image from random noise. The generation process is stable and controllable, and the output image quality is significantly better than that of traditional generative adversarial networks. It can generate high-quality Chu culture pattern images with rich details and clear textures. Multimodal feature fusion: The multimodal pattern feature representation, which integrates visual and semantic features, takes into account both the external visual aesthetics and the internal cultural semantics of the pattern. The generated pattern has the typical aesthetic features of traditional Chu culture patterns at the visual level and accurately carries the corresponding cultural connotation information at the semantic level.

[0019] This invention automates the entire process from knowledge graph construction to pattern generation and output, enabling knowledge-driven intelligent creation of patterns without human intervention. This significantly reduces the threshold and cost of innovative design of Chu culture patterns and improves the efficiency of pattern design and creation. User intent understanding: The natural language interaction process allows users to specify their generation needs through language description. The language encoder converts the user intent into semantic vectors to guide the generation process, realizing intelligent pattern customization based on cultural semantic understanding and meeting the personalized pattern generation needs of different application scenarios. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.

[0021] Figure 1 is a flowchart of the main process of the Chu culture pattern diffusion generation method that integrates knowledge graphs according to the present invention.

[0022] Figure 2 is a flowchart of the sub-process of constructing the Chu culture pattern knowledge graph of the present invention.

[0023] Figure 3 is a flowchart of the adversarial training iterative optimization sub-process of the present invention. Detailed Implementation

[0024] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the embodiments of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0025] The following is in conjunction with the appendix Figures 1-3 The embodiments of the present invention will be described in detail below.

[0026] Example 1: This embodiment discloses a method for generating Chu culture patterns by fusing knowledge graphs. The method achieves a structured expression of pattern knowledge by constructing a Chu culture pattern knowledge graph, generates a multimodal pattern feature representation that integrates visual and semantic elements by combining multimodal feature extraction and fusion techniques, constructs a conditional generation model based on a denoising diffusion probability mechanism, and encodes semantic enhancement vectors through knowledge graph embedding technology. Finally, the generation quality is optimized through an adversarial training strategy to output a high-quality pattern image that conforms to the semantic features and aesthetic style of Chu culture.

[0027] Specifically as follows: Step 1: Construct a knowledge graph of Chu culture patterns; organize and express the multi-dimensional knowledge of Chu culture decorative patterns in a structured form to provide semantic support and knowledge constraints for subsequent pattern generation. The knowledge graph uses the classic triple form to represent pattern knowledge, namely, Head-Relation-Tail, where the head is the pattern entity node, the predicate represents the semantic relationship between entities, and the tail is another related entity node or a specific attribute value.

[0028] In the specific implementation, step 1: Construct a knowledge graph of Chu culture patterns. The constructed knowledge graph covers the knowledge system of decorative patterns from different historical periods and artifact types of the Chu state. The core pattern entities include the phoenix pattern series, covering Chu phoenix patterns, coiled phoenix patterns, and variant forms of phoenix patterns with outstretched wings; the dragon pattern series, covering dragon head patterns, dragon body coiled patterns, and dragon and phoenix combined patterns; the cloud pattern series, covering cloud and thunder patterns, swirling cloud patterns, and flowing cloud patterns; the geometric pattern series, covering rhombus patterns, triangle patterns, and circle patterns; and the animal pattern series, covering tiger patterns, deer patterns, and monster patterns. The above pattern entity nodes are interconnected through a preset number of semantic relationship edges to form a complete knowledge network structure.

[0029] The relationship types in the knowledge graph include the following four main dimensions: pattern association, which describes the morphological relationship and evolution between different patterns, such as the combination relationship between phoenix patterns and dragon patterns, and the relationship between geometric patterns as background patterns and main patterns; craftsmanship association, which describes the correspondence between patterns and production techniques, such as the relationship between the coiled dragon pattern on bronzes and the casting process, and the relationship between painted patterns on lacquerware and the lacquering process; historical context association, which describes the historical period and cultural background in which the patterns originated, such as the evolution from the simple style of patterns in the early Warring States period to the elaborate style of patterns in the middle and late Warring States period; and application scenario association, which describes the application position and function of patterns on actual objects, such as the main decorative patterns on the body of bronze vessels and the auxiliary decorative patterns on the feet of the vessels.

[0030] The entity recognition process uses the BiLSTM-CRF algorithm based on bidirectional long short-term memory network as the core technical framework.

[0031] In the specific implementation, the following steps are taken: First, ancient textual records, unearthed artifact reports, and textual materials of pattern illustrations related to Chu culture patterns are collected and organized, along with an image annotation dataset of Chu culture patterns. After preprocessing steps including word segmentation, part-of-speech tagging, and named entity recognition, the text data is labeled with pattern entities according to the BIO annotation system, where B-PAT represents the starting position of the pattern entity, I-PAT represents the continuation position, and O represents a non-entity part. The labeled sequence data is input into a bidirectional long short-term memory network (LSTM). The network contains a predetermined number of hidden layer units, each with three gating mechanisms: a forget gate, an input gate, and an output gate, effectively capturing long-distance dependencies in the sequence data. The bidirectional structure allows the network to utilize both forward and backward contextual information, improving the accuracy of entity recognition. A conditional random field layer is located above the output layer of the bidirectional LSTM network, ensuring the sequence validity and label consistency of the entity recognition results by modeling the transition probabilities between label sequences. Once the entity recognition accuracy reaches the predetermined precision requirement, the recognition results are converted into entity nodes in a knowledge graph.

[0032] The relation extraction process employs an attention-based graph convolutional network (Graph Convolutional Network) to automatically extract multi-level semantic relationships between pattern entities. In the specific implementation, the identified pattern entities are treated as nodes in the graph, and an initial graph structure is constructed based on co-occurrence and syntactic dependency relationships in the text corpus. The graph convolutional network iteratively updates node features through multiple layers of graph convolution operations. Each convolution operation can be represented as:

[0033] In the formula, Indicates the first Layer nodes Feature representation; Represents a node The set of neighboring nodes; For the first Layer weight matrix; For the first Layer bias vector; It is a non-linear activation function; These are attention weight coefficients, calculated using a self-attention mechanism, used to measure the neighboring nodes. For the central node The extent of their contribution.

[0034] After the multi-layer graph convolution operation is completed, the node features aggregate information from multi-hop neighbor nodes, which can effectively capture the deep semantic relationships between pattern entities. Combined with the preset relation classification threshold, the extracted relation types are mapped to the above four relation dimensions, and finally the predicate edges in the knowledge graph are constructed.

[0035] Step 2: Extract pattern feature vectors; the goal is to extract visual and semantic feature vectors from Chu culture pattern images and knowledge graphs, respectively, and generate a multimodal pattern feature representation that integrates visual and semantic elements through a feature fusion module. This multimodal pattern feature representation serves as the conditional information input for the subsequent diffusion generation model, determining the visual form and cultural connotation of the generated pattern.

[0036] Visual feature extraction employs a Residual Network (ResNet) as its backbone network structure. The Residual Network effectively addresses the vanishing and exploding gradient problems during deep network training by introducing residual connections, enabling the extraction of rich multi-scale visual features from Chu culture pattern images. Specifically, using an RGB format Chu culture pattern image as input, the image first undergoes initial feature extraction through a convolutional layer with a kernel size of 7×7 and a stride of 2, outputting the preset initial number of channels. Subsequently, a max-pooling layer is used for downsampling to reduce the feature map size. The main feature extraction network consists of a preset number of stacked residual blocks. Each residual block contains two convolutional layers with a kernel size of 3×3, followed by a batch normalization layer and a ReLU activation function. Residual connections within the residual blocks directly sum the inputs to the output. The total depth of the Residual Network is determined based on preset hyperparameters, using either a ResNet-50 or ResNet-101 architecture. The network can output a multi-scale depth feature map with 2048 channels. To obtain a fixed-dimensional visual feature vector, the depth feature map is processed by a global average pooling layer and a fully connected layer, and the output is a depth feature vector with a preset dimension. This vector encodes the shape contour, texture details, and color distribution visual information of the pattern image.

[0037] Semantic feature extraction employs a Graph Attention Network (GraphAttentionNetwork) for message passing and feature aggregation within the knowledge graph. The GraphAttention Network calculates the contribution weights of different neighboring nodes to the central node through an attention mechanism, effectively capturing heterogeneous relationships and complex topological structures within the knowledge graph. In its implementation, the set of entity nodes in the knowledge graph is represented as... Each entity node The initial features are those of the entity randomly initialized in the knowledge graph embedding space. 1D vector. The entity node features are input into a multi-layer graph attention network. The computation process of each layer includes the following steps: First, compute node pairs... Attention coefficient between:

[0038] In the formula, For nodes Input feature representation; For nodes Input feature representation; For shared linear transformation matrices; This represents the weight vector for the attention mechanism; This represents a vector concatenation operation; LeakyReLU is a linear rectified activation function with leakage.

[0039] Then, the attention coefficients are normalized using the Softmax function to obtain the normalized attention weights. Finally, the features of neighboring nodes are weighted and aggregated according to the normalized attention weights to update the representation of the center node:

[0040] In the formula, Representative node The set of neighboring nodes; To normalize attention weights; Represents a non-linear activation function; To share the linear transformation matrix.

[0041] After message passing and feature aggregation through a multi-layer graph attention network, the feature representation of each entity node aggregates the information of its multi-hop neighbor nodes, and outputs a semantic embedding vector with a preset dimension. This vector encodes the semantic association information of the pattern entity in the knowledge graph.

[0042] The feature fusion module employs a cross-modal attention mechanism to achieve deep fusion of visual and semantic features. This mechanism calculates the similarity matrix between visual and semantic features to uncover the correspondence and complementary information between the two modalities.

[0043] In the specific implementation, let the visual feature vector be... The semantic feature vector is First, two vectors are mapped to a common space of the same dimension using linear projection, resulting in... and Then, calculate the cross-modal attention weights:

[0044] In the formula, As a characteristic dimension of public space; The visual features after projection; These are the semantic features after projection.

[0045] Visual to semantic attention weights This represents the degree to which semantic features contribute to visual features, and the attention weight from semantics to vision. Obtained through symmetric computation. The final multimodal pattern feature representation is achieved by fusing visual features, semantic features, and cross-modal attention information using bilinear interpolation.

[0046] In the formula, The visual features after projection; These are the semantic features after projection; For visual-to-semantic cross-modal attention weights; Semantic-to-visual cross-modal attention weights; These are learnable fusion weight coefficients; This represents the multimodal pattern features.

[0047] Output a multimodal pattern feature representation with a preset dimension. , Simultaneously encodes the visual morphological information of the pattern and the semantic association information in the knowledge graph.

[0048] Step 3: Construct a conditional diffusion generation model; Construct a conditional generation model for Chu culture patterns based on the denoising diffusion probabilistic mechanism. This model gradually adds Gaussian noise to the real pattern image through a forward diffusion process, and gradually recovers the target pattern image from the noise through a reverse denoising process, thereby achieving the generation of high-quality pattern images.

[0049] The overall time step of the diffusion model is set to a preset number of steps. The forward diffusion process involves adding fixed Gaussian noise. Given a real image of a Chu culture pattern... The forward process is in Time to Image Adding Gaussian noise yields This process can be represented as:

[0050] In the formula, For the first The noise variance scheduling parameter at time intervals, whose value starts from a preset initial value. Gradually increase to The noise variance scheduling employs a cosine scheduling strategy to ensure the stability of the diffusion process. It is an identity matrix.

[0051] go through After the first diffusion step, at any time Image You can start directly from the initial image The sampling yielded:

[0052] In the formula, ; ; It is an identity matrix.

[0053] The reverse denoising process uses a learned denoising network. From noisy images Gradually restore a near-realistic pattern image ,in For conditional information, This is equivalent to step 2 extracting the pattern feature vector; the output is a multimodal pattern feature representation. The denoising network adopts a U-Net structure, which, through an encoder-decoder architecture and a skip connection mechanism, can effectively capture multi-scale feature information and achieve accurate noise prediction.

[0054] The specific configuration of the U-shaped network is as follows: The encoder part contains a preset number of downsampling stages. Each downsampling stage consists of two convolutional layers and a downsampling operation. The convolutional layers use a 3×3 kernel with padding of 1, combined with group normalization and SiLU activation function, gradually doubling the number of feature channels from a preset initial value to a preset maximum value. The downsampling operation uses a 3×3 kernel with a stride of 2 to halve the feature map size. The decoder part contains a corresponding number of upsampling stages. Each upsampling stage consists of an upsampling operation, skip connections, and two convolutional layers. The upsampling operation uses nearest neighbor interpolation or transposed convolution to double the feature map size. Skip connections are used between the encoder and decoder to pass multi-scale feature information, ensuring that the network can simultaneously utilize shallow detailed features and deep semantic features. Residual connections are used within each convolutional block to stabilize the training process, directly adding the input to the convolutional output. In addition, a self-attention mechanism module is added between corresponding layers of the encoder and decoder to capture long-distance dependencies in the feature maps.

[0055] The conditional information injection mechanism is a key technology in conditional diffusion generative models, determining the alignment between the generated result and the conditional information. Conditional information is injected into each stage of the denoising network through the Adaptive Layer Normalization module. Adaptive Layer Normalization embeds the conditional information... As affine transformation parameters, the conditional embeddings are converted into layer-normalized scaling parameters through a learned mapping network. and offset parameters :

[0056] In the formula, This is the scale-mapped weight matrix; This is the offset mapping weight matrix; This is the scale mapping bias vector; For offset mapping bias vector; This represents the multimodal pattern features.

[0057] The layer normalization operation is modified as follows:

[0058] In the formula, Features of the intermediate layer; The mean of the intermediate layer features; The variance of the intermediate layer features; To prevent division by zero of small constants; , These are the affine transformation parameters.

[0059] In addition, the temporal embedding and conditional embedding are jointly input into the temporal conditional module of the network after dimensional concatenation. This module maps the concatenated vector to the network's temporal-related features through a multilayer perceptron.

[0060] The fundamental loss function of the diffusion model is defined as the mean square error between the predicted noise output by the decoder and the actual noise:

[0061] In the formula, The actual Gaussian noise added during the forward pass; This is the noise prediction output of the denoising network; To obtain from a uniform distribution The sampling time step; This represents the multimodal pattern feature representation.

[0062] Step 4: Generate conditional semantic enhancement; Encode entity embeddings and relational information in the knowledge graph into semantic enhancement vectors, embed them into the noise prediction network of the diffusion model, enhance the consistency between generated patterns and cultural semantics, and realize knowledge-driven innovative generation of patterns.

[0063] The construction process of semantically enhanced vectors is based on Knowledge Graph Embedding (KPE). The goal of KPE is to map entities and relations in a knowledge graph to a low-dimensional dense vector space, enabling semantic relations in the graph to be expressed and reasoned about through vector operations. Specifically, TransE or its variants are used as the KPE algorithm. The core idea of ​​this algorithm is to make triples... Head entity embedding With relational embedding The sum is close to the tail entity embedding ,Right now The loss function is defined as:

[0064] In the formula, The set of positive sample triples; This is a set of negative sample triples generated by randomly replacing the head or tail entity; For interval parameters; For a marginal-based scoring function, TransE .

[0065] By optimizing the loss function through end-to-end training, vector representations of all entities and relations in the knowledge graph are obtained.

[0066] For each pattern category to be generated, the corresponding entity embedding is first retrieved from the knowledge graph. and related attribute embedding sets The entity embedding and attribute embedding are weighted and fused to generate the initial semantic enhancement vector:

[0067] In the formula, The learnable projective weight matrix; For learnable projection bias vectors; This represents a vector concatenation operation; For the first The fusion weight of each attribute; Embed the pattern entity; For the first Embedded related attributes; This is a semantic enhancement vector.

[0068] The semantic enhancement vector and noisy image features are concatenated and fused along the channel dimension. Specifically, the input features of the denoising network... With semantic enhancement vectors Perform splicing along the channel dimension:

[0069] In the formula, Characterize the noise input to the network; For semantic enhancement vectors; Represents the splicing of channel dimensions; This is a splicing and fusion feature.

[0070] The splicing and fusion feature has a channel dimension length equal to the sum of the original number of channels and the dimension of the semantic enhancement vector. A convolutional layer then restores the number of channels to their original value, achieving deep fusion of semantic information and image features. The splicing and fusion feature is input to each layer of the denoising network, guiding the network to generate patterns that conform to cultural semantic features during the denoising process.

[0071] Step 5: Iterative optimization and pattern output; an adversarial training strategy is adopted to jointly optimize the generation quality of the diffusion model. The realism and cultural feature matching degree of the generated patterns are evaluated by a pre-trained discriminator. The model parameters are iteratively updated until convergence, and finally, a high-quality pattern image that conforms to the semantic features and aesthetic style of Chu culture is output.

[0072] Adversarial training employs a relativistic discriminator structure. This discriminator not only determines whether the input is a real or generated pattern, but also assesses the difference in realism probability between the generated pattern and the real pattern. The output of the relativistic discriminator is defined as:

[0073] In the formula, To enable the discriminator to identify the true pattern and generate patterns The relative realism assessment; For discriminator networks; Use the Sigmoid activation function; This represents the expectation of the generated pattern distribution.

[0074] The goal of the discriminator is to maximize its ability to distinguish between real and generated patterns, while the goal of the generator is to simultaneously minimize the realism of its own patterns and maximize the realism of the real patterns.

[0075] The generator loss consists of three components: diffusion reconstruction loss. Semantic consistency loss and combat losses The diffusion reconstruction loss, also known as the basic diffusion loss defined in step 3: constructing the conditional diffusion generation model, is the diffusion reconstruction loss. This is used to ensure the noise prediction accuracy of the diffusion model. The semantic consistency loss is determined by calculating the cosine similarity between the generated pattern features and the target semantic embedding.

[0076] In the formula, For pre-trained pattern feature extraction network; To generate a pattern image; The target semantic embedding vector; This represents the vector dot product operation.

[0077] Combating losses Defined as the relative hinge loss form:

[0078] In the formula, This is a real pattern sample; Generate pattern samples for the model; This is the output of the relativistic discriminator.

[0079] The generator's total loss function is a weighted sum of three components:

[0080] In the formula, These are weighting coefficients, which are adjusted according to the training phase and task requirements. This represents the total loss of the generator.

[0081] The training process employs a strategy of alternating optimization of the generator and discriminator. In each training iteration, the generator parameters are first fixed, and the discriminator parameters are updated a preset number of times; then the discriminator parameters are fixed, and the generator parameters are updated a preset number of times. The model parameters are iteratively updated until the diffusion reconstruction loss and semantic consistency loss converge. At this point, the generator can stably output high-quality pattern images that conform to the semantic features and aesthetic style of Chu culture.

[0082] The pattern post-processing step iteratively optimizes and outputs the pattern from step 5; the output pattern image is then optimized. This step includes two sub-steps: super-resolution reconstruction and color correction. Super-resolution reconstruction employs an end-to-end super-resolution network based on convolutional neural networks to upscale the low-resolution generated pattern image to a preset resolution. The network's feature extraction stage uses a residual dense block structure, fully utilizing feature information from each level through skip connections to achieve high-quality reconstruction of pattern details. The color correction sub-step uses an end-to-end color mapping network based on convolutional neural networks to map the generated pattern image to the typical color system of Chu culture. By learning the color distribution patterns of typical Chu culture patterns, the color mapping network automatically adjusts the hue, saturation, and brightness of the generated pattern to better conform to the visual aesthetic characteristics of Chu culture.

[0083] The quality assessment process evaluates the generated patterns from two dimensions: cultural accuracy and artistic aesthetics. Cultural accuracy assessment includes two metrics: pattern type recognition accuracy and semantic matching degree. Pattern type recognition accuracy is determined by classifying the generated patterns using a pre-trained pattern classifier, judging the consistency between the classification result and the target pattern type. Semantic matching degree is quantitatively assessed by calculating the cosine similarity between the semantic feature vector of the generated pattern and the target semantic embedding. Artistic aesthetics assessment employs a pre-trained pattern aesthetic evaluation network, trained on a large-scale pattern aesthetic annotation dataset, capable of outputting aesthetic scores for the patterns.

[0084] The overall quality score of the generated pattern is obtained by weighted integration of the evaluation results from two dimensions: cultural accuracy and artistic aesthetics.

[0085] In the formula, For overall quality scoring; Rate the cultural accuracy; Rate the aesthetic appeal of the artwork; and The weighting coefficients are and satisfy the following conditions: .

[0086] Patterns whose overall quality score is below a preset threshold will be marked and fed back to the generation model for regeneration.

[0087] The user interaction steps allow users to specify the cultural theme, application scenario, and style preference of the generated pattern through natural language descriptions. The natural language descriptions input by users first undergo text preprocessing, including word segmentation, named entity recognition, and intent recognition operations.

[0088] Natural language input is converted into dense semantic vectors by a pre-trained language encoder (such as BERT). These semantic vectors serve as additional conditional information and are combined with the pattern feature vectors extracted in step 2. The output multimodal pattern feature representations are then concatenated, and the concatenated results are fed into the diffusion generation model. The language encoder accurately captures key semantic information in the user's description, such as "phoenix pattern," "bronze style," and "red as the main color," and encodes this semantic information into high-dimensional semantic vectors. During the generation process, the diffusion model simultaneously considers the semantic constraints of the pattern knowledge graph and the multimodal pattern feature representations. It satisfies the user's specified generation requirements at the semantic level and maintains the typical aesthetic features of Chu culture patterns at the visual level, achieving intelligent pattern customization based on cultural semantic understanding.

[0089] In a specific implementation, the method of the present invention is illustrated by taking the generation of a three-phoenix pattern with the style of bronze ware from the Chu state in the mid-Warring States period as an example.

[0090] First, based on the user input "generate a three-phoenix pattern in the style of Chu bronze ware from the mid-Warring States period," the user interaction steps convert the natural language description into a semantic vector. This vector encodes the user's specific requirements for the pattern theme (three phoenixes), the historical period (mid-Warring States period), and the type of artifact (bronze).

[0091] Step 1: Construct a knowledge graph of Chu culture patterns; retrieve relevant pattern entity nodes based on semantic vectors. The search results show that the triple phoenix pattern is associated with the following entities in the knowledge graph: the phoenix pattern as the core pattern entity; its historical background is associated with the prosperous period of bronze ware in the mid-Warring States period; its craftsmanship is associated with casting and gold and silver inlay techniques; and its application scenario is associated with the decoration of bronze ritual vessels. Based on the search results, construct a sub-knowledge graph of the triple phoenix pattern, containing a preset number of entity nodes and a preset number of relational edges.

[0092] Step 2: Extract pattern feature vectors; process two inputs simultaneously: one is a reference pattern image uploaded by the user or a typical three-phoenix pattern image retrieved from the knowledge graph, which is processed by a residual network to obtain a 512-dimensional visual feature vector; the other is the entities and relations of the sub-knowledge graph, processed by a graph attention network to output a 512-dimensional semantic embedding vector. The visual feature vector and the semantic embedding vector are fused through a cross-modal attention mechanism to generate a 512-dimensional multimodal pattern feature representation. .

[0093] Step 3: Construct a conditional diffusion generation model; represent it using multimodal pattern features. As conditional information, a diffusion time step of 1000 steps is used for generation. The U-shaped network injects conditional information through adaptive layer normalization in each denoising stage, gradually recovering the target pattern image from the random noise image. The visual features in the conditional information guide the network to generate the shape and outline of the phoenix pattern, while the semantic features guide the network to generate the phoenix shape and decorative details that conform to the aesthetics of Chu culture.

[0094] Step 4: Generate conditional semantic enhancement; Based on the entity embedding of the triple phoenix pattern, generate a 512-dimensional semantic enhancement vector. The semantic enhancement vector and the features of the noisy image are concatenated and fused at the channel dimension to obtain the concatenated and fused features, which are then input into the denoising network to enhance the consistency of the semantic features of the generated pattern with the phoenix pattern of the mid-Warring States period, ensuring that the generated triple phoenix pattern is consistent with the typical style of that period in terms of form.

[0095] Step 5: Iterative optimization and pattern output; a pre-trained Chu culture pattern discriminator is used to evaluate the generation quality. The discriminator determines the probability of realism of the generated three-phoenix pattern relative to real bronze artifact patterns, and calculates the semantic consistency loss to ensure the alignment of the generated pattern with the target semantics. The generator and discriminator are optimized alternately until the generated three-phoenix pattern reaches a preset threshold in the overall quality score.

[0096] The post-processing step reconstructs the generated pattern using a 4x super-resolution method, increasing the image resolution from 512×512 to 2048×2048, significantly enhancing the clarity of pattern details. A color correction network adjusts the hue distribution of the pattern according to the typical color system of mid-Warring States period bronzes, enhancing the dominant patina of the copper-green color.

[0097] The final output image of the three phoenixes combines the typical stylistic features of bronzes from the mid-Warring States period with knowledge-driven innovative design elements, achieving a high level of quality in both cultural accuracy and artistic aesthetics.

[0098] Example 2: This example describes an alternative implementation of a conditional diffusion generation method based on graph convolutional network enhancement, applicable to the task of generating geometric patterns in Chu culture, as follows: In Step 1, constructing the Chu culture pattern knowledge graph, a graph isomorphism network is used instead of a graph convolutional network for relation extraction, taking into account the characteristics of geometric patterns. The graph isomorphism network aggregates neighbor node features through a multilayer perceptron, avoiding the bias caused by the graph structure on node features in graph convolutional networks. This allows for more accurate extraction of the structural relationships between basic and composite units in geometric patterns. For example, the basic unit of a rhombus pattern is a single rhombus, and the composite unit is a complex pattern formed by rotating, translating, and mirroring rhombuses. The graph isomorphism network can effectively learn the semantic connections corresponding to these transformation relationships.

[0099] In step 2, extracting the pattern feature vector, EfficientNet is used as the backbone network for visual feature extraction. EfficientNet adjusts the network depth, width, and input resolution simultaneously through a composite scaling strategy, improving feature extraction capabilities while maintaining model efficiency. The visual feature vector output by EfficientNet is fused with the semantic feature vector output by the graph attention network through a gating attention mechanism.

[0100] In the formula, The gated vector is constrained to the interval between 0 and 1 by the Sigmoid activation function; This represents element-wise multiplication. Represents vector concatenation; This is the weight matrix of the gated network; For visual feature vectors; It is a semantic feature vector; These are the features after fusion.

[0101] Gating mechanisms can adaptively adjust the contribution ratio of visual features and semantic features in the fused features.

[0102] In step 3, constructing the conditional diffusion generative model, the encoder of the U-shaped network employs a Squeeze-and-Excitation module to enhance the channel attention mechanism. The Squeeze-and-Excitation module compresses the spatial dimension through global average pooling, utilizes fully connected layers to learn the dependencies between channels, and adaptively adjusts the weights of each channel's features. This mechanism helps the model focus on key structural features within the geometric pattern.

[0103] In step 4: generating conditional semantic enhancement, a symmetry constraint loss is introduced to address the symmetry characteristics of geometric patterns.

[0104] In the formula, This indicates an operation to flip the image along a specified axis; Output pattern images for the model; The loss is due to symmetry constraints.

[0105] The symmetry constraint loss and semantic consistency loss are jointly optimized to ensure that the generated geometric patterns maintain good symmetry features.

[0106] This embodiment uses the application scenario of generating a geometric pattern from Chu culture as an example to verify the effectiveness of the alternative technical solution. The user inputs "generate a geometric pattern from Chu culture for silk fabric decoration," and the system automatically retrieves the geometric pattern entities and their association with the silk fabric application scenario from the knowledge graph. The generation model combines the structured feature constraints of the geometric pattern with the actual needs of the silk fabric application to output a geometric pattern suitable for silk fabric decoration. The symmetry score and visual quality score of the generated pattern both meet the preset requirements, verifying the applicability of the alternative solution in the geometric pattern generation task.

[0107] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0108] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating Chu culture patterns by integrating knowledge graphs, characterized in that, Includes the following steps: Step 1: Construct a knowledge graph of Chu culture patterns; perform entity recognition and relationship extraction on Chu culture decorative patterns, and construct a structured knowledge graph containing multi-dimensional semantic relationships including pattern type, cultural connotation, craft characteristics, and historical evolution; Step 2: Extract pattern feature vectors; Based on convolutional neural networks, perform deep feature extraction on Chu culture pattern images to obtain visual feature vectors of the patterns. At the same time, combine the semantic feature vectors in the knowledge graph, and generate a multimodal pattern feature representation that integrates visual and semantic features through the feature fusion module. Step 3: Construct a conditional diffusion generation model; Construct a conditional generation model based on a denoising diffusion probability mechanism. The model encoder receives multimodal pattern feature representations as conditional information. In the forward diffusion process, Gaussian noise is gradually added to the target pattern image. In the reverse process, the target pattern image is gradually restored through a learned denoising network. The denoising network adopts a U-shaped network structure, which includes multiple downsampling and upsampling stages. Each stage uses residual connections and group normalization to stabilize the training process. The conditional information is injected into each stage of the denoising network through an adaptive layer normalization module. The conditional embedding and the temporal step embedding are input into the network after dimensional concatenation. The mean square error between the predicted noise and the real noise output by the decoder is used as the basic loss function of the diffusion model. Step 4: Generate conditional semantic enhancement; Encode entity embeddings and relational information in the knowledge graph into semantic enhancement vectors, embed them into the noise prediction network of the diffusion model, and splice and fuse the semantic enhancement vectors with the noise image features in the channel dimension to enhance the consistency between the generated patterns and cultural semantics, and realize the knowledge-driven innovative generation of patterns. Step 5: Iterative optimization and pattern output; an adversarial training strategy is adopted to jointly optimize the generation quality of the diffusion model. The realism and cultural feature matching degree of the generated patterns are evaluated by a pre-trained discriminator. The model parameters are iteratively updated until convergence, and high-quality pattern images that conform to the semantic features and aesthetic style of Chu culture are output.

2. The method for generating Chu culture patterns by integrating knowledge graphs according to claim 1, characterized in that, Step 1: Constructing a knowledge graph of Chu culture patterns; The construction of the knowledge graph adopts a triplet form to represent pattern knowledge, with the subject being the pattern entity, the predicate being the relationship between entities, and the object being the associated entity or attribute value. The knowledge graph covers the knowledge system of decorative patterns of different historical periods and different types of artifacts in the Chu state.

3. The method for generating Chu culture patterns by integrating knowledge graphs according to claim 1, characterized in that, Step 1: Construct a knowledge graph of Chu culture patterns; entity recognition uses a conditional random field algorithm based on bidirectional long short-term memory network for sequence labeling to realize the recognition of pattern entities; relation extraction uses a graph convolutional network with attention mechanism enhancement to realize the automatic extraction of multi-level semantic relations between pattern entities.

4. The method for generating Chu culture patterns by integrating knowledge graphs according to claim 1, characterized in that, Step 2: Extracting pattern feature vectors; visual feature extraction uses a residual network as the backbone network, semantic feature extraction uses a graph attention network to perform message passing and feature aggregation on the knowledge graph, and feature fusion. The module employs a cross-modal attention mechanism to fuse visual and semantic features, generating multimodal pattern feature representations.

5. The method for generating Chu culture patterns by integrating knowledge graphs according to claim 1, characterized in that, Step 3: Construct a conditional diffusion generation model; the conditional information is injected into each stage of the denoising network through the adaptive layer normalization module. The conditional embedding and the time step embedding are input into the network after being concatenated by the dimension. The mean square error between the predicted noise and the real noise output by the decoder is used as the basic loss function of the diffusion model.

6. The method for generating Chu culture patterns by integrating knowledge graphs according to claim 1, characterized in that, Step 4: Generating conditional semantic enhancement; The construction of semantic enhancement vectors is based on knowledge graph embedding technology, which maps entity identifiers to a low-dimensional dense vector space, maps relation types to a learnable relation matrix, and splices and fuses semantic enhancement vectors with noisy image features in the channel dimension.

7. The method for generating Chu culture patterns by integrating knowledge graphs according to claim 1, characterized in that, Step 5: Iterative optimization and pattern output; the adversarial training adopts a relativistic discriminator structure. The discriminator judges the probability of the generated pattern's realism relative to the real pattern. The generator loss includes three components: diffusion reconstruction loss, semantic consistency loss, and adversarial loss. The semantic consistency loss is determined by calculating the cosine similarity between the generated pattern features and the target semantic embedding. The total generator loss is the weighted sum of the three components.

8. The method for generating Chu culture patterns by integrating knowledge graphs according to claim 1, characterized in that, It also includes a pattern post-processing step; the pattern post-processing step performs super-resolution reconstruction on the output pattern image of step 5: iterative optimization and pattern output, improves the image resolution to the preset resolution, and at the same time uses a color correction algorithm to adjust the visual presentation of the pattern to make it more in line with the typical color system of Chu culture. The color correction adopts an end-to-end color mapping network based on convolutional neural network.

9. The method for generating Chu culture patterns by integrating knowledge graphs according to claim 1, characterized in that, It also includes a quality assessment step; the quality assessment step evaluates the pattern image output by the pattern post-processing step, and evaluates the quality of the generated pattern from two dimensions: cultural accuracy and artistic aesthetics. The cultural accuracy assessment indicators include pattern type recognition accuracy and semantic matching degree. The artistic aesthetics assessment adopts a pre-trained pattern aesthetics evaluation network. The comprehensive quality score of the generated pattern is obtained by weighted fusion of the assessment results of the two dimensions.

10. The method for generating Chu culture patterns by integrating knowledge graphs according to claim 1, characterized in that, It also includes a user interaction step; the user interaction step allows users to specify the cultural theme, application scenario and style preference of the generated pattern through natural language description. The natural language input is converted into a semantic vector through a pre-trained language encoder. The semantic vector, as additional conditional information, is input into the diffusion model together with the multimodal pattern feature representation output from step 2, guiding the diffusion model to generate a pattern image that conforms to the user's intention.

Citation Information

Patent Citations

  • Printed fabric (digital prints of bronze ware patterns from Chu culture)

    CN303163920S

  • Book cover (A Study on the Genealogy of Decorative Patterns in Chu Culture)

    CN306047494S