Intuitively Driven Cross-Modal Generative Design Method, Equipment, and Media for Ceramic Teapots
Patent Information
- Application Number
- CN202610668481.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-05-15
AI Technical Summary
旨在解决现有技术中陶瓷设计整体协调性不足、情感偏好难以精确映射以及二维设计稿无法直观展示三维立体结构等问题,为非遗工艺的智能创新设计与文化传承提供高效的技术支撑
1、实现了跨模态的高效自动化设计: 本发明打通了从“用户非结构化情感需求解析”、“二维多模态要素协同生成”到“三维立体结构重建”的全链路自动化设计路径。相较于传统纯人工建模或单一模态生成,极大地缩短了产品迭代周期,提高了设计的整体效率。
Smart Images

Figure CN122197651B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and industrial design, and in particular to a sensory-driven cross-modal generative design method, device and medium for ceramic teapots. Background Technology
[0002] Ceramic teapots hold a unique position in the Chinese arts and crafts system, their shapes and decorations embodying cultural changes and design innovations. However, current ceramic teapot design practices still heavily rely on the implicit experience of individual artisans, making it difficult to accurately and quantitatively translate abstract user emotional preferences (Kansei) into tangible design elements. This design model lacks a synergistic optimization mechanism between morphological parameters and surface patterns, failing to meet the modern consumer market's urgent demands for efficient product iteration and personalized emotional experiences.
[0003] Although Kansei Engineering has been introduced into the field of ceramic design, existing technologies still have two significant limitations: First, most traditional Kansei generation algorithms can only output single-dimensional shape outlines and lack the ability to synchronously and intelligently extrapolate glaze colors and patterns, resulting in design integration and visual presentation still requiring a large amount of manual labor. Second, current generative AI-generated design schemes are only presented as two-dimensional images. This planar expression cannot realistically assess the geometric proportions and material texture of three-dimensional products in space, which seriously weakens the practical value of generative technology in traditional physical manufacturing processes.
[0004] To address the technical challenges of existing ceramic design, such as its heavy reliance on human experience, lack of a collaborative generation mechanism for form and glaze patterns, and difficulty in accurately assessing the spatial texture of three-dimensional products based solely on two-dimensional presentation, there is an urgent need for a cross-modal automated design method that can convert emotional semantics into high-fidelity three-dimensional product forms and patterns. Summary of the Invention
[0005] The purpose of this invention is to provide a sensory-driven cross-modal generative design method, device, and medium for ceramic teapots. It aims to address issues in existing ceramic designs such as insufficient overall coordination, difficulty in accurately mapping emotional preferences, and the inability of two-dimensional design drafts to intuitively represent three-dimensional structures, thus providing efficient technical support for the intelligent innovative design and cultural inheritance of intangible cultural heritage crafts.
[0006] In a first aspect, the present invention provides a sensory-driven cross-modal generative design method for ceramic teapots, comprising the following steps: Step S1: Obtain structured comment text data containing user evaluation elements; preprocess the structured comment text data to construct an original corpus; extract and determine a candidate set of emotional vocabulary from the original corpus based on a bidirectional long short-term memory network and a TF-IDF weighting mechanism. Step S2: Extract the component morphological features of historical ceramic teapot products and encode the feature sequence. Obtain the user sentiment preference data corresponding to the candidate emotional vocabulary set. Use the POA-CNN-BiGRU model to train based on the feature sequence encoding and the user sentiment preference data, and output the optimal teapot morphological parameter combination that matches the candidate emotional vocabulary set. Step S3: Call the pre-trained large language model to convert the candidate sensory vocabulary set into a pattern text guide word sequence; input the pattern text guide word sequence into the low-rank fine-tuned image diffusion model to generate candidate ceramic flower pattern images that are semantically consistent with the candidate sensory vocabulary set; Step S4: Call the pre-trained large language model to generate landscape image prompts that reflect the candidate sensory vocabulary set, input the landscape image prompts into the image generation engine to obtain a color reference map; perform clustering calculation on the pixel channels of the color reference map, and extract the main color value matrix to construct a glaze color scheme; Step S5: Based on the cross-attention mechanism, the morphological structural features corresponding to the optimal teapot shape parameter combination, the pattern features of the candidate ceramic floral pattern image, and the color semantic features of the glaze color scheme are jointly conditionally encoded; a redrawing diffusion model with norm constraint adjustment mechanism is used for sampling decoding to output planar ceramic teapot effect image data; a 3D reconstruction network is called to perform 3D mapping processing on the planar ceramic teapot effect image data to output a 3D ceramic teapot mesh model with surface texture features; Step S6: Obtain the subjective evaluation dataset of the evaluation object for the three-dimensional ceramic teapot mesh model, construct a decision matrix based on triangular fuzzy numbers; output the optimal ceramic teapot design scheme with the highest comprehensive gray relational degree by calculating the Euclidean distance and gray relational coefficient of the standardized fuzzy positive ideal sequence.
[0007] In a second aspect, embodiments of this application provide an electronic device, which includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of the method described in the first aspect.
[0008] Thirdly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0009] Compared with the prior art, the present invention has the following beneficial effects: 1. Achieves highly efficient automated design across modalities: This invention establishes a fully automated design path from "analysis of users' unstructured emotional needs," "collaborative generation of two-dimensional multimodal elements," to "three-dimensional structural reconstruction." Compared to traditional purely manual modeling or single-modal generation, it significantly shortens the product iteration cycle and improves overall design efficiency.
[0010] 2. A multi-element precise coupling mechanism for form, pattern, and glaze color was constructed: This breaks through the bottleneck of traditional emotional design, which can only predict a single outline. This invention innovatively introduces the POA-CNN-BiGRU deep learning model to map physical space and emotional space, and combines it with a redraw diffusion engine with norm constraint adjustment. Under the premise of ensuring that the structural proportions are not distorted, it achieves precise and natural integration of glaze color and texture features onto the specified ceramic surface.
[0011] 3. Solved the problem of preserving multiple cultural images and presenting them in three dimensions: In the diffusion generation stage, a low-rank fine-tuning matrix specific to the ceramic context was introduced to overcome the distortion problem of general large models in the generation of traditional cultural patterns; at the same time, it directly connects to the Tripo 3D reconstruction module to realize the automatic extension from planar vision to three-dimensional expression, providing a truly implementable design blueprint reference for traditional intangible cultural heritage crafts.
[0012] 4. Provides a highly objective and robust solution decision-making evaluation system: In the final solution selection stage, fuzzy grey relational analysis is adopted to convert subjective Likert scale scores into triangular fuzzy number operations, which effectively absorbs the perceptual fluctuations and subjective ambiguities of users in the evaluation process, and significantly improves the scientificity and stability of the optimal solution decision results. Attached Figure Description
[0013] Figure 1 This is a flowchart of the sensory-driven cross-modal generative design method for ceramic teapots according to the present invention. Figure 2 This invention describes the POA-CNN-BiGRU model architecture and prediction process. Figure 3 This is a typical architectural diagram of the diffusion model in this invention; Figure 4 This is a diagram of the Qwen-Image architecture in this invention; Figure 5 These are 200 common teapot examples from this invention; Figure 6 This is a table analyzing the morphology of the ceramic teapot in this invention; Figure 7 This invention presents the convergence curve of the POA and a comparison between the prediction and the actual user rating. Figure 8This is a comparison of the prediction results of the POA-CNN-BiGRU model in this invention across four sensory dimensions; Figure 9 This is the optimal combination of the five components of the teapot under the four sensory dimensions in this invention; Figure 10 This is a portion of the LoRA ceramic pattern training dataset used in this invention; Figure 11 These are the 48 ceramic floral patterns generated in this invention; Figure 12 This is the ceramic design line drawing of the morphological fusion pattern in this invention; Figure 13 These are the 40 extracted glaze color matching schemes in this invention; Figure 14 These are the effects of 40 sets of glaze-patterned teapots in this invention; Figure 15 These are 3D renderings of some of the ceramic teapots in this invention; Figure 16 This is the optimal ceramic teapot design scheme under the four emotional dimensions in this invention. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0016] Example This invention provides an intuition-driven cross-modal intelligent generative design method for ceramic teapots, with a detailed explanation of the entire process using the design of a ceramic teapot as an example.
[0017] Please see Figure 1 This is a flowchart illustrating a sensory-driven cross-modal generative design method for ceramic teapots, provided by an embodiment of the present invention. The method includes the following steps: Step S1: Obtain structured comment text data containing user evaluation elements, preprocess the structured comment text data to construct an original corpus; extract and determine candidate emotional vocabulary sets from the original corpus based on bidirectional long short-term memory network and TF-IDF weighting mechanism.
[0018] Specifically, step S1 involves extracting emotional vocabulary from user online comments and predicting the teapot shape parameter combination that best matches the target emotion based on an improved BiGRU model. (1) User Review Collection. This invention uses JD.com, a major domestic e-commerce platform, as the data source, focusing on keywords such as "ceramic teapot," "purple clay teapot," and "handmade teaware." A web crawler using the Python Scrapy framework was developed to collect structured review text data containing user evaluation elements from the past year. To ensure data quality, only samples with ≥20 characters were retained, resulting in 12,843 valid reviews. After preprocessing to remove advertising language and non-Chinese characters, the original corpus was formed.
[0019] (2) BERT-BiLSTM-TFIDF Extraction of Sentimental Vocabulary. To accurately identify sentimental vocabulary from unstructured comments, this invention first uses a pre-trained Chinese BERT model to perform context-aware vector encoding on each word, capturing BERT word vectors that have emotional tendencies in specific contexts; secondly, the BERT word vector sequence is input into a Bidirectional Long Short-Term Memory (BiLSTM) network to model the contextual dependencies between words, outputting a context-enhanced representation of each word: The model consists of a forward Long Short-Term Memory (LSTM) network and a backward Long Short-Term Memory network, forming a bidirectional structure. The hidden state of the forward LSTM at time t encodes the positive semantic information of the text. The backward LSTM encodes the reverse semantic information of the text in the hidden state at time t. The hidden state after bidirectional concatenation at time t is integrated with the complete contextual semantics.
[0020] The initial probability of each sentiment word belonging to the "emotional word" is calculated using a fully connected layer and a softmax function. A TF-IDF weighting mechanism is introduced to suppress interference from high-frequency meaningless adjectives (such as "good" and "very good"), and the initial probabilities are corrected. in, This represents the probability of the emotional word after TF-IDF weighting. This represents the probability of the intuitive word predicted by the model in its original form. The text words corresponding to time t.
[0021] Ultimately, selected The top 50 adjectives were used as a candidate set of emotional words, and were manually verified in conjunction with the opinions of ceramic design experts to eliminate ambiguous or irrelevant words (such as "fast delivery"), and finally four core emotional words were determined: "durable", "comfortable", "minimalist" and "exquisite", which formed the semantic input basis for the subsequent Kansei mapping model.
[0022] Step S2: Extract the component morphological features of historical ceramic teapot products and encode the feature sequence. Obtain the user sentiment preference data corresponding to the candidate emotional vocabulary set. Use the POA-CNN-BiGRU model to train based on the feature sequence encoding and the user sentiment preference data, and output the optimal teapot morphological parameter combination that matches the candidate emotional vocabulary set.
[0023] To establish the mapping relationship between component morphology and user sentiment preferences, this invention constructs a POA-CNN-BiGRU model (Figure 2). This model uses POA to globally optimize key hyperparameters of the CNN-BiGRU network, employing the root mean square error (RMSE) as the fitness function to achieve multivariate sentiment regression prediction.
[0024] First, let the input discrete sequence after encoding the five major component shapes be: in For the embedded vector dimension, This represents the morphological feature encoding vector of the i-th component. This represents the input feature sequence after encoding the shapes of the five main components of a teapot.
[0025] Discrete component labels are mapped to continuous dense vectors through an embedding layer: in For learnable embedding matrices; It is a continuous dense vector of component features after being mapped by the embedding layer.
[0026] This yields user sentiment preference data corresponding to the candidate set of emotional words; This is the component encoding vector.
[0027] Secondly, the processing of the POA-CNN-BiGRU model includes: (1) CNN local feature extraction To capture the local coupling relationships between adjacent components, a one-dimensional convolution operation is introduced: in This represents the convolution operation. Let σ be the size of the one-dimensional convolution kernel. It is a non-linear activation function specific to convolutional layers, used only for non-linear enhancement of local features; , For convolution parameters, This represents the local feature output at position i. The resulting local feature sequence is: in It is a local feature sequence; Indicates the total length of the feature sequence (consistent with the length of the component sequence).
[0028] (2) BiGRU long-range dependency modeling To model the overall proportional coordination between components, the convolutional output is input into a bidirectional gated recurrent unit (GRU) network. The GRU unit calculation process is as follows.
[0029] Update Gate: in, This is a GRU update gate with a value range of (0,1), controlling the hidden state of the previous time step. How much information is retained up to the current moment; of which The sigmoid activation function outputs values in the range of 0 to 1. To update the gate's weight matrix (learnable parameters), used for the current input. Perform a linear transformation; The input features at the current time step (i.e., the local feature sequence output by the preceding CNN layers); To update the weight matrix of the gate; It is the hidden state of the previous time step, storing the sequence in t. Historical information for step 1.
[0030] Reset Door: in Reset the door's value range (0,1) to control its hidden state from the previous moment. How much information has been "forgotten / reset"? It is the weight matrix (learnable parameters) of the reset gate, which is applied to the current input. ; The weight matrix (learnable parameter) of the reset gate is applied to the hidden state of the previous time step. .
[0031] Candidate hidden state: in, Represented as a GRU candidate hidden state, storing the new information at the current time, it is a... and after reset The integration; It is the hyperbolic tangent activation function, which maps the input to ( The 1,1) interval enhances the model's nonlinear expressive power; It is the weight matrix (learnable parameters) of the candidate states, which is applied to the current input. ; It is the weight matrix (learnable parameters) of the candidate state, which is applied to the reset historical state; It is the Hadamard product, which is the element-wise multiplication of matrices to achieve gated weighting. Use the reset door Regarding the historical state Weighted: The closer to 0, the more the historical information at the corresponding position is "forgotten / reset"; the closer to 1, the more the information is retained.
[0032] Final hidden state: in, It is the final hidden state at the current moment, which combines historical information and new candidate information; Invert the update gate to control historical information. The retention ratio; Historical information to be preserved: The closer to 0, The closer it is to 1, the more historical information is retained; For the newly added candidate information: The closer it is to 1, the more new information is incorporated.
[0033] The two-way structure is represented as: For the forward GRU cell, perform the forward calculations of formulas (7)-(10); For backward GRU cells, perform the reverse calculations of formulas (7)-(10); Input feature sequence.
[0034] The final output is in concatenated form: in, Let be the hidden state of the forward GRU at time t; This represents the hidden state of the backward GRU at time t. This represents the final hidden state after bidirectional GRU splicing.
[0035] (3) Fully connected regression output The final hidden state is fed into the fully connected layer to obtain a four-dimensional sentiment prediction vector: in This is a four-dimensional sentiment prediction vector; These are the output layer weights; For output layer bias; It is the GRU hidden state at the last moment.
[0036] in: This represents the predicted values for the four Kansei sentiment dimensions (durable / comfortable / minimalist / refined).
[0037] (4) Loss function and fitness function The model training uses mean squared error (MSE) as the loss function: in Indicates the number of samples. This indicates the actual rating; To predict scores. To evaluate model performance, root mean square error (MSE) is introduced: in, Root mean square error is the model accuracy index.
[0038] Using RMSE as the fitness function in the POA optimization process: in The set of hyperparameters to be optimized; Optimize the objectives for POA; represents the learning rate, the number of hidden nodes in the GRU layer, the number of hidden nodes in the BiGRU layer, and the L2 regularization coefficient, respectively.
[0039] (5) POA hyperparameter optimization process Let the population size be The maximum number of iterations is Each individual represents a set of hyperparameter vectors: In the l-th iteration, POA updates the individual through a global search and location update mechanism: in For the first The second iteration Individual; This is the current globally optimal solution. A random coefficient (also called step size coefficient / exploration coefficient) with values between 0 and 1 is used to balance the algorithm's global exploration capability and local exploration capability. Through iterative optimization, the optimal combination of teapot shape parameters matching the candidate set of sensory vocabulary is obtained. in, The optimal hyperparameters obtained by optimizing POA; finally, using Retrain the CNN-BiGRU network to construct a complete POA-CNN-BiGRU perceptual mapping model.
[0040] Step S3: Call the pre-trained large language model to convert the candidate sensory vocabulary set into a pattern text guide word sequence; input the pattern text guide word sequence into the low-rank fine-tuned image diffusion model to generate candidate ceramic flower pattern images that are semantically consistent with the candidate sensory vocabulary set.
[0041] Specifically, the implementation of using a pre-trained Large Language Model (LLM) to transform the candidate set of sensory words into a sequence of pattern text guide words, and then driving the fine-tuned Qwen-Image model to generate ceramic floral patterns consistent with cultural context, is as follows: (1) Establish a dataset of ceramic floral patterns: To support subsequent pattern generation and fine-tuning based on Qwen-Image, this invention first constructs a high-quality image dataset consisting of standard ceramic floral pattern line drawings. Original images were collected through sources such as ceramic pattern research books and underwent standardized preprocessing to improve data consistency and model generalization ability.
[0042] (2) The Qwen-Image diffusion model driven by sensory semantics generates candidate ceramic flower pattern images that are semantically consistent with the candidate sensory vocabulary set.
[0043] DM performs well in text-to-image generation. It gradually adds Gaussian noise through iterative calculations, making the image gradually approach pure noise as t increases.
[0044] in, This represents the joint probability distribution of the forward noise addition process in the diffusion model. Represents the original, clear image data; Indicates the number of time steps Data; This represents all noisy data sequences from t=1 to t=T; The noise figure represents the number of time steps t and controls the magnitude of the noise. Represents the identity matrix. This represents a Gaussian distribution. Non-iterative calculations can be performed using formula (14): in, This represents the noise figure minus the t-th step, and the signal preservation coefficient. The cumulative signal retention coefficients are the product of the previous t steps of α. This represents standard Gaussian noise.
[0045] DM employs a U-net network structure, defining the reverse process as a Markov chain composed of a series of Gaussian distributions parameterized by a neural network. It is defined as the mean parameter of the model prediction. As shown in formulas (15) and (16): in The joint probability distribution representing the backward denoising process of the diffusion model. This indicates that the neural network predicts from reduction The conditional probability; express, This represents the learnable parameters of the model; This represents the denoised mean of the model's predictions; The denoised variance of the model prediction.
[0046] Traditional DM relies on CLIP encoders and U-Net, which have limited ability to understand complex semantics. While Qwen-Image employs a different technical approach, it struggles to accurately reflect cultural styles and pattern norms when directly generating traditional Chinese ceramic floral patterns. Therefore, LoRA is introduced for efficient fine-tuning. The core idea of LoRA is to optimize the original weight matrix... Superimpose a low-rank decomposition The incremental term, its core formula is: in, , Let be two low-rank matrices, and For low-rank numbers, only update during training. and Freeze the original weights By iteratively updating the low-rank matrix in the incremental term, a low-rank fine-tuned weight network under a specific cultural context is obtained. The low-rank fine-tuned weight network is injected into the inverse denoising Markov chain of the image diffusion model, and the pattern text guide word sequence is used as a conditional guidance constraint. After multiple steps of iterative denoising operation, an image diffusion model with low-rank fine-tuning is obtained, and candidate ceramic flower pattern images are output, thereby achieving domain adaptation with extremely low computational overhead.
[0047] Step S4: Call the pre-trained large language model to generate landscape image prompts that reflect the candidate emotional vocabulary set, input the landscape image prompts into the image generation engine to obtain a color reference map; perform clustering calculation on the pixel channels of the color reference map, and extract the main color value matrix to construct a glaze color scheme.
[0048] Specifically, the process involves using LLM to generate emotional landscape image prompts that embody the emotional semantics, generating a color reference image through an image generation engine (such as Midjourney), extracting the main color value matrix using K-means clustering, and finally integrating the shape outline, floral pattern, and glaze color scheme in the ComfyUI workflow. (1) Landscape images generated based on Midjourney: This invention aims to transform abstract, emotional semantics into concrete visual expression by combining ChatGPT-4o and Midjourney to construct an emotional reference image generation process. Based on emotional words extracted from user comments, ChatGPT-4o first generates 10 standardized prompts with distinct styles for each word. These prompts are then input into Midjourney to generate candidate images, resulting in a total of 160 color reference images. After two rounds of manual selection by the design team based on the emotional expressiveness of color and compositional aesthetics, 10 representative images are selected for each emotional word as input materials for subsequent K-means clustering color extraction.
[0049] (2) K-means clustering for glaze color scheme extraction: Flatten the color values of all pixels in the landscape image into an N×3 matrix (N is the total number of pixels, and 3 corresponds to the three channels of the Lab color space), and use this matrix as input for K-means clustering. The goal of K-means is to divide n observation points into... There are several clusters that minimize the Within-Cluster Sum of Squares (WCSS): in, The K-means clustering loss function (in-cluster squared error and WCSS); The number of clusters (5 in this paper, extracting 5 primary colors); Indicates the first A color cluster; The pixel color vector belonging to the cluster; Let μ1 be the cluster center (dominant color value) of the i-th cluster. After the algorithm converges, it obtains 5 cluster centers {μ1, μ2, ..., μ5}, which represent the 5 most prominent colors in the image. This method preserves the color atmosphere of the original emotional reference image and provides accurate color input for generative design.
[0050] Step S5: Based on the cross-attention mechanism, the morphological structural features corresponding to the optimal teapot morphological parameter combination, the pattern features of the candidate ceramic floral pattern image, and the color semantic features of the glaze color scheme are jointly conditionally encoded; a redrawing diffusion model with norm constraint adjustment mechanism is used for sampling decoding to output planar ceramic teapot effect image data; a 3D reconstruction network is called to perform 3D mapping processing on the planar ceramic teapot effect image data to output a 3D ceramic teapot mesh model with surface texture features.
[0051] (1) In the multimodal condition modeling stage, ceramic line drawing images Reference image for glaze color Each is mapped to the latent feature space via a visual encoder: in Represents the visual feature encoding function. This represents the morphological and structural features corresponding to the optimal combination of teapot morphological parameters. Indicates the pattern features of the candidate ceramic floral pattern image; text prompt. Converted into semantic embedding vectors by a language encoder It is used to carry additional global prior constraints such as design style, aesthetic emotion, and craftsmanship.
[0052] (2) Subsequently, the morphological and structural features were integrated. Pattern characteristics Color visual features and global semantics of text Multimodal information is used to achieve cross-modal feature alignment and deep fusion through a cross-attention mechanism, generating a unified joint conditional latent representation: in Multimodal joint conditional features; This is a multimodal feature fusion function based on a cross-attention mechanism, ultimately yielding joint conditional features. This serves as the guiding input condition for subsequent redrawing of the diffusion model. To enhance conditional consistency and avoid structural scaling shifts, this invention introduces a classifier-free guidance (CFG) adjustment mechanism during the sampling process. First, the basic guidance vector between conditional and unconditional predictions is defined: in For conditional prediction results, This is an unconditional prediction result. The standard CFG form is: in This represents the final noise prediction after CFG weighting; s is the CFG guiding intensity coefficient. To avoid... When the value is too large, it leads to structural distortion. This invention uses a norm constraint strategy to rescale the basic guidance vector: in Represents the L2 norm; Minimal constants used for numerical applications; stable This is for noise prediction after norm constraints. After diffusion sampling is completed, the final latent variables are obtained. The image shows a reconstructed planar ceramic teapot using a VAE decoder. in Indicates the decoding function; This is the final output of a 2D ceramic teapot rendering; The final latent variable after denoising is completed. The generated 2D ceramic teapot rendering is then attached to the surface of the 3D model. Using a 3D reconstruction network (such as Tripo SR), the visual mapping of planar patterns to 3D form is achieved, outputting a 3D ceramic teapot mesh model with surface texture features. Through the above process, a collaborative generation mechanism that unifies morphological constraints, pattern semantic expression, and glaze color coordination is constructed, providing a controllable and interactive technical path for intelligent design of ceramic products.
[0053] Step S6: Obtain the subjective evaluation dataset of the evaluation object for the three-dimensional ceramic teapot mesh model, construct a decision matrix based on triangular fuzzy numbers; output the optimal ceramic teapot design scheme with the highest comprehensive gray relational degree by calculating the Euclidean distance and gray relational coefficient of the standardized fuzzy positive ideal sequence.
[0054] Specifically, the method for ranking user multidimensional questionnaire evaluation results to select design solutions that resonate with user emotions is as follows: To select the ceramic teapot design with the best glaze performance under four emotional themes, this invention uses fuzzy GRA (Fuzzy Number Decision Matrix) for comprehensive ranking based on subjective evaluation data of user visual perception. By converting Likert scale scores into triangular fuzzy number modeling, the stability of the evaluation results is improved.
[0055] (1) Constructing the fuzzy decision matrix First, calculate the 3D ceramic teapot mesh model at the 1st... The first scheme Under each indicator Arithmetic mean of ratings from evaluators : Where M represents the total number of people giving evaluations. Indicates the first The evaluator's opinion on the first The first option, the first The original scores of each indicator ( =1,2,…,M); For the first The first option, the first The average score of each indicator is used to eliminate individual evaluation bias and construct an objective evaluation benchmark.
[0056] Secondly, based on the mean Constructing triangular fuzzy numbers The left and right boundaries are each extended by half a level (0.5) from the mean to construct the lower bound of the triangular fuzzy number. Median and the Upper Realm To simulate the range of perceived fluctuations in the evaluation.
[0057] Finally, the initial fuzzy decision matrix is constructed: (2) Fuzzy standardization processing In this invention, all emotion sub-indicators are benefit-type indicators (the larger the value, the stronger the emotion expression). To eliminate the influence of unit dimensions and map the data to... The interval is standardized using the fuzzy linear scaling transformation method.
[0058] Let the first Individual indicators The maximum upper bound is Then the standardized triangular fuzzy number The calculation is as follows: (3) Construct a fuzzy ideal reference sequence and calculate the distance, fuzzy positive ideal reference sequence The maximum value of each indicator in the standardized matrix represents the perfect performance under this sentiment theme: Calculate the standardized fuzzy numbers for each scheme With ideal sequence Euclidean distance between To quantify the degree of deviation: in, express plan The Euclidean distance between the index and the ideal solution. At this point, Convert to clear real numbers. Based on the distance values of all schemes, determine the minimum difference between the two levels. and the maximum difference between the two levels .
[0059] (4) Calculate the fuzzy grey relational coefficient Using resolution coefficient (This invention takes) Calculate the fuzzy grey relational coefficients for each scheme under each indicator. : Finally, assuming that the four sub-indicators under the same sentiment theme have equal weights ( ), calculate the comprehensive grey relational degree of each scheme. : in, Indicates the first The solution incorporates the grey relational analysis. Indicator weights (indicators 4 in this paper are equally weighted, w=0.25). The higher the value, the closer the glaze scheme is to the ideal emotional expression effect.
[0060] Example verification The method of the present invention will be further explained below using a ceramic teapot as an example.
[0061] Step S10, design of the ceramic teapot's shape; First, user reviews of ceramic teapots from JD.com and Taobao platforms from month X to month Y of a given year were collected, yielding 10,001 raw data entries. After filtering, reviews containing clear descriptions of aesthetics or user experience (such as mentions of "good feel," "simple design," and "exquisite workmanship") were retained, resulting in 8,620 valid text entries. Subsequent basic cleaning was performed: links, emoticons, advertising slogans, and irrelevant content were removed. Jieba word segmentation tool, combined with a ceramics-related dictionary, was used for word segmentation to ensure accurate recognition of technical terms.
[0062] To systematically identify the emotional attributes that users focus on, we employ a BERT-BiLSTM-TFIDF fusion method to analyze the cleaned comments: First, TF-IDF is used to initially screen high-frequency emotional words; then, the Chinese BERT model (RoBERTa-wwm-ext) is used to obtain the contextual semantic representation of the words; and the BiLSTM network is used to model the contextual importance of words in sentences and dynamically adjust their weights. After combining TF-IDF and attention scores for ranking, and in conjunction with the manual review by two ceramic designers, words with strong subjectivity or irrelevant to design are eliminated. Finally, four representative KE words are determined: "durable," "comfortable," "minimalist," and "refined." These four words correspond to product lifespan, human-computer experience, styling, and craftsmanship quality, respectively, serving as the core evaluation dimensions in the KE framework of this invention and guiding the subsequent form combination of design schemes as well as the generation and evaluation of patterns and glazes (Kansei word results are shown in Table 1). Table 1. Results of BERT-BiLSTM-TFIDF for Kansei vocabulary Step S20, Morphology generation based on POA-CNN-BiGRU This invention constructs a CNN-BiGRU multivariate regression model based on POA optimization, and uses it in a case study to identify and verify the teapot shape that best meets the user's emotional needs.
[0063] To construct a mapping relationship between the shape features of ceramic teapots and user emotional preferences, this invention first collected 200 images of common ceramic teapot products as a basic sample. Figure 5 Combining traditional ceramic design theory, each teapot is broken down into five core components: body, spout, handle, knob, and base. Seven typical morphological characteristics are summarized for each component type, forming a morphological analysis diagram. Figure 6Subsequently, morphological features were annotated for each sample image. Simultaneously, to obtain users' subjective emotional evaluation of the teapot's design, this invention invited 100 participants with teaware usage experience to rate 200 teapot samples using a 5-point Likert scale, employing four Kansei words: "durable," "comfortable," "minimalist," and "exquisite." The average results were taken to eliminate the influence of individual differences, ultimately forming the KE evaluation matrix (Table 2).
[0064] Table 2. KE Evaluation Matrix POA-CNN-BiGRU Multivariate Regression Model Construction and Training To achieve accurate mapping of teapot-shaped component features to user emotional evaluation, this invention constructs a POA-CNN-BiGRU multivariate regression model. The model input is the morphological feature encoding sequence of the five main components of the teapot, which is converted into continuous vectors through a learnable embedding layer to enhance the expressive power of discrete morphological labels. Subsequently, a one-dimensional convolutional neural network (1D-CNN) extracts local features from the embedding sequence, capturing the neighborhood coupling relationships between components; BiGRU is used to model the long-range dependencies of the overall shape, thereby understanding the underlying structural semantics such as proportional harmony and visual balance. To highlight the contribution of key components to emotional expression, the model introduces a POA mechanism to dynamically allocate attention weights. The output performs simultaneous regression prediction on four emotional dimensions through a two-layer fully connected network.
[0065] The dataset contains 200 teapot examples, with 160 used for model training and 40 as a test set. Input user sentiment ratings were normalized using Min–Max to unify the units. The mean squared error (MSE) was used as the loss function during model training, and the Adam optimizer was used to update parameters. An Early-Stopping strategy was implemented during training to prevent overfitting. The POA optimization algorithm successfully found the optimal hyperparameter combination within 20 iterations. The validation set RMSE stabilized during iterations, with a minimum value close to 0.0021, demonstrating that the optimization strategy effectively improved model performance. On the test set, the overall performance was evaluated using the "durable" dimension. The predicted values were highly correlated with actual user ratings, with a determination coefficient R² of 0.954 (0.960 on the training set), indicating good generalization ability and no significant overfitting (Figure 7). The model's predictions across all four sentiment dimensions showed high consistency. Figure 8This verifies the effectiveness of POA-CNN-BiGRU in multivariate emotional regression tasks. Based on the trained model, this invention inputs the test set into the model for regression prediction, obtaining the optimal combination of the five components of the teapot under four emotional dimensions, thereby generating a complete ceramic teapot form that meets the user's emotional needs. Figure 9 ) Step S30: Generating glaze patterns based on LLM and DM; To achieve a consistent aesthetic expression between ceramic patterns and glaze colors, this invention constructs a unified prompt generation mechanism based on GPT-4o. Compared to manual prompt writing that relies on designer experience, the large language model possesses stronger semantic understanding and emotional feature abstraction capabilities, enabling it to accurately map ambiguous emotional appeals into executable visual generation instructions.
[0066] Since patterns are common and representative in traditional ceramics, this invention uses typical ceramic patterns (including scroll patterns, cloud patterns, geometric continuous patterns, and plant imagery patterns) as the main generation objects. Four core Key Emotional Indicators (KEs) extracted from preliminary user research are used as semantic constraints for the prompt words, and the semantic reasoning capabilities of GPT-4o are combined for multi-round iterative optimization of the prompt words. The prompt word construction focuses on controlling two types of structural attributes: ① line style (thickness variation, curvature tendency, symmetry and balance); ② spatial organization (density distribution, repetition rhythm, and compositional proportion). Finally, three pattern prompt words are generated for each KE emotion that can stably express its visual imagery, resulting in a total of 12 high-quality prompt words (Table 3).
[0067] In terms of generating glaze color cue words, this invention also uses GPT-4o as the core tool and utilizes landscape color images as visual references. Landscape colors can intuitively reflect different emotional images (e.g., warm-toned sunsets convey a sense of "comfort"), while natural light and shadow can provide color layers. Using LLMs, four Kansei indices are mapped to color semantic constraints, generating 10 cue words based on each emotion, for a total of 40 high-quality cue words (Table 4). This work provides a unified and controllable semantic input for subsequent collaborative generation of patterns and glaze colors based on Qwen-Image.
[0068] Table 3. Keywords for generating ceramic patterns Table 4. Landscape color image cues that match emotional imagery To achieve controllable generation of ceramic pattern line drawings, this invention trains a dedicated LoRA algorithm based on Qwen-Image. The dataset consists of 200 manually selected high-quality ceramic floral pattern line drawings (partial data can be found in...). Figure 10The data sources include traditional ceramic pattern materials and manually redrawn line drawings. All images are uniformly set to 512×512 resolution, and adaptive threshold denoising is used to ensure clear edges of the line drawings, thus meeting the quality requirements of the Qwen-Image model for training data.
[0069] The trained LoRA model was combined with the Qwen-Image model to generate patterns. The parameters were set as follows: 30 sampling steps, CFG (Classifier-Free Guidance) coefficient 4.0, and diffusion sampling using the Euler scheduler. A fixed random seed and a LoRA weight of 1.0 were used to ensure coherent lines, clear structure, and stable and reliable generation results. The 12 pattern generation prompts provided in the previous stage were input one by one, generating four 512×512 black and white pattern line drawings for each prompt, resulting in 48 candidate ceramic patterns. Subsequently, using a focus group interview method, 10 experts with 3 years of ceramic design experience from 20 groups were invited to select the ceramic pattern that best reflected Kansei's emotions from each group of candidate designs (the 48 generated ceramic floral patterns are shown in Figure 11). The ceramic designer integrated the most frequently selected design from each group of patterns with the optimal combination of the five main components of the teapot obtained in the previous stage, producing four ceramic design line drawings that combined form and pattern. Figure 12 ).
[0070] The generation and extraction of the glaze color scheme consisted of two main steps. First, the 40 Kansei-driven glaze color cues generated earlier were input into the Midjourney model for image generation. Each cue generated four 1024×1024 landscape images, resulting in a total of 160 candidate images. Subsequently, the images were manually screened, removing those with low resolution, uneven colors, or inaccurate imagery. Finally, the 40 high-quality landscape images that best matched the original emotional imagery were retained as the basis for subsequent color matching analysis.
[0071] Secondly, to extract the emotional color scheme represented by each landscape image, this invention uses the K-means clustering method in a Python environment for analysis. Forty high-quality landscape images generated and filtered by Midjourney are imported into Python, and the RGB color matrix of each image is extracted. To reduce the impact of noise on the clustering results, pixel colors are normalized. Subsequently, the K-means algorithm is used to cluster the pixels into five color classes, with the center RGB value of each cluster as the primary color. To ensure the stability of the clustering results, each cluster initialization is run 10 times, and the result with the least squares error is selected. Finally, each landscape image yields a five-color scheme (Figure 13), which accurately reflects the emotional imagery in the image and can be directly used for mapping the glaze color of ceramic teapots.
[0072] To achieve unified generation of the shape, pattern, and glaze of ceramic teapots, this invention constructs a cross-modal fusion workflow based on Comfy UI. First, two types of image information are loaded at the Comfy UI input: one is a 1024×1024 resolution line drawing of the teapot shape, used to provide structural proportions and contour constraints; the other is a color scheme extracted by K-means in the previous stage, used to provide glaze color matching and color distribution information. Both types of images enter the TextEncodeQwenImageEditPlus encoding node, where the morphological structural features and color semantic features are encoded into a unified conditional vector for subsequent diffusion generation control.
[0073] Secondly, regarding model configuration, this invention employs the Qwen Image Edit series models, including U-Net, CLIP, and VAE components, which belong to the image editing generation mode, i.e., controlled redrawing under existing image constraints. To enhance the expressive effect of the glaze patterns, two LoRA modules are loaded in the process, with their intensity parameters both set to 1.00, to strengthen the generation capability of decorative details.
[0074] In the diffusion sampling stage, the KSampler node is used for image generation, with the main parameters set as follows: 8 sampling steps, CFG value of 1.0, Euler sampling algorithm, and a simple linear scheduler. The AuraFlow sampling enhancement module is also introduced to improve generation stability. The lower sampling steps improve generation efficiency, while the lower CFG value makes the generation process more dependent on image conditional input, thereby strengthening the fusion consistency between the morphological line drawing and the glaze image.
[0075] The generated result, after being decoded by VAE, outputs a 1024×1024 resolution 2D teapot rendering. Figure 14Experimental results show that, while maintaining the original line drawing structure and proportions, the generated image achieves a reasonable distribution of the glaze color scheme in the areas of the body, spout, and handle of the pot, and forms a decorative focal point in the center of the body, indicating that structural and color information are synergistically expressed during the diffusion process.
[0076] After the 2D image is generated, the Tripo image-to-3D model module is used for 3D reconstruction. Texture preservation and automatic smoothing are enabled to achieve automatic conversion from 2D visual results to 3D mesh models (see some 3D renderings of ceramic teapots). Figure 15 This step extends the two-dimensional fusion effect to three-dimensional structural expression, allowing the design outcome to transition from a planar visual presentation to a perceptible three-dimensional form expression.
[0077] Specifically, step S40, Fuzzy GRA sentiment assessment; To select the optimal ceramic teapot design based on user visual and emotional evaluation under four emotional themes: "durable," "comfortable," "minimalist," and "refined," this section conducts an optimization experiment. In the questionnaire design, four emotional evaluation indicators were established for each emotional theme. These indicators were broken down based on the core connotations of each emotional theme to ensure the comprehensiveness and scientific rigor of the evaluation system. A five-point Likert scale was used for the questionnaire survey, inviting 30 designers with over five years of teapot design experience to independently score the 40 glaze pattern teapot renderings produced in the previous stage. The evaluation data was quantified and ranked using the Fuzzy GRA method to ultimately determine the optimal glaze scheme for each emotional theme. The core objective of this experiment is to select the design scheme that best fits each of the 10 glaze pattern ceramic teapot designs corresponding to the four emotional themes.
[0078] First, the questionnaire data was organized, and the mean scores of each group's solutions under the four dimensions and corresponding indicators were calculated to form a raw mean table. This table was then converted into triangular fuzzy numbers (based on a fluctuation of 0.5 around the mean) to construct a triangular fuzzy decision matrix to preserve the evaluation fuzziness. Subsequently, the matrix was standardized without dimensions to obtain a standardized triangular fuzzy matrix, laying the foundation for subsequent analysis.
[0079] Table 5. Standardized triangular fuzzy decision matrix The third step is to determine the ideal solution for each dimension, and use the average Euclidean distance at the vertices to calculate the distance between each group of schemes and the ideal solution, forming a distance matrix to intuitively reflect the differences in the performance of the schemes.
[0080] Table 6. Distance Matrix The fourth step is to calculate the grey relational coefficient based on the distance matrix, take the mean value to obtain the grey relational degree of each group of schemes (the larger the value, the better the performance), and sort them according to the relational degree to obtain the final optimal result.
[0081] Table 7. Grey Relational Coefficient Table 8. Relevance Ranking The final results showed that designs A9, B6, C3, and D6 were selected as the ceramic teapot designs that best met user needs across the four emotional dimensions.
[0082] Optionally, embodiments of this application also provide an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the various processes of the above-described embodiment of the inductively driven cross-modal generative design method for ceramic teapots and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0083] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described embodiment of the inductively driven cross-modal generative design method for ceramic teapots and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0084] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0085] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0086] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0087] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A sensory-driven cross-modal generative design method for ceramic teapots, characterized in that, Includes the following steps: Step S1: Obtain structured comment text data containing user evaluation elements; preprocess the structured comment text data to construct an original corpus; extract and determine a candidate set of emotional vocabulary from the original corpus based on a bidirectional long short-term memory network and a TF-IDF weighting mechanism. Step S2: Extract the component morphological features of historical ceramic teapot products and encode the feature sequence. Obtain the user sentiment preference data corresponding to the candidate emotional vocabulary set. Use the POA-CNN-BiGRU model to train based on the feature sequence encoding and the user sentiment preference data, and output the optimal teapot morphological parameter combination that matches the candidate emotional vocabulary set. The processing steps of the POA-CNN-BiGRU model include: obtaining the input discrete sequence after encoding the shapes of the five major components of the teapot; mapping the input discrete sequence into a continuous dense vector sequence through a learnable embedding matrix; extracting local features from the continuous dense vector sequence using a one-dimensional convolutional kernel to capture the local coupling relationship between adjacent components and outputting a local feature sequence; inputting the local feature sequence into the BiGRU network, calculating candidate hidden states through update and reset gates, and outputting the final hidden state in a concatenated form to model the overall proportional coordination relationship between components; inputting the final hidden state into a fully connected layer to output a multi-dimensional sentiment prediction vector; using the root mean square error as the fitness function, dynamically optimizing the learning rate, number of hidden layer nodes, and regularization coefficient of the CNN-BiGRU model in the global search and position update mechanism using the POA algorithm; updating the model weights using the optimal hyperparameter combination; and outputting the optimal teapot shape parameter combination after inputting target test data. Step S3: Call the pre-trained large language model to convert the candidate sensory vocabulary set into a pattern text guide word sequence; input the pattern text guide word sequence into the low-rank fine-tuned image diffusion model to generate candidate ceramic flower pattern images that are semantically consistent with the candidate sensory vocabulary set; Step S4: Call the pre-trained large language model to generate landscape image prompts that reflect the candidate sensory vocabulary set, input the landscape image prompts into the image generation engine to obtain a color reference map; perform clustering calculation on the pixel channels of the color reference map, and extract the main color value matrix to construct a glaze color scheme; Step S5: Based on the cross-attention mechanism, the morphological structural features corresponding to the optimal teapot shape parameter combination, the pattern features of the candidate ceramic floral pattern image, and the color semantic features of the glaze color scheme are jointly conditionally encoded; a redrawing diffusion model with norm constraint adjustment mechanism is used for sampling decoding to output planar ceramic teapot effect image data; a 3D reconstruction network is called to perform 3D mapping processing on the planar ceramic teapot effect image data to output a 3D ceramic teapot mesh model with surface texture features; Step S6: Obtain the subjective evaluation dataset of the evaluation object for the three-dimensional ceramic teapot mesh model, construct a decision matrix based on triangular fuzzy numbers; output the optimal ceramic teapot design scheme with the highest comprehensive gray relational degree by calculating the Euclidean distance and gray relational coefficient of the standardized fuzzy positive ideal sequence.
2. The method according to claim 1, characterized in that, In step S1, the specific steps for extracting and determining the candidate set of sensory vocabulary from the original corpus based on the bidirectional long short-term memory network and TF-IDF weighting mechanism include: The original corpus is input into a pre-trained BERT model for context vector encoding to obtain an initial word vector sequence with sentiment bias. The initial word vector sequence is input into a bidirectional long short-term memory network to model the dependencies between words. By concatenating the hidden states of the forward network and the hidden states of the backward network, a context-enhanced representation vector is output. The initial probability value of the context-enhanced representation vector belonging to emotional words is calculated by using a fully connected layer and a softmax function. A TF-IDF weighting mechanism is introduced to correct the initial probability values, and the final probability of the emotional words is calculated. The emotional words are then sorted in descending order according to their probabilities, and the head data is extracted to obtain the candidate emotional word set.
3. The method according to claim 1, characterized in that, In step S3, the specific steps of inputting the pattern text guide word sequence into the low-rank fine-tuned image diffusion model to generate candidate ceramic flower pattern images that are semantically consistent with the candidate sensory vocabulary set include: Construct a high-quality image training dataset consisting of standard ceramic floral pattern line drawings; The incremental term matrix of low-rank decomposition is superimposed on the original weight matrix of the image diffusion model; During the fine-tuning training phase, the original weight matrix is frozen, and the low-rank matrix in the incremental term is updated iteratively to obtain a low-rank fine-tuned weight network for a specific cultural context. The low-rank fine-tuned weight network is injected into the inverse denoising Markov chain of the image diffusion model, and the pattern text guide word sequence is used as a conditional guiding constraint. After multiple iterative denoising operations, the candidate ceramic flower pattern image is output.
4. The method according to claim 1, characterized in that, In step S4, the specific steps for performing clustering calculations on the pixel channels of the color reference image and extracting the main color value matrix to construct the glaze color scheme include: Extract all pixel values from the candidate landscape image and flatten its color space to construct a two-dimensional feature input matrix containing the total number of pixels and the number of color channels; The feature input matrix is input into an adaptive clustering algorithm, with the objective function being to minimize the intra-cluster squared error. After multiple rounds of iterative clustering, the observation points are divided into a preset number of color clusters. After the extraction algorithm converges, the center color channel value of each color cluster is used as the main color, and the main color value matrix corresponding to the glaze color scheme is constructed by splicing them together.
5. The method according to claim 1, characterized in that, In step S5, the specific steps for jointly conditionally encoding the morphological structural features corresponding to the optimal teapot morphological parameter combination, the pattern features of the candidate ceramic floral pattern image, and the color semantic features of the glaze color scheme based on the cross-attention mechanism include: The optimal teapot shape parameter combination is extracted and converted into a ceramic line drawing image matrix. This matrix, along with the glaze reference image matrix, is then input into a visual encoder for latent feature space mapping. This process yields the morphological and structural features corresponding to the optimal teapot shape parameter combination, as well as the pattern features of the candidate ceramic floral pattern images. The text cues describing the candidate set of sensory words are converted into a sequence of semantic embedding vectors by a language encoder; The attention weights of the morphological features, the pattern features, and the semantic embedding vector sequence are calculated through a cross-attention mechanism network, and a joint conditional representation matrix of multimodal information fusion is output.
6. The method according to claim 5, characterized in that, In step S5, the specific steps for sampling and decoding using the redrawing diffusion model with norm constraint adjustment mechanism to output the planar ceramic teapot effect image data include: Obtain the conditional and unconditional prediction noise results of the redraw diffusion model during the sampling process, and define the basic guidance vector between conditional and unconditional predictions; To prevent structural distortion caused by large guidance strength coefficients, a norm constraint strategy is introduced to rescale the basic guidance vector. The norm value of the conditional prediction noise result is calculated and normalized to adjust it, and the final latent variable sequence after rescaling correction is output. The final latent variable sequence is input into the decoder of the variational autoencoder, and the pixel array is reconstructed through feature deconvolution operation to output the planar ceramic teapot effect image data.
7. The method according to claim 1, characterized in that, The specific steps of step S6 include: Calculate the arithmetic mean of the scores for each 3D ceramic teapot mesh model under multiple evaluation metrics; Based on the arithmetic mean of the score, a preset level difference is extended to the left and right boundaries respectively to construct the lower bound, median and upper bound of the triangular fuzzy number, which are combined to form the initial fuzzy decision matrix; The fuzzy linear scaling transformation method is used to eliminate the dimensional influence of the initial fuzzy decision matrix, and the data is mapped to a standard interval to obtain a standardized fuzzy decision matrix. The maximum upper bound of each index in the standardized fuzzy decision matrix is selected to construct a fuzzy positive ideal reference sequence; the Euclidean distance between the standardized fuzzy number of each evaluation scheme and the fuzzy positive ideal reference sequence is calculated; The fuzzy grey relational coefficients of each evaluation scheme under each indicator are calculated based on the resolution coefficient. By summing the fuzzy grey relational coefficients of all evaluation indicators with equal weights, the comprehensive grey relational degree value of each scheme is calculated, and the three-dimensional ceramic teapot mesh model with the highest value is output as the optimal ceramic teapot design scheme.
8. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein when the program or instructions are executed by the processor, they implement the steps of a sensory-driven cross-modal generative design method for ceramic teapots as described in any one of claims 1-7.
9. A readable storage medium, characterized in that, The program or instructions are stored on the readable storage medium, and when the program or instructions are executed by the processor, they implement the steps of the inductively driven cross-modal generative design method for ceramic teapots as described in any one of claims 1-7.
Citation Information
Patent Citations
VMD-PCA-LSTM ultra-short-term wind power prediction method based on pelicans optimization algorithm
CN119272598A
Ceramic censer design method and system based on perceptual engineering
CN120562241A