An artificial intelligence-based art design assisting method and system
By constructing an art design logic structure diagram and introducing a user feedback mechanism, combined with Householder reflection and Bayesian optimization, the problem of generated results not conforming to complex design semantics and user expectations in existing technologies has been solved, achieving high-quality and personalized art design assistance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 山东电子职业技术学院
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies struggle to generate high-quality artistic design results that conform to complex design semantics within limited iterations, and lack feedback mechanisms consistent with user perception, resulting in a mismatch between the generated images in terms of practicality and aesthetics.
By constructing an art design logic structure diagram for encoding and mapping, latent vectors are generated. Dimensionality reduction and decoding reconstruction are performed by combining Householder reflection and Bayesian optimization. User feedback mechanisms and hybrid algorithms are introduced to optimize the image generation process and ensure that the generated results meet user expectations.
It improves the quality and diversity of art and design assistance systems, enhances user satisfaction and design process flexibility, and can quickly respond to user needs and generate reference assistance images that meet expectations.
Smart Images

Figure CN122113629A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and computer-aided design technology, and in particular to an art design assistance method and system based on artificial intelligence. Background Technology
[0002] In recent years, with the rapid development of artificial intelligence technology, especially the breakthroughs in visual content generation by deep learning and generative models, AI-based art design assistance technology has become a research hotspot at the intersection of computer graphics and creative tools. Early computer-aided design tools mainly relied on parametric modeling and rule bases, which, while improving design efficiency, had limited flexibility and creative expression capabilities. Subsequently, generative AI methods such as generative adversarial networks, variational autoencoders, and diffusion models have been widely applied to image synthesis and style transfer, making it possible to automatically generate visual content from text or sketches. In particular, the introduction of multimodal pre-trained models such as CLIP has achieved unified alignment of image and text semantic spaces, laying the foundation for semantically driven design generation. At the same time, the development of graph neural networks and graph variational autoencoders has made it possible to encode and generate non-Euclidean data structures, further promoting the structuring and interpretability of the design process.
[0003] Existing technologies typically perform optimization and search directly in the original high-dimensional latent space, resulting in high computational complexity and slow convergence. This makes it difficult to obtain high-quality results that conform to the semantics of complex designs within a limited number of iterations. Furthermore, most of them rely on a single objective function (such as reconstruction loss or adversarial loss) for optimization, lacking a fine-grained alignment mechanism for multimodal design intentions (such as image and text fusion). This leads to semantic discrepancies between the generated content and user expectations. Secondly, the lack of a feedback mechanism consistent with user perception often results in a "disharmony" between the generated images in terms of practicality and aesthetics. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides an art design assistance method and system based on artificial intelligence, which solves the problem that existing technologies are unable to obtain high-quality results that conform to complex design semantics within a limited number of iterations and lack a feedback mechanism consistent with user perception, resulting in generated images that are often "disproportionate" in terms of practicality and aesthetics.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides an art design assistance method based on artificial intelligence, comprising,
[0008] The art design logic structure diagram is constructed, encoded, and mapped to generate latent vectors. The target design reference information is preprocessed and semantically extracted by user input to generate target embedding vectors.
[0009] Based on latent vectors, dimensionality reduction is performed through Householder reflection. After constructing a low-dimensional Bayesian optimized vector, zero-padding, inverse mapping, and decoding reconstruction are performed to obtain candidate images and candidate embedding vectors. The target consistency loss value is calculated by combining the target embedding vector and converted into a maximization scoring function. Gaussian process modeling and sampling are then performed to generate the optimal result and construct a pairwise preference probability function. Finally, a ranking likelihood function is constructed for parameter updates and iterations to obtain the final image.
[0010] A preliminary display is performed based on the final image, and the user judges the image satisfaction. If the user is not satisfied, the image parameters are set, and the image parameters are optimized by combining the final image and the art design logic structure diagram to obtain the final image parameters and generate an adjustment reference auxiliary image.
[0011] Based on the adjusted reference auxiliary image, the final visualization is then performed.
[0012] As a preferred embodiment of the AI-based art design assistance method of the present invention, the method involves: performing dimensionality reduction based on latent vectors through Householder reflection, constructing low-dimensional Bayesian optimized vectors, and then performing zero-padding, inverse mapping, and decoding reconstruction to obtain candidate images, including:
[0013] Based on the latent vector, a random orthogonal rotation matrix is generated using Householder reflection. This random orthogonal rotation matrix is then used to perform a rotation transformation on the latent vector, resulting in a rotated latent vector. The previous state is then truncated. The vectors are concatenated along their dimensions to construct a low-dimensional Bayesian optimized vector. Zero-padding is then performed to restore the vectors to the same dimension as the rotated latent vectors. A random orthogonal rotation matrix is used to inversely map the zero-padding low-dimensional Bayesian optimized vector, generating a high-dimensional Bayesian optimized vector. This vector is then input into a pre-trained GraphVAE variational encoder for decoding, resulting in a decoded art design logic structure diagram. Finally, an SVG rule engine is applied to render this decoded art design logic structure diagram as a candidate image.
[0014] As a preferred embodiment of the AI-based art design assistance method of the present invention, the steps of obtaining candidate embedding vectors, calculating the target consistency loss value by combining it with the target embedding vector and converting it into a maximization scoring function, performing Gaussian process modeling and sampling, generating the optimal result and constructing a pairwise preference probability function, and then constructing a ranking likelihood function for parameter updates and iterations to obtain the final image include:
[0015] Based on candidate images, semantic vectors of candidate images are obtained through a pre-trained CLIP network model and then normalized to obtain candidate embedding vectors. Normalization is then performed on these vectors. Combined with the candidate embedding vectors, the target consistency loss value is calculated and converted into a Bayesian optimization maximization scoring function. The scoring function value is then used as the label of the low-dimensional Bayesian optimized vector. Gaussian process modeling is then performed to generate a Gaussian process surrogate model for initialization. Expectation boosting is used to sample within the 3σ effective domain of the low-dimensional Bayesian optimized vectors, resulting in a sequence of low-dimensional Bayesian optimized vectors. The label of each low-dimensional Bayesian optimized vector in the sequence is obtained through the Gaussian process surrogate model, and then sorted in descending order. The low-dimensional Bayesian optimized vector with the largest label is selected as the candidate vector. Zero-padding is then performed to restore it to the same dimension as the rotated latent vector, and this process continues until a new candidate image is rendered, resulting in a new candidate image set containing the new candidate image and its corresponding label. The new candidate image with the largest label is selected from the candidate image set as the optimal result.
[0016] Extract the labels of the optimal result and each new candidate image. Combine the labels of the optimal result with the labels of each other new candidate image to form preference pairs, construct a pairwise preference probability function, calculate the pairwise preference probability value of each preference pair, construct a ranking likelihood function, and calculate the variance of the score function value of the kernel function. Length scale The mean function of the Gaussian process model is updated, and the low-dimensional Bayesian optimization vector with the largest output score function value is zero-padding to restore it to the same dimension as the rotated latent vector, until it is rendered into the final image.
[0017] As a preferred embodiment of the AI-based art design assistance method of the present invention, the following steps are included: performing a preliminary display based on the final image, followed by a user-assessed image satisfaction judgment; if unsatisfied, setting image parameters, optimizing the image parameters by combining the final image and the art design logic structure diagram, obtaining the final image parameters, and generating an adjustment reference auxiliary image, including:
[0018] The final image is displayed to the user via HTML5 Canvas, where the user rates the final image with 0 or 1. If the user is not satisfied, the user sets the required image parameters, including contrast, height, and color. The pixel values and image parameters of the final image are recorded. The image parameters set by the user and the image parameters of the final image are concatenated into a parameter vector, including a graphic parameter vector and a user parameter vector.
[0019] Based on the recorded pixel values, calculate the mean square error of pixel values between the final image and the design logic structure diagram;
[0020] Based on the parameter vector, the feedback loss between the graphics parameter vector and the user parameter vector is calculated, and a weighted fusion formula is constructed by combining the mean square error.
[0021] The user parameter vector is used as an individual in the hybrid algorithm. After initializing the population by randomly generating it, the weighted fusion formula is used as the fitness function of the individual. The fitness function value of each individual is calculated. The differential evolution algorithm is used to randomly select three individuals from the population to perform mutation operations, generating a new user parameter vector. The new user parameter vector is crossed with each dimension of the user parameter vector to obtain the cross user parameter vector. The fitness function value is recalculated, and the user or cross user parameter vector is retained as the optimal individual.
[0022] Based on the optimal individual, the velocity and position of the optimal individual are updated through the echolocation operation of the bat optimization algorithm. After the maximum number of iterations, the final individual is output and the parameter vector is extracted.
[0023] Based on the extracted parameter vector, the user uses OpenCV to perform item-by-item transformation operations on the final image to generate an adjustment reference auxiliary image.
[0024] As a preferred embodiment of the AI-based art design assistance method of the present invention, the step of performing the final visualization assistance display based on adjusting the reference assistance image includes:
[0025] The reference auxiliary image is adjusted and used as an auxiliary image during the user's art design. After it is displayed synchronously with the art design logic structure diagram, a click response area is set in the image area using the HTML5Canvas image display framework. When the user hovers the mouse over the click response area, the node in the click response area is highlighted and the design operation type of the node is displayed.
[0026] As a preferred embodiment of the AI-based art design assistance method of the present invention, the step of constructing an art design logic structure diagram for encoding and mapping to generate latent vectors includes:
[0027] Design diagram samples are obtained from open-source design databases using web crawling technology and parsed to obtain design operations. Each operation in the design operation is taken as a node, and a directed acyclic graph is constructed according to the execution order relationship between the design operations. The directed acyclic graph is then uploaded to the storage as an art design logic structure diagram.
[0028] A pre-trained GIN network and a GraphVAE variational autoencoder are added to the memory side. The art design logic structure graph is input into the pre-trained GIN network, and the output hidden state is used to perform mean aggregation to generate a graph-level embedding representation. Then, it is input into the GraphVAE variational autoencoder to perform spatial mapping and output latent space distribution parameters, including mean vector and standard deviation vector. Latent vectors are generated through reparameterization sampling technology.
[0029] As a preferred embodiment of the AI-based art design assistance method of the present invention, the step of preprocessing and semantic extraction of target design reference information input by the user to generate a target embedding vector includes:
[0030] The user inputs target design reference information into the terminal, including target reference images and target design text descriptions. The target design reference information is preprocessed to obtain input data.
[0031] The input data is used as input to the pre-trained CLIP network model in OpenAI to obtain image and text semantic vectors respectively. Weighted fusion is then performed to generate target embedding vectors.
[0032] Secondly, this invention provides an art design assistance system based on artificial intelligence, comprising,
[0033] The initial execution module is used to construct an art design logic structure diagram for encoding and mapping, generate latent vectors, perform preprocessing and semantic extraction based on user-input target design reference information, generate target embedding vectors for dimensionality reduction, construct low-dimensional Bayesian optimized vectors, and then perform zero-padding, inverse mapping, and decoding reconstruction.
[0034] The update module is used to obtain candidate embedding vectors, calculate the target consistency loss value by combining it with the target embedding vector, convert it into a maximization scoring function, perform Gaussian process modeling and sampling, generate the optimal result and construct the pair preference probability function, and then construct the ranking likelihood function for parameter update and iteration.
[0035] The feedback optimization module is used to perform an initial display and allow the user to judge the image satisfaction. If the user is not satisfied, the image parameters are set and optimized by combining the final image and the art design logic structure diagram.
[0036] The visualization module is used to perform visual aids based on the final image.
[0037] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the artificial intelligence-based art design assistance method as described in the first aspect of the present invention.
[0038] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the artificial intelligence-based art design assistance method as described in the first aspect of the present invention.
[0039] The beneficial effects of this invention are as follows: This invention performs rotation and dimensionality reduction operations on latent vectors through the Householder reflection mechanism, combined with Bayesian optimization, which effectively improves the efficiency and stability of the optimization process. Furthermore, it combines GraphVAE variational autoencoder for decoding and reconstruction, which further optimizes the generation quality and diversity of design graphs. The introduction of user feedback mechanism and hybrid algorithm not only improves the personalization of image generation results and user satisfaction, but also makes the art design assistance process more adaptable and flexible, thereby enabling it to quickly respond to user needs and generate reference assistance images that meet expectations. Attached Figure Description
[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Fig. 1 This is a flowchart of the AI-based art design assistance method in Example 1.
[0042] Fig. 2 This is a structural diagram of the AI-based art design assistance system in Example 1.
[0043] Fig. 3 This is a flowchart of generating candidate images in Example 1. Detailed Implementation
[0044] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0045] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0046] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0047] Example 1, referring to Figs. 1-3 This is the first embodiment of the present invention, which provides an art design assistance method based on artificial intelligence, including the following steps:
[0048] S1. Construct an art design logic structure diagram for encoding and mapping, generate latent vectors, and preprocess and extract semantic information based on user input of target design reference information to generate target embedding vectors;
[0049] Specifically, an art design logic structure diagram is constructed for encoding and mapping to generate latent vectors, including:
[0050] By using web scraping technology, design drawings are obtained from open-source design databases and parsed to obtain design operations (including geometric transformations, symmetry constraints, scaling, color mapping, or texture overlay).
[0051] Each operation in the design process is treated as a node, and a directed acyclic graph (without forming a cycle) is constructed according to the execution order of the design operations. The directed acyclic graph is then uploaded to the memory as the logical structure diagram of the art design.
[0052] A pre-trained GIN network and a GraphVAE variational autoencoder are added to the memory side. The art design logic structure graph is input into the pre-trained GIN network, and the output hidden states are subjected to mean aggregation to generate a graph-level embedding representation. This representation is then input into the GraphVAE variational autoencoder to perform spatial mapping, and the output latent space distribution parameters, including the mean vector and standard deviation vector, are generated. Through reparameterization sampling, latent vectors are generated, as shown in the formula:
[0053]
[0054] In the formula, Represents the latent vector. Represents the mean vector. Represents the standard deviation vector. Represents a random vector generated by sampling from a standard normal distribution. , Represents the standard normal distribution. (Represents the identity matrix).
[0055] It should be noted that the GIN network is used as the backbone of the GraphVAE variational autoencoder, and the GraphVAE variational autoencoder includes an encoder and a decoder. The construction and training of the GraphVAE variational autoencoder are as follows: the variational lower bound (ELBO) is used as the objective function, the Adam optimizer is used to iteratively optimize the model parameters, and the parameters are propagated through the backpropagation algorithm. When the objective function value no longer decreases significantly, the iteration stops, and the trained GraphVAE variational autoencoder is output.
[0056] By automating the acquisition of design graph samples from open-source design databases using web crawling technology, the diversity and timeliness of the samples are ensured, providing rich training data for the system. Secondly, the design operations are transformed into a directed acyclic graph (DAG) structure, giving the design process a clear order and dependencies, ensuring the logic and rationality of the generated designs. Furthermore, through a GIN network encoder and a GraphVAE variational autoencoder, the system can efficiently map the semantic information of the design graph samples to the latent space, generating latent design representations and providing a precise foundation for subsequent image generation and optimization. Reparameterized sampling technology generates latent vectors through a standard normal distribution and combines mean and standard deviation vector adjustments, ensuring the stability and diversity of the latent space generation, thereby avoiding overfitting or mode collapse problems.
[0057] Furthermore, target design reference information input by the user is preprocessed and semantically extracted to generate a target embedding vector, including:
[0058] The user inputs target design reference information into the terminal, including target reference images and target design text descriptions;
[0059] The target design reference information is preprocessed to obtain the input data;
[0060] Preprocessing needs to be performed on the image and text separately. For example, after resizing the image to a uniform size and performing pixel value normalization, color space and tensor conversion are then performed.
[0061] After performing word segmentation and standardization on the text (e.g., lowercase, punctuation removed), the SentencePiece word segmenter is used for encoding, and the excess parts are truncated and then converted into word vectors.
[0062] The input data is used as input to a pre-trained CLIP network model in OpenAI (the pre-trained CLIP network model includes an image encoding network and a text encoding network, and performs semantic space alignment using a shared semantic space alignment mechanism). Image and text semantic vectors are obtained separately, and weighted fusion is performed to generate the target embedding vector, as shown in the formula:
[0063]
[0064] In the formula, Represents the target embedding vector. Modal weights (when the user only provides an image as input) represent modal weights. Setting it to 1 allows for full utilization of image semantic vectors, and modal weights will be adjusted when the user provides only text as input. Setting it to 0 allows for full use of the text semantic vector, while when images and text coexist, the value ranges from 0.4 to 0.7, and the modal weights... When the weight is less than 0.4, the image information weight is insufficient, and the generated image is prone to "instructional composition" caused by text overfitting (i.e., the image degenerates into text interpretation). This is also true for modal weights. When the modality weight is less than 0.7, the text control is weak, and users cannot constrain structural changes through language. Therefore, in this invention, when images and text coexist, a modality weight of 0.6 can be used as the modality weight. (default value) Represents the semantic vector of an image. Represents a text semantic vector. This indicates the element-wise multiplication operation.
[0065] It should be noted that the image encoding network and text encoding network in the pre-trained CLIP network model are constructed using the Vision Transformer and Transformer architectures, respectively. The InfoNCE loss function is used as the loss function, and the AdamW optimizer is used to iteratively optimize the parameters of the CLIP network model. During the iteration process, when the loss value of the loss function no longer decreases significantly, the iteration stops, and the trained CLIP network model is output.
[0066] By using a pre-trained CLIP network model to align the semantic space of images and text, images and text can complement and be understood each other within the same semantic space. This is achieved by setting appropriate modality weights. This invention allows for flexible adjustment of the influence of images and text during the generation process. Furthermore, by resizing, normalizing pixel values, and converting color spaces, it ensures the consistency of image data, providing stable input for subsequent model processing. Text is encoded using a SentencePiece tokenizer, with excess text truncated, ensuring that the input data is standardized and adapted to the needs of the CLIP network. Secondly, by generating target embedding vectors and combining them with reparameterized sampling techniques, the system can sample latent vectors from the latent space and optimize them based on target consistency loss.
[0067] S2. Based on latent vectors, dimensionality reduction is performed through Householder reflection. After constructing low-dimensional Bayesian optimized vectors, zero-padding, inverse mapping, and decoding reconstruction are performed to obtain candidate images and candidate embedding vectors. The target consistency loss value is calculated by combining the target embedding vector and converted into a maximization scoring function. Gaussian process modeling and sampling are then performed to generate the optimal result and construct a pairwise preference probability function. Finally, a ranking likelihood function is constructed for parameter updates and iterations to obtain the final image.
[0068] Specifically, based on latent vectors, dimensionality reduction is performed using Householder reflection. After constructing low-dimensional Bayesian optimized vectors, zero-padding, inverse mapping, and decoding reconstruction are performed to obtain candidate images, including:
[0069] Based on latent vectors, a random orthogonal rotation matrix is generated through Householder reflection. (During Householder reflection, a random vector is first generated using a standard normal distribution as the reflection vector, and then the random orthogonal rotation matrix is calculated.) , Represents a random orthogonal rotation matrix. This represents the identity matrix in the standard normal distribution. Represents the reflection vector. (This represents the transpose operation). A random orthogonal rotation matrix is used to perform a rotation transformation on the latent vector (i.e., multiplying the random orthogonal rotation matrix with the latent vector). After obtaining the rotated latent vector, the first part is truncated. Each dimension (e.g., take) To balance convergence speed and structural expressiveness, the vectors are concatenated to construct a low-dimensional Bayesian optimized vector. This vector is then zero-pasted to restore it to the same dimension as the rotated latent vector. A random orthogonal rotation matrix is used to inversely map the zero-pasted low-dimensional Bayesian optimized vector, generating a high-dimensional Bayesian optimized vector. This vector is then input into a pre-trained GraphVAE variational encoder for decoding, resulting in a decoded art design logic structure diagram. Finally, an SVG rule engine is used to render the decoded art design logic structure diagram into candidate images.
[0070] Traditional Bayesian optimization often faces slow convergence and local optima when optimizing high-dimensional latent spaces. Conventional methods (such as PCA and t-SNE) rely on linear transformations for dimensionality reduction, which can lead to a loss of structural information in high-dimensional spaces, resulting in a lack of diversity in the generated results and an inability to fully capture the complex features of the data. This invention, however, generates a random orthogonal rotation matrix using Householder reflection. This method not only preserves the geometric structure of the latent vector space but also rotates the data with high numerical stability. This rotation mechanism ensures that the data is effectively preserved in the low-dimensional space, allowing the optimization process to proceed efficiently while retaining the essential features of the data. This solves the common problems of slow convergence and local optima in high-dimensional spaces, improving the convergence speed and global search capability of Bayesian optimization. Furthermore, after rotating the latent vectors using the random orthogonal rotation matrix, the latent vector space is effectively explored randomly, allowing the optimization search to extend beyond local regions and perform a global search in a broader latent design space. This method introduces dynamic rotation during the optimization process, increasing the diversity of the search space and ensuring the global nature of the optimization results. Furthermore, the construction of low-dimensional Bayesian optimization vectors and the zero-padding and inverse mapping of rotated latent vectors further enhance the accuracy of the optimization process and the quality of the generated results. Through zero-padding and inverse mapping, the system can restore the same dimension as the initial latent space, ensuring that the generated design is both consistent with the data and operable. Secondly, the integration with the GraphVAE encoder makes the generated design not just an image, but a structured graph containing design logic, which provides designers with great interpretability. Finally, the application of the SVG rule engine ensures high-quality output of the design results while providing flexibility for design modifications.
[0071] Further, candidate embedding vectors are obtained, and the target consistency loss value is calculated by combining it with the target embedding vector and converted into a maximization scoring function. Gaussian process modeling and sampling are then performed to generate the optimal result and construct a pairwise preference probability function. Finally, a ranking likelihood function is constructed for parameter updates and iterations to obtain the final image, including:
[0072] Based on the candidate images, the semantic vectors of the candidate images are obtained through a pre-trained CLIP network model, and then normalized to obtain the candidate embedding vectors, as shown in the formula:
[0073]
[0074] In the formula, Representing candidate images candidate embedding vectors, Representing candidate images Image semantic vectors, Represents the L2 norm;
[0075] The target embedding vector is normalized, and combined with the candidate embedding vectors, the target consistency loss value is calculated and converted into a Bayesian optimization maximization scoring function. After calculating the scoring function value as the label of the low-dimensional Bayesian optimization vector, Gaussian process modeling is performed to generate a Gaussian process surrogate model for initialization. Expected Improvement (EI) is used to sample within the 3σ effective domain of the low-dimensional Bayesian optimization vector to obtain a sequence of low-dimensional Bayesian optimization vectors. After obtaining the label (i.e., the maximum scoring function value) of each low-dimensional Bayesian optimization vector in the sequence through the Gaussian process surrogate model, it is sorted in descending order. The low-dimensional Bayesian optimization vector with the largest label is selected as the candidate vector. Then, zero-padding is performed to restore it to the same dimension as the rotated latent vector, and this process is repeated until a new candidate image is rendered, resulting in a new candidate image set containing the new candidate image and its corresponding label. The new candidate image with the largest label is selected from the candidate image set as the optimal result.
[0076] The formula for calculating the target consistency loss value is as follows:
[0077]
[0078] In the formula, This represents the loss value related to target consistency.
[0079] The formula for calculating the scoring function value is as follows:
[0080]
[0081] In the formula, This represents the value of the scoring function;
[0082] The Gaussian process modeling is described by the following formula:
[0083]
[0084] In the formula, It indicates obedience to symbols. Represents the Gaussian process function. Represents the mean function, Represents the kernel function. Represent another low-dimensional Bayesian optimization vector;
[0085] The specific form of the kernel function is:
[0086]
[0087] In the formula, This represents the variance of the scoring function values. Represents an exponential function. Represents Euclidean distance. This represents a length scale, and the length scale This can be used as a decay control factor for the correlation between input points and function values in the kernel function. Its value ranges from 0.2 to 1. When the value is small, it is more sensitive to input changes and easy to fit small differences, but it may lead to overfitting. When the value is large, it tends to be smoother, but it may lose the ability to capture local structures. Therefore, this invention uses 0.6 as the length scale. The default value;
[0088] Extract the labels of the optimal result and each new candidate image. Combine the labels of the optimal result with the labels of each other new candidate image to form preference pairs, construct a pairwise preference probability function, and calculate the pairwise preference probability value of each preference pair. The formula is as follows:
[0089]
[0090] In the formula, This represents the pairwise preference probability value. Indicates the optimal result Low-dimensional Bayesian optimized vectors Symbols representing preference relations. Indicates the first Low-dimensional Bayesian optimized vectors for new candidate images The cumulative distribution function represents the standard normal distribution. Indicates the optimal result The tag, Indicates the first Labels for new candidate images This represents the preference noise (the preference noise) This represents the degree of uncertainty in subjectively judging the quality of art and design, with a value ranging from 0.05 to 0.3. (When there is a preference for noise...) When the value is large, the preference judgment becomes more ambiguous and less sensitive to differences; when preference noise increases... When the value is small, the preference judgment is more stringent and sensitive to small differences; therefore, 0.15 can be taken as the preference noise. (default value);
[0091] The specific form of the cumulative distribution function of the standard normal distribution is:
[0092]
[0093] In the formula, Indicates the maximum number of points. The normalization constant term represents the normalization constant of the standard normal distribution. The base of the natural logarithm. Represents the exponent term. Represents the integral variable. Representing the integral variable Infinitesimal change;
[0094] Construct a ranking likelihood function (maximize), and evaluate the variance of the kernel function's score function value. Length scale The mean function of the Gaussian process model is updated (the update stops when the maximum number of iterations is reached), and the formula is:
[0095]
[0096] In the formula, This represents the sorting likelihood function value. Indicates belonging to, Represents a set of preference pairs (statistically, all preference pairs generated). Represent the natural logarithm function;
[0097] The low-dimensional Bayesian optimized vector with the largest output score function value is zero-padding to restore it to the same dimension as the rotated latent vector, and then rendered into the final image.
[0098] By using the CLIP model, images and text are mapped to the same shared semantic space, achieving cross-modal understanding of images and user text descriptions. This allows the generated images to not only maintain semantic consistency with the target image but also to be optimized in real time based on user feedback. By introducing target consistency loss and Bayesian optimization, not only is the quality of the generated images ensured, but optimization can also be performed according to specific user needs (such as design style, color preferences, etc.). Furthermore, an image generation and optimization strategy combining Bayesian optimization and Gaussian process modeling is employed. Through surrogate modeling of Gaussian processes, the complexity of the potential image optimization space can be effectively captured, and the optimal balance between global search and local optimization can be found using expectation boosting (EI) sampling. Compared to traditional optimization methods, this invention can converge to the optimal solution faster in high-dimensional complex spaces and effectively avoid overfitting. Subsequently, the concept of pairwise preference probability is introduced, and by modeling user preferences for the generated images, the generation process is adjusted based on user feedback. By constructing a ranking likelihood function and combining it with Bayesian optimization, image features can be flexibly adjusted according to user preferences during the optimization process, achieving personalized and customized image optimization. This feedback mechanism not only improves the visual quality of the images, but also makes the optimization process more in line with the user's creative requirements.
[0099] S3. Perform a preliminary display based on the final image, and have the user judge the image satisfaction. If the user is not satisfied, set the image parameters, optimize the image parameters by combining the final image and the art design logic structure diagram, obtain the final image parameters, and generate an adjustment reference auxiliary image.
[0100] Specifically, a preliminary display is performed based on the final image, and the user judges the image satisfaction. If dissatisfied, image parameters are set, and the image parameters are optimized by combining the final image and the art design logic structure diagram to obtain the final image parameters and generate an adjustment reference auxiliary image, including:
[0101] The final image is displayed to the user via HTML5 Canvas, where the user rates the final image with 0 or 1 (0 for satisfied and 1 for dissatisfied). When the user is dissatisfied, the user sets the required image parameters, including contrast, height, and color. The pixel values and image parameters of the final image are recorded. The image parameters set by the user and the image parameters of the final image are concatenated into a parameter vector, including a graphic parameter vector and a user parameter vector.
[0102] Based on the recorded pixel values, the mean square error of the pixel values between the final image and the design logic structure diagram is calculated using the following formula:
[0103]
[0104] In the formula, This represents the mean square error value. This represents the total number of pixel values in the final image. Indicates the first in the final image pixel value, Represents the first in the design logic structure diagram Each pixel value;
[0105] It should be noted that when calculating the mean square error, the design logic structure diagram needs to be rasterized. That is, the artistic design logic structure diagram is rasterized into a bitmap using the SVG rendering library librsvg or Inkscape, and the pixel matrix of the rasterized design logic structure diagram is obtained. Then, the mean square error of the pixel values is calculated with the final image.
[0106] Based on the parameter vector, the feedback loss between the graphics parameter vector and the user parameter vector is calculated, and a weighted fusion formula is constructed by combining the mean square error.
[0107] The formula for calculating the feedback loss between the final image and the image parameters set by the user is:
[0108]
[0109] In the formula, This represents the feedback loss value. Represents Euclidean distance. Represents an image parameter vector. Represents a user parameter vector;
[0110] The weighted fusion formula is as follows:
[0111]
[0112] In the formula, This represents the weighted fusion value. The weights representing the mean square error values. The weights represent the feedback loss values;
[0113] and The values of are all between 0 and 1, and based on the optimization effect of balance, can be and The default value is set to 0.5;
[0114] The user parameter vector is used as an individual in a hybrid algorithm (differential evolution and bat optimization). After initializing the population by randomly generating it, the weighted fusion formula is used as the fitness function of the individual. The fitness function value of each individual is calculated. The differential evolution algorithm is used to randomly select three individuals from the population to perform mutation operations, generating a new user parameter vector. The new user parameter vector is then crossed with each dimension of the user parameter vector to obtain a cross user parameter vector. The fitness function value is recalculated, and the user or cross user parameter vector is retained as the optimal individual.
[0115] The differential evolution algorithm is used to randomly select three individuals from the population to perform mutation operations, and the formula is as follows:
[0116]
[0117] In the formula, Indicates the first New user parameter vectors for each individual , and These represent the user parameter vectors of the 1st, 2nd, and 3rd randomly selected individuals, respectively. This represents the scaling factor for the mutation operation. It can be set via cross-validation, and the value of this parameter ranges from 0 to 2, with 1 as the default value;
[0118] The formula for performing the cross operation between the new user parameter vector and the user parameter vector is as follows:
[0119]
[0120] In the formula, Indicates the first The first individual generated through crossover operation dimensional components, Indicates the first The first of the new user parameter vectors of each individual dimensional components, This represents a random number (generated by a random number generator). This represents the crossover probability (which can be set through cross-validation, and the value of this parameter ranges from 0.5 to 1; in this invention, 0.8 can be used as the default value). Indicates the first The first individual's user parameter vector dimensional components, Indicates otherwise;
[0121] The formula for retaining the new or cross-referenced user parameter vector as the optimal individual is as follows:
[0122]
[0123] In the formula, Indicates the first The optimal individual (i.e., the new or crossover user parameter vector) is generated by the crossover mutation operation of each individual. Indicates the first The cross-user parameter vector of each individual, Indicates the first The fitness function value of the cross-user parameter vector of each individual. Indicates the first The fitness function value of the user parameter vector of each individual;
[0124] Based on the optimal individual, the velocity and position of the optimal individual are updated through the echolocation operation of the bat optimization algorithm. After the maximum number of iterations, the final individual is output and the parameter vector is extracted.
[0125] The formula for updating the velocity and position of the optimal individual is as follows:
[0126]
[0127]
[0128] In the formula, Indicates the first The update speed of each individual Indicates the first The initial velocity of each individual, Indicates the first The individual in the first Position during iteration Indicates the first The position of each individual when the optimal individual is generated through crossover and mutation operations. Indicates the first The step size factor for each individual (which can be generated by a random number generator). Indicates the first The updated position of each individual;
[0129] Based on the extracted parameter vector, the user uses OpenCV to perform item-by-item transformation operations on the final image to generate an adjustment reference auxiliary image.
[0130] By incorporating real-time user feedback, a personalized adjustment process is added. When users rate the generated image and point out areas of dissatisfaction, this invention can instantly adjust the generated image according to user preferences, providing a fully dynamic and interactive image optimization process. Users can not only set image parameters such as brightness and color, but also precisely adjust advanced features such as style and structure through feedback. This dynamic optimization mechanism based on real-time user feedback effectively improves the personalization and satisfaction of the design. Furthermore, it combines mean squared error (MSE) with user feedback loss for weighted optimization. This method can simultaneously optimize image quality and user needs, ensuring a balance between the two in the final image. Moreover, the setting of weight coefficients allows the optimization process to consider the influence of image quality and user feedback simultaneously. This not only improves the pixel-level accuracy of the image but also ensures a high degree of consistency in the image in terms of user preferences, avoiding the limitations of a single optimization objective. Secondly, it combines two optimization algorithms: Differential Evolution (DE) and Bat Optimization (BAT). Differential Evolution is used for local optimization of image parameters, using mutation and crossover operations to quickly explore the solution space; while Bat Optimization provides powerful global search capabilities, avoiding getting trapped in local optima. By combining the advantages of both approaches, this invention achieves the goal of both global exploration and local fine-grained optimization, making it particularly suitable for complex image design optimization tasks. Furthermore, by converting user-defined parameters (such as color, brightness, and contrast) into vectors and then processing them as individual components of the optimization algorithm, a new optimization strategy is formed. This strategy not only fully reflects user needs but also integrates closely with the image generation process, further enhancing the personalization of image generation. Finally, the combination of crossover and mutation operations is optimized using a hybrid strategy of Differential Evolution (DE) and Bat Optimization (BAT). This method ensures that the algorithm avoids local optima while gradually converging to the global optimum during multiple optimization processes.
[0131] S4. Based on the adjusted reference auxiliary image, perform the final visualization auxiliary display;
[0132] Specifically, based on the adjusted reference auxiliary image, the final visualization auxiliary display is performed, including:
[0133] The reference auxiliary image is adjusted and used as an auxiliary image when the user is designing the artwork. After the image is displayed synchronously with the artwork design logic structure diagram, the click response area is set in the image area using the HTML5Canvas image display framework.
[0134] When a user hovers the mouse over a click response area (such as a repeating pattern area), the nodes in the click response area are highlighted and the design operation type of the node is displayed (e.g., "symmetric transformation, angle 45°").
[0135] By displaying candidate images synchronously with the design logic structure diagram, designers can intuitively see the relationship between the images and the underlying design operations, thus making design decisions more effectively. Secondly, the introduction of a click-responsive area function allows designers to highlight design nodes in real time when the mouse hovers over or clicks, and display the specific operation type of the node (such as "symmetric transformation, angle 45°"), effectively reducing misoperations and improving design accuracy.
[0136] This embodiment also provides an art design assistance system based on artificial intelligence, including:
[0137] The initial execution module is used to construct an art design logic structure diagram for encoding and mapping, generate latent vectors, perform preprocessing and semantic extraction based on user-input target design reference information, generate target embedding vectors for dimensionality reduction, construct low-dimensional Bayesian optimized vectors, and then perform zero-padding, inverse mapping, and decoding reconstruction.
[0138] The update module is used to obtain candidate embedding vectors, calculate the target consistency loss value by combining it with the target embedding vector, convert it into a maximization scoring function, perform Gaussian process modeling and sampling, generate the optimal result and construct the pair preference probability function, and then construct the ranking likelihood function for parameter update and iteration.
[0139] The feedback optimization module is used to perform an initial display and allow the user to judge the image satisfaction. If the user is not satisfied, the image parameters are set and optimized by combining the final image and the art design logic structure diagram.
[0140] The visualization module is used to perform visual aids based on the final image.
[0141] This embodiment also provides a computer device applicable to the case of an art design assistance method based on artificial intelligence, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the art design assistance method based on artificial intelligence as proposed in the above embodiment.
[0142] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0143] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the artificial intelligence-based art design assistance method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0144] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An art design assistance method based on artificial intelligence, characterized in that: include, The art design logic structure diagram is constructed, encoded, and mapped to generate latent vectors. The target design reference information is preprocessed and semantically extracted by user input to generate target embedding vectors. Based on latent vectors, dimensionality reduction is performed through Householder reflection. After constructing a low-dimensional Bayesian optimized vector, zero-padding, inverse mapping, and decoding reconstruction are performed to obtain candidate images and candidate embedding vectors. The target consistency loss value is calculated by combining the target embedding vector and converted into a maximization scoring function. Gaussian process modeling and sampling are then performed to generate the optimal result and construct a pairwise preference probability function. Finally, a ranking likelihood function is constructed for parameter updates and iterations to obtain the final image. A preliminary display is performed based on the final image, and the user judges the image satisfaction. If the user is not satisfied, the image parameters are set, and the image parameters are optimized by combining the final image and the art design logic structure diagram to obtain the final image parameters and generate an adjustment reference auxiliary image. Based on the adjusted reference auxiliary image, the final visualization is then performed.
2. The art design assistance method based on artificial intelligence as described in claim 1, characterized in that: The process involves using latent vectors to perform dimensionality reduction via Householder reflection, constructing low-dimensional Bayesian optimized vectors, and then performing zero-padding, inverse mapping, and decoding reconstruction to obtain candidate images, including: Based on the latent vector, a random orthogonal rotation matrix is generated using Householder reflection. This random orthogonal rotation matrix is then used to perform a rotation transformation on the latent vector, resulting in a rotated latent vector. The previous state is then truncated. The vectors are concatenated along their dimensions to construct a low-dimensional Bayesian optimized vector. Zero-padding is then performed to restore the vectors to the same dimension as the rotated latent vectors. A random orthogonal rotation matrix is used to inversely map the zero-padding low-dimensional Bayesian optimized vector, generating a high-dimensional Bayesian optimized vector. This vector is then input into a pre-trained GraphVAE variational encoder for decoding, resulting in a decoded art design logic structure diagram. Finally, an SVG rule engine is applied to render this decoded art design logic structure diagram as a candidate image.
3. The art design assistance method based on artificial intelligence as described in claim 2, characterized in that: The process involves obtaining candidate embedding vectors, calculating the target consistency loss value by combining it with the target embedding vector, converting it into a maximization scoring function, performing Gaussian process modeling and sampling, generating the optimal result and constructing a pairwise preference probability function, and then constructing a ranking likelihood function for parameter updates and iterations to obtain the final image, including: Based on candidate images, semantic vectors of candidate images are obtained through a pre-trained CLIP network model and then normalized to obtain candidate embedding vectors. Normalization is then performed on these vectors. Combined with the candidate embedding vectors, the target consistency loss value is calculated and converted into a Bayesian optimization maximization scoring function. The scoring function value is then used as the label of the low-dimensional Bayesian optimized vector. Gaussian process modeling is then performed to generate a Gaussian process surrogate model for initialization. Expectation boosting is used to sample within the 3σ effective domain of the low-dimensional Bayesian optimized vectors, resulting in a sequence of low-dimensional Bayesian optimized vectors. The label of each low-dimensional Bayesian optimized vector in the sequence is obtained through the Gaussian process surrogate model, and then sorted in descending order. The low-dimensional Bayesian optimized vector with the largest label is selected as the candidate vector. Zero-padding is then performed to restore it to the same dimension as the rotated latent vector, and this process continues until a new candidate image is rendered, resulting in a new candidate image set containing the new candidate image and its corresponding label. The new candidate image with the largest label is selected from the candidate image set as the optimal result. Extract the labels of the optimal result and each new candidate image. Combine the labels of the optimal result with the labels of each other new candidate image to form preference pairs, construct a pairwise preference probability function, calculate the pairwise preference probability value of each preference pair, construct a ranking likelihood function, and calculate the variance of the score function value of the kernel function. Length scale The mean function of the Gaussian process model is updated, and the low-dimensional Bayesian optimization vector with the largest output score function value is zero-padding to restore it to the same dimension as the rotated latent vector, until it is rendered into the final image.
4. The art design assistance method based on artificial intelligence as described in claim 3, characterized in that: The process involves performing a preliminary display based on the final image, followed by a user-assessed image satisfaction judgment. If the user is not satisfied, image parameters are set, and the image parameters are optimized by combining the final image with the art design logic structure diagram to obtain the final image parameters and generate an adjustment reference auxiliary image. This includes: The final image is displayed to the user via HTML5 Canvas, where the user rates the final image with 0 or 1. If the user is not satisfied, the user sets the required image parameters, including contrast, height, and color. The pixel values and image parameters of the final image are recorded. The image parameters set by the user and the image parameters of the final image are concatenated into a parameter vector, including a graphic parameter vector and a user parameter vector. Based on the recorded pixel values, calculate the mean square error of pixel values between the final image and the design logic structure diagram; Based on the parameter vector, the feedback loss between the graphics parameter vector and the user parameter vector is calculated, and a weighted fusion formula is constructed by combining the mean square error. The user parameter vector is used as an individual in the hybrid algorithm. After initializing the population by randomly generating it, the weighted fusion formula is used as the fitness function of the individual. The fitness function value of each individual is calculated. The differential evolution algorithm is used to randomly select three individuals from the population to perform mutation operations, generating a new user parameter vector. The new user parameter vector is crossed with each dimension of the user parameter vector to obtain the cross user parameter vector. The fitness function value is recalculated, and the user or cross user parameter vector is retained as the optimal individual. Based on the optimal individual, the velocity and position of the optimal individual are updated through the echolocation operation of the bat optimization algorithm. After the maximum number of iterations, the final individual is output and the parameter vector is extracted. Based on the extracted parameter vector, the user uses OpenCV to perform item-by-item transformation operations on the final image to generate an adjustment reference auxiliary image.
5. The art design assistance method based on artificial intelligence as described in claim 4, characterized in that: The process of performing the final visualization based on the adjusted reference auxiliary image includes: The reference auxiliary image is adjusted and used as an auxiliary image during the user's art design. After it is displayed synchronously with the art design logic structure diagram, a click response area is set in the image area using the HTML5Canvas image display framework. When the user hovers the mouse over the click response area, the node in the click response area is highlighted and the design operation type of the node is displayed.
6. The art design assistance method based on artificial intelligence as described in claim 5, characterized in that: The construction of the art design logic structure diagram is encoded and mapped to generate latent vectors, including: Design diagram samples are obtained from open-source design databases using web crawling technology and parsed to obtain design operations. Each operation in the design operation is taken as a node, and a directed acyclic graph is constructed according to the execution order relationship between the design operations. The directed acyclic graph is then uploaded to the storage as an art design logic structure diagram. A pre-trained GIN network and a GraphVAE variational autoencoder are added to the memory side. The art design logic structure graph is input into the pre-trained GIN network, and the output hidden state is used to perform mean aggregation to generate a graph-level embedding representation. Then, it is input into the GraphVAE variational autoencoder to perform spatial mapping and output latent space distribution parameters, including mean vector and standard deviation vector. Latent vectors are generated through reparameterization sampling technology.
7. The art design assistance method based on artificial intelligence as described in claim 6, characterized in that: The step of preprocessing and semantic extraction of target design reference information input by the user to generate a target embedding vector includes: The user inputs target design reference information into the terminal, including target reference images and target design text descriptions. The target design reference information is preprocessed to obtain input data. The input data is used as input to the pre-trained CLIP network model in OpenAI to obtain image and text semantic vectors respectively. Weighted fusion is then performed to generate target embedding vectors.
8. An art design assistance system based on artificial intelligence, based on the art design assistance method based on artificial intelligence as described in any one of claims 1 to 7, characterized in that: include, The initial execution module is used to construct an art design logic structure diagram for encoding and mapping, generate latent vectors, perform preprocessing and semantic extraction based on user-input target design reference information, generate target embedding vectors for dimensionality reduction, construct low-dimensional Bayesian optimized vectors, and then perform zero-padding, inverse mapping, and decoding reconstruction. The update module is used to obtain candidate embedding vectors, calculate the target consistency loss value by combining it with the target embedding vector, convert it into a maximization scoring function, perform Gaussian process modeling and sampling, generate the optimal result and construct the pair preference probability function, and then construct the ranking likelihood function for parameter update and iteration. The feedback optimization module is used to perform an initial display and allow the user to judge the image satisfaction. If the user is not satisfied, the image parameters are set and optimized by combining the final image and the art design logic structure diagram. The visualization module is used to perform visual aids based on the final image.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the artificial intelligence-based art design assistance method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the artificial intelligence-based art design assistance method according to any one of claims 1 to 7.