Art image generation method based on semantic graph guidance and style prior constraint
By employing semantic graph guidance and style prior constraints, and utilizing the Stable Diffusion model and feature fusion technology, this approach addresses the shortcomings in the quality and controllability of artistic image generation in existing technologies, resulting in the generation of high-quality, controllable artistic images.
Patent Information
- Application Number
- CN202610178244.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-08
- Publication Date
- 2026-05-29
AI Technical Summary
Existing artistic image generation methods struggle to fully and accurately control the content and style of images, resulting in insufficient quality and controllability of the generated results.
By using semantic graph guidance and style prior constraints, features are extracted using a pre-trained Stable Diffusion model, and feature fusion and matching are performed. Combined with a semantic feature prediction network and style difference measurement, details in the generated results are gradually corrected to generate an artistic image that combines the target semantic structure and painting style.
It significantly improves the quality and controllability of artistic image generation, making the generated images closer to the level of human expert artworks, and achieving high-quality and highly controllable artistic image generation.
Smart Images

Figure CN122115624A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and image generation technology, specifically relating to an artistic image generation method based on semantic graph guidance and style prior constraints. Background Technology
[0002] Artistic image generation is a current research hotspot in computer vision and generative artificial intelligence. It aims to automatically generate artistic images with specific styles, themes, or that conform to certain aesthetic rules through algorithms and models. This involves fundamental theories such as probabilistic modeling, statistical learning, and information theory, as well as key technologies such as image understanding, deep learning, and multimodal learning, and has significant academic research value. In recent years, scholars both domestically and internationally have conducted extensive and in-depth research on image generation, proposing various types of generative models, including generative adversarial networks, flow models, variational autoencoders, and diffusion models. Among these, diffusion models, with their excellent flexibility, diversity, stability, and extremely high generation quality, have become one of the most advanced generative models currently available. However, artistic images differ from ordinary images, typically possessing unique semantic structures and distinct artistic styles. To more accurately control the generation of artistic images, the desired artistic image can be described from two levels: content and style. Content refers to elements such as objects, scenes, layouts, and structures in the image, which determine the specific information of the subject or object expressed by the image. Style refers to elements such as color, texture, brushstrokes, and lighting in the image, reflecting the creative characteristics, painting techniques, and unique visual expression of the image. Current methods typically use text to describe the target content and style, but due to its complexity, text struggles to fully and accurately convey this information. These issues limit the practicality of current models and the quality of generated results, representing one of the bottlenecks in the current intelligent generation of artistic images. Summary of the Invention
[0003] The purpose of this invention is to provide an artistic image generation method based on semantic graph guidance and style prior constraints. On the one hand, semantic graphs are used to accurately control the content structure of the generated image, and on the other hand, style prior information in existing paintings is fully utilized, thereby effectively promoting the technological development of the generation model in terms of high quality, high controllability, and artistry.
[0004] The technical solution for achieving the purpose of this invention is: an artistic image generation method based on semantic graph guidance and style prior constraints, comprising:
[0005] Step 1, extract the target semantic map ,painting and its corresponding source semantic graph Input a pre-trained Stable Diffusion model and use the encoder E in the model to extract... , and The characteristics were obtained. , as well as ;
[0006] Step 2, feature and By cascading along the channel dimension, style fusion features are obtained. To fully utilize the structural information in the semantic graph; and simultaneously incorporate features By cascading it with itself along the channel dimension, content fusion features are obtained. , in order to Maintain consistency across dimensions;
[0007] Step 3, according to and The semantic correlation between them is used to reorganize style features in space through feature matching and exchange operations, so as to initially realize the fusion of content structure information and painting style information in the target semantic map;
[0008] Step 4: Construct a semantic feature prediction network to provide semantic guidance for the model during the back diffusion process, correct the semantic details in the initial fusion results, and avoid inconsistencies with the target semantic map at the region boundaries;
[0009] Step 5, during the back diffusion process, measure the implicit coding in the model. and painting characteristics The differences in style, and then the guidance of prior information on painting style, gradually correct the stylistic details in the aforementioned preliminary fusion results;
[0010] Step 6: After the reverse diffusion process is completed, the obtained denoised image features are... The decoder D in the input model generates a semantic graph that also contains the target semantics. Content structure and drawing Artistic images in a distinctive style.
[0011] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described above.
[0012] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the above-described method.
[0013] A computer program product includes a computer program that, when executed by a processor, implements the above-described method.
[0014] The art image generation method based on semantic graph guidance and style prior constraints provided by this invention can render the style of a painting onto the corresponding area of the semantic graph according to the semantic correspondence, thereby generating a new art image. Its beneficial effects are:
[0015] (1) Compared with other art image generation methods, this invention fully explores and utilizes the style prior information in existing paintings, combines the advantages of semantic graphs in precise control of content structure, and relies on the powerful generation capability of diffusion models, thereby significantly improving the quality and controllability of generated art images.
[0016] (2) Compared with other artistic image generation methods, this invention proposes a semantically aware feature matching and exchange method. Based on the semantic correlation between content features and style features, style features are reorganized in space, thereby achieving effective alignment between content features and style features.
[0017] (3) Compared with other artistic image generation methods, this invention solves the problem of predicting semantic structure at any time step during the inverse process of the diffusion model. Therefore, guided by the semantic structure in the semantic graph, it can gradually correct the semantic details in the generated results, significantly improving the performance of the generated results in terms of content preservation.
[0018] (4) Compared with other artistic image generation methods, this invention proposes a fine-grained style optimization method, which can use the style prior in painting as a constraint to gradually correct the style details in the generated result during the back diffusion process, significantly improving the performance of the generated result in style learning. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the background art, the accompanying drawings used in the embodiments of the present invention or the background art will be described below.
[0020] Figure 1 This is a flowchart illustrating an art image generation algorithm based on semantic graph guidance and style prior constraints.
[0021] Figure 2 This is a flowchart illustrating the feature matching and exchange strategy in one embodiment;
[0022] Figure 3 This is a flowchart illustrating a semantic detail correction strategy guided by a semantic feature prediction network in one embodiment.
[0023] Figure 4 This is a flowchart illustrating a style detail correction strategy under style prior constraints in one embodiment. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0025] This invention proposes an art image generation method based on semantic graph guidance and style prior constraints. Under the condition that the semantic graph represents the target content and the painting represents the target style, the method renders the style of the painting into the corresponding area of the semantic graph according to the semantic correspondence between the two, thereby generating a new art image. Figure 1 As shown, the method includes:
[0026] S1, target semantic map ,painting and its corresponding source semantic graph Input a pre-trained Stable Diffusion model and use the encoder E in the model to extract... , and The characteristics were obtained. , as well as ;
[0027] S2, features and By cascading along the channel dimension, style fusion features are obtained. To fully utilize the structural information in the semantic graph; and simultaneously incorporate features By cascading it with itself along the channel dimension, content fusion features are obtained. , in order to Maintain consistency across dimensions;
[0028] S3, according to and The semantic correlation between them is used to reorganize style features in space through feature matching and exchange operations, so as to initially realize the fusion of content structure information and painting style information in the target semantic map;
[0029] S4. Construct a semantic feature prediction network P to provide semantic guidance for the model during the reverse diffusion process, correct the semantic details in the aforementioned preliminary fusion results, and avoid inconsistencies with the target semantic map at key locations such as region boundaries.
[0030] S5, during the back diffusion process, the implicit coding in the measurement model and painting characteristics The differences in style, and then the guidance of prior information on painting style, gradually correct the stylistic details in the aforementioned preliminary fusion results;
[0031] S6, After the reverse diffusion process is completed, the obtained denoised image features The decoder D in the input model generates a semantic graph that also contains the target semantics. Content structure and drawing Artistic images in a distinctive style.
[0032] The proposed method for generating artistic images based on semantic graph guidance and style prior constraints starts from the different characteristics of content and style attributes. It uses semantic graphs to represent the target content, allowing for more controllable and precise specification of the required generation scene. Simultaneously, it employs painting to represent the target style, utilizing style prior information such as color scheme, brushstrokes, texture, and lighting. Through these methods, the effectiveness of artistic image generation can be significantly improved, resulting in higher-quality images that more closely resemble works of art created by human experts.
[0033] Optionally, a pre-trained Stable Diffusion model can be used as the base model. Stable Diffusion, based on a pre-trained autoencoder (containing an image encoder E and a decoder D), transfers the diffusion process, originally performed in the high-dimensional pixel space, to the low-dimensional latent space to reduce computational resource consumption. Furthermore, the backbone of Stable Diffusion is a U-Net denoising network responsible for predicting noise at each time step.
[0034] Optionally, the art image generation method based on semantic graph guidance and style prior constraints mainly includes three inputs: the target semantic graph. ,painting and the source semantic graph corresponding to the painting Semantic graphs can be obtained through manual annotation or by using semantic segmentation technology. Users can edit semantic graphs according to their own needs, thus this generation method has high interactivity.
[0035] Optionally, the core idea of the art image generation method based on semantic graph guidance and style prior constraints is to first achieve the alignment of semantic features and style features in the latent space of the autoencoder, and then gradually make detailed corrections during the back diffusion process.
[0036] Several alternative methods are provided below, but they are not intended as additional limitations on the overall solution above. They are merely further additions or optimizations. Provided there are no technical or logical contradictions, each alternative method can be combined individually with respect to the overall solution above, or multiple alternative methods can be combined with each other.
[0037] Specifically, in S1, the pre-trained Stable Diffusion is used as the base model for construction; the encoder E in the model is used to extract the target semantic map. ,painting and its corresponding source semantic graph Features:
[0038]
[0039] Specifically, in S2, for features and To fully utilize the structural information in the semantic graph:
[0040]
[0041] in, This represents cascading at the channel level. Unlike... Features of the target semantic graph Since there is no corresponding artwork image, it is used as the initial artwork image for feature fusion:
[0042]
[0043] and These represent stylistic integration and content integration characteristics, respectively.
[0044] Specifically, the feature matching and exchange strategy in S3 is as follows: Figure 2 As shown, the following feature matching and swapping operations are performed to transfer the style patterns in the painting to the corresponding positions in the target semantic map:
[0045] S301, respectively and Cut into several Size of feature blocks and ,in and Represents the number of feature blocks;
[0046] S302, for each target feature block Based on the normalized cross-correlation algorithm, find The source feature block that best matches it ;
[0047] S303, using Feature blocks in To replace Feature blocks in In order to Refactor;
[0048] S304, repeating the operations in S302 and S303 continuously, finally obtains the complete reconstructed features. .
[0049] Thus, the structural information in the target semantic map and the stylistic information in the painting can be initially integrated.
[0050] On the other hand, the semantic detail correction strategy guided by the semantic feature prediction network in S4 is as follows: Figure 3 As shown, it includes:
[0051] S401, to address the potential feature mismatch issue during the initial fusion process, a semantic feature prediction network P is constructed to provide semantic guidance to the model during the back-diffusion process. The specific construction process is as follows: Given an image... and its corresponding semantic graph First, encoder E is used to map both to the feature space. For image features... Noise is added during the forward diffusion process to obtain the features. , where t represents the time step. Then... The U-Net network is input to t during the reverse diffusion process, and features from each layer are extracted. After scaling them to the same scale, they are concatenated to obtain the feature vectors. Next, the features Input a semantic feature prediction network P at time step t to predict the value at that time. The corresponding semantic structure.
[0052] S402, semantic graph features Using the output of the semantic feature prediction network P as the true label, the following loss function is constructed:
[0053]
[0054] P is trained under the constraints of the aforementioned loss function to achieve accurate prediction of semantic features.
[0055] S403, during the reverse diffusion process, the noise features at each time step t are measured based on the semantic feature prediction network P trained in the previous step. With target semantic features Differences in semantic structure:
[0056]
[0057] S404, in the semantic difference measurement function Under the constraints, the semantic details in the aforementioned preliminary fusion results are corrected to avoid inconsistencies with the target semantic map at key locations such as region boundaries.
[0058] On the other hand, the style detail correction strategy under the style prior constraint in S5 is as follows: Figure 4 As shown, in order to further refine the stylistic details in the preliminary fusion results, painting was used during the reverse diffusion process. The style prior information is used to further constrain the style in the generated image, including:
[0059] S501, noise characteristics and painting characteristics Divided into different feature blocks;
[0060] S502, regarding noise characteristics Each feature block in By combining semantic information, we can find the characteristics of paintings. The feature block that best matches it ;
[0061] S503, construct the following style difference measurement function. To calculate and Differences in style:
[0062]
[0063] Where N represents the number of feature blocks. and These represent the mean and standard deviation, respectively.
[0064] S504, in the style difference measurement function Under the constraints, the style details in the aforementioned preliminary fusion results are corrected to avoid inconsistencies between the generated results and the input painting style.
[0065] Specifically, the denoised image features in S6 This is achieved as follows: To optimize both semantic and stylistic details simultaneously, the semantic difference metric function is... and style difference measurement function Combining these, we obtain the following overall objective loss function:
[0066]
[0067] in, This is a hyperparameter used to adjust the importance of the preceding and following functions. Finally, using... The gradient is used to guide the generation process of the diffusion model:
[0068]
[0069] in, It is a hyperparameter used to control The update step size. In the loss function Chinese variables The gradient. After optimization, the result was obtained. Then the diffusion model was applied to Denoising is performed to obtain the noise characteristics of the next time step. And continue repeating the above process. Thus, in the continuous iterative process, It will gradually approach the generated target. Finally, the decoder D in the model will convert the denoised features... Convert to artistic image.
[0070] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating artistic images based on semantic graph guidance and style prior constraints, characterized in that, include: Step 1, extract the target semantic map ,painting and its corresponding source semantic graph Input a pre-trained Stable Diffusion model and use the encoder E in the model to extract... , and The characteristics were obtained. , as well as ; Step 2, feature and By cascading along the channel dimension, style fusion features are obtained. To fully utilize the structural information in the semantic graph; and simultaneously incorporate features By cascading it with itself along the channel dimension, content fusion features are obtained. , in order to Maintain consistency across dimensions; Step 3, according to and The semantic correlation between them is used to reorganize style features in space through feature matching and exchange operations, so as to initially realize the fusion of content structure information and painting style information in the target semantic map; Step 4: Construct a semantic feature prediction network to provide semantic guidance for the model during the back diffusion process, correct the semantic details in the initial fusion results, and avoid inconsistencies with the target semantic map at the region boundaries; Step 5, during the back diffusion process, measure the implicit coding in the model. and painting characteristics The differences in style, and then the guidance of prior information on painting style, gradually correct the stylistic details in the aforementioned preliminary fusion results; Step 6: After the reverse diffusion process is completed, the obtained denoised image features are... The decoder D in the input model generates a semantic graph that also contains the target semantics. Content structure and drawing Artistic images in a distinctive style.
2. The method for generating artistic images based on semantic graph guidance and style prior constraints as described in claim 1, characterized in that, The model is built based on a pre-trained Stable Diffusion model. First, its encoder E is used to extract the target semantic map. ,painting and its corresponding source semantic graph Features: 。 3. The method for generating artistic images based on semantic graph guidance and style prior constraints as described in claim 1, characterized in that, Features and To fully utilize the structural information in the semantic graph: ; in, This represents cascading at the channel dimension; unlike... Features of the target semantic graph Since there is no corresponding artwork image, it is used as the initial artwork image for feature fusion. 。 4. The method for generating artistic images based on semantic graph guidance and style prior constraints as described in claim 1, characterized in that, To transfer stylistic features from the painting to their corresponding positions in the target semantic map, the following feature matching and exchange operations are performed: (1) Each and Cut into several Size of feature blocks and ,in and Represents the number of feature blocks; (2) For each target feature block Based on the normalized cross-correlation algorithm, find The source feature block that best matches it ; (3) Use Feature blocks in To replace Feature blocks in In order to Refactor; (4) Repeat the operations in the second and third steps continuously to finally obtain the complete reconstructed features. .
5. The method for generating artistic images based on semantic graph guidance and style prior constraints as described in claim 1, characterized in that, A semantic feature prediction network P is constructed to provide semantic guidance to the model during the backpropagation process; the specific construction process is as follows: given an image and its corresponding semantic graph First, encoder E is used to map the two to the feature space; for Image features Noise is added during the forward diffusion process to obtain the features. Where t represents the time step; The U-Net network is input to t during the reverse diffusion process, and features from each layer are extracted. After scaling them to the same scale, they are concatenated to obtain the feature vectors. ; Features Input a semantic feature prediction network P at time step t to predict the value at that time. The corresponding semantic structure, and simultaneously semantic graph features As the true labels, the following loss function can be constructed: ; The network P is trained under the constraints of the above loss function to achieve accurate prediction of semantic features; During the reverse diffusion process, a pre-trained network P is used to measure the noise features at each time step t. With target semantic features Differences in content structure: ; It is used to guide the correction of semantic details during the reverse diffusion process, avoiding inconsistencies with the target semantic map at key locations such as region boundaries.
6. The method for generating artistic images based on semantic graph guidance and style prior constraints as described in claim 1, characterized in that, Using painting during the reverse diffusion process The style prior information in the model constrains the style in the generated image; for noise features... Each feature block in By combining semantic information, we can find the characteristics of paintings. The feature block that best matches it And calculate the stylistic differences between the two: ; Where N represents the number of feature blocks. and These represent the mean and standard deviation, respectively.
7. The method for generating artistic images based on semantic graph guidance and style prior constraints as described in claim 1, characterized in that, To optimize both semantic and stylistic details simultaneously, the above loss function is... and Combining these, we obtain the following overall objective loss function: ; in, It is a hyperparameter used to adjust the importance of the preceding and following functions; utilizing The gradient is used to guide the generation process of the diffusion model: ; in, It is a hyperparameter used to control Update step size; In the loss function Chinese variables The gradient; After optimization, the result was obtained. Then the diffusion model was applied to Denoising is performed to obtain the noise characteristics of the next time step. And continue to repeat the above process; thus, in the continuous iterative process, The model will gradually approach the target; finally, the decoder D in the model will convert the denoised features into a more accurate representation. Convert to artistic image.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.