A text-driven progressive neural painting method

By combining multi-level text guidance and three-stage progressive optimization, the semantic inconsistency and lack of detail in the generated images of scalar stroke neural painting methods in the prior art when text input is used are solved, and high-quality text-driven scalar stroke image generation is achieved.

CN122089880BActive Publication Date: 2026-07-28XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2026-04-22
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing scalar stroke neural painting methods struggle to generate semantically consistent, spatially well-structured, and locally detailed scalar stroke images when only receiving text input.

Method used

A text-driven progressive neural painting method is adopted, which uses multi-level text guidance conditions and a three-stage progressive optimization strategy to determine the structure, apply color and enhance details respectively. It combines a pre-trained large language model, a text-to-image diffusion model and a contrastive language-image pre-trained model for semantic parsing, spatial guidance and semantic alignment. The three-stage progressive optimization strategy of drafting-coloring-refining decouples the optimization goals of structure, color and detail, and outputs the final image.

Benefits of technology

It significantly improves the semantic consistency and visual quality of scalar stroke neural painting under text input conditions, generating images that are semantically consistent with the input text, have a reasonable structure, and are rich in detail.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089880B_ABST
    Figure CN122089880B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of text-driven progressive neural painting methods, comprising: the text information obtained is respectively carried out semantic analysis processing, space guiding processing and semantic alignment processing, obtain multi-level text guide condition;Adopt draft-coloring-fine three-stage progressive optimization strategy, combined with multi-level text guide condition, in turn, the optimization of structure determination, color paving and detail enhancement is carried out to brush parameter, obtain optimized foreground brush parameter;Optimized background brush parameter and optimized foreground brush parameter are rendered to the same canvas in turn by renderer, output is consistent with the final image of text information semantics and detail enhancement.It can significantly improve the semantic consistency and visual quality of scalar brush neural painting, fill the technical blank of text-driven scalar brush generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence content generation technology, specifically relating to a text-driven progressive neural painting method. Background Technology

[0002] In recent years, neural painting techniques based on deep generative models have made significant progress. These techniques aim to generate artistically expressive digital images by optimizing a set of renderable brush stroke parameters. Depending on the form of the brush stroke parameters, they can be divided into vector curve-based neural painting and scalar brush stroke-based neural painting. Among them, scalar brush strokes have attracted widespread attention due to their simple parameter form, brush stroke elements that closely approximate realistic human brushstrokes, and flexible and variable types.

[0003] Existing scalar stroke neural painting methods typically require a target image as input, and then iteratively optimize stroke parameters to reconstruct that target image. To improve generation efficiency, some methods employ feedforward network structures to directly predict stroke parameters from the target image; to enhance generation quality, some methods introduce reinforcement learning mechanisms, focusing stroke parameter optimization on local regions of the current canvas. However, these methods are highly dependent on the target image input, limiting their application in scenarios where only text descriptions are provided.

[0004] For scenarios with only text input, existing technologies attempt to introduce text-guided mechanisms, utilizing pre-trained contrastive language-image pre-trained models or diffusion models to provide semantic constraints for brushstroke parameter optimization. However, due to the limited expressive power and high optimization difficulty of scalar brushstroke parameters, existing methods struggle to generate semantically accurate and visually high-quality scalar brushstroke images when only text input is received. Furthermore, existing methods generally employ a single-stage end-to-end optimization strategy, lacking progressive planning for structure, color, and details during the painting process, resulting in problems such as unreasonable spatial layout and blurred local details in the generated results.

[0005] In summary, the main problem with existing technologies is that existing scalar stroke neural painting methods, when only receiving text input, struggle to generate scalar stroke images that are semantically consistent, spatially well-arranged, and rich in local details. Summary of the Invention

[0006] To address the aforementioned problems in the existing technology, this invention provides a text-driven progressive neural painting method. The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a text-driven progressive neural painting method, comprising: S1: Perform semantic parsing, spatial guidance, and semantic alignment on the acquired text information to obtain multi-level text guidance conditions; S2: Adopting a three-stage progressive optimization strategy of drafting-coloring-refinement, combined with the multi-level text guidance conditions, the brush stroke parameters are optimized in sequence by determining the structure, laying the color and enhancing the details, to obtain the optimized foreground brush stroke parameters; S3: Obtain the optimized background stroke parameters, and render the optimized background stroke parameters and the optimized foreground stroke parameters onto the same canvas through the renderer in sequence, outputting a final image that is semantically consistent with the text information and has enhanced details.

[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: To address the problem that existing scalar stroke neural painting methods struggle to generate semantically consistent, spatially reasonable, and locally detailed scalar stroke images when only receiving text input, this invention provides a text-driven progressive neural painting method. It constructs multi-level text guidance conditions through S1, and utilizes a pre-trained large language model, a text-to-image diffusion model, and a contrastive language-image pre-trained model to perform semantic parsing, spatial guidance, and semantic alignment processing on the input text. This transforms natural language into semantic objects, colors, hierarchical structures, attention maps, and various loss calculation bases, providing multi-dimensional support for subsequent stroke optimization and solving the problem that text information cannot effectively drive stroke parameters. This invention addresses the challenge of scalar stroke optimization. Through S2, a three-stage progressive optimization strategy of drafting, coloring, and refinement is employed. This strategy decouples structure determination, color application, and detail enhancement into independent iterative optimizations at different stages. Furthermore, the transfer of control conditions between stages enables synergistic enhancement of structure, color, and detail, avoiding interference between different objectives in single-stage optimization. This effectively improves the spatial layout stability, color semantic accuracy, and local detail richness of the generated image. S3 renders the optimized stroke parameters into a scalar stroke image, and combined with independent optimization and smoothing of the background image, further enhances the visual salience of the foreground subject. The final output is a scalar stroke image that is semantically consistent with the input text, structurally sound, and rich in detail. This invention, through the organic integration of multi-level text guidance and three-stage progressive optimization, significantly improves the semantic consistency and visual quality of scalar stroke neural painting under text-only input conditions, filling the technological gap in text-driven scalar stroke generation. Attached Figure Description

[0008] Figure 1 This is a flowchart illustrating the text-driven progressive neural painting method provided in an embodiment of the present invention. Figure 2 This is an example diagram of the three-stage progressive optimization process of drafting, coloring, and refinement provided in the embodiments of the present invention; Figure 3 This is a schematic diagram of multi-object saliency extraction provided in an embodiment of the present invention; Figure 4This is an example diagram of the data processing flow during the drafting stage provided in an embodiment of the present invention; Figure 5 This is an example diagram of the data processing flow during the coloring stage provided in an embodiment of the present invention; Figure 6 This is an example diagram of the data processing flow during the refinement stage provided in an embodiment of the present invention. Detailed Implementation

[0009] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0010] The present invention will now be described in detail with reference to the accompanying drawings, a text-driven progressive neural painting method proposed in this invention.

[0011] Figure 1 This is a flowchart illustrating the text-driven progressive neural painting method provided in an embodiment of the present invention. Figure 1 As shown, the method includes S1-S3, specifically: To effectively guide the optimization of pen stroke parameters through text information, this invention first constructs multi-level text guidance conditions in S1, which serve as the input and constraint basis for subsequent optimization processes.

[0012] S1: Perform semantic parsing, spatial guidance, and semantic alignment on the acquired text information to obtain multi-level text guidance conditions.

[0013] Specifically, S1 includes: acquiring a pre-trained large language model, a pre-trained text-to-image diffusion model, and a pre-trained contrastive language-image pre-trained model; using the pre-trained large language model to perform semantic parsing on the text information, obtaining a set of semantic objects, the RGB color parameters corresponding to each semantic object, the hierarchical order between semantic objects, and background description information; using the pre-trained text-to-image diffusion model to perform spatial guidance processing on the text information, obtaining a global self-attention map, a cross-attention map corresponding to each semantic object, and the conditional basis for calculating fractional distillation sampling loss or interval fractional matching loss; using the pre-trained contrastive language-image pre-trained model to perform semantic alignment processing on the text information, establishing the computational basis for language-image alignment loss, so as to measure the semantic consistency between the rendered image and the input text in the subsequent optimization process; and summarizing the results obtained by the pre-trained large language model, the pre-trained text-to-image diffusion model, and the pre-trained contrastive language-image pre-trained model to form a multi-level text guidance condition.

[0014] In other words, the multi-level text guidance conditions include at least: a set of semantic objects, the RGB color parameters corresponding to each semantic object, the hierarchical order between semantic objects, background description information, a global self-attention map, a cross-attention map corresponding to each semantic object, the conditional basis for calculating fractional distillation sampling loss or interval fractional matching loss, and the computational basis for language-image alignment loss.

[0015] It should be noted that the pre-trained large language model, the pre-trained text-to-image diffusion model, and the pre-trained contrastive language-image pre-trained model are all publicly available models that have been pre-trained on large-scale data. This invention directly calls their inference interfaces or loads their weight parameters, without requiring additional training for this task. Furthermore, the pre-trained text-to-image diffusion model is not only used for spatial guidance processing in S1, but also for calculating fractional distillation sampling loss or interval fractional matching loss in the drafting, coloring, refinement, and background optimization stages. Similarly, the pre-trained contrastive language-image pre-trained model is not limited to S1, but is also repeatedly called in each optimization stage to calculate the language-image alignment loss, ensuring that the rendered image in each stage (drafting, coloring, and refinement) maintains semantic consistency with the input text.

[0016] It's important to clarify that the conditional basis used to calculate the fractional distillation sampling loss or interval fractional matching loss refers to the inference capability of the pre-trained text-to-image diffusion model (and its optional control module)—that is, the ability to receive noisy images, time steps, text conditions (and structural control conditions), and output predicted noise components or perform DDIM inversion / multi-step sampling. This is a function mapping capability, not a set of fixed values. Similarly, the computational basis for the language-image alignment loss obtained using the pre-trained contrastive language-image pre-trained model refers to the feature extraction and similarity calculation capability of the CLIP model—that is, the ability to receive images and text, output feature vectors respectively, and calculate cosine similarity. Again, this is a function, not specific numerical values. During optimization, these "computational bases" are repeatedly invoked, with different current images and texts input each time, resulting in different loss values / gradients output. Therefore, they are essentially reusable model inference interfaces, rather than pre-stored static numerical values.

[0017] After obtaining multi-level text guidance conditions in S1, optimization processing is performed in S2. This step is the core of this invention, and its design concept is to simulate the cognitive and execution process of human painting: first, determine the picture structure, then lay out semantic colors, and finally add detailed brushstrokes. By decoupling the optimization objectives of structure, color, and detail to different stages and passing control conditions between stages, this invention effectively solves the problem that existing scalar brushstroke neural painting methods struggle to simultaneously achieve spatial layout rationality, color semantic accuracy, and local detail richness when only text is input. Specifically: S2: Adopting a three-stage progressive optimization strategy of drafting-coloring-refinement, combined with multi-level text guidance conditions, the brush stroke parameters are optimized in sequence through structure determination (drafting stage), color application (coloring stage), and detail enhancement (refinement stage) to obtain the optimized foreground brush stroke parameters.

[0018] Figure 2 This is an example diagram illustrating the three-stage progressive optimization process of drafting, coloring, and refinement provided in an embodiment of the present invention. Figure 2 As shown, S2 includes: S2.1: In the drafting stage, the parameters of the first stroke are iteratively optimized multiple times to generate a sketch image and complete the structural determination; S2.2: In the coloring stage, the second stroke parameters are iteratively optimized multiple times, and in each iteration, the structural control conditions are input into the pre-trained text with control module to the image diffusion model to guide the generation of a coloring image with semantic color distribution and complete the color laying. S2.3: In the refinement stage, a detail-enhanced version is generated based on the colored image to initialize the third stroke parameter. Then, the initialized third stroke parameter is iteratively optimized multiple times to obtain the optimized foreground stroke parameter, thus completing the detail enhancement. Here, the stroke parameter refers to the first stroke parameter in the drafting stage, the second stroke parameter in the coloring stage, and the third stroke parameter in the refinement stage.

[0019] Here, the brush stroke parameters for different types or stages all include position parameters, shape parameters, color parameters, and transparency parameters. That is, the brush stroke parameters for different types (the first brush stroke parameters in the drafting stage, the second brush stroke parameters in the coloring stage, and the third brush stroke parameters in the refinement stage) as well as the background brush stroke parameters all include position parameters, shape parameters, color parameters, and transparency parameters.

[0020] Before introducing the specific working principle of S2, it is necessary to explain that, Figure 2 In this context, MOS stands for Multi-Object Saliency. Its purpose is to generate structured stroke sampling probability maps for different semantic objects based on the input text, thereby achieving object-aware stroke initialization.

[0021] The following is an example of inputting the text message "a gray cat sitting on a table" (hereinafter referred to as symbol). Taking (referencing) as an example, combined with Figure 3 The process of extracting saliency from multiple objects is explained. Figure 3 This is a schematic diagram of multi-object saliency extraction provided in an embodiment of the present invention.

[0022] like Figure 3 As shown, multi-object saliency extraction first utilizes a pre-trained large language model (this pre-trained large language model can also use symbolic representation). (Indicates) the input text Perform semantic parsing to extract a set of semantic objects. In the formula, Refers to the first The name of a semantic object, It is the total number of semantic objects extracted. And infer their hierarchical order and representative colors; In the formula, yes RGB color parameters. Figure 3 The model identifies two semantic objects: "cat" and "table," and infers their hierarchical order—the cat is located above the table, therefore the cat's mask has higher priority than the table's in subsequent processing. Simultaneously, the large language model also predicts a representative color for each object; for example, the cat is predicted as gray, and the table as brown.

[0023] Meanwhile, the system utilizes a pre-trained text-to-image diffusion model to extract two types of attention maps during text-guided denoising: cross-attention maps and self-attention maps. For the objects "cat" and "table," the diffusion model generates corresponding cross-attention maps, where high-response regions indicate the possible positions of the cat and table in the image, respectively. Simultaneously, a global self-attention map is extracted to capture the overall structural relationships of the image (such as the relative positions and spatial proportions of the cat and table). Next, the system applies Otsu's automatic thresholding method to the cross-attention maps of each object, automatically calculating the threshold for segmenting the foreground and background to obtain a binary mask for each semantic object. , It is the first Binary masks for semantic objects. For example, the cross-attention map of a cat generates an initial mask for the cat after thresholding, and the cross-attention map of a table generates an initial mask for the table. Since the cat and the table may overlap spatially (the cat sits on the table), their initial masks will conflict in the overlapping area. At this time, the system uses the hierarchical order inferred from the pre-trained large language model (the cat is on the upper level) to deduplicate the masks, and prioritizes the allocation of pixels in the overlapping area to the cat, thus obtaining unique masks (i.e., unambiguous masks) for both the cat and the table.

[0024] like Figure 3 As shown, the system also introduces an "optional" product operation. Specifically, after obtaining the unique mask of each object, the system multiplies the global self-attention map element-wise with the unique mask of each object to generate a structured sampling probability map within each object's region. This product operation ensures that the sampling probability is not only constrained by the object mask but also incorporates the overall compositional information contained in the global self-attention map. This allows the brushstroke layout to maintain the semantic regions of the objects while better conforming to the overall structural relationships of the image. For example, within the cat's mask region, areas with higher responses in the global self-attention map (such as the cat's head and body outline) will receive higher sampling probabilities, guiding brushstrokes to be placed preferentially in key structural areas. Finally, the system merges the structured sampling probability maps of each object to form a multi-object structured sampling probability map. This probability map is used for brushstroke position initialization in the subsequent drafting stage, ensuring that the brushstrokes in the first brushstroke parameter can be reasonably laid out according to the spatial distribution, hierarchical relationship, and global structure of the semantic objects, providing a good starting point for subsequent structural optimization.

[0025] Here, we continue with the explanation of each sub-step in S2. The drafting, coloring, and refinement stages in S2 are all independent iterations, and all three stages run within the same optimization framework, sharing the same differentiable renderer, optimizer, and pre-trained model resources. The brush stroke parameter instances in the three stages are independent of each other, and each completes its initialization and multiple iterations of optimization independently within its respective stage. Through this "same framework, independent instances" design, this invention achieves progressive enhancement of structure, color, and detail while ensuring decoupling between stages.

[0026] For brevity, this invention employs a gradient-based iterative optimization algorithm. In each iteration, the image is first rendered based on the current brush stroke parameters. Then, the gradients corresponding to each loss term are calculated. Finally, the brush stroke parameters are updated via an optimizer. For terms with closed-form loss values, the loss function is directly defined; for terms where only gradient approximation can be obtained, gradient update rules are provided.

[0027] S2.1 (drafting stage) includes: S2.1.1: Utilize the global self-attention map and the cross-attention map corresponding to each semantic object in the multi-level text guidance conditions to generate a multi-object structured sampling probability map.

[0028] Here, S2.1.1 specifically includes: using the cross-attention map corresponding to each semantic object, extracting the binary mask of each semantic object through the Otsu automatic thresholding method; using the hierarchical order between semantic objects to remove duplicates from the binary masks of each semantic object, obtaining the unambiguous mask of each semantic object; performing masking operations between the global self-attention map and the unambiguous mask of each semantic object respectively, and outputting the sampling probability distribution within the region of each semantic object; merging the sampling probability distributions within the region of each semantic object to generate a multi-object structured sampling probability map.

[0029] It should be noted that this multi-object structured sampling probability map specifically employs... Figure 3 The multi-object saliency extraction (MOS) method is shown in the diagram. For the sake of brevity, it will not be described in detail here.

[0030] S2.1.2: Sample multiple initial positions within each semantic object region based on the multi-object structured sampling probability map, and summarize them as the first stroke parameter; wherein, the color parameter in the first stroke parameter is initialized to white, and the canvas is set to black.

[0031] Here, each semantic object region refers to the unique mask range corresponding to each semantic object after deduplication (such as the cat region or the table region). Within each semantic object region, based on probability distribution, several coordinate points are extracted from each region and aggregated as the positional attribute of each stroke in the first stroke parameter.

[0032] For example, 300 position coordinate points are sampled in the cat's area (mainly in high-probability areas such as the head and back), and 200 position coordinate points are evenly sampled in the table's area (the entire desktop area). These 500 position coordinate points are used as the position attribute array of the first stroke parameter.

[0033] It should be noted that the color parameter of the first stroke is not updated during the drafting stage; it remains white (i.e., the RGB value of all strokes is set to (255, 255, 255)). The drafting stage focuses on structure formation, and fixing the color to white avoids color information interfering with structure optimization. At the same time, white strokes on a black background make it easier to extract structural control conditions in subsequent stages (such as morphological processing, edge detection, etc.).

[0034] Here, the shape and transparency parameters of the first stroke parameter can be set to an initial value according to actual needs. For example, the initial size of all stroke parameters can be set to 8 pixels; or dynamically allocated according to the stroke position (larger strokes in the center area and smaller strokes at the edges), and the initial preset transparency of all stroke parameters can be opaque.

[0035] S2.1.3: Perform multiple iterations to optimize the optimizable parameters in the first stroke parameters other than the color parameter. Each iteration includes: rendering the current first stroke parameters into the current first image using a differentiable renderer; calculating the current first fractional distillation sampling loss and the current first language-image alignment loss based on the current first image; and updating the optimizable parameters in the current first stroke parameters other than the color parameter based on the current first fractional distillation sampling loss and the current first language-image alignment loss.

[0036] Figure 4 This is an example diagram of the data processing flow during the drafting stage provided in an embodiment of the present invention. Figure 4 Yes Figure 3 Text information in Further processing is required. For example... Figure 4 As shown, in each iteration, let the parameter of the current first stroke be... Differentiable renderer Render as the current first image Noise sampled from Gaussian distribution In the formula It is an identity matrix, used to add noise to the current first image through a diffusion process. The noise-added current first image is represented as follows. ,in, It is a time step The cumulative product of the corresponding noise intensity scheduling parameters, , This refers to the time step. From 0 A uniform distribution of positive integers. It is preset Maximum value in the range of values From standard Gaussian distribution Random noise sampled in the middle is used to add forward diffusion noise to the currently rendered image.

[0037] The current first image after adding noise With the addition of a specific prefix (e.g., "a sketch of" + After) Inputting pre-trained text into an image diffusion model to predict current noise components, this pre-trained text-to-image diffusion model can use symbolic representations. Thus, the current noise component is represented as For computational efficiency reasons, the pre-trained text-to-image diffusion model is ignored. The Jacobian matrix of U-Net is used to calculate the image gradient to measure the difference between diffuse noise and sampling noise, and this is used as the fractional distillation sampling loss (SDS).

[0038] Here, the gradient calculation expression for the current first fractional distillation sampling loss in each iteration is: ; In the formula, It is about finding the expected value function. It is dependent on The weight function, It is a partial derivative operation. This indicates the parameters of the current first stroke. Find the gradient. This is the current first fractional distillation sampling loss. Refers to the current noise component With real noise The difference between them.

[0039] In one possible implementation, the calculation expression for the current first fractional distillation sampling loss in each iteration is rewritten as the current first image. With pre-trained text-to-image diffusion models Predicted results The specific formula for the gradient corresponding to the current first fractional distillation sampling loss in each iteration, based on the deviation between them, is: ; In the formula, It is a time step The weight function, , , It is a pre-trained text-to-image diffusion model Based on the current first image after noise addition and Denoising estimated image based on single-step prediction.

[0040] To ensure consistency between the rendered image and the input text in the semantic space, a pre-trained contrastive language-image pre-trained model (CLIP) is used to simultaneously compute the language-image alignment loss.

[0041] By using the current first image Image encoder input to CLIP model Extracting image feature vectors and extracting text information Text encoder for input CLIP model The text feature vector is extracted; the negative value of the least cosine similarity between the image features and the text features is used as the current first language-image alignment loss. Its calculation expression is: ; In the formula, This refers to the operation of taking the magnitude of a vector.

[0042] The drafting total loss is established based on the current first fractional distillation sampling loss and the current first language-image alignment loss, and its expression is: ; In the formula, This refers to the first stroke parameter to be optimized. This refers to the current first fractional distillation sampling loss calculated under certain constraints. This refers to the current first language-image alignment loss. This refers to the current first image. and All are preset weights.

[0043] S2.1.4: After reaching the first preset number of iterations, the optimized first stroke parameters are obtained.

[0044] S2.1.5: Render the optimized first stroke parameters as a sketch image. For example... Figure 4 As shown, the sketch image is represented as .

[0045] After completing the iterative optimization of the drafting phase, we proceeded to the optimization of the coloring phase.

[0046] S2.2 (coloring stage) includes: S2.2.1: For the sketch image Corrosion and expansion operations were performed sequentially to obtain the structural control conditions. .

[0047] Here, the purpose of operation S2.2.1 is to transform the sketch image generated during the drafting stage into structured guidance information suitable for input control modules. The erosion operation can eliminate the sketch image. Small, redundant strokes are removed, isolated noise is eliminated, and the dilation operation improves the connectivity of strokes, connecting broken structural areas into a cohesive whole.

[0048] S2.2.2: Using structural control conditions, a pre-trained text-to-image diffusion model with a control module is used to re-acquire the cross-attention maps corresponding to each semantic object. Multiple initial positions are uniformly collected within the re-acquired cross-attention maps corresponding to each semantic object and summarized as the second stroke parameters.

[0049] For example, to input text For example, in S1, a pre-trained diffusion model without a control module has been used to generate cross-attention maps corresponding to the semantic objects "cat" and "table," which are then stored as multi-level text guidance conditions. In the coloring stage, the structural control conditions of the sketch image after erosion and dilation are first obtained. (i.e., the outline and connected regions of the sketch); subsequently, a pre-trained text-to-image diffusion model with control modules is used to control the structural conditions. As additional input, 200 initial positions are uniformly sampled within the "cat" region indicated by the cross-attention map, and another 200 initial positions are uniformly sampled within the "table" region, resulting in 400 stroke positions, which serve as the positional attributes of the second stroke parameter. During this process, the control module ensures that the sampled positions not only lie within the semantic object region but also preferentially cover the contours and connected regions emphasized by the sketch structure, providing structural priors for subsequent color tiling.

[0050] S2.2.3: Initialize the color parameters in the second stroke parameters to the RGB color parameters corresponding to each semantic object in the multi-level text guidance conditions.

[0051] S2.2.4: Perform multiple iterative optimizations on the second stroke parameters, where each iteration includes: rendering the current second stroke parameters into the current color image using a differentiable renderer; based on the current color image, structural control conditions, and text information, utilizing a pre-trained text-to-image diffusion model with a control module (this pre-trained text-to-image diffusion model with a control module can use symbolic representations). (This indicates that) the current first interval score matching loss is calculated, and the current second language-image alignment loss is calculated using the contrastive language-image pre-trained model. Based on the current first interval score matching loss, the current second language-image alignment loss, and the current color constraint, the current second stroke parameter is updated. The current color constraint is calculated in real time during each iteration of optimization based on the difference between the current stroke color parameter and the initial color parameter of the second stroke parameter (i.e., the RGB color parameters corresponding to each semantic object in the multi-level text guidance conditions).

[0052] Figure 5 This is an example diagram of the data processing flow during the coloring stage provided in an embodiment of the present invention. Figure 5 Yes Figure 4 sketch image in Further processing is carried out based on this. For example... Figure 5 As shown, in each iteration, a color constraint term is designed for each semantic object, and its calculation expression is: ; In the formula, It represents the total number of pixels within the cross-attention map corresponding to each semantic object. It is the first Within the cross-attention graph corresponding to the semantic object, the first... The hue of each pixel It is the first Within the cross-attention graph corresponding to the semantic object, the first... The saturation value of each pixel. It is the first The initial RGB hue of the semantic object, i.e., the first one in the multi-level text guidance condition. The RGB hue of a semantic object, It is the first The initial RGB saturation of a semantic object, and All are preset weights. Represents the cyclic distance L1 on the unit circle. , .

[0053] Here, the current first interval score matching loss (Controlled Interval Score Matching, CISM) is an extension of interval score matching (ISM) guided by the control module. Its core idea is to use denoising diffusion implicit models (DDIM) inversion and multi-step sampling to make the current color image structurally conform to the structural control conditions. Alignment while maintaining semantic guidance of the text.

[0054] Set the current color image as The second stroke parameter is In each iteration, the current color image is first inverted using DDIM. Convert to noise latent code This process is guided solely by textual information and does not involve structural control conditions. This is done to faithfully preserve the structural information of the current color image. Subsequently, the noise latent code obtained from the inversion... Starting from this point, and considering the constraints, a pre-trained text-to-image diffusion model with a control module is used for multi-step denoising sampling to generate pseudo-ground values. Here, during the multi-step denoising sampling, the structural control conditions are obtained simultaneously using S2.2.3. and text information , The prefix "aDSLR photo of" + The structure guides the noise reduction direction through a control module. It should be noted that this uses... Or, in other words, change the text information. The prefix is ​​to guide the diffusion model from structural optimization to color texture optimization.

[0055] Here, noise latent code Represented as: ; ; In the formula, From time step The denoised estimated image predicted from the noise latent code is calculated using a single-step prediction formula. It is an empty text condition. It is to use noise latent code and empty text conditions Input pre-trained text to image diffusion model The noise components predicted in the middle, It is a time step The corresponding noise latent code, , It is a time step The weight function, It is to use noise latent code and empty text conditions Input pre-trained text to image diffusion model The noise components predicted in the middle, It is the first The noise latent code corresponding to the time step, It is to display the current color image and empty text conditions Input pre-trained text to image diffusion model The noise components predicted in the middle, It is a time step The corresponding noise latent code represents an earlier latent code. It is the noise latent code and empty text conditions Input pre-trained text to image diffusion model The noise components predicted in the middle, It is a time step The weight function, .

[0056] The control module provides structural constraints during the multi-step sampling process, guiding the alignment of the sampled structure with the sketch. The control module guides the sampling process from the noise latent code... Initially, pseudo-true values ​​are generated gradually through multiple denoising steps, and in each denoising step, the control module receives structural control conditions. This is encoded into a U-Net feature injection diffusion model, allowing noise prediction to be influenced by textual information simultaneously. and structural control conditions The constraints. Specifically, the noise latent code on the denoising path obtained through multi-step denoising using a pre-trained text-to-image diffusion model with a control module can be represented as: ; In the formula, It is a noise latent code Starting with the image, a denoised estimated image is obtained after multiple denoising steps. During the process, after a single step time interval The intermediate state is used as the starting point for subsequent multi-step sampling. It is a time step The cumulative product of the corresponding noise intensity scheduling parameters, From time step The denoised estimated image obtained by single-step prediction of the corresponding noisy latent code. It is a time step The weight function, It is to use noise latent codes Text information and structural control conditions Input pre-trained text with control module into image diffusion model The noise component predicted in the data.

[0057] Following the model of ISM, the denoised estimated image obtained through multiple denoising steps using a pre-trained text-to-image diffusion model with a control module can be uniformly represented as: ; In the formula, It is the time step obtained after inverting the current color image using DDIM. intermediate state Text information and structural control conditions Input pre-trained text with control module into image diffusion model The noise component predicted in the middle; It is the time step obtained by inverting the current color image using DDIM. intermediate state Text information and structural control conditions Input pre-trained text with control module into image diffusion model The noise component predicted in the middle; It is the time step obtained by inverting the current color image using DDIM. intermediate state Text information and structural control conditions Input pre-trained text with control module into image diffusion model The noise component predicted in the data.

[0058] Based on this, the gradient calculation expression for the current first interval score matching loss is: ; In the formula, It is to seek the desired operation.

[0059] It should be noted that the current color image is shown here. and text information The image encoders input to the CLIP model respectively and text encoder Calculate the current second language-image alignment loss.

[0060] Based on this, we construct the current first interval score matching loss, the current second language-image alignment loss, and the total loss for the coloring stage based on the current color constraint term. The expression for calculating the total loss of the coloring stage at each iteration is: ; In the formula, , and All are preset weights. This refers to the current first interval score matching loss. This refers to the current second language-image alignment loss. This refers to the current color constraint.

[0061] S2.2.5: After reaching the second preset number of iterations, the optimized second stroke parameters are obtained.

[0062] S2.2.6: Render the optimized second stroke parameters into a colored image. .

[0063] After completing the iterative optimization of the coloring stage, continue with the optimization of the refinement stage.

[0064] S2.3 (refinement stage) includes: S2.3.1: Obtain a restricted error map based on the detail-enhanced version and the colored image, and sample multiple initial positions according to the restricted error map as the third stroke parameter. The color parameter in the third stroke parameter is initialized to the color value of the corresponding position in the detail-enhanced version.

[0065] Here, S2.3.1 specifically includes: using a first classifier-free guided scale to process the colorized image. Perform denoised diffusion probability model inversion to obtain the inversion latent code. A second classifier-free guided scale is used to retrieve the latent code. Resampling is performed to generate a version with enhanced details. The second classifier-free guidance scale is larger than the first classifier-free guidance scale; the colorized image is calculated. With enhanced details The pixel-level absolute error values ​​between them are used to obtain the error map. Using a semantic segmentation model, taking a set of semantic objects as input, from the colorized image... Extract the mask of each semantic object to transform the error map. The error map is obtained by limiting it to the mask formed by the masks of each semantic object; the error map is normalized to a sampling probability distribution; multiple initial positions are sampled from the sampling probability distribution and summarized as the third stroke parameter.

[0066] Error maps can reflect the spatial distribution of detail enhancement. Areas with larger errors indicate more significant detail enhancement and are key areas where detailed brushstrokes need to be added.

[0067] S2.3.2: Perform multiple iterations to optimize the third stroke parameters. Each iteration includes: rendering the current third stroke parameters into the current detail map using a differentiable renderer; calculating the current second interval score matching loss and the current third language-image alignment loss based on the current detail map and some conditions in the multi-level text guidance conditions; and updating the current third stroke parameters according to the current second interval score matching loss and the current third language-image alignment loss.

[0068] Figure 6 This is an example diagram of the data processing flow during the refinement stage provided in an embodiment of the present invention. Figure 6 Yes Figure 5 Colored images in Further processing is carried out based on this. For example... Figure 6 As shown, the current detail image is set to... The third stroke parameter is .

[0069] In each iteration, the input text is text information. The current second interval fractional matching loss is calculated using the same logic as the interval fractional matching loss in the coloring stage, but without the structural control conditions (i.e., without being guided by the control module). The standard interval fractional matching loss (ISM) is applied directly. Similarly, the current detail map is used... The process of inverting DDIM into a noisy latent code relies solely on textual information. To guide this process, we then start with the noise latent code and utilize a pre-trained text-to-image diffusion model. Multi-step denoising sampling is performed to generate pseudo-true values; structural control conditions are not used in this process.

[0070] And in the current detailed image and text information The image encoders input to the CLIP model respectively and text encoder Calculate the current third language-image alignment loss.

[0071] Unlike the coloring stage, the refinement stage no longer uses color constraints. This is because the colors are largely determined in the coloring stage, and this stage focuses on detail enhancement, requiring no additional constraints on the colors. Therefore, the total loss in the refinement stage is simply a weighted average of the interval score matching loss and the language-image alignment loss, specifically calculated as follows: ; In the formula, and All are preset weights. It is the current second interval score matching loss. It is the current third language-image alignment loss.

[0072] S2.3.3: After reaching the third preset number of iterations, the optimized third stroke parameters are obtained, and together with the optimized second stroke parameters, they constitute the optimized foreground stroke parameters.

[0073] S3: The optimized background stroke parameters and the optimized foreground stroke parameters are rendered sequentially onto the same canvas using the renderer, outputting a final image that is semantically consistent with the text information and has enhanced details.

[0074] Here, after completing the three-stage progressive optimization of the foreground—drafting, coloring, and refining—this invention further optimizes the background image independently to enhance the visual prominence of the foreground subject. The generation of the background image is executed in parallel or sequentially with the three foreground stages. Its optimization goal is to simulate a blurred background effect through a smooth transition while maintaining consistency with the text semantics, thereby highlighting the foreground content through visual contrast. The following first describes the process of acquiring the background image, and then elaborates on the composite output of the foreground and background.

[0075] This invention utilizes the background description information inferred from the large language model in S1 to independently construct background stroke parameters and perform multiple iterative optimizations. Specifically, the optimized background image is obtained through the following methods: generating background stroke parameters using the background description information; performing multiple iterative optimizations on the background stroke parameters, wherein each iteration includes: rendering the current background stroke parameters into the current background image using a differentiable renderer; calculating the current second fractional distillation sampling loss, the current fourth language-image alignment loss, and the total variation loss based on the current background image; updating the current background stroke parameters based on the current second fractional distillation sampling loss, the current fourth language-image alignment loss, and the total variation loss; obtaining the optimized background stroke parameters after reaching a fourth preset number of iterations; and rendering the optimized background image based on the optimized background stroke parameters.

[0076] Here, the total background loss, calculated based on the current second fractional distillation sampling loss, the current fourth language-image alignment loss, and the total variation loss, is expressed as follows: ; In the formula, These are the background stroke parameters that need optimization. , and All are preset weights. This is the current second fractional distillation sampling loss. It is the current fourth language-image alignment loss. It is the total variation loss. This is the current background image.

[0077] After multiple iterations, optimized background stroke parameters are obtained. After obtaining the optimized foreground and background stroke parameters, the optimized background stroke parameters are first rendered onto the canvas using a differentiable renderer, and then the optimized foreground stroke parameters are rendered onto the same canvas using the same differentiable renderer, thus compositing the final output image.

[0078] To address the problem that existing scalar stroke neural painting methods struggle to generate semantically consistent, spatially reasonable, and locally detailed scalar stroke images when only receiving text input, this invention provides a text-driven progressive neural painting method. Through a pre-trained large language model, a text-to-image diffusion model, and a contrastive language-image pre-trained model, semantic parsing, spatial guidance, and semantic alignment of the input text are performed, constructing multi-level text guidance conditions to provide semantic objects, colors, hierarchical structures, attention maps, and various loss calculation foundations for stroke optimization. A three-stage progressive optimization strategy of drafting, coloring, and refinement is adopted, sequentially optimizing stroke parameters by determining the structure, applying color, and enhancing details. In the drafting stage, object-aware stroke initialization and fractional distillation sampling combined with language-image alignment loss generate a structurally clear sketch. In the coloring stage, under the control of the sketch structure, precise alignment of structure and color is achieved through interval fractional matching loss, language-image alignment loss, and color constraints guided by a control module. In the refinement stage, detailed stroke initialization is guided by an error map, and local details are enhanced using interval fractional matching loss and language-image alignment loss. The three-stage independent iteration, with layer-by-layer decoupling, achieves progressive collaborative enhancement through image-level control conditions. The background image is independently optimized, and a total variation loss is introduced to achieve smooth blurring, highlighting the foreground subject. The final output is a scalar brushstroke image that is semantically consistent with the input text, has a stable spatial layout, and rich local details. This invention, through the fusion of multi-level text guidance and three-stage progressive optimization, effectively improves the semantic consistency and visual quality of scalar brushstroke neural painting under text-only input conditions.

[0079] First, simulation experiments are used to verify the performance of the text-driven progressive neural painting method provided in this embodiment of the invention.

[0080] The experiments used the SGD optimizer and were all performed on a single 24GB NVIDIA GeForce RTX 3090 graphics processing unit. The hardware used in the application can be adjusted according to the requirements of computational efficiency and computational load. All hyperparameters in the experiments were manually adjusted based on empirical observations of performance on five different text inputs to achieve a balanced contribution of each loss component; to mitigate the impact of randomness, each input was run at least three times. The specific weights used in the application can be adjusted according to the type of input image or the desired stylization effect. The method of this invention was compared and evaluated with four representative baseline methods: Stylized Neural Painting (SNP), Stable Diffusion (SD), CLIPDraw, and DiffSketcher. These baselines cover key technological nodes in the fields of neural painting and text-to-image generation: SNP pioneered a neural painting method based on scalar strokes; SD represents the current state-of-the-art text-to-image generation technology based on diffusion models; CLIPDraw was the first to introduce semantic guidance from CLIP into text-driven vector stroke generation tasks; and DiffSketcher was the first to integrate fractional distillation sampling into text-driven vector stroke generation. This invention selects CLIPDraw and DiffSketcher as comparison objects because they are pioneers of the technology used in this invention; it does not compare with other text-driven vector stroke generation methods because this invention aims to fill the gap in text-driven scalar stroke generation methods, and introducing vector methods would confuse the task boundaries and affect the objective evaluation of this invention.

[0081] To ensure a fair comparison under a uniform text-driven stylized image generation setting, the experiment configured the baselines as follows: (1) For SD, add the style suffix "in the style of an oilpainting, thick impasto brushstrokes, vibrant colors, painterly style" after each content prompt to generate stylized output; (2) For SNP, a reference image generated by SD using the original cue words (without style enhancement) is provided as input; (3) For DiffSketcher, a loss configuration with fractional distillation sampling loss is adopted, and the manual word segmentation markers in its stroke initialization are replaced with a large language model, which is consistent with the method in the multi-object saliency extraction of this invention.

[0082] All experiments used 512×512 resolution images generated from 150 random text prompts. All methods were evaluated on the same set of text prompts to ensure consistency. For fair comparison, all methods used a total of no more than 500 strokes. The three-stage optimization of this invention employed 500, 300, and 300 iterations respectively; DiffSketcher underwent 1500 iterations, ensuring its total runtime exceeded that of this invention to guarantee fairness in comparison; SNP and CLIPDraw did not use fractional distillation sampling, resulting in significantly shorter iteration times, so their official settings were followed. All experiments followed the above configuration, and the average performance of each method on all samples in the test set was recorded for quantitative evaluation.

[0083] The experiment used the following indicators for quantitative evaluation: (1) CLIP Score: measures the semantic consistency between the text prompt and the generated image; (2) Fréchet Inception Distance (FID): Evaluates the distributional distance between the generated image and the “Impressionist” subset in WikiArt to assess the stylization quality; (3) Human Preference Score (HPS) v2.1: Reflects the quality of synthesized images that meet textual expectations under human preferences; (4) Aesthetic Score v2.5: Evaluates the aesthetic quality of an image, with a score range of 1 to 10.

[0084] The experiment used CLIP Score to evaluate both the foreground and full results of the invention, but FID, HPS, and Aesthetic Score were calculated for the full results, as these metrics are intended to evaluate the generation quality of the full image.

[0085] In addition, the experiment conducted a user study, inviting 34 participants with diverse backgrounds in art and computer vision to be presented with images generated by different methods for 20 text inputs randomly selected from the quantitative assessment. Participants were asked to evaluate the images from the following two dimensions: (1) Semantic consistency: Which image best matches the given text prompt? (2) Artistic quality: Which image is closest to a human-painted oil painting with rich brushstroke texture?

[0086] Each participant evaluated all five methods under all 20 prompts to ensure consistency and comprehensiveness of the comparison.

[0087] The quantitative experimental results are shown in Tables 1 and 2, which present comparative data on semantic consistency, stylization quality, and user preference for each method. Table 1 shows the quantitative comparison of semantic consistency and stylization quality among different methods; Table 2 shows the user research preference results for different methods. As shown in Table 1, this invention achieves a CLIP Score of 0.2974, close to Stable Diffusion and Stylized Neural Painting, indicating good consistency between its generated results and text semantics. In the user preference-related HPS index, this invention achieves 0.189, superior to DiffSketcher and CLIPDraw; and in the Aesthetic Score, it scores 4.00, higher than Stylized Neural Painting, DiffSketcher, and CLIPDraw. Table 2 shows the user research preference results for different methods. As shown in Table 2, this invention receives 40.15% of the votes for semantic consistency preference and 49.63% for artistic quality preference, both significantly better than other comparative methods, indicating that the images generated by this invention are more consistent with text semantics and have higher artistic quality in human subjective evaluation.

[0088] Table 1

[0089] Table 2

[0090] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A text-driven progressive neural painting method, characterized in that, include: S1: Perform semantic parsing, spatial guidance, and semantic alignment on the acquired text information to obtain multi-level text guidance conditions; S2: Employing a three-stage progressive optimization strategy of drafting, coloring, and refinement, and combining the aforementioned multi-level text guidance conditions, the brush stroke parameters are sequentially optimized through structural determination, color application, and detail enhancement to obtain optimized foreground brush stroke parameters. S2 includes: S2.1: In the drafting stage, the first brush stroke parameters are iteratively optimized multiple times to generate a sketch image, completing the structural determination; S2.2: In the coloring stage, the structural control conditions obtained from the morphological processing of the sketch image are acquired, and the second brush stroke parameters are iteratively optimized multiple times, with the structural control conditions being input in each iteration. S2.3: In the refinement stage, a detail-enhanced version is generated based on the colorized image to initialize the third stroke parameter, and then the initialized third stroke parameter is iteratively optimized multiple times to obtain the optimized foreground stroke parameter, thus completing the detail enhancement; wherein, the stroke parameter refers to the first stroke parameter in the drafting stage, the second stroke parameter in the coloring stage, and the third stroke parameter in the refinement stage. S3: Obtain the optimized background stroke parameters, and render the optimized background stroke parameters and the optimized foreground stroke parameters onto the same canvas using a renderer, outputting a final image that is semantically consistent with the text information and has enhanced details; wherein, obtaining the optimized background stroke parameters includes: generating background stroke parameters using background description information; performing multiple iterations to optimize the background stroke parameters, wherein each iteration includes: rendering the current background stroke parameters into a current background image using a differentiable renderer, calculating the current second fractional distillation sampling loss, the current fourth language-image alignment loss, and the total variation loss based on the current background image, and updating the current background stroke parameters according to the current second fractional distillation sampling loss, the current fourth language-image alignment loss, and the total variation loss; after reaching a fourth preset number of iterations, the optimized background stroke parameters are obtained.

2. The text-driven progressive neural painting method according to claim 1, characterized in that, S1 includes: Obtain a pre-trained large language model, a pre-trained text-to-image diffusion model, and a pre-trained contrastive language-image pre-trained model; Using the pre-trained large language model, the text information is semantically parsed to obtain a set of semantic objects, the RGB color parameters corresponding to each semantic object, the hierarchical order between semantic objects, and background description information. Using the pre-trained text-to-image diffusion model, spatial guidance processing is performed on the text information to obtain a global self-attention map, cross-attention maps corresponding to each semantic object, and a conditional basis for calculating fractional distillation sampling loss or interval fractional matching loss. Using the pre-trained contrastive language-image pre-trained model, semantic alignment processing is performed on the text information to establish the computational basis for language-image alignment loss, so as to measure the semantic consistency between the rendered image and the input text in the subsequent stroke parameter optimization process. The results obtained after processing the pre-trained large language model, the pre-trained text-to-image diffusion model, and the pre-trained contrastive language-image pre-trained model are combined to form the multi-level text guidance conditions.

3. The text-driven progressive neural painting method according to claim 2, characterized in that, S2.1 includes: Using the global self-attention map and the cross-attention map corresponding to each semantic object in the multi-level text guidance conditions, a multi-object structured sampling probability map is generated; Multiple initial positions are sampled within each semantic object region based on the multi-object structured sampling probability map, and the results are aggregated as the first stroke parameter; wherein, the color parameter in the first stroke parameter is initialized to white, and the canvas is set to black; The optimizable parameters other than the color parameter in the first stroke parameters are iteratively optimized multiple times, wherein each iteration includes: rendering the current first stroke parameters into a current first image through a differentiable renderer, calculating the current first fractional distillation sampling loss and the current first language-image alignment loss based on the current first image, and updating the optimizable parameters other than the color parameter in the current first stroke parameters according to the current first fractional distillation sampling loss and the current first language-image alignment loss; After reaching the first preset number of iterations, the optimized first stroke parameters are obtained; The optimized first stroke parameters are rendered into the sketch image using the differentiable renderer.

4. The text-driven progressive neural painting method according to claim 3, characterized in that, The step of generating a multi-object structured sampling probability map by utilizing the global self-attention map and the cross-attention maps corresponding to each semantic object in the multi-level text guidance conditions includes: Using the cross-attention maps corresponding to each semantic object, the binary mask of each semantic object is extracted by the Otsu automatic thresholding method; By utilizing the hierarchical order among the semantic objects, the binary masks of each semantic object are deduplicated to obtain the unambiguous mask of each semantic object; The global self-attention map is masked with the unambiguous mask of each semantic object, and the sampling probability distribution within the region of each semantic object is output. The sampling probability distributions within each semantic object region are merged to generate the multi-object structured sampling probability map.

5. The text-driven progressive neural painting method according to claim 2, characterized in that, S2.2 includes: The sketch image is subjected to erosion and dilation operations in sequence to obtain the structural control conditions; Using the structural control conditions, the pre-trained text-to-image diffusion model with control module is used to re-acquire the cross-attention map corresponding to each semantic object, and multiple initial positions are uniformly collected within the re-acquired cross-attention map corresponding to each semantic object, and summarized as the second stroke parameter. The color parameter in the second stroke parameter is initialized to the RGB color parameter corresponding to each semantic object in the multi-level text guidance condition; The second stroke parameter is iteratively optimized multiple times, wherein each iteration includes: rendering the current second stroke parameter into the current color image using a differentiable renderer; calculating the current first interval score matching loss using the pre-trained text-to-image diffusion model with a control module based on the current color image, the structural control conditions, and the text information; simultaneously calculating the current second language-image alignment loss using the contrastive language-image pre-trained model; and updating the current second stroke parameter according to the current first interval score matching loss, the current second language-image alignment loss, and the current color constraint term; the current color constraint term is calculated in real time during each iteration based on the difference between the current stroke color parameter and the initial color parameter of the second stroke parameter. After reaching the second preset number of iterations, the optimized second stroke parameters are obtained; The optimized second stroke parameters are rendered as the colored image.

6. The text-driven progressive neural painting method according to claim 2, characterized in that, S2.3 includes: Based on the enhanced detail version and the colored image, a restricted error map is obtained, and multiple initial positions are sampled according to the restricted error map as the third stroke parameter, wherein the color parameter in the third stroke parameter is initialized to the color value of the corresponding position in the enhanced detail version; The third stroke parameter is iteratively optimized multiple times, wherein each iteration includes: rendering the current third stroke parameter into a current detail map through a differentiable renderer; calculating the current second interval score matching loss and the current third language-image alignment loss based on the current detail map and some conditions in the multi-level text guidance conditions; and updating the current third stroke parameter according to the current second interval score matching loss and the current third language-image alignment loss. After reaching the third preset number of iterations, the optimized third stroke parameters are obtained, and together with the optimized second stroke parameters, they constitute the optimized foreground stroke parameters.

7. The text-driven progressive neural painting method according to claim 6, characterized in that, The process involves obtaining a limited error map based on the enhanced detail version and the colored image, and sampling multiple initial positions based on the limited error map as the third brushstroke parameters, including: Using a first classifier-free guided scale, the colorized image is inverted using a denoising diffusion probability model to obtain the inverted latent code; The inverted latent code is resampled using a second classifier-free guided scale to generate the enhanced detail version; wherein the second classifier-free guided scale is larger than the first classifier-free guided scale. Calculate the pixel-level absolute error between the colored image and the detail-enhanced version to obtain an error map; Using a semantic segmentation model, with the set of semantic objects as input, the mask of each semantic object is extracted from the colorized image to restrict the error map to the range of the mask formed by the masks of each semantic object, thus obtaining the restricted error map; The restricted error map is normalized into a sampling probability distribution; Multiple initial positions are sampled from the sampling probability distribution and summarized as the third stroke parameter.

8. The text-driven progressive neural painting method according to claim 1, characterized in that, Also includes: The optimized background image is rendered based on the optimized background stroke parameters.

9. The text-driven progressive neural painting method according to claim 1, characterized in that, The brush stroke parameters for different types or stages include position parameters, shape parameters, color parameters, and transparency parameters.