Component-based text-guided color SVG (scalable vector graphics) image generation method and device
Through the component-based text guidance method, SVG images are generated using the text features input by the user and the preset component library, which solves the problem of large calculation consumption and incomplete generation paths of the existing SVG generation method, and realizes efficient and accurate SVG image generation.
Patent Information
- Application Number
- CN202411950753.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-06-06
AI Technical Summary
The existing SVG generation methods have problems such as high calculation and time consumption, difficulty in processing complex data sets, and incomplete generation paths. They especially show insufficient performance when generating complex SVGs using autoregressive language models.
The text-guided color SVG image generation method based on component is adopted, and text features are obtained through the boot text encoding input by the user, the base point component is determined and retrieved from the preset component library, and the components are gradually spliced to generate SVG images.
It significantly reduces redundant calculations during the generation process, improves the overall generation efficiency of the model, and ensures high quality and accuracy of the generated results, especially showing significant advantages when generating complex and fine SVG graphics.
Smart Images

Figure CN120107378A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer graphics, and in particular to a component-based text-guided color SVG image generation method and device. Background Art
[0002] Among the current SVG generation methods, one of the most straightforward approaches is to use a text-to-image generation model to first generate a bitmap and then convert it to SVG via vectorization techniques. Although this approach can generate high-quality results, it is computationally and time-consuming. Another approach is optimization-based methods, which start from a random SVG path and iteratively optimize to match the target SVG. Although these methods are able to generate SVG, they have significant disadvantages such as long processing time and high computational cost. In addition, some studies focus on unsupervised random generation of SVG paths, which poses challenges for practical applications due to the lack of guided output. Recently, there has been interest in generating SVG directly from text descriptions using autoregressive language models. However, SVG can contain a large amount of code, which poses a significant challenge to autoregressive language models when generating complex SVGs. The length and complexity of the code increase the difficulty of generating accurate and efficient SVG outputs, often resulting in poor performance in complex design tasks. In addition, generating SVG based on a single tag easily leads to incomplete or invalid paths, further complicating the generation process. Therefore, creating complex SVGs using these models remains a significant computational and methodological challenge.
[0003] Specifically, the field of generating images from text has made significant progress, going through several important stages involving generative adversarial networks (GANs) and diffusion models. Generative adversarial networks (GANs) have played an important role in this process, using a generator to generate images and a discriminator to evaluate the images, aligning the output with the text description through iterative optimization. Although text-conditioned GANs have made significant progress, they still face problems of scalability and stability when dealing with complex datasets. Recently, diffusion models have become the preferred method for text-to-image generation. These models start with Gaussian noise and gradually refine it to generate coherent images. Text-guided diffusion models guide the image generation process by introducing text embeddings directly or through cross-attention mechanisms. However, these studies have mainly focused on generating raster images with fixed resolution.
[0004] In the field of SVG generation, currently, the widely used method is an optimization-based method, which involves randomly initializing some SVG paths and then optimizing them using a differentiable rasterizer. However, this method is time-consuming, and it may take more than 20 minutes to generate an SVG graphic containing 24 SVG paths. Some methods generate SVG purely through random initialization without relying on text guidance. Recent studies have explored autoregressive methods for generating SVG, but the lengthy code of SVG graphics poses challenges when using this method to generate complex SVG.
[0005] In summary, although current SVG generation methods have made some progress in generation quality, they generally face problems such as computational and time consumption, difficulty in processing complex data sets, and incomplete generation paths. Text-to-image generation models and optimization methods have significant deficiencies in efficiency and computational cost, and autoregressive language models also show insufficient performance when processing complex SVG codes. Summary of the invention
[0006] In order to overcome the defects of the above-mentioned prior art, the present invention provides a component-based text-guided color SVG image generation and device, which can improve the generation quality and efficiency of SVG images.
[0007] An embodiment of the present invention provides a component-based text-guided color SVG image generation method, comprising the following steps:
[0008] According to the guide text input by the user, the text features are encoded;
[0009] Determine a base point component according to the text feature, and retrieve the base point component from a preset component library into a preset blank image to obtain an initial image;
[0010] Several stitching components are sequentially retrieved from a preset component library and stitched into the initial image, and an intermediate image is obtained after each stitching, until it is determined that the stitching components are no longer needed, and a color SVG image is output; wherein the stitching components retrieved for the first time are determined based on the text features and the initial image, and the stitching components retrieved each time thereafter are determined based on the text features and the intermediate image obtained from the previous stitching.
[0011] Furthermore, encoding the guide text input by the user to obtain text features specifically includes:
[0012] The guide text is input into a preset text encoder so that the preset text encoder extracts semantic information of the guide text and generates the text features.
[0013] Further, determining a base point component according to the text feature, and retrieving the base point component from a preset component library to a preset blank image to obtain an initial image specifically includes:
[0014] Inputting the preset blank image into a preset image encoder to obtain blank image features;
[0015] According to the text features and the blank image features, a base point component sequence is predicted by a preset SVG decoder; wherein the base point component sequence includes an identifier, a translation offset, a scaling factor, and a color mark corresponding to the base point component;
[0016] According to the identifier corresponding to the base point component, the base point component is retrieved from the preset component library, and according to the translation offset and scaling factor corresponding to the base point component, the base point component is moved to the preset blank image, and finally the base point component is colored according to the color mark corresponding to the base point component to obtain the initial image.
[0017] Preferably, the training process of the preset SVG decoder is specifically as follows:
[0018] Randomly extract an SVG file from a preset SVG training set, and rasterize the SVG file into a raster image, then input the raster image into a preset title generator to obtain training text, and generate training text features based on the training text;
[0019] Input the training text features and the blank image features into the preset SVG decoder, predict a training base point component sequence, and according to the training base point component sequence, retrieve the training base point components from a preset component library and splice them into a preset blank image to obtain a training initial image;
[0020] According to the training text features and the training initial image, a training component sequence corresponding to a first training splicing component is predicted by the preset SVG decoder, and according to the training component sequence, the first training splicing component is retrieved from a preset component library and spliced into the initial image to obtain a training intermediate image;
[0021] Repeatedly inputting the training intermediate image generated by each splicing into the preset SVG decoder, so that the preset SVG decoder continuously predicts the next training splicing component according to the training text features and the training intermediate image, until the preset SVG decoder determines that the training splicing component is no longer needed according to the latest generated training intermediate image and the training text features, and outputs the latest generated training intermediate image as the training SVG image, completing one iteration of training;
[0022] When the number of iterations exceeds a preset number threshold, the iterative training is terminated to obtain a trained preset SVG decoder.
[0023] Furthermore, the construction process of the preset component library is specifically as follows:
[0024] Decompose all SVGs in a preset SVG data set into separate paths, and determine each of the paths as a component;
[0025] Removing the colors of all the components, and scaling the components after the colors are removed to a preset component unit size to obtain image components;
[0026] All the image components are redundantly merged, and the preset component library is formed according to the redundantly merged image components.
[0027] Preferably, the redundant merging of all the image components specifically includes:
[0028] According to the paths of the image components, the image components with completely the same paths are merged in the same items;
[0029] Convert the image components after merging the same items into grayscale images of preset pixel size;
[0030] Calculating the Jaccard similarity index between the grayscale images respectively, and when the Jaccard similarity index between two grayscale images exceeds a preset similarity threshold, determining that the image components corresponding to the two grayscale images are a pair of similar components;
[0031] All the similar components are merged respectively until the remaining image components are no longer similar to each other, thus completing the redundant merging.
[0032] Further, the method sequentially retrieves a plurality of splicing components from the preset component library and splices them into the initial image, and obtains an intermediate image after each splicing, until it is determined that the splicing component is no longer needed, and then outputs a color SVG image, which specifically includes:
[0033] Inputting the initial image into the preset image encoder to obtain initial image features;
[0034] According to the text features and the initial image features, predicting by the preset SVG decoder a splicing component sequence corresponding to the first splicing component;
[0035] According to the splicing component sequence, the first splicing component is retrieved from a preset component library and spliced into the initial image to obtain an intermediate image;
[0036] Repeatedly input the intermediate image generated by each splicing into the preset SVG decoder, so that the preset SVG decoder continuously predicts the next splicing component according to the text features and the intermediate image, until the preset SVG decoder determines that the splicing component is no longer needed according to the latest generated intermediate image and the text features, and outputs the latest generated intermediate image as the color SVG image.
[0037] Another embodiment of the present invention provides a component-based text-guided color SVG image generation device, including: a text module, a base point module, and a splicing module;
[0038] The text module is used to encode text features according to the guide text input by the user;
[0039] The base point module is used to determine the base point component according to the text feature, and retrieve the base point component from the preset component library to the preset blank image to obtain the initial image;
[0040] The splicing module is used to sequentially retrieve several splicing components from a preset component library and splice them into the initial image, obtaining an intermediate image after each splicing, until it is determined that the splicing component is no longer needed, and then outputting a color SVG image; wherein the splicing component retrieved for the first time is determined based on the text features and the initial image, and the splicing component retrieved each time thereafter is determined based on the text features and the intermediate image obtained from the previous splicing.
[0041] Furthermore, the text module is used to encode text features according to the guide text input by the user, specifically including:
[0042] The guide text is input into a preset text encoder so that the preset text encoder extracts semantic information of the guide text and generates the text features.
[0043] Furthermore, the base point module is used to determine the base point component according to the text feature, and retrieve the base point component from the preset component library to the preset blank image to obtain the initial image, which specifically includes:
[0044] Inputting the preset blank image into a preset image encoder to obtain blank image features;
[0045] According to the text features and the blank image features, a base point component sequence is predicted by a preset SVG decoder; wherein the base point component sequence includes an identifier, a translation offset, a scaling factor, and a color mark corresponding to the base point component;
[0046] According to the identifier corresponding to the base point component, the base point component is retrieved from the preset component library, and according to the translation offset and scaling factor corresponding to the base point component, the base point component is moved to the preset blank image, and finally the base point component is colored according to the color mark corresponding to the base point component to obtain the initial image.
[0047] Compared with the prior art, the beneficial effects of the present invention are:
[0048] By componentizing SVG components and combining text guidance and autoregressive generation strategies, redundant calculations in the generation process are significantly reduced, and the overall generation efficiency of the model is improved. At the same time, the autoregressive generation strategy enables the model to gradually refine the generation process to ensure high quality and accuracy of the generated results. This feature enables the present invention to show significant advantages in practical applications, especially in scenarios where complex and fine SVG graphics need to be generated. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 A flowchart of a component-based text-guided color SVG image generation method provided by an embodiment of the present invention.
[0050] Figure 2 A flowchart for constructing a preset component library provided by an embodiment of the present invention.
[0051] Figure 3 A schematic diagram of a process for generating a color SVG image provided by an embodiment of the present invention.
[0052] Figure 4 A schematic diagram for comparing SVG images generated by the method of the present invention and other models provided in an embodiment of the present invention.
[0053] Figure 5 A schematic diagram of a portion of an SVG image generated by the method of the present invention is provided as an embodiment of the present invention.
[0054] Figure 6 A structural schematic diagram of a component-based text-guided color SVG image generation device provided by another embodiment of the present invention. DETAILED DESCRIPTION
[0055] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;
[0056] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0057] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0058] Reference Figure 1 , is a flowchart of a component-based text-guided color SVG image generation method provided by an embodiment of the present invention, comprising the following steps:
[0059] S1: Encode the text features according to the guide text input by the user;
[0060] S2: determining a base point component according to the text feature, and retrieving the base point component from a preset component library to a preset blank image to obtain an initial image;
[0061] S3: sequentially retrieve several stitching components from a preset component library and stitch them into the initial image, obtaining an intermediate image after each stitching, until it is determined that the stitching components are no longer needed, and then outputting a color SVG image; wherein the stitching components retrieved for the first time are determined based on the text features and the initial image, and the stitching components retrieved each time thereafter are determined based on the text features and the intermediate image obtained from the previous stitching.
[0062] For step S1, specifically, encoding the guide text input by the user to obtain text features specifically includes:
[0063] The guide text is input into a preset text encoder so that the preset text encoder extracts semantic information of the guide text and generates the text features.
[0064] In a preferred embodiment, the guide text input by the user needs to first be passed through a preset text encoder to generate a high-dimensional feature vector (i.e., the text features). These text features capture the semantic information in the text and provide a basis for subsequent generation steps. In this process, the guide text is not limited to simple labels, but can also be a complex description, such as "a piece of emerald green leaves and bright flowers". This can prompt the model to generate richer and more expressive SVG.
[0065] For step S2, specifically, determining a base point component according to the text feature, and retrieving the base point component from a preset component library to a preset blank image to obtain an initial image specifically includes:
[0066] Inputting the preset blank image into a preset image encoder to obtain blank image features;
[0067] According to the text features and the blank image features, a base point component sequence is predicted by a preset SVG decoder; wherein the base point component sequence includes an identifier, a translation offset, a scaling factor, and a color mark corresponding to the base point component;
[0068] According to the identifier corresponding to the base point component, the base point component is retrieved from the preset component library, and according to the translation offset and scaling factor corresponding to the base point component, the base point component is moved to the preset blank image, and finally the base point component is colored according to the color mark corresponding to the base point component to obtain the initial image.
[0069] In a preferred embodiment, after obtaining the text features, the decoding and generation phase of SVG can be entered. During the decoding process, the preset SVG decoder predicts components from the component library based on the text features, generates a sequence to gradually build the SVG image. First, the model starts from the BOS tag and predicts the first component. Then, the translation offset and scaling factor of the component are gradually predicted to ensure that it is accurately placed in the appropriate position in the image. Finally, the model generates appropriate RGB color values according to the input text prompt, assigns the component its color, and obtains the initial image.
[0070] Preferably, the training process of the preset SVG decoder is specifically as follows:
[0071] Randomly extract an SVG file from a preset SVG training set, and rasterize the SVG file into a raster image, then input the raster image into a preset title generator to obtain training text, and generate training text features based on the training text;
[0072] Input the training text features and the blank image features into the preset SVG decoder, predict a training base point component sequence, and according to the training base point component sequence, retrieve the training base point components from a preset component library and splice them into a preset blank image to obtain a training initial image;
[0073] According to the training text features and the training initial image, a training component sequence corresponding to a first training splicing component is predicted by the preset SVG decoder, and according to the training component sequence, the first training splicing component is retrieved from a preset component library and spliced into the initial image to obtain a training intermediate image;
[0074] Repeatedly inputting the training intermediate image generated by each splicing into the preset SVG decoder, so that the preset SVG decoder continuously predicts the next training splicing component according to the training text features and the training intermediate image, until the preset SVG decoder determines that the training splicing component is no longer needed according to the latest generated training intermediate image and the training text features, and outputs the latest generated training intermediate image as the training SVG image, completing one iteration of training;
[0075] When the number of iterations exceeds a preset number threshold, the iterative training is terminated to obtain a trained preset SVG decoder.
[0076] In a preferred embodiment, during the training process, the original SVG training set provides fewer textual cues, usually limited to category labels such as "flowers" without further description of the SVG. To address this limitation, the preferred embodiment introduces a caption generator that is only used during training and not during inference.
[0077] First, a random SVG file is extracted from the training set and rasterized into an image. The raster image is input into the title generator, which generates a description such as "a bouquet of flowers and a pink ribbon." This text description is then input into the text encoder to generate text features. At the same time, image features are generated by inputting a blank image into the image encoder.
[0078] The image features output by the image encoder are fused with the text features to obtain combined features, which are then input into a preset SVG decoder to enable the preset SVG decoder to predict the components required for the final generated SVG image.
[0079] As the components are progressively generated and concatenated, the preset SVG decoder also takes as input the SVG decoder sequence to generate the next tag. The partially generated SVG is rasterized into an image and then fed back into the image encoder. This process is repeated until the end tag is generated or the maximum sequence length is reached.
[0080] The image encoder in this process allows the model to understand semantics from a textual or path perspective as well as a visual perspective. By introducing the image encoder, the model can recognize and understand the components selected at each step, including their layout and color. This dual learning mechanism ensures that the model understands the specific visual attributes and spatial arrangement of elements, thereby improving its overall semantic understanding and performance in tasks involving textual and visual data.
[0081] For step S2, further, the construction process of the preset component library is specifically as follows:
[0082] Decompose all SVGs in a preset SVG data set into separate paths, and determine each of the paths as a component;
[0083] Removing the colors of all the components, and scaling the components after the colors are removed to a preset component unit size to obtain image components;
[0084] All the image components are redundantly merged, and the preset component library is formed according to the redundantly merged image components.
[0085] Preferably, the redundant merging of all the image components specifically includes:
[0086] According to the paths of the image components, the image components with completely the same paths are merged in the same items;
[0087] Convert the image components after merging the same items into grayscale images of preset pixel size;
[0088] Calculating the Jaccard similarity index between the grayscale images respectively, and when the Jaccard similarity index between two grayscale images exceeds a preset similarity threshold, determining that the image components corresponding to the two grayscale images are a pair of similar components;
[0089] All the similar components are merged respectively until the remaining image components are no longer similar to each other, thus completing the redundant merging.
[0090] In a preferred embodiment, referring to Figure 2 , is a flowchart of building a preset component library provided by an embodiment of the present invention. Figure 2 (a) It can be seen that all SVGs in the preset SVG dataset are first decomposed into individual paths, each of which forms a unique component. These isolated components have various different appearances and need to be thoroughly reorganized.
[0091] Next, if Figure 2 As shown in (b), these components need to be decolored and normalized. Initially, the geometric centers of these components are not aligned with the origin, which makes the scaling and normalization process difficult. To solve this problem, the geometric center of each component needs to be calculated and translated to the origin. Then, the components are scaled to 100 units in the longest dimension, ensuring that the expansion in both the positive and negative axes is 50 units. This process reorganizes and normalizes the components, making them ready for subsequent model use.
[0092] The initial component library contains many duplicate components. For example, the shapes of flowers in a bouquet may be exactly the same except for the color. Therefore, it is necessary to merge these redundant components. Figure 2As shown in (c), we first merge paths that are exactly the same in SVG. This initial merging reduces the number of components but does not eliminate all redundancy. To further optimize the component library, all components are converted into 100×100 pixel images, and these images are converted into grayscale images. In these grayscale images, the area enclosed by the path is represented by 1 (white) and the background is represented by 0 (black). We then calculate the Jaccard similarity index for each pair of images and sort them in descending order of similarity scores. By using a union-find data structure, a merging threshold is set and the similarity scores are checked sequentially. If the similarity between two components meets or exceeds the threshold, they are merged when they are determined to be a pair of similar components. The root nodes of these merged groups are considered the final components in the library. This process generates a refined and integrated component library that ensures minimal redundancy and optimal component organization for subsequent use.
[0093] When it is necessary to restore standardized components, such as Figure 2 As shown in (d), it needs to be placed in the correct position, similar to how people arrange elements in drawing software. This involves scaling the component according to the scale factor and translating it according to the offset values on the X and Y axes (Offset X and Offset Y). This ensures accurate placement of the component. Finally, the component is colored using the generated RGB markup to restore its original appearance.
[0094] For step S3, specifically, the step of sequentially retrieving a plurality of splicing components from the preset component library and splicing them into the initial image, obtaining an intermediate image after each splicing, and outputting a color SVG image after determining that the splicing component is no longer needed, specifically includes:
[0095] Inputting the initial image into the preset image encoder to obtain initial image features;
[0096] According to the text features and the initial image features, predicting by the preset SVG decoder a splicing component sequence corresponding to the first splicing component;
[0097] According to the splicing component sequence, the first splicing component is retrieved from a preset component library and spliced into the initial image to obtain an intermediate image;
[0098] Repeatedly input the intermediate image generated by each splicing into the preset SVG decoder, so that the preset SVG decoder continuously predicts the next splicing component according to the text features and the intermediate image, until the preset SVG decoder determines that the splicing component is no longer needed according to the latest generated intermediate image and the text features, and outputs the latest generated intermediate image as the color SVG image.
[0099] In a preferred embodiment, referring to Figure 3 , is a schematic diagram of a generation process of a color SVG image provided by an embodiment of the present invention. Figure 3 It can be seen that when the first splicing component is generated, the model continues to predict the next splicing component and repeats the same process until all splicing components are generated, and finally a complete SVG image is formed. During the entire generation process, the model continuously refers to text features and combines the intermediate image generated in the previous step to dynamically adjust the subsequent generation steps to ensure that the final generated color SVG image is highly consistent with the input text prompt.
[0100] In order to further improve the quality of generation, the preferred embodiment introduces some optimization techniques in the generation process. The preset SVG decoder can determine the layout of subsequent paths by viewing the partially generated SVG paths or intermediate images, thereby avoiding overlap between components and generating a more natural and coordinated image.
[0101] In summary, the embodiment of the present invention successfully generates high-quality color SVG images by accurately encoding text prompts, gradually decoding SVG components, and integrating visual and text features. This generation process not only improves the semantic understanding ability of the model, but also significantly enhances the visual consistency and complexity of the generated images, making it perform well in diverse and complex SVG generation tasks.
[0102] In order to evaluate the effect of the component-based text-guided color SVG image generation method provided by the present invention, this preferred embodiment compares the performance of the method of the present invention and other commonly used models.
[0103] In the comparison, the preferred embodiment uses Fréchet Inception Distance (FID) to quantify the distance between the image features of the generated SVG image and the real SVG image. In addition, the preferred embodiment also uses two CLIPScore indicators: CLIPScore-T2I, which is used to measure the similarity between the generated SVG rasterized image and the text hint used for generation; CLIPScore-I2I, which is used to measure the similarity between the generated SVG rasterized image and the real SVG rasterized image.
[0104] In addition, the preferred embodiment also evaluates the "uniqueness" and "novelty" of the generated SVG, which are derived from SkexGen. "Uniqueness" refers to the proportion of generated data that appears only once in all generated results, while "novelty" refers to the proportion of generated data that does not exist in the training set. In addition, the preferred embodiment introduces the Human Preference Score (HPS) to evaluate the user's satisfaction with the generated results. Together, these indicators provide a comprehensive evaluation of the quality, originality, and performance of our SVG generation method. Finally, the preferred embodiment measures the generation speed of each SVG to evaluate the efficiency of the method.
[0105] In this comparison, the preferred embodiment compares and evaluates the method of the present invention with various SOTA open source models on the ColorSVG-100K dataset. VectorFusion uses a diffusion model trained on image pixel representation to generate SVG without the need for a large number of SVG datasets with titles. By optimizing the differentiable vector graphics rasterizer, VectorFusion extracts semantic knowledge from the pre-trained diffusion model to generate high-quality vector graphics of various styles. CLIPDraw uses the pre-trained CLIP language-image encoder to synthesize new drawings by optimizing the similarity between the generated drawings and the given description. It processes vector strokes rather than pixel images and can generate simple and easily recognizable human shapes. DiffSketcher optimizes a set of Bézier curves based on a pre-trained diffusion model to create vectorized hand-drawn sketches that retain the subject structure and visual details. We re-implemented IconShop using Flan-T5 as the backbone, and used an autoregressive Transformer to sequence and tokenize SVG paths and text descriptions, thereby improving the icon synthesis capability. In addition, this comparison also explores the capabilities of GPT-3.5 and GPT-4 in SVG generation, demonstrating the progress of large language models (LLMs) in this task. These methods are all text-guided, allowing this preferred embodiment to evaluate the performance and effectiveness of the method described in the present invention compared to the current SOTA methods in the field of SVG generation. The comparison results are shown in the following table (the bold values in the table represent the overall best performance among all models):
[0106]
[0107]
[0108] As can be seen from the above table, CLIPScore is divided into T2I (representing CLIPScore between generated SVG image and text description) and I2I (representing CLIPScore between generated SVG image and real SVG image). It can be clearly seen from the table that the optimization-based models have lower FID scores, which indicates that the images they generate are closer to the distribution of real images in the test set. In contrast, except for the method described in the present invention, the language-based models generally have higher FID scores, which indicates that the previous methods failed to accurately fit the distribution of the dataset.
[0109] Regarding CLIPScore, optimization-based methods generally achieve higher T2I scores. This may be because the images generated by these models are more consistent with the image distribution that the CLIP model was originally trained on. The CLIP model may not have been extensively trained on the SVG dataset, so it has difficulty capturing the planar 2D nature of SVG graphics in the CLIPScore-T2I score. In contrast, CLIPScore-I2I measures the similarity between generated images and real SVG images and is therefore not affected by this limitation. It is worth noting that the method described in the present invention also achieves the highest CLIPScore among language-based models.
[0110] Again, since HPS uses the same CLIP model, it tends to favor the optimized SVG generation model from bitmap images. Regarding uniqueness, the optimization-based model achieved a score of 100% due to its randomness, while the language-based model showed slightly lower uniqueness due to the possibility of generating similar outputs from similar inputs. However, the method described in the present invention achieved the highest uniqueness score among the language-based models. All models achieved a score of 100% in terms of novelty.
[0111] In terms of the generation speed of each SVG, the value in brackets indicates the speedup factor relative to VectorFusion. Obviously, it takes longer to generate an SVG based on the optimized model, with VectorFusion taking about 13.69 minutes, and even the fastest CLIPDraw taking 2.92 minutes. In contrast, the language-based model only takes a few seconds, greatly reducing the generation time. The method of the present invention achieves a generation time of 1.36 seconds, which is 604 times faster than VectorFusion, highlighting the potential and advantages of language-based models in SVG generation.
[0112] Furthermore, the preferred embodiment also analyzes the impact of different SVG complexities on model performance. First, the SVGs are sorted in ascending order according to the number of paths and divided into four complexity levels: 0-25%, 25-50%, 50-75%, and 75-100%. Among them, the 0-25% range represents simple SVGs with the least paths, indicating the lowest complexity, while the 75-100% range covers SVGs with the most paths, indicating the highest complexity. Then, the FID indicator is used to evaluate the performance of different models at these complexity levels, and the evaluation results are shown in the following table:
[0113]
[0114] As can be seen from the above table, the method described in the present invention shows the best results in all complexity levels. It is observed that as the complexity of SVG increases, the performance of the optimization-based model does not deteriorate, while the performance of the language-based model decreases. This may be because randomly initialized bitmaps are usually more complex than SVG.
[0115] Finally, this preferred embodiment compares and analyzes SVGs generated by different models to evaluate the advantages of the method of the present invention over the prior art.
[0116] Reference Figure 4 , which is a schematic diagram of a comparison of SVG images generated by the method of the present invention and other models provided in an embodiment of the present invention. Figure 4 It can be seen that the results generated by DiffSketcher are visually excellent and accurately capture the original shape and color of the object, which may be due to the application of the diffusion model. However, some strange line artifacts appeared during the optimization process from bitmap to SVG. The output generated by GPT-4 matches the object in color and meets the outline requirements to a certain extent, but the outline is difficult to interpret without additional text. IconShop generates accurate outlines, with poor performance except for the taxi example. This difference may be due to the fact that the complexity of the SVG in the dataset used by the method described in the present invention is higher than the dataset used by IconShop, indicating that it is more suitable for simple SVG generation tasks.
[0117] The results generated by the method described in the present invention are consistent with the objects in terms of color and outline. However, the "avocado" example is affected by the seeds, which affects the overall tone, and the "taxi" looks more abstract. However, overall, the method described in the present invention is closer to the expression based on the optimization model, showing its potential.
[0118] Finally, refer to Figure 5 , which is a schematic diagram of a portion of an SVG image generated by the method of the present invention provided in one embodiment of the present invention.
[0119] Reference Figure 6 , is a structural diagram of a component-based text-guided color SVG image generation device provided by another embodiment of the present invention, comprising: a text module 101, a base point module 102 and a splicing module 103;
[0120] The text module 101 is used to encode the guide text input by the user to obtain text features;
[0121] The base point module 102 is used to determine the base point component according to the text feature, and retrieve the base point component from the preset component library to the preset blank image to obtain the initial image;
[0122] The stitching module 103 is used to sequentially retrieve a number of stitching components from a preset component library and stitch them into the initial image, obtaining an intermediate image after each stitching, until it is determined that the stitching component is no longer needed, and then outputting a color SVG image; wherein the stitching component retrieved for the first time is determined based on the text features and the initial image, and the stitching component retrieved each time thereafter is determined based on the text features and the intermediate image obtained from the previous stitching.
[0123] Furthermore, the text module 101 is used to encode the guide text input by the user to obtain text features, which specifically include:
[0124] The guide text is input into a preset text encoder so that the preset text encoder extracts semantic information of the guide text and generates the text features.
[0125] Furthermore, the base point module 102 is used to determine the base point component according to the text feature, and retrieve the base point component from the preset component library to the preset blank image to obtain the initial image, which specifically includes:
[0126] Inputting the preset blank image into a preset image encoder to obtain blank image features;
[0127] According to the text features and the blank image features, a base point component sequence is predicted by a preset SVG decoder; wherein the base point component sequence includes an identifier, a translation offset, a scaling factor, and a color mark corresponding to the base point component;
[0128] According to the identifier corresponding to the base point component, the base point component is retrieved from the preset component library, and according to the translation offset and scaling factor corresponding to the base point component, the base point component is moved to the preset blank image, and finally the base point component is colored according to the color mark corresponding to the base point component to obtain the initial image.
[0129] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A component-based text-guided color SVG image generation method, characterized in that: The steps include: According to the guide text input by the user, the text features are encoded; Determine a base point component according to the text feature, and retrieve the base point component from a preset component library into a preset blank image to obtain an initial image; Several stitching components are sequentially retrieved from a preset component library and stitched into the initial image, and an intermediate image is obtained after each stitching, until it is determined that the stitching components are no longer needed, and a color SVG image is output; wherein the stitching components retrieved for the first time are determined based on the text features and the initial image, and the stitching components retrieved each time thereafter are determined based on the text features and the intermediate image obtained from the previous stitching.
2. The component-based text-guided color SVG image generation method according to claim 1, characterized in that: The encoding of the guide text input by the user to obtain text features specifically includes: The guide text is input into a preset text encoder so that the preset text encoder extracts semantic information of the guide text and generates the text features.
3. The component-based text-guided color SVG image generation method according to claim 1, characterized in that: The step of determining a base point component according to the text feature, and retrieving the base point component from a preset component library to a preset blank image to obtain an initial image specifically includes: Inputting the preset blank image into a preset image encoder to obtain blank image features; According to the text features and the blank image features, a base point component sequence is predicted by a preset SVG decoder; wherein the base point component sequence includes an identifier, a translation offset, a scaling factor, and a color mark corresponding to the base point component; According to the identifier corresponding to the base point component, the base point component is retrieved from the preset component library, and according to the translation offset and scaling factor corresponding to the base point component, the base point component is moved to the preset blank image, and finally the base point component is colored according to the color mark corresponding to the base point component to obtain the initial image.
4. The component-based text-guided color SVG image generation method according to claim 3, characterized in that: The training process of the preset SVG decoder is specifically as follows: Randomly extract an SVG file from a preset SVG training set, and rasterize the SVG file into a raster image, then input the raster image into a preset title generator to obtain training text, and generate training text features based on the training text; Input the training text features and the blank image features into the preset SVG decoder, predict a training base point component sequence, and according to the training base point component sequence, retrieve the training base point components from a preset component library and splice them into a preset blank image to obtain a training initial image; According to the training text features and the training initial image, a training component sequence corresponding to a first training splicing component is predicted by the preset SVG decoder, and according to the training component sequence, the first training splicing component is retrieved from a preset component library and spliced into the initial image to obtain a training intermediate image; Repeatedly inputting the training intermediate image generated by each splicing into the preset SVG decoder, so that the preset SVG decoder continuously predicts the next training splicing component according to the training text features and the training intermediate image, until the preset SVG decoder determines that the training splicing component is no longer needed according to the latest generated training intermediate image and the training text features, and outputs the latest generated training intermediate image as the training SVG image, completing one iteration of training; When the number of iterations exceeds a preset number threshold, the iterative training is terminated to obtain a trained preset SVG decoder.
5. The component-based text-guided color SVG image generation method according to claim 1, characterized in that: The construction process of the preset component library is specifically as follows: Decompose all SVGs in a preset SVG data set into separate paths, and determine each of the paths as a component; Removing the colors of all the components, and scaling the components after the colors are removed to a preset component unit size to obtain image components; All the image components are redundantly merged, and the preset component library is formed according to the redundantly merged image components.
6. The component-based text-guided color SVG image generation method according to claim 5, characterized in that: The redundant merging of all the image components specifically includes: According to the paths of the image components, the image components with completely the same paths are merged in the same items; Convert the image components after merging the same items into grayscale images of preset pixel size; Calculating the Jaccard similarity index between the grayscale images respectively, and when the Jaccard similarity index between two grayscale images exceeds a preset similarity threshold, determining that the image components corresponding to the two grayscale images are a pair of similar components; All the similar components are merged respectively until the remaining image components are no longer similar to each other, thus completing the redundant merging.
7. The component-based text-guided color SVG image generation method according to claim 3, characterized in that: The method sequentially retrieves a plurality of splicing components from the preset component library and splices them into the initial image, and obtains an intermediate image after each splicing, until it is determined that the splicing component is no longer needed, and then outputs a color SVG image, specifically including: Inputting the initial image into the preset image encoder to obtain initial image features; According to the text features and the initial image features, predicting by the preset SVG decoder a splicing component sequence corresponding to the first splicing component; According to the splicing component sequence, the first splicing component is retrieved from a preset component library and spliced into the initial image to obtain an intermediate image; Repeatedly input the intermediate image generated by each splicing into the preset SVG decoder, so that the preset SVG decoder continuously predicts the next splicing component according to the text features and the intermediate image, until the preset SVG decoder determines that the splicing component is no longer needed according to the latest generated intermediate image and the text features, and outputs the latest generated intermediate image as the color SVG image.
8. A component-based text-guided color SVG image generation device, characterized in that: include: Text module, base point module and splicing module; The text module is used to encode text features according to the guide text input by the user; The base point module is used to determine the base point component according to the text feature, and retrieve the base point component from the preset component library to the preset blank image to obtain the initial image; The splicing module is used to sequentially retrieve several splicing components from a preset component library and splice them into the initial image, obtaining an intermediate image after each splicing, until it is determined that the splicing component is no longer needed, and then outputting a color SVG image; wherein the splicing component retrieved for the first time is determined based on the text features and the initial image, and the splicing component retrieved each time thereafter is determined based on the text features and the intermediate image obtained from the previous splicing.
9. The component-based text-guided color SVG image generation device according to claim 8, characterized in that: The text module is used to encode the guide text input by the user to obtain text features, specifically including: The guide text is input into a preset text encoder so that the preset text encoder extracts semantic information of the guide text and generates the text features.
10. The component-based text-guided color SVG image generation device according to claim 8, characterized in that: The base point module is used to determine the base point component according to the text feature, and retrieve the base point component from the preset component library to the preset blank image to obtain the initial image, which specifically includes: Inputting the preset blank image into a preset image encoder to obtain blank image features; According to the text features and the blank image features, a base point component sequence is predicted by a preset SVG decoder; wherein the base point component sequence includes an identifier, a translation offset, a scaling factor, and a color mark corresponding to the base point component; According to the identifier corresponding to the base point component, the base point component is retrieved from the preset component library, and according to the translation offset and scaling factor corresponding to the base point component, the base point component is moved to the preset blank image, and finally the base point component is colored according to the color mark corresponding to the base point component to obtain the initial image.