A SVG generation system and method based on large language model
By introducing multi-agent collaboration mechanism and iterative generation optimization technology into the SVG generation system, the problem of poor SVG generation quality in the existing technology is solved, and more accurate and detailed SVG image generation is achieved.
Patent Information
- Application Number
- CN202510169521.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Existing methods for generating SVG images using large language models are poor in quality, difficult to meet actual needs, and the generation results are deviated from expectations, especially when the user's description needs are blurred or simplified.
Using the SVG generation system based on a large language model, the multi-agent collaboration mechanism of the requirements analysis module, component design module, SVG design module and quality inspection module is used to decompose complex SVG generation tasks to multiple component generation tasks, and the results are generated through iterative generation and two-dimensional evaluation.
It significantly improves the quality of SVG generation, ensures that the generation results meet user needs more accurately, retain detailed characteristics, and continuously adjust through feedback mechanisms, reducing the update workload.
Smart Images

Figure CN119668582B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image text generation, and in particular to a SVG generation system and method based on a large language model. Background Art
[0002] As a modern image format, SVG (Scalable Vector Graphics) images have demonstrated their excellence in a variety of application scenarios with their unique advantages. SVG images based on XML code (SVG code) can be scaled losslessly to ensure that the image clarity is maintained at any resolution. Due to the small file size, it can improve web page loading speed and optimize user experience, so it is also favored in fields such as web development and image design.
[0003] Traditional SVG generation methods rely almost entirely on manual design, and their efficiency and cost cannot meet the increasingly complex needs. With the rapid development of artificial intelligence technology, especially the Large Language Model (LLM), the use of large language models to generate SVG images can significantly improve the design efficiency of SVG images and reduce labor costs by leveraging the model's powerful language understanding and code understanding and generation capabilities.
[0004] At present, the method of using a large language model to generate SVG images includes: the user describes the user's needs in a natural language text format, inputs the demand text into the large language model, and the large language model uses its powerful language understanding and code generation capabilities to generate XML code that is adapted to the user's needs, and then converts it into an SVG image. Compared with the manual design method, it has the characteristics of fast generation speed, and can create vector components on a large scale, which can significantly expand the vector resource library.
[0005] However, the SVG images generated by this method are often of poor quality and difficult to meet actual needs. The main reason is that the complexity of the target SVG exceeds the direct understanding ability of the large language model. The large language model directly generates an overall target SVG from a macro perspective based on the demand text, resulting in the lack of details and deviation from expectations. In addition, this method is directly affected by the richness of the demand text. When the user's demand description is vague or the demand description is relatively brief, it is difficult for the large language model to accurately grasp the user's needs, causing the generation result to deviate from expectations. In addition, when the generated result of this method does not meet the requirements, it can only regenerate a new complete SVG image based on the demand document. Only when the entire target SVG that meets the requirements is generated at one time, it has practical value. Summary of the invention
[0006] The object of the present invention is to provide an SVG generation system and method based on a large language model to improve the accuracy and practicality of generating SVG images using a large language model in order to address all or part of the above-mentioned problems.
[0007] The technical solution adopted by the present invention is as follows:
[0008] An SVG generation system based on a large language model includes a demand analysis module, a component design module, an SVG design module and a quality inspection module, wherein the demand analysis module, the component design module, the SVG design module and the quality inspection module all perform configuration operations based on the large language model; wherein:
[0009] The demand analysis module is configured to: formulate a product design document based on user requirements, the product design document indicating design requirements for a target SVG;
[0010] The component design module is configured to: based on the product design document, perform component decomposition on the target SVG and formulate a component disassembly document, wherein the component disassembly document indicates a disassembly result of the target SVG;
[0011] The SVG design module is configured to: design each component to obtain the target SVG based on the product design document and the component disassembly document; and adjust the corresponding component according to the modification suggestions fed back by the quality inspection module;
[0012] The quality inspection module is configured to: based on the product design document and the component disassembly document, review the target SVG generated by the SVG design module from two dimensions: SVG code and SVG image, and generate modification suggestions based on the review results and feed them back to the SVG design module.
[0013] To solve the above problem, the present invention further provides a method for generating SVG based on a large language model, the method comprising: guiding the large language model to perform the following operations:
[0014] Formulate a product design document based on user requirements, wherein the product design document indicates design requirements for a target SVG;
[0015] Based on the product design document, the target SVG is decomposed into components and a component disassembly document is prepared, wherein the component disassembly document indicates a disassembly result of the target SVG;
[0016] Based on the product design document and the component disassembly document, each component is designed to obtain the target SVG; and the corresponding components are adjusted according to the feedback modification opinions;
[0017] Based on the product design document and the component disassembly document, the generated target SVG is reviewed from two dimensions, SVG code and SVG image, and modification suggestions are generated and fed back according to the review results.
[0018] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0019] Based on the large language model, this application constructs four agents (artificial intelligence agents) for the SVG image generation task through prompt, namely, a demand analysis module, a component design module, an SVG design module and a quality inspection module. The demand analysis module conducts a comprehensive analysis of user needs and specifies a detailed product design document to illustrate the design requirements for the target SVG; the component design module refers to the product design document, decomposes the elements mentioned in it, writes a component disassembly document and sends it to the SVG design module; the SVG design module then generates the target SVG in an iterative generation manner according to the instructions of the product design document and the instructions of the component disassembly document, and after each iteration, the generated target SVG is passed to the quality inspection module, which reviews the target SVG in two dimensions, SVG code and SVG image, according to the instructions of the product design document and the instructions of the component disassembly document, and generates modification opinions based on the review results and feeds them back to the SVG design module for adjustment guidance for the next round of iteration. This application adopts a multi-agent collaboration mechanism, and decomposes and implements complex SVG image generation tasks through component disassembly and iterative generation. While meeting user needs, it retains more detailed features, and continuously tunes the target SVG through a feedback mechanism, significantly improving the quality of SVG generation, so that it can meet actual use requirements. In addition, this application decomposes complex SVG generation tasks into multiple components for implementation, simplifies many problems faced by directly generating target SVG, such as the difficulty in considering the collection shape, style, and independent identifiability of elements, and decomposes SVG generation tasks that are difficult to complete in one go into multiple independent component generation tasks, and draws them one by one in an iterative recursive manner, which improves the independent identifiability of components. And each component is independent of each other, and can be iteratively optimized independently, without the need to re-render the entire target SVG every time, reducing the workload of updating the target SVG. In addition, in the optimization stage, this application examines the target SVG from two dimensions, SVG code and SVG image, and comprehensively considers the visual and logical characteristics of the target SVG, further ensuring the accuracy of the target SVG generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present invention will now be described by way of example with reference to the accompanying drawings, in which:
[0021] Figure 1 It is a structural diagram of the SVG generation system based on the large language model provided in an embodiment of the present application.
[0022] Figure 2 It is a schematic diagram of the format of the product design document in the embodiment of the present application.
[0023] Figure 3 It is a schematic diagram of the format of the component disassembly document in the embodiment of the present application.
[0024] Figure 4 It is a schematic diagram of the format of the component necessity assessment document in the embodiment of the present application.
[0025] Figure 5 It is a comparison chart of the target SVG before and after the quality inspection module provided in the embodiment of the present application.
[0026] Figure 6 It is a review flowchart configured in one implementation manner of the quality inspection module provided in an embodiment of the present application.
[0027] Figure 7 It is a review flowchart configured in another implementation manner of the quality inspection module provided in the embodiment of the present application.
[0028] Figure 8 It is a flowchart of the SVG generation method based on the large language model provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] All features disclosed in this specification, or steps in all methods or processes disclosed, except mutually exclusive features and / or steps, can be combined in any manner.
[0030] Any feature disclosed in this specification (including any additional claims and abstract), unless otherwise stated, may be replaced by other alternative features that are equivalent or have similar purposes. That is, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.
[0031] In view of the problems that traditional manual design of SVG images is inefficient and costly, and that the existing use of large language models to generate SVG images has poor quality and low practical value, the embodiments of the present application provide an SVG generation system and method based on a large language model, aiming to improve the generation quality of SVG images and ensure the accuracy and practicality of SVG images.
[0032] In some embodiments, Figure 1As shown, the SVG generation system based on the large language model provided by the present application defines four agents, namely, the requirement analysis module, the component design module, the SVG design module and the quality inspection module, for the SVG generation task through prompts. That is, the requirement analysis module, the component design module, the SVG design module and the quality inspection module all perform configuration operations based on the large language model, and restrict the output content of each agent through prompts, so as to guide each agent to generate structured output results that meet the respective task requirements and goals. The four agents work collaboratively in accordance with the SOP (Standard Operating Procedure) idea of the SVG design task. The large language model used in the present application is, for example, the multimodal large model Claude-3.5-sonnet. Of course, it can also be other large language models with powerful reasoning capabilities.
[0033] The demand analysis module is configured to: formulate a product design document based on user requirements, wherein the product design document indicates design requirements for a target SVG.
[0034] User demand, i.e., the user's requirements for the SVG image (target SVG) to be generated, is described in natural language and input into the system in the form of user instructions in text format.
[0035] The demand analysis module guides the large language model to analyze and improve user needs through prompts. Relying on the thrust capability of the large language model and the knowledge learned from history, it supplements the unclear parts of user needs and improves the description text of user needs.
[0036] As an optional implementation, the product design document clearly indicates the design task, display target, application scenario and design style of the target SVG to be generated.
[0037] 1) Design tasks: Clarify the SVG design tasks specified in the user instructions, and rely on the reasoning ability of the large language model to supplement the unclear parts of the input user requirements and improve the user requirements.
[0038] 2) Display goal: Make it clear whether the SVG image is more focused on displaying the overall effect or the detailed effect, for example, whether it is displayed as a general appearance illustration or a technical detail illustration.
[0039] 3) Application scenarios: Clarify the specific application scenarios of SVG images, such as UI / UX design, brand and logo design, etc.
[0040] 4) Design style: Clarify the design style requirements of SVG images, such as simple, retro, realistic, etc.
[0041] In some specific embodiments, the product design document is Figure 2 The format shown clearly indicates the design requirements for the target SVG, which is then obtained by the component design module for component decomposition.
[0042] The component design module is configured to: based on the above-mentioned product design document, perform component decomposition on the target SVG planned to be generated, that is, decompose the SVG image generation task into several interrelated component generation tasks, and formulate a component disassembly document, which indicates the disassembly result of the target SVG planned to be generated.
[0043] In some embodiments, the component design module performs component decomposition on the target SVG according to the following configuration:
[0044] 1) Determine the target object in the target SVG and the viewing angle of the target object based on the product design document. Different viewing angles will cause different visual effects for the same target object. For example, it is difficult to observe the details on the back of the object from a frontal angle, and some components on the right side of the object will be hidden from a left-side viewing angle.
[0045] 2) According to the MEMC (Mutually Exclusiv Collectively Exhaustive; mutually independent and completely exhaustive) principle, the determined target object is disassembled into multiple components. In this way, the original SVG generation task as a whole is decomposed into multiple component generation tasks, and then the accuracy of the target object details is improved as much as possible while reducing the complexity of SVG generation, thereby improving the overall visual effect. The disassembly of the target object is the disassembly of the target object in three-dimensional space.
[0046] Under the premise of simplifying the complexity of the SVG image generation task, in an optional implementation, the present application also considers the visual consistency between the disassembled components and the target object. To this end, the component design module disassembles the target object in two steps according to the following configuration to ensure the coordination of the components and the target object in the overall visual effect as much as possible:
[0047] Based on the product design documents, the target object is macroscopically disassembled to identify the main structural units of the target object. The so-called main structural units are the main components of the target object. For example, if the target object is a car, its main components include the frame, engine, seat, center console, wheel, chassis, etc.
[0048] Based on the principle of visual independence or functional independence between components, each main structural unit is disassembled separately. For the disassembly of the main structural unit, each main structural unit is disassembled into smaller units. For a certain main structural unit, the further disassembled units are visually or functionally independent of each other. For example, for the wheel, it is visually or functionally disassembled into tires, wheels, etc.; for the central control, it is visually or functionally disassembled into display screens, instrument panels, button systems, air conditioning systems, etc. In this way, the target object can be disassembled into multiple components that are visually or functionally independent of each other, and the components of each part are interrelated under the same main structural unit, and the main structural units are interrelated, so as to ensure the correlation between the generation tasks of each component and the consistency of the construction between the components.
[0049] 3) Mark the attribute information of each component separately. The attribute information of the component is the physical attribute of the component. In some embodiments, by combining the relevant description of the target SVG in the product design document, such as the description of the appearance and interactive functions of the target object, and the model common sense of each component of the target object, detailed annotations are added to the physical attributes of each component (such as the color, shape, function, position, etc. of the component) to obtain the attribute information of each component. The attribute information is used to instruct the large language model to generate the SVG code of the component with the corresponding attribute, and the corresponding SVG image can be generated through the SVG code.
[0050] 4) Based on the perspective of the target object, the necessity of the disassembled components is marked to mark the necessary components and / or non-essential components in the disassembled components, where the necessary components are components that must be generated, and the non-essential components are components that do not have to be generated, that is, non-essential components can be generated or not. The necessity of the component refers to whether the component is necessary for the SVG generation task. As mentioned above, under different perspectives, some details of the target object will be hidden. If the SVG image is displayed as a two-dimensional image, the hidden components do not need to be generated. Therefore, the visible components can be marked for drawing, or the invisible components can be marked for screening.
[0051] In some optional implementations, the component design module marks the necessity of the disassembled components according to the following configurations:
[0052] Based on the design requirements for the target SVG indicated by the product design document, the necessary components of the target object are identified, and the remaining components are marked as non-essential components; among the identified necessary components, the visibility is judged based on the viewing angle of the target object, and the invisible components are marked as non-essential components. This method is to mark the non-essential components for screening. According to the instructions of the product design document, each component is judged to be a necessary component of the target SVG, that is, whether it is essential to the overall display effect of the target SVG. For example, for the display target of the general appearance illustration, it is necessary to ignore the details of the object, so the main component is a necessary component, and the detail component is not a necessary component; for the technical detail illustration, the focus is on the detail display, and at this time, the detail component is also a necessary component. Whether it is a main component or a detail component can be known through the attribute information of each component. If the component is not a necessary component, it is marked as a non-essential component. On the basis of identifying the necessary components, each necessary component is determined to be visible at the viewing angle according to the viewing angle of the target object. For the invisible necessary components, they can be marked as non-essential components.
[0053] In another method, the necessity of each component is marked by marking the necessary components. The method for necessary components and non-essential components in this method is the same as the previous marking method, the only difference is that the previous marking method marks non-essential components, while this method marks necessary components. This method identifies the necessary components of the target object based on the design requirements for the target SVG indicated by the product design document; among the identified necessary components, the visibility is judged based on the viewing angle of the target object, the invisible components are eliminated, and the remaining necessary components are marked.
[0054] In another method, the essential components and the non-essential components are marked separately. This method is based on any of the above two methods, and the unmarked components are further marked with opposite necessity.
[0055] As an example, in some specific embodiments, the format of the component disassembly document generated by disassembling the target object is as follows: Figure 3 shown.
[0056] For the components disassembled from the target object, in order to ensure the efficiency of generating the subsequent SVG design module, in some optional implementations, a separate component necessity assessment document is used to mark the necessity of each component. The component necessity assessment document adopts the last marking method mentioned above, that is, the necessity of each component is marked. In some specific embodiments, the format of the component necessity assessment document is as follows: Figure 4 As shown, the necessity of the component is marked by the parameter "Must Draw", "Yes" represents a necessary component, and "No" represents a non-essential component.
[0057] In a further preferred feasible implementation manner, the component design module may be further configured as follows based on marking the necessity of each component:
[0058] 5) Filter out non-essential components from all components disassembled from the target object. The filtering operation is performed differently for different necessity marking methods. For the method of marking non-essential components, the components with the mark are filtered out; for the method of marking necessary components, the unmarked components are filtered out; for the method of marking both non-essential components and necessary components, the components marked as non-essential components are filtered out, for example, Figure 4 In the method shown, components whose "Must Draw" is "No" are filtered out.
[0059] The SVG design module is configured to: design each component to obtain the target SVG based on the above product design document and component disassembly document; and adjust the corresponding components according to the modification suggestions fed back by the quality inspection module. Therefore, the SVG design module of this application includes two parts of work: one is to generate the corresponding components respectively and construct the target SVG; the other is to adjust the corresponding components according to the modification suggestions to optimize the target SVG.
[0060] As an optional implementation, the SVG design module designs each component according to the following configuration to obtain the target SVG:
[0061] 1) Based on the design requirements for the target SVG indicated in the product design document, formulate a generation strategy for each necessary component.
[0062] For example, the product design document clearly indicates the various design requirements for the target SVG, such as the design tasks, display goals, application scenarios and design styles described in the previous embodiments. Therefore, in some embodiments, the SVG design module formulates the generation strategy of each necessary component according to the design requirements for display goals, application scenarios and design styles indicated by the product design document to ensure that the target SVG finally generated meets the expected design specifications. For example, the display goal clarifies whether the SVG image focuses on displaying the overall effect or the detail effect, and whether it is displayed as a general appearance illustration or a technical detail illustration. For general appearance illustrations, the SVG design module will focus on using simple shapes and contrasting colors to construct each component, avoiding too many complex details; while for technical detail illustrations, the SVG design module focuses on the accuracy and functionality of the target SVG, and will use fine lines, paths and text annotations to express clear object details. The application scenario clarifies the specific application scenario of the SVG image and determines the actual function and interactivity of the target SVG. If the target SVG is used for UI / UX design, the SVG design module will consider the responsive design of the target SVG and design certain interactive functions in the interactive design; if the target SVG is used for brand and content design, the SVG design module will focus on drawing clear, concise and highly recognizable graphics to ensure that the target SVG can convey the brand concept. The design style clarifies the design style requirements of the SVG image and indicates the expression form of the target SVG. For example, for minimalist design requirements, the SVG design module will use simple geometric shapes and single tones; for retro design requirements, the SVG design module will use soft colors, gradient effects and complex decorative elements to create a nostalgic atmosphere; and for realistic design requirements, the SVG design module will use more detailed outlines, shadows and gradient effects to enhance the realism of the target SVG.
[0063] 2) Generate basic components among necessary components.
[0064] The so-called basic component is the first component of the target SVG, and subsequent components will be generated around this component. Therefore, the rationality of the selection of the basic component will affect the difficulty of generating subsequent components and even the visual effect of the final target SVG.
[0065] As an optional implementation, the SVG design module generates basic components among the necessary components according to the following configuration:
[0066] Based on the design requirements for the target SVG indicated in the product design document, the component with the greatest visual impact on the target SVG is identified from the necessary components as the basic component; and the SVG code of the basic component is generated. As the component with the greatest visual impact on the target SVG, it can be the largest component, or the component that best reflects the content expressed by the target SVG, or other components with outstanding characteristics. Different basic components may be obtained according to different guidance strategies for the large language model.
[0067] In addition, as an optional implementation, before generating the basic components, the SVG design module first divides the components (necessary components) into modules, and relies on the reasoning ability of the large language model to combine components that are closely related in physical structure into the same module, and divide components that are relatively independent in physical structure into different modules. Then, from the perspective of the module, consider the module that has the greatest visual impact on the target SVG, and select the component with the greatest visual impact from the module as the basic component. In addition, if the computing power is sufficient, all components of the module can be used as basic components to generate the SVG code of the basic components, that is, iteratively generate the target SVG in modules.
[0068] 3) Based on the basic components, other necessary components that are most relevant to the generated components are generated one by one in an iterative manner.
[0069] The so-called iterative method is to generate other components one by one, and the newly generated components are most closely related to the already generated components in terms of physical structure. For example, referring to the above embodiment, the newly generated components and the already generated components belong to the same module, or the newly generated components and the already generated components belong to different modules but are closest to each other. In this way, the newly generated components can be seamlessly integrated with the already generated components on the SVG code.
[0070] In addition, as an optional implementation, in the above embodiment of generating components in modules, based on the basic components, other modules that are most related in physical structure (such as the closest distance) are gradually added in an iterative manner to ensure that the SVG code of the newly added module is seamlessly integrated with the existing SVG code.
[0071] The SVG design module is also responsible for adjusting the corresponding components according to the modification suggestions. The modification suggestions fed back by the quality inspection module are adjustments to the attribute information of some or all of the generated components, such as adjustments in shape, color, position, etc. The modification suggestions can be the modified attribute information or the correction amount of the attribute information. The SVG design module adjusts the components involved in the adjustment according to the modification suggestions.
[0072] The quality inspection module is configured to review the target SVG generated by the SVG design module from two dimensions, SVG code and SVG image, based on the product design document and the component disassembly document, and generate modification suggestions based on the review results and feed them back to the SVG module.
[0073] As an optional implementation, the quality inspection module reviews the SVG code of the target SVG generated by the SVG design module according to the following configuration:
[0074] Identify the relative position relationship between each component from the SVG code of the target SVG. The SVG code uses XML language to describe the relevant attribute information of the SVG image. This attribute information includes not only the color, style, and interactive attributes of the component, but also the coordinates, size, and stacking order of the component. With this attribute information, the quality inspection module can identify the relative position relationship between each component from the SVG code. By locating the SVG code segment of the corresponding component in the SVG code (which can be located by the component name), the attribute information of the corresponding component can be quickly obtained.
[0075] According to the design requirements for the target SVG indicated in the product design document, the relative position relationship of the layout errors is identified, and the correction amount is calculated and written into the modification opinions. The product design document indicates the design requirements for the design tasks, display goals, application scenarios and design styles of the target SVG. According to these relevant requirements, the quality inspection module identifies the incorrect layouts with incorrect relative position relationships between components (such as overlap, occlusion, uneven proportions, etc.), and calculates the correction amount to adjust to the correct relative position relationship based on the identified relative position relationship of the layout errors, and writes it into the modification opinions. The correction amount is, for example, the pixel deviation amount, the size change amount, etc. Modification opinions based on quantification make the correction process of the SVG design module more systematic and data-based, and reduce correction errors. Figure 5 The following is a comparison diagram of the target SVG before and after the quality inspection module performs quality inspection. Figure 5 (a) is the SVG image initially generated by the SVG design module (compiled from the SVG code), Figure 5 (b) is the SVG image after the quality inspection module reviews the SVG code, proposes modification suggestions, and then adjusts the SVG design module. In this example, the pseudo code of the quality inspection module review process is as follows:
[0076]
[0077] The present application also examines the target SVG from the dimension of the SVG image. As an optional implementation, the quality inspection module examines the SVG image of the target SVG generated by the SVG design module according to the following configuration:
[0078] The SVG code of each necessary component is obtained respectively (the SVG code contains the attribute information of the component) to generate the corresponding component image. That is, each component is extracted separately to generate an independent component image.
[0079] Each component image is compared with the target SVG image generated by the SVG design module to identify component images with insufficient independent recognition. The so-called independent recognition describes whether the component is easy to identify. It refers to whether the component can be easily identified by its unique shape, color, detail features and other visual elements after it is separated from the target object, without being confused with other components. For example, the window component is extracted separately and compared with the entire car to determine whether the window component can be clearly identified. Usually, when the contrast with other components or the background is not obvious enough, for example, when the contrast between the window and the door is very small (that is, the color is similar), it will be difficult to visually distinguish the boundary between the door and the window, which is easy to cause misidentification of the component. Therefore, component images with insufficient independent recognition can be identified by analyzing the contrast between components that overlap or are adjacent to each other visually, and between components and the background.
[0080] An adjustment strategy that is conducive to improving the independent recognition of the identified component image is written into the modification opinion. The method of improving independent recognition can be achieved by improving the color contrast of the component, optimizing the shape structure of the component, or adding some visual elements that are conducive to recognition. Therefore, for a component image with insufficient independent recognition, at least one of the above adjustment strategies that is conducive to improving its independent recognition can be written into the modification opinion to guide the SVG design module to perform visual optimization on the component.
[0081] In some feasible implementations, after the SVG design module generates all components (necessary components), the quality inspection module conducts a review of all components, and the SVG design module makes corresponding adjustments based on the review modification suggestions, and then the quality inspection module reviews again, and the cycle continues. Figure 6 In other feasible implementations, after the SVG design module generates a component (necessary component) of a module, the quality inspection module reviews all components of the module, and the SVG design module makes corresponding adjustments to the components of the module according to the review modification opinions, and then the quality inspection module reviews again, and repeats this process until the components of the module do not need to be adjusted, and then the SVG design module generates the next module, and the quality inspection module reviews the components of the next module, and so on. Figure 7 shown.
[0082] As an optional implementation, the quality inspection module will generate modification opinions after each round of review, but the modification opinions may not contain suggestions for component adjustments (such as the amount of component correction or the adjustment strategy to improve the independent identification of components). When there are suggestions for component adjustments in the modification opinions, the SVG design module adjusts the corresponding components according to the adjustment suggestions. When there are no suggestions for component adjustments in the modification opinions, it means that the quality inspection module has passed the review, that is, the target SVG currently generated by the SVG design module is the final version.
[0083] In addition, the demand analysis module, component design module, SVG design module and quality inspection module in this application use an information sharing mechanism for communication. That is, the documents output by each module are stored centrally, and each module retrieves the required documents from the centrally stored documents according to its own needs. As an optional implementation method, this application designs the above-mentioned agents based on the AutoGen and LangChain frameworks. Through the agent communication mechanism provided by AutoGen, each agent can send and receive messages containing task information and data, thereby realizing collaborative work and information sharing among agents.
[0084] This application innovatively introduces a multi-agent collaboration mechanism and a dual-dimensional SVG evaluation and iterative optimization mechanism. Compared with the method of directly using a large language model to generate the target SVG as a whole at one time, it has the following outstanding features:
[0085] 1) Introduction of multi-agent collaboration mechanism: This application utilizes multiple agents to collaborate to perform SVG generation tasks. Through the design of multiple prompts, it defines a requirement analysis module, a component design module, an SVG design module, and a quality inspection module for the SVG generation task. It enables the agents to output structured documents to achieve standardized information transmission and collaboration among multiple agents, realize the decomposition and implementation of complex SVG generation tasks, and improve the quality of SVG generation.
[0086] 2) Introduction of component decomposition concept and iterative SVG image generation method: For complex SVG images, directly generating the entire image will face many challenges, such as the complexity of geometric shapes, the diversity of style attributes, etc., which will affect the effect of large language models directly generating target SVG. This application decomposes complex SVG images into multiple components, decomposes the SVG generation task that is difficult to complete in one go into multiple smaller component generation tasks, and after generating the SVG code of the basic component, it continues to gradually draw the SVG code of the next component / module based on the existing components, and iteratively constructs the entire SVG image. In addition, the componentized design makes the various elements (components) in the SVG image independent of each other. The SVG design module can modify the SVG code of a component separately without affecting other parts, without re-rendering the entire SVG image, reducing the complexity of SVG image review and optimization.
[0087] 3) Quantization-based correction mechanism: This application further proposes a quantitative adjustment mechanism, which uses the reasoning ability of a large language model to automatically analyze the relative position, size, and z-order of each component in the SVG code, so as to accurately calculate the relative offset, size change, and layer order adjustment between components, ensuring that the correction of each layout error problem can be accurately controlled at the pixel level.
[0088] 4) SVG dual-dimensional evaluation and iterative optimization mechanism: Since SVG has the dual-modal characteristics of code text and image, this application conducts collaborative review of SVG code text and image vision, breaking through the limitation of traditional SVG generation tasks that only focus on single-modal evaluation. And through the multi-agent collaboration mechanism, real-time feedback and optimization in the SVG generation process are ensured.
[0089] In some embodiments, Figure 8 As shown, the SVG generation method based on the large language model provided by the present application includes: guiding the large language model to perform the following operations:
[0090] S1. Formulate a product design document based on user requirements, where the product design document indicates design requirements for the target SVG;
[0091] S2. Based on the product design document, the target SVG is decomposed into components and a component disassembly document is prepared, where the component disassembly document indicates the disassembly result of the target SVG;
[0092] S3. Based on the component disassembly document, design each component separately to obtain the target SVG; and adjust the corresponding components according to the feedback modification suggestions;
[0093] S4. Based on the product design documents and component disassembly documents, the generated target SVG is reviewed from two dimensions: SVG code and SVG image. Modification suggestions are generated and feedback is provided based on the review results.
[0094] The operations performed in the above steps S1-S4 can be referred to for detailed features in the configurations of the demand analysis module, component design module, SVG design module and quality inspection module in the previous embodiment of the SVG generation system based on the large language model, so the optional implementation methods of each step will not be described separately here.
[0095] The present invention is not limited to the above-mentioned specific embodiments, but extends to any new features or any new combination disclosed in this specification, as well as any new method or process steps or any new combination disclosed.
Claims
1. A SVG generation system based on a large language model, characterized in that: It includes a demand analysis module, a component design module, an SVG design module and a quality inspection module, wherein the demand analysis module, the component design module, the SVG design module and the quality inspection module all perform configuration operations based on a large language model; wherein: The demand analysis module is configured to: formulate a product design document based on user requirements, the product design document indicating design requirements for a target SVG; The component design module is configured to: based on the product design document, perform component decomposition on the target SVG and formulate a component decomposition document, wherein the component decomposition document indicates a decomposition result of the target SVG; the component design module performs component decomposition on the target SVG according to the following configuration: Determine a target object in the target SVG and a viewing angle of the target object based on the product design document; disassemble the target object into multiple components according to the MEMC principle; mark the attribute information of each component respectively; mark the necessity of the disassembled components based on the viewing angle of the target object to mark necessary components and / or non-essential components in the disassembled components, wherein the necessary components are components that must be generated, and the non-essential components are components that do not have to be generated; The SVG design module is configured to: design each component to obtain the target SVG based on the product design document and the component disassembly document; and adjust the corresponding components according to the modification suggestions fed back by the quality inspection module; the SVG design module designs each component to obtain the target SVG according to the following configurations: formulate generation strategies for each necessary component based on the design requirements for the target SVG indicated by the product design document; generate basic components among the necessary components; based on the basic components, iteratively generate other necessary components that are most relevant to the generated components one by one; The quality inspection module is configured to: based on the product design document and the component disassembly document, review the target SVG generated by the SVG design module from two dimensions: SVG code and SVG image, and generate modification suggestions based on the review results and feed them back to the SVG design module.
2. The SVG generation system based on a large language model as claimed in claim 1, characterized in that: The product design document clearly indicates the design task, display target, application scenario and design style of the target SVG.
3. The SVG generation system based on a large language model as claimed in claim 1, characterized in that: The component design module disassembles the target object into multiple components according to the following configuration: Based on the product design document, the target object is macroscopically disassembled to identify the main structural units of the target object; Based on the principle of visual independence or functional independence between components, each main structural unit is disassembled separately.
4. The SVG generation system based on a large language model as claimed in claim 1, characterized in that: The component design module marks the necessity of disassembled components according to the following configurations: Based on the design requirements for the target SVG indicated by the product design document, necessary components of the target object are identified, and the remaining components are marked as non-essential components; among the identified necessary components, visibility is determined based on the viewing angle of the target object, and invisible components are marked as non-essential components; Alternatively, based on the design requirements for the target SVG indicated by the product design document, necessary components of the target object are identified; among the identified necessary components, visibility is determined based on the viewing angle of the target object, invisible components are removed, and the remaining necessary components are marked; Alternatively, based on the design requirements for the target SVG indicated by the product design document, the necessary components of the target object are identified, and the remaining components are marked as non-essential components; among the identified necessary components, the visibility is judged based on the viewing angle of the target object, and the invisible components are marked as non-essential components, while the visible components are marked as necessary components.
5. The SVG generation system based on a large language model as claimed in claim 1, characterized in that: The SVG design module generates basic components among the necessary components according to the following configuration: Based on the design requirements for the target SVG indicated by the product design document, identifying the component with the greatest visual impact on the target SVG from the necessary components as the basic component; Generate SVG code for the base component.
6. The SVG generation system based on a large language model according to any one of claims 1-2, characterized in that: The quality inspection module examines the SVG code of the target SVG generated by the SVG design module according to the following configuration: Identify the relative position relationship between components from the SVG code of the target SVG; According to the design requirements for the target SVG indicated by the product design document, the relative position relationship of the layout error is identified, and the correction amount is calculated and written into the modification suggestion.
7. The SVG generation system based on a large language model as claimed in claim 1, characterized in that: The quality inspection module examines the SVG image of the target SVG generated by the SVG design module according to the following configuration: Get the SVG code of each necessary component respectively and generate the corresponding component image; Comparing each component image with the target SVG image generated by the SVG design module, respectively, to identify component images that are insufficiently identifiable independently; An adjustment strategy that helps improve the independent recognizability of the identified component images is written into the modification opinion.
8. A method for generating SVG based on a large language model, characterized in that: The method includes: guiding the large language model to perform the following operations: Formulate a product design document based on user requirements, wherein the product design document indicates design requirements for a target SVG; Based on the product design document, the target SVG is decomposed into components, and a component disassembly document is prepared, wherein the component disassembly document indicates the disassembly result of the target SVG; the target SVG is decomposed into components according to the following method: based on the product design document, a target object in the target SVG and a viewing angle of the target object are determined; according to the MEMC principle, the target object is disassembled into multiple components; attribute information of each component is marked respectively; based on the viewing angle of the target object, the necessity of the disassembled components is marked to mark the necessary components and / or non-necessary components in the disassembled components, wherein the necessary components are components that must be generated, and the non-necessary components are components that are not required to be generated; Based on the component disassembly document, each component is designed to obtain the target SVG; and, the corresponding components are adjusted according to the feedback modification opinions; each component is designed to obtain the target SVG according to the following method: based on the design requirements for the target SVG indicated by the product design document, a generation strategy for each necessary component is formulated; basic components among the necessary components are generated; based on the basic components, other necessary components that are most relevant to the generated components are generated one by one in an iterative manner; Based on the product design document and the component disassembly document, the generated target SVG is reviewed from two dimensions, SVG code and SVG image, and modification suggestions are generated and fed back according to the review results.
Citation Information
Patent Citations
Application page starting method and device, equipment and medium
CN114546534A