Controllable schematic diagram generation method and system based on large language model and storage medium
Through the controllable schematic generation method based on large language models, the challenges of schematic acquisition and understanding are solved, the schematic generation process is optimized, the training efficiency and performance of the model are improved, and strong support is provided for the development of cross-media intelligence.
Patent Information
- Application Number
- CN202510276662.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively solve the challenges of acquisition and understanding of schematic diagrams, especially when model training needs to be completed in low-resource scenarios with data constrained, which affects the training effect and generalization ability of the model.
The controllable diagram generation method based on a large language model is adopted, and the serialization process and image generation process are optimized by constructing a schematic generation task data set, pre-establishing the layout planning rules for schematic diagrams, serialization processing and positioning the application area of text expression form in the image, and combining multimodal embedding representation and fine-tuning diffusion model, the serialization processing process and image generation process are optimized.
Effectively respond to the challenges of high-level semantic complexity and underlying visual diversity in the generation process of schematic diagrams, significantly improving the training efficiency and performance of schematic diagram understanding models, and providing strong support for the development of cross-media intelligence.
Smart Images

Figure CN120219558A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of schematic diagram generation, and particularly relates to a controllable schematic diagram generation method, system and storage medium based on a large language model. Background Art
[0002] A schematic diagram is a highly abstract knowledge carrier, and its main function is to depict the structural principle or mechanism of things. Such images not only widely exist on digital knowledge platforms such as MOOC (Massive Open Online Course) websites, open knowledge bases, and technical forums, but also play a key role in promoting cross-media intelligent understanding. The effective analysis and understanding of schematic diagrams not only form the cornerstone of knowledge-intensive applications such as knowledge base construction and intelligent question answering, but also are crucial for improving the efficiency and accuracy of information retrieval. Nevertheless, there are obvious challenges in obtaining and understanding schematic diagrams compared with natural images. In particular, the scarcity of schematic diagrams and their copyright issues significantly limit the research and application of such images. These challenges are particularly prominent when constructing a neural network model for schematic diagram understanding, because model training needs to be completed in a low-resource scenario with limited data, thus affecting the training effect and generalization ability of the model. To address this problem, it is particularly crucial to explore the technology of generating schematic diagrams through natural language. This method can not only effectively reduce the difficulty of obtaining schematic diagrams, expand the available data set of schematic diagrams, but also significantly improve the training efficiency and performance of the schematic diagram understanding model, providing strong support for the further development of cross-media intelligence.
[0003] In recent years, the text-to-image generation task has achieved remarkable results. This task aims to generate images that are faithfully consistent with the described content by accepting text inputs that describe images. Current research on image generation algorithms mainly focuses on natural images, where early research relied on conditional generative adversarial networks. Researchers have made significant progress recently in autoregressive methods and diffusion model-based techniques by expanding the dataset and increasing the model scale. However, natural images and schematic diagrams exhibit profound differences in both underlying visual and high-level semantic features. From the perspective of underlying visual features, schematic diagrams have a diversity of underlying visual information and usually exhibit rich and more fine-grained visual objects. From the perspective of high-level semantic features, schematic diagrams have high-level semantic complexity, where the visual objects in the diagram are not only interrelated but also exhibit complex logical relationships. Given these significant differences, directly applying natural image generation methods to schematic diagrams will encounter major obstacles, mainly because they cannot fully capture the unique visual details and semantic complexity of schematic diagrams. Recently, large language models have proven their practical value in numerous language generation tasks, and language input has become an indispensable part of many vision-language tasks. With the powerful generalization ability demonstrated by large language models, recent research has begun to attempt to use these models to handle multi-modal tasks. The visual input is innovatively integrated with large language models to enhance the efficiency of complex reasoning in multi-modal scenarios. Summary of the Invention
[0004] An object of the present invention is to provide a controllable schematic diagram generation method, system, and storage medium based on a large language model, aiming at the above problems in the existing technology, optimizing the schematic diagram generation process, effectively coping with the challenges of high-level semantic complexity and underlying visual diversity in the schematic diagram generation process, and using the large language model to improve the overall quality of schematic diagram understanding technology.
[0005] To achieve the above object, the present invention has the following technical solutions:
[0006] In a first aspect, a controllable schematic diagram generation method based on a large language model is provided, including:
[0007] Construct a dataset for the schematic diagram generation task based on the large language model;
[0008] According to the image characteristics of the schematic diagram, pre-establish layout planning rules for the schematic diagram, and convert the dataset for the schematic diagram generation task based on the large language model into schematic diagrams with a set layout through the layout planning rules for the schematic diagram;
[0009] Serialize the layout planning rules for the schematic diagram, convert them into a text expression form, and locate the application area of the layout planning rules in the text expression form in the image;
[0010] The application area of the corresponding layout planning rules in the image, combined with the multi-modal embedding representation and the fine-tuned diffusion model, optimizes the serialization processing process and the image generation process to generate a schematic diagram that meets the requirements.
[0011] As a preferred solution, the construction of the schematic diagram generation task dataset based on the large language model includes the following steps:
[0012] Sort the existing schematic diagrams by category, and select schematic diagrams according to the principles of subject classification and data balance; use the large language model to perform fine-grained annotation on the selected schematic diagrams to form a text-to-schematic diagram generation dataset with annotation information.
[0013] As a preferred solution, the establishment of the layout planning rules of the schematic diagram according to the image characteristics of the schematic diagram, and the conversion of the schematic diagram generation task dataset based on the large language model into a schematic diagram with a set layout includes the following steps:
[0014] The planner uses the large language model to generate a layout, converting the input data from text into a schematic diagram layout that meets the requirements;
[0015] The auditor uses the large language model to give feedback on modification suggestions for the schematic diagram layout to the planner;
[0016] The planner updates the schematic diagram layout according to the auditor's modification suggestions and submits it to the auditor again;
[0017] During the interaction between the planner and the auditor, the schematic diagram layout is iteratively updated until the auditor determines that there is no room for modification or the upper limit of the predetermined number of modification times is reached, at which time the iterative update of the schematic diagram layout stops.
[0018] As a preferred solution, in the step of iterative update of the schematic diagram layout, according to the generative pre-training model, analyze the relationship between the gradient descent of the linear layer and the attention mechanism in the Transformer structure. The gradient update part of the linear layer is expressed by F(x)=(W0 + ΔW)x, where ΔW is the sum of the outer products of the input and the error signal, that is From this, it can be obtained that:
[0019]
[0020] In the linear attention model, x is used as the query, x′ is used as the key, and e i As the value;
[0021] In the context learning mechanism of the large language model, the attention calculation expression of a module is as follows:
[0022]
[0023] In the formula, ICL is In-Context Learning;
[0024] On this basis, when removing the activation function softmax and the loss function parameters, it is simplified to:
[0025]
[0026] In the formula, W V X′ is called the meta-feature, W V X (W K X) T Indicates that the zero-shot learning parameter fixed area is ΔW;
[0027] Introduce a generative pre-trained model as a meta-optimizer; generate meta-features in the forward propagation according to the context; through the attention mechanism, the meta-gradient is applied to the large language model to form an in-context learning model.
[0028] As a preferred solution, the application area of the layout planning rule for positioning the text expression form in the image includes:
[0029] Effectively fuse the position information describing two visual objects in the layout planning rule. The specific fusion method is as follows:
[0030] Bounding Box = (min(x 1,0 , x 2,0 ), min(y 1,0 , y 2,0 ), max(x 1,1 , x 2,1 ), max(y 1,1 , y 2,1 ))
[0031] In the formula, (x 1,0 , y 1,0 , x 1,1 , y 1,1 ) represents the coordinates of the first visual object, and (x 2,0 , y 2,0 , x 2,1 , y 2,1 ) represents the coordinates of the second visual object; the application area of the layout planning rule in the image is the fusion of the two visual object areas.
[0032] As a preferred solution, the optimization of the serialization process and the image generation process for the diffusion model combining multi-modal embedding representation and fine-tuning includes the following steps:
[0033] The multi-modal embedding representation supplements the semantics in text-to-image generation by introducing positional tokens, receives and processes spatial position data represented by four floating-point numbers for the upper-left and lower-right coordinates of each region, quantifies and converts the coordinates into discrete positional tokens; by combining the positional tokens and text descriptions, it allows users to control specific regions of the image and processes visual and semantic information as expected, achieving an understanding and application of the layout planning rules of the established schematic diagrams.
[0034] The fine-tuned diffusion model jointly optimizes the model to generate schematic diagrams that meet requirements through feature degradation in forward propagation, image reconstruction in reverse inference, and fine-tuning for the characteristics of schematic diagrams.
[0035] As a preferred solution, the feature degradation in forward propagation includes the following steps:
[0036] For a real image x0, assuming the corresponding distribution is q(x), the diffusion model cumulatively introduces Gaussian noise with zero mean in the forward process, generating a series of images x1, x2,..., x T ; define a sequence of Gaussian distribution variance parameters β t ∈(0, 1), regard the forward process as a Markov chain, and at each time step t, the image transforms from the previous moment t - 1:
[0037]
[0038] In this continuous process, as the time step t increases, the image x t is transformed into pure noise; when T approaches infinity, x T becomes an image completely composed of Gaussian noise; the parameter β is continuously selected, β1 < β2 < … < β T ; the noise is added to the image in a continuous and controllable manner;
[0039] The image reconstruction in reverse inference includes the following steps:
[0040] Achieve the reverse generation from the pure noise image t-1 |x t to the clear image x0 through the backward distribution q(x ); if q(x t-1 |x t ) is a Gaussian distribution and remains unchanged, then (x t-1 |x t ) also follows a Gaussian distribution;
[0041] Introduce the method of variational lower bound, and combine the structure of U-Net and attention mechanism to predict the backward distribution p θ :
[0042]
[0043]
[0044] Even if the actual backward distribution q(x t-1 |x t ) cannot be directly used, given the known starting point x0, q(x t-1 |x t , x0) can be derived through Bayes' formula:
[0045]
[0046] The fine-tuning for the characteristics of the schematic diagram includes the following steps:
[0047] The diffusion model is constructed based on the Stable Diffusion model architecture and consists of an autoencoder, a U-Net noise estimation module, and a CLIP ViT-L / 14 text encoder combined;
[0048] In the diffusion model, the encoder E converts the image x into a latent representation z = E(x) through eight-fold downsampling, and this representation is used in the diffusion process, while the decoder D is used to reconstruct the latent representation z into an image
[0049] The conditioning process of U-Net is based on the denoising time step and the text condition τ θ (y(T)), where y(T) is the input text query containing the text token T; in terms of the training strategy, the diffusion model extends the original text query y(T) to the input query y(P, T) that combines the text word T and the position token P;
[0050] The fine-tuning follows the loss function of latent diffusion modeling:
[0051]
[0052] In the formula, ε θ and τ θ represent the fine-tuned network modules;
[0053] During the fine-tuning process, the training data is obtained by cropping the image regions according to the annotated bounding boxes using the image descriptions; during fine-tuning, the model adjusts the short side of the image to the set number of pixels and randomly crops out a square region as the input image x.
[0054] In a second aspect, a controllable schematic diagram generation system based on a large language model is provided, including:
[0055] A dataset construction module for constructing a dataset for the schematic diagram generation task based on a large language model;
[0056] A layout planning rule conversion module for pre - establishing the layout planning rules of the schematic diagram according to the image characteristics of the schematic diagram, and converting the dataset for the schematic diagram generation task based on the large language model into a schematic diagram with a set layout through the layout planning rules of the schematic diagram;
[0057] A rule serialization processing and positioning module for serializing the layout planning rules of the schematic diagram, converting them into a text expression form, and positioning the application area of the layout planning rules in the text expression form in the image;
[0058] A multi - modal embedding representation and diffusion model fine - tuning module for optimizing the serialization processing process and the image generation process in combination with the multi - modal embedding representation and the fine - tuned diffusion model corresponding to the application area of the layout planning rules in the image, and generating a schematic diagram that meets the requirements.
[0059] In a third aspect, there is provided an electronic device, including:
[0060] A memory storing at least one instruction; and a processor for executing the instruction stored in the memory to implement the controllable schematic diagram generation method based on the large language model.
[0061] In a fourth aspect, there is provided a computer - readable storage medium storing at least one instruction, and the at least one instruction is executed by a processor in an electronic device to implement the controllable schematic diagram generation method based on the large language model.
[0062] Compared with the prior art, the present invention has at least the following beneficial effects:
[0063] By constructing a dataset for schematic diagram generation tasks based on large language models, the problems of unbalanced image categories in existing schematic diagram datasets and the lack of fine-grained annotations for generation tasks are solved. Aiming at the complexity of the high-level semantics of schematic diagrams, a layout planning rule for schematic diagrams is designed according to the image characteristics of schematic diagrams, and a method for layout planning of visual objects in schematic diagrams based on large language models is proposed. The input text is parsed through semantic analysis and logical reasoning of the large language model, optimizing the schematic diagram generation process. Aiming at the diversity of the underlying visual features of schematic diagrams, a semantic controllable visual content generation strategy for schematic diagrams based on diffusion models is designed to precisely control the layout of each element in the image and the semantic relationship. Through the serialization process of the schematic diagram layout planning rule, the integrity of semantic information is ensured and the requirements of the generation model are adapted, enabling the model to accurately execute layout instructions and achieve high-precision image rendering. Combining multi-modal embedding representation and a specially fine-tuned diffusion model, the input sequence processing and image generation processes are optimized, significantly improving the accuracy and quality of the image. The method of the present invention takes large language models as the core, focuses on the automatic generation and understanding of schematic diagrams, and effectively improves the overall quality of schematic diagram understanding technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and those of ordinary skill in the art can obtain other relevant drawings without creative efforts based on these drawings.
[0065] Figure 1 is a flowchart of the controllable schematic diagram generation method based on large language models in the embodiments of the present invention;
[0066] Figure 2 is a framework diagram for generating the layout planning rule of schematic diagrams based on large language models in the embodiments of the present invention;
[0067] Figure 3 is a process diagram for the serialization process of the layout planning rule of schematic diagrams in the embodiments of the present invention;
[0068] Figure 4 is a schematic diagram of the principle of generating semantic controllable visual content of schematic diagrams in the embodiments of the present invention;
[0069] Figure 5 is a process diagram of the multi-modal embedding representation in the embodiments of the present invention;
[0070] Figure 6 is a process diagram of the forward propagation and reverse reasoning of the diffusion model in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, those of ordinary skill in the art can also obtain other embodiments without creative efforts.
[0072] Please refer to Figure 1 , an embodiment of the present invention proposes a method for generating controllable schematic diagrams based on a large language model, which effectively addresses the problems of high-level semantic complexity and low-level visual diversity challenges in schematic diagram generation, mainly including the following steps:
[0073] S1. Construct a schematic diagram generation task dataset based on a large language model;
[0074] S2. Pre-establish layout planning rules for schematic diagrams according to the image characteristics of schematic diagrams, and convert the schematic diagram generation task dataset based on the large language model into schematic diagrams with a set layout through the layout planning rules of schematic diagrams;
[0075] S3. Serialize the layout planning rules of schematic diagrams, convert them into text expression forms, and locate the application areas of the layout planning rules in text expression forms in the image;
[0076] S4. Corresponding to the application areas of the layout planning rules in the image, combine multi-modal embedding representations and fine-tuned diffusion models to optimize the serialization process and the image generation process, and generate schematic diagrams that meet the requirements.
[0077] Step S1 of the embodiment of the present invention for constructing a schematic diagram generation task dataset based on a large language model includes the following steps:
[0078] Sort the existing schematic diagrams, and select schematic diagrams according to the principles of subject classification and data balance; use the large language model to perform fine-grained annotation on the selected schematic diagrams to form a text-to-schematic diagram generation dataset with annotation information.
[0079] Through in-depth classification and optimization of the existing schematic diagram dataset, the embodiment of the present invention creates a large-scale schematic diagram dataset with rich annotations to support subsequent research and applications.
[0080] In a possible implementation manner, for the layout planning generation problem in the schematic diagram generation task, traditional direct visual content generation algorithms often ignore the high-level semantic relationships of schematic diagrams, resulting in insufficient relevance between the generated schematic diagrams and the input text. To solve this problem, the present invention proposes a method for generating schematic diagram layout planning based on a large language model. This method utilizes the powerful semantic analysis ability of the large language model to address the challenges of high-level semantic complexity of schematic diagrams.
[0081] Please refer to Figure 2 Figure 2 , the present invention utilizes the powerful semantic analysis ability of the large language model to cope with the challenges of the high-level semantic complexity of the schematic diagram. Step S2 pre-establishes the layout planning rules of the schematic diagram according to the image characteristics of the schematic diagram, and converts the schematic diagram generation task data set based on the large language model into a schematic diagram with a set layout mainly including:
[0082] The planner uses the large language model to generate the layout, and converts the input data from text into a schematic diagram layout that meets the requirements;
[0083] The auditor uses the large language model to put forward modification opinions on the schematic diagram layout and feedback them to the planner;
[0084] The planner updates the schematic diagram layout according to the modification opinions of the auditor and submits it to the auditor again;
[0085] During the interaction process between the planner and the auditor, the schematic diagram layout is iteratively updated until the auditor determines that there is no room for modification, or the upper limit of the predetermined number of modification times is reached. At this time, the iterative update of the schematic diagram layout is aborted.
[0086] The specific implementation process of step S2 mainly includes three major components:
[0087] 1) Planner: Large language model layout generation
[0088] The planner plays a crucial role in the schematic diagram generation task. By using the large language model to generate the layout, the planner can convert the input text into a schematic diagram layout that meets the requirements. In this process, the planner first receives the generation rules and algorithms of the layout planning as input, and then generates a preliminary layout planning draft of the schematic diagram according to the input text. This draft is not just a simple conversion of the text, but is obtained after the in-depth understanding and semantic inference of the large language model. After generating the draft, the planner communicates with the auditor to obtain feedback and modification suggestions, so as to continuously optimize the layout rules and submit the optimized layout plan. This process is a cyclic iterative process, and the planner continuously adjusts the layout rules until the termination condition is reached. Through the powerful semantic understanding ability of the large language model, the planner can effectively convert the input text into a schematic diagram with a good layout, providing important support for the smooth progress of the schematic diagram generation task.
[0089] 2) Auditor: Large language model layout modification
[0090] Auditors play a crucial role in the layout planning of the schematic diagram. Use large language models to propose modification suggestions to ensure the accuracy and semantic expression ability of the layout planning. Auditors first receive the generation rules of the layout rules as input, and based on the input text content and the preliminary layout planning generated by the planner, put forward detailed modification suggestions. These modification suggestions may involve adjustments to aspects such as the layout structure, the position and relationship of visual elements, etc. Subsequently, the planner updates the layout planning according to the auditor's suggestions and submits it to the auditor again. The auditor continuously reviews and puts forward further modification suggestions, iterating multiple rounds until the auditor determines that there is no further room for modification or reaches the upper limit of the predetermined number of modification times, at which point the iteration stops. The work of the auditor ensures the quality and integrity of the layout planning and provides a solid foundation for the generation of the final schematic diagram.
[0091] 3) Iterative modification
[0092] During the interaction between the planner and the auditor, the layout planning of the schematic diagram is gradually iteratively updated. The in-context learning of large language models is essentially a meta-optimization process, which deeply explores the interaction between in-context learning and fine-tuning with the help of the GPT model. When analyzing the connection between the gradient descent of the linear layer and the attention mechanism in the Transformer structure, the gradient update part of the linear layer is expressed as F(x)=(W0 + ΔW)x, where ΔW is the sum of the outer products of the input and the error signal, that is It can be obtained from this that:
[0093]
[0094] In the linear attention model, x is used as the query, x′ is used as the key, and e i is used as the value;
[0095] In the in-context learning mechanism of large language models, the attention calculation expression of a module is as follows:
[0096]
[0097] where ICL is in-context learning;
[0098] On this basis, when removing the activation function softmax and the loss function parameters, it is simplified to:
[0099]
[0100] On this basis, further approximating the activation function, the self-attention of in-context learning is actually equal to a new operation of a cross-section. Where W V X′ is called the meta-feature, W V X (WK X) T Indicates that the zero-shot learning parameter fixed region is ΔW.
[0101] It has actually evolved into the concept of zero-shot learning. Generally speaking, the process of in-context learning is a meta-optimization process: introducing the pre-trained large language model GPT as a meta-optimizer; generating meta-features in the forward pass according to the context; through the attention mechanism, the meta-gradient is applied to the language model to form an in-context learning model.
[0102] In a possible implementation, due to the complexity of schematic layout planning, especially in the processing of text vocabulary, position markers, and semantic relationships. These elements are often not necessary to consider in traditional image generation tasks, but are crucial in schematic generation. Therefore, it becomes particularly critical to develop a method that can comprehensively represent and process these layout planning rules. An effective serialization processing method can not only simplify the complexity of schematic layout planning, but also provide strong support for subsequent visual content generation models. Such a method should be able to accurately capture and express the formal descriptions and positional relationships in the schematic, so as to ensure that the generated images are not only visually appealing but also semantically accurate. To achieve this goal, the method of the embodiment of the present invention proposes a serialization processing technology that uses a comprehensive encoding mechanism to uniformly encode various elements in the schematic, such as text descriptions, position data, and semantic connections. In this way, the generation model can better understand and execute the instructions in the layout planning, and achieve highly accurate and dynamic image rendering.
[0103] Please refer to Figure 3 , in the process of converting the layout planning of the schematic into a semantically controllable text sequence, the following two steps are mainly followed: 1. Textualization of layout rules: To effectively serialize the layout rules and convert them into text inputs that the model can understand, the present invention carefully defines the specific text meanings of each visual relationship. This process includes concretizing the abstract layout rules into text expression forms so that the model can accurately understand and execute these rules. The specific conversion strategies and details are shown in Table 1. 2. Region determination of visual rules: To achieve controllable semantic visual content generation, not only the layout rules need to be serialized into text form, but also the application regions of these rules in the image need to be accurately located. For example, it is necessary to clearly indicate which regions in the image space contain specific relationships between visual objects, such as connecting lines or directional elements. The goal of this processing step is to effectively fuse the position information of the two visual objects described in the layout planning rules. The specific fusion method is as follows:
[0104] Bounding Box=(min(x 1,0 ,x 2,0 ),min(y1,0 , y 2,0 ), max(x 1,1 , x 2,1 ), max(y 1,1 , y 2,1 ))
[0105] In the formula, (x 1,0 , y 1,0 , x 1,1 , y 1,1 ) represents the coordinates of the first visual object, and (x 2,0 , y 2,0 , x 2,1 , y 2,1 ) represents the coordinates of the second visual object; the application area of the layout planning rule in the image is the fusion of the two visual object areas.
[0106] Table 1
[0107]
[0108] In a possible implementation manner, for the existing schematic diagram layout planning criteria, due to the extremely diverse visual features of the schematic diagram and the inclusion of complex high-level semantics, there is still a large room for improvement in directly using the existing image generation model for the optimization of visual object content generation in the future. To overcome the limitations of the existing model in visual content optimization, the present invention introduces a new semantic controllable schematic diagram visual content generation method based on the diffusion model. This method utilizes the powerful ability of the diffusion model in visual content rendering to effectively address the challenges of the diversity of schematic diagram content. By introducing multi-modal embedding representations, semantic controllable visual content generation is achieved, which more accurately matches the semantic requirements of the schematic diagram description text.
[0109] Please refer to Figure 4 , the semantic controllable schematic diagram visual content generation method based on the diffusion model in the embodiments of the present invention mainly includes two major components: multi-modal embedding representation and schematic diagram generation diffusion model.
[0110] 1) Multi-modal embedding representation
[0111] The multi-modal embedding representation is generated by embedding the model input sequence, which can finely distinguish the semantics of different layout plans, thereby generating a more accurate representation that conforms to the original intention of the image description. In addition, the multi-modal embedding representation also enhances the adaptability of the model, enabling it to be flexibly applied in various different application scenarios, thereby optimizing the overall image generation process.
[0112] In the multi-modal embedding representation of a semantically controllable schematic visual content generation model, the problem of unclear semantics in text-to-image generation is improved by introducing position tokens. By adding instances of multi-modal embedding representation, semantic and spatial information can be combined to precisely control image generation. The model not only processes traditional text inputs but can also receive and process spatial position data represented by four floating-point numbers, i.e., the upper-left and lower-right coordinates of each region. These coordinates are quantized and converted into discrete position tokens, such as the position token <x1> , <y1> , <x2> <y2>。Such a sequence arrangement is similar to a short natural language sentence, making the overall input more natural and semantically clear. The specific operation of the multimodal embedding representation is as Figure 5 shown.
[0113] The motivation for multimodal embedding representation is to solve the problems in traditional text-to-image generation, where text descriptions may be difficult to precisely specify specific regions of an image due to ambiguity and verbosity. This problem is particularly severe in schematic diagram generation because the diversity of visual features and high-level semantic complexity involved in schematic diagrams require the generation model to be able to precisely capture and implement these detailed layout and content specifications. By combining positional markers and text descriptions, the semantically controllable schematic diagram visual content generation model not only allows users to precisely control specific regions of the image but also can more conformably process highly complex visual and semantic information. The model utilizes the powerful capabilities of large-scale pre-trained text-to-image generation technology and further enhances the processing ability for complex scenes by introducing the embedding of spatial information. Especially in the schematic diagram generation task, the method of the embodiment of the present invention can effectively understand and apply the schematic diagram layout planning rules generated previously, thus ensuring that the generated image is not only visually appealing but also semantically rich and accurate. In summary, the semantically controllable schematic diagram visual content generation model provides a more precise and controllable image generation method by combining the embedding of text and position information. This method makes the generated image more conform to the input description and layout planning and is particularly suitable for situations where complex scenes need to be created.
[0114] 2) Schematic diagram generation diffusion model
[0115] Specific fine-tuning of the diffusion model pre-trained on a large-scale image dataset is performed on the schematic diagram dataset to better adapt to and optimize the unique characteristics of schematic diagram generation. Through the feature degradation of forward propagation, the image reconstruction of reverse inference, and the detailed fine-tuning for the characteristics of schematic diagrams, the model is jointly optimized to precisely generate schematic diagrams that meet the requirements.
[0116] Regarding the training of the diffusion model in the semantically controllable schematic diagram visual content generation method, it mainly includes the following three steps: forward propagation of the diffusion model, reverse inference, and diffusion-based fine-tuning for schematic diagrams.
[0117] (1) Forward propagation of the diffusion model
[0118] In the diffusion model, the forward propagation process refers to the process of gradually adding noise to an image. Although this step itself does not directly generate an image, it is crucial for understanding how the diffusion model works and constructing the characteristics of its training set. Specifically, for a real image x0 with an assumed corresponding distribution q(x), the diffusion model cumulatively introduces Gaussian noise with a mean of zero during the forward process, generating a series of images x1, x2,..., x T , as Figure 6 shown. To carry out this process, a sequence of Gaussian distribution variance parameters β t ∈(0,1) is defined, and the forward process is regarded as a Markov chain. At each time step t, the image transforms from the previous time step t - 1:
[0119]
[0120] In this continuous process, as the time step t increases, the image x t is transformed into pure noise; theoretically, when T approaches infinity, x T becomes an image composed entirely of Gaussian noise; in actual operation, the parameter β is continuously selected, β1 < β2 < … < β T ; and the noise is added to the image in a continuous and controllable manner.
[0121] (2) Reverse Inference of the Diffusion Model
[0122] The forward propagation process of the diffusion model involves adding noise to an image, while reverse inference is the process of removing this noise.
[0123] Through the backward distribution q(x t-1 |x t ), the reverse generation from a pure noise image to a clear image x0 is achieved; if q(x t-1 |x t ) is a Gaussian distribution and remains unchanged, then (x t-1 |x t ) also follows a Gaussian distribution; however, directly applying the ordinary (x t-1 |x t ) is not feasible. Therefore, the method of variational lower bound is introduced, and a structure combining U-Net and attention mechanism is used to predict the reverse distribution p θ :
[0124]
[0125]
[0126] Even if the actual backward distribution q(x t-1 |x t ), given the known starting point x0, q(x is derived through Bayes' formula t-1 |x t , x0):
[0127]
[0128] (3) Fine-tuning of the diffusion model generated for the schematic diagram
[0129] The diffusion model for semantically controllable schematic diagram visual content generation is constructed based on the Stable Diffusion model architecture and is composed of an autoencoder, a U-Net noise estimation module, and a CLIP ViT-L / 14 text encoder;
[0130] In the diffusion model, the encoder E converts the image x into a latent representation z = E(x) through eight-fold downsampling, and this representation is used in the diffusion process, while the decoder D is used to reconstruct the latent representation z into an image
[0131] The conditioning process of U-Net is based on the denoising time step and the text condition τ generated by the text encoder θ (y(T)), where y(T) is the input text query containing the text token T; in terms of the training strategy, the diffusion model extends the original text query y(T) to the input query y(P, T) that combines the text word T and the position token P;
[0132] The fine-tuning follows the loss function of latent diffusion modeling:
[0133]
[0134] In the formula, ε θ and τ θ represent the fine-tuned network modules.
[0135] Except for the position token embedding E P other than that, all model parameters are inherited from the pre-trained Stable Diffusion model. During fine-tuning, in order to improve the adaptability to the visual diversity and complex high-level semantics of the schematic diagram, image descriptions and specific descriptions of each region are required for fine-tuning. The training data is obtained by processing the image regions cropped by the annotated bounding boxes using an advanced image description generation model. During fine-tuning, the model adjusts the short side of the image to 512 pixels and randomly crops a square region as the input image x to optimize the processing of the specific regions and text of the schematic diagram during training. This fine-tuning method not only exploits the powerful underlying performance of the diffusion model but also finely adjusts to meet the semantic requirements of complex scenarios.
[0136] Another embodiment of the present invention further provides a controllable schematic diagram generation system based on a large language model, including:
[0137] A dataset construction module for constructing a schematic diagram generation task dataset based on a large language model;
[0138] A layout planning rule conversion module for pre - establishing layout planning rules for schematic diagrams according to the image characteristics of schematic diagrams, and converting the schematic diagram generation task dataset based on the large language model into schematic diagrams with a set layout through the layout planning rules of schematic diagrams;
[0139] A rule serialization processing and positioning module for serializing the layout planning rules of schematic diagrams, converting them into a text expression form, and positioning the application area of the layout planning rules in the text expression form in the image;
[0140] A multi - modal embedding representation and diffusion model fine - tuning module for optimizing the serialization processing process and the image generation process in correspondence with the application area of the layout planning rules in the image, combining multi - modal embedding representation and a fine - tuned diffusion model, and generating schematic diagrams that meet the requirements.
[0141] Another embodiment of the present invention further provides an electronic device, including:
[0142] A memory storing at least one instruction; and a processor for executing the instruction stored in the memory to implement the controllable schematic diagram generation method based on the large language model.
[0143] Another embodiment of the present invention further provides a computer - readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the controllable schematic diagram generation method based on the large language model.
[0144] Exemplarily, the instruction stored in the memory can be divided into one or more modules / units. The one or more modules / units are stored in the computer - readable storage medium and executed by the processor to complete the controllable schematic diagram generation method based on the large language model of the present invention. The one or more modules / units can be a series of computer - readable instruction segments capable of completing specific functions, and this instruction segment is used to describe the execution process of the computer program in the server.
[0145] The electronic device can be a computing device such as a smart phone, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the electronic device may further include more or fewer components, or combine certain components, or different components. For example, the electronic device may further include input - output devices, network access devices, a bus, etc.
[0146] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0147] The memory may be an internal storage unit of the server, such as the hard disk or memory of the server. The memory may also be an external storage device of the server, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the server. Further, the memory may also include both the internal storage unit and the external storage device of the server. The memory is used to store the computer-readable instructions and other programs and data required by the server. The memory may also be used to temporarily store the data that has been output or will be output.
[0148] It should be noted that, regarding the information interaction, execution process, etc. between the above-mentioned module units, since they are based on the same concept as the method embodiments, for their specific functions and the technical effects brought, reference may be specifically made to the method embodiment part, and details will not be elaborated here.
[0149] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the foregoing method embodiments, and details will not be elaborated here.
[0150] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.
[0151] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0152] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application. < / x2> < / y1> < / x1>
Claims
1. A controllable schematic diagram generation method based on a large language model, characterized in that: include: Construct a dataset for diagram generation tasks based on a large language model; Pre-establishing a layout planning rule for the schematic diagram according to the image characteristics of the schematic diagram, and converting a schematic diagram generation task data set based on a large language model into a schematic diagram with a set layout through the layout planning rule for the schematic diagram; Serializing the layout planning rules of the schematic diagram, converting them into a textual expression form, and locating the application area of the layout planning rules in the textual expression form in the image; Corresponding to the application area of the layout planning rules in the image, the serialization process and the image generation process are optimized by combining the multimodal embedding representation and the fine-tuned diffusion model to generate a schematic diagram that meets the requirements.
2. The controllable schematic diagram generation method based on a large language model according to claim 1, characterized in that: The construction of a diagram generation task dataset based on a large language model comprises the following steps: The existing schematic diagrams are categorized and selected according to subject classification and data balance principles. The selected schematic diagrams are annotated in a fine-grained manner using a large language model to form a text-to-schematic generation dataset with annotated information.
3. The controllable schematic diagram generation method based on a large language model according to claim 1, characterized in that: The method of pre-establishing a layout planning rule of the schematic diagram according to the image characteristics of the schematic diagram, and converting the schematic diagram generation task data set based on the large language model into a schematic diagram with a set layout according to the layout planning rule of the schematic diagram comprises the following steps: Planners use large language models for layout generation, converting input data from text to a schematic layout that meets the requirements; Auditors use large language models to suggest changes to the schematic layout and give feedback to planners; The planner updates the schematic layout based on the auditor’s revisions and submits it to the auditor again; During the interaction between the planner and the auditor, the schematic layout is iteratively updated until the auditor determines that there is no room for modification, or reaches a predetermined upper limit on the number of modifications, at which point the iterative update of the schematic layout is terminated.
4. The controllable schematic diagram generation method based on a large language model according to claim 3 is characterized in that: In the step of iterative updating of the schematic layout, the gradient descent of the linear layer is analyzed according to the generative pre-training model, and the gradient update part of the linear layer is expressed by F(x)=(W0+ΔW)x, where ΔW is the sum of the outer products of the input and error signals, that is, It follows that: In the linear attention model, x is the query and x ′ As a key, e i as value; In the contextual learning mechanism of the large language model, the attention calculation expression of a module is as follows: Where, ICL is context learning; On this basis, removing the activation function softmax and the loss function parameters, the simplification is: Where W V X ′ is called meta-feature, W V X(W K X) T Indicates that the zero-shot learning parameter fixed region is ΔW; Introducing a generative pre-trained model as a meta-optimizer; Generate meta-features in the forward pass based on the context; Through the attention mechanism, meta-gradients are applied to the large language model to form a contextual learning model.
5. The controllable schematic diagram generation method based on a large language model according to claim 1, characterized in that: The application areas of the layout planning rules of the positioning text expression form in the image include: The position information of the two visual objects described in the layout planning rules is effectively fused. The specific fusion method is as follows: Bounding Box=(min(x 1,0 ,x 2,0 ),min(y 1,0 ,y 2,0 ),max(x 1,1 ,x 2,1 ),,max(y 1,1 ,y 2,1 )) In the formula, (x 1,0 ,y 1,0 ,x 1,1 ,y 1,1 ) represents the first visual object coordinate, (x 2,0 ,y 2,0 ,x 2,1 ,y 2,1 ) represents the coordinates of the second visual object; the application area of the layout planning rule in the image is the fusion of the two visual object areas.
6. The controllable schematic diagram generation method based on a large language model according to claim 1, characterized in that: The method of combining multimodal embedding representation and fine-tuning diffusion model to optimize the serialization process and image generation process includes the following steps: The multimodal embedding representation supplements the semantics in text-to-image generation by introducing position tags. The spatial position data represented by four floating-point numbers are received and processed for the upper left corner and lower right corner coordinates of each region, and the coordinates are quantified and converted into discrete position tags. By combining position tags and text descriptions, users are allowed to control specific areas of the image and process visual and semantic information as expected, so as to understand and apply the layout planning rules of the established schematic diagram. The fine-tuned diffusion model jointly optimizes the model to generate diagrams that meet the requirements through feature degradation of forward propagation, image reconstruction of reverse reasoning, and fine-tuning of diagram characteristics.
7. The controllable schematic diagram generation method based on a large language model according to claim 6, characterized in that: The feature degradation of the forward propagation comprises the following steps: For a real image x0, assuming the corresponding distribution is q(x), the diffusion model cumulatively introduces Gaussian noise with zero mean in the forward process to generate a series of images x1, x2, ..., x T ; Define a Gaussian distribution variance parameter sequence β t ∈(0,1), the forward process is regarded as a Markov chain. At each time step t, the image is transformed from the previous moment t-1: In this continuous process, as the time step t increases, the image x t turns into pure noise; when T tends to infinity, x T becomes an image composed entirely of Gaussian noise; the parameter β is selected continuously, β1<β2<…<β T ;Add noise to the image in a continuous and controllable way; The image reconstruction by reverse reasoning comprises the following steps: Through the backward distribution q(x t-1 |x t ) to achieve pure noise image To the reverse generation of the clear image x0; if q(x t-1 |x t ) is Gaussian and remains unchanged, then (x t-1 |x t ) also follows a Gaussian distribution; The variational lower bound method is introduced, and the structure of U-Net and attention mechanism is combined to predict the inverse distribution p θ : Even if we cannot directly use the actual backward distribution q(x t-1 |x t ), when the starting point x0 is known, q(x t-1 |x t ,x0): The fine-tuning of the schematic diagram characteristics includes the following steps: The diffusion model is built on the Stable Diffusion model architecture, which is composed of an autoencoder, a U-Net noise estimation module, and a CLIP ViT-L / 14 text encoder. In the diffusion model, the encoder E converts the image x into a latent representation z = E(x) by eight-fold downsampling, which is used in the diffusion process, and the decoder D is used to reconstruct the latent representation z into an image The conditioning of U-Net is based on the denoising time step and the text condition τ generated by the text encoder. θ (y(T)), where y(T) is the input text query containing the text tag T; in terms of training strategy, the diffusion model expands the original text query y(T) to the input query y(P,T) that combines the text word T and the position tag P; Fine-tuning follows the loss function of latent diffusion modeling: In the formula, ε θ and τ θ Represents the fine-tuned network module; During fine-tuning, the training data is obtained by cropping image regions according to the annotated bounding boxes using image descriptions; during fine-tuning, the model adjusts the short side of the image to the set pixels and randomly crops a square area as the input image x.
8. A controllable schematic diagram generation system based on a large language model, characterized in that: include: The dataset construction module is used to construct a dataset for diagram generation tasks based on a large language model; A layout planning rule conversion module is used to pre-establish the layout planning rules of the schematic diagram according to the image characteristics of the schematic diagram, and convert the schematic diagram generation task data set based on the large language model into a schematic diagram with a set layout through the layout planning rules of the schematic diagram; A rule serialization processing and positioning module is used to serialize the layout planning rules of the schematic diagram, convert them into textual expressions, and locate the application area of the layout planning rules in the textual expressions in the image; The multimodal embedding representation and diffusion model fine-tuning module is used to correspond to the application area of the layout planning rules in the image, combine the multimodal embedding representation and the fine-tuned diffusion model, optimize the serialization processing process and the image generation process, and generate a schematic diagram that meets the requirements.
9. An electronic device, characterized in that: include: A memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the controllable schematic diagram generation method based on a large language model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in an electronic device to implement the controllable schematic diagram generation method based on a large language model as described in any one of claims 1 to 7.
Citation Information
Cited By
Text-to-image generation method and device, equipment, storage medium and product
CN121330091A
Ancient furniture decoration pattern generation method based on multi-modal large model
CN121414902A