Drawing large model adjusting method and system based on multi-modal knowledge driving
By deeply parsing the text prompts input by the user and using a multimodal knowledge-driven method, the problem of existing image generation models generating errors or hallucinations in complex tasks is solved, achieving higher quality and controllable image generation.
Patent Information
- Application Number
- CN202510824992.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing image generation models have difficulty accurately capturing users' high-level semantics and factual details when dealing with complex, sophisticated, or knowledge-required generation tasks, resulting in frequent content errors or hallucinations in the generated images.
By obtaining the text prompts of user input for deep analysis and extracting structured intent representation, and combining multimodal knowledge retrieval and knowledge-attention translation modules, rich multimodal knowledge is introduced to provide accurate semantic information and factual basis for the generation process, and dynamically modulate the attention layer parameters of the large drawing model.
It achieves a more accurate understanding of user intent, generates images that are highly consistent with external knowledge and have more precise details, and greatly improves the quality and controllability of generated images.
Smart Images

Figure CN120747299A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of large model adjustment, and more specifically, to a large drawing model adjustment method and system based on multimodal knowledge driving. Background Art
[0002] Artificial intelligence has made significant progress in image generation. Large generative models, particularly those based on diffusion models, are now capable of generating high-quality and diverse images based on textual instructions. Through complex training processes, these models are able to learn the semantics, style, and structure of images from massive amounts of text and image data and transform them into visual representations. However, despite the remarkable potential of existing technologies in image generation, their inherent limitations and performance in specific application scenarios still present numerous challenges.
[0003] Existing mainstream text-to-image generation models primarily rely on encoding user-entered text prompts and using these encoded results as conditions to guide the image generation process. While this approach is intuitive and easy to use, its shortcomings become increasingly prominent when handling complex, sophisticated, or knowledge-intensive generation tasks. First, the semantic expression of the text prompt itself is often inherently ambiguous and incomplete. Users may find it difficult to accurately describe all their intentions with a limited text vocabulary, resulting in the results generated by the model potentially deviating from their actual expectations. For example, when a user needs to generate an image that contains a specific historical event, scientific principle, or cultural symbol, a simple text description often fails to convey sufficient information, making it difficult for the model to accurately capture these high-level semantics and factual details. This can lead to the generated image containing common sense errors, distorted details, or hallucinations that are inconsistent with objective facts. Summary of the Invention
[0004] In order to solve the above technical problems, the present application is proposed. The embodiments of the present application propose a method and system for adjusting a large drawing model based on multimodal knowledge drive, which aims to overcome the inherent limitations of text prompts in existing image generation models and the lack of explicit knowledge integration.
[0005] According to one aspect of the present application, a method for adjusting a large drawing model driven by multimodal knowledge is provided, comprising: obtaining a text prompt input by a user; performing text parsing and semantic embedding encoding on the text prompt to obtain a text prompt structured intent representation and a text prompt standard condition embedding; performing multimodal knowledge retrieval based on the text prompt structured intent representation to obtain a multimodal knowledge representation; inputting the text prompt standard condition embedding and the multimodal knowledge representation into a knowledge-attention translation module to obtain attention modulation parameters; inputting the text prompt standard condition embedding and a randomly generated initial noise latent representation into a large drawing model to obtain a final denoised latent representation, wherein the attention modulation parameters are used to influence the parameters of the attention layer of the denoising U-Net network of the large drawing model; and performing feature decoding on the final denoised latent representation to obtain a generated image.
[0006] According to another aspect of the present application, a multimodal knowledge-driven drawing large model adjustment system is provided, which is used to execute the above-mentioned multimodal knowledge-driven drawing large model adjustment method, including: a text prompt acquisition module, used to obtain a text prompt input by a user; a text semantic parsing and embedding module, used to perform text parsing and semantic embedding encoding on the text prompt to obtain a text prompt structured intent representation and a text prompt standard condition embedding; a multimodal intent-driven retrieval module, used to perform multimodal knowledge retrieval based on the text prompt structured intent representation to obtain a multimodal knowledge representation; a multi-source attention modulation module, used to input the text prompt standard condition embedding and the multimodal knowledge representation into a knowledge-attention translation module to obtain attention modulation parameters; a text-guided attention denoising module, used to input the text prompt standard condition embedding and a randomly generated initial noise latent representation into a drawing large model to obtain a final denoised latent representation, wherein the attention modulation parameters are used to influence the parameters of the attention layer of the denoising U-Net network of the drawing large model; and a denoising representation feature decoding module, used to feature decode the final denoised latent representation to obtain a generated image.
[0007] Compared with the prior art, the multimodal knowledge-driven drawing large model adjustment method and system provided by the present application first obtains the text prompt input by the user, and deeply analyzes it to extract a structured representation of the user's intention. Subsequently, based on this structured intention, it actively retrieves and integrates external multimodal knowledge, thereby introducing richer and more accurate semantic information and factual basis for the generation process. Then, a knowledge-attention translation module is introduced, which can convert this rich multimodal knowledge into refined attention modulation parameters. These parameters will directly and dynamically affect the working mode of the attention mechanism inside the drawing large model, realizing knowledge-driven fine-grained feature generation control. In this way, complex user intentions can be understood more accurately, the "hallucination" phenomenon can be avoided, and images that are highly consistent with external knowledge and have more accurate details can be generated, greatly improving the quality and controllability of the generated images. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0009] Figure 1 The figure illustrates a schematic flow chart of a large drawing model adjustment method based on multimodal knowledge driving according to an embodiment of the present application.
[0010] Figure 2 The figure illustrates a schematic flow chart of step S2 in the method for adjusting a large drawing model driven by multimodal knowledge according to an embodiment of the present application.
[0011] Figure 3 The figure illustrates a schematic flow chart of step S3 in the multimodal knowledge-driven drawing large model adjustment method according to an embodiment of the present application.
[0012] Figure 4 The figure shows a schematic flow chart of step S5 in the large drawing model adjustment method driven by multimodal knowledge according to an embodiment of the present application.
[0013] Figure 5 The figure shows a schematic block diagram of a large drawing model adjustment system based on multimodal knowledge driving according to an embodiment of the present application. DETAILED DESCRIPTION
[0014] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0015] Figure 1 FIG2 shows a schematic flow chart of a method for adjusting a large drawing model based on multimodal knowledge-driven drawing according to an embodiment of the present application. Figure 1 As shown, the present application provides a drawing large model adjustment method driven by multimodal knowledge, including: S1, obtaining a text prompt input by a user; S2, performing text parsing and semantic embedding encoding on the text prompt to obtain a text prompt structured intention representation and a text prompt standard condition embedding; S3, performing multimodal knowledge retrieval based on the text prompt structured intention representation to obtain a multimodal knowledge representation; S4, inputting the text prompt standard condition embedding and the multimodal knowledge representation into a knowledge-attention translation module to obtain attention modulation parameters; S5, inputting the text prompt standard condition embedding and the randomly generated initial noise latent representation into the drawing large model to obtain a final denoised latent representation, wherein the attention modulation parameters are used to influence the parameters of the attention layer of the denoising U-Net network of the drawing large model; S6, performing feature decoding on the final denoised latent representation to obtain a generated image.
[0016] For example, in step S1, a text prompt input by the user is obtained. It should be understood that natural language text is widely accepted as an effective medium for expressing complex ideas and describing image content and style. Through text prompts, users can freely outline the image concept in their minds in the language they understand, whether it is a specific entity, an abstract scene, a specific emotional atmosphere, or detailed requirements for visual effects. Therefore, this step ensures that the model can receive the user's clear intention instructions, making the generation process directional and goal-oriented, and is the basis for achieving personalized and customized image generation.
[0017] In a specific embodiment, a user enters the desired text description through a graphical user interface, such as a text input box in a web application, a text input area in desktop software, or a virtual keyboard in a mobile application. When the user completes the input and triggers the generation action, for example, by clicking a "Generate" button or pressing the Enter key, the text string is captured on the client. Subsequently, the captured text data is securely transmitted to the drawing model service backend deployed on a remote server in the form of a request body, such as form data or JSON payload in a POST request, via a standard network communication protocol. On the server side, an application program interface or service module that receives the request listens to and receives the text data stream from the client.
[0018] For example, in step S2, the text prompt is subjected to text parsing and semantic embedding encoding to obtain a structured intent representation of the text prompt and a standard conditional embedding of the text prompt. It should be understood that when a user uses a text to describe a desired image, for example: a red sports car driving on a tree-lined path on a rainy night, in the style of a watercolor painting, with city neon lights reflected on the car windows, it contains many discrete semantic units that need to be accurately identified. For example, the sports car is the subject, red is its color attribute, the rainy night and the tree-lined path are scene constraints, driving is the subject action, the watercolor painting is the style, and the city neon lights are the specific content of the reflection. If this description is directly input into the model without rigorous text parsing, the model will find it difficult to accurately identify these independent and important information fragments, let alone understand the complex hierarchical relationships and interactions between them. Traditional text-to-image models, when dealing with such complex, delicate, or knowledge-required generation tasks, often fail to fully and structure the original text information, resulting in a deviation from the user's true intent and even hallucinations. Therefore, the text prompt is subjected to text parsing and semantic embedding encoding to obtain a text prompt structured intent representation and a text prompt standard condition embedding. Among them, through text parsing, the unstructured natural language text is converted into a machine-readable, clear and logically hierarchical text prompt structured intent representation, aiming to convert the user's vague intention into precise instructions. This structured intent representation solves the ambiguity problem of natural language, allowing the subsequent multimodal knowledge retrieval link to use these structured key entities and attributes as highly accurate query conditions.
[0019] At the same time, considering that deep learning models cannot directly process original text strings, they require numerical, high-dimensional vector input. Therefore, the text prompt is encoded to compress the semantic information of the text and map it into a continuous vector space. In this vector space, semantically similar texts are usually geometrically represented by vectors with close distances, while texts with large semantic differences have vectors with longer distances. This numerical representation enables text information to be effectively calculated and utilized by neural networks. Specifically, in this application, CLIP Text Encoder is selected as a pre-trained text encoder. CLIP Text Encoder is pre-trained by comparative learning on a large-scale image-text pair dataset, so that the text embedding learned by its text encoder and the image embedding learned by the image encoder can be in the same shared, semantically aligned latent space. This means that the standard conditional embeddings of text prompts obtained by CLIP Text Encoder not only capture the pure semantics of the text itself, but more importantly, these embeddings are naturally integrated with visual semantic perception capabilities. When these text embeddings rich in visual semantic information are input as conditions into the drawing model, the model can more effectively and intuitively associate them with the visual features of the image, thereby accurately guiding the denoising U-Net network to generate visual content that is highly consistent with the text description.
[0020] That is, in one embodiment, if Figure 2 As shown, the text prompt is subjected to text parsing and semantic embedding encoding to obtain a text prompt structured intent representation and a text prompt standard condition embedding, including: S21, performing text parsing on the text prompt to extract key entities therefrom, and constructing a text prompt structured intent representation in JSON format based on the key entities; S22, inputting the text prompt into a pre-trained text encoder to obtain the text prompt standard condition embedding, and the pre-trained text encoder is CLIP Text Encoder.
[0021] Specifically, the text prompt is parsed to extract key entities, and a structured intent representation of the text prompt in JSON format is constructed based on these key entities. This involves: First, the original text undergoes a series of preprocessing steps, such as removing irrelevant symbols and standardizing text case. Then, a tokenizer is used to break it down into its smallest semantic units, namely, words. Next, part-of-speech tagging is performed to identify each word after tokenization, such as noun, verb, or adjective. This provides fundamental linguistic information for subsequent entity recognition and relation extraction. Pre-trained NER models are then used to identify entities with specific meanings in the text, such as specific objects like "sports car" and "city neon," colors like "red," scene elements like "tree-lined path" and "rainy night," and styles like "watercolor painting." These models, trained on a large amount of data, can accurately extract these important concepts from unstructured text. After identifying the key entities, adjectives or adverbs that modify them are further identified. These attributes are used to enrich the entity description. Furthermore, relation extraction and dependency syntactic analysis are performed to understand the logical relationships and syntactic structures between entities in the text. For example, in the sentence "The sports car is driving on the tree-lined path", the system needs to recognize that "sports car" is the subject, "tree-lined path" is the location, "driving" is the action occurring between the two, and the color of "sports car" is "red". This processing can be done by parsing the sentence structure through dependency syntactic analysis, or by using a specialized relation extraction model to identify predefined relation types such as "located in", "having" or "perform". Finally, after all key entities, attributes and relationships are extracted, a domain ontology is generated based on the predefined image, and all this discrete information is organized and mapped into a specific JSON format to form a clear, unambiguous and hierarchical text prompt structured intent representation. For example, for the previously mentioned text prompt "A red sports car driving on a tree-lined path on a rainy night, in the style of a watercolor painting, with city neon lights reflected on the windows", a structured JSON representation can be constructed that clearly depicts the core object in the user's intent, its attributes, the environment it is in, the action being performed, and the overall artistic style, even including the details of the reflection, thereby completely eliminating the ambiguity of natural language.
[0022] The complete original text prompt string entered by the user is fed directly into the pre-trained CLIPText Encoder. The CLIP Text Encoder includes a specialized tokenizer, which typically uses algorithms such as byte pair encoding to break the original text into a series of tokens from a predefined vocabulary. To meet the input requirements of the Transformer model, special start, end, and padding tokens are automatically added to ensure that all input sequences have uniform lengths and that the beginning and end of the sequence can be identified. After each token is broken down, its corresponding initial word embedding vector is retrieved through the model's lookup table. Given that the Transformer architecture itself lacks the ability to handle sequence order, these word embeddings are combined with positional encodings to explicitly incorporate information about the relative or absolute position of words in the text, ensuring that the model understands the impact of word order on semantics. This sequence of word embeddings, incorporating positional information, is then fed into the CLIP Text Encoder's deep neural network, consisting of multiple stacked Transformer blocks. Each Transformer block typically includes a multi-head self-attention mechanism and a feedforward network, working together. By iteratively processing these layers, the model is able to compute the complex interrelationships between different words in the sequence and recursively update word embeddings based on these relationships. This means that ultimately, each word's embedding vector carries information about its contextual semantics within the entire sentence. Next, a pooling strategy is employed to obtain a single, fixed-dimensional vector representation of the entire text prompt. Finally, the embedding vector corresponding to the special end token in the sequence is directly extracted as a representation of the entire text prompt, which is the standard conditional embedding of the text prompt.
[0023] Exemplarily, in step S3, multimodal knowledge retrieval is performed based on the text prompt structured intent representation to obtain a multimodal knowledge representation. It should be understood that existing image generation methods, especially those diffusion models that rely purely on text conditions and implicit knowledge, although they perform well in generating diverse and high-quality images, rely solely on text prompts, and the model's control over the generated content is relatively coarse. When the user wants a local feature or overall style of the generated image to be highly consistent with an existing visual paradigm, traditional methods are difficult to accurately achieve through simple text prompts. The introduction of multimodal knowledge retrieval can provide the model with more specific visual anchors and references, thereby achieving more precise control over the generation process. Specifically, multimodal knowledge retrieval aims to make up for the lack of implicit knowledge of the model by accessing an external, structured, and explicit knowledge base, providing the model with richer and more accurate semantic and visual information, thereby improving the accuracy, detail richness, and controllability of the generated image, and ultimately producing high-quality images that meet user expectations and objective facts.
[0024] In one embodiment, Figure 3 As shown, multimodal knowledge retrieval is performed based on the text prompt structured intention representation to obtain the multimodal knowledge representation, including: S31, using the text prompt structured intention representation as a query, retrieving knowledge fragments related to the query in the multimodal knowledge base, the knowledge fragments including structured related knowledge facts and related visual feature vectors; S32, aggregating the knowledge fragments to obtain the multimodal knowledge representation.
[0025] First, using the structured intent representation of the text prompt as a query, a multimodal knowledge base is retrieved for knowledge fragments related to the query. These knowledge fragments include structured relevant knowledge facts and related visual feature vectors. It should be understood that a multimodal knowledge base is a large database integrating multimodal information, containing various structured knowledge, image / video libraries, visual feature embeddings, and multimodal association indexes. For ease of understanding, we will further explain that structured knowledge refers to the various entities, attributes, and their interrelationships stored in the multimodal knowledge base. For example, "Einstein" is a "physicist" whose "major achievement" is "relativity"; "sports cars" have "streamlined bodies" and the "brand" is "Ferrari." These facts are typically stored in the form of triples (entity-relationship-entity) or key-value pairs and are indexed for query purposes. The image / video library refers to the image and video data stored in the multimodal knowledge base, containing various entities and scenes, and is associated with entities in the structured knowledge. Visual feature embedding involves extracting high-dimensional visual feature vectors for each image or video frame in the library using a powerful pre-trained visual encoder. These vectors are dense representations of the image content, capturing the semantics and style of the image in the latent space. These visual feature vectors are typically stored in a vector database for efficient approximate nearest neighbor search. Multimodal association indexing means that the knowledge base establishes association indexes between entities, text descriptions, images, and their visual feature vectors, so that when a query for an entity is made, the relevant structured facts, text descriptions, and visual references can be retrieved simultaneously.
[0026] First, the core entities (such as "sports car", "tree-lined path") and attributes (such as "red") in the structured intent representation of the text prompt are identified. Using these entities and attributes as keywords, queries are performed in the structured knowledge to retrieve all related structured knowledge facts (for example, about the general shape, common brands, and performance characteristics of sports cars; about the plant types and lighting conditions of tree-lined paths, etc.). These are indexed knowledge fragments. Based on the identified entities and attributes, semantic matching and visual feature similarity searches are performed in the image / video part of the multimodal knowledge base to obtain relevant visual feature vectors. In a specific embodiment, a pre-trained text encoder is used to encode the key descriptions in the structured intent representation of the text prompt again into a query vector. Then, an approximate nearest neighbor search is performed in the pre-calculated visual feature vector library to find visually related images and their corresponding visual feature vectors. For example, "red sports car" is encoded as a query vector, and visually similar "red sports car" pictures and their feature vectors are retrieved in the image vector database.
[0027] Next, considering that retrieved knowledge fragments are often discrete and diverse, potentially containing multiple facts, multiple images, and multiple visual feature vectors, aggregation is necessary to integrate this scattered information into a unified, concise representation that can be used by subsequent modules. The goal of aggregation is to generate a high-dimensional vector that captures the core semantics and visual cues of all relevant knowledge fragments.
[0028] Specifically, for the retrieved structured relevant knowledge facts (for example, (sports car, has attributes, streamlined)), they first need to be encoded into vector representations. This can be done by converting the facts into natural language phrases (for example, "sports car has a streamlined body") and then encoding them through a text encoder such as BERT or Sentence-BERT. If the nodes and edges in the knowledge graph themselves have pre-trained embeddings, these embeddings can be extracted directly. Then, the vectors of all these facts can be averaged, summed, or fused through a Transformer layer to obtain a vector representing all structured knowledge. At the same time, multiple relevant visual feature vectors may be retrieved. For these visual feature vectors, the following methods can be used for fusion, for example: Average pooling: take the average of all relevant visual feature vectors to obtain a vector representing the overall visual clues. Attention weighting: introduce a small attention network to assign weights to each visual feature vector based on its relevance to the original query, and then perform weighted summation. This enables the model to focus more on the visual references that are most relevant to the current intent.
[0029] Finally, the fused text fact vector and the fused visual feature vector are further combined to form the final multimodal knowledge representation. In one specific embodiment, the fused text fact vector and the fused visual feature vector can be directly concatenated to form a longer vector. Alternatively, the concatenated vector can be nonlinearly transformed through one or more fully connected layers to learn the complex interactions between different modalities.
[0030] Exemplarily, in step S4, the text prompt standard conditional embedding and the multimodal knowledge representation are input into the knowledge-attention translation module to obtain attention modulation parameters. It should be understood that although the text prompt standard conditional embedding and the multimodal knowledge representation each provide valuable guidance for the drawing model, they themselves are not in a format that can be directly used to adjust the parameters of the attention layer within the model. The denoising U-Net network of the drawing model plays a core role in the image generation process, in which the cross-attention mechanism is responsible for accurately integrating external conditions (such as text features) into the layer-by-layer denoising process of the image potential features. However, if the text embedding is directly used as the source of the key and value vectors of the cross-attention layer, this traditional approach can provide rough conditional guidance, but it is powerless when faced with the need for fine knowledge control, handling complex scenes, or correcting hallucinations that may occur in the model. Traditional text diffusion models often find it difficult to accurately capture details that are not explicitly mentioned in the text but can be inferred through knowledge, or are prone to misunderstandings when generating specific, uncommon entities.
[0031] Based on this, in the technical solution of this application, the text prompt standard conditional embedding and the multimodal knowledge representation are input into the knowledge-attention translation module to obtain attention modulation parameters. In other words, the text prompt standard conditional embedding and the multimodal knowledge representation are received and deeply integrated, understood, and transformed to produce a set of parameters specifically used to modulate the behavior of the attention layer within the large drawing model, namely, attention modulation parameters. These parameters are not direct attention weights, but rather provide a learnable, dynamic bias or scaling factor. They will affect the generation of key and value vectors in the attention layer, or directly intervene in the calculation of attention scores, thereby inducing the model to more selectively and knowledge-awarely focus on the features and regions in the latent space that are most relevant to the user's intent and retrieved external knowledge. For example, a user wishes to generate an image of a "rare butterfly with a specific wing pattern." Based solely on the text embedding, the model may only be able to generate a general butterfly. However, if the multimodal knowledge retrieval process successfully obtains visual references and structured descriptions of the wing pattern of this specific butterfly species, the knowledge-attention translation module can convert this unique and precise knowledge information into attention modulation parameters. During the denoising process, these parameters will guide the model to focus on and draw areas that match the pattern of the specific butterfly's wings, rather than arbitrary patterns, thereby significantly improving the knowledge correctness, detail richness and fit of the generated images with user expectations.
[0032] In one embodiment, the text prompt standard condition embedding and the multimodal knowledge representation are input into a knowledge-attention translation module to obtain attention modulation parameters, including: inputting the text prompt standard condition embedding and the multimodal knowledge representation into a Transformer network to obtain the attention modulation parameters.
[0033] In one specific embodiment, a Transformer network is composed of multiple identical layers stacked together. Each layer incorporates a multi-head self-attention mechanism, a feedforward network, residual connections, and layer normalization. The multi-head self-attention mechanism is the core of the Transformer. It allows the model to simultaneously weigh the importance of different parts of the sequence when processing an input vector sequence and learn the complex relationships between them. By converting the input vector into query, key, and value representations and calculating the similarity between them, the self-attention mechanism computes a weighted sum for each position in the sequence, creating a more context-aware fused representation. The key role of the self-attention mechanism is that it learns to effectively fuse the abstract intent expressed in the text prompt with the specific details and facts contained in multimodal knowledge. For example, if the text prompt refers to the general term "bird," if the multimodal knowledge representation contains detailed knowledge about "a specific species of bird with distinctive blue feathers and a long beak," the self-attention mechanism can flexibly allow the semantics of the text embedding to "focus" and incorporate the key visual features of "blue feathers" and "long beak" in the knowledge representation, thereby forming a unified and more precise semantic understanding. The multi-head setup allows the model to simultaneously fuse information from different perspectives and focus points, improving its expressive power. Next, the output of the self-attention mechanism is further fed into a feedforward network with a fully connected layer (typically two layers with activation functions). This network independently applies nonlinear transformations to the representation of each position in the sequence to capture more complex feature interactions and enhance its expressive power. To avoid the vanishing gradient problem of deep networks and accelerate training, the output of each sublayer is added to the input via a residual connection, followed by layer normalization. This helps stabilize training and improve the model's generalization.
[0034] After iterative processing through multiple Transformer layers, the original standard conditional embedding of the text prompt and the initial information of the multimodal knowledge representation are continuously fused, refined, and semantically enhanced to produce a vector output at the final layer. This output vector cannot directly serve as the parameters for the attention layer of the large drawing model's denoising U-Net network; it must be decoded into attention modulation parameters. This is achieved through one or more linear projection layers, which are essentially fully connected neural networks. These linear layers map the Transformer output vector to a space that matches the dimensions and form required by the attention modulation parameters.
[0035] Exemplarily, in step S5, the text prompt standard conditional embedding and the randomly generated initial noise latent representation are input into the large drawing model to obtain the final denoised latent representation. The attention modulation parameters are used to influence the parameters of the attention layer of the denoising U-Net network of the large drawing model. It should be understood that the operating principle of the large drawing model, and in particular the image generation system based on the diffusion model, is essentially an iterative denoising process. This process begins with a completely disordered, seemingly random initial noise latent representation. This random noise can be viewed as a starting point full of infinite possibilities, with the potential to generate any image, but lacking directionality and specific semantics. Therefore, the introduction of the randomly generated initial noise latent representation is the starting point for the diffusion model to generate diverse images. It provides randomness and exploration space in the generation process, ensuring that each generation is not exactly the same, but rather possesses a certain degree of creativity and randomness under given conditions. However, in order to guide this infinite possibility towards the specific image content desired by the user, the model must be subject to clear conditional constraints to gradually shape the characteristics of the target image during the denoising process.
[0036] Here, we introduce standard conditional embeddings based on textual prompts to provide macro-semantic guidance for the denoising process, ensuring that while removing noise, the model gradually aligns image features with the initial semantics described by the text. For example, when the text prompt is "a cute cat," the text embedding can guide the model to generate an image of a cat, rather than a dog or a car.
[0037] However, traditional large-scale drawing models, guided solely by standard conditional embeddings based on textual prompts, still have significant limitations. They often fail to accurately capture details not explicitly mentioned in the text but requiring background knowledge to understand. They are also prone to hallucinations when encountering complex, concrete entities, generating images that are inconsistent with reality or common sense. For example, if a user asks to generate "a red sports car driving on a tree-lined path on a rainy night, in the style of a watercolor painting, with city neon lights reflected in the car windows," a model relying solely on generalized textual embeddings may struggle to correctly understand the complex physical relationship of reflection and the specific artistic style of watercolor painting. The model may mistakenly render neon lights as stickers on the car windows or body, rather than as distorted, dynamic light spots on the wet, curved glass. Furthermore, it may fail to harmoniously integrate the hard metallic texture of the sports car with the soft, diffuse qualities of the watercolor painting style, resulting in a visually disjointed image. This hallucination, stemming from a lack of knowledge of physical optics and artistic principles, not only reduces the authenticity and artistic value of the generated image but also makes it difficult for users to obtain images that meet their specific creative needs through simple textual prompts. In order to remove the hallucination phenomenon and improve the control granularity, the attention modulation parameter is introduced to influence the attention layer of the denoising U-Net network of the large drawing model.
[0038] Specifically, the core of the large-scale drawing model is a powerful denoising U-Net network, which plays a crucial role in the reverse process of the diffusion model. The entire image generation is an iterative process, starting with a completely disordered and seemingly random latent representation of an initial noisy image. At each step, the U-Net predicts the noise present in the current noisy image and then subtracts the predicted noise from the image, gradually making the image clearer, ultimately converging to the latent representation of the target image.
[0039] The U-Net network, named for its unique shape, consists of an encoder (also known as the downsampling path) and a decoder (also known as the upsampling path), connected by sophisticated skip connections, allowing them to work together. The encoder is responsible for gradually extracting image features. It consists of a series of convolutional layers and downsampling layers (such as max pooling or strided convolution). At each layer, the spatial resolution of the feature map decreases, for example, from 64x64 to 32x32, to 16x16, and even lower, while the number of feature channels increases. This design enables the encoder to capture local details and increasingly abstract high-level semantic information at different scales, transforming the image's raw pixel information into a more compact and meaningful feature representation. The decoder is responsible for gradually restoring the spatial details of the image based on the features extracted by the encoder. It consists of a series of upsampling layers (such as transposed convolution or interpolation followed by convolution) and convolutional layers. At each layer, the spatial resolution of the feature map increases, for example, from low-resolution features to high-resolution features, while the number of feature channels decreases to reconstruct the image's visual structure. Skip connections are a key design feature of the U-Net. At each level of the decoder path, feature maps from the encoder path at the corresponding resolution are directly connected (typically concatenated or added) to the decoder path's input. These skip connections are crucial because they ensure that the high-resolution, fine details captured by the encoder in earlier stages of image restoration are not lost during downsampling, allowing them to be effectively reused in the decoder, resulting in clearer, more textured, and more accurately detailed images.
[0040] At each level of the U-Net network, external conditions need to be injected into it to guide the denoising process and ensure that the generated results meet the user's specific requirements. In this application, text conditional injection is performed, where the file condition refers to the text prompt standard conditional embedding, which is achieved through the cross-attention layer embedded in the U-Net. Specifically, the denoising U-Net network of the large drawing model includes an attention layer. In the encoder and decoder paths of the U-Net network, especially in its deep and middle layers, multiple attention blocks are strategically embedded. These attention blocks usually include self-attention mechanisms and cross-attention mechanisms. Cross-attention is responsible for introducing external conditions, that is, text prompt standard conditional embedding, into the calculation of image potential features.
[0041] In one embodiment, Figure 4 As shown, the text prompt standard condition embedding and the randomly generated initial noise potential representation are input into the drawing large model to obtain the final denoised potential representation, including: in an iterative process of denoising using the denoising U-Net network of the drawing large model: S51, extracting intermediate potential features from the previous layer of the denoising U-Net network; S52, feature flattening the intermediate potential features to obtain an intermediate potential feature flattened vector; S53, projecting the intermediate potential feature flattened vector to the query space to obtain an intermediate potential feature query vector; S54, mapping the text prompt standard condition embedding to the key space and the value space to obtain a text prompt key vector and a text prompt value vector; S55, modulating the text prompt key vector and the text prompt value vector with the attention modulation parameter as a bias item to obtain a modulated text prompt key vector and a modulated text prompt value vector; S56, calculating the attention weight based on the intermediate potential feature query vector, the modulated text prompt key vector and the modulated text prompt value vector; S57, weighted summing the modulated text prompt value vector based on the attention weight to obtain the output feature.
[0042] Specifically, at each denoising time step, starting with a randomly generated initial noisy latent representation, this highly noisy latent representation is input into the U-Net's encoder path. This latent representation is progressively processed through the various layers of the U-Net, where the U-Net attempts to predict the noise given the current noise level. Simultaneously, the standard conditional embedding of the text prompt serves as important conditional information and is fed into multiple cross-attention layers within the U-Net. In these cross-attention layers, the intermediate latent features generated by the U-Net itself, representing the current state of the image, serve as query vectors. The generation of the key and value vectors is determined by external conditions. Traditionally, the key and value vectors are directly derived from linear projections of the standard conditional embedding of the text prompt.
[0043] In the present application, the attention modulation parameter plays a core role here. The attention modulation parameter generated by the knowledge-attention translation module according to the text prompt standard condition embedding and multimodal knowledge representation will act in the form of a bias term on the modulated text prompt key vector and the modulated text prompt value vector generated by the text prompt standard condition embedding after linear projection. Specifically, the attention modulation parameter is used as a bias term to modulate the text prompt key vector and the text prompt value vector to obtain the modulated text prompt key vector and the modulated text prompt value vector, including: using the attention modulation parameter as a bias term to modulate the text prompt key vector and the text prompt value vector according to the following formula, the formula is:
[0044] K modulated =K std +Δ k
[0045] V modulated =V std +Δ k
[0046] Among them, K std Represents the text prompt key vector, V std represents the text prompt value vector, Δ k represents the attention modulation parameter, K modulated Represents the modulation text prompt key vector, V modulated Represents a vector of modulated text prompt values.
[0047] This means that the key and value vectors generated from the standard conditional embedding of the text prompt are no longer fixed, but are instead adjusted by a precise attention modulation parameter that acts as a bias term. This bias term carries precise details from external multimodal knowledge. When the U-Net's attention layer calculates the attention weights, the intermediate latent feature query vector is dot-producted with the modulated text prompt key vector. This attention weight is then applied to the modulated text prompt value vector. This modulation process enables the attention layer to calculate the correlation between image features and text conditions, rather than relying solely on the generalized semantics of the text. Instead, it can precisely direct attention to specific visual features or concepts that are explicitly guided by multimodal knowledge. This effectively places a label on the model's attention focus, allowing the model to pay more attention to knowledge-related regions or attributes, thereby prioritizing and accurately reproducing these knowledge details during the generation process.
[0048] Specifically, calculating the attention weight based on the intermediate potential feature query vector, the modulated text prompt key vector, and the modulated text prompt value vector includes: calculating the attention weight based on the intermediate potential feature query vector, the modulated text prompt key vector, and the modulated text prompt value vector using the following formula, wherein the formula is:
[0049]
[0050] Wherein, a represents the attention weight, Q represents the intermediate potential feature query vector, d is the dimension of the modulated text prompt key vector, Softmax represents the Softmax function, and T represents the vector transpose.
[0051] Here, after mapping the text prompt standard conditional embedding to the key space and the value space to obtain the text prompt key vector and the text prompt value vector, bias is further performed based on the attention modulation parameters obtained through context-related regression based on the text prompt standard conditional embedding and the multimodal knowledge representation. It can be seen that the attention modulation parameters, as a common bias benchmark for knowledge-related coupling, essentially provide a benchmark framework for the distribution morphology measurement of the text prompt key vector and the text prompt value vector, thereby realizing the introduction of external knowledge. However, only through this simple linear bias operation, although the signal of external knowledge is injected into the key vector and the value vector, it may further lead to key-value space topology mismatch. In the context of the attention mechanism, the functions and topological structures of the key vector space and the value vector space are essentially different. "The key vector space is mainly responsible for responding to queries and calculating similarity scores, while the value vector space carries the content information to be aggregated. When a unified attention modulation parameter generated by external knowledge is applied equally to these two functionally heterogeneous spaces, it may destroy the original subtle correspondence between them learned in pre-training. For example, when generating the complex scene of "a red sports car driving on a tree-lined path on a rainy night, in the style of a watercolor painting, with city neon reflected on the car window", the "city neon" introduced by external knowledge is not necessarily the same as the original one. The modulation signal may effectively adjust the key vector associated with the "window surface", but when adjusting its corresponding value vector, it may not be able to accurately retain the subtle representation of the smooth paint of the "red sports car", resulting in a visually inconsistent "stitching" of the generated neon lights and the car body, as if the neon lights were stickers affixed to the car body, rather than light and shadow reflected on the glass. This is a concrete manifestation of the key-value space topology mismatch problem, that is, the key vector has moved to a new semantic position, but its associated value vector has not undergone a structural adjustment to match it in terms of content expression. Therefore, it is preferred to optimize the modulated text prompt key vector and the modulated text prompt value vector.
[0052] In one embodiment, the attention weight is calculated based on the intermediate potential feature query vector, the modulated text prompt key vector and the modulated text prompt value vector, including: performing key-value space topology matching optimization on the modulated text prompt key vector and the modulated text prompt value vector to obtain an optimized modulated text prompt key vector and an optimized modulated text prompt value vector; and calculating the attention weight based on the intermediate potential feature query vector, the optimized modulated text prompt key vector and the optimized modulated text prompt value vector.
[0053] Specifically, performing key-value space topology matching optimization on the modulated text prompt key vector and the modulated text prompt value vector to obtain an optimized modulated text prompt key vector and an optimized modulated text prompt value vector includes:
[0054] First, calculate the spatial topological gradient field matrix between the modulated text prompt key vector and the modulated text prompt value vector. Specifically, based on the modulated text prompt key vector V a Each eigenvalue a of i and the modulated text prompt value vector V b Each eigenvalue b of j , based on the Euclidean space distance representation to obtain the spatial topological gradient field, that is:
[0055]
[0056] Among them, a i Represents each eigenvalue of the modulated text prompt key vector, b j Represents each eigenvalue of the modulated text prompt value vector, m i,j Represents the eigenvalue of the (i, j) position in the spatial topological gradient field matrix.
[0057] That is, by introducing the explicit metric representation of the spatial topological gradient in the Euclidean space, the spatial topological gradient field matrix M is constructed based on the spatial distribution characteristics of the text prompt key vector and the text prompt value vector, and m i,j ∈M. The spatial topological gradient field matrix M is represented by the gradient of the partial derivative of the spatial distance, providing a spatial spectral topological measurement reference framework that includes both spatial curvature and spatial resolution relationships. Specifically, it can be understood that the spatial topological gradient field matrix describes how a small change in a point in one space causes a change in the corresponding point in another space, capturing the distribution pattern and local spatial curvature between the two. For the aforementioned "sports car on a rainy night" scene, this gradient field matrix can depict how the concept of "neon reflection" affects the key-value correspondence of the various components of the "sports car" (such as windows, paint, and tires), providing a global topological reference framework for subsequent calibration.
[0058] Based on the modulation text prompt key vector V a and the modulated text prompt value vector V b The variance σ a and σ b , the spatial topology gradient field matrix M is subjected to eigenvalue granularity spatial pedigree cross reconstruction to obtain the spatial topology pedigree reconstruction gradient field matrix, namely:
[0059]
[0060] Where M represents the spatial topological gradient field matrix, σ a represents the variance of the modulated text prompt key vector, σ b represents the variance of the modulated text prompt value vector, and M' represents the spatial topology spectrum reconstruction gradient field matrix.
[0061] However, simply having a spatial topological gradient field matrix is not enough, because the high-dimensional feature space itself is sparse, and the spatial topological gradient field matrix may be biased due to data noise or feature discreteness. Therefore, based on the variance of the modulated text prompt key vector and the modulated text prompt value vector, the above spatial topological gradient field matrix is reconstructed at the eigenvalue granularity of the spatial spectral cross-reconstruction to obtain a spatial topological spectral reconstructed gradient field matrix. It should be understood that in order to solve the measurement inaccuracy problem caused by the discreteness of the feature space, a statistical importance consideration is introduced. The variance of a vector can be regarded as a reflection of the amount of information it contains. Feature dimensions with large variances generally carry more critical semantic information. By incorporating the variance information of the key and value vectors into the reconstruction process of the gradient field matrix, it is equivalent to performing an information content-based correction on the spatial topological gradient field matrix, so that the construction of the topological relationship focuses more on the core feature dimensions that have a significant impact on the final image generation. Specifically, this means that the model will pay more attention to the correct topological association between high-variance features such as the sharp outline of "sports car", the bright colors of "neon lights", and the transparent texture of "car windows", while smoothing out some minor low-variance details that may cause noise (such as random brushstrokes in watercolor style), thereby obtaining a more robust and accurate topological relationship representation.
[0062] Finally, the modulated text prompt key vector and the modulated text prompt value vector are respectively mapped to the feature space of the spatial topology spectrum reconstruction gradient field matrix to obtain the optimized modulated text prompt key vector and the optimized modulated text prompt value vector, that is:
[0063] V' a =M'×V a
[0064] V'b=M'×V b
[0065] Among them, V a Represents the modulated text prompt key vector, V b represents the modulated text prompt value vector, M' represents the spatial topology spectrum reconstruction gradient field matrix, V' a represents the optimized modulation text prompt key vector, V' b Represents the optimized modulated text prompt value vector.
[0066] By mapping the original modulated text prompt key vector and the modulated text prompt value vector respectively into the feature space of the reconstructed gradient field matrix, the original modulated text prompt key vector and the modulated text prompt value vector are made to follow this newly established, ideal topological correspondence. Through this mapping, the key-value mismatch problem that may have existed originally is corrected, and cross-dimensional mismatch compensation of the interaction between the global space spectrum topology and the eigenvalue granularity is achieved. This means that the optimized modulated text prompt key vector and the optimized modulated text prompt value vector after mapping will be more consistent in semantics and structure, thereby improving the key-value space topological mismatch problem of the modulated text prompt key vector and the modulated text prompt value vector, and improving the calculation accuracy of the attention weight.
[0067] This cross-attention mechanism with modulation is repeated at different levels of the U-Net throughout the entire image denoising process. In the U-Net encoder path, the attention mechanism helps the model more effectively extract features relevant to the conditions, particularly knowledge details, during downsampling, ensuring that key knowledge information is preserved and enhanced even in the abstract latent space. In the decoder path, it guides the model to fine-tune and restore image content consistent with the conditions (including knowledge details) during upsampling and feature fusion, ensuring that the final image accurately reflects knowledge at the pixel level. Through this iterative and layered modulation, the denoising process is no longer blind but instead knowledge-driven. For example, if the knowledge modulation parameter indicates the need for "wheel details of a specific sports car model" or "a specific gradient color of a rare bird's feathers," the U-Net's attention mechanism during the denoising process will be guided by this bias term, preferring to render wheel or feather regions with a design and texture specific to that model, or a color gradient unique to that bird, rather than a generic wheel or random feather color. This fine-grained control allows the model to avoid hallucinations and produce images that are highly accurate both visually and intellectually.
[0068] Finally, after multiple rounds of iterative denoising, each step is guided by highly precise text conditioning and knowledge modulation parameters. This process of U-Net predicting noise and subtracting it from the current latent representation is repeated many times until the preset minimum noise level is reached. The initial random noise latent representation is gradually transformed into a final denoised latent representation. This final denoised latent representation has condensed all the user's intentions and accurately reflects the specific details in the external multimodal knowledge, preparing for subsequent image decoding. This implementation method, by using knowledge as a learnable and dynamic bias to fine-tune the model's attention calculation level, solves the problem of external knowledge being difficult to effectively inject into the core generation process of the model in traditional methods, significantly improving the knowledge consistency, detail accuracy and user intention matching of the generated images, and represents an important breakthrough in the field of text-to-image generation.
[0069] Exemplarily, in step S6, the final denoised latent representation is feature-decoded to obtain a generated image. It should be understood that while the final denoised latent representation processed and manipulated in the latent space contains all the semantic and visual information required to generate the image, it is not a pixel-level image, but rather a high-dimensional, abstract numerical matrix. The human visual system cannot directly understand this abstract representation; therefore, it must be converted back into a displayable image form, such as an RGB pixel matrix, through a decoder before it can be viewed, evaluated, and applied by the user.
[0070] In a specific embodiment, the final denoised latent representation is feature-decoded using a specialized decoder network to produce the generated image. The decoder network is configured to receive a low-resolution, high-channel-count latent representation and progressively upsample it to convert it into a standard high-resolution, low-channel-count image. In an embodiment of the present application, the decoder network is the decoder portion of a pre-trained variational autoencoder.
[0071] Specifically, the decoder portion of the pre-trained variational autoencoder comprises a series of upsampling and convolutional layers. The data processing process involves first inputting the final denoised latent representation into the initial layer of the VAE decoder. This latent representation typically has lower spatial resolution but a higher number of feature channels. The decoder network then gradually increases the spatial resolution of the latent representation through multiple layers of upsampling. Upsampling can be implemented using transposed convolutions, also known as fractionally strided convolutions, or by increasing the resolution through methods such as nearest neighbor interpolation and bilinear interpolation before performing convolution. These upsampling layers not only expand the spatial size of the feature map but also map the abstract features in the latent representation to more concrete visual elements through the learned convolution kernels. With each upsampling layer, the number of feature channels typically decreases as the spatial resolution increases. For example, starting with a latent representation with four channels, after several layers of processing, the final result is a feature map with three channels, corresponding to the red, green, and blue color channels of an RGB image. The last layer of the decoder is usually a convolutional layer with 3 output channels, and uses an activation function to limit the pixel value range to a specific interval to match the pixel representation range of the image. This final convolutional layer converts high-dimensional feature information into pixel values that can be directly displayed on the screen.
[0072] In summary, the present application provides a method for adjusting a large drawing model based on multimodal knowledge-driven drawing. It first obtains the text prompts input by the user, performs in-depth analysis on them, and extracts a structured representation of the user's intention. Subsequently, based on this structured intention, it actively retrieves and integrates external multimodal knowledge, thereby introducing richer and more accurate semantic information and factual basis for the generation process. Then, a knowledge-attention translation module is introduced, which can convert this rich multimodal knowledge into refined attention modulation parameters. These parameters will directly and dynamically affect the working mode of the attention mechanism inside the large drawing model, realizing knowledge-driven fine-grained feature generation control. In this way, complex user intentions can be understood more accurately, the "hallucination" phenomenon can be avoided, and images that are highly consistent with external knowledge and have more accurate details can be generated, greatly improving the quality and controllability of the generated images.
[0073] The present application also provides a drawing large model adjustment system based on multimodal knowledge driving, which is used to execute the above-mentioned drawing large model adjustment method based on multimodal knowledge driving, such as Figure 5As shown, the multimodal knowledge-driven drawing model adjustment system 500 includes: a text prompt acquisition module 510, which is used to obtain a text prompt input by a user; a text semantic parsing and embedding module 520, which is used to perform text parsing and semantic embedding encoding on the text prompt to obtain a text prompt structured intention representation and a text prompt standard condition embedding; a multimodal intention-driven retrieval module 530, which is used to perform multimodal knowledge retrieval based on the text prompt structured intention representation to obtain a multimodal knowledge representation; a multi-source attention modulation module 540, which is used to input the text prompt standard condition embedding and the multimodal knowledge representation into a knowledge-attention translation module to obtain attention modulation parameters; a text-guided attention denoising module 550, which is used to input the text prompt standard condition embedding and the randomly generated initial noise latent representation into the drawing model to obtain a final denoised latent representation, wherein the attention modulation parameters are used to influence the parameters of the attention layer of the denoising U-Net network of the drawing model; and a denoising representation feature decoding module 560, which is used to feature decode the final denoised latent representation to obtain a generated image.
[0074] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to necessarily being implemented using the above specific details.
Claims
1. A method for adjusting a large drawing model based on multimodal knowledge driving, characterized in that: include: Get the text prompt entered by the user; Performing text parsing and semantic embedding encoding on the text prompt to obtain a text prompt structured intent representation and a text prompt standard condition embedding; Performing multimodal knowledge retrieval based on the text prompt structured intent representation to obtain a multimodal knowledge representation; Inputting the text prompt standard condition embedding and the multimodal knowledge representation into a knowledge-attention translation module to obtain attention modulation parameters; Inputting the text prompt standard condition embedding and the randomly generated initial noise latent representation into the drawing large model to obtain a final denoised latent representation, wherein the attention modulation parameter is used to influence the parameters of the attention layer of the denoising U-Net network of the drawing large model; Feature decoding is performed on the final denoised latent representation to obtain a generated image.
2. The method for adjusting a large drawing model based on multimodal knowledge driving according to claim 1 is characterized in that: The text prompt is subjected to text parsing and semantic embedding encoding to obtain a text prompt structured intent representation and a text prompt standard condition embedding, including: Performing text parsing on the text prompt to extract key entities therefrom, and constructing a structured intent representation of the text prompt in JSON format based on the key entities; The text prompt is input into a pre-trained text encoder to obtain the text prompt standard conditional embedding, and the pre-trained text encoder is CLIP Text Encoder.
3. The method for adjusting a large drawing model based on multimodal knowledge driving according to claim 1 is characterized in that: Performing multimodal knowledge retrieval based on the text prompt structured intent representation to obtain a multimodal knowledge representation includes: Using the text prompt structured intent representation as a query, searching a multimodal knowledge base for knowledge fragments related to the query, wherein the knowledge fragments include structured relevant knowledge facts and relevant visual feature vectors; The knowledge fragments are aggregated to obtain the multimodal knowledge representation.
4. The method for adjusting a large drawing model based on multimodal knowledge drive according to claim 3 is characterized in that: The text prompt standard condition embedding and the multimodal knowledge representation are input into a knowledge-attention translation module to obtain attention modulation parameters, including: The text prompt standard condition embedding and the multimodal knowledge representation are input into the Transformer network to obtain the attention modulation parameters.
5. The method for adjusting a large drawing model based on multimodal knowledge driving according to claim 4 is characterized in that: Inputting the text prompt standard conditional embedding and the randomly generated initial noise latent representation into the drawing large model to obtain a final denoised latent representation, including: in an iterative process of denoising using a denoising U-Net network of the drawing large model: Extracting intermediate latent features from the previous layer of the denoising U-Net network; performing feature flattening on the intermediate latent feature to obtain an intermediate latent feature flattened vector; Projecting the intermediate latent feature flattened vector into a query space to obtain an intermediate latent feature query vector; Embedding and mapping the text prompt standard condition into a key space and a value space to obtain a text prompt key vector and a text prompt value vector; Modulating the text prompt key vector and the text prompt value vector using the attention modulation parameter as a bias term to obtain a modulated text prompt key vector and a modulated text prompt value vector; Calculating an attention weight based on the intermediate latent feature query vector, the modulated text prompt key vector, and the modulated text prompt value vector; The modulated text prompt value vector is weightedly summed based on the attention weight to obtain an output feature.
6. The method for adjusting a large drawing model based on multimodal knowledge drive according to claim 5, characterized in that: The method comprises: using the attention modulation parameter as a bias item to modulate the text prompt key vector and the text prompt value vector to obtain a modulated text prompt key vector and a modulated text prompt value vector, comprising: using the attention modulation parameter as a bias item to modulate the text prompt key vector and the text prompt value vector according to the following formula, wherein the formula is: K modulated =K std +D k V modulated =V std +Δ k Among them, K std Represents the text prompt key vector, V std represents the text prompt value vector, Δ k represents the attention modulation parameter.
7. The method for adjusting a large drawing model based on multimodal knowledge drive according to claim 6, characterized in that: Calculating an attention weight based on the intermediate latent feature query vector, the modulated text prompt key vector, and the modulated text prompt value vector, comprising: Performing key-value space topology matching optimization on the modulation text prompt key vector and the modulation text prompt value vector to obtain an optimized modulation text prompt key vector and an optimized modulation text prompt value vector; The attention weight is calculated based on the intermediate potential feature query vector, the optimized modulated text prompt key vector and the optimized modulated text prompt value vector.
8. The method for adjusting a large drawing model based on multimodal knowledge drive according to claim 7, characterized in that: Performing key-value space topology matching optimization on the modulated text prompt key vector and the modulated text prompt value vector to obtain an optimized modulated text prompt key vector and an optimized modulated text prompt value vector, comprising: Calculating a spatial topological gradient field matrix between the modulated text prompt key vector and the modulated text prompt value vector; Based on the variance of the modulated text prompt key vector and the modulated text prompt value vector, performing eigenvalue granularity spatial pedigree cross reconstruction on the spatial topology gradient field matrix to obtain a spatial topology pedigree reconstruction gradient field matrix; The modulated text prompt key vector and the modulated text prompt value vector are respectively mapped to the feature space of the spatial topology spectrum reconstructed gradient field matrix to obtain the optimized modulated text prompt key vector and the optimized modulated text prompt value vector.
9. The method for adjusting a large drawing model based on multimodal knowledge drive according to claim 8, characterized in that: Calculating an attention weight based on the intermediate latent feature query vector, the modulated text prompt key vector, and the modulated text prompt value vector, comprising: Based on the intermediate potential feature query vector, the modulated text prompt key vector and the modulated text prompt value vector, the attention weight is calculated using the following formula: Wherein, a represents the attention weight, Q represents the intermediate potential feature query vector, d is the dimension of the modulated text prompt key vector, Softmax represents the Softmax function, and T represents the vector transpose.
10. A large-scale drawing model adjustment system driven by multimodal knowledge, characterized in that: include: A text prompt acquisition module is used to acquire the text prompt input by the user; A text semantic parsing and embedding module, configured to perform text parsing and semantic embedding encoding on the text prompt to obtain a structured intent representation of the text prompt and a standard condition embedding of the text prompt; A multimodal intent-driven retrieval module, configured to perform multimodal knowledge retrieval based on the textual prompt structured intent representation to obtain a multimodal knowledge representation; a multi-source attention modulation module, configured to input the text prompt standard condition embedding and the multimodal knowledge representation into a knowledge-attention translation module to obtain attention modulation parameters; a text-guided attention denoising module, configured to input the text prompt standard condition embedding and the randomly generated initial noise latent representation into the drawing model to obtain a final denoised latent representation, wherein the attention modulation parameters are used to influence the parameters of the attention layer of the denoising U-Net network of the drawing model; The denoising representation feature decoding module is used to perform feature decoding on the final denoised latent representation to obtain a generated image.