Continuous visual concept learning method based on scene graph extension
By constructing a scene graph knowledge base and designing a new attention mechanism, the problem of insufficient generation quality of literary graph models in complex concept generation tasks is solved, and the continuous expansion of semantic knowledge and the diversity of image generation is achieved, ensuring the logical coherence and visual accuracy of the generated results.
Patent Information
- Application Number
- CN202510365826.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
When facing complex generation tasks involving multiple concepts, the existing literary and biographical graph generation model is difficult to ensure the diversity and consistency of generation, and the lack of effective semantic knowledge continuous expansion mechanism, resulting in a decline in the quality of generation.
By building a knowledge base based on scene graphs, the semantic relationships and context information of images are captured dynamically, and the scene graph assists in prompt word design, combined with a new attention mechanism to guide the generation process, the continuous expansion of semantic knowledge and the diversity of image generation is achieved.
It improves the personalized concept fidelity and semantic richness of the generated images, ensures the logical coherence and visual accuracy of the generated results, and enhances the diversity and semantic consistency of the generated results.
Smart Images

Figure CN120297412A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and generation technology, and is mainly a continuous learning text-to-image method that uses scene graphs to improve the quality of personalized visual concept generation. Background Art
[0002] The text-to-image generation technology has developed rapidly in recent years and has become an important branch in the field of generative artificial intelligence. It combines natural language descriptions with visual generation models to achieve the ability to generate high-quality images from text prompts. These models can generate rich scenes and diverse objects that match the user's description by learning from large datasets of image-text pairs. This technology has shown broad application prospects in multiple fields, such as film and television entertainment, education and training, e-commerce, virtual reality, game development, and medical imaging, greatly improving the efficiency and quality of content creation. The core advantage of the text-to-image generation technology lies in its flexibility and controllability. Users can generate visual content that meets expectations simply by entering natural language. However, although existing methods can generate high-quality single concepts or simple combined scenes, it is often difficult to ensure the diversity and consistency of the models when facing complex tasks involving multiple concept generations.
[0003] At the same time, text-to-image models also face the challenge of continual learning in practical applications. Continual learning requires the model to retain the knowledge learned previously when learning new tasks without losing its original capabilities due to catastrophic forgetting. This problem is particularly prominent in generative models because the learning of new concepts is usually accompanied by significant distribution changes, and the model needs to maintain both the visual features of old concepts and the semantic associations of new concepts. In the absence of appropriate mechanisms, the introduction of new knowledge often significantly weakens the model's generation ability for old knowledge, resulting in a decline in generation quality and diversity.
[0004] Traditional continuous learning methods have made significant progress in tasks such as classification and recognition. These methods usually alleviate the problem of catastrophic forgetting by strategies such as fixing some model parameters, introducing knowledge distillation, or memorizing samples. However, their application in concept generation tasks is relatively limited, mainly because generation tasks need to handle more complex interactions between visual features and semantic information, rather than just maintaining class information. In recent years, continuous learning methods for generation tasks have also begun to develop gradually, such as C-LoRA and L2DM, which attempt to alleviate the forgetting problem of generative models through strategies such as parameter-efficient fine-tuning (e.g., low-rank adaptation) and extended memory banks. However, the limitation of these methods is that they usually focus on the maintenance of visual features and ignore the continuous expansion of semantic knowledge, while semantic knowledge is crucial for improving generation quality, maintaining logical consistency of relationships between concepts, and diversity.
[0005] In addition, the design of prompts has an important impact on the quality of text-to-image generation results. High-quality prompts can not only improve the alignment between the generated image and the input text, but also significantly enhance the diversity and fidelity of the generated results. References: A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in Proceedings of International Conference on Machine Learning, 2021, pp. 8821–8831. However, existing methods usually have a high dependence on prompts. The lack of semantic description of the relationship between scenes or objects may lead to the generated image not meeting expectations or even having logical errors. References: J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al., “Chain-of-thought prompt in gelicits reasoning in large language models,” Advances in neural information processing systems, vol. 35, pp. 24824–24837, 2022. Therefore, the present invention proposes to use an ever-expanding scene graph to store the semantic knowledge of images and use this scene graph to assist in the design of user prompts, so as to significantly improve the diversity of diffusion model generation. Summary of the Invention
[0006] The present invention is a method for continuous visual concept learning based on scene graph expansion, which uses an expanding scene graph to supplement and improve the generated prompt words, thereby enhancing the diversity of generated images. This method introduces the rich semantic information of the scene graph into the generation model through a large language model and guides the generation process of the model through a new attention mechanism to achieve higher personalized concept fidelity and text fidelity.
[0007] The present invention is based on a diffusion model as the basic framework. By continuously expanding the scene graph, high-quality image generation in continuous personalized generation tasks is achieved. First, this method constructs a large-scale scene graph using the Visual Genome dataset, and then extracts the relationships from the image-text pairs in the training dataset and expands the extracted relationships on the existing scene graph, thereby continuously increasing the semantic information of the scene graph. In the generation stage, a subgraph is retrieved from the scene graph based on the prompt words input by the user, and this subgraph is input into the large language model to obtain the scene layout required for model generation. Finally, through the attention mechanism designed by this method, image generation corresponding to the scene layout is achieved.
[0008] From the perspective of continuous personalized generation learning, this method mainly accomplishes the following tasks: In this method, we use a new continuous learning framework to enable the model to transition from visual-level learning to semantic knowledge expansion. First, we construct a knowledge base using the scene graph to dynamically capture the semantic relationships and context information of the images. The scene graph is expanded through the image-text pairs of the dataset, thereby achieving continuous semantic knowledge expansion while the model learns new concept visual information. This knowledge base can automatically adjust the user input according to the model's preferences, thereby helping the user improve the design of the prompt words and ensuring that the generated images are not only visually accurate but also semantically rich and logically coherent. In addition, to further improve the quality and diversity of the generated images, we also design a novel attention mechanism that can use the semantic knowledge in the knowledge base to guide the model generation process.
[0009] To improve the readability of this invention, some terms are defined and explained herein
[0010] Definition 1: Diffusion model. The diffusion model defines two completely opposite iterative processes to add noise and denoise images, thereby fitting the entire data distribution. The forward diffusion process gradually adds Gaussian noise to the training images within a limited number of steps, converting the complex data distribution into a simple and tractable Gaussian distribution. The backward denoising process is the inverse process of the forward diffusion process, and the purpose of this process is to gradually remove the noise added to the training images during the forward process.
[0011] Definition 2: Scene Graph. It is used to describe the objects, entities in an image or video and the relationships between them. Each scene graph consists of a set of nodes and edges, where the nodes represent objects or entities, and the edges represent the relationships or interactions between the nodes. When constructing a scene graph, the objects and relationships in the image can be represented by a set of triples (subject, relationship, object). By representing semantic information in this structured format, semantic relationships can be effectively captured and organized, and it is widely used in tasks such as knowledge management, reasoning, and retrieval.
[0012] Definition 3: Large Language Model (LLM). It is a natural language processing (NLP) model based on deep learning technology and the Transformer architecture. Compared with traditional language models, large language models usually have hundreds of millions or even trillions of parameters and are capable of handling complex language tasks such as text generation, text understanding, question answering, and translation.
[0013] Definition 4: CLIP (Contrastive Language-Image Pre-training). CLIP is a pre-training model based on contrastive text-image pairs. CLIP includes a text encoder and an image encoder, where the text encoder is used to extract text features and the image encoder is used to extract image features. The core idea of CLIP is to map images and text into the same shared embedding space through contrastive learning methods, so as to achieve mutual understanding between images and text. The pre-training process of the CLIP model is based on a large-scale image-text pair dataset, which can learn the deep relationship between vision and language and can be applied to various downstream tasks.
[0014] Definition 5: U-Net. It acts as a denoising network in the diffusion model, and realizes high-quality image generation by recovering the original image distribution from the noisy image. U-Net includes a symmetric encoder and decoder part: the encoder consists of multiple convolutional layers and pooling layers, gradually reducing the spatial resolution of the image while extracting higher-level features; the decoder consists of multiple transposed convolutional layers or upsampling layers, gradually restoring the low-resolution features to a high-resolution image.
[0015] Definition 6: Low-Rank Adaptation. Low-rank adaptation is a way to fine-tune a pre-trained model. Its core idea is that considering that the pre-trained model shows a lower intrinsic dimension when adapting to downstream tasks, that is, the weight changes during the adaptation process have an inherent low-rank property; by introducing low-rank matrix factorization in some layers of the model (such as linear layers or attention layers), the original parameter update is replaced with a low-rank matrix update, so as to achieve more efficient adaptive adjustment without changing the original structure of the model.
[0016] Definition 7: Transformer. Transformer has achieved long-range dependence modeling and efficient parallel computing through self-attention mechanism, and has become the basis for many cutting-edge models (such as BERT, GPT series, etc.). Transformer is also widely used in multiple fields such as computer vision and speech recognition, and has become a core technology in the fields of modern natural language processing and computer vision.
[0017] Definition 8: Attention mechanism. The self-attention mechanism improves the processing ability of long sequences by modeling global dependencies within the sequence; the cross-attention mechanism realizes efficient information fusion across sequences by combining the information of the query sequence and the context sequence. The attention module uses three linear transformations to convert the input X into query (Query), key (Key), and value (Value) matrices, calculates the attention scores through Q and K, and performs weighted combination through V: To capture the relationships in different subspaces, the self-attention mechanism is usually extended to multi-head attention. Multiple attention heads perform parallel computing, and the results are concatenated and then linearly transformed to enhance the model's expressive ability.
[0018] The present invention is a continuous visual concept learning method based on scene graph expansion, and the method includes:
[0019] Step 1: Data preprocessing;
[0020] The dataset includes three datasets: DreamBooth, CustomConcept101, and Visual Genome; the DreamBooth dataset contains 30 concepts, the CustomConcept101 dataset contains 101 concepts, and each concept consists of 3 to 8 images; 10 concepts (such as dogs, cats, backpacks, etc.) are selected to construct a sub-dataset for concept continuous learning; each concept is guided by a text description (such as "A photo of V1 dog") to generate the process; to prevent overfitting, 200 regularization images are generated using a pre-trained text-to-image diffusion model, and data augmentation (such as rotation, cropping, scaling, etc.) is performed; the Visual Genome dataset contains 108,077 images, and 178 object categories and 45 relationship types are selected to construct a scene graph;
[0021] Step 2: Continuously expand the scene graph;
[0022] First, based on the Visual Genome dataset, construct a large-scale scene graph G:
[0023] G = ((o i , r, o j ) │ o i , oj (∀o ∈ O, r ∈ R)
[0024] where O represents the object set, o i , o j represents the i-th and j-th objects, R represents the relation set, and r represents the relation between the corresponding two objects; during the continuous learning process, text annotation is performed on each personalized concept image in the dataset, for example: "There is a V dog on a stone"; use a pre-trained language model (RoBERTa-large in the present invention) to extract the object relation triples of the annotated text: subject, relation, object; for example, given the text "A V dog is on a stone", the extracted triple will be (V dog , on, stone); for each personalized concept in each new task, add the triples in its annotated text to the existing scene graph to achieve continuous expansion of the scene graph;
[0025]
[0026] where f PLM (·) represents the pre-trained language model, h(·) represents the softmax classifier, and c represents the text annotation; the predicted relation finally forms a triple together with the object: object, relation, object; r i represents an element in the relation set R, represents the relation predicted by the model, and argmax represents the value function of the independent variable that makes the subsequent function obtain the maximum value;
[0027] Step 3: Construct a neural network;
[0028] The constructed neural network includes three sub-networks: CLIP text encoder, variational autoencoder, and U-Net; where U-Net is the core network; the CLIP text encoder receives the text string and outputs the text encoding, and the variational autoencoder receives the image and outputs the low-dimensional latent features; the specific process is: the input image first generates latent features through the pre-trained variational autoencoder, then performs noise addition processing, and inputs it into U-Net;
[0029] Step 4: Low-rank mixture of experts module;
[0030] Construct the low-rank adaptation expert ΔW = AB, Initialized as a Gaussian distribution matrix; a new low-rank adaptation expert is introduced for each personalized concept. A represents the Gaussian initialization matrix, B represents the zero initialization matrix, and D1 and D2 are different dimensions of the weight W. For the t-th task prompt (e.g., "photo of V*dog"), its embedding representation encoded by CLIP is where n represents the number of prompt tokens and d represents the dimension of each token; a gating function g t (·): R n×d →R t is used to combine the existing low-rank adaptation experts. → represents the function mapping, and the final parameter matrix W t is expressed as:
[0031]
[0032] where W0 is the pre-trained parameter matrix, ΔW τ is the low-rank adaptation expert matrix specific to the τ-th task, t is the number of tasks learned so far, and α t,τ represents the combination coefficient of the expert parameter matrix ΔW τ at the t-th task; the trainable gating function g t (·) dynamically selects the expert most relevant to the given prompt, and the function is:
[0033]
[0034] The superscript represents the transpose;
[0035] Step 5: Define the network loss function;
[0036] Step 6: Generate the scene graph layout;
[0037] When the user inputs a prompt word containing a single concept, first identify the concept object o in the large-scale scene graph s ; iteratively expand the subgraph by retrieving the objects and relationships related to the concept o s ; the expansion process only includes adjacent nodes and edges and stops when a predefined constraint is reached in terms of node count; the resulting subgraph G s = ((o i , r, o j ) │ o i , o j ∈O s , r ∈ R s ) is used as the input to the LLM;
[0038] Predict the scene graph layout where each Denote the bounding box of the k-th object, where there are four dimensions corresponding to the upper left and lower right coordinates of the bounding box;
[0039] Step 7: Layout-guided attention enhancement:
[0040] Modify the cross-attention module of the models obtained in Steps 3, 4, and 5; Given K bounding boxes The corresponding self-attention masks where the values of the bounding box regions within the mask are 0 and the values of other regions are -∞, denote the self-attention mask of object k, h represents the length dimension of the image, and w represents the width dimension; These masks are applied before the softmax operation of the attention mechanism to ensure that the final attention scores are still valid after introducing spatial constraints; The attention map before softmax The calculation formula is:
[0041]
[0042] where, are the query matrix and the key matrix respectively, and hw represents the number of patches; The attention map A contains hw row vectors: where a i denotes the attention vector of the i-th patch, and d represents the dimension of each patch; To interact with the spatial mask First, is flattened into a one-dimensional vector, denoted as For each patch i, the attention vector a i is updated to:
[0043] If patch i is within the bounding box b k The updated attention vector Together form the modified attention map Finally, is used in subsequent self-attention calculations; This method only applies self-attention enhancement in the τ steps before the diffusion step to ensure smooth transitions between different regions; This process is as Figure 3 shown;
[0044] Step 8: Layout-guided multi-expert combination;
[0045] Given K bounding boxes Generate the corresponding cross-attention masks denote the cross-attention mask of object k, where the values of the bounding box regions within the mask are 1 and the values of other regions are 0;
[0046] During the inference stage, a region-based guidance mechanism is introduced to optimize the generation process, as Figure 4 shown; First, the prompt composed of scene Figure 3 tuples is input into the CLIP encoder, and the embedding vector E of each object k is extracted respectively k ; Then, a module is used to identify the identifier V* of the personalized concept. If the identifier is recognized, in the cross-attention layer, the mixture-of-experts mechanism (as Figure 4 shown at the bottom) is used to select a suitable low-rank adaptation expert for the object region according to the embedding vector:
[0047]
[0048]
[0049] represents the combination coefficient of the τ expert on the object k, E k represents the text encoding of the object k. If the identifier is not recognized, the region is generated by the pre-trained weight W0; Finally, the output is based on the region mask to ensure that the attention is concentrated on specific regions of interest; The final output of the cross-attention layer is calculated as follows:
[0050]
[0051] In each layer, F(Z in , E; W) is defined as the mapping function in the cross-attention layer; Specifically, F accepts the input feature and maps it to the output feature
[0052] Through the above steps, the mechanism of the region mask combined with the low-rank adaptation expert further improves the attention to specific regions and the generation quality during the generation process;
[0053] Furthermore, the U-Net structure in step 3 is as Figure 2 shown, and successively includes: a convolutional layer, three cross-attention downsampling modules, a downsampling module, a cross-attention middle block, an upsampling module, three cross-attention upsampling modules, and finally outputs Gaussian noise with the same dimension as the input through the GSC module; The three cross-attention downsampling modules perform data jumping to the corresponding three cross-attention upsampling modules. In order, the first cross-attention downsampling module corresponds to the third cross-attention upsampling module, the second cross-attention downsampling module corresponds to the second cross-attention upsampling module, and the third cross-attention downsampling module corresponds to the first cross-attention upsampling module.
[0054] Furthermore, the specific method of step 5 is:
[0055] Apply a regularization constraint to the gating function g t (·) to prevent catastrophic forgetting; this regularization term is:
[0056]
[0057] where g t (E t ) [1:t-1] represents the first t - 1 elements of the output of the gating function; for the update of the low - rank adaptation matrix, the reconstruction loss of the latent diffusion model (LDM) is used This loss is defined as follows:
[0058]
[0059] represents the expectation of the model loss, ∈ represents the added Gaussian distributed noise, ∈ θ (z t ,t,ψ θ (c)) represents the model - predicted noise, z t represents the model input feature at time step t, ψ θ (c) represents the text encoding of the prompt c;
[0060] The final total loss is defined as follows:
[0061]
[0062] where β is a hyperparameter that controls the trade - off between the two losses; by incorporating this low - rank mixture - of - experts mechanism, the model can continuously expand its capabilities while retaining the knowledge of previous tasks, thus addressing the challenges of continuous learning.
[0063] Furthermore, an additional step 9 is added for the image diversity evaluation metric;
[0064] In step 7, after modifying the model attention module for image generation, the generated image is used with ViT - B to extract the features of the image; the variance of these features is used as a diversity metric; a higher variance indicates that the generated image has greater diversity, reflecting the model's ability to produce diverse and rich results.
[0065] The innovations of the present invention:
[0066] 1) From the perspective of dynamic semantic knowledge expansion, the present invention proposes a continuous learning framework from visual feature learning to semantic knowledge expansion, which solves the problem of the lack of continuous growth of semantic data in existing continuous learning methods. By constructing a knowledge base based on scene graphs, the semantic relationships and context information of images are dynamically captured, effectively maintaining the semantic richness of new concepts. Using image-text pair data to construct scene graphs and seamlessly integrating image and text information into the knowledge base enables the model to dynamically expand semantic knowledge when learning new concepts. This design enables the model to avoid semantic information loss during the continuous learning process.
[0067] 2) The present invention designs a knowledge base-assisted prompt optimization method that automatically adjusts the user input prompt content through a dynamically generated knowledge base of scene graphs to match the model preferences. This optimization mechanism helps users design prompts with greater semantic depth, ensuring that the generated images are significantly improved in terms of visual accuracy, semantic richness, and logical consistency, thus overcoming the problem of difficult user prompt design.
[0068] 3) The present invention designs a novel attention mechanism that uses the scene layout integrating semantic information to guide the model's generation process. This mechanism not only improves the quality of the generated images but also enhances the diversity and semantic consistency of the generated results, solving the problem of insufficient diversity of generated images in existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is the overall framework and flowchart of the method of the present invention.
[0070] Figure 2 It is the U-Net structure diagram of the method of the present invention.
[0071] Figure 3 It is the schematic diagram of layout-guided attention enhancement of the method of the present invention.
[0072] Figure 4 It is the schematic diagram of layout-guided multi-expert combination of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0073] Under the diffusion model framework, the present invention conducts continuous learning of model capabilities and continuous expansion of scene graphs to achieve the retention of visual knowledge and semantic knowledge. In the inference stage, a scene layout with rich semantic information is obtained from the existing scene graphs, and the attention mechanism proposed by the present invention is used to guide the model's generation process, thereby obtaining image generation results with high diversity, high fidelity, and high text matching degree. The framework and process of the method of the present invention are shown in Figure 1 .
[0074] Step 1: Experimental data preprocessing;
[0075] The experimental data of the present invention is based on two personalized generated image datasets and a scene graph dataset: DreamBooth and CustomConcept101 datasets. The DreamBooth dataset contains 30 concepts, and the CustomConcept101 dataset contains 101 concepts. These concepts include natural color images such as pets, toys, scenes, wearable devices, etc. Each concept contains 3 to 8 images. Select images of ten concepts including dogs, duck toys, cats, backpacks, teddy bears, cars, flowers, table lamps, shoes, and bicycles to construct a sub-dataset for the continuous learning task of concepts. For each concept, construct a simple text description without background information, which includes category information corresponding to the image, such as "A photo of V1 dog", where V1 is an identifier specific to the concept. The text description will be used as a guide in the model generation process. To avoid overfitting, use a pre-trained text-to-image diffusion model to generate 200 regularized images according to the text description of each concept. These regularized images perform the same data augmentation operations as the images in the dataset during training, including random rotation, cropping, scaling, etc. The scene graph dataset Visual Genome consists of 108,077 images with scene graph annotations. Select 178 object categories and 45 relationship types that appear at least 2,000 times and 500 times respectively to construct a large-scale scene graph;
[0076] Step 2: Continuous expansion of the scene graph;
[0077] First, based on the Visual Genome dataset processed in Step 1, construct a large-scale scene graph where represents the object set, represents the relationship set. During the continuous learning process, perform text annotation on each personalized concept image in the dataset. For example: "There is a V dog on a stone". Use a pre-trained language model (RoBERTa-large in the present invention) to extract the object relationship triples (subject, relationship, object) of the annotated text. For example, given the text "A V dog is on a stone", the extracted triple will be (V dog , on, stone). For each personalized concept in each new task, we will add the triples in its annotated text to the existing scene graph to achieve the continuous expansion of the scene graph.
[0078]
[0079] where f PLM(·) represents a pre-trained language model, h(·) represents a softmax classifier, and c represents text annotation. The predicted relationship finally forms a triple (object, relationship, object) together with the object;
[0080] Step 3: Construct a neural network;
[0081] The constructed neural network consists of three sub-networks, namely: CLIP text encoder, variational autoencoder, and U-Net, where U-Net is the core network. The input of the CLIP text encoder is a text string, and the output is a text encoding; the input of the variational autoencoder is an image, and the output is a low-dimensional latent feature. The specific process is as follows: The input image first generates a low-dimensional latent feature through a pre-trained variational autoencoder, and then this latent feature is added with noise during the forward process and input into the U-Net. The structure of the U-Net is as Figure 2 shown, including: The first layer is a convolutional layer, followed by three cross-attention downsampling modules, a downsampling module, a cross-attention intermediate block, an upsampling module, and three cross-attention upsampling modules, and finally a Gaussian noise with the same dimension as the input is output through a GSC module;
[0082] Step 4: Low-rank mixture of experts module;
[0083] Construct a low-rank adaptation expert ΔW = AB, initialized as a Gaussian distribution matrix. Each personalized concept introduces a new low-rank adaptation expert. For the t-th task prompt (e.g., "photo of V*dog"), its embedding representation encoded by CLIP is where n represents the number of prompt tokens, and d represents the dimension of each token. Use a gating function g t (·):R n×d →R t to combine the existing low-rank adaptation experts, and finally the parameter matrix W t is expressed as:
[0084]
[0085] where W0 is the pre-trained parameter matrix, Δw τ is the low-rank adaptation expert matrix specific to the τ-th task, and t is the number of tasks learned so far. The trainable gating function g t (·) dynamically selects the expert most relevant to the given prompt, and this function is defined as follows:
[0086]
[0087] Step 5: Define the network loss function;
[0088] For the gating function gt (·) Apply regularization constraints to prevent catastrophic forgetting. The regularization term is defined as:
[0089]
[0090] where g t (E τ ) [1:t-1] represents the first t - 1 elements of the output of the gate function. For the update of the low - rank adaptation matrix, the reconstruction loss of the latent diffusion model (LDM) is adopted, and the loss is defined as follows:
[0091]
[0092] The final total loss is defined as follows:
[0093]
[0094] where β is a hyperparameter that controls the trade - off between the two losses. By incorporating this low - rank mixture - of - experts mechanism, the model can continuously expand its capabilities while retaining the knowledge of previous tasks, thus addressing the challenges faced by continuous learning;
[0095] Step 6: Scene graph layout generation;
[0096] When the user inputs a prompt word containing a single concept, first identify the concept object o in the large - scale scene graph s . Iteratively expand the sub - graph by retrieving the objects and relationships related to the concept o s . The expansion process only includes adjacent nodes and edges and stops when a predefined constraint on the node count is reached. The resulting sub - graph is used as the input to the LLM to predict the scene graph layout where each represents the bounding box of the k - th object, with four dimensions corresponding to the coordinates of the upper - left and lower - right corners of the bounding box;
[0097] Step 7: Layout - guided attention enhancement:
[0098] Modify the cross - attention module of the model obtained in steps 3, 4, and 5. Given K bounding boxes and their corresponding self - attention masks where the values within the bounding - box regions of the masks are 0 and the values in other regions are - ∞. These masks are applied before the softmax operation of the attention mechanism to ensure that the final attention scores are still valid after introducing spatial constraints. The attention map before softmax is calculated as:
[0099]
[0100] Among them, are the query matrix and the key matrix respectively, and hw represents the number of patches. The attention map A contains hw row vectors: where a i represents the attention vector of the i-th patch. To interact with the spatial mask first, is flattened into a one-dimensional vector, denoted as For each patch i, the attention vector a i is updated to:
[0101] If patch i is within the bounding box b k inside
[0102] The updated attention vectors together form the modified attention map Finally, is used in the subsequent self-attention calculation. This method only applies self-attention enhancement in the τ steps before the diffusion step to ensure smooth transitions between different regions. This process is as Figure 3 shown.
[0103] Step 8: Layout-guided multi-expert combination
[0104] Given K bounding boxes generate the corresponding cross-attention masks where the values of the regions inside the masks within the bounding boxes are 1 and the values of other regions are 0. In the inference stage, a region-based guiding mechanism is introduced to optimize the generation process, as Figure 4 shown. First, the prompt consisting of the scene Figure 3 tuples is input into the CLIP encoder, and the embedding vectors E k of each object k are extracted respectively. Then, a module is used to identify the identifier V* of the personalized concept. If the identifier is recognized, in the cross-attention layer, a low-rank adaptation expert suitable for the object region is selected according to the embedding vectors using the mixture of experts mechanism (as Figure 4 shown at the bottom) (here g T and the number of tasks T in
[0105]
[0106] are omitted): Ensure that the attention is focused on a specific region of interest. The final output of the cross-attention layer is calculated as follows:
[0107]
[0108] In each layer, define F(Z in , E; W) as the mapping function in the cross-attention layer. Specifically, F takes the input feature and maps it to the output feature
[0109] Through the above steps, the regional mask combined with the mechanism of low-rank adaptation experts further improves the attention to specific regions and the generation quality during the generation process.
[0110] Step 9: Image diversity evaluation metrics;
[0111] In Step 7 and 8, after modifying the model attention module, image generation is performed. The generated images are used to extract the features of the images using ViT-B. The variance of these features is used as a diversity metric. A higher variance indicates that the generated images have greater diversity, reflecting the ability of the model to produce diverse and rich results. The test results are shown in Table 1.
[0112] Step 10: Comparison of different continuous learning strategies on test metrics;
[0113] For each personalized concept, 4 different generated scenario graphs are selected. For each scenario graph, 25 images are generated. All the generated images are used to calculate the CLIP score with the text description corresponding to the scenario graph, denoted as TA, so as to evaluate the ability of the model of the present invention to generate concepts in different semantic environments; use "A photo of V * <concept>"The prompt generates 100 images, and the CLIP score is calculated for the training images, denoted as IA. Thus, the ability of the model of the present invention to generate target concepts is evaluated; the forgetting rate is calculated based on the CLIP-IA score results of the model in different task stages. The test results are shown in Table 2. The present invention verifies the superiority of the continuous visual concept generation method based on scene graph expansion on the used dataset. The experimental results show that the proposed method can well learn to improve the concept generation fidelity and text matching degree, and at the same time improve the diversity of the image generation results.
[0114] Table 1 and Table 2 are the experimental results of the method of the present invention.
[0115] Table 1 shows the experimental results of the diversity of the method of the present invention in 10 tasks. The higher the Diversity, the better the generation diversity of the model.
[0116] Table 1
[0117]
[0118]
[0119] Table 2 shows the experimental results of the diversity of the method of the present invention in 10 tasks. The higher the text alignment TA and image alignment IA, and the lower the forgetting rate Forgetting, the better the generation diversity of the model.
[0120] Table 2
[0121]
[0122] The test is carried out under different continuous learning methods. From the results in Table 1, the method proposed by the present invention obtains the highest diversity score, indicating that using scene graph knowledge to enrich the input prompt can effectively enhance the richness of the generated images. From the evaluation results of indicators such as text alignment, image alignment, and forgetting in Table 2, it can be seen that this method is significantly superior to other methods in text alignment, which indicates that the attention mechanism proposed by the present invention effectively retains the context relationship in the prompt, ensuring a strong consistency between the generated images and the scene layout, thereby improving the rationality of the generation. In addition, while achieving competitive image alignment with other continuous learning methods, our method obtains the lowest forgetting score, verifying the effectiveness of the gated function regularization loss proposed by the present invention.< / concept>
Claims
1. A continuous visual concept learning method based on scene graph expansion, the method comprising: Step 1: Data preprocessing; The dataset includes three datasets: DreamBooth, CustomConcept101, and Visual Genome; 10 concepts are selected to construct a sub-dataset for concept continuous learning; each concept is guided by a text description to generate the process; to prevent overfitting, a pre-trained text-to-image diffusion model is used to generate regularization images and data augmentation is performed; Step 2: Continuously expand the scene graph; First, based on the Visual Genome dataset, a large-scale scene graph G is constructed: G = ((o i , r, o j ) │ o i , o j ∈ O, r ∈ R) where O represents the set of objects, o i , o j represents the i-th and j-th objects, R represents the set of relationships, and r represents the relationship between the corresponding two objects; During the continuous learning process, text annotation is performed on each personalized concept image in the dataset, and a pre-trained language model is used to extract the object relationship triples of the annotation text: subject, relationship, object; for each personalized concept in each new task, the triples in its annotation text are added to the existing scene graph to achieve continuous expansion of the scene graph; Among them, f PLM (·) represents the pre-trained language model, h(·) represents the softmax classifier, and c represents the text annotation; the predicted relationship finally forms a triple with the object: object, relationship, object; r i represents an element in the relationship set R, represents the relationship predicted by the model, and argmax represents the independent variable value function that makes the subsequent function obtain the maximum value; Step 3: Construct a neural network; The constructed neural network includes three sub-networks: a CLIP text encoder, a variational autoencoder, and a U-Net; where the U-Net is the core network; the CLIP text encoder receives a text string and outputs a text encoding, and the variational autoencoder receives an image and outputs a low-dimensional latent feature; the specific process is: the input image first generates a latent feature through a pre-trained variational autoencoder, then noise is added, and it is input into the U-Net; Step 4: Low-rank mixture of experts module; Construct a low-rank adaptation expert Δw = AB, Initialize it as a Gaussian distribution matrix; introduce a new low-rank adaptation expert for each personalized concept, A represents the Gaussian initialization matrix, B represents the zero initialization matrix, D1 and D2 are different dimensions of the weight W respectively; for the t-th task prompt, its embedding representation encoded by CLIP is where n represents the number of prompt tokens, d represents the dimension of each token; utilize a gating function g t (·): R n×d →R t to combine the existing low-rank adaptation experts, → represents the function mapping, and the final parameter matrix W t is expressed as: where W0 is the pre-trained parameter matrix, ΔW τ is the low-rank adaptation expert matrix specific to the τ-th task, t is the number of tasks learned so far, α t,τ represents the combination coefficient of the expert parameter matrix ΔW τ at the t-th task; the trainable gate function g t (·) dynamically selects the expert most relevant to the given prompt, and the function is: The superscript T represents transpose; Step 5: Define the network loss function; Step 6: Scene graph layout generation; When the user inputs a prompt containing a single concept, first identify the concept object o in the large-scale scene graph s ; iteratively expand the subgraph by retrieving objects and relationships related to the concept o s ; the expansion process only includes adjacent nodes and edges and stops when a predefined constraint is reached in terms of node count; the resulting subgraph G s = ((o i , r, o j ) | o i , o j ∈ O s , r ∈ R s ) as the input of the LLM; Predict the layout of the scene graph where each represents the bounding box of the k-th object, with four dimensions corresponding to the coordinates of the upper left and lower right corners of the bounding box; Step 7: Layout-guided attention enhancement: Modify the cross-attention module of the models obtained in steps 3, 4, and 5; given K bounding boxes corresponding self-attention masks where the values of the bounding box regions within the mask are 0 and the values of other regions are -∞, denotes the self-attention mask of object k, h represents the length dimension of the image, and w represents the width dimension; these masks are applied before the softmax operation of the attention mechanism to ensure that the final attention scores remain valid after introducing spatial constraints; the attention map before softmax is calculated as follows: Among them, are the query matrix and the key matrix respectively, and hw represents the number of patches; the attention map A contains hw row vectors: A = [a1; a2; …; a hw T , where a i represents the attention vector of the i-th patch, and d represents the dimension of each patch; in order to interact with the spatial mask , first is flattened into a one-dimensional vector, denoted as For each patch i, the attention vector a i is updated to: If patch i is within the bounding box b k inside Updated attention vector Together they form the modified attention map Finally It is used in subsequent self-attention calculations; this method only applies self-attention enhancement in the τ steps before the diffusion step to ensure smooth transitions between different regions; Step 8: Layout-guided multi-expert combination; Given K bounding boxes Generate corresponding cross-attention masks Denote the cross-attention mask of object k, where the values of the bounding box region inside the mask are 1 and the values of other regions are 0; During the inference stage, a region-based guidance mechanism is introduced to optimize the generation process; first, the prompt composed of scene graph triples is input into the CLIP encoder, and the embedding vector E of each object k is extracted separately k ; then, a module is used to identify the identifier V* of the personalized concept. If the identifier is identified, in the cross-attention layer, the mixture-of-experts mechanism is used to select a suitable low-rank adaptation expert for the object region according to the embedding vector: Denotes the combination coefficient of the τ expert on object k, Denotes the text encoding of object k. If no identifier is recognized, this area is generated by the pre-trained weight W0; finally, the output is based on the regional mask Ensures that attention is focused on specific regions of interest; the final output of the cross-attention layer is calculated as follows: In each layer, define F(Z in , E; W) as the mapping function in the cross-attention layer; specifically, F takes the input feature and maps it to the output feature 2. A continuous visual concept learning method based on scene graph expansion according to claim 1, characterized in that The U-Net structure in the said Step 3 sequentially includes: a convolutional layer, three cross-attention downsampling modules, a downsampling module, a cross-attention intermediate block, an upsampling module, three cross-attention upsampling modules, and finally outputs Gaussian noise with the same dimension as the input through the GSC module; the three cross-attention downsampling modules perform data jumping for the corresponding three cross-attention upsampling modules, and in order, the first cross-attention downsampling module corresponds to the third cross-attention upsampling module, the second cross-attention downsampling module corresponds to the second cross-attention upsampling module, and the third cross-attention downsampling module corresponds to the first cross-attention upsampling module.
3. A continuous visual concept learning method based on scene graph expansion according to claim 1, characterized in that The specific method of the said Step 5 is: For the gating function g t (·) impose a regularization constraint to prevent catastrophic forgetting; this regularization term is as follows: Among them, g t (E t ) [1:t-1] represents the first T - 1 elements of the output of the gate function; for the update of the low - rank adaptation matrix, the reconstruction loss of the latent diffusion model is adopted This loss is defined as follows: Denote the expectation of the model loss, ∈ denote the added Gaussian distributed noise, ∈ θ (z t ,t,ψ θ (c)) denote the model predicted noise, z t denote the model input features at time step t, ψ θ (c) denote the text encoding of the prompt c; The final total loss is defined as follows: where β is a hyperparameter that controls the trade-off between the two losses.
4. A continuous visual concept learning method based on scene graph expansion according to claim 1, characterized in that An additional Step 9 is added for the image diversity evaluation metric; In Step 7 and 8, after modifying the model attention module, image generation is performed, and the generated image is used with ViT-B to extract the features of the image; The variance of these features is used as a diversity metric.
Citation Information
Cited By
Double-branch diffusion three-dimensional scene generation method based on multi-modal semantic graph
CN120833445A
Image generation method and device based on role consistency
CN121458830A