Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

142 results about "Image diffusion" patented technology

License plate recognition system and method based on image technology and medium

The invention relates to the technical field of image recognition, in particular to a license plate recognition system and method based on an image technology and a medium. The method comprises the following steps: acquiring area sensing data and a camera image set, and performing deformation effect compensation to obtain an environment compensation image set; performing image diffusion reverse enhancement on the environment compensation image set to obtain a license plate area enhanced image set; performing character region high-dimensional topological mapping based on the license plate region enhanced image set to obtain a character segmentation matrix; extracting character morphological characteristics according to the character segmentation matrix, and performing character recognition on the character morphological characteristics to obtain a character recognition result; and carrying out cross-character semantic compensation on the character recognition result to obtain a semantic compensation license plate character vector, and carrying out multi-target cross verification on the semantic compensation license plate character vector to obtain a license plate recognition result. According to the invention, the accuracy and robustness of license plate recognition can be improved.
Owner:SHENZHEN YUNBO IND CO LTD

Learning continuous control for 3d-aware image generation on text-to-image diffusion models

A method, apparatus, non-transitory computer readable medium, and system for image processing include obtaining a text prompt describing an element and an attribute value for a continuous attribute of the element, embedding the text prompt to obtain a text embedding in a text embedding space, embedding the attribute value to obtain an attribute embedding in the text embedding space, and generating a synthetic image based on the text embedding and the attribute embedding, where the synthetic image depicts the continuous attribute of the element based on the attribute value.
Owner:ADOBE INC

Text generation image diffusion model enhancement method based on multi-target preference optimization

The invention discloses a text generation image diffusion model enhancement method based on multi-target preference optimization. The method comprises the following steps: firstly, determining a plurality of reward models, and constructing a sample pair training set comprising positive and negative samples; and then, generating a loss weight of each sample pair in the sample pair training set, performing fine tuning training on the text map diffusion model by using the sample pair training set, and in the fine tuning training process, calculating a loss function value of each sample pair in combination with the loss weight of each sample pair until the training is completed, thereby obtaining an aligned text map diffusion model. According to the method provided by the invention, manual data annotation is not needed, the problems of preference inconsistency and over-optimization in a multi-reward scene are effectively solved, and the image quality, the text alignment capability and the multi-target optimization performance of the text-to-image generation model are remarkably improved. The method is superior to an existing optimization method under single-reward and multi-reward setting, shows higher generation quality and robustness, and can be seamlessly applied to various picture generation models.
Owner:ZHEJIANG UNIV

Training-free consistent text-to-video generation

Embodiments of the present disclosure relate to training-free consistent text-to-image generation. A pre-trained text-to-image diffusion model is leveraged to generate images depicting a consistent subject for diverse prompts describing scenes. Inputs to the model are a text description of at least one subject with prompts (scene text descriptions) describing scenes, where each prompt is associated with a different generated image and the text description is used for all images that depict the subject. Internal activations (intermediate data) computed by the model during generation of the different images are shared for generation of the different images. A subject-driven shared attention block and correspondence-based feature injection are incorporated into the model to promote subject consistency within each image and / or between images. Additionally, layout diversity is encouraged while maintaining subject consistency. The model achieves state-of-the-art performance on subject consistency and text alignment, without requiring any optimization and naturally extends to multi-subject scenarios.
Owner:NVIDIA CORP

System and Method for Event-Driven Video Synthesis Using Textual Descriptions

A video generation framework that is controllable, unsupervised and based on events (CUBE) includes an event camera, which captures changes in light intensity at each pixel of a scene asynchronously and generates event camera data. A text-to-image diffusion model that is conditioned on textual descriptions integrates the event camera data to control video synthesis. Further, an edge extraction module translates event data into a format usable by the text-to-image diffusion model, whereby the diffusion model synthesizes detailed and contextually accurate videos based on textual prompts. Further, an improved system (CUBE Plus) includes a content frame identification module which selectively identifies and uses only the most information-rich event segments of the event camera data to drive cross-frame attention, and an event driven attention mechanism that allows the framework to focus on event-dense moments.
Owner:THE UNIVERSITY OF HONG KONG

Automatic quality control of image diffusion processing

A computer system that performs quality control (QC) on images associated with diffusion and structural magnetic resonance imaging (MRI) is described. This computer may include: a computation device that executes program instructions; and memory that stores the program instructions. During operation, the computer system may automatically perform a set of validation operations, where, when one or more of the validation operations fails, the images are rejected. Moreover, the set of validation operations may include: performing QC on brain-tissue segmentation; performing QC on diffusion MRI processing; and performing QC on bundles determined from the images using a tractometry technique.
Owner:IMEKA SOLUTIONS INC

Avatar Generation using Image Diffusion Models

A method of generating a 3-dimensional representation of a subject is provided. The method includes receiving one or more descriptions characterizing the subject. The method also includes inputting the one or more descriptions characterizing the subject into a first specialized network of a machine learning model to generate one or more images depicting the subject according to the one or more descriptions. The method further includes inputting the generated one or more images to a second specialized network of the machine learning model to generate the 3-dimensional representation of the subject according to the one or more descriptions characterizing the subject.
Owner:GOOGLE LLC

Method for embedding robust watermark in diffusion model generated image

The invention belongs to the field of image processing, and relates to a method for embedding a robust watermark in a diffusion model generated image, which comprises the following steps of: acquiring cue words, initial Gaussian noise and watermark information, and inputting the cue words, the initial Gaussian noise and the watermark information into a trained diffusion model based on watermark embedding to obtain a watermark-embedded image; the training process of the diffusion model comprises the following steps: acquiring cue words, initial Gaussian noise and watermark information, and inputting the cue words, the initial Gaussian noise and the watermark information into an encoder to obtain potential vectors; inputting the potential vector into a self-attention module to obtain an embedded position vector; embedding the watermark information into the potential vector according to the embedding position vector; inputting the potential vector embedded with the watermark information into a decoder to obtain an image embedded with the watermark; extracting watermark information from the image embedded with the watermark; updating parameters of a diffusion model according to the image embedded with the watermark and the extracted watermark information until a trained diffusion model is obtained; according to the method, the embedding position is selected by combining the potential of the diffusion model and the accuracy of the self-attention mechanism, so that efficient watermark embedding and extraction are realized.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Single view reconstruction and rendering method

The invention discloses a single view reconstruction and rendering method, which belongs to the field of view reconstruction and rendering, and comprises the following steps of: firstly, generating multi-view feature representation with strong geometric consistency from a single input image by introducing an image diffusion module of a cross attention mechanism; a point cloud reconstruction module with self-attention and cross-attention is utilized, and multi-view information is fused to reconstruct an accurate three-dimensional point cloud; and finally, constructing differentiable three-dimensional Gaussian representation based on the point cloud, rendering the differentiable three-dimensional Gaussian representation, and outputting a new view angle image, a normal map and a depth map. According to the system, a staged training strategy is adopted, and a composite loss function including multi-scale bidirectional consistency smooth loss and feature consistency loss is innovatively used for optimization. According to the method, the problems of low geometric accuracy, multi-view inconsistency, detail missing and the like in single-view reconstruction are effectively solved, and the method can be widely applied to the fields of virtual reality, digital twinning, cultural heritage digitization and the like.
Owner:北京渲光科技有限公司

Text-to-image diffusion model rearchitecture

Described is a system for improving machine learning models. In some cases, the system improves such models by identifying a performance characteristic for machine learning model blocks in an iterative denoising process of a machine learning model, connecting a prior machine learning model block with a subsequent machine learning model block of the machine learning model blocks within the machine learning model based on the identified performance characteristic, identifying a prompt of a user, the prompt indicative of an intent of the user for generative images, and analyzing data corresponding to the prompt using the machine learning model to generate one or more images, the machine learning model trained to generate images based on data corresponding to prompts.
Owner:SNAP INC

Defect image diffusion generation method and device, equipment and storage medium

The invention discloses a defect image diffusion generation method and device, equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: determining a target defect type and a defect size parameter corresponding to a target defect according to a natural language prompt text input by a user; building a defect physical vector based on the target defect type and the defect size parameter; adjusting the diffusion noise amplitude in the diffusion process based on the defect physical vector, and generating an initial defect image corresponding to the initial workpiece noise image; and performing physical consistency verification on the initial defect image, and outputting a target defect image according to a verification result. The diffusion noise amplitude in the diffusion process is adjusted based on the defect physical vector corresponding to the target defect, the initial defect image is generated, and the target defect image is output according to the verification result of the initial defect image, so that the problem that the defect image generation method in the prior art cannot pay attention to the physical size of the defect is solved. And the accuracy of the generated defect image is not high.
Owner:GUANGDONG MECHANICAL & ELECTRICAL COLLEGE

Method for generating aviation lifelike tree sample and identifying tree species by using image diffusion model

The invention discloses a method for generating an aviation lifelike tree sample and identifying a tree species by using an image diffusion model. The method comprises the following steps: constructing a multi-modal data set; training an image-text comparison model and a Unet denoising network model; outputting predicted noise in the trained Unet denoising network model, and finally generating final image reconstruction which is in semantic alignment with the input text description; the obtained single tree crown images are synthesized; and constructing a YOLOv11 detection model, and training the YOLOv11 detection model by using the synthesized forest image to realize tree species identification. According to the method, language semantics and a diffusion model are combined to generate a vivid tree sample for capturing specific features of species and seasonal phenological changes, and the vivid tree sample is synthesized into a high-fidelity forest image, so that tree species identification in an aerial image is enhanced, and effective crown detection and accurate tree species identification are realized.
Owner:NANJING FORESTRY UNIV

Sea surface temperature image diffusion completion method and system based on two-stage fusion constraint

The invention discloses a sea surface temperature image diffusion completion method and system based on dual-stage fusion constraint, and the method comprises the steps: firstly constructing sea area multi-source priori knowledge features to guide the reconstruction of sea surface temperature through integrating three types of heterogeneous data of geography, space-time and power; deep fusion is carried out on the sea area multi-source priori knowledge features through a cross-modal feature fusion module to obtain fused multi-source priori knowledge, and the fused multi-source priori knowledge is fused with structural knowledge to form comprehensive condition features used for guiding the generation process of a conditional diffusion model; and finally, realizing refined reconstruction of the missing region based on a conditional diffusion model, and outputting a completion result. And obtaining a final complete image according to the complementation result, the sea surface temperature image to be complemented and the missing mask. According to the invention, the accuracy of sea surface temperature image completion is improved.
Owner:OCEAN UNIV OF CHINA

Intelligent analysis method for wound image at applet side

The invention, which relates to the technical field of image identification, discloses an intelligent analysis method for a wound image at an applet end, comprising an image acquisition port and an auxiliary end set according to the applet end; according to the invention, the image acquisition port and the auxiliary end are set, the image diffusion models under different networks are constructed, the output image influence analysis value can quantify the influence degree of network factors on image analysis, and in combination with the preset image influence virtual mechanism, the real-time image data are reasonably supplemented and recovered, so that the real-time image analysis efficiency is improved. The method comprises the following steps: preliminarily extracting demand features from two-dimensional image data, constructing a wound three-dimensional dynamic model, performing dynamic matching according to a human body standard image, analyzing and judging the healing condition in a future stage, predicting a wound recovery volume value and actual wound image data, and obtaining the healing degree between wounds. And false analysis and judgment are carried out on the wound healing process, so that false images possibly appearing in the wound healing process can be identified, and the wound identification quality is improved.
Owner:FOURTH MILITARY MEDICAL UNIVERSITY

Image diffusion generation method and system based on multi-modal guidance and feedback closed loop

The invention provides an image diffusion generation method and system based on multi-modal guidance and feedback closed loop, and the method comprises the steps: obtaining multi-modal input information; performing feature extraction processing on the multi-modal input information, and determining structural features and semantic features; generating a guide vector according to the structural features and the semantic features; according to the guide vector, inputting the structural features and the semantic features into a preset diffusion generation network in stages, and determining an initial image; obtaining user feedback input information according to the initial image; inputting the user feedback input information into a preset feedback interpreter, and determining an updated guide vector; and according to the updated guide vector, inputting the structural features and the semantic features into a preset diffusion generation network in stages, and determining a target diffusion image. According to the method and the device, the multi-modal data is guided in a staged manner, a user feedback closed loop is realized, and the consistency, controllability and user interactivity of the generated image are improved.
Owner:SHANGHAI JIAOTONG UNIV

Compositional text-to-image generation with dense blob representations

Systems and methods are disclosed that generate dense blob representations such as blob parameters and blob descriptions, and use the dense blob representations to generate images. For example, embodiments of the present disclosure may decompose a scene into visual primitives (e.g., dense blob representations) and based on the blob representations, embodiments of the present disclosure develop a blob-grounded text-to-image diffusion model (BlobGEN) for compositional generation. For example, in some embodiments, a new masked cross-attention module may be introduced to disentangle the fusion between blob representations and visual features. In some embodiments, to leverage the compositionality of large language models (LLMs), a new in-context learning approach may be introduced to generate blob representations from text prompts.
Owner:NVIDIA CORP

Intelligent building texture repairing and beautifying method fusing deep learning model

The invention is suitable for the technical field of building model data processing, and provides an intelligent building texture repairing and beautifying method fusing a deep learning model, and the method comprises the steps: firstly, carrying out the recombination of the texture of a three-dimensional building model, and generating a to-be-processed wall surface and roof texture picture; a to-be-repaired area, needing to be repaired, of the building texture is extracted by combining low-rank matrix decomposition and an SAM image segmentation algorithm, and then whether an LAMA image repair model or a PatchMatch image repair algorithm is used for texture processing is determined according to the proportion of the to-be-repaired area in a whole texture image so as to achieve the maximum repair effect. And finally, carrying out texture distortion repair and material beautification by using an image diffusion model, a ControlNet control chart and cue words, and outputting beautified wall surface and roof textures. The problems of texture shielding, garland, noise, window distortion and the like can be solved at the same time, automatic detection and automatic texture repairing and beautifying of the texture shielding garland area are achieved, and the time consumption of texture editing processing in the building monotonized modeling process is remarkably shortened.
Owner:WUDA GEOINFORMATICS CO LTD

Deep hash image retrieval method based on diffusion model for power grid defect maintenance

The invention relates to the field of power grid defect retrieval, in particular to a diffusion model-based deep hash image retrieval method for power grid defect maintenance, which comprises the following steps of: 1, performing fusion coding by inputting text data and image data, and constructing an initial hash code generation model; 2, using a Pair-wise loss function to optimize the distribution of sample pairs in a hash space, introducing a quantization loss function, generating an efficient binary hash code, and generating a high-quality binary hash code; 3, constructing a Hash code-image latent diffusion model, performing diffusion generation by encoding and decoding the Hash code / image to a continuous latent space, enabling the Hash code to correspond to the image in a generative manner, and directly fitting spatial distribution; and 4, defining a loss function of the Hash code-image diffusion model, generating a high-quality Hash code and an image, and obtaining a power grid defect type in a mode of searching images by images. Auxiliary training is carried out through fusion of text features and image features, so that the semantic features understand the images more deeply.
Owner:STATE GRID SHANDONG ELECTRIC POWER CO JIMO POWER SUPPLY CO

Training-free consistent text-to-image generation

Embodiments of the present disclosure relate to training-free consistent text-to-image generation. A pre-trained text-to-image diffusion model is leveraged to generate images depicting a consistent subject for diverse prompts describing scenes. Inputs to the model are a text description of at least one subject with prompts (scene text descriptions) describing scenes, where each prompt is associated with a different generated image and the text description is used for all images that depict the subject. Internal activations (intermediate data) computed by the model during generation of the different images are shared for generation of the different images. A subject-driven shared attention block and correspondence-based feature injection are incorporated into the model to promote subject consistency within each image and / or between images. Additionally, layout diversity is encouraged while maintaining subject consistency. The model achieves state-of-the-art performance on subject consistency and text alignment, without requiring any optimization and naturally extends to multi-subject scenarios.
Owner:NVIDIA CORP

Driver driving state image data generation method based on conditional diffusion model

The invention belongs to the technical field of intelligent automobiles, and particularly relates to a driver driving state image data generation method based on a conditional diffusion model. Comprising the following steps: step 1, acquiring original driver driving state data and constructing camera parameters; 2, potential representation extraction and geometric control vector construction of an original image; step 3, potential representation generation of the target image; step 4, image reconstruction and post-processing; according to the method, an image diffusion generation network is taken as a core, a geometric transformation relation between an original camera and a target camera is combined, and generalization generation from a single image to a multi-view image is completed through a potential spatial modeling and condition guidance mechanism; the objective of the invention is to synthesize lifelike images under configuration of other visual angles and camera parameters by using a small amount of original images.
Owner:JILIN UNIVERSITY +1

Automobile marketing video generation method and system based on diffusion model

The invention discloses an automobile marketing video generation method and system based on a diffusion model. The method comprises the following steps: firstly, performing deep feature extraction on an official photo of a target vehicle, and constructing a brand landmark feature parameter set; generating a parameterized motion template based on a marketing script, and generating scene static graphs in batches through an image diffusion model; performing brand consistency closed-loop verification on the static graph by adopting a feature matching algorithm, and automatically regenerating if the brand consistency does not reach the standard; encoding the static image and the motion template which pass the verification, inputting the encoded static image and motion template into a video diffusion model, constraining time sequence consistency through a cross-frame attention mechanism, and generating a video clip; and based on scene similarity, adaptively selecting a transition mode to perform intelligent splicing, and adding brand elements to synthesize a final video. According to the invention, brand security, parameterized fine control and intelligent reuse of generated assets of the automobile marketing video are realized, and the generation efficiency and quality are remarkably improved.
Owner:TIANJIN AUTOHOME DATA INFORMATION TECH CO LTD

Image generation method and device, equipment and storage medium

The invention provides an image generation method and apparatus, a device and a storage medium. The method comprises the steps of obtaining target mask images corresponding to at least two target elements included in a to-be-processed image; for each target element, generating a cue word corresponding to the target element, the cue word comprising a target attribute value of the target element; and inputting the to-be-processed image, the target mask image corresponding to each target element and the cue word corresponding to each target element into an image diffusion model to obtain a target image corresponding to each target element generated by the image diffusion model, the attribute value of the corresponding target element in the target image is the target attribute value included in the cue word of the corresponding target element. According to the embodiment of the invention, a large batch of sample images can be quickly generated, and the generation efficiency of the sample images is improved.
Owner:JINAN BOGUAN INTELLIGENT TECH CO LTD

Extreme image coding and decoding method and device based on diffusion model

The invention discloses an extreme image coding and decoding method based on a diffusion model. The method comprises the following steps: S100, a compression and pre-decoding module at a coding end extracts, compresses and transmits image information by using a neural coding network; s200, a compression and pre-decoding module of the decoding end preliminarily decodes the received image information into a content variable aligned with the potential diffusion space by using a neural decoding network; and S300, a denoising module uses the content variable as a control condition and a text graph diffusion model as prior information, and reconstructs an image from random noise through step-by-step denoising. The method can utilize the strong generation capability of the pre-trained text image diffusion model to realize perception-friendly reconstruction consistent with the original image at an extremely low bit rate, has the characteristics of ultrahigh compression ratio and good image reconstruction effect, can carry out image transmission under an extremely low bandwidth condition, and has a wide application prospect. The method can be widely applied to the fields of smart phone satellite communication, short-wave communication, unmanned equipment remote operation and the like.
Owner:XI AN JIAOTONG UNIV

Printing-shooting process image degradation simulation method, device and equipment based on image-to-image diffusion model

The invention discloses a printing-shooting process image degradation simulation method, device and equipment based on a graph-to-graph diffusion model, and aims to solve the problems that an existing degradation simulation method is greatly different from a real physical process and cannot be accurately controlled. The method comprises the following steps: firstly, constructing a printing-shooting data set containing a plurality of printing parameters and shooting parameters; then coding the physical parameters into conditional control vectors, and injecting the conditional control vectors into a Unet network of a graph-to-graph diffusion model for training; when an image is generated, an innovative double-flow noise layer structure is adopted, the structure combines a first branch adopting a traditional digital simulation method and a second branch adopting the pre-training diffusion model in parallel, and one of the first branch and the second branch is selected to be output according to a preset probability. According to the method, a highly vivid degraded image can be generated, so that the robustness and accuracy of downstream tasks (such as deep watermarking and image recognition) in a real printing-shooting scene are remarkably improved.
Owner:CHANGSHA YIYUE TECHNOLOGY CO LTD

Three-Dimensional Diffusion Models

Provided are systems and methods to perform novel view synthesis of a three-dimensional (3D) scene with a machine-learned diffusion model. Example implementations of the proposed models may be referred to as “3D Diffusion Models” or 3DiM. The models described herein can be or include an image-to-image diffusion model that takes one or more (e.g., a single) reference views and one or more (e.g., a single) relative poses as input and generates the target view. Thus, the machine-learned diffusion models described herein can perform novel view synthesis from as few as a single image.
Owner:GOOGLE LLC

Synthetic license plate data generation

A method or system for enhancing vehicle identification accuracy. The system identifies a gap in a training dataset for training a license plate identification model. The gap represents underrepresented visual characteristics in misidentified license plates. A guidance prompt is generated based on the visual characteristics of the misidentified license plate. Condition embeddings are then generated from the guidance prompt, which are used to condition a diffusion model to create synthetic license plate images. The diffusion model is trained to receive a real license plate image, encode it into a vector, apply forward diffusion to add noise, and then apply reverse diffusion to remove the noise, resulting in a denoised vector that represents a synthetic license plate image conditioned by the condition embeddings. The denoised vector is then decoded to produce the synthetic license plate image. The synthetic images are then used to retrain the license plate identification model.
Owner:METROPOLIS IP HOLDINGS LLC

Generating objects of mixed concepts using text-to-image diffusion models

Generating an object using a diffusion model includes obtaining a first input and a second input, and synthesizing an output object from the first input and the second input. The synthesizing of the output object includes generating a layout of the output object from the first input, injecting the second input as a content conditioner to the layout of the output object, and de-noising the layout of the output object injected with the content conditioner to generate a content of the output object.
Owner:LEMON INC(GB)

Text-to-three-dimensional surface generation method and system based on two-dimensional Gaussian surface element

PendingCN121259176A3D-image rendering3D modellingGeometric consistencyStructure from motion
The invention discloses a text-to-three-dimensional surface generation method and system based on a two-dimensional Gaussian surface element, and belongs to the technical field of three-dimensional reconstruction, and the method comprises the steps: representing a three-dimensional object surface as a Gaussian surface element on a local tangent plane; receiving a text description, generating normal diagrams and texture diagrams of a plurality of visual angles based on a pre-training text-to-image diffusion model, and generating a preliminary background mask at the same time; determining the center of sphere and the maximum radius according to the extrinsic parameters of the camera, and randomly generating Gaussian surface elements in the spherical range; removing the surface elements falling in the preliminary background mask area; on the basis of a two-dimensional Gaussian dot drawing renderer, the normal graph and the texture graph are rendered from multiple perspectives, and curvature consistency regular loss and surface convergence constraint loss are calculated; geometric and texture parameters of Gaussian surface elements are optimized through back propagation, iteration is carried out until convergence, and a final three-dimensional model is output. According to the method, the high-fidelity three-dimensional surface can be generated from the text without SfM (Structural Recovery Motion) initialization, and the method has relatively high geometric consistency and generation speed.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

SAR image denoising method and system

The invention relates to the technical field of digital image processing, in particular to an SAR image denoising method and system. The method comprises the following steps: acquiring a noisy SAR image; obtaining a pre-trained diffusion model; the diffusion model comprises de-noising processes of T time steps; uniformly sampling K time steps from the T time steps; wherein K is equal to 1 / 40 T to 1 / 20 T; performing a de-noising process of K time steps on the SAR image with noise to obtain a de-noised SAR image; the diffusion model comprises a de-noising network established based on a U-Net architecture, and the de-noising network is used for predicting a noise component and variance required by reverse sampling; the denoising network comprises an encoder, and the encoder is used for extracting a multi-scale feature map from a noisy SAR image and decomposing any feature map into a low-frequency approximate sub-band and three high-frequency detail sub-bands based on Haar wavelet transform. By adopting the scheme, the de-noising precision, de-noising efficiency and reasoning stability of the diffusion model can be improved.
Owner:BEIJING INST OF TECH

Text generation image space guiding method and system based on vocabulary mapping graph

The invention discloses a text generation image space guiding method based on a vocabulary mapping graph. The method comprises the following steps: converting a serialized text prompt input by a user into a structured two-dimensional vocabulary map (Lexical Map) by utilizing a large language model, wherein the map explicitly encodes semantic and spatial position information of an object; in the inference process of the text map diffusion model, weight scores of self-attention and cross attention are adaptively adjusted according to the layout of the vocabulary mapping graph through an attention rearrangement mechanism, so that image generation is guided. According to the method, the diffusion model does not need to be additionally trained, the problem that an existing model is difficult to process complex spatial relation prompts is solved in a plug-and-play mode, the alignment precision of generated images and text description in spatial layout and quantity is remarkably improved, and meanwhile high generation quality and reasoning speed are kept.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS