Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

105 results about "Image diffusion" patented technology

System and Method for Event-Driven Video Synthesis Using Textual Descriptions

A video generation framework that is controllable, unsupervised and based on events (CUBE) includes an event camera, which captures changes in light intensity at each pixel of a scene asynchronously and generates event camera data. A text-to-image diffusion model that is conditioned on textual descriptions integrates the event camera data to control video synthesis. Further, an edge extraction module translates event data into a format usable by the text-to-image diffusion model, whereby the diffusion model synthesizes detailed and contextually accurate videos based on textual prompts. Further, an improved system (CUBE Plus) includes a content frame identification module which selectively identifies and uses only the most information-rich event segments of the event camera data to drive cross-frame attention, and an event driven attention mechanism that allows the framework to focus on event-dense moments.
Owner:THE UNIVERSITY OF HONG KONG

Automatic quality control of image diffusion processing

A computer system that performs quality control (QC) on images associated with diffusion and structural magnetic resonance imaging (MRI) is described. This computer may include: a computation device that executes program instructions; and memory that stores the program instructions. During operation, the computer system may automatically perform a set of validation operations, where, when one or more of the validation operations fails, the images are rejected. Moreover, the set of validation operations may include: performing QC on brain-tissue segmentation; performing QC on diffusion MRI processing; and performing QC on bundles determined from the images using a tractometry technique.
Owner:IMEKA SOLUTIONS INC

Avatar Generation using Image Diffusion Models

A method of generating a 3-dimensional representation of a subject is provided. The method includes receiving one or more descriptions characterizing the subject. The method also includes inputting the one or more descriptions characterizing the subject into a first specialized network of a machine learning model to generate one or more images depicting the subject according to the one or more descriptions. The method further includes inputting the generated one or more images to a second specialized network of the machine learning model to generate the 3-dimensional representation of the subject according to the one or more descriptions characterizing the subject.
Owner:GOOGLE LLC

Method for embedding robust watermark in diffusion model generated image

The invention belongs to the field of image processing, and relates to a method for embedding a robust watermark in a diffusion model generated image, which comprises the following steps of: acquiring cue words, initial Gaussian noise and watermark information, and inputting the cue words, the initial Gaussian noise and the watermark information into a trained diffusion model based on watermark embedding to obtain a watermark-embedded image; the training process of the diffusion model comprises the following steps: acquiring cue words, initial Gaussian noise and watermark information, and inputting the cue words, the initial Gaussian noise and the watermark information into an encoder to obtain potential vectors; inputting the potential vector into a self-attention module to obtain an embedded position vector; embedding the watermark information into the potential vector according to the embedding position vector; inputting the potential vector embedded with the watermark information into a decoder to obtain an image embedded with the watermark; extracting watermark information from the image embedded with the watermark; updating parameters of a diffusion model according to the image embedded with the watermark and the extracted watermark information until a trained diffusion model is obtained; according to the method, the embedding position is selected by combining the potential of the diffusion model and the accuracy of the self-attention mechanism, so that efficient watermark embedding and extraction are realized.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Single view reconstruction and rendering method

The invention discloses a single view reconstruction and rendering method, which belongs to the field of view reconstruction and rendering, and comprises the following steps of: firstly, generating multi-view feature representation with strong geometric consistency from a single input image by introducing an image diffusion module of a cross attention mechanism; a point cloud reconstruction module with self-attention and cross-attention is utilized, and multi-view information is fused to reconstruct an accurate three-dimensional point cloud; and finally, constructing differentiable three-dimensional Gaussian representation based on the point cloud, rendering the differentiable three-dimensional Gaussian representation, and outputting a new view angle image, a normal map and a depth map. According to the system, a staged training strategy is adopted, and a composite loss function including multi-scale bidirectional consistency smooth loss and feature consistency loss is innovatively used for optimization. According to the method, the problems of low geometric accuracy, multi-view inconsistency, detail missing and the like in single-view reconstruction are effectively solved, and the method can be widely applied to the fields of virtual reality, digital twinning, cultural heritage digitization and the like.
Owner:北京渲光科技有限公司

Text-to-image diffusion model rearchitecture

Described is a system for improving machine learning models. In some cases, the system improves such models by identifying a performance characteristic for machine learning model blocks in an iterative denoising process of a machine learning model, connecting a prior machine learning model block with a subsequent machine learning model block of the machine learning model blocks within the machine learning model based on the identified performance characteristic, identifying a prompt of a user, the prompt indicative of an intent of the user for generative images, and analyzing data corresponding to the prompt using the machine learning model to generate one or more images, the machine learning model trained to generate images based on data corresponding to prompts.
Owner:SNAP INC

Method for generating aviation lifelike tree sample and identifying tree species by using image diffusion model

The invention discloses a method for generating an aviation lifelike tree sample and identifying a tree species by using an image diffusion model. The method comprises the following steps: constructing a multi-modal data set; training an image-text comparison model and a Unet denoising network model; outputting predicted noise in the trained Unet denoising network model, and finally generating final image reconstruction which is in semantic alignment with the input text description; the obtained single tree crown images are synthesized; and constructing a YOLOv11 detection model, and training the YOLOv11 detection model by using the synthesized forest image to realize tree species identification. According to the method, language semantics and a diffusion model are combined to generate a vivid tree sample for capturing specific features of species and seasonal phenological changes, and the vivid tree sample is synthesized into a high-fidelity forest image, so that tree species identification in an aerial image is enhanced, and effective crown detection and accurate tree species identification are realized.
Owner:NANJING FORESTRY UNIV

Sea surface temperature image diffusion completion method and system based on two-stage fusion constraint

The invention discloses a sea surface temperature image diffusion completion method and system based on dual-stage fusion constraint, and the method comprises the steps: firstly constructing sea area multi-source priori knowledge features to guide the reconstruction of sea surface temperature through integrating three types of heterogeneous data of geography, space-time and power; deep fusion is carried out on the sea area multi-source priori knowledge features through a cross-modal feature fusion module to obtain fused multi-source priori knowledge, and the fused multi-source priori knowledge is fused with structural knowledge to form comprehensive condition features used for guiding the generation process of a conditional diffusion model; and finally, realizing refined reconstruction of the missing region based on a conditional diffusion model, and outputting a completion result. And obtaining a final complete image according to the complementation result, the sea surface temperature image to be complemented and the missing mask. According to the invention, the accuracy of sea surface temperature image completion is improved.
Owner:OCEAN UNIV OF CHINA

Deep hash image retrieval method based on diffusion model for power grid defect maintenance

The invention relates to the field of power grid defect retrieval, in particular to a diffusion model-based deep hash image retrieval method for power grid defect maintenance, which comprises the following steps of: 1, performing fusion coding by inputting text data and image data, and constructing an initial hash code generation model; 2, using a Pair-wise loss function to optimize the distribution of sample pairs in a hash space, introducing a quantization loss function, generating an efficient binary hash code, and generating a high-quality binary hash code; 3, constructing a Hash code-image latent diffusion model, performing diffusion generation by encoding and decoding the Hash code / image to a continuous latent space, enabling the Hash code to correspond to the image in a generative manner, and directly fitting spatial distribution; and 4, defining a loss function of the Hash code-image diffusion model, generating a high-quality Hash code and an image, and obtaining a power grid defect type in a mode of searching images by images. Auxiliary training is carried out through fusion of text features and image features, so that the semantic features understand the images more deeply.
Owner:STATE GRID SHANDONG ELECTRIC POWER CO JIMO POWER SUPPLY CO

Driver driving state image data generation method based on conditional diffusion model

The invention belongs to the technical field of intelligent automobiles, and particularly relates to a driver driving state image data generation method based on a conditional diffusion model. Comprising the following steps: step 1, acquiring original driver driving state data and constructing camera parameters; 2, potential representation extraction and geometric control vector construction of an original image; step 3, potential representation generation of the target image; step 4, image reconstruction and post-processing; according to the method, an image diffusion generation network is taken as a core, a geometric transformation relation between an original camera and a target camera is combined, and generalization generation from a single image to a multi-view image is completed through a potential spatial modeling and condition guidance mechanism; the objective of the invention is to synthesize lifelike images under configuration of other visual angles and camera parameters by using a small amount of original images.
Owner:JILIN UNIVERSITY +1

Automobile marketing video generation method and system based on diffusion model

The invention discloses an automobile marketing video generation method and system based on a diffusion model. The method comprises the following steps: firstly, performing deep feature extraction on an official photo of a target vehicle, and constructing a brand landmark feature parameter set; generating a parameterized motion template based on a marketing script, and generating scene static graphs in batches through an image diffusion model; performing brand consistency closed-loop verification on the static graph by adopting a feature matching algorithm, and automatically regenerating if the brand consistency does not reach the standard; encoding the static image and the motion template which pass the verification, inputting the encoded static image and motion template into a video diffusion model, constraining time sequence consistency through a cross-frame attention mechanism, and generating a video clip; and based on scene similarity, adaptively selecting a transition mode to perform intelligent splicing, and adding brand elements to synthesize a final video. According to the invention, brand security, parameterized fine control and intelligent reuse of generated assets of the automobile marketing video are realized, and the generation efficiency and quality are remarkably improved.
Owner:TIANJIN AUTOHOME DATA INFORMATION TECH CO LTD

Image generation method and device, equipment and storage medium

The invention provides an image generation method and apparatus, a device and a storage medium. The method comprises the steps of obtaining target mask images corresponding to at least two target elements included in a to-be-processed image; for each target element, generating a cue word corresponding to the target element, the cue word comprising a target attribute value of the target element; and inputting the to-be-processed image, the target mask image corresponding to each target element and the cue word corresponding to each target element into an image diffusion model to obtain a target image corresponding to each target element generated by the image diffusion model, the attribute value of the corresponding target element in the target image is the target attribute value included in the cue word of the corresponding target element. According to the embodiment of the invention, a large batch of sample images can be quickly generated, and the generation efficiency of the sample images is improved.
Owner:JINAN BOGUAN INTELLIGENT TECH CO LTD

Three-Dimensional Diffusion Models

Provided are systems and methods to perform novel view synthesis of a three-dimensional (3D) scene with a machine-learned diffusion model. Example implementations of the proposed models may be referred to as “3D Diffusion Models” or 3DiM. The models described herein can be or include an image-to-image diffusion model that takes one or more (e.g., a single) reference views and one or more (e.g., a single) relative poses as input and generates the target view. Thus, the machine-learned diffusion models described herein can perform novel view synthesis from as few as a single image.
Owner:GOOGLE LLC

Synthetic license plate data generation

A method or system for enhancing vehicle identification accuracy. The system identifies a gap in a training dataset for training a license plate identification model. The gap represents underrepresented visual characteristics in misidentified license plates. A guidance prompt is generated based on the visual characteristics of the misidentified license plate. Condition embeddings are then generated from the guidance prompt, which are used to condition a diffusion model to create synthetic license plate images. The diffusion model is trained to receive a real license plate image, encode it into a vector, apply forward diffusion to add noise, and then apply reverse diffusion to remove the noise, resulting in a denoised vector that represents a synthetic license plate image conditioned by the condition embeddings. The denoised vector is then decoded to produce the synthetic license plate image. The synthetic images are then used to retrain the license plate identification model.
Owner:METROPOLIS IP HOLDINGS LLC

Generating objects of mixed concepts using text-to-image diffusion models

Generating an object using a diffusion model includes obtaining a first input and a second input, and synthesizing an output object from the first input and the second input. The synthesizing of the output object includes generating a layout of the output object from the first input, injecting the second input as a content conditioner to the layout of the output object, and de-noising the layout of the output object injected with the content conditioner to generate a content of the output object.
Owner:LEMON INC(GB)

Text-to-three-dimensional surface generation method and system based on two-dimensional Gaussian surface element

PendingCN121259176A3D-image rendering3D modellingGeometric consistencyStructure from motion
The invention discloses a text-to-three-dimensional surface generation method and system based on a two-dimensional Gaussian surface element, and belongs to the technical field of three-dimensional reconstruction, and the method comprises the steps: representing a three-dimensional object surface as a Gaussian surface element on a local tangent plane; receiving a text description, generating normal diagrams and texture diagrams of a plurality of visual angles based on a pre-training text-to-image diffusion model, and generating a preliminary background mask at the same time; determining the center of sphere and the maximum radius according to the extrinsic parameters of the camera, and randomly generating Gaussian surface elements in the spherical range; removing the surface elements falling in the preliminary background mask area; on the basis of a two-dimensional Gaussian dot drawing renderer, the normal graph and the texture graph are rendered from multiple perspectives, and curvature consistency regular loss and surface convergence constraint loss are calculated; geometric and texture parameters of Gaussian surface elements are optimized through back propagation, iteration is carried out until convergence, and a final three-dimensional model is output. According to the method, the high-fidelity three-dimensional surface can be generated from the text without SfM (Structural Recovery Motion) initialization, and the method has relatively high geometric consistency and generation speed.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

SAR image denoising method and system

The invention relates to the technical field of digital image processing, in particular to an SAR image denoising method and system. The method comprises the following steps: acquiring a noisy SAR image; obtaining a pre-trained diffusion model; the diffusion model comprises de-noising processes of T time steps; uniformly sampling K time steps from the T time steps; wherein K is equal to 1 / 40 T to 1 / 20 T; performing a de-noising process of K time steps on the SAR image with noise to obtain a de-noised SAR image; the diffusion model comprises a de-noising network established based on a U-Net architecture, and the de-noising network is used for predicting a noise component and variance required by reverse sampling; the denoising network comprises an encoder, and the encoder is used for extracting a multi-scale feature map from a noisy SAR image and decomposing any feature map into a low-frequency approximate sub-band and three high-frequency detail sub-bands based on Haar wavelet transform. By adopting the scheme, the de-noising precision, de-noising efficiency and reasoning stability of the diffusion model can be improved.
Owner:BEIJING INST OF TECH

Text generation image space guiding method and system based on vocabulary mapping graph

The invention discloses a text generation image space guiding method based on a vocabulary mapping graph. The method comprises the following steps: converting a serialized text prompt input by a user into a structured two-dimensional vocabulary map (Lexical Map) by utilizing a large language model, wherein the map explicitly encodes semantic and spatial position information of an object; in the inference process of the text map diffusion model, weight scores of self-attention and cross attention are adaptively adjusted according to the layout of the vocabulary mapping graph through an attention rearrangement mechanism, so that image generation is guided. According to the method, the diffusion model does not need to be additionally trained, the problem that an existing model is difficult to process complex spatial relation prompts is solved in a plug-and-play mode, the alignment precision of generated images and text description in spatial layout and quantity is remarkably improved, and meanwhile high generation quality and reasoning speed are kept.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Detecting defects using a text-to-image diffusion model

Methods for fine-tuning convolutional neural networks of text-to-image diffusion models in the context of identifying defects of manufactured products within images of those products are disclosed. Images of manufactured images having various scratches, dings, or other defects are provided to the model along with words or phrases indicating the presence of a defect. The model then learns to identify portions of the overall image that include the defect. Learning of this type of task is based on using segmentation masks corresponding to the images, which are then used along with cross-attention maps of the model in order to compute an average defect mask loss parameter of the model. By computing this parameter and applying it in updating the weights of the model, the model can be fine-tuned to detect defects of manufactured products.
Owner:ROBERT BOSCH GMBH

Generative machine learning models for generating roof damage images

Techniques are described herein for generating synthetic images of damaged roofs using generative machine learning (ML) models. In various examples, a generative ML system receives text input describing roof attributes and / or damage characteristics for a synthetic image, which may be processed by a text encoder to determine a set of text embeddings. The text embeddings may be used as conditioning data for a generative ML model, such as an image diffusion model, to produce realistic synthetic images of damaged roofs. These generated images can be used to augment training datasets for additional ML models focused on roof damage detection and assessment, addressing the challenges of limited real-world training data.
Owner:STATE FARM MUTAL AUTOMOBILE INSURANCE COMPANY

Image diffusion enhancement method and system based on decoupled guidance and reprojection refinement

The present application relates to the technical field of image enhancement processing, in particular to an image diffusion enhancement method and system based on decoupling guidance and re-projection refinement. The method comprises: extracting mutually orthogonal physical degradation latent variables and semantic content latent variables through a two-way decoupling encoder; constructing a diffusion inverse process generation model containing a double conditional time sequence perception U-Net network, introducing physical consistency loss through a physical modulation module and a hierarchical cross-attention mechanism, generating a preliminary enhanced image, inputting the preliminary enhanced image into a re-projection refinement network for latent space residual correction; constructing a semantic-guided detail gain network to generate a detail gain map; performing detail enhancement through wavelet transform and the detail gain map to obtain a detail enhanced image output and generate a visual analysis report. The decoupling guidance ensures that the image generation process is strictly constrained, and the re-projection refinement corrects the deviation, enhances the image diffusion enhancement effect, and improves the image quality.
Owner:GUIZHOU INST OF TECH +1

Generating memes and enhanced content in electronic communication

A method and system for generating and displaying candidate memes within electronic communications. The method includes accessing a portion of an electronic communication, determining a textual and visual response, and generating a candidate meme that includes these responses. The meme is then provided as a selectable option within the communication. Upon selection, the meme is displayed within the communication. The visual response may be based on a text-to-image diffusion model, context, user profile, or meme template. The method includes monitoring the communication for changes in topic or end of communication and resets parameters accordingly. The method can also include incorporating information about live events or specific subjects relevant to the communication.
Owner:ADEIA GUIDES INC

Text-to-image diffusion model reconstruction

A system for improving a machine learning model is described. In some cases, the system improves such a model by identifying performance characteristics of machine learning model blocks during iterative denoising of the machine learning model; connecting a preceding one of the machine learning model blocks within the machine learning model to a following one of the machine learning model blocks based on the identified performance characteristics; identifying a prompt of the user, the prompt indicating the user's intent to generate the image; and analyzing the data corresponding to the cues using a machine learning model to generate one or more images, the machine learning model being trained to generate the images based on the data corresponding to the cues.
Owner:SNAP INC

Video generation method and training method of video generation model

The invention provides a video generation method and a training method of a video generation model, and relates to the technical field of image processing, in particular to the technical field of artificial intelligence and deep learning. According to the specific implementation scheme, a first image sequence and a second image sequence are obtained, the first image sequence comprises a first image frame needing to be referenced in video generation, and the second image sequence comprises a second image frame generated by pure noise; splicing the first image sequence and the second image sequence along a time axis to obtain a third image sequence; and generating a target video through the target image diffusion model according to the prompt information and the third image sequence.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Personalized image generation method and system based on maximum difference anchoring and mask attention guidance

The invention discloses a personalized image generation method and system based on maximum difference anchoring and mask attention guidance. In order to solve the problems of concept excessive generalization and background interference distortion of an existing text-to-image diffusion model in a user-defined concept and general concept combined scene, an end-to-end dynamic constraint framework is constructed. The method comprises the following steps: firstly, generating a candidate set related to a user concept through a multi-modal large model; a plurality of anchor point concepts with the maximum semantic difference are selected through K-means clustering; secondly, designing a double-stage training mechanism, wherein in the first stage, attention loss guided by an SAM mask is adopted to suppress background interference, and in the second stage, maximum differentiation anchor point constraint loss is introduced to optimize embedded distribution; and finally, a high-quality image fusing the custom concept and the general concept is generated. According to the method provided by the invention, unique features of a user-specified concept can be better reserved in the generated image, meanwhile, the deficiency or distortion of a general concept is avoided, and the consistency and the structural accuracy of the generated image are improved.
Owner:ZHEJIANG GONGSHANG UNIVERSITY

Video generation method, electronic device and computer-readable storage medium

The present disclosure relates to the technical fields of video processing technology and computers. Disclosed are a video generation method, an electronic device and a computer-readable storage medium. The method comprises: generating a first video on the basis of prompt text and noise data, wherein the prompt text is used for describing video content of a target video to be generated, and the data dimension of the noise data is the same as the video data dimension of the target video; and performing video generation processing on the first video on the basis of a target video generation model, so as to obtain the target video, wherein a video diffusion model and an image diffusion model are integrated into the target video generation model, and the initial time step of the video diffusion model is the same as the initial time step of the image diffusion model. The present disclosure solves the technical problem of the quality of videos generated by video generation models in the related art being relatively poor.
Owner:ALIBABA (CHINA) CO LTD

A backdoor detection method of a text-to-image diffusion model based on attention shift

The application discloses a backdoor detection method of a text-to-graph diffusion model based on attention transfer, comprising the following steps: obtaining an input text; inputting the input text into a text encoder of the text-to-graph diffusion model to obtain a last layer text embedding and a final text embedding; determining a first backdoor result according to a self-attention score matrix between tokens obtained from the last layer text embedding and a starting token; inputting the final text embedding into a conditional diffusion module of the text-to-graph diffusion model, obtaining a cross-attention score matrix according to the final text embedding based on a cross-attention mechanism, and determining a second backdoor result according to the cross-attention score matrix. The backdoor detection method has universality and high efficiency, can reduce the consumption of computing resources while ensuring the detection accuracy, and can provide solid technical support for the safe application of the text-to-graph diffusion model.
Owner:XIDIAN UNIV +1

Method and device for automatically identifying benign and malignant kidney cystic lesions based on magnetic resonance image

The invention relates to a method and a device for automatically identifying benign and malignant kidney cystic lesions based on magnetic resonance images. The method comprises the following steps: S1, preprocessing and standardizing a T2 weighted image, a diffusion weighted image, an apparent diffusion coefficient image, a T1 weighted image, a skin medullary phase image, a parenchyma phase image and an excretion phase image; s2, respectively training automatic segmentation models corresponding to different images in a targeted manner, and predicting a focus by using the automatic segmentation models; s3, extracting morphological features, first-order features and textural features of the lesions from all the lesions, wherein the morphological features, the first-order features and the textural features comprise features of capsule walls, partitions and nodules of the lesions; screening the extracted features, and constructing a classification model by using the features with good robustness; and S4, preprocessing and standardizing the image of the current patient, respectively inputting the image into each corresponding automatic segmentation model, and operating the segmentation model and the classification model to realize benign and malignant recognition based on image recognition. According to the invention, integrated and automatic benign and malignant accurate diagnosis of kidney cystic lesions is realized.
Owner:THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL

Dynamic cartoon generation method, device and equipment based on large model

The invention relates to the technical field of video generation, particularly provides a dynamic cartoon generation method, device and equipment based on a large model, and aims to solve the problem of low dynamic cartoon generation efficiency. The method comprises the following steps: acquiring an image material, an image description, a script and a scene description; according to the image material and the image description, training a preset image diffusion model to obtain a personalized image diffusion model; according to the script and the scene description, cue words are obtained; obtaining a key image according to the cue word and the personalized image diffusion model; obtaining a video clip according to the prompt word and the key image; and editing the video clip to obtain the dynamic cartoon. And the generation efficiency of the dynamic cartoon is improved.
Owner:SHANGHAI YUNCHONG ENTERPRISE DEV CO LTD

Dressing pedestrian re-identification method based on image diffusion

The invention relates to the technical field of computer vision, and discloses a clothes changing pedestrian re-identification method based on image diffusion, comprising: training an initial image generation model based on a pedestrian image set and a training clothes changing text to obtain a target image generation model, the initial image generation model adopting a truncation diffusion strategy; inputting the pedestrian image and the corresponding target clothes changing text into a target image generation model to obtain a clothes changing pedestrian image; identity consistency filtering is carried out based on the clothes changing pedestrian image and the corresponding pedestrian image, and a training image set is screened; and training a clothes changing pedestrian re-identification model based on the training image set, and identifying the to-be-identified image by using the trained re-identification model to obtain an identification result. According to the method, external prior information is not needed, training data are expanded only by using the pedestrian image and the training clothes changing text, the data acquisition cost is reduced, and the efficiency, accuracy and robustness of clothes changing pedestrian re-identification are improved through truncation diffusion, reverse denoising and identity filtering.
Owner:PENG CHENG LAB +1