Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

21 results about "Image diffusion" patented technology

Detecting defects using a text-to-image diffusion model

PendingCN122175854AImage enhancementImage analysisImage diffusionComputational model
Methods for fine-tuning convolutional neural networks of text-to-image diffusion models in the context of identifying defects of manufactured products within images of those products are disclosed. Images of manufactured images having various scratches, dings, or other defects are provided to the model along with words or phrases indicating the presence of a defect. The model then learns to identify portions of the overall image that include the defect. Learning of this type of task is based on using segmentation masks corresponding to the images, which are then used along with cross-attention maps of the model in order to compute an average defect mask loss parameter of the model. By computing this parameter and applying it in updating the weights of the model, the model can be fine-tuned to detect defects of manufactured products.
Owner:ROBERT BOSCH GMBH

Lidar point cloud generation method and system based on pre-trained image diffusion model

This invention discloses a method and system for generating LiDAR point clouds based on a pre-trained image diffusion model. This invention enables unconditionally controlled generation of LiDAR point clouds in real-world scenes, conditionally controlled generation of LiDAR point clouds based on camera images, and LiDAR point cloud completion under zero-shot learning conditions. By introducing a dual-spatial consistency optimization strategy at both the 2D and 3D levels, the pre-trained Stable Diffusion 3 model can perceive spatial information and geometric consistency. A physically guided control network (ControlNet) is designed to endow the model with the ability to generate point clouds for specific scenes, thereby supporting the conversion task from camera images to LiDAR point clouds. Extensive experiments in a 64-line LiDAR scenario demonstrate the significant effectiveness and application potential of the proposed method in both conditional and unconditional generation tasks.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Backdoor attack methods, systems, and media based on text-to-image diffusion models with multi-object semantic coexistence.

PendingCN122090449AImproved visual concealmentImprove attack success rateBiological modelsCharacter and pattern recognitionData setImage diffusion
This invention discloses a backdoor attack method, system, and medium based on multi-object semantic coexistence in text-to-image diffusion models. This method utilizes the perspective of multi-object semantic coexistence in text-to-image diffusion models to develop MOBA backdoor attack schemes. First, by constructing a trigger alignment dataset and optimizing backdoor implantation and semantic preservation in parallel, the attack success rate and visual concealment are effectively improved, reducing the impact of semantic corruption on attack effectiveness. Second, during model training, an attention-based decoupling backdoor enhancement mechanism is adopted. By decoupling the attention regions of different objects, the semantic integrity of the input prompt is maintained while the backdoor is activated and the attention distribution is reasonably adjusted to reduce the impact of generation bias on attack concealment. The method of this invention ensures both visual concealment and model performance under benign input while achieving efficient backdoor attacks.
Owner:HUNAN UNIV OF SCI & TECH SANYA RES INST

A dementia classification method based on knowledge-guided image diffusion reasoning

PendingCN122337641AImage diffusionNeuro-degenerative disease
This invention relates to the field of multimodal data fusion and neurodegenerative disease classification, specifically a dementia classification method based on knowledge-guided image diffusion reasoning. First, initial features are extracted from magnetic resonance imaging (MRI) data using a 3D neural network, and a conditional diffusion model is used to generate category-enhanced image features. Simultaneously, a dementia knowledge graph is constructed from clinical metadata to integrate domain prior knowledge. Then, the enhanced image features and the knowledge graph embedding representation are input into a multimodal learning framework. This framework aligns and fuses information from the two modalities through bimodal consistency learning and modality-specific learning achieved through contrastive learning. Finally, the fused features are fed into a classifier for dementia classification. This invention provides a novel knowledge-guided deep learning framework for dementia classification, effectively overcoming the challenges of limited data and poor interpretability faced by traditional models, and significantly improving the accuracy and robustness of classification.
Owner:ZHEJIANG SCI-TECH UNIV

Method for creating generative time series dataset for change detection in remote sensing

PendingUS20260154950A1Image enhancementImage analysisData setImage diffusion
A system, method and non-transitory computer readable medium for generating validated remote sensing change images that includes a user input device for selecting high-resolution satellite images, and processing circuitry to generate a depth map and a semantic map from a static image. A change simulator determines candidate areas for change simulation and generates a change depth map and change mask focusing on objects removed from the static image. An image diffusion neural network applies a control network and stable diffusion to generate pre-change and post-change image tiles. Validation processing circuitry iterates through a validation process to validate the pair of change tiles to obtain a validated pair of change tiles and a validated change mask.
Owner:ELM INC

System, method, and non-transitory computer-readable media using generative artificial intelligence to optimize product search queries

ActiveUS12664576B22D-image generationCommerceImage diffusionEngineering
Methods and systems are provided for using generative AI to optimize product search queries. In embodiments described herein, product descriptions and product images for a plurality of products are obtained. A multi-modal style classification model classifies each product into a corresponding style of a plurality of styles based on the product's product description and product image. Relationships of each product to other products in the plurality of products are stored in a knowledge graph based on the corresponding style of each product and the corresponding product description of each product. An image is generated by a text-to-image diffusion model with a set of products of the plurality of products based on the relationships of each product of the plurality of products to other products in the plurality of products.
Owner:ADOBE INC

Medical image processing method, model training method, electronic device, and storage medium

PendingCN122453628AImaging processingImage diffusion
Embodiments of the present application provide a medical image processing method, a model training method, an electronic device and a storage medium. The medical image processing method comprises: predicting an sub-optimal noise center based on a source modality medical image and a reference target modality medical image corresponding to the source modality medical image; projecting the sub-optimal noise center to a predetermined hypersphere to obtain a region center of a sub-optimal noise region; screening noise in a predetermined noise image based on the region center to obtain guide noise; and performing image diffusion generation processing on the source modality medical image based on the guide noise to obtain a target modality medical image. The present scheme can improve the processing effect of cross-modality generation of medical images.
Owner:ALIBABA DAMOYUAN (BEIJING) TECH CO LTD

Image diffusion model processing method and device, electronic equipment and storage medium

PendingCN122368234APattern recognitionImage diffusion
This disclosure provides a method, apparatus, electronic device, and storage medium for processing text-based image diffusion models. The method includes: determining a target text-based image diffusion model corresponding to a target text-based image task, wherein the target text-based image diffusion model includes a first noise predictor but does not include a second noise predictor, and the target text-based image diffusion model satisfies the requirement that, during reverse diffusion training, the denoising effect of the first noise predictor can be trained to approach the denoising effect of the second noise predictor; determining target text prompt information corresponding to the target text-based image task, wherein the target text prompt information includes text information describing the content of the image to be generated in the target text-based image task; and generating an image based on the target text prompt information using the target text-based image diffusion model. This solution enables the text-based image diffusion model to significantly improve image generation speed while maintaining high-quality image generation.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Method, device and medium for generating extraterrestrial building image

This disclosure relates to a method, apparatus, device, and medium for generating images of extraterrestrial architecture, comprising: acquiring an input statement describing extraterrestrial architecture; converting the input statement into first prompts matching the extraterrestrial environment using a pre-set structured prompt framework; inputting a geospatial image of the extraterrestrial surface and the first prompts into a pre-trained image diffusion model; generating multiple sets of initial architectural images using the image diffusion model; generating second prompts describing architectural attributes using the structured prompt framework; and generating multiple sets of target architectural images using the image diffusion model based on the second prompts and the multiple sets of initial architectural images. This disclosure can effectively achieve automated and controllable image generation from macro-planning to micro-detailing, effectively improving the technical difficulties of poor professional adaptability, low generation efficiency, and limited solutions in traditional methods for visualizing extraterrestrial architecture.
Owner:ARCHITECTURAL DESIGN & RES INST OF TSINGHUA UNIV

Video Editing Timing Consistency Enhancement Method and System

This invention relates to the field of video editing technology and discloses a method and system for enhancing temporal consistency in video editing. The method includes: processing the original video frame sequence in the original video stream data using an image diffusion model to obtain a set of frame-by-frame editing results; processing the set of frame-by-frame editing results using both the image diffusion model and the video diffusion model to obtain the frame-by-frame latent variables output by the image diffusion model and the video latent variables output by the video diffusion model, and performing feature processing to obtain initial noise for denoising; inputting the initial noise into the video diffusion model to obtain approximate video frames; calculating the optical flow field between adjacent frames of the approximate video frames using a pre-trained optical flow network; determining the fusion latent variables based on the optical flow field; and decoding the fusion latent variables to obtain the final edited video frame sequence. This invention solves the problems of temporal instability and low processing efficiency in existing temporal consistency enhancement methods by using a multi-agent collaborative mechanism to perform closed-loop organization and iterative optimization of the generation, evaluation, and parameter adjustment processes.
Owner:HUNAN UNIV

A backdoor defense method for text image diffusion model based on Ollivier-Ricci curvature

The application discloses a backdoor defense method of text generation graph diffusion model based on Ollivier-Ricci curvature, and belongs to the technical field of artificial intelligence security. The method constructs a concept database to store text encoder features corresponding to predefined concepts and intermediate layer features of a denoising network; a to-be-detected text prompt is input into a diffusion model to extract text encoder features and intermediate layer features of the denoising network, and reference features are obtained based on semantic similarity matching. A K nearest neighbor graph is constructed based on the text encoder features, Ollivier-Ricci curvature is calculated, and abnormality is determined. If no abnormality is detected, global features are constructed at multiple time steps, curvature is calculated in the same way, and an abnormality proportion is counted. When abnormality is detected at any stage, fine-grained curvature analysis is performed on local intermediate layer features at multiple time steps, and whether a backdoor trigger is contained is determined according to the abnormality proportion, so that backdoor defense is achieved. The method is suitable for backdoor detection of various diffusion models and samplers.
Owner:BEIHANG UNIV

A method for blurry image diffusion restoration with background-guided constraints

This invention relates to the field of artificial intelligence image processing technology, specifically to a blurred image diffusion restoration method incorporating background-guided constraints. The method includes integrating background-guided constraints into the restoration loss and combining it with a pre-trained diffusion generation module to restore blurred images. To improve the accuracy and realism of the restoration, the similarity between the generated image and the background of the noisy image is used to guide the restoration process, forming background-guided constraints. This addresses the problem of unrealistic restored images caused by overly smooth backgrounds. To preserve more details and contour information in the image, the pre-trained diffusion generation module is used again to generate more image details based on the restored image. Finally, a blurred image restoration network incorporating background-guided constraints is proposed, which can effectively alleviate image degradation caused by noise, shooting conditions, and other factors, improving image quality. Image restoration can make the main subject of the image clearer.
Owner:HENAN UNIVERSITY +1

Text-to-image and image-to-text dual diffusion model

A computing system is provided, including one or more processing devices configured to, during an inferencing phase, receive an input image at a dual diffusion model. The one or more processing devices are further configured to process the input image at the dual diffusion model to compute output text. The one or more processing devices are further configured to output the output text. The dual diffusion model has been computed during a training phase by finetuning a text-to-image (T2I) diffusion model, the finetuning having been performed using a finetuning dataset that includes a plurality of image-text pairs. The finetuning has further utilized a loss function that includes an image distribution loss term and a text distribution loss term.
Owner:LEMON INC(GB)

An image anomaly detection method based on a text-to-image diffusion model

ActiveCN117218413BSemantic representationImage diffusion
This invention discloses an image anomaly detection method based on a text-to-image diffusion model. It utilizes a pre-trained text-to-image diffusion model to extract semantic representations of images, and simultaneously constructs an implicit description generator to generate text embeddings describing the images, which are then injected into the text-to-image diffusion model to enhance the semantic representations. A compression module is then used to compress the semantic representations to obtain the final image features. In the detection phase, the distance between the features of the image to be detected and the normal image features used during training is calculated as the basis for anomaly detection. Experimental results show that this invention achieves good results in image anomaly detection.
Owner:NANJING UNIV OF SCI & TECH

Watermarking method based on parameter-merging graph diffusion model

PendingCN122367704AAlgorithmWatermark method
The application provides a watermarking method based on a parameter-merging text-to-image diffusion model, which maps a watermark message to a bias offset through a message-conditioned bias modulator and acquires a weight compensation irrelevant to the watermark message, so that the watermark information can exist in the form of parameter offset independently of the parameters of the text-to-image diffusion model. Since the bias offset is dynamically generated by the message, the watermark message can be flexibly switched by only replacing the bias offset without retraining the model, solving the problem of fixed watermark information and inability to dynamically update in the prior art. By merging the bias offset and the weight compensation into the original parameters of the variational autoencoder decoder in the deployment stage, the network structure of the text-to-image diffusion model is not changed, and no additional watermark encoding module needs to be introduced or the inference path needs to be changed, so that zero inference overhead and high concealment are realized, solving the problem that the existing watermark technology is easily detected or increases the deployment burden due to structural changes.
Owner:MACAO POLYTECHNIC INST

Text-to-image diffusion model for generalizable mesh generation

PendingCN122349650APattern recognitionAlgorithm
Certain aspects of the present disclosure provide techniques and apparatuses for improved machine learning. In an example method, a multi-view latent tensor is generated based on processing a text input using a diffusion machine learning model, where the multi-view latent tensor corresponds to a plurality of orthographic projections corresponding to the text input. A tri-plane latent tensor is generated based on the multi-view latent tensor using a transformation machine learning model, and a three-dimensional mesh is generated based on processing the tri-plane latent tensor using a decoder machine learning model.
Owner:QUALCOMM INC

A method and system for fine efficient motion-guided four-dimensional content generation

The application discloses a kind of fine efficient motion guide four-dimensional content generation method and system, the method of the present application receives motion prompt information through interactive user interface;Using multi-view image diffusion model generates four-view static image according to fixed pitch angle and preset orthogonal azimuth angle;Motion prompt information is converted into trajectory image, and model input characteristics are constructed;Model input characteristics are sent into four-dimensional Gaussian reconstruction model, and complete three-dimensional Gaussian attribute set is predicted to initial frame, and time sequence consistent three-dimensional Gaussian sequence is generated;Appearance loss function, rigid body loss function as far as possible and vector consistency loss function are calculated on the generated three-dimensional Gaussian sequence, to jointly supervise the optimization of model parameter, realize physical reasonable dynamic geometric modeling.The application can complete the entire four-dimensional generation process through single forward propagation, and end-to-end training does not need additional prompt processing module, while ensuring high throughput in realizing fine-grained control.
Owner:TSINGHUA UNIVERSITY

System and Method for Automatic Creation of Product Photoshoots that Seamlessly Combine Real-Life Image Portions and Artificial Intelligence (AI) Generated Portions

Automatic creation of product photoshoots that seamlessly combine real-life image portions and Artificial Intelligence (AI) generated portions. An input photograph of a product is received, and a background-removed version is generated. The image is fed into a Vision and Language Model (VLM) that generates textual attributes that pertain to the product. The textual attributes are fed into a Large Language Model (LLM), that generates a proposed textual prompt that will command an Image Diffusion unit to generate via Generative AI an image of a scenery that would be appropriate for showcasing that product. The AI-generated image is fed into an AI-based unit that detects and cures visual abnormalities or visual distortions, with AI-based blending and refinement. The method generates a high-definition abnormality-free and distortion-free output image that depicts the real-world product blended seamlessly within the Generative-AI scenery. A similar process performs virtual staging of a room or a house or other venue.
Owner:CLOUDINARY LTD

Multi-concept adaptor learning of multi-modal LLM for image diffusion model

A method, apparatus, non-transitory computer readable medium, and system for image processing include obtaining an input image and a text prompt, wherein the input image depicts a first image element and the text prompt describes a second image element, generating a multimodal embedding based on the input image and the text prompt, wherein the multimodal embedding represents the first image element and the second image element in a multimodal embedding space, generating a guidance embedding based on the multimodal embedding, wherein the guidance embedding represents the first image element and the second image element in a guidance embedding space different from the multimodal embedding space, and generating a synthetic image based on the guidance embedding, wherein the synthetic image depicts the first image element and the second image element.
Owner:ADOBE INC

Text-to-image diffusion model with component locking and rank-one editing

A text-to-image machine learning model takes a user input text and generates an image matching the given description. While text-to-image models currently exist, there is a desire to personalize these models on a per-user basis, including to configure the models to generate images of specific, unique user-provided concepts (via images of specific objects or styles) while allowing the user to use free text “prompts” to modify their appearance or compose them in new roles and novel scenes. Current personalization solutions either generate images with only coarse-grained resemblance to the provided concept(s) or require fine tuning of the entire model which is costly and can adversely affect the model. The present description employs component locking and / or rank-one editing for personalization of text-to-image diffusion models, which can improve the fine-grained details of the concepts in the generated images, reduce the memory footprint update of the underlying model instead of full fine-tuning, and reduce adverse effects to the model.
Owner:NVIDIA CORP