Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

266 results about "Image editing" patented technology

Image editing encompasses the processes of altering images, whether they are digital photographs, traditional photo-chemical photographs, or illustrations. Traditional analog image editing is known as photo retouching, using tools such as an airbrush to modify photographs, or editing illustrations with any traditional art medium. Graphic software programs, which can be broadly grouped into vector graphics editors, raster graphics editors, and 3D modelers, are the primary tools with which a user may manipulate, enhance, and transform images. Many image editing programs are also used to render or create computer art from scratch.

Image inversion and editing using rectified flow neural networks

Systems and methods for performing image modification. In particular, the system can, using a rectified flow neural network, perform an image inversion and image editing process to generate a modified image that has been modified according to a conditioning input received by the system.
Owner:GOOGLE LLC

Face image reconstruction method based on semantic identity feature decoupling and consistency retention of diffusion model

The invention discloses a face image reconstruction method based on semantic identity feature decoupling and consistency reservation of a diffusion model, and the method comprises the steps: 1, obtaining and preprocessing a face image set of identity labeling, and generating a face feature point distribution diagram, a semantic mask diagram and a description text; 2, multi-modal features are extracted and fused through a semantic identity extraction network; 3, carrying out noise adding and de-noising processing by utilizing a diffusion model, and combining a reconstructed network and semantic identity loss optimization; and 4, face image reconstruction is completed. According to the method, in the face image reconstruction process, the driving requirements of semantic information such as texts for image editing can be accurately captured, fine-grained semantic features and identity features are decoupled, the core identity features of the face can be effectively reserved, and loss of identity consistency caused by semantic editing is avoided; therefore, technical support is provided for application scenes with high requirements on face identity accuracy in the field of computer vision, and the reliability and practicability of face image reconstruction are improved.
Owner:ANHUI UNIV

Image editing with generative artificial intelligence

A computer-implemented method includes receiving a request for a type of output image and a prompt from a user that describes an output image. The method further includes selecting, based on the type of output image and the prompt, a machine-learning model from a set of machine-learning models. The method further includes providing the request and the prompt as input to the selected machine-learning model. The method further includes generating, by the selected machine-learning model, the output image that satisfies the request and the prompt.
Owner:GOOGLE LLC

Complex instruction image editing method based on multi-modal large language model

The invention discloses a complex instruction image editing method based on a multi-modal large language model, and the method comprises the steps: firstly constructing a data set expansion strategy based on a data set of multi-round image editing, carrying out the two-stage preprocessing, and constructing a complex instruction image editing data set; secondly, constructing a context prompt template, decoupling an editing instruction in a complex instruction image editing data set by using a multi-modal large language model to obtain a sub-editing instruction and an editing area, injecting a diffusion model with space-time perception background enhancement, including a space-time perception cross attention module and a background enhancement module, and carrying out weighted fusion on the obtained features to obtain an editing result; and an edited image is decoupled, output and edited through the variational auto-encoder. And finally, carrying out fine tuning by adopting a two-stage training strategy, and continuously training until the whole model is converged. According to the method, the instruction and background consistency of complex instruction image editing is remarkably improved, and the problems that complex instructions are neglected, background information is lost, and non-intention editing exists in an existing method are solved.
Owner:HANGZHOU DIANZI UNIV

Utilizing machine learning models to generate image editing directions in a latent space

The present disclosure relates to systems, non-transitory computer-readable media, and methods for utilizing machine learning models to generate modified digital images. In particular, in some embodiments, the disclosed systems generate image editing directions between textual identifiers of two visual features utilizing a language prediction machine learning model and a text encoder. In some embodiments, the disclosed systems generated an inversion of a digital image utilizing a regularized inversion model to guide forward diffusion of the digital image. In some embodiments, the disclosed systems utilize cross-attention guidance to preserve structural details of a source digital image when generating a modified digital image with a diffusion neural network.
Owner:ADOBE INC

Spatial self-adaptive plug-and-play watermarking method and device for image editing traceability

The invention discloses a spatial self-adaptive plug-and-play watermarking method and device for image editing traceability, and the method comprises the steps: obtaining the hidden space feature representation of an edited image through an image editing model according to an input original image and an editing condition; constructing a structured watermark grayscale image used for bearing image editing traceability core information, and encoding the structured watermark grayscale image into watermark potential features through a watermark encoder; based on the watermark potential features, watermark embedding of content and structure perception is realized in the submerged space, and the submerged space feature representation of the edited image with the watermark is acquired; and inputting the hidden space feature representation of the edited image into an image decoder, decoding the edited image into an edited image with a watermark, inputting the edited image with the watermark into a watermark decoder, extracting an embedded watermark grayscale image, and restoring complete traceability metadata information based on the watermark grayscale image. The device comprises a processor and a memory. According to the invention, high visual quality of the image is maintained while reliable traceability is realized.
Owner:TIANJIN UNIV

Method and apparatus for automating personalized artificial intelligence image editing

Provided is a method and apparatus for automating personalized artificial intelligence (AI) image editing, which can generate an output into which a personal style of a user is incorporated when an image is edited by incorporating editing requirements of the user. The method includes an input step of receiving an input image and an user edit instruction from a user terminal, a personalized text encoding step of converting the user edit instruction into a personalized image editing command into which personal characteristics and preference have been incorporated, a personalized denoising step of generating an output image by editing the input image by incorporating the personalized image editing command, and an output step of outputting an edited output image.
Owner:ELECTRONICS & TELECOMM RES INST

Advertisement material generation method and device based on multi-modal large model, equipment and storage medium

The embodiment of the invention relates to the technical field of artificial intelligence, and discloses an advertisement material generation method and device based on a multi-modal large model, equipment and a storage medium, and the method comprises the steps: receiving an advertisement material generation command, inputting title data into a large language model according to the command, carrying out the core semantic reservation and compression processing, and outputting a prompt title; if the commodity main image has no background, inputting the commodity main image and the commodity category into a large-scale visual language model to generate a scenarized background prompt word; inputting the commodity main image and the background cue word into an image editing model for semantic fusion and synthesis to obtain a background synthesis main image; and loading the image-text advertisement template and the template information, and generating batch advertisement materials according to the image-text advertisement template, the template information, the background synthesis main image, the prompt title and the price information. Through the mode, the method is adaptive to diversified commodity characteristics, and collaborative understanding and fusion are carried out on image and text multi-modal information, so that the advertisement quality and the batch generation efficiency are improved.
Owner:SHENZHEN YOUYOU INTERNET TECH CO LTD

Image editing with a selected machine-learning model

A computer-implemented method includes receiving an initial image and an original prompt from a user, wherein the original prompt includes a request to modify the initial image. The method further includes selecting, based on the original prompt, a machine-learning model from a set of machine-learning models. The method further includes providing the original prompt and the initial image as input to a large language model (LLM). The method further includes receiving, from the LLM and based on the original prompt and the initial image, a rewritten prompt. The method further includes selecting, based on the rewritten prompt, a machine-learning model from a set of machine-learning models. The method further includes generating, by the selected machine-learning model, an output image that satisfies the rewritten prompt.
Owner:GOOGLE LLC

3d-consistent image inpainting with diffusion models

The present disclosure relates to image editing or inpainting techniques leveraging a generator model conditionally trained on one or more in-context images during a reverse diffusion process. The generator model performs inpainting of an image at inference by accessing a set of images varying in context that depicts a same or similar scene. A masked version of the image may be generated by obscuring portions of the image using a masking technique. After masking, a noisy image may be generated by iteratively introducing noise to the masked version of the image based on a noise schedule. The noisy image may act as a starting point for the subsequent reverse process leveraging the generator model configured to receive an iterated version of the image and the one or more in-context images. Based on the generator model, a transformed version of the image may be generated by iteratively denoising the noisy image.
Owner:NAVER CORP

Image editing using prompt-aware content segmentation masks and mask-aware content-generation

Methods and systems are provided for image editing using prompt-aware content segmentation masks and mask-aware content generation. In embodiments described herein, an image, prompt, and selection to replace a selected type of content in the image with generated content is received. An image-generating model generates a generated image based on the prompt and image. A content mask extraction model extracts a first content mask from the image and a second content mask from the generated image based on the selected type of content. A refined content mask is generated by geometrically transforming the second content mask with respect to the first content mask and combing the two content masks. The image, prompt, and refined content mask are applied to a mask-aware content generating model to generate content within the refined content mask. The input image with the generated content within the refined content mask is displayed.
Owner:ADOBE INC

Data set construction method oriented to multi-task image editing and related device

The invention provides a multi-task image editing-oriented data set construction method and a related device, and the method comprises the steps: inputting an original image into a multi-modal large language model for semantic feature analysis, and determining an adaptive image editing task type from a predefined task type set based on an analysis result; based on the image editing task type, calling a multi-modal large language model to perform multi-round editing deduction on the original image, generating an editing task instruction, selecting a corresponding diffusion model according to the task type in the editing task instruction, inputting the original image, the natural language instruction and the auxiliary information into the corresponding diffusion model, and generating an edited target image; a plurality of image editing pairs are evaluated by adopting a multi-dimensional evaluation index, a data set meeting a preset requirement is screened out, the image editing pairs comprise an original image, an edited target image and an editing task instruction involved in deduction, and the problem of insufficient task diversity in the prior art is solved; and the data set generation quality is effectively improved.
Owner:BEIJING ZERO-1000 TECHNOLOGY CO LTD

Portrait restoration method and system for solving illusion generation of large model

The invention discloses a portrait restoration method and system for solving large model illusion generation, and the method comprises the steps: taking an original face image, an edited image and a corresponding text description as the input, and achieving the unification and smooth fusion of double-ID latent variable tracks through latent space alignment; a staged hybrid sampler is used to consider both identity geometry maintenance and local attribute detail enhancement; and a gating mechanism is introduced into a self-attention and cross-attention layer to selectively inject target attributes, so that semantic drift of a non-target area is avoided. The method directly acts on the pre-training diffusion model, extra training is not needed, deployment is convenient, and reasoning is efficient. Experimental results show that the method is remarkably superior to an existing method in the aspects of indexes such as identity similarity, attribute consistency and human preference, and stable performance is still kept in complex scenes such as multiple persons, shielding and non-front-face scenes.
Owner:WUHAN UNIV

Retrieving digital images in response to search queries for search-driven image editing

Systems, methods, and non-transitory computer-readable media implements related image search and image modification processes using various search engines and a consolidated graphical user interface. For instance, one or more embodiments involve receiving an input digital image and search input and further modify the input digital image using the image search results retrieved in response to the search input. In some cases, the search input includes a multi-modal search input having multiple queries (e.g., an image query and a text query), and one or more embodiments involve retrieving the image search results utilizing a weighted combination of the queries. Some implementations involve generating an input embedding for the search input (e.g., the multi-modal search input) and retrieving the image search results using the input embedding.
Owner:ADOBE INC

Image classification robustness test enhancement method and system based on artificial intelligence model

The invention discloses an artificial intelligence model-based image classification robustness test enhancement method and system, and the method comprises the steps: carrying out the deep semantic analysis of the image content based on a multi-modal artificial intelligence model, automatically recognizing and extracting the class related features and non-class related features in an image, and generating a structured feature analysis report; the method comprises the following steps: generating a Keep Strategy and a Replace Strategy which are complementary to each other, and generating a Keep Strategy and a Replace Strategy which are complementary to each other; performing image editing operation based on the generated strategy to generate a corresponding image sample; performing automatic quality verification on the edited image sample, filtering out low-quality samples which do not conform to expectation, and generating a test sample; and comprehensively evaluating the target classification model based on the generated test samples, calculating performance indexes of the model on different types of test samples, generating a detailed robustness analysis report, and identifying weak links and improvement directions of the model. According to the scheme, a complete'generation-test-evaluation 'closed-loop system can be established, a standardized and quantifiable robustness evaluation index system is formed, and the model robustness can be scientifically and quantitatively evaluated.
Owner:THE THIRD RES INST OF MIN OF PUBLIC SECURITY

Text-guided image editing by learning guidance scales via reinforcement learning

Certain aspects of the present disclosure provide techniques and apparatus for improved machine learning. In an example method, a first latent tensor generated during a first iteration of processing data using a denoising backbone of a diffusion machine learning model is accessed. A guidance scale is generated based on processing the first latent tensor using a guidance machine learning model. A second latent tensor is generated during a second iteration of processing data using the denoising backbone based on the first latent tensor and the first guidance scale, and an output from the diffusion machine learning model is generated based at least in part on the second latent tensor.
Owner:QUALCOMM INC

Image processing method and device

The embodiment of the invention provides an image processing method and device.The image processing method comprises the steps that in response to an editing instruction for an original image, the original image is updated into an edited image through an image editing model; inputting the original image and the edited image into an image evaluation model for processing to obtain image evaluation information corresponding to a plurality of image evaluation dimensions; dividing the editing image into an editing area and a non-editing area, and determining editing evaluation information corresponding to the editing area and non-editing evaluation information corresponding to the non-editing area; and according to the region information of the editing region, fusing the image evaluation information, the editing evaluation information and the non-editing evaluation information to obtain target evaluation information corresponding to the editing image.
Owner:HANGZHOU ALIBABA INT INTERNET IND CO LTD

Diffusion model-oriented remote sensing image reverse editing method

The invention relates to a remote sensing image reverse editing method for a diffusion model, and belongs to the field of computer vision. The strong generation and editing capability of the diffusion model brings a series of potential safety hazards, and the image reverse editing technology aims at adding a layer of protective noise to an image, so that the diffusion model is difficult to output an ideal editing result. In order to solve the problems that an existing anti-editing technology is poor in effect on a remote sensing image and generally faces low efficiency and obvious protection traces, the invention provides an efficient and hidden protection scheme which comprises the following three steps: step 1, generating anti-editing disturbance based on hidden space high-frequency component suppression; step 2, noise generation network training based on single-time forward transmission; and step 3, concealment enhancement based on noise perception graph constraint. The method can effectively resist malicious image editing based on a diffusion model, is suitable for the field of remote sensing images, and gives consideration to safety, efficiency and image quality.
Owner:BEIHANG UNIV

Synthetic data generation of image training data

Systems and methods to generate synthetic images for use in a machine learning training set. The process begins with accessing a database of real-world 3-D images of equipment in a power grid, the 3-D images of equipment include 3-D measurements to create a dimensionally accurate and photorealistic model of the equipment. Optionally, the 3-D images could be aged or weathered using imaging editing software. Next, a database of real-world photographs of scenes in which the equipment is installed is accessed. Optionally, the identical scenes can be captured at different times of day, different times of the year, and at different perspectives. Next, using image editing software, the 3-D images of equipment is inserted into at least one of the scenes to form a synthetic image based on a combination of the equipment and the scene in which each of the equipment and the scene were previously captured independently of each other.
Owner:FLORIDA POWER & LIGHT CO

Graphical user interface for image editing of an electronic device

1. The name of the design product: the graphical user interface of an electronic device for image editing. 2. The use of the design product: an electronic device. 3. The design points of the design product: the graphical user interface. 4. The picture or photo that best indicates the design points: design 1 interface change state figure 1. 5. The design points not involved in the appearance are omitted, the left view is omitted, the right view is omitted, the top view is omitted, and the bottom view is omitted. 6. Design 1 is designated as the basic design. 7. The use of the graphical user interface: the graphical user interface is used for interactive operation and information display for the user to edit the photos or videos taken by the thermal imaging device through the graphical user interface. 8. The human-computer interaction mode of the graphical user interface: the product can edit the photos or videos taken by the thermal imaging device through human-computer interaction, and different keys can enter different function interfaces; touching the 3D editing button of design 1 main view can enter design 1 interface change state figure 1, in which the coordinate system is the image to be edited, and the buttons below the coordinate system display the views of the image to be edited in different directions, and the user can perform related operations according to the interface prompts; touching the calibration button of design 1 interface change state figure 1 can enter design 1 interface change state figure 2, and the user can perform related operations according to the interface prompts; touching the isothermal button of design 1 interface change state figure 2 can enter design 1 interface change state figure 3, and the user can perform related operations according to the interface prompts; touching the pseudo-color button of design 2 main view can enter design 2 interface change state figure, and the user can perform related operations according to the interface prompts.
Owner:深圳鼎匠科技有限公司

Fine-grained image editing method based on multi-modal thinking chain reasoning

The invention discloses a fine-grained image editing method based on multi-modal thinking chain reasoning, and belongs to the technical field of computer vision and image processing. The invention aims to solve the problem that the existing image editing method cannot meet the requirements of controllability and refined editing at the same time in a complex editing scene. The method comprises the following steps: firstly, generating text chain thinking reasoning according to an editing instruction and an input image by utilizing a multi-modal generation-understanding unified model so as to determine a target object referred by a user; and on this basis, a pixel-level visual positioning image corresponding to the target is generated. Secondly, the model generates semantic reasoning of an editing description and an editing result according to a multi-modal positioning clue, and executes local region editing to generate an accurate edited image; in the training process, positioning enhancement is realized through a multi-modal thinking chain alignment mechanism and auxiliary mask supervision, so that semantic consistency and positioning accuracy between an inference chain and an actual editing area are ensured. According to the method, the editing capability with high interpretability, accurate space alignment and interactivity is realized.
Owner:HARBIN INST OF TECH

Image editing method and device, electronic equipment and storage medium

The application provides an image editing method and device, electronic equipment and storage medium, and relates to the technical field of artificial intelligence. First, the image to be edited and user instructions are acquired, then the geometric structure features in the image to be edited are extracted, and finally, the structure editing area in the image to be edited is determined by using the user instructions and the geometric structure features, and the structure editing area is edited to obtain the target image after structure editing. The method extracts the geometric structure features in the image to be edited in a display mode, can realize the decoupling of content and structure, avoid the destruction of the original texture when editing the structure, and ensure the physical rationality of the target image after physical editing. Moreover, the method combines the user instructions and the geometric structure features to more accurately guide the structure editing process, thereby reducing the artifact problems such as breakage and distortion caused by structure misalignment from the source, and improving the quality of the target image and the satisfaction of the user.
Owner:CHINA MOBILE COMM LTD RES INST +1

Image editing and fusing method based on structured state space sequence model

The invention provides an image editing and fusing method based on a structured state space sequence model, and the method comprises the steps: S1, obtaining at least one input image, and carrying out the time sequence feature extraction of the input image through employing the structured state space sequence model, and obtaining a background feature; s2, background features are separated through mask operation, foreground features output by a conditional generation model are introduced, and latent space feature coding is performed on the background features and the foreground features through a variational auto-encoder to obtain coded background features and coded foreground features; and S3, inputting the encoded background features and the encoded foreground features into a Mamba fusion network for feature fusion to obtain a fusion image, and dynamically adjusting fusion weights of the encoded background features and the encoded foreground features by using a selective scanning mechanism in a fusion process to perform feature space alignment. The method has the beneficial effects that the dynamic feature modeling capability can be realized, the feature fusion accuracy is improved, and the calculation efficiency is improved.
Owner:NINGBO MUNICIPAL ENG CONSTR GROUP

Image editing method and device, computer readable medium and electronic equipment

The embodiment of the invention provides an image editing method and device, a computer readable medium and electronic equipment. The method comprises the steps that an original image, a single-color image, first description information of a first object in the original image and second description information of a second object in a target image are acquired; encoding the original image into a first feature, encoding the single-color image into a second feature, and constructing a text feature according to the first description information and the second description information; splicing the first feature and the second feature, and adding the splicing result and the noise graph to obtain a first hidden space feature; performing forward diffusion operation based on the first hidden space feature to obtain a second hidden space feature; performing denoising operation of multiple time steps on the second hidden space feature based on the text feature to obtain a third hidden space feature; and generating a target image matched with the second description information according to the third hidden space feature. According to the scheme provided by the embodiment of the invention, the image editing efficiency and the image editability can be greatly improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A scribble face image editing method based on a multi-modal interaction module

The present application relates to the technical field of face image editing, in particular to a kind of face image editing method based on multi-modal interactive module of scribble. Through two-way multi-modal interactive module, cross attention graph is calculated from position dimension and channel dimension respectively, the cross attention graph obtained and scribble vector are used to carry out latent space latent reflection on face image latent vector iterative modification, complete scribble and image target editing position semantic content alignment, and scribble is embedded into corresponding latent space;Through single-path multi-modal interactive module, save the original face identity feature after editing, the texture supplement vector obtained carries out latent space latent reflection on face image latent vector iterative modification, and finally generates face image editing result.The present application can more intuitively and fully express the editing intention of user, realize the real feeling editing effect that meets user's expectation, and the present application has superiority in editing effect, editing real feeling and face identity information saving of face image editing.
Owner:HUNAN UNIV

Methods and systems for image editing

Embodiments relate to methods, systems and computer-readable media for editing images. An embodiment of a method for editing an image includes accessing a source image for editing, accessing a reference image, generating a mapping, wherein the mapping relates at least one source attribute value to at least one target attribute value, and modifying the source image to generate an edited image by applying the generated mapping, the edited image having at least one attribute value based on the attribute values of the reference image. Embodiments also relate to methods, systems and computer-readable media for training machine learning models. An embodiment of a method for training a machine learning model to edit an image, includes accessing a source image for editing, accessing a reference image, generating a source histogram for the source image and a reference histogram for the reference image, inputting the source histogram and the reference histogram into an encoder to generate a source histogram vector and a target histogram vector, concatenating the source histogram vector and the reference histogram vector into a concatenated vector, inputting the concatenated vector into the machine learning model, generate a mapping using the machine learning model, applying the mapping to the source image, calculating at least one loss function, and updating the parameters of the machine learning model on the basis of the at least one loss function.
Owner:CANVA PTY LTD

Image editing device, image editing method, program, and recording medium

Provided are an image editing device, an image editing method, a program, and a recording medium that make it possible to perform editing within an appropriate range when editing any image in a composite image generated using a plurality of images. An image editing device according to one embodiment of the present invention comprises a processor. The processor acquires a composite image including a plurality of images, and first editing information for a first image among the plurality of images in the composite image, calculates a degree of unsuitability of the first editing information on the basis of the relationship between the first image and a second image other than the first image among the plurality of images in the composite image in which the first image has been edited on the basis of the first editing information, and generates second editing information in which a degree of editing is more restricted than that in the first editing information when the degree of unsuitability satisfies a restriction condition.
Owner:FUJIFILM CORP

Electronic device and method for image editing in the same

An electronic device is provided. The electronic device includes memory storing instructions and including one or more storage media and at least one processor including processing circuitry, the at least one processor communicatively coupled to the memory, wherein the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to activate a search function comprising a circle-to-search function, based on a configuration of the electronic device, after the search function is activated, determine an area in which an object is to be recognized, based on a user input for a display, recognize an object displayed within the determined area, display the recognized object in a first area of the display, display a search result of the recognized object in a second area of the display, and store an image of the recognized object and information on the search result of the recognized object in a predetermined space of the memory, and wherein the circle-to-search function includes a function of designating a specific area in the display, based on an input of a user, and performing a search for an image of the predetermined area.
Owner:SAMSUNG ELECTRONICS CO LTD