Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

596 results about "Image editing" patented technology

Image editing encompasses the processes of altering images, whether they are digital photographs, traditional photo-chemical photographs, or illustrations. Traditional analog image editing is known as photo retouching, using tools such as an airbrush to modify photographs, or editing illustrations with any traditional art medium. Graphic software programs, which can be broadly grouped into vector graphics editors, raster graphics editors, and 3D modelers, are the primary tools with which a user may manipulate, enhance, and transform images. Many image editing programs are also used to render or create computer art from scratch.

Picture editing method and system based on multi-modal condition adaptation

The invention provides a picture editing method and system based on multi-modal condition adaptation, and the method comprises the steps: obtaining a first text vector: obtaining a picture description corresponding to a picture through processing, processing the picture description through an editor, and obtaining a first text vector; obtaining a first text vector based on the picture information; a second text vector acquisition step: processing the used editing instruction to obtain a second text vector based on the editing instruction; a fusion feature acquisition step: fusing the first text vector and the second text vector through weight to obtain a fusion feature; an image potential feature code acquisition step; in the image editing step, the injected condition information is received, meanwhile, the received potential features and potential noise of the image are denoised, and the image desired by the user is gradually generated in the iterative denoising process under the guidance of the received condition information; and an image restoration step. According to the invention, the stability, controllability and accuracy of image editing can be improved.
Owner:SHENZHEN EMDOOR DIGITAL TECH

Multi-modal image editing

Systems and methods for multi-modal image editing are provided. In one aspect, a system and method for multi-modal image editing includes identifying an image, a prompt identifying an element to be added to the image, and a mask indicating a first region of the image for depicting the element. The system then generates a partially noisy image map that includes noise in the first region and image features from the image in a second region outside the first region. A diffusion model generates a composite image map based on the partially noisy image map and the prompt. In some cases, the composite image map includes the target element in the first region that corresponds to the mask.
Owner:ADOBE INC

Text-guided image editing by learning guidance scales via reinforcement learning

Certain aspects of the present disclosure provide techniques and apparatus for improved machine learning. In an example method, a first latent tensor generated during a first iteration of processing data using a denoising backbone of a diffusion machine learning model is accessed. A guidance scale is generated based on processing the first latent tensor using a guidance machine learning model. A second latent tensor is generated during a second iteration of processing data using the denoising backbone based on the first latent tensor and the first guidance scale, and an output from the diffusion machine learning model is generated based at least in part on the second latent tensor.
Owner:QUALCOMM INC

Modal language model image editing technology fusing non-perpetual culture elements

The invention discloses a modal language model image editing technology fusing non-perpetual culture elements. The modal language model image editing technology comprises the steps that a model is selected and trained, an LLaMA model is selected as a basis, LoRA is introduced for adaptive fine adjustment, and in this way, the model is subjected to adaptive adjustment under the condition that original parameter freezing is kept; according to the method, the understanding and reasoning capabilities in instruction editing are enhanced by combining MLLM (such as LLaVA), the MLLM can perform collaborative learning across text and image modalities and extract deep semantic information, so that the model can not only process basic instructions, but also understand complex non-perpetual culture elements, and the method has great significance in improving the understanding of the model on the non-perpetual culture elements. An enhanced bidirectional interaction mechanism is designed, and the mechanism realizes deep interaction between image and text features through a cross attention mechanism, so that the image features can be used as query and key value pairs to perform bidirectional communication with the text features, and the expression of a model in a complex non-perpetual scene is improved.
Owner:BEIJING JIAOTONG UNIV

Image editing method and system based on diffusion model inversion and attention optimization

The invention discloses an image editing method and system based on diffusion model inversion and attention optimization, and relates to the technical field of image editing, and the method comprises the steps: mapping a source image to a potential space through a pre-trained automatic encoder, and obtaining an initial noise feature; an EF noise space inversion algorithm is adopted to process the initial noise features, a high-variance noise graph is generated, and an intermediate result under each time step is obtained through a decoder; constructing an improved U-Net network based on edit-friendly feature reweighting and an edit-friendly attention mechanism, fusing the target prompt information into a denoising process of the improved U-Net network through a cross attention mechanism, and performing feature optimization based on the improved U-Net network; and reconstructing the potential features through a U-Net decoder, and generating an edited image conforming to the target prompt information. On the low-cost premise that complex model fine tuning and retraining are not needed, the method is superior in text-guided controllable editing tasks, and the image generation quality is improved.
Owner:ZHEJIANG NORMAL UNIV +2

Picture editing method, related equipment and system

The embodiment of the invention provides a picture editing method, related equipment and a picture editing system. Editing types (such as deletion, dragging, replacement, addition, color adjustment and the like) and editing parameters (such as replaceable contents, dragging target positions, newly added contents, color adjustment numerical values and the like) are recommended to a user according to various features such as mask features, depth features, contour features, color features and the like of an editing area selected by the user, and an editing thought is provided for the user; moreover, whether the editing instruction is reasonably matched with the editing area selected by the user can be judged, so that the user is prompted that the editing instruction cannot be executed, the user can be recommended to select a new editing area, or the user is helped to modify the editing instruction, and the user experience is improved. The problem that the editing effect does not conform to the common sense logic due to the fact that the user randomly inputs the editing instruction is solved, so that the trial and error times when the user edits the picture can be reduced, and the picture output efficiency is remarkably improved.
Owner:HUAWEI TECH CO LTD

Image editing method and apparatus, device, and storage medium

Embodiments of the present disclosure provide an image editing method and apparatus, a device, and a storage medium. The image editing method comprises: acquiring an area to be edited in an original image and a target prompt word corresponding to said area, wherein the target prompt word is text information used for describing an expected effect of image editing; determining a target mask image on the basis of said area, and adding preset noise to said area on the basis of the target mask image by means of an image editing model to obtain a local noise image; and on the basis of the target prompt word and by means of the image editing model, performing noise prediction processing and image generation processing on an area to be edited in the local noise image, and outputting a target image on the basis of a noise prediction result and an image generation result. The embodiments of the present disclosure solve the problem that image editing methods in the related art cannot accurately modify and edit local areas, thereby improving image-text consistency and image generation quality.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Discovering and mitigating biases in large pre-trained multimodal based image editing

A method, apparatus, non-transitory computer readable medium, apparatus, and system for image processing include obtaining a text prompt and an input image depicting a person, generating a latent code based on the text prompt and the input image, wherein the latent code is optimized by an identity preserving loss, and generating, using an image generator of a machine learning model, a synthetic image based on the latent code, wherein the synthetic image includes an element of the text prompt and preserves an identity of the person in the input image.
Owner:ADOBE INC

Image editing through utilization of large language model

Some implementations are directed to editing a source image based on a user request to edit the source image. The source image and the user request to edit the source image can be processed, using an image-editing system, to generate one or more image editing instructions. The one or more image editing instructions can indicate an image mask that edit (or preserves) one or more portions of the source image and / or can indicate a target object to be present in the edited image to replace a source object in the source image. Based on the one or more image editing instructions and source image, an edited image that shares the one or more portions with the source image and that differs from the source image by replacing the source object in the source image with the target object can be generated.
Owner:GOOGLE LLC

Interactive point-based image editing

Embodiments of the disclosure relate to interactive point-based image editing. According to example embodiments of the present disclosure, a user edit input for a source image is obtained to indicate at least one handle point and at least one target point in the source image. A feature map is extracted from the source image using a diffusion model at an iteration step of an inverse denoising diffusion process performed on the source image. The feature map is then updated based on the user edit input. Then a target image is generated based on the updated feature map using the diffusion model through a denoising diffusion process performed on the updated feature map.
Owner:LEMON INC(GB)

Image inversion and editing using rectified flow neural networks

Systems and methods for performing image modification. In particular, the system can, using a rectified flow neural network, perform an image inversion and image editing process to generate a modified image that has been modified according to a conditioning input received by the system.
Owner:GOOGLE LLC

Proxy-guided image editing

A method, apparatus, non-transitory computer readable medium, and system for image processing include obtaining an input image and an input mask, wherein the input mask indicates a region of the input image to be modified and generating, using a first image generation model, an intermediate result based on the input image and the input mask, wherein the intermediate result modifies the region of the input image indicated by the input mask. A second image generation model generates a synthetic image based on the input image and the intermediate result, wherein the synthetic image depicts the input image with content from the modified region at a higher level of detail than the intermediate result.
Owner:ADOBE INC

Condition-based image editing

A computer system and a computer-implement method include obtaining a source image and a modification input that indicates a target edit to the source image and generating a modification encoding representing the target edit. An image generation model generates an output image that depicts the source image with the target edit based on the source image and the modification encoding. The image generation model is trained to perform a pose modification task and a part replacement task.
Owner:ADOBE INC

Image editing method and device based on neural radiation field, equipment and storage medium

The invention provides an image editing method and device based on a neural radiation field, equipment and a storage medium, and the method comprises the steps: obtaining a plurality of to-be-processed images, and carrying out the sampling processing of the to-be-processed images, so as to obtain the pixel information of a plurality of sampling points; pixel information of the sampling points is input into a preset neural radiation field optimization model, feature information corresponding to the sampling points is output through the preset neural radiation field optimization model, and radiation rotation invariant constraints are added into the preset neural radiation field optimization model; generating a target neural renderer according to the feature information corresponding to the sampling point; generating a three-dimensional neural point cloud segmentation mask according to the target neural renderer; in response to an editing request of a user, generating an edited image according to the three-dimensional neural point cloud segmentation mask; the technical effect of improving the image editing efficiency and editing controllability based on the neural radiation field is achieved.
Owner:BEIHANG UNIV

Face image reconstruction method based on semantic identity feature decoupling and consistency retention of diffusion model

The invention discloses a face image reconstruction method based on semantic identity feature decoupling and consistency reservation of a diffusion model, and the method comprises the steps: 1, obtaining and preprocessing a face image set of identity labeling, and generating a face feature point distribution diagram, a semantic mask diagram and a description text; 2, multi-modal features are extracted and fused through a semantic identity extraction network; 3, carrying out noise adding and de-noising processing by utilizing a diffusion model, and combining a reconstructed network and semantic identity loss optimization; and 4, face image reconstruction is completed. According to the method, in the face image reconstruction process, the driving requirements of semantic information such as texts for image editing can be accurately captured, fine-grained semantic features and identity features are decoupled, the core identity features of the face can be effectively reserved, and loss of identity consistency caused by semantic editing is avoided; therefore, technical support is provided for application scenes with high requirements on face identity accuracy in the field of computer vision, and the reliability and practicability of face image reconstruction are improved.
Owner:ANHUI UNIV

Adaptive diffusion image editing method and system based on concept attention

The invention discloses a self-adaptive diffusion image editing method and system based on concept attention, and the method comprises the following steps: constructing a paired data set; analyzing the editing instruction, and extracting a key concept; a pre-trained T5 language model is utilized to convert the key concept into text embedding, and the text embedding is mapped to an image feature space; modifying a diffusion model based on a Transform architecture, embedding a concept attention module in an attention layer of a multi-modal diffusion converter, calculating an attention score between image features and concept embedding, and generating a concept saliency map; in the denoising process, the weight of the target area is adjusted by using the concept saliency map so as to realize accurate editing. According to the method, under the condition that the global image quality is not affected, the editing precision can be improved, interference to a non-target area is reduced, and meanwhile, reinforcement learning and real-time feedback are combined, so that the model can be adaptively optimized, and an editing result better meeting the user requirement is generated.
Owner:NANJING UNIV OF POSTS & TELECOMM

Methods and systems for preserving image features during image editing

Described embodiments generally relate to a computer-implemented method for editing an image. The method includes accessing an image; identifying at least a first area of the image and a second area of the image; configuring a model to generate an edited image based on the first area of the image and the second area of the image, wherein the edited image comprises a first area of the edited image and a second area of the edited image; wherein the model is configured to generate the edited image such that the first area of the edited image differs from the first area of the image less than the second area of the edited image differs from the second area of the image.
Owner:CANVA PTY LTD

Operation interaction method and system applied to camera image editing

The invention discloses an operation interaction method and system applied to camera image editing, and relates to the technical field of image processing, and the method comprises the steps: obtaining a voice semantic heat map, a pointing intensity map, a touch confidence map and a gazing confidence map based on a multi-modal interaction data packet, and calculating an image feature matrix at the same time; fusing into a multi-modal evidence graph through a normalized scale; performing semantic segmentation according to the image feature matrix to obtain a semantic segmentation first draft and a pixel-by-pixel category confidence coefficient, and performing position correlation weighting on the pixel-by-pixel category confidence coefficient by taking the multi-modal evidence graph as a confidence coefficient modulation factor to generate a candidate object mask sequence; and performing highlight display on the candidate object mask sequence, and performing conflict resolution and priority rearrangement in combination with the multi-mode evidence graph to generate a target object mask. According to the method, deep fusion of the interaction intention and image segmentation is realized, the precision and consistency of candidate region detection are improved, and the stability of real-time rendering and the reliability of an editing result are improved.
Owner:SHENZHEN XUJING DIGITAL TECH CO LTD

Image processing apparatus and image processing method

In an image processing apparatus 100, an image acquiring portion 11 acquires image information. A position acquiring portion 12 acquires an image capturing position and an image capturing direction of the image information. An image-capturing propriety information acquiring portion 13 acquires image-capturing propriety information indicating that a position of an information terminal 200 corresponds to image-capturing propriety setting information. An image editing portion 14 performs a masking process to a range based on the information terminal 200 in the image information acquired by the image acquiring portion 11, in accordance with image-capturing propriety setting information of the image-capturing propriety information indicating that the position of the information terminal 200 in the image-capturing propriety information corresponds to the image capturing range based on the image capturing position and the image capturing direction acquired by the position acquiring portion 12.
Owner:MAXELL LTD

Editing digital images using executable code generated by large language models from natural language input

The present disclosure relates to systems, methods, and non-transitory computer-readable media that perform text-to-image editing using executable code generated from natural language text input. For instance, in one or more embodiments, the disclosed systems receive, from a client device, a digital image and natural language text input providing instructions for modifying the digital image. The disclosed systems also generate, using a large language model, executable action code for modifying the digital image in accordance with the instructions of the natural language text input, the executable action code being compatible with an editing application. The disclosed systems further modify the digital image by executing the executable action code via the editing application and provide the modified digital image for display via a graphical user interface of the client device.
Owner:ADOBE INC

Implementing drag-based image editing

The present disclosure describes techniques for implementing drag-based image editing. Feature maps are generated based on latent representations of an image by a first sub-model of a machine learning model. The first sub-model is configured to preserve an identity of the image. Embeddings corresponding to at least one pair of points are generated by a second sub-model of the machine learning model. Each pair of points comprises a handle point and a target point. The handle point identifies an area of the image. The target point indicates a target location to which the area is to be relocated. The feature maps and the embeddings are injected into a third sub-model of the machine learning model to guide a process of generating a target image by the third sub-model. The target image depicts the area of the image relocated at the target location.
Owner:LEMON INC(GB)

Utilizing cross-attention guidance to preserve content in diffusion-based image modifications

The present disclosure relates to systems, non-transitory computer-readable media, and methods for utilizing machine learning models to generate modified digital images. In particular, in some embodiments, the disclosed systems generate image editing directions between textual identifiers of two visual features utilizing a language prediction machine learning model and a text encoder. In some embodiments, the disclosed systems generated an inversion of a digital image utilizing a regularized inversion model to guide forward diffusion of the digital image. In some embodiments, the disclosed systems utilize cross-attention guidance to preserve structural details of a source digital image when generating a modified digital image with a diffusion neural network.
Owner:ADOBE INC

Proxy-guided image editing

The embodiment of the invention relates to agent-guided image editing. A method, apparatus, non-transitory computer readable medium and system for image processing includes obtaining an input image and an input mask, wherein the input mask indicates a region in the input image to be modified; and generating an intermediate result based on the input image and the input mask using the first image generation model, where the intermediate result modifies a region indicated by the input mask in the input image. The second image generation model generates a composite image based on the input image and the intermediate result, where the composite image depicts the input image at a higher level of detail with content from the modified region than the intermediate result.
Owner:ADOBE INC

Three-dimensional scene video editing method based on point cloud guidance

The invention provides a three-dimensional scene video editing method based on point cloud guidance, and the method comprises the steps: obtaining an original video of a three-dimensional scene, and estimating the three-dimensional point cloud of the scene in a specified frame of the video and camera parameters of each video frame; determining an editing reference image of the specified frame according to the image of the specified frame, the pixel-level mask and the editing area description text; estimating the edited depth of the specified frame according to the edited reference image to obtain an edited three-dimensional point cloud corresponding to the specified frame; according to the mask of the specified frame and the pre-edit depth map and the post-edit depth map corresponding to the image of the specified frame, constructing a three-dimensional grid model used for surrounding an edit area, and transmitting the mask of the specified frame to the view angle of other frames by using the three-dimensional grid model to obtain masks of other frames; and obtaining a point cloud rendering image of each frame rendered according to the edited three-dimensional point cloud and the camera parameters of each frame, generating an image editing result of each frame according to the point cloud rendering image of each frame, the image, the mask and the editing reference image, and splicing the image editing result into an edited video.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Method and electronic device for automatic generation of high quality data for image editing applications

A method for generating an image editing dataset is provided. The method may include obtaining a candidate image for inputting an AI model from among at least one candidate image. The method may include determining an editing operation to be performed by the AI model on the candidate image from among at least one editing operation. The method may include determining a base prompt and a control prompt based on inputting of the candidate image and the editing operation into the AI model, wherein the base prompt comprises base instructions for detection of at least one object within the candidate image, and the control prompt comprises control instructions relevant to the editing operation to be performed by the AI model on the candidate image. The method may include generating the image editing dataset based on the base prompt and the control prompt.
Owner:SAMSUNG ELECTRONICS CO LTD

Picture editing method, electronic equipment and storage medium

The invention provides a picture editing method, electronic equipment and a storage medium. The first electronic equipment receives editing operation of a user for a first picture and displays a first user interface, an editing interface of the first picture comprises the first picture and a first option, the first option comprises a first node and a second node, the first node is related to a first picture editing algorithm, and the second node is related to a second picture editing algorithm; the first picture is obtained based on a first image editing algorithm and a second image editing algorithm on the basis of the third picture; the first electronic device receives a first operation of the user for the first node and displays a second user interface, the second user interface comprises a second picture, and the second picture is obtained on the basis of the third picture on the basis of the first image editing algorithm. Through the method, the electronic equipment can re-edit the picture stored in the picture library application, so that the stored picture is returned to any editing node, and the picture editing efficiency is improved.
Owner:HUAWEI TECH CO LTD

Multi-condition guide text image generation method based on decoupling and multi-domain guide strategy

The invention provides a multi-condition guide text image generation method based on decoupling and a multi-domain guide strategy. According to the method, an image meeting text description and spatial alignment at the same time can be generated according to the text and any spatial condition. Specifically, structure representation and appearance representation in the image generation process are decoupled, and two independent guide branches, namely an appearance guide branch and a structure guide branch, are designed. The two branches guide the generation process to be highly aligned with the input structure of the guide branch while guiding the appearance content in the accurate expression text through a classifier guide strategy. Besides, in order to realize better structural consistency, the method provides a multi-domain guide strategy, and more comprehensive structural supervision is realized by combining a spatial domain and a frequency domain. According to the method, the text generation image guided by any space condition can be realized, the method can be used in various generative models in a plug-and-play manner, and common downstream tasks such as image deblurring, image coloring, image restoration and image editing can be completed.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Image editing with generative artificial intelligence

A computer-implemented method includes receiving a request for a type of output image and a prompt from a user that describes an output image. The method further includes selecting, based on the type of output image and the prompt, a machine-learning model from a set of machine-learning models. The method further includes providing the request and the prompt as input to the selected machine-learning model. The method further includes generating, by the selected machine-learning model, the output image that satisfies the request and the prompt.
Owner:GOOGLE LLC

Diffusion model-based training-free reference image guided image editing method

A diffusion model-based training-free reference image guided image editing method comprises the following steps: an image preprocessing stage: carrying out resolution standardization on an original image and a reference image, generating a text cue word and extracting a binary mask of a target object; in the image inversion stage, the image is coded into an initial noise potential variable through a deterministic noise adding algorithm; in the image generation stage, an improved self-attention calculation module is used for guiding generation of an edited image in a staged mode, and the method comprises the steps that the layout initialization stage retains original features, the editing stage dynamically estimates masks and fuses reference image key value features to achieve target injection, and the detail synthesis stage optimizes region consistency. The method can automatically position the editing area and inject the reference object features without training, can retain the features of the original image, and overcomes the defects that a user needs to assign the editing area through smearing and needs to spend a large amount of cost to train a model in a traditional method.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Background reserved image editing method and device, electronic equipment and program product

The invention provides a background-reserved image editing method and device, electronic equipment and a program product. The method comprises the following steps: dividing an input image into a foreground image and a background image through a mask; wherein the foreground image comprises a foreground mark, and the background image comprises a background mark; executing an inversion step of inverting the input image into a noise space and executing a de-noising step; wherein in the inversion step, a K value and a V value of a background mark are stored in each time step and an attention layer, in the denoising step, only a foreground mark is processed, and the K value and the V value of the foreground mark are connected with the cached K value and the cached V value of the background mark; and connecting the foreground image output by the de-noising step with the background image before the execution of the inversion step to generate an edited image with a reserved background. According to the background-reserved image editing method and device, the electronic equipment and the program product provided by the invention, background-consistent image editing is realized through a non-training method, the image quality is improved, and the image editing cost is reduced.
Owner:BEIJING XIAOBING YUEDONG TECHNOLOGY CO LTD