Image processing method and device and computer equipment
By processing the prompt vector and mask diagram in the latent space by the first denoising module and the second denoising module of the diffusion model, the problem of low local processing efficiency of image is solved and efficient local modification of image is achieved.
Patent Information
- Application Number
- CN202510389845.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the local image processing efficiency is low, the labor cost is high, the local editing algorithm is computationally large, and the image detail processing is not smooth enough.
The first denoising module using the diffusion model denoising in the first potential space, and fuses the prompt vector corresponding to the first prompt word, determines the region of interest and obtains the mask map, uses the first denoising module to process the first prompt word, and combines the second denoising module of the diffusion model to process the second prompt word, and generates a second image.
It improves the efficiency of local image modification, avoids destroying non-interested areas, and meets users' needs for local image modification.
Smart Images

Figure CN120259125A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer vision, involves image processing technology, and particularly relates to an image processing method, apparatus, and computer device. Background Art
[0002] With the development of image processing technology, images that meet user requirements can be generated, and local modifications can also be made to images. Currently, professional image processing tools are usually used to make local modifications to images, but such methods have a high labor cost and low image editing efficiency. In addition, related technologies have also proposed using local editing algorithms to perform local processing on images, but such local editing algorithms have a large amount of calculation and low efficiency of local image modification. Moreover, local editing algorithms rely too much on the similarity of local pixels when editing local areas (such as human hair color), resulting in insufficiently smooth image detail processing and affecting image processing efficiency. Summary of the Invention
[0003] Embodiments of this application provide an image processing method, apparatus, and computer device, which can solve the technical problem of low efficiency of local processing of images in related technologies.
[0004] Embodiments of this application provide an image processing method, which includes: denoising in a first latent space using a first denoising module of a preset diffusion model, and fusing a first prompt vector corresponding to a preset first prompt to generate a first image corresponding to the first prompt; determining a region of interest in the first image, and obtaining a mask image corresponding to the region of interest; processing the first prompt using the first denoising module, processing a second prompt for modifying the region of interest using a second denoising module of the diffusion model, and determining a second image in combination with the mask image.
[0005] In some embodiments of this application, the step of denoising in a first latent space using a first denoising module of a preset diffusion model and fusing a first prompt vector corresponding to a preset first prompt to generate a first image corresponding to the first prompt includes: converting the first prompt into the first prompt vector; starting from a preset random noise, fusing the first prompt vector during the denoising process of the first latent space using the first denoising module to generate a latent variable corresponding to the first prompt vector; and decoding the latent variable to obtain the first image.
[0006] In some embodiments of the present application, the processing of the first prompt using the first denoising module and the processing of the second prompt for modifying the region of interest using the second denoising module of the diffusion model, and determining the second image in combination with the mask image, includes: converting the first prompt into a first prompt vector, and converting the second prompt into a second prompt vector; during the denoising process of the first denoising module in the first latent space, fusing the first prompt vector and obtaining a first feature map in combination with the mask image, where the first feature map is the feature map corresponding to the non-region of interest; during the denoising process of the second denoising module in the second latent space, fusing the second prompt vector and obtaining a second feature map in combination with the mask image, where the second feature map is the feature map corresponding to the region of interest; fusing the first feature map and the second feature map to obtain the second image.
[0007] In some embodiments of the present application, the fusing of the first feature map and the second feature map to obtain the second image includes: using the second denoising module to fuse the first feature map and the second feature map to obtain the second image.
[0008] In some embodiments of the present application, the using of the first denoising module to fuse the first prompt vector during the denoising process in the first latent space and obtaining the first feature map in combination with the mask image includes: starting from random noise, fusing the first prompt vector during the denoising process in the first latent space, obtaining the i-th layer feature output by the i-th layer of the first denoising module, where i belongs to positive integers; based on the i-th layer feature and the mask image, obtaining the first feature map output by the i-th layer of the first denoising module.
[0009] In some embodiments of the present application, the using of the second denoising module to fuse the second prompt vector during the denoising process in the second latent space and obtaining the second feature map in combination with the mask image includes: starting from the random noise, fusing the second prompt vector during the denoising process in the second latent space, obtaining the intermediate feature generated after the i-th layer of the second denoising module processes the output result of the (i - 1)-th layer; based on the intermediate feature and the mask image, obtaining the second feature map determined by the i-th layer of the second denoising module.
[0010] In some embodiments of the present application, the fusing the first feature map and the second feature map to obtain the second image includes: using the second denoising module to fuse the first feature map output by the i-th layer of the first denoising module and the second feature map determined by the i-th layer of the second denoising module to obtain the output result of the i-th layer of the second denoising module; determining the second image based on the output result of the last layer of the second denoising module.
[0011] In some embodiments of the present application, the determining the region of interest in the first image includes: segmenting the first image based on a preset image segmentation model to determine the region of interest; or, in response to a user's annotation operation on the first image, determining the region of interest.
[0012] Embodiments of the present application provide an image processing apparatus, including: an image generation module, configured to perform denoising in a first latent space using a first denoising module of a preset diffusion model and fuse a first prompt vector corresponding to a preset first prompt word to generate a first image corresponding to the first prompt word; an image acquisition module, configured to determine a region of interest in the first image and acquire a mask image corresponding to the region of interest; an image modification module, configured to process the first prompt word using the first denoising module, process a second prompt word for modifying the region of interest using a second denoising module of the diffusion model, and determine a second image in combination with the mask image.
[0013] Embodiments of the present application further provide a computer device, including: a memory, and a processor, where the processor executes computer-readable instructions stored in the memory to implement the image processing method described above.
[0014] Embodiments of the present application further provide a computer-readable storage medium, where the computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the image processing method described above is implemented.
[0015] In the image processing method provided in the embodiments of the present application, the first denoising module of the preset diffusion model is used to denoise in the first latent space, and the first prompt vector corresponding to the preset first prompt word is fused to generate the first image corresponding to the first prompt word, and the initial image expected by the user is generated using the diffusion model. Further, for local modification of the image on the initial image, the region of interest in the first image is determined, and the mask image corresponding to the region of interest is obtained. After determining the mask image, in order to avoid damaging the part of the first image corresponding to the first prompt word outside the region of interest, the first denoising module is used to process the first prompt word. In addition, the second denoising module of the diffusion model is used to process the second prompt word for modifying the region of interest. Combining the processing of the first prompt word, the processing of the second prompt word, and the mask image, the second image is generated. The embodiments of the present application can meet the modification of the local image of the image generated based on the diffusion model without retraining other diffusion models, improving the modification efficiency of the local image of the image generated based on the diffusion model. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is an application scenario diagram of the image processing method provided in an embodiment of the present application.
[0018] Figure 2 It is a flowchart of the image processing method provided in an embodiment of the present application.
[0019] Figure 3 It is a processing schematic diagram of the diffusion model provided in an embodiment of the present application.
[0020] Figure 4 They are the first image and the second image provided in an embodiment of the present application.
[0021] Figure 5 It is a flowchart of the generation of the second image provided in an embodiment of the present application.
[0022] Figure 6 It is a processing schematic diagram of the diffusion model provided in another embodiment of the present application.
[0023] Figure 7 It is a schematic diagram of the image processing device provided in an embodiment of the present application. Detailed Embodiments
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following provides a detailed description of this application with reference to the accompanying drawings and specific embodiments.
[0025] It should be noted that in this application, "at least one" means one or more, and "a plurality" means two or more than two. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The terms "first", "second", "third", "fourth", etc. (if any) in the description, claims, and drawings of this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0026] In the embodiments of this application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0027] With the development of image processing technology, images that meet user requirements can be generated, and local modifications can also be made to images. Currently, professional image processing tools are usually used to make local modifications to images, but such methods have a high labor cost and low image editing efficiency. In addition, related technologies have also proposed using local editing algorithms to perform local processing on images, but such local editing algorithms have a large amount of calculation and low efficiency of local image modification. Moreover, local editing algorithms rely too much on the similarity of local pixels when editing local areas (such as human hair color), resulting in insufficiently smooth image detail processing and affecting image processing efficiency.
[0028] To improve the local processing efficiency of images, the embodiments of this application provide an image processing method, apparatus, and computer device. By deploying a first denoising module and a second denoising module that run in parallel in a diffusion model and fusing the processing results of the first denoising module and the second denoising module, the efficiency of local modification of images is improved. The following first introduces the application scenario.
[0029] Figure 1It is an application scenario diagram of the image processing method provided by an embodiment of the present application. The image processing method provided by the embodiment of the present application is applied to the computer device 10. The user can access the online store 210 running in the e-commerce platform 20 through the computer device 10, so as to obtain the product images in the online store 210. Among them, the e-commerce platform 20 can be implemented as cloud computing services, software as a service (SaaS), infrastructure as a service (IaaS), platform as a service (PaaS), desktop as a service (DaaS), management software as a service (MSaaS), mobile backend as a service (MBaaS), information technology management as a service (ITMaaS), and so on. In some embodiments, the elements of the e-commerce platform 20 can be implemented to operate on various platforms and operating systems, such as iOS, Android, on the Web, etc. (for example, the administrator 220 is implemented in multiple instances of a given online store for iOS, Android, and for the Web, and each instance has similar functions).
[0030] The e-commerce platform 20 can provide a centralized system for providing online resources and facilities for merchants to manage their businesses. The facilities described herein can be partially or fully deployed by a machine that executes computer software, modules, program code, and / or instructions on a part of the e-commerce platform 20 or on one or more processors external to the e-commerce platform 20. Merchants can utilize the e-commerce platform 20 to manage commerce with consumers, such as by interacting with consumers via the communication 230 of the e-commerce platform 20, or any combination thereof.
[0031] The online store 210 can represent a multi-tenant facility including multiple virtual storefronts. In an embodiment, a merchant can manage one or more storefronts in the online store 210, such as through a computer device 10 (for example, a laptop, a mobile computing device, etc.). A merchant can upload images (such as images of products), videos, content, data, etc. to the e-commerce platform 20, such as for the system to store (for example, store as data 240). In addition, the e-commerce platform 20 may be configured with a business management engine 250 for content management, task automation, and data management to support and serve multiple online stores 210 (such as independent website stores). The business management engine 250 includes the basic or "core" functions of the e-commerce platform 20 (e.g., common for most online store activities, such as cross-channel, administrator interface, merchant location, industry, product type, etc.), and can be reused across the online stores 210 (e.g., functions that can be reused / modified across the core functions). In one example, the business management engine 250 can be used to extract the original images of each product in the online store 210 and store the original images in the data 240 in an associated manner.
[0032] A neural network model can be deployed on the computer device 10. The computer device 10 can obtain the original images corresponding to one or more stores in the online store 210 by retrieving the data in the data 240, and perform image processing on the original images through the neural network model. In addition, the neural network model deployed on the computer device 10 includes a diffusion model. The diffusion model is used to generate product images that meet the user's needs, and can also perform local modification on the generated product images to improve the image editing efficiency. The product images processed by the diffusion model are stored in the data 240 for the online store 210 to be published on the browsing interface of the store.
[0033] Please continue to refer to Figure 1 , the computer device 10 may include, but is not limited to, a communication module 110, a memory 120, a processor 130, an input / output (I / O) interface 140, and a bus. The processor 130 is coupled to the communication module 110, the memory 120, and the I / O interface 140 through the bus respectively.
[0034] Those skilled in the art can understand that the schematic diagram is only an example of the computer device 10, and does not constitute a limitation on the computer device 10. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the computer device 10 may further include a network access device, etc.
[0035] The communication module 110 may include a wired communication module and / or a wireless communication module. The wired communication module may provide one or more of the solutions for wired communication such as Universal Serial Bus (USB), Controller Area Network (CAN), etc. The wireless communication module may provide one or more of the solutions for wireless communication such as Wireless Fidelity (Wi-Fi), Bluetooth (BT), mobile communication network, Frequency Modulation (FM), near field communication (NFC), Infrared (IR) technology, etc.
[0036] The memory 120 can be used to store computer-readable instructions and / or modules. The processor 130 realizes various functions of the computer device 10 by running or executing the computer-readable instructions and / or modules stored in the memory 120, and by calling the data stored in the memory 120. The memory 120 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the computer device 10. The memory 120 may include non-volatile and volatile memories, such as: hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other storage devices.
[0037] The memory 120 can be an external memory and / or an internal memory of the computer device 10. Further, the memory 120 can be a memory in a physical form, such as a memory stick, a Trans-flash Card (TF card), etc. The processor 130 can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The processor 130 is the computing core and control center of the computer device 10, connecting all parts of the entire computer device 10 through various interfaces and lines, and executing the operating system of the computer device 10 and various installed application programs, program codes, etc. The I / O interface 140 is used to provide a channel for user input or output. For example, the I / O interface 140 can be used to connect various input and output devices, such as a mouse, keyboard, touch device, display screen, etc., enabling the user to input information or visualize information.
[0038] The bus is at least used to provide a communication channel for mutual communication between the communication module 110, memory 120, processor 130, and I / O interface 140 in the computer device 10.
[0039] It can be understood that the application scenarios illustrated in the embodiments of the present application do not constitute specific limitations. In some embodiments of the present application, the computer device 10 may further include a network access device, etc., and the present application does not limit this.
[0040] Figure 2 is a flowchart of an image processing method provided by an embodiment of the present application. As Figure 2 shown, the image processing method provided by the embodiments of the present application is applied in a computer device (such as Figure 1 the computer device 10). According to different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted. As Figure 2 shown, it includes the following steps: Step S201: Denoise in the first latent space using the first denoising module of the preset diffusion model, and fuse the first prompt vector corresponding to the preset first prompt word to generate the first image corresponding to the first prompt word.
[0041] In some embodiments of the present application, the diffusion model can be the Stable Diffusion (SD) model. The diffusion model is an image generation model that gradually applies noise to an image in the forward stage until the image is destroyed and becomes completely Gaussian noise, and learns the process of restoring from Gaussian noise to the original image in the reverse stage. The diffusion model includes a first denoising module, and the first denoising module can be a Unet module. The Unet module is the core component of the diffusion model and is used for denoising and multi-scale feature fusion in the Latent Space. The structure of the Unet module is an encoder-decoder structure.
[0042] In addition, the diffusion model further includes a text encoder and a Variational Autoencoder (VAE) decoder. Among them, the text encoder usually adopts a pre-trained natural language processing model (such as the text encoder of CLIP, BERT, etc.) to convert natural language prompt words into high-dimensional semantic vectors (Embeddings). These high-dimensional semantic vectors capture the semantic information of the prompt words and serve as the guiding signals for the subsequent generation process. Among them, the prompt word is the key information source for guiding the diffusion model to generate specific images. The prompt word is a natural language text description designed to convey various features, attributes, scenes, styles, etc. of the image that the user expects to generate. The diffusion model generates an image that conforms to the content of the prompt word by understanding and analyzing the prompt word and converting this semantic information into a guiding signal in the image generation process. The VAE decoder is used to decode the output result of the Unet module to obtain the generated image.
[0043] In some embodiments of the present application, when the computer device detects an operation by the user to wake up the display screen of the computer device, the computer device controls the display screen to display a display interface, and multiple controls are presented on the display interface. Different controls can correspond to different functions. For example, each control corresponds to an application (APP). The user can change the position of the control on the display interface by dragging, and can also click on the control to enter the application interface corresponding to the control. Multiple display areas can be presented in the application interface, and by selecting a display area, the information, pictures, etc. corresponding to the display area can be jumped and presented.
[0044] In one example, an image generation control is displayed on the display interface. By clicking on this image generation control, the user enters the application interface of the application indicated by the image generation control. On this application interface, a text input box can be displayed for the user to enter a prompt. The prompt can describe the desired image subject in a simple descriptive manner. For example, "a cat", "a castle". The prompt can also describe the desired image in detail from multiple aspects in a detailed descriptive manner. For example, "a cat with golden fur and blue eyes, sitting on the green grass, with a castle with a red roof behind it". The prompt can also describe the desired image style in a stylized descriptive manner. For example, "a painting in the impressionist style, showing a cat playing in the garden".
[0045] The above-described prompts are all presented in Chinese. In actual applications, they can also be presented in non-Chinese forms. This application does not limit the language type of the prompts. For example, the prompt "A closeup of a man", the prompt "A closeup of a man with robot head".
[0046] In some embodiments of the present application, the first prompt is converted into a first prompt vector by using a text encoder in a preset diffusion model. Since the computer device cannot directly understand the first prompt, it is necessary to use the text encoder to convert the first prompt into a vector form that the computer device can understand, that is, the first prompt vector. The first prompt vector contains the semantic information of the first prompt and plays a role in guiding the diffusion model to generate an image that conforms to the description of the first prompt in the subsequent process.
[0047] After determining the first prompt vector, starting from a preset random noise, the first denoising module fuses the first prompt vector during the denoising process in the first latent space. Through multiple steps of iteration, the noise is gradually removed in the first latent space and a latent variable corresponding to the first prompt vector is generated. The latent variable can be understood as an intermediate representation containing image feature information.
[0048] When the first denoising module generates the latent variable, the latent variable is transmitted to the VAE decoder. The VAE decoder uses the learned mapping relationship to decode the information in the latent variable into a high-resolution image that conforms to the semantic description, that is, the first image.
[0049] In other embodiments of the present application, the first prompt can be a sentence for generating a product image, which is used to generate a product image in the e-commerce field.
[0050] Step S202, determine the region of interest in the first image, and obtain the mask image corresponding to the region of interest.
[0051] In some embodiments of the present application, the region of interest (ROI) of an image refers to a specific region in the image that has specific significance and requires key attention or processing. After the computer device generates the first image using the diffusion model, it can call a preset image segmentation model to segment the first image, thereby determining the region of interest in the first image. Among them, the image segmentation model is a pre-trained model, which can be a convolutional neural network (CNN), an encoder-decoder architecture, or a combination of one or more of these models.
[0052] In one example, the first image is a portrait. Using the image segmentation model, the portrait can be segmented into multiple different regions, and each region represents a semantic object or a part with similar features. The multiple different regions include the head region and the non-head region of the portrait. The multiple different regions can be displayed on the display interface of the computer device, and the user can select at least one region from the multiple different regions as the region of interest.
[0053] In some embodiments of the present application, in addition to using the image segmentation model to determine the region of interest, a target detection model can also be used to identify the first image, thereby detecting the region of interest from the first image. Among them, the target detection model can be constructed based on one or more of a convolutional neural network (CNN), a region proposal network (RPN), and a feature pyramid network (FPN).
[0054] Specifically, the computer device reads the first image and converts the read first image into the input format required by the target detection model. Generally, the target detection model requires the input first image to be of a specific size and data type (such as in tensor form). The adjustment of the first image is achieved by operations such as resizing the image and normalizing the pixel values to a specific range (such as [0, 1] or [-1, 1]). The preprocessed first image is input into the loaded target detection model. The target detection model will perform feature extraction and analysis on the first image to identify the region of interest.
[0055] In addition, after generating the first image, the computer device can also display the first image on its display interface. The user can perform annotation operations such as box selection and clicking on the first image on the display interface. The computer device responds to the user's annotation operation on the first image, thereby determining the region of interest on the first image.
[0056] In some embodiments of the present application, after determining the region of interest, a mask image corresponding to the region of interest can be obtained. A mask image is a binary image (usually composed of black and white, corresponding to numerical values 0 and 1 or other specific values to represent different states), which is used to specify which parts of the image need to be processed, concerned, or retained, and which parts should be ignored or masked. If an image segmentation model or an object detection model is used to determine the region of interest of the first image, the mask image corresponding to the region of interest can be directly obtained from the output result of the image segmentation model or the object detection model. If the region of interest is determined through the user's annotation operation, an image processing tool (such as Photoshop) can be used to create a mask image with the region of interest being black (pixel value 0) and the non - region of interest being white (pixel value 255).
[0057] The above are just examples, and the present application does not limit the determination method of the region of interest, which can be determined according to the actual application.
[0058] Step S203: Process the first prompt word using the first denoising module, process the second prompt word for modifying the region of interest using the second denoising module of the diffusion model, and determine the second image in combination with the mask image.
[0059] In some embodiments of the present application, the diffusion model further includes a second denoising module, and the second denoising module can be a Unet module. The first denoising module and the second denoising module can be two modules with the same structure and can operate in cooperation. To avoid damaging the local image corresponding to the non - region of interest during the generation of the second image, the first denoising module is used to process the first prompt word. Specifically, the first prompt word is converted into a first prompt vector using the text encoder of the diffusion model. The first denoising module fuses the first prompt vector during the denoising process in the first latent space and combines it with the mask image to obtain a first feature map, which is the feature map corresponding to the non - region of interest.
[0060] The computer device receives the second prompt word for modifying the region of interest, and converts the second prompt word into a second prompt vector using the text encoder of the diffusion model. The second denoising module fuses the second prompt vector during the denoising process in the second latent space and combines it with the mask image to obtain a second feature map, which is the feature map corresponding to the region of interest.
[0061] After the computer device determines the first feature map and the second feature map, it uses the second denoising module to fuse the first feature map and the second feature map to obtain the second image. For the detailed process of the diffusion model generating the second image, please refer to the following Figure 5 as shown.
[0062] Next, in combination with Figure 3 Describe the interaction process of the first denoising module, the second denoising module, and the information exchange module for storing the mask map. The diffusion model receives the first prompt word, for example, "A closeup of a man"; and receives the second prompt word, for example, "A closeup of a man with robot head". Input the first prompt word and the second prompt word into the text encoder (Text Encoder) to obtain the first prompt vector and the second prompt vector. As Figure 3 shown, different text encoders are used to process the first prompt word and the second prompt word respectively. In another example, a single text encoder can also be used to process the first prompt word and the second prompt word simultaneously. The present application is not limited thereto.
[0063] Starting from random noise (Noise), during the denoising process of the first denoising module in the first latent space, fuse the first prompt vector, and combine it with the mask map in the information exchange module to obtain the first feature map. In one example, the first denoising module can call the mask map in the information exchange module to fuse with the first feature generated by the first denoising module, so as to generate the first feature map in the first denoising module, and store the first feature map in the information exchange module. In another example, after the first denoising module fuses the first prompt vector to generate the first feature during the denoising process in the first latent space, it can send the first feature to the information exchange module, and in the information exchange module, fuse the first feature and the mask map to generate the first feature map.
[0064] Starting from random noise, during the denoising process of the second denoising module in the second latent space, fuse the second prompt vector, and combine it with the mask map in the information exchange module to obtain the second feature map. Input the first feature map in the information exchange module into the second denoising module, and the second denoising module fuses the first feature map and the second feature map. Encode the latent variable generated by fusing the first feature map and the second feature map output by the last layer of the second denoising module to obtain the second image. As Figure 3 shown, the diffusion model also outputs the first image and the second image simultaneously. The man in the first image is not wearing a mechanical mask, and the man in the second image is wearing a mechanical mask.
[0065] To better present the image effects of the embodiments of the present application, the following will be combined with Figure 4 to further describe the first image and the second image. As Figure 4As shown, the first image can be an effect diagram generated based on the first prompt "a little boy wearing a blue T-shirt on the beach by the sea". If the user is not satisfied with the hairstyle of the little boy in the first image, then the hairstyle of the little boy needs to be locally modified. The area where the hair is located is extracted as the region of interest. Based on the first prompt "a little boy wearing a blue T-shirt on the beach by the sea", the second prompt "a curly-haired little boy wearing a blue T-shirt on the beach by the sea", and the mask image corresponding to the hair, the first denoising module and the second denoising module of the diffusion module generate the first image and the second image as shown in Figure 4 shown. The above embodiments meet the user's need for local modification of images.
[0066] Through the above embodiments, the first denoising module of the preset diffusion model is used to denoise in the first latent space, and the first prompt vector corresponding to the preset first prompt is fused to generate the first image corresponding to the first prompt. The initial image expected by the user is generated using the diffusion model. Further, for local modification of the image on the initial image, the region of interest in the first image is determined, and the mask image corresponding to the region of interest is obtained. After determining the mask image, in order to avoid damaging the part of the first image corresponding to the first prompt that is not the region of interest, the first denoising module is used to process the first prompt. In addition, the second denoising module of the diffusion model is used to process the second prompt for modifying the region of interest. Combining the processing of the first prompt, the processing of the second prompt, and the mask image, the second image is generated. The embodiments of the present application can meet the local image modification of the image generated based on the diffusion model without retraining other diffusion models, improving the local image modification efficiency of the image generated based on the diffusion model.
[0067] Figure 5 is the flowchart for generating the second image provided by an embodiment of the present application. As shown in Figure 5 shown, it includes the following steps: Step S501, convert the first prompt into a first prompt vector, and convert the second prompt into a second prompt vector.
[0068] In some embodiments of the present application, the diffusion model includes a text encoder. After receiving the first prompt and the second prompt, the diffusion model calls the text encoder to extract the first prompt vector of the first prompt and the second prompt vector of the second prompt. Among them, the first prompt vector and the second prompt vector are vector representations that the diffusion model can understand, facilitating subsequent processing. As shown in the structure of the diffusion model in Figure 3 shown, the diffusion model can include a text encoder to process multiple prompts simultaneously. The diffusion model can also include multiple text encoders to process a single prompt separately. The present application does not limit the number of text encoders in the diffusion model.
[0069] In step S502, during the denoising process of the first denoising module in the first latent space, the first hint vector is fused, and a first feature map is obtained in combination with the mask map.
[0070] In some embodiments of the present application, the structure of the first denoising module is an encoder-decoder structure. The first denoising module includes multiple levels, namely an encoder level and a decoder level. Among them, the encoder level performs step-by-step downsampling operations. The encoder level is generally composed of multiple convolutional blocks (Convolutional Block) and a pooling layer (such as max pooling MaxPooling). Each convolutional block contains several convolutional layers, activation functions (such as ReLU), and non-linear transformations for learning data features. The pooling layer reduces the width and height of the feature map through downsampling, increasing the receptive field, so as to capture context information in a larger range. As the number of network layers increases, the encoder continuously integrates local features to form a more global and abstract feature representation. The decoder level performs step-by-step upsampling operations. The decoder level is usually composed of multiple transposed convolutional blocks (TransposedConvolutional Block or Upsampling Block). The transposed convolutional layer (or the upsampling layer combined with the convolutional layer) is used to increase the size of the feature map, and each transposed convolutional block is also followed by a convolutional layer and an activation function for further processing and refining the upsampled features.
[0071] In some embodiments of the present application, assume that the current level is the i-th level. In the first latent space, the dimension of the first hint vector is expanded to match the number of channels of the latent space features corresponding to the i-th level. The first hint vector with the expanded dimension is fused with the latent space features corresponding to the i-th level to obtain the i-th level feature output by the i-th level in the multiple levels of the first denoising module, where i belongs to positive integers. The dimension of the mask map is updated so that the dimension of the mask map is the same as the dimension of the i-th level feature output by the i-th level. The mask map with the updated dimension is fused with the i-th level feature output by the i-th level to obtain the first feature map output by the i-th level of the first denoising module. It is expressed by the formula as fi×(1 - mask), where fi is the i-th level and mask is the mask map.
[0072] In step S503, during the denoising process of the second denoising module in the second latent space, the second hint vector is fused, and a second feature map is obtained in combination with the mask map.
[0073] In some embodiments of the present application, the structure of the second denoising module is an encoder-decoder structure. The second denoising module includes multiple levels, namely an encoder level and a decoder level. Among them, the encoder level performs step-by-step downsampling operations. The encoder level is generally composed of multiple convolutional blocks and a pooling layer (such as max pooling). Each convolutional block contains several convolutional layers, activation functions (such as ReLU), and non-linear transformations for learning data features. The pooling layer reduces the width and height of the feature map through downsampling, increasing the receptive field, thereby capturing context information in a larger range. As the number of network layers increases, the encoder continuously integrates local features to form more global and abstract feature representations. The decoder level performs step-by-step upsampling operations. The decoder level is usually composed of multiple transposed convolutional blocks (Transposed Convolutional Block or Upsampling Block). The transposed convolutional layer (or the upsampling layer combined with the convolutional layer) is used to increase the size of the feature map. Each transposed convolutional block is also followed by a convolutional layer and an activation function for further processing and refining the upsampled features.
[0074] In some embodiments of the present application, assume that the current level is the i-th level. In the second latent space, the i-th level receives the output result of the (i - 1)-th level, converts the result of the (i - 1)-th level into the latent space features corresponding to the i-th level. The second prompt vector is dimensionally expanded to match the number of channels of the latent space features corresponding to the i-th level. The dimensionally expanded second prompt vector is fused with the latent space features corresponding to the i-th level to obtain intermediate features. The dimension of the mask map is updated so that the dimension of the mask map is the same as that of the intermediate features. The mask map with the updated dimension is fused with the intermediate features to obtain the second feature map determined by the i-th level of the second denoising module. It is expressed by the formula fi×mask, where fi is the i-th level and mask is the mask map.
[0075] Step S504, fuse the first feature map and the second feature map to obtain the second image.
[0076] In some embodiments of the present application, since the first feature map contains the feature representations corresponding to the non-interested regions and the second feature map contains the feature representations corresponding to the interested regions, the second denoising module is used to fuse the first feature map and the second feature map.
[0077] In one example, the mask image is a binary image, usually represented by 0 and 1. 0 can represent the non - region of interest, and 1 can represent the region of interest. Suppose the first feature map is the fusion result of the i - th layer feature output by the i - th layer in the first denoising module and the mask image (such as the positions of 0 in the mask image), and the second feature map is the fusion result of the intermediate feature generated after processing the output result of the (i - 1) - th layer by the i - th layer in the second denoising module and the mask image (such as the positions of 1 in the mask image). The first feature map and the second feature map are fused in the i - th layer of the second denoising module as the input of the (i + 1) - th layer of the second denoising module. The output result of the last layer of the second denoising module is used as the second image. This not only retains the non - region of interest in the first image but also updates the region of interest in the first image, improving the efficiency of local modification of the image.
[0078] To describe the process of generating the second image more clearly, the following is described in combination with Figure 6 the schematic diagram shown. As Figure 6 shown, the diffusion model receives the first prompt "A closeup of a man" and the second prompt "A closeup of a man with robot head". The diffusion model converts the first prompt "A closeup of a man" into a first prompt vector through the text encoder. The diffusion model converts the second prompt "A closeup of a man with robot head" into a second prompt vector through the text encoder.
[0079] The same levels of the first denoising module and the second denoising module can be executed simultaneously. Starting from random noise, during the denoising process of the first denoising module in the first latent space, the first prompt vector is fused, and the first feature map is obtained by combining the mask image in the information exchange module. The first feature is generated by fusing the latent space feature corresponding to the i - th layer in the first denoising module and the first prompt vector. The first feature map is generated by fusing the first feature and the mask image. The i - th layer of the second denoising module converts the result of the (i - 1) - th layer into the latent space feature corresponding to the i - th layer. The intermediate feature is generated by fusing the latent space feature corresponding to the i - th layer in the second denoising module and the second prompt vector. The second feature map is generated by fusing the intermediate feature and the mask image. The first feature map and the second feature map are fused as the input of the (i + 1) - th layer of the second denoising module.
[0080] The above process is the processing process for the encoder levels in the first denoising module and the second denoising module. The decoder levels in the first denoising module and the second denoising module are similar to the encoder levels. As Figure 6As shown, starting from random noise, the first denoising module fuses the first prompt vector during denoising in the first latent space, and combines with the mask image in the information exchange module to obtain the first feature map. The latent space feature corresponding to the j-th layer of the first denoising module is fused with the first prompt vector to generate the first feature. The first feature is fused with the mask image to generate the first feature map. The j-th layer of the second denoising module converts the result of the (j - 1)-th layer into the latent space feature corresponding to the j-th layer. The latent space feature corresponding to the j-th layer of the second denoising module is fused with the second prompt vector to generate an intermediate feature. The intermediate feature is fused with the mask image to generate the second feature map. The first feature map and the second feature map are fused as the input of the (j + 1)-th layer of the second denoising module.
[0081] Through the above embodiments, two Unet modules are deployed in the diffusion module. In a collaborative manner, the first denoising module is used to retain the non-interested regions, and the second denoising module is used to modify the interested regions. It can maximize the satisfaction of user needs and improve the local modification efficiency of the image.
[0082] Figure 7 is a schematic diagram of an image processing device provided by an embodiment of the present application. As Figure 7 is a functional embodiment of the image processing method of the present application. The image processing device includes an image generation module 701, an image acquisition module 702, and an image modification module 703, where: The image generation module 701 is configured to perform denoising in the first latent space using the first denoising module of a preset diffusion model, and fuse the first prompt vector corresponding to a preset first prompt word to generate the first image corresponding to the first prompt word. The image acquisition module 702 is configured to determine the interested region in the first image and acquire the mask image corresponding to the interested region. The image modification module 703 is configured to process the first prompt word using the first denoising module, process the second prompt word for modifying the interested region using the second denoising module of the diffusion model, and determine the second image in combination with the mask image.
[0083] Based on any embodiment of the present application, the step of performing denoising in the first latent space using the first denoising module of a preset diffusion model and fusing the first prompt vector corresponding to a preset first prompt word to generate the first image corresponding to the first prompt word includes: converting the first prompt word into the first prompt vector; starting from a preset random noise, fusing the first prompt vector during the denoising process of the first denoising module in the first latent space to generate a latent variable corresponding to the first prompt vector; and decoding the latent variable to obtain the first image.
[0084] Based on any embodiment of the present application, the process of processing the first prompt using the first denoising module, processing the second prompt for modifying the region of interest using the second denoising module of the diffusion model, and determining the second image in combination with the mask image includes: converting the first prompt into a first prompt vector, and converting the second prompt into a second prompt vector; during the denoising process of the first denoising module in the first latent space, fusing the first prompt vector and obtaining a first feature map in combination with the mask image, where the first feature map is the feature map corresponding to the non-region of interest; during the denoising process of the second denoising module in the second latent space, fusing the second prompt vector and obtaining a second feature map in combination with the mask image, where the second feature map is the feature map corresponding to the region of interest; fusing the first feature map and the second feature map to obtain the second image.
[0085] Based on any embodiment of the present application, the process of fusing the first feature map and the second feature map to obtain the second image includes: using the second denoising module to fuse the first feature map and the second feature map to obtain the second image.
[0086] Based on any embodiment of the present application, the process of fusing the first prompt vector during the denoising process of the first denoising module in the first latent space and obtaining the first feature map in combination with the mask image includes: starting from random noise, fusing the first prompt vector during the denoising process in the first latent space, obtaining the i-th layer feature output by the i-th layer among multiple layers of the first denoising module, where i belongs to positive integers; based on the i-th layer feature and the mask image, obtaining the first feature map output by the i-th layer of the first denoising module.
[0087] Based on any embodiment of the present application, the process of fusing the second prompt vector during the denoising process of the second denoising module in the second latent space and obtaining the second feature map in combination with the mask image includes: starting from the random noise, fusing the second prompt vector during the denoising process in the second latent space, obtaining the intermediate feature generated after the i-th layer of the second denoising module processes the output result of the (i - 1)-th layer; based on the intermediate feature and the mask image, obtaining the second feature map determined by the i-th layer of the second denoising module.
[0088] Based on any embodiment of the present application, the fusion of the first feature map and the second feature map to obtain the second image includes: using the second denoising module to fuse the first feature map output by the i-th layer of the first denoising module and the second feature map determined by the i-th layer of the second denoising module to obtain the output result of the i-th layer of the second denoising module; determining the second image based on the output result of the last layer of the second denoising module.
[0089] Based on any embodiment of the present application, the determination of the region of interest in the first image includes: segmenting the first image based on a preset image segmentation model to determine the region of interest; or, in response to a user's annotation operation on the first image, determining the region of interest.
[0090] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions, and the method implemented when the program instructions are executed can refer to the methods in the above various embodiments of the present application. Among them, the computer-readable storage medium may be the internal memory of the electronic device in the above embodiment, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0091] In some embodiments, the computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the electronic device.
[0092] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0093] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0094] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division, and there can be other division methods in actual implementation.
[0095] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] In addition, in each embodiment of the present application, the functional modules can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0097] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be construed as limiting the claimed rights.
[0098] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices can also be implemented by one unit or device through software or hardware. Words such as first, second, etc. are used to represent names and do not represent any specific order.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. An image processing method, characterized in that, The method includes: Performing denoising in a first latent space using a first denoising module of a preset diffusion model, and fusing a first prompt vector corresponding to a preset first prompt to generate a first image corresponding to the first prompt; Determining a region of interest in the first image, and obtaining a mask image corresponding to the region of interest; Processing the first prompt using the first denoising module, processing a second prompt for modifying the region of interest using a second denoising module of the diffusion model, and determining a second image in combination with the mask image.
2. The image processing method according to claim 1, wherein The performing denoising in a first latent space using a first denoising module of a preset diffusion model, and fusing a first prompt vector corresponding to a preset first prompt to generate a first image corresponding to the first prompt includes: Converting the first prompt into the first prompt vector; Starting from a preset random noise, fusing the first prompt vector during the denoising process of the first denoising module in the first latent space to generate a latent variable corresponding to the first prompt vector; Decoding the latent variable to obtain the first image.
3. The image processing method according to claim 1, wherein The processing the first prompt using the first denoising module, processing a second prompt for modifying the region of interest using a second denoising module of the diffusion model, and determining a second image in combination with the mask image includes: Converting the first prompt into a first prompt vector, and converting the second prompt into a second prompt vector; Fusing the first prompt vector during the denoising process of the first denoising module in the first latent space, and obtaining a first feature map in combination with the mask image, where the first feature map is a feature map corresponding to a region other than the region of interest; Fusing the second prompt vector during the denoising process of the second denoising module in a second latent space, and obtaining a second feature map in combination with the mask image, where the second feature map is a feature map corresponding to the region of interest; Fusing the first feature map and the second feature map to obtain the second image.
4. The image processing method according to claim 3, wherein The fusing the first feature map and the second feature map to obtain the second image includes: Fusing the first feature map and the second feature map using the second denoising module to obtain the second image.
5. The image processing method according to claim 3, wherein The fusing the first prompt vector during the denoising process of the first denoising module in the first latent space, and obtaining a first feature map in combination with the mask image includes: Starting from random noise, fusing the first prompt vector during the denoising process in the first latent space, and obtaining an i-th layer feature output by the i-th layer of a plurality of layers of the first denoising module, where i belongs to positive integers; Based on the i-th layer feature and the mask image, obtaining a first feature map output by the i-th layer of the first denoising module.
6. The image processing method according to claim 5, characterized in that The fusing the second prompt vector during the denoising process of the second denoising module in a second latent space, and obtaining a second feature map in combination with the mask image includes: Starting from the random noise, during the denoising process in the second latent space, fuse the second prompt vector to obtain the intermediate features generated after processing the output result of the (i - 1)-th layer by the i-th layer among multiple layers of the second denoising module; Based on the intermediate features and the mask image, obtain the second feature map determined by the i-th layer of the second denoising module.
7. The image processing method according to claim 6, wherein The fusing the first feature map and the second feature map to obtain the second image includes: Using the second denoising module to fuse the first feature map output by the i-th layer of the first denoising module and the second feature map determined by the i-th layer of the second denoising module to obtain the output result of the i-th layer of the second denoising module; Based on the output result of the last layer of the second denoising module, determine the second image.
8. The image processing method according to claim 1, wherein The determining the region of interest in the first image includes: Segmenting the first image based on a preset image segmentation model to determine the region of interest; or, In response to a user's annotation operation on the first image, determine the region of interest.
9. An image processing apparatus, characterized in that, The image processing apparatus includes: An image generation module, configured to perform denoising in the first latent space using the first denoising module of a preset diffusion model and fuse the first prompt vector corresponding to a preset first prompt word to generate the first image corresponding to the first prompt word; An image acquisition module, configured to determine the region of interest in the first image and obtain the mask image corresponding to the region of interest; An image modification module, configured to process the first prompt word using the first denoising module, process the second prompt word for modifying the region of interest using the second denoising module of the diffusion model, and combine the mask image to determine the second image.
10. A computer device, characterized in that, including: A memory; and A processor; the processor executes the computer-readable instructions stored in the memory to implement the image processing method according to any one of claims 1 to 8.