Image processing method and image expansion method
By generating mask images and extracting, fusing, and adding noise to the image generation information, and using the guidance information from multiple image generation for denoising, the problem of inconsistent styles in the expanded regions in the expanded image model is solved, thereby improving the quality of the expanded images and the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-27
AI Technical Summary
Existing image expansion models based on diffusion models often result in inconsistent content and poor integration between the expanded and original images, leading to a reduced user experience.
By generating a mask image corresponding to the initial image, image generation information is extracted, fused, and noise is added. Multiple image generation guidance information is used to denoise the fused noisy image and generate the target image.
Ensure that the content of the expanded area is consistent with the style of the initial image to improve the quality of the generative image expansion results and enhance the user experience.
Smart Images

Figure CN121746191A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of computer technology, and in particular to an image processing method and an image expansion method. Background Technology
[0002] With the widespread use of mobile phones, cameras, and other devices, a large number of images are generated, and these images can be edited. Image expansion is an important function in image editing, which can meet users' needs for content creation and other aspects.
[0003] However, existing diffusion-based image expansion models, when used for generative image expansion, suffer from issues such as inconsistent style between the expanded area and the original image, poor integration with the original image, and a sense of disharmony, which degrade the user experience. Summary of the Invention
[0004] In view of this, embodiments of this specification provide an image processing method and an image expansion method. One or more embodiments of this specification also relate to an image processing apparatus, an image expansion apparatus, a computing device, a computer-readable storage medium, and a computer program product, to solve the technical defects in the prior art, such as inconsistencies between the expanded image content and the original image style, poor integration with the original image, and a sense of disharmony, when expanding an image.
[0005] According to a first aspect of the embodiments of this specification, an image processing method is provided, comprising: An initial image is input into an image processing model, and the image processing model is used to generate a mask image corresponding to the initial image, wherein the size of the mask image is larger than the size of the initial image; Image generation information is extracted from the initial image to determine multiple image generation guidance information; The mask image and the initial image are fused and noise-added to obtain a fused noise image; Based on the guidance information generated from the multiple images, the fused noisy image is denoised to obtain the target image output by the image processing model.
[0006] According to a second aspect of the embodiments of this specification, an image processing apparatus is provided, comprising: The generation module is configured to input an initial image into an image processing model and use the image processing model to generate a mask image corresponding to the initial image, wherein the size of the mask image is larger than the size of the initial image; The determination module is configured to extract image generation information from the initial image and determine multiple image generation guidance information; The fusion module is configured to fuse and add noise to the mask image and the initial image to obtain a fused noise image; The acquisition module is configured to generate guiding information based on the plurality of images to denoise the fused noisy image and obtain the target image output by the image processing model.
[0007] According to a third aspect of the embodiments of this specification, an image augmentation method is provided, comprising: The image to be expanded is input into the image processing model, and the image processing model is used to generate an expanded mask image corresponding to the image to be expanded, wherein the size of the expanded mask image is larger than the size of the image to be expanded; Image generation information is extracted from the image to be expanded to determine multiple image generation guidance information; The extended mask image and the image to be extended are fused and noise-added to obtain a fused noise image; Based on the guidance information generated from the multiple images, the fused noisy image is denoised to obtain the extended image output by the image processing model.
[0008] According to a fourth aspect of the embodiments of this specification, an image expansion device is provided, comprising: The generation module is configured to input the image to be expanded into an image processing model, and use the image processing model to generate an expanded mask image corresponding to the image to be expanded, wherein the size of the expanded mask image is larger than the size of the image to be expanded; The determination module is configured to extract image generation information from the image to be expanded and determine multiple image generation guidance information; The fusion module is configured to fuse and add noise to the extended mask image and the image to be expanded to obtain a fused noise image; The acquisition module is configured to generate guiding information based on the multiple images to denoise the fused noisy image and obtain the extended image output by the image processing model.
[0009] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, they implement the steps of the above-described image processing method and image expansion method.
[0010] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions, which, when executed by a processor, implement the steps of the above-described image processing method and image expansion method.
[0011] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described image processing method and image expansion method.
[0012] An embodiment of this specification provides an image processing method that generates a mask image corresponding to an initial image through an image processing model. The size of the mask image is larger than that of the initial image, thereby expanding the image based on the mask image. The mask image and the initial image are then fused and denoised to obtain a fused noise image. Denoising the fused noise image yields a clear target image expanded from the initial image. By determining multiple image generation guidance information for the initial image, information about the style, content, and features of the initial image is obtained. This information is then used to denoise the fused noise image during the denoising process. Guided by this information, the content of the expanded region in the generated target image is made consistent with the style and content of the initial image, avoiding problems such as style inconsistency, poor fusion, and disharmony between the expanded region and the initial image, thus improving the quality of the generative image expansion result. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of a scene illustrating an image processing method provided in one embodiment of this specification; Figure 2 This is a flowchart illustrating an image processing method provided in one embodiment of this specification; Figure 3 This is a flowchart of an image expansion method provided in one embodiment of this specification; Figure 4a This is a flowchart illustrating the processing procedure of an image expansion method provided in one embodiment of this specification; Figure 4b This is a schematic diagram of the architecture of an image expansion system provided in one embodiment of this specification; Figure 4c This is a schematic diagram illustrating the noise reduction and color adjustment process in an image expansion method provided in one embodiment of this specification; Figure 4d This is a schematic diagram of the guided denoising process in an image expansion method provided in one embodiment of this specification; Figure 5This is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification; Figure 6 This is a schematic diagram of the structure of an image expansion device provided in one embodiment of this specification; Figure 7 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0014] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0017] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0018] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0019] Generative image expansion: refers to using a diffusion model to expand an image outward and supplement the content of the expanded area.
[0020] Diffusion models are a type of latent variable model, which is a Markov chain trained using variational estimation. Diffusion models learn the latent structure of a dataset by modeling how data points diffuse in the latent space. After training, randomly sampled noise can be fed into the diffusion model, and data can be generated by learning a denoising process.
[0021] This specification provides an image processing method and an image expansion method. It also relates to an image processing apparatus, an image expansion apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0022] See Figure 1 , Figure 1 A schematic diagram of a scene illustrating an image processing method provided in one embodiment of this specification is shown. Specifically, this image processing method is implemented using an edge device 102 and a server 104. The edge device 102 sends an initial image to the server 104, such as a "close-up image of a person". An image processing model is trained in the server 104. When the server 104 receives the initial image sent by the edge device 102, it inputs the initial image into the image processing model and uses the model to generate a mask image corresponding to the initial image. The size of the mask image is larger than that of the initial image. Image generation information is extracted from the initial image to determine multiple image generation guidance information. The mask image and the initial image are fused and denoised to obtain a fused noise image. The fused noise image is denoised according to the multiple image generation guidance information to obtain the target image output by the image processing model. This target image can be an "image that expands the background of the close-up image", and the target image is returned to the edge device 102.
[0023] The edge device 102 may include a browser, an app (application), or a web application such as an H5 (Hypertext Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. The edge device can be developed based on a software development kit (SDK) provided by the server, such as a real-time communication (RTC) SDK. The edge device can be deployed in an electronic device and depends on the device's operation or certain apps within the device to run. The electronic device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured in the electronic device, such as human-computer interaction applications, model training applications, image processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.
[0024] Server 104 can be understood as a server providing various services, including physical servers and cloud servers. Examples include servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It's important to note that Server 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. Server 104 can also be a server in a distributed system, or a server integrated with blockchain. Server 104 can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0025] It is worth noting that the image processing method provided in the embodiments of this specification can be executed by the server 104. In other embodiments of this specification, the image processing model can be deployed in the end device 102, so that the end device 102 can also have similar functions to the server 104, thereby executing the image processing method provided in the embodiments of this specification. In other embodiments, the image processing method provided in the embodiments of this specification can also be jointly executed by the end device 102 and the server 104.
[0026] An embodiment of this specification provides an image processing method that generates a mask image corresponding to an initial image through an image processing model. The size of the mask image is larger than that of the initial image, thereby expanding the initial image based on the mask image. By determining multiple image generation guidance information of the initial image, information about the style, content, and features of the initial image is obtained. The mask image and the initial image are fused and denoised to obtain a fused noise image. By denoising the fused noise image, a clear target image expanded from the initial image is obtained. In the denoising process, when the fused noise image is denoised using multiple image generation guidance information, the content of the expanded area in the generated target image can be guided to be consistent with the style and content of the initial image, avoiding problems such as inconsistent style, poor fusion, and disharmony between the expanded area and the initial image in the target image, thus improving the quality of the generative image expansion result.
[0027] See Figure 2 , Figure 2 A flowchart of an image processing method provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0028] Step 202: Input the initial image into the image processing model, and use the image processing model to generate a mask image corresponding to the initial image.
[0029] Wherein, the size of the mask image is larger than the size of the initial image; the initial image can be understood as the image to be processed, and subsequent image expansion processing is performed on the initial image; the mask image can be understood as the mask image of the expanded area when performing generative image expansion on the initial image, that is, the expanded image content of the initial image is generated on the mask image.
[0030] Of course, in practical applications, a mask image can be obtained outside of the image processing model. For example, the mask image can be manually drawn using image processing software (such as Photoshop) or programming tools (such as OpenCV), and then input into the image processing model along with the initial image.
[0031] In practical applications, before inputting the initial image into the image processing model and generating the mask image corresponding to the initial image using the image processing model, the process further includes: The initial image is received from the client, wherein the initial image is sent by the client in response to an interactive operation command from the user interface.
[0032] Specifically, the initial image can be an image to be processed sent by the user through the client. For example, the user can perform interactive operations such as uploading an image and confirming the sending on the user interface of the client, so that the client can send the initial image uploaded by the user to the server based on the interactive operation instructions of the user interface.
[0033] Of course, multiple images can also be displayed on the client's user interface. Users can select one image from the multiple images as the initial image for confirmation and sending; there are no restrictions on this.
[0034] The image processing method provided in the embodiments of this specification allows users to conveniently and quickly determine and send the initial image through the user interface provided by the client, thereby improving the user's interactive experience.
[0035] In one or more embodiments of this specification, the image processing model includes a mask generation unit. Therefore, an initial image can be input into the mask generation unit, and the mask generation unit can be used to obtain a mask image corresponding to the initial image. Specific implementation methods are as follows: The step of generating a mask image corresponding to the initial image using the image processing model includes: The mask image is generated using the mask generation unit to obtain the mask image corresponding to the initial image.
[0036] The mask generation unit is used to generate a mask image of the input image. For example, the mask generation unit can be a machine learning model. Through model training, the size of the generated mask image can be larger than the size of the initial image. The size of the image can be represented by pixel values in the width and height directions, or by length in the width and height directions (e.g., in centimeters).
[0037] When the size of an image is represented by the pixel values in the width and height directions, the generated mask image can have more pixel values in both the width and height directions than the initial image when the mask generation unit generates the mask image. This expands the initial image in both the width and height directions. For example, if the size of the initial image is 400x300 pixels, the size of the mask image is 800x600 pixels.
[0038] Of course, in practical applications, the generated mask image can also be used to expand the initial image in the width or height direction. For example, if the size of the initial image is 400x300 and the size of the mask image is 800x300, the initial image can be expanded in the width direction. Or, if the size of the initial image is 400x300 and the size of the mask image is 400x600, the initial image can be expanded in the height direction.
[0039] The image processing method provided in the embodiments of this specification can quickly and efficiently generate a mask image through the mask generation unit of the image processing model, generate extended image content on the mask image, and use the mask image to realize the extension of the initial image, providing a data foundation for the subsequent extension of the initial image.
[0040] In one or more embodiments of this specification, when there is no specific requirement for the size of the mask image, the initial image can be directly input into the mask generation unit to generate the mask image, and the size of the mask image is random. However, in practical applications, the size of the mask image can be specified using the target extended size to achieve personalized requirements. Specific implementation methods are described below: Before inputting the initial image into the image processing model and generating the mask image corresponding to the initial image using the image processing model, the method further includes: In response to an image processing instruction, the initial image carried in the image processing instruction and the target expansion size for the initial image are determined; The step of inputting the initial image into the image processing model and using the image processing model to generate a mask image corresponding to the initial image includes: The initial image and the target extended size are input into the image processing model. The image processing model is used to extend the initial image by masking it according to the target extended size, thereby generating a mask image corresponding to the initial image.
[0041] In practical applications, before inputting the initial image into the image processing model and generating the mask image corresponding to the initial image using the image processing model, the process further includes: The initial image and image processing prompt text are determined; The step of inputting the initial image into the image processing model and using the image processing model to generate a mask image corresponding to the initial image includes: The initial image and the image processing prompt text are input into the image processing model. The image processing model is used to expand the initial image by masking it according to the target expansion size contained in the image processing prompt text, thereby generating a mask image corresponding to the initial image.
[0042] The target expansion size can be understood as the specific size of the mask image. For example, if the image processing prompt text is "Expand to generate an image with a size of 800x600", the target expansion size is 800x600 pixels. The target expansion size can also be understood as the size of the expansion based on the initial image. For example, if the image processing prompt text is "Expand the width by 200 pixels based on the input image", the target expansion size is 200 pixels in the width direction.
[0043] The target expansion size is determined by image processing instructions or image processing prompt text. Then, a mask image of a specific size is generated by expanding the initial image according to the target expansion size.
[0044] In practical applications, the client's user interface may include an image upload control and a target extended size selection control. Responding to user interactions with these controls, the client sends an image processing instruction to the server. This instruction includes the initial image uploaded by the user and the target extended size for that initial image. Alternatively, the client's user interface may include an image upload control and a text input control. The user can interact with the input control to enter image processing prompt text, which includes the target extended size.
[0045] The image processing method provided in the embodiments of this specification enables the image processing model to generate a mask image of a specific size based on the initial image and the target expansion size by using various methods such as carrying the target expansion size in the image processing instructions and including the target expansion size in the image processing prompt text, thereby meeting the user's personalized needs.
[0046] Step 204: Extract image generation information from the initial image to determine multiple image generation guidance information.
[0047] Image generation guidance information can be understood as information used to guide image generation. For example, image generation guidance information includes image style information, image content information, image feature information, etc. Through image generation guidance information, the style and content of the generated image can be guided during the image generation process.
[0048] Specifically, by extracting image generation information from the initial image, information such as image style, image content, and image features can be obtained. In the subsequent image generation process, this information can be used as a guide to ensure that the generated image matches the style and content of the initial image, thereby improving the image quality of generative image expansion.
[0049] In one or more embodiments of this specification, the image processing model includes an extended guidance unit, and the image generation guidance information includes text description guidance information and image feature guidance information; in the extended guidance unit, the text description guidance information and image feature guidance information of the initial image are obtained. Specific implementation methods are as follows: The step of extracting image generation information from the initial image and determining multiple image generation guidance information includes: Using the extended guidance unit, image generation information is extracted from the initial image to obtain the text description guidance information and the image feature guidance information of the initial image.
[0050] The text description guidance information can include image style information and image content information of the initial image. This image style information and image content information can be determined using the image text description of the initial image. The image feature guidance information can be understood as the feature information of the image content of the initial image in the latent space. The latent space is a multi-dimensional vector space learned by the model. In this space, the complex features of the original data are represented as a set of simpler variables, which are usually continuously distributed. These variables capture the key characteristics of the input data and often have certain semantic meaning.
[0051] Specifically, by extending the guidance unit, image generation information is extracted from the initial image to obtain textual description guidance information including the image style information and image content information of the initial image, as well as the feature information of the image content in the latent space of the initial image; the textual description guidance information and the image feature guidance information can be data represented by feature vectors through feature extraction.
[0052] In practical applications, the extended guidance unit includes a multimodal processing module and an image encoding module; the multimodal processing module obtains textual description guidance information of the initial image, and the image encoding module obtains image feature guidance information of the initial image. Specific implementation details are as follows: The step of using the extended guidance unit to extract image generation information from the initial image to obtain the text description guidance information and the image feature guidance information of the initial image includes: The multimodal processing module is used to perform text description processing on the initial image to obtain the text description guidance information; The initial image is processed using the image encoding module to obtain the image feature guidance information.
[0053] The multimodal processing module can be understood as an Image Tagger, which is used to add descriptive labels to the content in an image. In practical applications, this multimodal processing module can be implemented using a multimodal large model. That is, by inputting the initial image into the multimodal large model, the multimodal large model is used to perform text description on the initial image to obtain the image text description of the initial image and determine the text features of the image text description (i.e., text description guidance information).
[0054] The image encoding module can be understood as a module implemented using an image encoder. It inputs an initial image into the image encoder, which then performs image encoding processing on the initial image to obtain the image features (i.e., image feature guidance information) of the initial image.
[0055] The image processing method provided in the embodiments of this specification can not only obtain the text description guidance information corresponding to the initial image through the multimodal processing module, but also obtain the image feature guidance information of the initial image through the image encoding module. When using multiple guidance information to guide image generation, it helps to improve the quality of the generated target image.
[0056] Step 206: Perform fusion and noise addition processing on the mask image and the initial image to obtain a fused noise image.
[0057] Specifically, the image processing model includes an image generation unit; firstly, the mask image and the initial image are fused, and then noise is added to the fused image to obtain a fused noise image. The specific implementation is as follows: The process of fusing and adding noise to the mask image and the initial image to obtain a fused noise image includes: The image generation unit is used to fuse the mask image and the initial image to obtain a fused image, and then the fused image is subjected to noise processing to obtain a fused noise image.
[0058] The image generation unit can be understood as a diffusion model, which uses the diffusion model to fuse the mask image and the initial image and add noise.
[0059] Specifically, fusing the mask image and the initial image can be understood as superimposing the mask image and the initial image to obtain a fused image. For example, if the width and height of the initial image need to be expanded by a factor of two, the initial image can be placed in the middle area of the mask image. On this basis, other areas in the mask image can be filled with the edge color of the initial image to obtain a fused image. The other areas in the mask image are the areas to be expanded corresponding to the initial image.
[0060] Random noise, such as Gaussian noise, is introduced into the fused image to obtain a fused noisy image. This allows for the generation of image content that matches the initial image in other areas of the mask image when the fused noisy image is subsequently denoised, thus obtaining a target image that is expanded based on the initial image.
[0061] The image processing method provided in the embodiments of this specification obtains a fused noisy image by fusing and adding noise to a mask image and an initial image, which provides the data basis for subsequent denoising to generate a target image. That is, the target image is obtained by denoising the fused noisy image.
[0062] Step 208: Generate guiding information based on the multiple images to denoise the fused noisy image and obtain the target image output by the image processing model.
[0063] The target image can be understood as an expanded image derived from the initial image, and the size of the target image is the same as the size of the mask image.
[0064] Specifically, multiple image generation guidance information is injected into the denoising process of the fused noisy image, so that the target image output by the image processing model can be obtained under the guidance of multiple image generation guidance information, or with reference to multiple image generation guidance information.
[0065] In one or more embodiments of this specification, in the image generation unit, multiple image generation guidance information is used to denoise the fused noisy image to obtain a target image after expanding the initial image. In practical applications, the image generation unit includes a feature denoising module, and the multiple image generation guidance information includes text description guidance information and image feature guidance information. Specific implementation methods are as follows: The step of generating guiding information based on the multiple images to denoise the fused noisy image and obtain the target image includes: The feature denoising module is used to extract features from the fused noise image to obtain the noise image features of the fused noise image; Using the feature denoising module, through a cross-attention mechanism, the noisy image features are denoised based on the text description guidance information and the image feature guidance information to obtain denoised image features; The target image is obtained by using the feature denoising module based on the denoised image features.
[0066] In the case where the image generation unit is a diffusion model (which can generate a clear image by gradually reducing noise in the image), the feature denoising module can be understood as the denoising module of the diffusion model. The denoising module is used to achieve gradual denoising of the fused noisy image. The denoised image features can be understood as the remaining information features in the fused noisy image after a series of denoising processes. These features have removed noise components as much as possible and are closer to the features of the clear target image, which is an abstract representation that can describe the content of the target image.
[0067] Specifically, feature extraction is performed on the fused noisy image. Key feature representations are extracted from the noisy fused noisy image to obtain noisy image features, which will be used for subsequent denoising processing. A cross-attention mechanism is used to combine text description guidance information and image feature guidance information to denoise the noisy image features. Specifically, text description guidance information provides semantic information such as the style and content of the initial image, while image feature guidance information provides information at the image structure level. Therefore, through the cross-attention mechanism, text description guidance information and image feature guidance information are injected into the diffusion model in the denoising process of the fused noisy image, so that the diffusion model can consider both aspects of information at the same time, more accurately guide the denoising process, obtain denoised image features, and obtain the target image based on the denoised image features.
[0068] The image processing method provided in the embodiments of this specification injects text description guidance information and image feature guidance information into the diffusion model during the denoising process of the fused noisy image through a cross-attention mechanism, thereby guiding the diffusion model to obtain a target image that is more consistent with the style and content of the initial image.
[0069] In one or more embodiments of this specification, the feature denoising module includes multiple denoising network layers. When using a cross-attention mechanism to inject text description guidance information and image feature guidance information into the denoising process, the text description guidance information and image feature guidance information can be injected into the denoising network layers respectively. This allows for cross-attention processing between the text description guidance information and the noise image features of each denoising network layer, thereby denoising the noise image features of each denoising network layer. Similarly, cross-attention processing between the image feature guidance information and the noise image features of the target denoising network layer is performed to denoise the noise image features of the target denoising network layer. Specific implementation methods are described below: The feature denoising module, through a cross-attention mechanism, denoises the noisy image features based on the text description guidance information and the image feature guidance information to obtain denoised image features, including: The text description guidance information is cross-attention processed with the noise image features in each of the multiple denoising network layers, and the image feature guidance information is cross-attention processed with the noise image features in the target denoising network layer of the multiple denoising network layers, so as to denoise the noise image features and obtain the denoised image features output by the last denoising network layer in the multiple denoising network layers.
[0070] The denoising module of the diffusion model includes multiple denoising network layers. These layers learn features from the data layer by layer to gradually remove noise. The roles of each denoising network layer are not entirely the same; that is, each layer learns different features. Specifically, the target denoising network layer focuses on learning the style information of the image. For example, in the shallow network layers (the initial stage of the diffusion process, when the noise level is high), the diffusion model mainly learns the basic structure and content information of the noise-fused image. In the deep network layers (the later stages of the diffusion process, when the noise level gradually decreases and the image becomes clearer), the focus is more on learning the style, texture, and detail features of the image.
[0071] For example, the denoising module includes 11 denoising network layers. The 7th denoising network layer is determined as the target denoising network layer. This involves cross-attention processing between the text description guidance information and the noise image features in each denoising network layer, and cross-attention processing between the image feature guidance information and the noise image features of the 7th denoising network layer. This achieves denoising of the noise image features corresponding to the 7th denoising network layer. The noise image features input to the last denoising network layer, i.e., the 11th denoising network layer, are then determined as the denoised image features.
[0072] In practical implementation, the specific process of denoising the noisy fused image to generate the target image can be viewed in the diffusion model as a stepwise reverse diffusion process, whereby the noisy image (denoted as X) is denoised. T Denoising is performed until the original, clear image, unaffected by noise, is restored (denoted as X0); while gradually reducing X... T In the case of noise reduction, each step uses the noise reduction module to reduce X t Denoising is X t-1 In each step, the noise reduction module will denoise X. t Denoising is X t-1 At this time, the text description guidance information is extracted through the corresponding text cross-attention module and the noise image features of each denoising network layer in the denoising module. The output of the text cross-attention module is used as input and merged with the noise image features of that level in the denoising module. The merged result is then used as input to the next level of the denoising network layer. Following the previous example, the image feature guidance information will be injected into the 7th denoising network layer. That is, the image feature guidance information is extracted through the corresponding image cross-attention module and the noise image features of the 7th denoising network layer in the denoising module. The output of the image cross-attention module is used as input and merged with the noise image features of the 7th level in the denoising module. The merged result is then used as input to the 8th denoising network layer.
[0073] When text description guidance information and image feature guidance information are injected into the denoising network layer of the denoising module in the corresponding manner, the noisy image features output by the last denoising network layer among multiple denoising network layers will be determined as denoised image features.
[0074] The image processing method provided in the embodiments of this specification can ensure the consistency of the style of the generated target image with the style of the initial image by injecting text description guidance information into each denoising network layer and injecting image feature guidance information into the target denoising network layer.
[0075] In one or more embodiments of this specification, the image processing model includes an image generation unit, which includes a feature denoising module and a feature color adjustment module. To ensure that the generated target image has the same tone as the initial image, the feature color adjustment module performs color adjustment processing on the denoised image features output by the feature denoising module. Specific implementation methods are as follows: The step of generating guiding information based on the multiple images to denoise the fused noisy image and obtaining the target image output by the image processing model includes: Using the feature denoising module, the fused noisy image is denoised based on the guidance information generated from the multiple images to obtain denoised image features, wherein the denoised image features include channel image features of multiple channels; Using the feature color adjustment module, the channel image features of the target channel among the multiple channels are color adjusted to obtain the target image features; The target image is obtained based on the target image features.
[0076] Each denoised image feature can be viewed as a combination of one or more channel-specific image features, which contain abstract features of the noise-fused image in different dimensions and from different angles.
[0077] For example, taking a denoised image feature comprising four channels as an example, channel 0 is the luminance channel, affecting the brightness of the generated image, and the feature color adjustment module does not adjust it; channel 3 is the feature channel, affecting the content of the generated image, and the feature color adjustment module does not adjust it; channel 1 is the cyan / red color channel, affecting whether the generated image is cyan or red; specifically, the larger the value, the more cyan it is, and the smaller the value, the more red it is; channel 2 is the yellow / purple color channel, affecting whether the generated image is yellow or purple; specifically, the larger the value, the more yellow it is, and the smaller the value, the more purple it is; the feature color adjustment module performs color adjustment processing on the channel image features of channels 1 and 2; that is, in this embodiment of the specification, channels 1 and 2, which affect the color of the generated image, are determined as target channels.
[0078] In specific implementation, after obtaining denoised image features using the feature denoising module, these features are input into the feature color adjustment module. The feature color adjustment module performs color adjustment on the first and second channels of the denoised image features to obtain the target image features. Using these target image features, the target image is obtained. It should be noted that since the diffusion model performs stepwise denoising on the fused noisy image, each denoising step (where X...)... t Denoising is X t-1 All images will be processed by the feature denoising module and the feature color adjustment module. Therefore, the feature output by the feature color adjustment module in the last denoising step (denoising X1 to X0) is determined as the feature of the target image.
[0079] The image processing method provided in the embodiments of this specification can perform feature color adjustment on the denoised image features output by the feature denoising module through the feature color adjustment module, so that the generated target image can maintain the same tone as the initial image, avoiding the problem that the extended area in the target image is inconsistent with the tone of the initial image and the fusion with the initial image is poor.
[0080] In one or more embodiments of this specification, a color adjustment amplitude value is obtained by calculating the channel image features of the target channel. Then, the channel image features of the target channel are color-adjusted using this color adjustment amplitude value and the channel image features of the target channel. Specific implementation methods are described below: The step of using the feature-based color adjustment module to perform color adjustment processing on the channel image features of the target channel among the multiple channels to obtain the target image features includes: Using the feature-based color adjustment module, the vector mean of the channel image features of the target channel is calculated to obtain the mean channel image features, and the color adjustment amplitude value is calculated based on the mean channel image features; The target image features are obtained based on the channel image features of the target channel and the color adjustment amplitude value.
[0081] Specifically, the mean vector value of the channel image features of the target channel is calculated to obtain the mean channel image features representing the overall color level of the target channel. The color adjustment amplitude value is calculated based on the mean channel image features. The color adjustment amplitude value is then subtracted from the channel image features of the target channel to obtain the color-adjusted target image features.
[0082] In practical applications, the image features of the mean channel are calculated using the following formula: , in, In this round of denoising, the channel image feature of the i-th channel in the denoised image features output by the denoising module. In the embodiments of this specification, i represents the channel number, which can be 1 or 2. In To find the mean function, The mean channel image features are shown in the above embodiments.
[0083] The color correction range value is calculated using the following formula: , in, This ensures that even when tensor_mean is 0, the result will not be negative infinity (because the domain of the logarithmic function is positive). Furthermore, logarithmic transformation is typically used to compress data ranges, especially when the data has a long tail distribution, as it reduces the impact of extreme values. Therefore, this logarithmic function amplifies smaller mean differences in a non-linear way, while having less impact on larger mean differences. Dividing by 8 is a scaling factor (i.e., the constant value is not limited and can be flexibly set according to actual conditions). Constant division is used to scale the logarithmic value to fit a specific range, such as standardizing it to a smaller interval for further processing or display. In the embodiments of this specification, it is used to adjust the intensity of the color correction amplitude. This is the calculated color adjustment amplitude value for the i-th channel.
[0084] The color-corrected target image features are obtained using the following formula: , in, This is the output vector of the i-th channel after passing through the feature color adjustment module.
[0085] The image processing method provided in the embodiments of this specification accurately performs color adjustment processing on the channel image features of the target channel by calculating the obtained color adjustment amplitude value and the channel image features of the target channel, thereby obtaining the target image features.
[0086] In one or more embodiments of this specification, after generating guiding information based on the plurality of images to denoise the fused noisy image and obtain the target image output by the image processing model, the method further includes: The target image is returned to the client for display on the client's user interface.
[0087] The image processing method provided in the embodiments of this specification returns the target image to the client and displays the generated target image on the client's user interface, making it easier for users to view and improving the user experience.
[0088] The image processing method provided in this specification addresses the problem of inconsistent tones between the extended region of the target image and the initial image, resulting in poor fusion. The feature-based color adjustment module effectively avoids this issue by adjusting the color of the channel image features of the target channel. This helps maintain consistency between the extended region and the initial image's tones. For the problem of inconsistent styles between the extended region and the initial image, the extension guidance unit can obtain multiple image generation guidance information for the initial image. This guidance information includes various aspects such as the initial image's style, content, and features. This information is used to guide the diffusion process, ensuring that the style of the extended region matches the initial image, and that the generated content of the extended region matches the content of the initial image, thus avoiding any sense of incongruity.
[0089] See Figure 3 , Figure 3 A flowchart of an image expansion method provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0090] Step 302: Input the image to be expanded into the image processing model, and use the image processing model to generate an expanded mask image corresponding to the image to be expanded.
[0091] Wherein, the size of the extended mask image is larger than the size of the image to be extended; the image to be extended can be understood as the initial image in the above embodiments; the extended mask image can be understood as the mask image in the above embodiments; for specific implementation, please refer to the above embodiments, which will not be repeated here.
[0092] Step 304: Extract image generation information from the image to be expanded and determine multiple image generation guidance information.
[0093] Step 306: Perform fusion and noise addition processing on the extended mask image and the image to be extended to obtain a fused noise image.
[0094] Step 308: Generate guiding information based on the multiple images to denoise the fused noisy image and obtain the extended image output by the image processing model.
[0095] In specific implementation, the image processing model includes an image generation unit, which includes a feature denoising module and a feature color adjustment module; The step of generating guiding information based on the multiple images to denoise the fused noisy image and obtain the extended image output by the image processing model includes: Using the feature denoising module, the fused noisy image is denoised based on the guidance information generated from the multiple images to obtain denoised image features, wherein the denoised image features include channel image features of multiple channels; Using the feature color adjustment module, the channel image features of the target channel in the multiple channels are color adjusted to obtain extended image features; The extended image is obtained based on the extended image features.
[0096] The extended image can be understood as the target image in the above embodiments; for specific implementation, please refer to the above embodiments, which will not be repeated here.
[0097] The image expansion method provided in this specification uses a feature color adjustment module to adjust the color of the channel image features of the target channel. This helps to ensure that the expanded area of the expanded image maintains the same tone as the initial image, avoiding inconsistencies between the expanded area and the initial image's tone, and preventing poor fusion with the initial image. The expansion guidance unit can obtain multiple image generation guidance information for the initial image. When this guidance information includes various aspects such as the style, content, and features of the initial image, it guides the diffusion process, ensuring that the style of the expanded area remains consistent with the initial image, and guiding the generated image content in the expanded area to match the image content of the initial image, thereby avoiding any sense of disharmony.
[0098] See Figure 4a , Figure 4a A flowchart illustrating the processing steps of an image expansion method according to an embodiment of this specification is shown, specifically including the following steps.
[0099] Specifically, this image augmentation method is applied to an image augmentation system, see [link to relevant documentation]. Figure 4b , Figure 4b This document illustrates a schematic diagram of the architecture of an image expansion system according to an embodiment of this specification. An input image (i.e., the initial image in the above embodiment) and a mask of the region to be expanded (i.e., the mask image in the above embodiment) are input together into an image processing model. The image processing model includes an expansion style guidance unit (i.e., the expansion guidance unit in the above embodiment) and a color-corrected diffusion completion model (i.e., the image generation unit in the above embodiment). The expansion style guidance unit obtains text description guidance information and image feature guidance information of the initial image. These two guidance information, along with the input image and the mask of the region to be expanded, are then fed into the color-corrected diffusion completion model to obtain the expanded image result (output image, i.e., the target image in the above embodiment).
[0100] Step 402: Determine the mask image.
[0101] The mask image is determined based on the target expansion size, and then the mask image is used to perform masking processing on the input image to obtain the masked input image; for example, when the input image is expanded by 2 times in length and width, the masked input image is an image in which the input image is placed in the middle area on the canvas of the doubled input image, and the other areas are filled with the edge color of the input image.
[0102] Step 404: Extracting guiding information.
[0103] By utilizing the extended style guidance unit of the image processing model, style guidance information of the input image is obtained: image description text (including content information and style information) and image features (including feature information of image content in the latent space).
[0104] Specifically, the input image is sent to the extended style guidance unit, and two types of guidance information are obtained through two guidance information extraction modules, namely Image Tagger and Image encoder. Among them, Image Tagger refers to the image content description module, and its output information is the image description text of the input image (which can be understood as the text description guidance information in the above embodiment); Image encoder refers to the image encoding module, and its output information is the image features (which can be understood as the image feature guidance information in the above embodiment).
[0105] In practical applications, labeling models such as multimodal visual language pre-trained models can be used to describe the content of input images with text.
[0106] Step 406: Fusion and noise reduction processing.
[0107] The input image and the mask image are fused and noise-added to obtain a fused noisy image. The fused noisy image is then denoised and color-corrected.
[0108] See Figure 4c , Figure 4c This diagram illustrates the denoising and color correction process in an image augmentation method according to an embodiment of this specification; specifically, X T That is, to fuse noisy images, perform stepwise denoising on the fused noisy images, at each X... t -X t-1 Regarding the denoising step size, by adjusting X t Denoising to obtain X t-1Specifically, at each denoising step, denoising and color adjustment are performed by the denoising module (i.e., the feature denoising module in the above embodiment) and the latent space color adjustment module (i.e., the feature color adjustment module in the above embodiment). The features output by the feature color adjustment module in the last denoising step (denoising X1 to X0) are determined as the target image features, so that a clear and expanded target image can be obtained by using the target image features.
[0109] Step 408: Guided noise reduction.
[0110] See Figure 4d , Figure 4d This diagram illustrates a guided denoising process in an image augmentation method provided in one embodiment of this specification.
[0111] The denoising module consists of 11 layers of denoising network. When injecting style guidance information into the denoising module, the text description guidance information obtained by the Image Content Description Module (Image Tagger) is sent to the denoising network layers of each layer of the denoising module. The text cross-attention module performs cross-attention feature extraction with the features of each layer in the denoising module. The result is used as input, concatenated with the features of the denoising module at that layer, and used as input for the next layer.
[0112] The image feature guidance information obtained by the image encoder is fed into the 7th layer of the denoising network used for style generation. Similarly, the image cross-attention module performs cross-attention feature extraction with the 7th layer features in the denoising module. The result is concatenated with the features of the denoising module at that layer as input and used as input for the next layer.
[0113] Step 410: Feature color adjustment.
[0114] After the denoising process in step 408, the latent space color adjustment module is used to adjust the color of the denoised image features output by the denoising module. The latent space denoised image features include 4 channels, where the first channel is the cyan / red color channel, which affects the generated image to be more cyan or more red, and the second channel is the yellow / purple color channel, which affects the generated image to be more yellow or more purple. Therefore, the first and second channels of the denoised image features are color adjusted.
[0115] By color-correcting the features, the final target image features are obtained, and the output image is then obtained using these target image features.
[0116] For specific implementation details, please refer to the above embodiments, which will not be repeated here.
[0117] The image expansion method provided in this specification can generatively complete the area to be expanded of the input image based on the mask of the area to be expanded, resulting in an expanded image with consistent tone, style, and no sense of incongruity. Specifically, style guidance information is used to control the generated content of the expanded area, which effectively improves the problems of inconsistent style and disharmonious content in the expanded area. The use of the latent space color adjustment module to control the tone of the expanded area to be consistent with the original image (input image) can effectively avoid the problems of poor integration and inconsistent tone with the original image.
[0118] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 5 A schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification is shown. Figure 5 As shown, the device includes: The generation module 502 is configured to input an initial image into an image processing model and use the image processing model to generate a mask image corresponding to the initial image, wherein the size of the mask image is larger than the size of the initial image; The determining module 504 is configured to extract image generation information from the initial image and determine multiple image generation guidance information; The fusion module 506 is configured to fuse and add noise to the mask image and the initial image to obtain a fused noise image; The acquisition module 508 is configured to generate guiding information based on the plurality of images to denoise the fused noisy image and obtain the target image output by the image processing model.
[0119] Optionally, the obtaining module 508 is further configured to: Using the feature denoising module, the fused noisy image is denoised based on the guidance information generated from the multiple images to obtain denoised image features, wherein the denoised image features include channel image features of multiple channels; Using the feature color adjustment module, the channel image features of the target channel among the multiple channels are color adjusted to obtain the target image features; The target image is obtained based on the target image features.
[0120] Optionally, the obtaining module 508 is further configured to: Using the feature-based color adjustment module, the vector mean of the channel image features of the target channel is calculated to obtain the mean channel image features, and the color adjustment amplitude value is calculated based on the mean channel image features; The target image features are obtained based on the channel image features of the target channel and the color adjustment amplitude value.
[0121] Optionally, the generation module 502 is further configured to: The mask image is generated using the mask generation unit to obtain the mask image corresponding to the initial image.
[0122] Optionally, the determining module 504 is further configured to: Using the extended guidance unit, image generation information is extracted from the initial image to obtain the text description guidance information and the image feature guidance information of the initial image.
[0123] Optionally, the determining module 504 is further configured to: The multimodal processing module is used to perform text description processing on the initial image to obtain the text description guidance information; The initial image is processed using the image encoding module to obtain the image feature guidance information.
[0124] Optionally, the fusion module 506 is further configured to: The image generation unit is used to fuse the mask image and the initial image to obtain a fused image, and then the fused image is subjected to noise processing to obtain a fused noise image.
[0125] Optionally, the obtaining module 508 is further configured to: The feature denoising module is used to extract features from the fused noise image to obtain the noise image features of the fused noise image; Using the feature denoising module, through a cross-attention mechanism, the noisy image features are denoised based on the text description guidance information and the image feature guidance information to obtain denoised image features; The target image is obtained based on the denoised image features.
[0126] Optionally, the obtaining module 508 is further configured to: The text description guidance information is cross-attention processed with the noise image features in each of the multiple denoising network layers, and the image feature guidance information is cross-attention processed with the noise image features in the target denoising network layer of the multiple denoising network layers, so as to denoise the noise image features and obtain the denoised image features output by the last denoising network layer in the multiple denoising network layers.
[0127] The apparatus further includes a response module configured to, in response to an image processing instruction, determine the initial image carried in the image processing instruction and a target extended size for the initial image.
[0128] Optionally, the generation module 502 is further configured to: The initial image and the target extended size are input into the image processing model. The image processing model is used to extend the initial image by masking it according to the target extended size, thereby generating a mask image corresponding to the initial image.
[0129] The device further includes an image and text determination module configured to determine the initial image and image processing prompt text.
[0130] Optionally, the generation module 502 is further configured to: The initial image and the image processing prompt text are input into the image processing model. The image processing model is used to expand the initial image by masking it according to the target expansion size contained in the image processing prompt text, thereby generating a mask image corresponding to the initial image.
[0131] The apparatus further includes: a receiving module configured to receive the initial image sent by a client, wherein the initial image is sent by the client in response to an interactive operation command from a user interface.
[0132] The apparatus further includes a sending module configured to return the target image to the client for displaying the target image on the client's user interface.
[0133] An image processing apparatus provided in one embodiment of this specification can generate a mask image corresponding to an initial image through an image processing model. The size of the mask image is larger than that of the initial image, thereby expanding the initial image based on the mask image. By determining multiple image generation guidance information of the initial image, information about the style, content, features, and other aspects of the initial image is obtained. The mask image and the initial image are fused and denoised to obtain a fused noise image. By denoising the fused noise image, a clear target image expanded from the initial image is obtained. In the denoising process, when the fused noise image is denoised using multiple image generation guidance information, the content of the expanded area in the generated target image can be guided to be consistent with the style and content of the initial image, avoiding problems such as inconsistent style, poor fusion, and disharmony between the expanded area and the initial image in the target image, thus improving the quality of the generative image expansion result.
[0134] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.
[0135] Corresponding to the above method embodiments, this specification also provides embodiments of image expansion devices. Figure 6 A schematic diagram of an image expansion device according to one embodiment of this specification is shown. Figure 6 As shown, the device includes: The generation module 602 is configured to input the image to be expanded into an image processing model, and use the image processing model to generate an expanded mask image corresponding to the image to be expanded, wherein the size of the expanded mask image is larger than the size of the image to be expanded; The determining module 604 is configured to extract image generation information from the image to be expanded and determine multiple image generation guidance information. The fusion module 606 is configured to fuse and add noise to the extended mask image and the image to be expanded to obtain a fused noise image. The acquisition module 608 is configured to generate guiding information based on the plurality of images to denoise the fused noisy image and obtain the extended image output by the image processing model.
[0136] Optionally, the obtaining module 608 is further configured to: Using the feature denoising module, the fused noisy image is denoised based on the guidance information generated from the multiple images to obtain denoised image features, wherein the denoised image features include channel image features of multiple channels; Using the feature color adjustment module, the channel image features of the target channel in the multiple channels are color adjusted to obtain extended image features; The extended image is obtained based on the extended image features.
[0137] The above is a schematic scheme of an image expansion device according to this embodiment. It should be noted that the technical solution of this image expansion device and the technical solution of the image expansion method described above belong to the same concept. For details not described in detail in the technical solution of the image expansion device, please refer to the description of the technical solution of the image expansion method described above.
[0138] Figure 7A structural block diagram of a computing device 700 according to one embodiment of this specification is shown. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and a database 750 is used to store data.
[0139] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0140] In one embodiment of this specification, the above-described components of the computing device 700 and Figure 7 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 7 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0141] The computing device 700 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 700 can also be a mobile or stationary server.
[0142] The processor 720 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-mentioned image processing method and image expansion method.
[0143] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the computing device embodiments are basically similar to the image processing method and image expansion method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the image processing method and image expansion method embodiments.
[0144] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described image processing method and image expansion method.
[0145] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the image processing method and image expansion method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the image processing method and image expansion method embodiments.
[0146] An embodiment of this specification also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the above-described image processing method and image expansion method.
[0147] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the technical solutions of the image processing method and the image expansion method described above. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solutions of the image processing method and the image expansion method described above.
[0148] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0149] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0150] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0151] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0152] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. An image processing method, comprising: An initial image is input into an image processing model, and the image processing model is used to generate a mask image corresponding to the initial image, wherein the size of the mask image is larger than the size of the initial image; Image generation information is extracted from the initial image to determine multiple image generation guidance information; The mask image and the initial image are fused and noise-added to obtain a fused noise image; Based on the guidance information generated from the multiple images, the fused noisy image is denoised to obtain the target image output by the image processing model.
2. The image processing method according to claim 1, wherein the image processing model includes an image generation unit, and the image generation unit includes a feature denoising module and a feature color adjustment module; The step of generating guiding information based on the multiple images to denoise the fused noisy image and obtaining the target image output by the image processing model includes: Using the feature denoising module, the fused noisy image is denoised based on the guidance information generated from the multiple images to obtain denoised image features, wherein the denoised image features include channel image features of multiple channels; Using the feature color adjustment module, the channel image features of the target channel among the multiple channels are color adjusted to obtain the target image features; The target image is obtained based on the target image features.
3. The image processing method according to claim 2, wherein the step of using the feature color adjustment module to perform color adjustment processing on the channel image features of the target channel among the plurality of channels to obtain the target image features includes: Using the feature-based color adjustment module, the vector mean of the channel image features of the target channel is calculated to obtain the mean channel image features, and the color adjustment amplitude value is calculated based on the mean channel image features; The target image features are obtained based on the channel image features of the target channel and the color adjustment amplitude value.
4. The image processing method according to claim 1, wherein the image processing model includes a mask generation unit; The step of generating a mask image corresponding to the initial image using the image processing model includes: The mask image is generated using the mask generation unit to obtain the mask image corresponding to the initial image.
5. The image processing method according to claim 1, wherein the image processing model includes an extended guidance unit, and the image generation guidance information includes text description guidance information and image feature guidance information; The step of extracting image generation information from the initial image and determining multiple image generation guidance information includes: Using the extended guidance unit, image generation information is extracted from the initial image to obtain the text description guidance information and the image feature guidance information of the initial image.
6. The image processing method according to claim 5, wherein the extended guidance unit comprises a multimodal processing module and an image encoding module; The step of using the extended guidance unit to extract image generation information from the initial image to obtain the text description guidance information and the image feature guidance information of the initial image includes: The multimodal processing module is used to perform text description processing on the initial image to obtain the text description guidance information; The initial image is processed using the image encoding module to obtain the image feature guidance information.
7. The image processing method according to claim 1, wherein the image processing model includes an image generation unit; The process of fusing and adding noise to the mask image and the initial image to obtain a fused noise image includes: The image generation unit is used to fuse the mask image and the initial image to obtain a fused image, and then the fused image is subjected to noise processing to obtain a fused noise image.
8. The image processing method according to claim 1, wherein the image processing model includes an image generation unit, the image generation unit includes a feature denoising module, and the plurality of image generation guidance information includes text description guidance information and image feature guidance information; The step of generating guiding information based on the multiple images to denoise the fused noisy image and obtain the target image includes: The feature denoising module is used to extract features from the fused noise image to obtain the noise image features of the fused noise image; Using the feature denoising module, through a cross-attention mechanism, the noisy image features are denoised based on the text description guidance information and the image feature guidance information to obtain denoised image features; The target image is obtained by using the feature denoising module based on the denoised image features.
9. The image processing method according to claim 8, wherein the feature denoising module comprises a plurality of denoising network layers; The feature denoising module, through a cross-attention mechanism, denoises the noisy image features based on the text description guidance information and the image feature guidance information to obtain denoised image features, including: The text description guidance information is cross-attention processed with the noise image features in each of the multiple denoising network layers, and the image feature guidance information is cross-attention processed with the noise image features in the target denoising network layer of the multiple denoising network layers, so as to denoise the noise image features and obtain the denoised image features output by the last denoising network layer in the multiple denoising network layers.
10. The image processing method according to claim 1, further comprising, before inputting the initial image into the image processing model and generating the mask image corresponding to the initial image using the image processing model: In response to an image processing instruction, the initial image carried in the image processing instruction and the target expansion size for the initial image are determined; The step of inputting the initial image into the image processing model and using the image processing model to generate a mask image corresponding to the initial image includes: The initial image and the target extended size are input into the image processing model. The image processing model is used to extend the initial image by masking it according to the target extended size, thereby generating a mask image corresponding to the initial image.
11. The image processing method according to claim 1, further comprising, before inputting the initial image into the image processing model and generating the mask image corresponding to the initial image using the image processing model: The initial image and image processing prompt text are determined; The step of inputting the initial image into the image processing model and using the image processing model to generate a mask image corresponding to the initial image includes: The initial image and the image processing prompt text are input into the image processing model. The image processing model is used to expand the initial image by masking it according to the target expansion size contained in the image processing prompt text, thereby generating a mask image corresponding to the initial image.
12. The image processing method according to claim 1, before inputting the initial image into the image processing model and generating the mask image corresponding to the initial image using the image processing model, further comprising: Receive the initial image sent by the client, wherein the initial image is sent by the client in response to an interactive operation command from the user interface; After generating guiding information based on the multiple images to denoise the fused noisy image and obtain the target image output by the image processing model, the method further includes: The target image is returned to the client for display on the client's user interface.
13. An image expansion method, comprising: The image to be expanded is input into the image processing model, and the image processing model is used to generate an expanded mask image corresponding to the image to be expanded, wherein the size of the expanded mask image is larger than the size of the image to be expanded; Image generation information is extracted from the image to be expanded to determine multiple image generation guidance information; The extended mask image and the image to be extended are fused and noise-added to obtain a fused noise image; Based on the guidance information generated from the multiple images, the fused noisy image is denoised to obtain the extended image output by the image processing model.
14. The image extension method according to claim 13, wherein the image processing model includes an image generation unit, and the image generation unit includes a feature denoising module and a feature color adjustment module; The step of generating guiding information based on the multiple images to denoise the fused noisy image and obtain the extended image output by the image processing model includes: Using the feature denoising module, the fused noisy image is denoised based on the guidance information generated from the multiple images to obtain denoised image features, wherein the denoised image features include channel image features of multiple channels; Using the feature color adjustment module, the channel image features of the target channel in the multiple channels are color adjusted to obtain extended image features; The extended image is obtained based on the extended image features.
15. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1-14.
16. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1-14.
17. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1-14.