Image processing method and image outpainting method

By generating a mask image of the initial image and performing fusion and denoising processing, the problem of inconsistent styles in the expanded regions in the expanded image model is solved, thereby improving the quality of the expanded image and the user experience.

WO2026066121A1PCT designated stage Publication Date: 2026-04-02ALIBABA (CHINA) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing image expansion models based on diffusion models often result in inconsistent content and poor integration between the expanded area and the original image, leading to a reduced user experience.

Method used

By generating a mask image corresponding to the initial image, image generation information is extracted, fused, and noise is added. Multiple image generation guidance information is used to denoise the fused noisy image, generating an extended image with a consistent style.

Benefits of technology

It improves the quality of generative image expansion results, ensuring that the content of the expanded area matches the style of the initial image, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025094423_02042026_PF_FP_ABST
    Figure CN2025094423_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an image processing method and an image outpainting method. The image processing method comprises: inputting an initial image into an image processing model to generate, by means of the image processing model, a mask image corresponding to the initial image, wherein the size of the mask image is larger than that of the initial image; performing image generation information extraction on the initial image to determine a plurality of pieces of image generation guidance information; performing fusion and noise addition processing on the mask image and the initial image to obtain a fused noise image; and denoising the fused noise image on the basis of the plurality of pieces of image generation guidance information to obtain a target image outputted by the image processing model. In the denoising process, when the plurality of pieces of image generation guidance information are used to denoise the fused noise image, picture content of an outpainted area in the generated target image can be guided to be consistent with the style of the initial image and to conform to picture content of the initial image.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method and image expansion method TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of computer, and particularly relate to an image processing method and an image expansion method. BACKGROUND

[0002] With the popularity of mobile phones, cameras and other devices, a large number of pictures are generated, and these pictures can be edited again. Picture frame expansion is an important function in picture editing, and picture frame expansion can meet the needs of users in content creation and other aspects.

[0003] However, when implementing picture frame expansion by using an existing diffusion model-based expansion model, the picture expansion content of the expansion area is inconsistent with the style of the original picture, the expansion area is not well integrated with the original picture, and there is a sense of discomfort, which reduces the user experience. SUMMARY

[0004] Therefore, embodiments of the present disclosure provide an image processing method and an image expansion method. One or more embodiments of the present disclosure also provide an image processing device, an image expansion device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects of the prior art, such as the inconsistency between the picture expansion content and the style of the original picture, the poor integration between the expansion area and the original picture, and the sense of discomfort when implementing picture frame expansion.

[0005] According to a first aspect of embodiments of the present disclosure, an image processing method is provided, including:

[0006] inputting an initial image into an image processing model, and generating a mask image corresponding to the initial image by using the image processing model, wherein the size of the mask image is greater than the size of the initial image;

[0007] extracting image generation information from the initial image to determine a plurality of image generation guide information;

[0008] fusing and adding noise to the mask image and the initial image to obtain a fused noise image;

[0009] de-noising the fused noise image according to the plurality of image generation guide information to obtain a target image output by the image processing model.

[0010] According to a second aspect of embodiments of the present disclosure, an image processing device is provided, including:

[0011] a generation module configured to input an initial image into an image processing model, and generate a mask image corresponding to the initial image by using the image processing model, wherein the size of the mask image is greater than the size of the initial image;

[0012] The determining module is configured to perform image generation information extraction on the initial image, and determine a plurality of image generation guide information;

[0013] The fusion module is configured to perform fusion and noise adding processing on the mask image and the initial image, and obtain a fusion noise image;

[0014] The obtaining module is configured to perform denoising on the fusion noise image according to the plurality of image generation guide information, and obtain a target image output by the image processing model.

[0015] According to a third aspect of an embodiment of the present disclosure, an image expansion method is provided, comprising:

[0016] inputting an image to be expanded into an image processing model, and generating an expansion mask image corresponding to the image to be expanded by using the image processing model, wherein the size of the expansion mask image is greater than the size of the image to be expanded;

[0017] performing image generation information extraction on the image to be expanded, and determining a plurality of image generation guide information;

[0018] performing fusion and noise adding processing on the expansion mask image and the image to be expanded, and obtaining a fusion noise image;

[0019] performing denoising on the fusion noise image according to the plurality of image generation guide information, and obtaining an expansion image output by the image processing model.

[0020] According to a fourth aspect of an embodiment of the present disclosure, an image expansion device is provided, comprising:

[0021] The generating module is configured to input an image to be expanded into an image processing model, and generate an expansion mask image corresponding to the image to be expanded by using the image processing model, wherein the size of the expansion mask image is greater than the size of the image to be expanded;

[0022] The determining module is configured to perform image generation information extraction on the image to be expanded, and determine a plurality of image generation guide information;

[0023] The fusion module is configured to perform fusion and noise adding processing on the expansion mask image and the image to be expanded, and obtain a fusion noise image;

[0024] The obtaining module is configured to perform denoising on the fusion noise image according to the plurality of image generation guide information, and obtain an expansion image output by the image processing model.

[0025] According to a fifth aspect of an embodiment of the present disclosure, a computing device is provided, comprising:

[0026] a memory and a processor;

[0027] The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, so as to implement the steps of the image processing method and the image expansion method.

[0028] According to a sixth aspect of an embodiment of the present disclosure, a computer readable storage medium is provided, which stores computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the steps of the image processing method and the image expansion method.

[0029] According to a seventh aspect of an embodiment of the present disclosure, a computer program product is provided, which includes computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the steps of the image processing method and the image expansion method.

[0030] The image processing method provided by one embodiment of the present disclosure can generate a mask image corresponding to the initial image through an image processing model, the size of the mask image is greater than the size of the initial image, so that the picture expansion of the initial image is realized on the basis of the mask image, the mask image and the initial image are fused and noise processing is performed, to obtain a fused noise image, and the fused noise image is denoised to obtain a target image which is clear and expanded from the initial image; the guide information of multiple images is determined, to obtain information about the style, picture content, features, and the like of the initial image, so that the fused noise image is denoised by using the guide information of multiple images in the denoising process; through the guidance of the guide information of multiple images, the picture content of the expanded area in the generated target image is consistent with the style of the initial image and consistent with the picture content of the initial image, and the problems of inconsistent style, poor fusion, and discomfort, and the like between the expanded area in the target image and the initial image are avoided, and the result quality of the generative expansion is improved. BRIEF DESCRIPTION OF DRAWINGS

[0031] FIG. 1 is a scene schematic diagram of an image processing method provided by one embodiment of the present disclosure;

[0032] FIG. 2 is a flowchart of an image processing method provided by one embodiment of the present disclosure;

[0033] FIG. 3 is a flowchart of an image expansion method provided by one embodiment of the present disclosure;

[0034] FIG. 4a is a flowchart of a processing process of an image expansion method provided by one embodiment of the present disclosure;

[0035] FIG. 4b is an architecture schematic diagram of an image expansion system provided by one embodiment of the present disclosure;

[0036] FIG. 4c is a process schematic diagram of denoising and color adjustment in an image expansion method provided by one embodiment of the present disclosure;

[0037] FIG. 4d is a schematic diagram of a process of guided denoising in an image expansion method according to an embodiment of the present disclosure;

[0038] FIG. 5 is a schematic diagram of a structure of an image processing apparatus according to an embodiment of the present disclosure;

[0039] FIG. 6 is a schematic diagram of a structure of an image expansion apparatus according to an embodiment of the present disclosure;

[0040] FIG. 7 is a structural block diagram of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0041] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, the present disclosure can be practiced without the specific details, which are not described in the present disclosure, and it is understood that the scope of the present disclosure is not limited to the details of the embodiments described herein. In other instances, well-known methods associated with computing, software development, and / or data analytics have not been described in detail in order to avoid unnecessarily obscuring aspects of the present disclosure.

[0042] The terminology used in one or more embodiments of the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present disclosure. As used in one or more embodiments of the present disclosure and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be understood that the terms "and / or," "at least one of," and "one or more of" as used herein refer to and encompass any one of the listed items, any combination of the listed items, and / or all of the listed items, as appropriate. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0043] It is to be understood that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It is to be understood that the terms "including," "comprising," "consisting" and "consisting essentially of" to the present disclosure are open-ended terms that do not preclude the addition of one or more components, integers, aspects, and / or steps to the compositions and / or methods described herein. Accordingly, the terms "comprising," "comprises" and "comprised of" as used herein are synonymous with "includes," "includes," and "including," respectively.

[0044] In addition, it should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0045] First, the noun terms related to one or more embodiments of the present disclosure are explained.

[0046] Generative diffusion map: refers to the use of a diffusion model to expand outwardly from an image frame, supplementing the content of the outwardly expanded area.

[0047] Diffusion model: a class of latent variable model, a Markov chain trained by variational estimation; the diffusion model learns the latent structure of the data set by modeling the diffusion mode of the data points in the latent space. After training, random noise can be transmitted into the diffusion model to generate data by learning the denoising process.

[0048] In the present disclosure, an image processing method and an image expansion method are provided. The present disclosure also relates to an image processing device, an image expansion device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail in the following embodiments.

[0049] Referring to FIG. 1, FIG. 1 shows a scene diagram of an image processing method according to an embodiment of the present disclosure.

[0050] Specifically, the image processing method is implemented by an end-side device 102 and a server 104. The end-side device 102 is configured to send an initial image to the server 104, for example, the initial image is a close-up image of a person. The server 104 is configured to train an image processing model, and in a case where the server 104 receives the initial image sent by the end-side device 102, the server 104 is configured to input the initial image into the image processing model, and generate a mask image corresponding to the initial image by using the image processing model, wherein the size of the mask image is greater than the size of the initial image. The image processing method is configured to extract image generation information from the initial image, and determine a plurality of image generation guide information. The image processing method is configured to fuse and add noise to the mask image and the initial image to obtain a fused noise image. The image processing method is configured to denoise the fused noise image according to the plurality of image generation guide information to obtain a target image output by the image processing model, and the target image can be an image in which the background of the close-up image is expanded, and the target image is returned to the end-side device 102.

[0051] The end-side device 102 can include a browser, an application (APP), or a web application such as a Hyper Text Markup Language 5 (H5) application, or a light application (also known as a small program, a lightweight application), or a cloud application, and the like. The end-side device can be developed based on a software development kit (SDK) of a corresponding service provided by the server, such as a real-time communication (RTC) SDK, and the like. The end-side device can be deployed in an electronic device, and needs to be run in dependence on a device or an APP in the device, and the like. The electronic device can have a display screen and support information browsing, and the like, and can be a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, and the like. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, image processing applications, web browser applications, shopping applications, search applications, instant communication tools, mailbox clients, social platform software, and the like.

[0052] The server 104 can be understood as a server providing various services, including a physical server, a cloud server, for example, a server providing communication services for multiple clients, for example, a server for background training supporting a model used on a client, for example, a server processing data sent by a client, and the like. It should be noted that the server 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server 104 can also be a server of a distributed system, or a server combined with a blockchain. The server 104 can also be a cloud server of cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, and the like. The server 104 can also be an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.

[0053] It should be noted that the image processing method provided in the embodiments of the present disclosure can be executed by the server 104. In other embodiments of the present disclosure, the image processing model can be deployed in the end-side device 102, so that the end-side device 102 can also have similar functions as the server 104, thereby executing the image processing method provided in the embodiments of the present disclosure; in other embodiments, the image processing method provided in the embodiments of the present disclosure can also be executed by the end-side device 102 and the server 104 together.

[0054] An embodiment of this disclosure provides an image processing method that generates a mask image corresponding to an initial image through an image processing model. The size of the mask image is larger than that of the initial image, thereby expanding the initial image based on the mask image. By determining multiple image generation guidance information of the initial image, information about the style, content, and features of the initial image is obtained. The mask image and the initial image are fused and denoised to obtain a fused noise image. By denoising the fused noise image, a clear target image expanded from the initial image is obtained. In the denoising process, when the fused noise image is denoised using multiple image generation guidance information, the content of the expanded area in the generated target image can be guided to be consistent with the style and content of the initial image, avoiding problems such as inconsistent style, poor fusion, and disharmony between the expanded area and the initial image in the target image, thus improving the quality of the generative image expansion result.

[0055] Referring to Figure 2, which shows a flowchart of an image processing method provided in an embodiment of the present disclosure, the method specifically includes the following steps.

[0056] Step 202: Input the initial image into the image processing model, and use the image processing model to generate a mask image corresponding to the initial image.

[0057] The size of the mask image is larger than the size of the initial image; the initial image can be understood as the image to be processed, and subsequent image expansion processing is performed on the initial image; the mask image can be understood as the mask image of the expanded area when the initial image is generatively expanded, that is, the expanded image content of the initial image is generated on the mask image.

[0058] Of course, in practical applications, a mask image can be obtained outside of the image processing model. For example, the mask image can be manually drawn using image processing software (such as Photoshop) or programming tools (such as OpenCV), and then input into the image processing model along with the initial image.

[0059] In practical applications, before inputting the initial image into the image processing model and generating the mask image corresponding to the initial image using the image processing model, the following steps are also included:

[0060] Receive an initial image sent by the client, wherein the initial image is sent by the client in response to an interactive operation command from the user interface.

[0061] Specifically, the initial image can be an image to be processed sent by the user through the client. For example, the user can perform interactive operations such as uploading an image and confirming the sending on the user interface of the client, so that the client can send the initial image uploaded by the user to the server based on the interactive operation instructions of the user interface.

[0062] Of course, multiple images can also be displayed on the user interaction interface of the client, and the user can select one of the multiple images as the initial image for confirmation and sending, which is not limited herein.

[0063] The image processing method provided by the embodiments of the present disclosure can enable the user to conveniently and quickly determine the initial image and send it through the user interaction interface provided by the client, thereby improving the user's interactive experience.

[0064] In one or more embodiments of the present disclosure, the image processing model includes a mask generation unit, so that the initial image can be input into the mask generation unit to obtain the mask image corresponding to the initial image by using the mask generation unit. The specific implementation is as follows:

[0065] The image processing model is used to generate the mask image corresponding to the initial image, including:

[0066] The mask image is generated by using the mask generation unit to obtain the mask image corresponding to the initial image.

[0067] The mask generation unit is used to generate the mask image of the input image, for example, the mask generation unit can be a machine learning model, and through model training, the size of the generated mask image can be greater than that of the initial image; the size of the image can be represented by the pixel value in the width and height directions, or by the length (such as in centimeters) in the width and height directions.

[0068] In the case where the size of the image is represented by the pixel value of the image in the width and height directions, when the mask image is generated by using the mask generation unit, the pixel value in the width and height directions of the generated mask image can be more than that of the initial image, realizing the expansion of the initial image in the width and height directions, for example, the size of the initial image is 400x300 pixels, and the size of the mask image is 800x600 pixels.

[0069] Of course, in actual applications, the generated mask image can also be used to expand the initial image in the width direction or the height direction, for example, the size of the initial image is 400x300, and the size of the mask image is 800x300, realizing the expansion of the initial image in the width direction, or the size of the initial image is 400x300, and the size of the mask image is 400x600, realizing the expansion of the initial image in the height direction.

[0070] The image processing method provided by the embodiment of the present disclosure can quickly and efficiently generate a mask image by a mask generation unit of an image processing model, generate extended picture content on the mask image, and realize extension of the initial image by using the mask image, thereby providing a data basis for subsequent extension of the initial image.

[0071] In one or more embodiments of the present disclosure, when there is no specific requirement for the size of the mask image, the initial image can be directly input into the mask generation unit for mask image generation, and the size of the mask image is random. In actual applications, the size of the mask image can be specified by using the target extension size to meet individualized requirements. The specific implementation is described as follows.

[0072] Before the initial image is input into the image processing model and the mask image corresponding to the initial image is generated by using the image processing model, the method further includes the following steps.

[0073] In response to the image processing instruction, the initial image carried in the image processing instruction and the target extension size for the initial image are determined.

[0074] The initial image is input into the image processing model, and the mask image corresponding to the initial image is generated by using the image processing model. The method includes the following steps.

[0075] The initial image and the target extension size are input into the image processing model, and the mask image corresponding to the initial image is generated by using the image processing model and performing extension mask on the initial image according to the target extension size.

[0076] In actual applications, before the initial image is input into the image processing model and the mask image corresponding to the initial image is generated by using the image processing model, the method further includes the following steps.

[0077] The initial image and the image processing prompt text are determined.

[0078] The initial image is input into the image processing model, and the mask image corresponding to the initial image is generated by using the image processing model. The method includes the following steps.

[0079] The initial image and the image processing prompt text are input into the image processing model, and the mask image corresponding to the initial image is generated by using the image processing model and performing extension mask on the initial image according to the target extension size contained in the image processing prompt text.

[0080] The target extension size can be understood as a specific size of the mask image. For example, if the image processing prompt text is "expand to generate an image with a size of 800x600", the target extension size is 800x600 pixels. The target extension size can also be understood as an extension size based on the initial image. For example, if the image processing prompt text is "extend the width by 200 pixels based on the input image", the target extension size is 200 pixels in the width direction.

[0081] The target extension size is determined through the image processing instruction or the image processing prompt text. Thus, a mask image with a specific size is generated by extending the mask based on the target extension size of the initial image.

[0082] In actual application, the user interaction interface of the client can have an upload control for uploading an image and a selection control for selecting a target extension size. The client sends an image processing instruction to the server in response to the user's interaction operation instruction on the user interaction interface and the upload control and the selection control. The image processing instruction carries the initial image uploaded by the user and the target extension size of the initial image. Alternatively, the user interaction interface of the client has an upload control for uploading an image and an input control for inputting text. The user can input an image processing prompt text through the interaction operation with the input control. The image processing prompt text includes the target extension size.

[0083] The image processing method provided by the embodiments of the present disclosure can make the image processing model generate a mask image with a specific size based on the initial image and the target extension size through diversified ways such as carrying the target extension size in the image processing instruction and including the target extension size in the image processing prompt text, thereby meeting the personalized needs of users.

[0084] Step 204: Extract image generation information from the initial image to determine a plurality of image generation guide information.

[0085] The image generation guide information can be understood as information for guiding image generation. For example, the image generation guide information includes image style information, image content information, image feature information, and the like. Through the image generation guide information, the style and content of the generated image can be guided in the image generation process.

[0086] Specifically, by extracting image generation information from the initial image, the image style information, image content information, image feature information, and the like of the initial image can be obtained. Thus, in the subsequent image generation process, these information can be used as a guide to make the generated image consistent with the style and content of the initial image, thereby improving the image quality of the generated image.

[0087] In one or more embodiments of the present disclosure, the image processing model comprises an extended guidance unit, the image generation guidance information comprises text description guidance information and image feature guidance information; in the extended guidance unit, the text description guidance information and the image feature guidance information of the initial image are obtained. The specific implementation is as follows:

[0088] The image generation information extraction is performed on the initial image, and a plurality of image generation guidance information is determined, comprising:

[0089] The image generation information extraction is performed on the initial image by using the extended guidance unit, and the text description guidance information and the image feature guidance information of the initial image are obtained.

[0090] The text description guidance information can comprise image style information and image content information of the initial image, and the image style information and the image content information can be determined by using the image text description of the initial image. The image feature guidance information can be understood as the feature information of the image content of the initial image in a hidden space. The hidden space is a multi-dimensional vector space learned by the model. In this space, the complex features of the original data are represented as a set of more concise and usually continuously distributed variables, which capture the key characteristics of the input data and often have certain semantic significance.

[0091] Specifically, the image generation information extraction is performed on the initial image by using the extended guidance unit, and the text description guidance information comprising the image style information and the image content information of the initial image and the feature information of the image content of the initial image in the hidden space are obtained. The text description guidance information and the image feature guidance information can be data represented by using a feature vector through feature extraction.

[0092] In actual application, the extended guidance unit comprises a multi-modal processing module and an image coding module; the text description guidance information of the initial image is obtained by using the multi-modal processing module, and the image feature guidance information of the initial image is obtained by using the image coding module. The specific implementation is as follows:

[0093] The image generation information extraction is performed on the initial image by using the extended guidance unit, and the text description guidance information and the image feature guidance information of the initial image are obtained, comprising:

[0094] The text description processing is performed on the initial image by using the multi-modal processing module, and the text description guidance information is obtained.

[0095] The image coding processing is performed on the initial image by using the image coding module, and the image feature guidance information is obtained.

[0096] The multi-modal processing module can be understood as an image tagger configured to add descriptive labels to the content in the image. In practical applications, the multi-modal processing module can be implemented using a multi-modal large model, i.e., by inputting the initial image into the multi-modal large model, using the multi-modal large model to generate a text description of the initial image, and determining the text features of the text description (i.e., the text description guide information).

[0097] The image encoding module can be understood as a module implemented using an image encoder. By inputting the initial image into the image encoder, the image encoder is used to perform image encoding processing on the initial image to obtain the image features of the initial image (i.e., the image feature guide information).

[0098] The image processing method provided by the embodiments of the present disclosure can not only obtain the text description guide information corresponding to the initial image through the multi-modal processing module, but also obtain the image feature guide information of the initial image through the image encoding module. In the case of using the multi-modal guide information to guide the generation of the image, the quality of the generated target image can be improved.

[0099] Step 206: fuse and add noise to the mask image and the initial image to obtain a fused noise image.

[0100] Specifically, the image processing model includes an image generation unit. First, the mask image and the initial image are fused, and then the fused image is added with noise to obtain a fused noise image. The specific implementation is as follows:

[0101] Fusing and adding noise to the mask image and the initial image to obtain a fused noise image includes:

[0102] Fusing and adding noise to the mask image and the initial image to obtain a fused noise image includes:

[0103] The image generation unit can be understood as a diffusion model, which is used to fuse and add noise to the mask image and the initial image.

[0104] Specifically, the fusion processing of the mask image and the initial image can be understood as superimposing the mask image and the initial image to obtain a fused image. For example, in the case of expanding the initial image by a factor of two in the width direction and the height direction, the initial image can be placed in the middle region of the mask image. On this basis, other regions of the mask image can be filled with the edge color of the initial image to obtain a fused image. The other regions of the mask image are the regions to be expanded corresponding to the initial image.

[0105] Introducing random noise in the fusion image, such as adding Gaussian noise, to obtain a fusion noise image, so that in the case of subsequent denoising of the fusion noise image, the picture content consistent with the initial image is generated in other areas of the mask image, and the target image expanded from the initial image is obtained.

[0106] The image processing method provided by the embodiments of the present disclosure obtains the fusion noise image by fusing and adding noise to the mask image and the initial image, and provides a data basis for generating the target image in the subsequent denoising, that is, the target image is obtained by denoising the fusion noise image.

[0107] Step 208: denoising the fusion noise image according to the plurality of image generation guide information to obtain the target image output by the image processing model.

[0108] The target image can be understood as an expanded image expanded from the initial image, and the size of the target image is consistent with the size of the mask image.

[0109] Specifically, the plurality of image generation guide information is injected into the denoising process of the fusion noise image, so that the target image output by the image processing model is obtained under the guidance of the plurality of image generation guide information, or in reference to the plurality of image generation guide information.

[0110] In one or more embodiments of the present disclosure, in the image generation unit, the fusion noise image is denoised by using the plurality of image generation guide information to obtain the target image expanded from the initial image. In actual application, the image generation unit includes a feature denoising module, and the plurality of image generation guide information includes text description guide information and image feature guide information. The specific implementation is as follows:

[0111] Denoising the fusion noise image according to the plurality of image generation guide information to obtain the target image, comprising:

[0112] Using the feature denoising module, the noise image features of the fusion noise image are obtained by feature extraction on the fusion noise image;

[0113] Using the feature denoising module, the noise image features are denoised according to the text description guide information and the image feature guide information by a cross-attention mechanism to obtain denoised image features;

[0114] Using the feature denoising module, the target image is obtained according to the denoised image features.

[0115] In a case where the image generation unit is a diffusion model (the diffusion model can generate a clear image by gradually reducing noise in an image), the feature denoising module can be understood as a denoising module of the diffusion model, and the step-by-step denoising of the fused noise image is implemented by using the denoising module; the denoised image feature can be understood as the information feature remaining in the fused noise image after a series of denoising processes, and the noise component has been removed as much as possible, and the feature is closer to the clear feature of the target image, which is an abstract representation capable of describing the content of the target image.

[0116] Specifically, the feature extraction is performed on the fused noise image to extract the key feature representation from the fused noise image containing noise, to obtain the noise image feature, which will be used for subsequent denoising processing; the cross-attention mechanism is used to combine the text description guide information and the image feature guide information to denoise the noise image feature. Specifically, the text description guide information provides semantic information of the initial image style, content, etc., and the image feature guide information provides information at the image structure level. Therefore, by using the cross-attention mechanism, the text description guide information and the image feature guide information are injected into the denoising process of the diffusion model on the fused noise image, so that the diffusion model can consider both aspects of information, more accurately guide the denoising process, obtain the denoised image feature, and obtain the target image according to the denoised image feature.

[0117] The image processing method provided by the embodiments of the present disclosure injects the text description guide information and the image feature guide information into the denoising process of the diffusion model on the fused noise image through the cross-attention mechanism, to guide the diffusion model to obtain a target image that is more consistent with the style and picture content of the initial image.

[0118] In one or more embodiments of the present disclosure, the feature denoising module includes a plurality of denoising network layers; in a case where the text description guide information and the image feature guide information are injected into the denoising process by using the cross-attention mechanism, the text description guide information and the image feature guide information can be injected into the denoising network layers respectively, to perform cross-attention processing between the text description guide information and the noise image feature of each denoising network layer, to denoise the noise image feature of each denoising network layer, and perform cross-attention processing between the image feature guide information and the noise image feature of the target denoising network layer, to denoise the noise image feature of the target denoising network layer. The specific implementation is as follows:

[0119] The feature denoising module is used to denoise the noise image feature according to the text description guide information and the image feature guide information by using the cross-attention mechanism, to obtain the denoised image feature, including:

[0120] The text description guide information is cross-attention processed with the noise image features in each of the plurality of denoising network layers, and the image feature guide information is cross-attention processed with the noise image features of the target denoising network layer in the plurality of denoising network layers, so as to denoise the noise image features, and obtain the denoised image features output by the last denoising network layer in the plurality of denoising network layers.

[0121] The denoising module of the diffusion model includes a plurality of denoising network layers, and the denoising network layers gradually remove noise by learning features in the data layer by layer. The functions of the denoising network layers are not completely the same, that is, the features learned by each layer in the plurality of denoising network layers are different. The target denoising network layer is used to focus on learning the style information of the image. For example, in the shallow network layer (the initial stage of the diffusion process, the noise level is high), the diffusion model mainly learns the basic structure and content information of the noise fusion image. In the deep network layer (the later stage of the diffusion process, the noise level gradually decreases, and the image gradually becomes clear), more attention is paid to learning the style, texture and detail features of the image.

[0122] For example, the denoising module includes 11 denoising network layers, and the seventh denoising network layer is determined as the target denoising network layer. That is, the text description guide information is cross-attention processed with the noise image features in each denoising network layer, and the image feature guide information is cross-attention processed with the noise image features of the seventh denoising network layer, so as to denoise the noise image features corresponding to the seventh denoising network layer. The noise image features input by the last denoising network layer, that is, the eleventh denoising network layer, are determined as the denoised image features.

[0123] In specific implementation, the specific implementation of denoising the noise fusion image to generate the target image can be regarded as, in the diffusion model, denoising the noise image (denoted as X T ) through a step-by-step reverse diffusion process until recovering into the original clear image (denoted as X0) not interfered by noise. In the process of denoising X T step by step, X t is denoised as X t-1 by the denoising module at each step, X t is denoised as X t-1When the text description guide information is injected into the corresponding denoising network layer of the denoising module, the text description guide information is cross-attention feature extracted with the noise image features of the denoising network layer in the denoising module through the corresponding text cross-attention module, and the result output by the text cross-attention module is merged with the noise image features of the denoising network layer in the denoising module as input, and the merged result is taken as the input of the next level denoising network layer. In the above example, the image feature guide information will be injected into the 7th denoising network layer, that is, the image feature guide information is cross-attention feature extracted with the noise image features of the 7th denoising network layer in the denoising module through the corresponding image cross-attention module, the result output by the image cross-attention module is merged with the noise image features of the 7th denoising network layer in the denoising module as input, and the merged result is taken as the input of the 8th denoising network layer.

[0124] In the case of injecting the text description guide information and the image feature guide information into the denoising network layer of the denoising module in a corresponding manner, the noise image features output by the last level denoising network layer in the plurality of denoising network layers will be determined as the denoised image features.

[0125] The image processing method provided by the embodiments of the present disclosure can ensure the consistency of the style of the generated target image with the style of the initial image by injecting the text description guide information into each denoising network layer and injecting the image feature guide information into the target denoising network layer.

[0126] In one or more embodiments of the present disclosure, the image processing model includes an image generation unit, and the image generation unit includes a feature denoising module and a feature toning module; to ensure that the generated target image is consistent in tone with the initial image, the feature toning module is used to perform toning processing on the denoised image features output by the feature denoising module. The specific implementation is as follows:

[0127] According to the plurality of image generation guide information, the fused noise image is denoised to obtain a target image output by the image processing model, including:

[0128] According to the plurality of image generation guide information, the fused noise image is denoised by using the feature denoising module to obtain denoised image features, wherein the denoised image features include channel image features of a plurality of channels;

[0129] The feature toning module is used to perform toning processing on the channel image features of the target channel in the plurality of channels to obtain target image features.

[0130] According to the target image features, a target image is obtained.

[0131] Each denoised image feature can be regarded as a combination of channel image features corresponding to one or more channels, and these channels contain abstract features of the noise fused image in different dimensions and different angles.

[0132] For example, taking the case of a denoised image feature including 4 channels, the 0th channel is a luminance channel, which affects the luminance of the generated image, and the feature toning module does not adjust it; the 3rd channel is a feature channel, which affects the content of the generated image, and no adjustment is made in the feature toning module; the 1st channel is a cyan / red color channel, which affects the cyan or red of the generated image; specifically, the larger the value, the more cyan, and the smaller the value, the more red; the 2nd channel is a yellow / purple color channel, which affects the yellow or purple of the generated image; specifically, the larger the value, the more yellow, and the smaller the value, the more purple; the feature toning module performs toning processing on the channel image features of the 1st channel and the 2nd channel; that is, in the embodiments of the present disclosure, the 1st channel and the 2nd channel, which affect the color of the generated image, are determined as the target channels.

[0133] In specific implementation, in the case of obtaining the denoised image feature by using the feature denoising module, the denoised image feature is input into the feature toning module, the feature toning module performs toning processing on the channel image features of the 1st channel and the 2nd channel of the denoised image feature, obtains the target image feature, and obtains the target image by using the target image feature. It should be noted that, since the diffusion model is to gradually denoise the fused noise image, each denoising step (X t denoising X t-1 ) will be processed by the feature denoising module and the feature toning module, so the feature output by the feature toning module in the last denoising step (denoising X1 to X0) is determined as the target image feature.

[0134] The image processing method provided by the embodiments of the present disclosure can perform feature toning on the denoised image feature output by the feature denoising module through the feature toning module, so that the generated target image can be consistent with the color tone of the initial image, avoiding the problem that the expansion area in the target image is inconsistent with the color tone of the initial image and does not blend well with the initial image.

[0135] In one or more embodiments of the present disclosure, by calculating the channel image feature of the target channel, a toning amplitude value is obtained, and the channel image feature of the target channel is toned by using the toning amplitude value and the channel image feature of the target channel. The specific implementation is as follows:

[0136] Using the feature toning module, the channel image feature of the target channel in the plurality of channels is toning processed to obtain a target image feature, including:

[0137] Using the feature toning module, the vector mean of the channel image feature of the target channel is calculated to obtain a mean channel image feature, and a toning amplitude value is calculated according to the mean channel image feature;

[0138] According to the channel image feature of the target channel and the toning amplitude value, a target image feature is obtained.

[0139] Specifically, a vector mean of the channel image features of the target channel is calculated to obtain a mean channel image feature representing the overall color level of the target channel, a recoloring amplitude value is calculated based on the mean channel image feature, and the recoloring amplitude value is subtracted from the channel image features of the target channel to obtain the recolored target image features.

[0140] In actual applications, the mean channel image feature is calculated by the following formula: tensor_mean = input_tensor[i].mean()

[0141] Wherein, input_tensor[i] is the channel image feature of the i-th channel in the denoised image feature output by the denoising module in the current round of denoising, and in the embodiments of the present disclosure, i represents the channel label and can be 1 or 2; mean() in input_tensor[i].mean() is a mean function, and tensor_mean is the mean channel image feature in the above embodiments.

[0142] The recoloring amplitude value is calculated by the following formula: recolor_tensor = log(1+|tensor_mean|) / 8

[0143] Wherein, log(1+|tensor_mean|) / 8 ensures that the result will not be negative infinity even when tensor_mean is 0 (because the domain of the logarithmic function is positive); and the logarithmic transformation is usually used to compress the data range, especially when the data has a long-tailed distribution, it can reduce the influence of extreme values, so this logarithmic function will amplify the smaller mean difference in a nonlinear way, while the larger mean difference will have less impact; dividing by 8 is a scaling factor (i.e. the constant to be divided is not limited to what, the constant can be flexibly set according to the actual situation), and the constant division is used to scale the logarithmic value to make it suitable for a specific range, such as normalized to a smaller interval, to facilitate further processing or display, and in the embodiments of the present disclosure, it is used to adjust the intensity of the recoloring amplitude. recolor_tensor is the recoloring amplitude value of the i-th channel calculated.

[0144] The recolored target image features are obtained by the following formula: output_tensor[i] = input_tensor[i]-recolor_tensor

[0145] Wherein, output_tensor[i] is the output vector of the i-th channel after the feature recoloring module.

[0146] The image processing method provided in the embodiments of the present disclosure can accurately perform color adjustment processing on the channel image features of the target channel by calculating the obtained color adjustment amplitude value and the channel image features of the target channel, and obtain target image features.

[0147] In one or more embodiments of the present disclosure, after the denoising of the fused noise image is performed according to the plurality of image generation guide information and the target image output by the image processing model is obtained, the method further includes:

[0148] The target image is returned to the client to display the target image on a user interaction interface of the client.

[0149] The image processing method provided in the embodiments of the present disclosure returns the target image to the client to display the generated target image on a user interaction interface of the client, which is convenient for users to view and improves user experience.

[0150] The image processing method provided in the embodiments of the present disclosure can effectively avoid the problem that the color tone of the extended region of the target image is inconsistent with that of the initial image and the fusion of the initial image is poor, the feature color adjustment module can help the color tone of the extended region to be consistent with that of the initial image by performing color adjustment on the channel image features of the target channel, the extended guide unit can obtain a plurality of image generation guide information of the initial image for the problem that the style of the extended region is inconsistent with that of the initial image, in the case that the plurality of image generation guide information includes information such as the style, picture content, and features of the initial image, the diffusion process is guided by using the information, so that the style of the extended region is consistent with that of the initial image, and the picture content of the extended region is consistent with that of the initial image, thereby avoiding the problem of discomfort.

[0151] Referring to FIG. 3, FIG. 3 shows a flowchart of an image expansion method according to an embodiment of the present disclosure, which specifically includes the following steps.

[0152] Step 302: input the image to be expanded into an image processing model, and generate an expansion mask image corresponding to the image to be expanded by using the image processing model.

[0153] The size of the expansion mask image is greater than that of the image to be expanded; the image to be expanded can be understood as the initial image in the above embodiments; the expansion mask image can be understood as the mask image in the above embodiments; for specific implementation, refer to the above embodiments, which will not be described here.

[0154] Step 304: perform image generation information extraction on the image to be expanded to determine a plurality of image generation guide information.

[0155] Step 306: fuse and add noise to the expansion mask image and the image to be expanded to obtain a fused noise image.

[0156] Step 308: denoising the fused noise image according to the plurality of image generation guide information to obtain an extended image output by the image processing model.

[0157] In a specific implementation, the image processing model includes an image generation unit, and the image generation unit includes a feature denoising module and a feature toning module.

[0158] The denoising the fused noise image according to the plurality of image generation guide information to obtain an extended image output by the image processing model includes:

[0159] The feature denoising module is configured to denoise the fused noise image according to the plurality of image generation guide information to obtain denoised image features, and the denoised image features include channel image features of a plurality of channels.

[0160] The feature toning module is configured to perform toning processing on the channel image features of a target channel in the plurality of channels to obtain extended image features.

[0161] The extended image is obtained according to the extended image features.

[0162] The extended image can be understood as the target image in the above embodiments. For details, refer to the above embodiments, which are not described here again.

[0163] The image extension method provided by the embodiments of the present disclosure uses the feature toning module to tone the channel image features of the target channel, which can help the extended region of the extended image to maintain the same color tone as the initial image, avoiding the problem that the extended region and the initial image do not have the same color tone and do not blend well. The use of the extension guide unit can obtain a plurality of image generation guide information of the initial image. In the case where the plurality of image generation guide information includes information such as the style, the picture content, and the features of the initial image, the diffusion process can be guided using these information, so that the style of the extended region is consistent with the initial image, and the picture content of the extended region is consistent with the picture content of the initial image, thereby avoiding the problem of discomfort.

[0164] Referring to FIG. 4a, FIG. 4a shows a processing process flowchart of an image extension method according to an embodiment of the present disclosure, which specifically includes the following steps.

[0165] Specifically, the image expansion method is applied to an image expansion system. Referring to FIG. 4b, FIG. 4b shows an architecture schematic diagram of an image expansion system according to an embodiment of the present disclosure. An input picture (i.e., the initial image in the above embodiment) and a region mask to be expanded (i.e., the mask image in the above embodiment) are input into an image processing model. The image processing model includes an expansion style guiding unit (i.e., the expansion guiding unit in the above embodiment) and a diffusion completion model with colorization (i.e., the image generation unit in the above embodiment). The text description guiding information and the image feature guiding information of the initial image are obtained by using the expansion style guiding unit. The two guiding information are input into the diffusion completion model with colorization together with the input picture and the region mask to be expanded, so as to obtain an expansion result (an output picture, i.e., the target image in the above embodiment).

[0166] Step 402: determining a mask image.

[0167] The mask image is determined according to the target expansion size, so as to realize mask processing on the input picture by using the mask image, and obtain a mask-processed input picture. For example, when the input picture is expanded by two times in length and width, the mask-processed input picture is a picture in which the input picture is placed in the middle region and the other regions are filled with the edge color of the input picture on the canvas of the two-time input picture.

[0168] Step 404: guiding information extraction.

[0169] The expansion style guiding unit of the image processing model is used to obtain the style guiding information of the input picture, including image description text (including content information and style information) and image features (including feature information of the image content in the hidden space).

[0170] Specifically, the input picture is input into the expansion style guiding unit, and two guiding information are obtained by two guiding information extraction modules. The guiding information extraction module includes an image content description module (Image Tagger) and an image encoding module (Image encoder). The Image Tagger outputs image description text (which can be understood as the text description guiding information in the above embodiment) of the input picture. The Image encoder outputs image features (which can be understood as the image feature guiding information in the above embodiment).

[0171] In actual application, a tagging model such as a multi-modal visual language pre-training model can be used to describe the picture content of the input picture in words.

[0172] Step 406: fusion and noise adding processing.

[0173] Fuse the input picture with the mask image and add noise processing to obtain a fused noise image, and denoise and color adjustment are performed on the fused noise image.

[0174] Referring to FIG. 4c, FIG. 4c shows a process diagram of denoising and color adjustment in an image expansion method provided by an embodiment of the present disclosure. Specifically, X T That is, the fused noise image is gradually denoised, and at each X t -X t-1 denoising step, X t is obtained by denoising X t-1 ; specifically, at each denoising step, the denoising and color adjustment are performed by the denoising module (i.e., the feature denoising module in the above embodiment) and the latent space color adjustment module (i.e., the feature color adjustment module in the above embodiment), the feature output by the feature color adjustment module in the last denoising step (denoising X1 to X0) is determined as the target image feature, so that a clear and expanded target image can be obtained by using the target image feature.

[0175] Step 408: guided denoising.

[0176] Referring to FIG. 4d, FIG. 4d shows a process diagram of guided denoising in an image expansion method provided by an embodiment of the present disclosure.

[0177] The denoising module includes 11 levels of denoising network layers. When injecting style guide information into the denoising module, the text description guide information obtained by the Image Tagger is sent into the denoising network layers of each level of the denoising module, wherein cross-attention feature extraction is performed on the features of each level in the denoising module through the text cross-attention module, and the result is used as input, spliced with the features of the level of the denoising module, and used as the input of the next level.

[0178] The image feature guide information obtained by the Image encoder is sent into the 7th level of the denoising network layer for style generation. Similarly, cross-attention feature extraction is performed on the 7th level features in the denoising module through the image cross-attention module, and the result is used as input and spliced with the features of the level of the denoising module, and used as the input of the next level.

[0179] Step 410: feature color adjustment.

[0180] After the denoising processing in step 408, the denoised image features output by the denoising module are processed by using the hidden space toning module. The denoised image features in the hidden space include four channels, wherein the first channel is a cyan / red color channel, which affects the generated picture to be cyan or red; and the second channel is a yellow / purple color channel, which affects the generated picture to be yellow or purple. Therefore, the first channel and the second channel of the denoised image features are processed.

[0181] By toning the features, the final target image features are obtained, so that the output picture is obtained by using the target image features.

[0182] The specific implementation method is described in the above embodiment, which will not be repeated here.

[0183] The image expansion method provided by the embodiments of the present disclosure can generate the expansion area of the input picture according to the mask of the expansion area to be expanded, so as to obtain an expansion completion result with consistent color tone and style and no sense of strangeness. Specifically, the style guide information is used to control the generated content of the expansion area, which effectively improves the problem of inconsistent style and inharmonious picture content of the expansion area. The hidden space toning module is used to control the expansion area to keep consistent with the color tone of the original picture (input picture), which can effectively avoid the problems of poor fusion and inconsistent color tone with the original picture.

[0184] Corresponding to the method embodiments described above, the present disclosure also provides image processing device embodiments. FIG. 5 shows a structural schematic diagram of an image processing device according to an embodiment of the present disclosure. As shown in FIG. 5, the device includes:

[0185] The generation module 502 is configured to input an initial image into an image processing model, and generate a mask image corresponding to the initial image by using the image processing model, wherein the size of the mask image is greater than the size of the initial image;

[0186] The determination module 504 is configured to extract image generation information from the initial image, and determine a plurality of image generation guide information;

[0187] The fusion module 506 is configured to fuse and add noise to the mask image and the initial image to obtain a fused noise image;

[0188] The obtaining module 508 is configured to denoise the fused noise image according to the plurality of image generation guide information to obtain a target image output by the image processing model.

[0189] Optionally, the obtaining module 508 is further configured to:

[0190] The feature denoising module is used to denoise the fused noise image according to the plurality of image generation guide information to obtain a denoised image feature, wherein the denoised image feature includes channel image features of a plurality of channels.

[0191] The feature toning module is used to perform toning processing on the channel image feature of the target channel in the multiple channels, to obtain a target image feature;

[0192] The target image is obtained according to the target image feature.

[0193] Optionally, the obtaining module 508 is further configured to:

[0194] The feature toning module is used to calculate a vector mean of the channel image feature of the target channel, to obtain a mean channel image feature, and calculate a toning amplitude value according to the mean channel image feature;

[0195] The target image feature is obtained according to the channel image feature of the target channel and the toning amplitude value.

[0196] Optionally, the generating module 502 is further configured to:

[0197] The mask image generation is performed by using the mask generation unit, to obtain a mask image corresponding to the initial image.

[0198] Optionally, the determining module 504 is further configured to:

[0199] The image generation information extraction is performed on the initial image by using the extended guidance unit, to obtain text description guidance information and image feature guidance information of the initial image.

[0200] Optionally, the determining module 504 is further configured to:

[0201] The text description processing is performed on the initial image by using the multi-modal processing module, to obtain the text description guidance information.

[0202] The image encoding processing is performed on the initial image by using the image encoding module, to obtain the image feature guidance information.

[0203] Optionally, the fusion module 506 is further configured to:

[0204] The fusion processing is performed on the mask image and the initial image by using the image generation unit, to obtain a fusion image, and the fusion image is subjected to noise adding processing, to obtain a fusion noise image.

[0205] Optionally, the obtaining module 508 is further configured to:

[0206] The feature extraction is performed on the fusion noise image by using the feature denoising module, to obtain a noise image feature of the fusion noise image.

[0207] The feature denoising module is used to denoise the noisy image features according to the text description guide information and the image feature guide information through the cross attention mechanism, and obtain denoised image features.

[0208] The target image is obtained according to the denoised image features.

[0209] Optionally, the obtaining module 508 is further configured to:

[0210] The text description guide information is processed by cross attention with the noisy image features in each of the plurality of denoising network layers, and the image feature guide information is processed by cross attention with the noisy image features in the target denoising network layer in the plurality of denoising network layers, so as to denoise the noisy image features and obtain the denoised image features output by the last denoising network layer in the plurality of denoising network layers.

[0211] The device further comprises a response module configured to determine an initial image carried in the image processing instruction and a target expansion size for the initial image in response to the image processing instruction.

[0212] Optionally, the generating module 502 is further configured to:

[0213] The initial image and the target expansion size are input into the image processing model, and the image processing model is used to expand the initial image according to the target expansion size to generate a mask image corresponding to the initial image.

[0214] The device further comprises a text and image determination module configured to determine the initial image and the image processing prompt text.

[0215] Optionally, the generating module 502 is further configured to:

[0216] The initial image and the image processing prompt text are input into the image processing model, and the image processing model is used to expand the initial image according to the target expansion size contained in the image processing prompt text to generate a mask image corresponding to the initial image.

[0217] The device further comprises a receiving module configured to receive an initial image sent by a client, wherein the initial image is sent by the client in response to an interactive operation instruction of a user interaction interface.

[0218] The device further comprises a sending module configured to return the target image to the client to display the target image on a user interaction interface of the client.

[0219] The image processing device provided by one embodiment of the present disclosure can generate a mask image corresponding to an initial image through an image processing model, the size of the mask image is greater than the size of the initial image, so that expansion of the initial image is realized on the basis of the mask image, guide information is determined through a plurality of image generation guide information, information about the style, picture content, features and the like of the initial image is obtained, the mask image and the initial image are fused and noise is added to obtain a fused noise image, and a clear target image after expansion of the initial image is obtained through denoising of the fused noise image. In the denoising process, the fused noise image is denoised by using the plurality of image generation guide information, so that the picture content of the expanded area in the generated target image is consistent with the style of the initial image and consistent with the picture content of the initial image, and problems such as inconsistent style, poor fusion and discomfort in the expanded area of the target image and the initial image are avoided, and the result quality of the generated image expansion is improved.

[0220] The above is a schematic scheme of the image processing device of the present embodiment. It should be noted that the technical scheme of the image processing device belongs to the same concept as the technical scheme of the image processing method described above, and the details of the technical scheme of the image processing device that are not described in detail can be referred to the description of the technical scheme of the image processing method.

[0221] Corresponding to the method embodiments described above, the present disclosure also provides image expansion device embodiments. FIG. 6 shows a structural schematic diagram of an image expansion device according to one embodiment of the present disclosure. As shown in FIG. 6, the device includes:

[0222] The generation module 602 is configured to input the image to be expanded into an image processing model, and generate an expansion mask image corresponding to the image to be expanded by using the image processing model, wherein the size of the expansion mask image is greater than the size of the image to be expanded;

[0223] The determination module 604 is configured to extract image generation information from the image to be expanded, and determine a plurality of image generation guide information;

[0224] The fusion module 606 is configured to fuse and add noise to the expansion mask image and the image to be expanded to obtain a fused noise image;

[0225] The obtaining module 608 is configured to denoise the fused noise image according to the plurality of image generation guide information to obtain an expansion image output by the image processing model.

[0226] Optionally, the obtaining module 608 is further configured to:

[0227] The feature denoising module is used to denoise the fusion noise image according to the guide information generated from the plurality of images to obtain denoised image features, wherein the denoised image features include channel image features of a plurality of channels;

[0228] The feature toning module is used to tone the channel image features of a target channel in the plurality of channels to obtain extended image features.

[0229] The extended image is obtained according to the extended image features.

[0230] The above is a schematic scheme of the image extension device of the embodiment. It should be noted that the technical scheme of the image extension device belongs to the same concept as the technical scheme of the image extension method described above, and the details of the technical scheme of the image extension device that are not described in detail can be referred to the description of the technical scheme of the image extension method.

[0231] FIG. 7 shows a structural block diagram of a computing device 700 according to one embodiment of the present disclosure. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 through a bus 730, and a database 750 is used to save data.

[0232] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 740 can include one or more of any type of network interface (e.g., a network interface card (NIC)) such as a IEEE 802.11 wireless local area network (WLAN) wireless interface, a worldwide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, near field communication (NFC), or the like.

[0233] In one embodiment of the present disclosure, the above-mentioned components of the computing device 700 and other components not shown in FIG. 7 can also be connected to each other, for example, through a bus. It should be understood that the computing device structure block diagram shown in FIG. 7 is merely for the purpose of example, and is not a limitation on the scope of the present disclosure. Other components can be added or replaced as needed by those skilled in the art.

[0234] The computing device 700 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 700 can also be a mobile or stationary server.

[0235] The processor 720 is configured to execute computer programs / instructions that implement the steps of the above-mentioned image processing method and image extension method.

[0236] Each of the embodiments in the present disclosure is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the computing device embodiment is described simply because it is basically similar to the image processing method and image extension method embodiments, and the relevant parts can be referred to the description of the image processing method and image extension method embodiments.

[0237] An embodiment of the present disclosure also provides a computer-readable storage medium storing computer programs / instructions that implement the steps of the above-mentioned image processing method and image extension method.

[0238] Each of the embodiments in the present disclosure is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the computer-readable storage medium embodiment is described simply because it is basically similar to the image processing method and image extension method embodiments, and the relevant parts can be referred to the description of the image processing method and image extension method embodiments.

[0239] An embodiment of the present disclosure also provides a computer program product including computer programs / instructions that implement the steps of the above-mentioned image processing method and image extension method.

[0240] The above is a schematic scheme of a computer program product of the embodiment. It should be noted that the technical scheme of the computer program product is the same as the technical scheme of the image processing method and the image extension method described above, and the technical scheme of the computer program product is not described in detail. The contents can be seen from the description of the technical scheme of the image processing method and the image extension method.

[0241] The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than those described in the embodiments and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0242] The computer instructions include computer program code, which can be in the form of source code, object code, executable code, or some intermediate form. The computer-readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice, for example, in some regions, according to the patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0243] It should be noted that for the foregoing method embodiments, in order to facilitate description, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the embodiments of the present disclosure are not limited by the order of the described actions, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the embodiments of the present disclosure.

[0244] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be seen from the related description of other embodiments.

[0245] The preferred embodiments of the present disclosure disclosed above are only used to help explain the present disclosure. The alternative embodiments do not describe all the details of the present disclosure, nor limit the present disclosure to the specific embodiments. Obviously, according to the content of the embodiments of the present disclosure, many modifications and variations can be made. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can well understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. An image processing method, comprising: inputting an initial image into an image processing model, and generating a mask image corresponding to the initial image by using the image processing model, wherein a size of the mask image is greater than a size of the initial image; extracting image generation information from the initial image to determine a plurality of image generation guide information; fusing and adding noise to the mask image and the initial image to obtain a fused noise image; de-noising the fused noise image according to the plurality of image generation guide information to obtain a target image output by the image processing model.

2. The image processing method of claim 1, wherein the image processing model comprises an image generation unit, and the image generation unit comprises a feature de-noising module and a feature toning module; and the de-noising the fused noise image according to the plurality of image generation guide information to obtain the target image output by the image processing model comprises: de-noising the fused noise image according to the plurality of image generation guide information by using the feature de-noising module to obtain a de-noised image feature, wherein the de-noised image feature comprises a plurality of channel image features; toning the channel image feature of a target channel in the plurality of channels by using the feature toning module to obtain a target image feature; and obtaining the target image according to the target image feature.

3. The image processing method of claim 2, wherein the toning the channel image feature of the target channel in the plurality of channels by using the feature toning module to obtain the target image feature comprises: calculating a vector mean of the channel image feature of the target channel by using the feature toning module to obtain a mean channel image feature, and calculating a toning amplitude value according to the mean channel image feature; and obtaining the target image feature according to the channel image feature of the target channel and the toning amplitude value.

4. The image processing method of claim 1, wherein the image processing model comprises a mask generation unit; and the generating the mask image corresponding to the initial image by using the image processing model comprises: generating the mask image corresponding to the initial image by using the mask generation unit.

5. The image processing method of claim 1, wherein the image processing model comprises an expansion guide unit, the image generation guide information comprises text description guide information and image feature guide information; and the extracting the image generation information from the initial image to determine the plurality of image generation guide information comprises: extracting the text description guide information and the image feature guide information of the initial image by using the expansion guide unit.

6. The image processing method of claim 5, wherein the expansion guide unit comprises a multi-modal processing module and an image encoding module; and the extracting the text description guide information and the image feature guide information of the initial image by using the expansion guide unit comprises: ​ ​ ​ ​ The multi-modal processing module is used to perform text description processing on the initial image to obtain the text description guide information. The image encoding module is used to perform image encoding processing on the initial image to obtain the image feature guide information.

7. The image processing method of claim 1, wherein the image processing model comprises an image generation unit. The fusion and noise adding processing on the mask image and the initial image to obtain a fusion noise image comprises: The image generation unit is used to perform fusion processing on the mask image and the initial image to obtain a fusion image, and perform noise adding processing on the fusion image to obtain the fusion noise image.

8. The image processing method of claim 1, wherein the image processing model comprises an image generation unit, the image generation unit comprises a feature denoising module, and the plurality of image generation guide information comprises text description guide information and image feature guide information. The denoising on the fusion noise image according to the plurality of image generation guide information to obtain the target image comprises: The feature denoising module is used to perform feature extraction on the fusion noise image to obtain noise image features of the fusion noise image. The feature denoising module is used to perform denoising on the noise image features according to the text description guide information and the image feature guide information through a cross-attention mechanism to obtain denoised image features. The feature denoising module is used to obtain the target image according to the denoised image features.

9. The image processing method of claim 8, wherein the feature denoising module comprises a plurality of denoising network layers. The feature denoising module is used to perform denoising on the noise image features according to the text description guide information and the image feature guide information through a cross-attention mechanism to obtain denoised image features, comprising: The text description guide information is processed through cross-attention with noise image features in each denoising network layer in the plurality of denoising network layers, and the image feature guide information is processed through cross-attention with noise image features of a target denoising network layer in the plurality of denoising network layers, so as to denoise the noise image features and obtain the denoised image features output by a last denoising network layer in the plurality of denoising network layers.

10. The image processing method of claim 1, wherein before the initial image is input into the image processing model and the mask image corresponding to the initial image is generated by using the image processing model, the method further comprises: In response to an image processing instruction, determining the initial image and a target expansion size for the initial image carried in the image processing instruction; The initial image is input into the image processing model, and the mask image corresponding to the initial image is generated by using the image processing model, comprising: The initial image and the target expansion size are input into the image processing model, and the initial image is expanded according to the target expansion size by using the image processing model to generate the mask image corresponding to the initial image.

11. The image processing method of claim 1, before the inputting the initial image into the image processing model and generating the mask image corresponding to the initial image by using the image processing model, further comprising: determining the initial image and image processing prompt text; and the inputting the initial image into the image processing model and generating the mask image corresponding to the initial image by using the image processing model comprises: inputting the initial image and the image processing prompt text into the image processing model, and generating the mask image corresponding to the initial image by using the image processing model and according to a target expansion size contained in the image processing prompt text.

12. The image processing method of claim 1, before the inputting the initial image into the image processing model and generating the mask image corresponding to the initial image by using the image processing model, further comprising: receiving the initial image sent by a client, wherein the initial image is sent by the client in response to an interactive operation instruction of a user interaction interface; and after the denoising the fusion noise image according to the plurality of image generation guide information to obtain the target image output by the image processing model, further comprising: returning the target image to the client to display the target image on a user interaction interface of the client.

13. An image expansion method, comprising: inputting an image to be expanded into an image processing model, and generating an expansion mask image corresponding to the image to be expanded by using the image processing model, wherein a size of the expansion mask image is greater than a size of the image to be expanded; extracting image generation information from the image to be expanded to determine a plurality of image generation guide information; fusing and adding noise to the expansion mask image and the image to be expanded to obtain a fusion noise image; and denoising the fusion noise image according to the plurality of image generation guide information to obtain an expanded image output by the image processing model.

14. The image expansion method of claim 13, wherein the image processing model comprises an image generation unit, and the image generation unit comprises a feature denoising module and a feature color adjustment module; and the denoising the fusion noise image according to the plurality of image generation guide information to obtain the expanded image output by the image processing model comprises: denoising the fusion noise image according to the plurality of image generation guide information by using the feature denoising module to obtain denoised image features, wherein the denoised image features comprise channel image features of a plurality of channels; adjusting a color of the channel image features of a target channel in the plurality of channels by using the feature color adjustment module to obtain expanded image features; and obtaining the expanded image according to the expanded image features.

15. A computing device, comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, and the computer programs / instructions, when executed by the processor, implement the steps of the method of any one of claims 1-14. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 16. A computer readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the method of any one of claims 1-14.

17. A computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the method of any one of claims 1-14.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and storage medium

    CN116993864A

  • Image generation method and device, and training method and device of generative model

    CN117196992A

  • Image expansion method and device, electronic equipment and storage medium

    CN118154438A

  • Artifact reduction for image style transfer

    US20180300850A1