Image processing method and device, equipment and storage medium
By introducing predetermined prompts and query representations specific to style feature extraction, the dual decoupled representation extraction technology is used to solve the problem of time-consuming and semantic conflicts in the prior art, and the rapid generation of images that conform to style and semantics is achieved.
Patent Information
- Application Number
- CN202410070352.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art takes a long time and requires manual participation when generating images with a specific style, and it is difficult to decouple the semantics of the reference image and the semantics of the input text, resulting in the possible semantic conflicts of the generated images.
Pre-determined prompts and query representations specific to style feature extraction are introduced. Style-related features of the reference image are extracted through the feature transformation model, and target images are generated by the input text. Double decoupled representation extraction (DDRE) technology is used to extract style-related features and semantic-related features respectively.
It realizes the rapid generation of images with reference image style and conforming to the semantics of input text, alleviates the problem of semantic inconsistency and improves the efficiency and accuracy of image generation.
Smart Images

Figure CN120339050A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for image processing. Background Art
[0002] In the field of computer vision (CV), various machine learning-based image generation technologies have been significantly developed and have a wide range of applications. For example, in many application scenarios such as social, gaming, and image editing, it is desired to generate and use images with a specific style. Machine learning-based image generation technologies can be used in such application scenarios to improve the image generation effect. Summary of the Invention
[0003] In a first aspect of the present disclosure, there is provided a method for image processing. The method includes: obtaining first reference image features of a first reference image; determining style-related features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query representation; and generating a first target image based on the style-related features of the first reference image and a first input text indicating the content of the target image, the first target image matching the image style and the target image content of the first reference image.
[0004] In a second aspect of the present disclosure, there is provided an apparatus for image processing. The apparatus includes: a first image feature extraction module configured to obtain first reference image features of a first reference image; a style-related feature module configured to determine style-related features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query representation; and a first target image generation module configured to generate a first target image based on the style-related features of the first reference image and a first input text indicating the content of the target image, the first target image matching the image style and the target image content of the first reference image.
[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium. A computer program is stored on the computer-readable storage medium and can be executed by a processor to implement the method of the first aspect.
[0007] It should be understood that the content described in this content section is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0009] Figure 1 FIG. shows a schematic diagram of an exemplary environment in which embodiments of the present disclosure can be implemented;
[0010] Figure 2 FIG. shows a schematic diagram of an exemplary architecture for generating an image with a reference style and target content according to some embodiments of the present disclosure;
[0011] Figure 3 FIG. shows a schematic diagram of an exemplary architecture for generating an image with a reference content and target style according to some embodiments of the present disclosure;
[0012] Figure 4 FIG. shows a schematic diagram of the application of a cross-attention mechanism according to some embodiments of the present disclosure;
[0013] Figure 5A FIG. shows a schematic diagram of a training task for style feature extraction according to some embodiments of the present disclosure;
[0014] Figure 5B FIG. shows a schematic diagram of a training task for semantic feature extraction according to some embodiments of the present disclosure;
[0015] Figure 5C FIG. shows a schematic diagram of a training task for image reconstruction according to some embodiments of the present disclosure;
[0016] Figure 6 FIG. shows a flowchart of a process for image processing according to some embodiments of the present disclosure;
[0017] Figure 7 FIG. shows a block diagram of an apparatus for image processing according to some embodiments of the present disclosure; and
[0018] Figure 8 FIG. shows a block diagram of a device capable of implementing multiple embodiments of the present disclosure. DETAILED DESCRIPTION
[0019] It is understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users in an appropriate manner and user authorization should be obtained in accordance with relevant laws and regulations.
[0020] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0021] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0022] It is understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0023] It is understood that the data involved in the technical solution of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.
[0024] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0025] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. In addition, the embodiments described in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different sections / subsections in any manner.
[0026] In this document, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.
[0027] In the description of the embodiments of the present disclosure, the term "including" and its like shall be understood as open inclusion, i.e., "including but not limited to". The term "based on" shall be understood as "at least partially based on". The term "one embodiment" or "the embodiment" shall be understood as "at least one embodiment". The term "some embodiments" shall be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter. The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0028] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, for a given input, a corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In this document, "model" may also be referred to as "machine learning model", "machine learning network" or "network", and these terms are used interchangeably herein. A model can also include different types of processing units or networks.
[0029] As used herein, that a target image matches an image style may refer to that the target image visually has a style consistent with or similar to that image style. That a target image matches image content may refer to that the target image visually contains the same or similar content as that target content.
[0030] Example environment
[0031] Figure 1 A schematic diagram of an example environment 100 in which the embodiments of the present disclosure can be implemented is shown. In environment 100, an image processing system 120, also simply referred to as system 120, is deployed in an electronic device 110. The image processing system 120 is configured to generate a target image 105 based on an input text 102 and a reference image 101.
[0032] The reference image 101 can be used to provide or indicate image elements expected to be included in the target image 105, such as image style, image content, etc. The input text can be used to indicate another image element expected to be included in the target image 105. As used herein, the term "image element" can indicate various explicit or implicit elements that make up an image. Exemplarily, image elements can include image style or image content. Examples of image styles include but are not limited to watercolor, crayon, sketch, comic, etc. Examples of image content can include but are not limited to at least a part of the foreground of the image, at least a part of the background of the image, etc.
[0033] In some embodiments, the image processing system 120 may generate an image having the same style as the reference image 101. Hereinafter, the image style of the reference image 101 is also referred to as the reference style. In such an embodiment, the input text 102 may indicate the target image content. The generated target image 105 may match the image style of the reference image 101 and the target image content. For example, the target image 105 may have the image style of the reference image 101 and include the target image content.
[0034] Alternatively or additionally, in some embodiments, the image processing system 120 may generate an image having the same content as the reference image 101. In such an embodiment, the input text 102 may indicate the target image style. The generated target image 105 may match the target image style and the image content in the reference image 101. For example, the target image 105 may include the image content in the reference image 101 and have the style indicated by the input text 102.
[0035] In the environment 100, the electronic device 110 may be any type of device with computing capabilities, including a terminal device or a server device. The terminal device may be any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server device may, for example, include a computing system / server, such as a mainframe, an edge computing node, an electronic device in a cloud environment, and so on.
[0036] It should be understood that the structure and function of the environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure.
[0037] As mentioned above, it is desired to generate an image having the reference style of the reference image. To this end, in some conventional solutions, all or part of the parameters of a basic image generation model (such a model can generate an image according to the input text) may be fine-tuned to generate a stylized image. However, this solution takes a long time (e.g., at least in minutes) and requires human participation.
[0038] In some other conventional solutions, a sampled two-stage encoder is used, which includes a fixed backbone network and a head network that needs to be trained, to extract the features of the reference image. The features obtained by the encoder can be further combined with a denoising network. However, in this convention, the training task is usually a reconstruction task, so what the encoder learns is a hybrid of content and style information. This results in a conflict between the semantics in the reference image and the semantics indicated by the input text, which may cause the generated image to contain the semantics of the reference image.
[0039] For this reason, embodiments of the present disclosure provide an improved solution for image generation. In the embodiments of the present disclosure, a predetermined prompt and query representation specific to style feature extraction are introduced to extract style-related features in the reference image. Then, based on the extracted style-related features and the input text indicating the content of the target image, a target image that matches the image style of the reference image and the content of the target image can be generated.
[0040] In the embodiments of the present disclosure, a predetermined prompt and query representation specific to style feature extraction are used for feature extraction to decouple the style features and semantic features of the reference image. This can alleviate the inconsistency problem between the semantics in the reference image and the semantics of the input text. Thus, an image that has both the style of the reference image and conforms to the semantics of the input text can be advantageously generated.
[0041] Example image generation architecture
[0042] In some embodiments, the input text 102 can indicate the content that the desired target image is to have, also referred to as the target image content or simply the target content. In such embodiments, what the reference image 101 indicates is the image style. The target image 105 has the reference style and the target content. Figure 2 A schematic diagram of an exemplary architecture 200 for generating an image with a reference style and a target content according to some embodiments of the present disclosure is shown. The architecture 200 can be implemented in the system 120. For the extraction of image features, the architecture 200 adopts a two-stage structure.
[0043] As Figure 2 shown, the reference style of the reference image 201 is Style A. The system 120 can obtain the image features of the reference image 201, also referred to as the reference image features 202. Exemplarily, the image encoder 220 can generate the reference image features 202 based on the reference image 201. The image encoder 220 can adopt any suitable network structure, and the embodiments of the present disclosure are not limited in this regard.
[0044] The reference image feature 202 can be used as the input of the feature transformation model 230. In addition, the feature transformation model 230 also has other inputs, including a predetermined prompt 204 (also referred to as the first predetermined prompt) for style feature extraction and a query representation 203 (also referred to as the first query representation) for style feature extraction. The predetermined prompt 204 can be any suitable text or prompt word. Showing the predetermined prompt 204 as the word "style" in this example is only exemplary and is not intended to be any limitation. In the embodiments of the present disclosure, the first predetermined prompt is specific to style feature extraction and is independent of the reference image. That is, the first predetermined prompt does not change with the reference image. The feature transformation model 230 can be implemented by a network with any suitable structure. By way of example and not limitation, the feature transformation model 230 can be a Query Transformer.
[0045] The query representation 203 is learnable. The query representation 203 can be obtained during the training of the feature transformation model 230. For example, the query representation 203 can be initialized, and during the training of the feature transformation model 230, the query representation 203 can be updated together until it is solidified upon completion of the training. Such a query representation 203 can be used to instruct the feature transformation model 230 to extract features related to the image style. As Figure 2 shown, the query representation 203 can be a vectorized representation of any suitable dimension. The query representation 203 is obtained for style feature extraction.
[0046] The feature transformation model 230 can generate style-related features 205 of the reference image 201 based on the reference image feature 202, the first predetermined prompt 204, and the first query representation 203. In other words, under the prompt or guidance of the first predetermined prompt 204 and the first query representation 203, the feature transformation model 230 can extract features related to the image style from the reference image 201.
[0047] Next, the system 120 can generate the target image 210 based on the style-related features 205 and the input text 209 indicating the content of the target image. The input text 209 can use any suitable characters to indicate the content expected to be included in the target image, such as but not limited to people, animals, items, landscapes, buildings, etc. In this example, the target content indicated by the input text 209 is "panda", but this is only exemplary and is not intended to be any limitation. As Figure 2 shown, the target image 210 has the style of the reference image 201 (i.e., style A) and contains the content indicated by the input text 209 (i.e., panda).
[0048] The system 120 can adopt any suitable algorithm or model to generate the target image. As Figure 2As shown, in some embodiments, the model 240, also referred to as the first model, can be employed. Any suitable network structure can be used to implement the model 240. Exemplarily, the model 240 can be a diffusion model, which can perform multiple denoising steps. The model 240 can also be of other types, such as a generative adversarial network.
[0049] Specifically, the text encoder 250 can generate the text features 207 of the input text 209. The text features 207 can be provided to the model 240 as text conditions. In addition, the model 240 can also receive style-related features 205. The model 240 can generate target image features 208 based on the text features 207 and the style-related features 205. The target image features 208 can then be used to generate the target image 210. For example, the system 120 can utilize an image decoder (not shown) to convert the target image features 208 into the target image 210. The embodiments of the present disclosure are not limited in how to convert the target image features into the target image. In addition, in each step (e.g., denoising step) performed by the model 240, the model 240 also receives the image features output from the previous step as input.
[0050] In some embodiments, the model 240 can include multiple processing layers, and the multiple processing layers are respectively used to process features of corresponding sizes. In such an embodiment, the style-related features 205 are provided to a first number of processing layers whose size is greater than a first threshold size. For example, the model 240 can be a denoising U-shaped network. Exemplarily, as Figure 2 shown, the processing layers close to the model input and output are used to process features of larger sizes, while the processing layers close to the middle of the model are used to process features of smaller sizes. Correspondingly, the style-related features 205 are provided to the processing layers close to the model input and output. By way of example and not intending any limitation, assume that the model 240 includes 16 processing layers, which are numbered 0, 1,..., 15 in sequence. The style-related features 205 can be provided to layers 0 to 3 and 9 to 15. The combination of the style-related features and the text features will be described below with reference to Figure 4 describe the combination of the style-related features and the text features.
[0051] Features of larger sizes include more details of the image and correspond to fine processing layers. It is generally considered that such fine processing layers are mainly responsible for generating style-related elements such as the color and structure of the image. Features of smaller sizes include the high-level semantic information of the image and correspond to coarser processing layers. It is generally considered that such coarser processing layers are mainly responsible for the semantic generation of the image. Therefore, by injecting the style-related features only into the fine processing layers and not into the coarser processing layers, the style and semantics of the reference image can be further decoupled. In such an embodiment, it can be advantageously made that the model 240 is more focused on the style of the reference image in image generation.
[0052] The above describes an example implementation of generating an image with a reference style and target content. The following describes an example implementation of generating an image with a target style and reference content.
[0053] In some embodiments, the input text 102 may indicate the style that the desired target image is to have, also referred to as the target image style or simply the target style. In such embodiments, what the reference image 101 indicates is the image content, also referred to as the reference content. The target image 105 has the reference content and the target style. Figure 3 A schematic diagram of an example architecture 300 for generating an image with a reference content and a target style according to some embodiments of the present disclosure is shown. The architecture 300 may be implemented in the system 120. For the extraction of image features, the architecture 300 employs a two-stage structure. Note that Figure 3 the text encoder 350 in Figure 2 may be the same as or different from the text encoder 250 in Figure 3 and the image encoder 320 in Figure 2 may be the same as or different from the image encoder 220 in
[0054] As Figure 3 shown, the reference style of the reference image 301 is Style A, and the image content is a smiling face. The system 120 may obtain the image features of the reference image 301, also referred to as the reference image features 302. Exemplarily, the image encoder 320 may generate the reference image features 302 based on the reference image 301. The image encoder 320 may employ any suitable network structure, and the embodiments of the present disclosure are not limited in this regard.
[0055] The reference image features 302 may be used as an input to the feature transformation model 330. In addition, the feature transformation model 330 also has other inputs, including a predetermined prompt 304 (also referred to as the second predetermined prompt) for semantic feature extraction and a query representation 303 (also referred to as the second query representation) for semantic feature extraction. The predetermined prompt 304 may be any suitable text or prompt word. The showing of the predetermined prompt 304 as the word "content" in this example is merely exemplary and is not intended to be limiting in any way. In the embodiments of the present disclosure, the second predetermined prompt is specific to semantic feature extraction and is independent of the reference image. That is, the second predetermined prompt does not change with the reference image. The feature transformation model 230 may be implemented with a network of any suitable structure. By way of example and not intended to be limiting in any way, the feature transformation model 230 may be a Query Transformer.
[0056] The query representation 303 is learnable. The query representation 303 can be obtained during the training of the feature transformation model 330. For example, the query representation 303 can be initialized, and during the training of the feature transformation model 330, the query representation 303 can be updated together until it is solidified upon completion of the training. Such a query representation 303 can be used to instruct the feature transformation model 330 to extract features related to the image content, that is, semantically related features. As Figure 3 shown, the query representation 303 can be a vectorized representation of any suitable dimension. The query representation 303 is obtained for semantic feature extraction.
[0057] Based on the reference image feature 302, the second predetermined prompt 304, and the second query representation 303, the feature transformation model 330 can generate semantically related features 305 of the reference image 301. In other words, under the prompt or guidance of the second predetermined prompt 304 and the second query representation 303, the feature transformation model 330 can extract features related to the image semantics from the reference image 301.
[0058] Next, the system 120 can generate the target image 310 based on the semantically related features 305 and the input text 309 indicating the style of the target image. The input text 309 can use any suitable characters to indicate the style that the target image is desired to have. In this example, the target style indicated by the input text 309 is "Style B", but this is merely exemplary and not intended to be limiting in any way. As Figure 3 shown, the target image 310 has the content of the reference image 301 (i.e., a smiling face) and includes the style indicated by the input text 309 (i.e., Style B).
[0059] Similar to that Figure 2 described above, the system 120 can adopt any suitable algorithm or model to generate the target image 310. As Figure 3 shown, in some embodiments, the model 240 can be adopted. Specifically, the text encoder 350 can generate the text feature 307 of the input text 309. The text feature 307 can be provided to the model 240 as a text condition. In addition, the model 240 can also receive the semantically related features 305. The model 240 can generate the target image feature 308 based on the text feature 307 and the semantically related features 305. The target image feature 308 can then be used to generate the target image 310. For example, the system 120 can utilize an image decoder (not shown) to convert the target image feature 308 into the target image 310. The embodiments of the present disclosure are not limited in how to convert the target image feature into the target image. In addition, in each step (e.g., the denoising step) performed by the model 240, the model 240 also receives the image feature output by the previous step as an input.
[0060] In some embodiments, the model 240 may include multiple processing layers, and the multiple processing layers are respectively used to process features of corresponding sizes. In such an embodiment, the semantically related features 305 are provided to a second number of processing layers whose sizes are smaller than a second threshold size. The second threshold size may be the same as or different from the first threshold size described above. For example, the model 240 may be a denoising U-shaped network. Exemplarily, as Figure 3 shown, the processing layers close to the model input and output are used to process features of larger sizes, while the processing layers close to the middle of the model are used to process features of smaller sizes. Accordingly, the semantically related features 305 are provided to the processing layers close to the middle of the model. By way of example and not limitation, assume that the model 240 includes 16 processing layers, which are sequentially numbered 0, 1,..., 15. The semantically related features 305 may be provided to layers 4 to 8. The combination of semantically related features and text features will be described below with reference to Figure 4 the description.
[0061] As described in reference to Figure 2 , features of larger sizes include more details of the image and correspond to fine processing layers. It is generally believed that such fine processing layers are mainly responsible for generating style-related elements such as the color and structure of the image. Features of smaller sizes include high-level semantic information of the image and correspond to coarser processing layers. It is generally believed that such coarser processing layers are mainly responsible for semantic generation of the image. Therefore, by injecting semantically related features only into coarser processing layers and not into fine processing layers, the style and semantics of the reference image can be further decoupled. In such an embodiment, the model 240 can be made more focused on the content of the reference image in image generation.
[0062] Combination of features of reference image and text features
[0063] As mentioned above, in some embodiments, style-related features and semantically related features may be provided to certain processing layers of the model 240. For illustrative purposes only below, style-related features and semantically related features are collectively referred to as image-related features. In such processing layers, the image features and the text features of the input text can be used as conditions for target image generation. In some embodiments, the model 240 may be based on a cross-attention mechanism. Accordingly, in the processing layers, cross-attention related to image-related features and text features can be applied.
[0064] Figure 4 shows a schematic diagram of the application of the cross-attention mechanism according to some embodiments of the present disclosure. In Figure 4 , the corresponding sizes are shown below each feature. As Figure 4As shown, in a certain processing layer, the input image features 411 of this processing layer (which are represented by Z and can come from the previous processing layer or the previous denoising step) are converted into query features 421. By converting the text features 412 (which are represented by c t ), and the image-related features 413 (which are represented by c j and can be semantic-related features or style-related features), key features 431 and value features 432 are generated. Then, based on the query features 421, key features 431, and value features 432, the output image features 450 of this processing layer can be generated. As Figure 4 shown, the dot product of the query features 421 and the key features 431 can obtain the attention map. Through the matrix multiplication of the attention map and the value features 432, the output image features 450 can be obtained.
[0065] In some embodiments, the parameters of such a processing layer can be obtained through different training modes, and different training modes can include, for example, pre-training and fine-tuning. As Figure 4 shown, the processing layer can include a first conversion unit 401 for converting the input image features 411 into query features 421. The processing layer can also include a second conversion unit 402, a third conversion unit 403, a fourth conversion unit 404, and a fifth conversion unit 405. The second conversion unit 402 is used to convert the text features 412 into the first intermediate text features 422 as part of the key features. The third conversion unit 403 is used to convert the text features 412 into the second intermediate text features 423 as part of the value features. The fourth conversion unit 404 is used to convert the image-related features 413 into the first intermediate image-related features 424 as part of the key features. The fifth conversion unit 405 is used to convert the image-related features 413 into the second intermediate image-related features 425 as part of the value features.
[0066] The first intermediate text features 422 and the first intermediate image-related features 424 can be combined into the key features 431. For example, they can be concatenated into the key features 431. Similarly, the second intermediate text features 423 and the second intermediate image-related features 425 can be combined into the value features 432. For example, they can be concatenated into the value features 432.
[0067] In some embodiments, the first conversion unit 401, the second conversion unit 402, and the third conversion unit 403 are obtained through a first training mode, and the fourth conversion unit 404 and the fifth conversion unit 405 are obtained through a second training mode different from the first training mode. For example, the first conversion unit 401, the second conversion unit 402, and the third conversion unit 403 may be obtained through pre-training, and the fourth conversion unit 404 and the fifth conversion unit 405 may be obtained through fine-tuning. In this way, additional units can be added to the pre-trained base model, so that the additional units can be fine-tuned without training the parts included in the base model. This can effectively reduce the training cost.
[0068] In the embodiments described above, the semantic information and the style information in the reference image are decoupled. This can be regarded as a Dual-Decoupled Representation Extraction (DDRE). Specifically, a predetermined prompt text specific to style feature extraction (e.g., "style") and a predetermined prompt text specific to semantic feature extraction (e.g., "content") are respectively input to the feature transformation model, so that the feature transformation model can obtain image features aligned with the prompt text. That is, style-related features and semantic-related features can be obtained respectively.
[0069] Training tasks and samples
[0070] To implement the above Dual-Decoupled Representation Extraction, various suitable training tasks can be performed during training.
[0071] In some embodiments, a training task for style feature extraction, also referred to as a Style Representation Extraction (STRE) task, can be performed. In the STRE task, paired different sample images with the same style are used for training. Figure 5A A schematic diagram of a training task 500A for style feature extraction according to some embodiments of the present disclosure is shown.
[0072] In training task 500A, the paired sample images 511 and the sample image 520 have the same style, such as style A. The sample image 511 is used as a reference, and the sample image 520 is used as a target. During training, the first query representation 203 is initialized. The feature transformation model 230 can generate the style-related feature 515 of the sample image 511 based on the first sample image feature of the sample image 511 (e.g., the output of the image encoder 220), the first predetermined prompt 204, and the first query representation 203. Taking the style-related feature 515 of the sample image 511 and the text 519 indicating the image content of the sample image 520 as conditions, the feature transformation model 230 and the initialized first query representation 203 are updated by adding noise to and denoising the sample image 520. When a predetermined condition is satisfied, the update of the feature transformation model 230 and the first query representation 203 is stopped. In some embodiments, if the model 240 includes a trainable conversion unit (e.g., as described in Figure 4 ), such a conversion unit is also updated during training.
[0073] The above-described STRE task is a non-reconstruction task. Through such a training task, it is beneficial to decouple the style and semantics of the reference image, and it can also ensure that the image information (style information in this example) is not too strong during the training process, so as not to "drown" the prompt information of the input text.
[0074] In some embodiments, a training task for semantic feature extraction, also referred to as a Semantics Representation Extraction (SERE) task, can be performed. In the SERE task, training is performed using paired different sample images with the same semantics. Figure 5B A schematic diagram of a training task 500B for semantic feature extraction according to some embodiments of the present disclosure is shown.
[0075] In training task 500B, the paired sample image 521 and the sample image 530 have the same content, such as musical notes. The sample image 521 is used as a reference, and the sample image 530 is used as a target. During training, the second query representation 303 is initialized. The feature transformation model 330 can generate the semantic-related feature 525 of the sample image 521 based on the sample image feature of the sample image 531 (e.g., the output of the image encoder 320), the second predetermined prompt 304, and the second query representation 303. Using the semantic-related feature 525 of the sample image 521 and the text 529 indicating the style of the sample image 530 as conditions, the feature transformation model 330 and the initialized second query representation 303 are updated by adding noise to and denoising the sample image 530. When a predetermined condition is satisfied, the update of the feature transformation model 330 and the second query representation 303 is stopped. In some embodiments, if the model 240 includes a trainable transformation unit (e.g., as described in Figure 4 ), such a transformation unit is also updated during training.
[0076] The SERE task described above is a non-reconstruction task. Through such a training task, it is beneficial to decouple the style and semantics of the reference image, and it can also ensure that the image information (semantic information in this example) during the training process is not too strong, so as not to "drown out" the prompt information of the input text.
[0077] In some embodiments, in order to avoid missing image information caused by the non-reconstruction task, a reconstruction task can be additionally performed. Figure 5C A schematic diagram of a training task 500C for image reconstruction according to some embodiments of the present disclosure is shown.
[0078] As Figure 5C shown, the upper half branch is the style branch. During training, the feature transformation model 230 can generate the style-related feature 533 of the sample image 531 based on the sample image feature of the sample image 531 (e.g., the output of the image encoder 220), the first predetermined prompt 204, and the first query representation 203.
[0079] In the semantic branch, the semantic-related feature 532 of the sample image 531 can be obtained. For example, the semantic-related feature 532 can be obtained by using the feature transformation model 330. As shown in the figure, the feature transformation model 330 can generate the semantic-related feature 532 of the sample image 531 based on the sample image feature of the sample image 531 (e.g., the output of the image encoder 320), the second predetermined prompt 304, and the second query representation 303.
[0080] In this way, the style-related feature 533 and the semantics-related feature 532 can be used as conditions to update the feature transformation model 230 and the first query representation 203 by reconstructing the sample image 531. Correspondingly, the feature transformation model 330 and the second query representation 303 are also updated. In some embodiments, if the model 240 includes a trainable transformation unit (e.g., as referred to in Figure 4 described), such a transformation unit is also updated along with the training based on the reconstruction task.
[0081] As described above, during the execution of the training tasks 500A, 500B, and 500C, the feature transformation model 230, the feature transformation model 330, the first query representation 203, and the second query representation 303 are trainable, the model 240 is partially trainable (as the partial described in reference to Figure 4 ), and the remaining models are frozen.
[0082] In some embodiments, to support the above training tasks, a corresponding sample set can be created. Exemplarily, a set of style words indicating different styles and a set of subject words indicating different contents can be determined. Through the combination of the style words and the subject words, multiple pieces of hint information can be obtained, for example, multiple hint words. Each piece of hint information indicates the style and content to be generated. For example, multiple sample images can be generated from the same piece of hint information. In some embodiments, any pair of the multiple sample images generated from the same piece of hint information can be used for the training task of style feature extraction. For example, Figure 5A the sample image 511 and the sample image 520 in are generated from the same hint information. Compared with the image pairs using the same style words but different subject words, the image pairs generated from the same hint information can obtain a better stylization effect.
[0083] In some embodiments, the sample image pairs with the same subject word and different style words can be used for the training task of semantic feature extraction. For example, Figure 5B the sample image 521 and the sample image 530 in can be generated based on the same subject word and different style words.
[0084] Example processes, apparatuses and devices
[0085] Figure 6 shows a flowchart of a process 600 for image processing according to some embodiments of the present disclosure. The process 600 can be implemented at the electronic device 110.
[0086] In block 610, the electronic device 110 obtains the first reference image feature of the first reference image.
[0087] At block 610, the electronic device 110 determines style-related features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query representation.
[0088] At block 610, the electronic device 110 generates a first target image based on the style-related features of the first reference image and a first input text indicating the target image content. The first target image matches the image style of the first reference image and the target image content.
[0089] In some embodiments, generating the first target image includes: obtaining first text features of the first input text; generating first target image features based on the style-related features and the first text features using a first model, where the first model includes a plurality of processing layers respectively for processing features of corresponding sizes, and the style-related features are provided to a first number of processing layers with a size greater than a first threshold size; and generating the first target image based on the first target image features.
[0090] In some embodiments, generating the first target image features using the first model includes: for a given processing layer among the first number of processing layers: converting the input image features of the given processing layer into query features; generating key features and value features by transforming the first text features and the style-related features; generating the output image features of the given processing layer based on the query features, the key features, and the value features.
[0091] In some embodiments, the input image features are converted into query features using a first conversion unit, and generating the key features and the value features includes: converting the first text features into a first intermediate text feature using a second conversion unit; converting the first text features into a second intermediate text feature using a third conversion unit; converting the style-related features into a first intermediate style feature using a fourth conversion unit; converting the style-related features into a second intermediate style feature using a fifth conversion unit; combining the first intermediate text feature and the first intermediate style feature into the key features; and combining the second intermediate text feature and the second intermediate style feature into the value features.
[0092] In some embodiments, the first conversion unit, the second conversion unit, and the third conversion unit are obtained through a first training mode, and the fourth conversion unit and the fifth conversion unit are obtained through a second training mode different from the first training mode.
[0093] In some embodiments, process 600 further includes: obtaining second reference image features of a second reference image; determining semantically related features of the second reference image based on the second reference image features, a second predetermined prompt for semantic feature extraction, and a second query representation; and generating a second target image based on the semantically related features and a second input text indicating the target image style, where the second target image matches the image content and the target image style of the second reference image.
[0094] In some embodiments, generating the second target image includes: obtaining second text features of the second input text; generating second target image features according to a first model based on the semantically related features and the second text features, where the first model includes a plurality of processing layers respectively configured to process features of corresponding sizes, and the semantically related features are provided to a second number of processing layers with sizes smaller than a second threshold size; and generating the second target image based on the second target image features.
[0095] In some embodiments, the style-related features are determined using a feature transformation model, the first query representation is obtained through training of the feature transformation model, and the training of the feature transformation model includes: obtaining a first sample image and a second sample image having the same image style; determining style-related features of the first sample image using the feature transformation model based on first sample image features of the first sample image, a first predetermined prompt, and the first query representation; and updating the feature transformation model and the first query representation by adding noise to and denoising the second sample image with the style-related features of the first sample image and text indicating the image content of the second sample image as conditions.
[0096] In some embodiments, the first sample image and the second sample image are generated based on the same prompt information indicating the style and content of the image to be generated.
[0097] In some embodiments, the training of the feature transformation model further includes: determining style-related features of a third sample image using the feature transformation model based on third sample image features of the third sample image, the first predetermined prompt, and the first query representation; obtaining semantically related features of the third sample image; and updating the feature transformation model and the first query representation by reconstructing the third sample image with the style-related features and the semantically related features as conditions.
[0098] Figure 7 FIG. shows a schematic structural block diagram of an apparatus 700 for image processing according to certain embodiments of the present disclosure. The apparatus 700 can be implemented as or included in an electronic device 110. Each module / component in the apparatus 700 can be implemented by hardware, software, firmware, or any combination thereof.
[0099] As shown in the figure, the apparatus 700 includes a first image feature extraction module 710 configured to obtain first reference image features of a first reference image. The apparatus 700 further includes a style-related feature module 720 configured to determine style-related features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query representation. The apparatus 700 further includes a first target image generation module 730 configured to generate a first target image based on the style-related features of the first reference image and a first input text indicating the content of the target image, where the first target image matches the image style and the target image content of the first reference image.
[0100] In some embodiments, the first target image generation module 730 is further configured to: obtain first text features of the first input text; generate first target image features based on the style-related features and the first text features using a first model, where the first model includes a plurality of processing layers respectively configured to process features of corresponding sizes, and the style-related features are provided to a first number of processing layers with a size greater than a first threshold size; and generate the first target image based on the first target image features.
[0101] In some embodiments, the first target image generation module 730 is further configured to, for a given processing layer among the first number of processing layers: convert the input image features of the given processing layer into query features; generate key features and value features by transforming the first text features and the style-related features; and generate output image features of the given processing layer based on the query features, the key features, and the value features.
[0102] In some embodiments, the input image features are converted into query features using a first conversion unit, and the first target image generation module 730 is further configured to: convert the first text features into a first intermediate text feature using a second conversion unit; convert the first text features into a second intermediate text feature using a third conversion unit; convert the style-related features into a first intermediate style feature using a fourth conversion unit; convert the style-related features into a second intermediate style feature using a fifth conversion unit; combine the first intermediate text feature and the first intermediate style feature into key features; and combine the second intermediate text feature and the second intermediate style feature into value features.
[0103] In some embodiments, the first conversion unit, the second conversion unit, and the third conversion unit are obtained through a first training mode, and the fourth conversion unit and the fifth conversion unit are obtained through a second training mode different from the first training mode.
[0104] In some embodiments, the apparatus 700 further includes: a second image feature extraction module configured to obtain second reference image features of a second reference image; a semantic relevance feature module configured to determine semantic relevance features of the second reference image based on the second reference image features, a second predetermined prompt for semantic feature extraction, and a second query representation; and a second target image generation module configured to generate a second target image based on the semantic relevance features and a second input text indicating the target image style, where the second target image matches the image content and the target image style of the second reference image.
[0105] In some embodiments, the second target image generation module is further configured to: obtain second text features of the second input text; generate second target image features according to a first model based on the semantic relevance features and the second text features, where the first model includes a plurality of processing layers respectively configured to process features of corresponding sizes, and the semantic relevance features are provided to a second number of processing layers with sizes smaller than a second threshold size; and generate the second target image based on the second target image features.
[0106] In some embodiments, the style relevance features are determined using a feature transformation model, the first query representation is obtained through training of the feature transformation model, and the training of the feature transformation model includes: obtaining a first sample image and a second sample image with the same image style; determining style relevance features of the first sample image using the feature transformation model based on first sample image features of the first sample image, a first predetermined prompt, and the first query representation; and updating the feature transformation model and the first query representation by adding noise to and denoising the second sample image with the style relevance features of the first sample image and text indicating the image content of the second sample image as conditions.
[0107] In some embodiments, the first sample image and the second sample image are generated based on the same prompt information, and the prompt information indicates the style and content of the image to be generated.
[0108] In some embodiments, the training of the feature transformation model further includes: determining style relevance features of a third sample image using the feature transformation model based on third sample image features of the third sample image, the first predetermined prompt, and the first query representation; obtaining semantic relevance features of the third sample image; and updating the feature transformation model and the first query representation by reconstructing the third sample image with the style relevance features and the semantic relevance features as conditions.
[0109] Figure 8 A block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that Figure 8The illustrated electronic device 800 is merely exemplary and should not impose any limitation on the functionality and scope of the embodiments described herein. Figure 8 The illustrated electronic device 800 can be used to implement Figure 1 the electronic device 110.
[0110] As Figure 8 shown, the electronic device 800 is in the form of a general-purpose electronic device. The components of the electronic device 800 can include, but are not limited to, one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processing unit 810 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing ability of the electronic device 800.
[0111] The electronic device 800 generally includes multiple computer storage media. Such media can be any accessible media that can be obtained by the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable medium and can include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 800.
[0112] The electronic device 800 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 8 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 820 can include a computer program product 825 having one or more program modules that are configured to perform the various methods or actions of the various embodiments of the present disclosure.
[0113] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 800 can be implemented with a single computing cluster or multiple computing machines that are capable of communicating via a communication link. Thus, the electronic device 800 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0114] The input device 850 can be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 860 can be one or more output devices such as a display, speaker, printer, etc. The electronic device 800 can also communicate with one or more external devices (not shown) as needed via the communication unit 840, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 800, or communicate with any device that enables the electronic device 800 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0115] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, where the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.
[0116] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0117] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing apparatus, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing apparatus, and / or other devices to operate in a specific manner, so that the computer-readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0118] Computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to generate a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0119] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0120] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled in the art to understand the various implementations disclosed herein.
Claims
1. An image processing method, comprising: Obtaining first reference image features of a first reference image; Determining style-related features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query representation; And Generating a first target image based on the style-related features of the first reference image and a first input text indicating target image content, the first target image matching the image style of the first reference image and the target image content.
2. The method according to claim 1, wherein generating the first target image comprises: Obtaining first text features of the first input text; Generating first target image features based on the style-related features and the first text features by using a first model, wherein the first model comprises a plurality of processing layers, the plurality of processing layers are respectively used for processing features of corresponding sizes, and the style-related features are provided to a first number of processing layers with sizes larger than a first threshold size; And Generating the first target image based on the first target image features.
3. The method according to claim 2, wherein generating the first target image features by using the first model comprises: For a given processing layer among the first number of processing layers: Converting input image features of the given processing layer into query features; Generating key features and value features by converting the first text features and the style-related features; Generating output image features of the given processing layer based on the query features, the key features, and the value features.
4. The method according to claim 3, wherein the input image features are converted into the query features by using a first conversion unit, and generating the key features and the value features comprises: Converting the first text features into first intermediate text features by using a second conversion unit; Converting the first text features into second intermediate text features by using a third conversion unit; Converting the style-related features into first intermediate style features by using a fourth conversion unit; Converting the style-related features into second intermediate style features by using a fifth conversion unit; Combining the first intermediate text features and the first intermediate style features into the key features; and Combining the second intermediate text features and the second intermediate style features into the value features.
5. The method according to claim 4, wherein the first conversion unit, the second conversion unit, and the third conversion unit are obtained through a first training mode, and the fourth conversion unit and the fifth conversion unit are obtained through a second training mode different from the first training mode.
6. The method according to claim 1, further comprising: Obtaining second reference image features of a second reference image; Determining semantic-related features of the second reference image based on the second reference image features, a second predetermined prompt for semantic feature extraction, and a second query representation; And Generate a second target image based on the semantic correlation features and a second input text indicating the target image style, where the second target image matches the image content of the second reference image and the target image style.
7. The method according to claim 6, wherein generating the second target image comprises: Obtain second text features of the second input text; Based on the semantic correlation features and the second text features, generate second target image features according to a first model, where the first model includes a plurality of processing layers, the plurality of processing layers are respectively configured to process features of corresponding sizes, and the semantic correlation features are provided to a second number of processing layers with a size smaller than a second threshold size; And Generate the second target image based on the second target image features.
8. The method according to claim 1, wherein the style correlation features are determined using a feature transformation model, the first query representation is obtained through training of the feature transformation model, and the training of the feature transformation model includes: Obtain a first sample image and a second sample image with the same image style; Based on the first sample image features of the first sample image, the first predetermined prompt, and the first query representation, determine the style correlation features of the first sample image using the feature transformation model; Using the style correlation features of the first sample image and the text indicating the image content of the second sample image as conditions, update the feature transformation model and the first query representation by adding noise to and denoising the second sample image.
9. The method according to claim 8, wherein the first sample image and the second sample image are generated based on the same prompt information, and the prompt information indicates the style and content of the image to be generated.
10. The method according to claim 8, wherein the training of the feature transformation model further includes: Based on the third sample image features of a third sample image, the first predetermined prompt, and the first query representation, determine the style correlation features of the third sample image using the feature transformation model; Obtain the semantic correlation features of the third sample image; and Using the style correlation features and the semantic correlation features as conditions, update the feature transformation model and the first query representation by reconstructing the third sample image.
11. An apparatus for image processing, comprising: A first image feature extraction module configured to obtain first reference image features of a first reference image; A style correlation feature module configured to determine the style correlation features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query representation; And A first target image generation module configured to generate a first target image based on the style correlation features of the first reference image and a first input text indicating the target image content, where the first target image matches the image style of the first reference image and the target image content.
12. An electronic device, comprising: At least one processing unit; And At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.