Method and apparatus for image processing, and device and storage medium

Through predetermined prompts and query representations specific to style feature extraction, the style and semantic features of the reference image are decoupled to generate target images that conform to the input text, solving the problem of semantic inconsistency in the prior art and achieving the matching of style and semantics.

WO2025152653A1PCT designated stage expired Publication Date: 2025-07-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/138045
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-17
Filing Date
2024-12-10
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

In the prior art, when generating images with a specific style, it is difficult to effectively decouple the semantics of the reference image from the semantics of the input text, resulting in the semantics of the generated image that may contain the reference image and the semantics indicated by the input text.

Method used

Pre-determined prompts and query representations specific to style feature extraction are used to extract style-related features of the reference image through the feature transformation model, and the target image is generated in combination with the input text to achieve decoupling of style and semantics.

Benefits of technology

The generated target image has both the style of the reference image and conforms to the semantics of the input text, solving the problem of semantic inconsistency and improving the effect of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138045_24072025_PF_FP_ABST
    Figure CN2024138045_24072025_PF_FP_ABST
Patent Text Reader

Abstract

According to the embodiments of the present disclosure, provided are a method and apparatus for image processing, and a device and a storage medium. The method comprises: acquiring a first reference image feature of a first reference image; determining a style-related feature of the first reference image on the basis of the first reference image feature, a first predetermined prompt for style feature extraction, and a first query representation; and generating a first target image on the basis of the style-related feature of the first reference image and first input text indicating target image content, wherein the first target image matches both the image style of the first reference image and the target image content. In this way, a style feature of a reference image can be decoupled from a semantic feature thereof, thus facilitating the generation of an image that has the style of the reference image and also meets the semantics of input text.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for image processing

[0001] This application claims priority to the Chinese invention patent application entitled “Method, apparatus, device and storage medium for image processing” and application number 202410070352.8 filed on January 17, 2024, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for image processing. Background Art

[0003] In the field of computer vision (CV), various image generation techniques based on machine learning have made significant progress and are widely used. For example, in many application scenarios such as social networking, gaming, and image editing, it is desirable to generate and use images with specific styles. Machine learning-based image generation techniques can be used in such applications to improve image generation performance. Summary of the Invention

[0004] In a first aspect of the present disclosure, an image processing method is provided. The method includes: obtaining first reference image features of a first reference image; determining style-related features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query expression; and generating a first target image based on the style-related features of the first reference image and first input text indicating target image content, wherein the first target image matches the image style and target image content of the first reference image.

[0005] In a second aspect of the present disclosure, a device for image processing is provided. The device includes: a first image feature extraction module configured to obtain first reference image features of a first reference image; a style-related feature module configured to determine style-related features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query representation; and a first target image generation module configured to generate a first target image based on the style-related features of the first reference image and first input text indicating target image content, wherein the first target image matches the image style and target image content of the first reference image.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of the present disclosure, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions, which when executed by a device cause the device to implement the method of the first aspect.

[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG2 shows a schematic diagram of an example architecture for generating an image with a reference style and a target content according to some embodiments of the present disclosure;

[0013] FIG3 shows a schematic diagram of an example architecture for generating an image with reference content and a target style according to some embodiments of the present disclosure;

[0014] FIG4 shows a schematic diagram of an application of a cross-attention mechanism according to some embodiments of the present disclosure;

[0015] FIG5A shows a schematic diagram of a training task for style feature extraction according to some embodiments of the present disclosure;

[0016] FIG5B shows a schematic diagram of a training task for semantic feature extraction according to some embodiments of the present disclosure;

[0017] FIG5C shows a schematic diagram of a training task for image reconstruction according to some embodiments of the present disclosure;

[0018] FIG6 shows a flowchart of an image processing process according to some embodiments of the present disclosure;

[0019] FIG7 shows a block diagram of an apparatus for image processing according to some embodiments of the present disclosure; and

[0020] FIG8 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION

[0021] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0022] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0023] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0025] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0026] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0027] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.

[0028] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0029] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.

[0030] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In this article, "model" may also be referred to as "machine learning model", "machine learning network" or "network", and these terms are used interchangeably in this article. A model can also include different types of processing units or networks.

[0031] As used herein, a target image matches an image style, which may refer to the target image having a visually consistent or similar style to the image style. A target image matches an image content, which may refer to the target image containing visually the same or similar content to the target content.

[0032] Sample Environment

[0033] 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In environment 100 , an image processing system 120 , also referred to as system 120 , is deployed in electronic device 110 . Image processing system 120 is configured to generate a target image 105 based on input text 102 and a reference image 101 .

[0034] The reference image 101 can be used to provide or indicate image elements that are desired to be included in the target image 105, such as image style, image content, etc. Input text can be used to indicate another image element that is desired to be included in the target image 105. As used herein, the term "image element" can refer to various explicit or implicit elements that constitute an image. For example, an image element can include image style or image content. Examples of image style include, but are not limited to, watercolor, crayon, sketch, comics, etc. Examples of image content can include, but are not limited to, at least a portion of the foreground of an image, at least a portion of the background of an image, etc.

[0035] In some embodiments, image processing system 120 can generate an image having the same style as reference image 101. Hereinafter, the image style of reference image 101 is also referred to as the reference style. In such embodiments, input text 102 can indicate target image content. The generated target image 105 can match the image style and target image content of reference image 101. For example, target image 105 can have the image style of reference image 101 and include the target image content.

[0036] Alternatively or additionally, in some embodiments, the image processing system 120 may generate an image having the same content as the reference image 101. In such embodiments, the input text 102 may indicate a target image style. The generated target image 105 may match the target image style and the image content in the reference image 101. For example, the target image 105 may contain the image content in the reference image 101 and have the style indicated by the input text 102.

[0037] In the environment 100, the electronic device 110 can be any type of device with computing capabilities, including a terminal device or a server device. The terminal device can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. The server device can include, for example, a computing system / server, such as a mainframe, an edge computing node, an electronic device in a cloud environment, and the like.

[0038] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0039] As mentioned above, it is desirable to generate an image with the reference style of a reference image. To this end, in some conventional solutions, all or some parameters of a base image generation model (such a model can generate images based on input text) can be fine-tuned to generate a stylized image. However, this solution is time-consuming (e.g., at least minutes) and requires manual intervention.

[0040] In other conventional approaches, a two-stage encoder is used, consisting of a fixed backbone network and a trained head network, to extract features from a reference image. The features obtained by the encoder can be further combined with a denoising network. However, in this conventional approach, the training task is usually reconstruction, so the encoder learns a mixture of content and style. This can lead to the semantics of the reference image conflicting with the semantics indicated by the input text, resulting in the generated image potentially incorporating the semantics of the reference image.

[0041] To this end, embodiments of the present disclosure provide an improved approach for image generation. In this embodiment, predetermined prompts and query representations specific to style feature extraction are introduced to extract style-related features from a reference image. Based on these extracted style-related features and input text indicating the target image's content, a target image can be generated that matches the reference image's image style and content.

[0042] In the disclosed embodiments, feature extraction is performed using predetermined cues and query representations specific to stylistic feature extraction to decouple the stylistic features of a reference image from its semantic features. This alleviates the inconsistency between the semantics of the reference image and the input text. This advantageously allows for the generation of an image that possesses both the style of the reference image and the semantics of the input text.

[0043] Example image generation architecture

[0044] In some embodiments, input text 102 may indicate the desired content of a target image, also referred to as target image content or simply target content. In such embodiments, reference image 101 indicates an image style. Target image 105 has a reference style and target content. FIG2 illustrates a schematic diagram of an example architecture 200 for generating an image having a reference style and target content, according to some embodiments of the present disclosure. Architecture 200 may be implemented in system 120. Architecture 200 employs a two-stage structure for extracting image features.

[0045] As shown in FIG2 , the reference style of reference image 201 is style A. System 120 can obtain image features of reference image 201, also referred to as reference image features 202. For example, image encoder 220 can generate reference image features 202 based on reference image 201. Image encoder 220 can employ any suitable network structure, and embodiments of the present disclosure are not limited in this respect.

[0046] The reference image features 202 can be used as inputs to the feature transformation model 230. In addition to this, the feature transformation model 230 also has other inputs, including a predetermined prompt 204 (also referred to as a first predetermined prompt) for style feature extraction and a query representation 203 (also referred to as a first query representation) for style feature extraction. The predetermined prompt 204 can be any suitable text or prompt word. In this example, the predetermined prompt 204 is shown as the text "style" for illustrative purposes only and is not intended to be limiting. In an embodiment of the present disclosure, the first predetermined prompt is specific to style feature extraction and is not related to the reference image. That is, the first predetermined prompt does not change with the reference image. The feature transformation model 230 can be implemented using a network of any suitable structure. As an example and not intended to be limiting, the feature transformation model 230 can be a query transformer.

[0047] Query representation 203 is learnable. It can be obtained during the training of feature transformation model 230. For example, query representation 203 can be initialized and, during the training of feature transformation model 230, updated together until the training is complete and solidified. Such query representation 203 can be used to instruct feature transformation model 230 to extract features related to image style. As shown in FIG2 , query representation 203 can be a vectorized representation of any suitable dimension. Query representation 203 is obtained for style feature extraction.

[0048] The feature transformation model 230 can generate style-related features 205 of the reference image 201 based on the reference image features 202, the first predetermined hint 204, and the first query representation 203. In other words, under the hint or guidance of the first predetermined hint 204 and the first query representation 203, the feature transformation model 230 can extract features related to the image style from the reference image 201.

[0049] Next, the system 120 can generate a target image 210 based on the style-related features 205 and the input text 209 indicating the content of the target image. The input text 209 can indicate the content that the target image is expected to contain in any suitable characters, such as but not limited to people, animals, objects, scenery, buildings, etc. In this example, the target content indicated by the input text 209 is "panda", but this is merely exemplary and not intended to be limiting. As shown in Figure 2, the target image 210 has the style of the reference image 201 (i.e., style A) and contains the content indicated by the input text 209 (i.e., panda).

[0050] System 120 can employ any suitable algorithm or model to generate the target image. As shown in FIG2 , in some embodiments, model 240, also referred to as a first model, can be employed. Model 240 can be implemented using any suitable network architecture. For example, model 240 can be a diffusion model that can perform multiple denoising steps. Model 240 can also be of other types, such as a generative adversarial network.

[0051] Specifically, the text encoder 250 can generate text features 207 of the input text 209. The text features 207 can be provided to the model 240 as text conditions. In addition, the model 240 can also receive style-related features 205. The model 240 can generate target image features 208 based on the text features 207 and the style-related features 205. The target image features 208 can then be used to generate the target image 210. For example, the system 120 can utilize an image decoder (not shown) to convert the target image features 208 into the target image 210. The embodiments of the present disclosure are not limited in how to convert the target image features into the target image. In addition, in each step (e.g., a denoising step) performed using the model 240, the model 240 also receives the image features output by the previous step as input.

[0052] In some embodiments, the model 240 may include multiple processing layers, each of which is used to process features of corresponding sizes. In such an embodiment, the style-related features 205 are provided to a first number of processing layers whose sizes are greater than a first threshold size. For example, the model 240 may be a denoising U-shaped network. Exemplarily, as shown in FIG2 , the processing layers close to the model input and output are used to process features of larger sizes, while the processing layers close to the middle of the model are used to process features of smaller sizes. Accordingly, the style-related features 205 are provided to the processing layers close to the model input and output. As an example and without any limitation, it is assumed that the model 240 includes 16 processing layers, which are numbered 0, 1, ..., 15 in sequence. The style-related features 205 can be provided to layers 0 to 3 and layers 9 to 15. The combination of style-related features and text features will be described below with reference to FIG4 .

[0053] Features with larger sizes include more details of the image and correspond to the fine processing layer. It is generally believed that such fine processing layers are mainly responsible for the generation of style-related elements such as color and structure of the image. Features with smaller sizes include high-level semantic information of the image and correspond to the coarser processing layer. It is generally believed that such coarser processing layers are mainly responsible for the semantic generation of the image. Therefore, by injecting style-related features only into the fine processing layer and not into the coarser processing layer, the style and semantics of the reference image can be further decoupled. In this embodiment, it is advantageous to enable the model 240 to focus more on the style of the reference image in image generation.

[0054] The above describes an example implementation of generating an image with a reference style and a target content. The following describes an example implementation of generating an image with a target style and a reference content.

[0055] In some embodiments, the input text 102 may indicate the style that the target image is expected to have, also referred to as the target image style or simply the target style. In this embodiment, the reference image 101 indicates the image content, also referred to as the reference content. The target image 105 has the reference content and the target style. Figure 3 shows a schematic diagram of an example architecture 300 for generating an image with reference content and a target style according to some embodiments of the present disclosure. The architecture 300 can be implemented in the system 120. For the extraction of image features, the architecture 300 adopts a two-stage structure. Note that the text encoder 350 in Figure 3 may be the same as or different from the text encoder 250 in Figure 2, and the image encoder 320 in Figure 3 may be the same as or different from the image encoder 220 in Figure 2. The embodiments of the present disclosure are not limited in this respect.

[0056] As shown in Figure 3, the reference style of reference image 301 is Style A, and the image content is a smiling face. System 120 can obtain image features of reference image 301, also referred to as reference image features 302. For example, image encoder 320 can generate reference image features 302 based on reference image 301. Image encoder 320 can employ any suitable network structure, and embodiments of the present disclosure are not limited in this respect.

[0057] The reference image features 302 can be used as inputs to the feature transformation model 330. In addition to this, the feature transformation model 330 also has other inputs, including a predetermined prompt 304 (also referred to as a second predetermined prompt) for semantic feature extraction and a query representation 303 (also referred to as a second query representation) for semantic feature extraction. The predetermined prompt 304 can be any suitable text or prompt word. In this example, the predetermined prompt 304 is shown as the text "content" for illustrative purposes only and is not intended to be limiting. In an embodiment of the present disclosure, the second predetermined prompt is specific to semantic feature extraction and is not related to the reference image. That is, the second predetermined prompt does not change with the reference image. The feature transformation model 230 can be implemented using a network of any suitable structure. As an example and not intended to be limiting, the feature transformation model 230 can be a query transformer.

[0058] Query representation 303 is learnable. Query representation 303 can be obtained during the training of feature transformation model 330. For example, query representation 303 can be initialized, and during the training of feature transformation model 330, query representation 303 can be updated together until the training is completed and solidified. Such query representation 303 can be used to instruct feature transformation model 330 to extract features related to image content, that is, semantically relevant features. As shown in Figure 3, query representation 303 can be a vectorized representation of any suitable dimension. Query representation 303 is obtained for semantic feature extraction.

[0059] Feature transformation model 330 can generate semantically relevant features 305 of reference image 301 based on reference image features 302, second predetermined hint 304, and second query representation 303. In other words, under the hint or guidance of second predetermined hint 304 and second query representation 303, feature transformation model 330 can extract features that are semantically relevant to the reference image 301.

[0060] Next, the system 120 can generate a target image 310 based on the semantically relevant features 305 and the input text 309 indicating the style of the target image. The input text 309 can indicate the desired style of the target image using any suitable characters. In this example, the target style indicated by the input text 309 is "Style B", but this is merely exemplary and not intended to be limiting. As shown in FIG3 , the target image 310 has the content of the reference image 301 (i.e., a smiley face) and includes the style indicated by the input text 309 (i.e., Style B).

[0061] Similar to what is described with reference to FIG2 , the system 120 can use any suitable algorithm or model to generate the target image 310. As shown in FIG3 , in some embodiments, the model 240 can be used. Specifically, the text encoder 350 can generate text features 307 of the input text 309. The text features 307 can be provided to the model 240 as text conditions. In addition, the model 240 can also receive semantically related features 305. The model 240 can generate target image features 308 based on the text features 307 and the semantically related features 305. The target image features 308 can then be used to generate the target image 310. For example, the system 120 can use an image decoder (not shown) to convert the target image features 308 into the target image 310. The embodiments of the present disclosure are not limited in how to convert the target image features into the target image. In addition, in each step (e.g., the denoising step) performed using the model 240, the model 240 also receives the image features output by the previous step as input.

[0062] In some embodiments, the model 240 may include multiple processing layers, each of which is used to process features of corresponding sizes. In this embodiment, the semantically relevant features 305 are provided to a second number of processing layers whose size is less than a second threshold size. The second threshold size may be the same as or different from the first threshold size described above. For example, the model 240 may be a denoising U-shaped network. Exemplarily, as shown in Figure 3, the processing layers close to the model input and output are used to process features with larger sizes, while the processing layers close to the middle of the model are used to process features with smaller sizes. Accordingly, the semantically relevant features 305 are provided to the processing layers close to the middle of the model. As an example and without any intention of limitation, it is assumed that the model 240 includes 16 processing layers, which are numbered 0, 1, ..., 15 in sequence. The semantically relevant features 305 can be provided to 4 to 8 layers. The combination of semantically relevant features and text features will be described below with reference to Figure 4.

[0063] As described with reference to FIG2 , features with larger sizes include more details of the image and correspond to the fine processing layer. It is generally believed that such a fine processing layer is mainly responsible for the generation of style-related elements such as color and structure of the image. Features with smaller sizes include high-level semantic information of the image and correspond to the coarser processing layer. It is generally believed that such a coarser processing layer is mainly responsible for the semantic generation of the image. Therefore, by injecting semantically relevant features only into the coarser processing layer and not into the fine processing layer, the style and semantics of the reference image can be further decoupled. In this embodiment, the model 240 can be made to focus more on the content of the reference image in image generation.

[0064] Combination of reference image features and text features

[0065] As mentioned above, in some embodiments, style-related features and semantic-related features can be provided to certain processing layers of model 240. Hereinafter, for the purpose of illustration only, style-related features and semantic-related features are collectively referred to as image-related features. In such processing layers, image features and text features of the input text can be used as conditions for target image generation. In some embodiments, model 240 can be based on a cross-attention mechanism. Accordingly, in the processing layers, cross-attention related to image-related features and text features can be applied.

[0066] FIG4 shows a schematic diagram of the application of the cross attention mechanism according to some embodiments of the present disclosure. In FIG4 , the corresponding size is shown below each feature. As shown in FIG4 , in a certain processing layer, the input image feature 411 of the processing layer (which is represented by Z and can come from the previous processing layer or the previous denoising step) is converted into a query feature 421. By converting the text feature 412 (which is represented by C t ) and image related features 413 (which are represented by C i The processing layer generates key features 431 and value features 432 based on the query features 421, key features 431, and value features 432. As shown in Figure 4, the dot product of query features 421 and key features 431 produces an attention map. The matrix multiplication of the attention map and value features 432 yields the output image features 450.

[0067] In some embodiments, the parameters of such a processing layer may be obtained through different training modes, which may include, for example, pre-training and fine-tuning. As shown in Figure 4, the processing layer may include a first conversion unit 401 for converting an input image feature 411 into a query feature 421. The processing layer may also include a second conversion unit 402, a third conversion unit 403, a fourth conversion unit 404, and a fifth conversion unit 405. The second conversion unit 402 is used to convert a text feature 412 into a first intermediate text feature 422 as part of a key feature. The third conversion unit 403 is used to convert a text feature 412 into a second intermediate text feature 423 as part of a value feature. The fourth conversion unit 404 is used to convert an image-related feature 413 into a first intermediate image-related feature 424 as part of a key feature. The fifth conversion unit 405 is used to convert an image-related feature 413 into a second intermediate image-related feature 425 as part of a value feature.

[0068] The first intermediate text feature 422 and the first intermediate image-related feature 424 may be combined into a key feature 431, for example, they may be concatenated into the key feature 431. Similarly, the second intermediate text feature 423 and the second intermediate image-related feature 425 may be combined into a value feature 432, for example, they may be concatenated into the value feature 432.

[0069] In some embodiments, the first conversion unit 401, the second conversion unit 402, and the third conversion unit 403 are obtained through a first training mode, and the fourth conversion unit 404 and the fifth conversion unit 405 are obtained through a second training mode different from the first training mode. For example, the first conversion unit 401, the second conversion unit 402, and the third conversion unit 403 can be obtained through pre-training, and the fourth conversion unit 404 and the fifth conversion unit 405 can be obtained through fine-tuning. In this way, additional units can be added to the pre-trained base model, so that the additional units can be fine-tuned without having to train the parts included in the base model. This can effectively reduce training costs.

[0070] In the embodiments described above, the semantic information and style information in the reference image are decoupled. This can be considered a dual decoupled representation extraction (DDRE). Specifically, predetermined prompt text specific to style feature extraction (e.g., "style") and predetermined prompt text specific to semantic feature extraction (e.g., "content") are input into the feature transformation model, allowing the feature transformation model to obtain image features aligned with the prompt text. In other words, style-related features and semantic-related features can be obtained separately.

[0071] Training tasks and samples

[0072] In order to achieve the above-mentioned dual-decoupled representation extraction, various suitable training tasks can be performed during training.

[0073] In some embodiments, a training task for style feature extraction, also known as a style representation extraction (STRE) task, can be performed. In the STRE task, training is performed using pairs of different sample images with the same style. Figure 5A shows a schematic diagram of a training task 500A for style feature extraction according to some embodiments of the present disclosure.

[0074] In training task 500A, a pair of sample images 511 and 520 have the same style, for example, style A. Sample image 511 serves as a reference, and sample image 520 serves as a target. During training, a first query representation 203 is initialized. The feature transformation model 230 can generate style-related features 515 for the sample image 511 based on first sample image features of the sample image 511 (e.g., the output of the image encoder 220), a first predetermined prompt 204, and the first query representation 203. Conditioned on the style-related features 515 of the sample image 511 and text 519 indicating the image content of the sample image 520, the feature transformation model 230 and the initialized first query representation 203 are updated by denoising and denoising the sample image 520. When the predetermined conditions are met, updating of the feature transformation model 230 and the first query representation 203 ceases. In some embodiments, if the model 240 includes a trainable transformation unit (e.g., as described with reference to FIG. 4 ), such a transformation unit is also updated during training.

[0075] The STRE task described above is a non-reconstruction task. This training task facilitates the decoupling of the style and semantics of the reference image, while also ensuring that the image information (in this case, the style information) is not too strong during training, thereby drowning out the input text.

[0076] In some embodiments, a training task for semantic feature extraction, also known as a semantic representation extraction (SERE) task, can be performed. In a SERE task, training is performed using pairs of different sample images with the same semantic meaning. Figure 5B shows a schematic diagram of a training task 500B for semantic feature extraction according to some embodiments of the present disclosure.

[0077] In training task 500B, a pair of sample images 521 and 530 have the same content, such as musical notes. Sample image 521 serves as a reference, and sample image 530 serves as a target. During training, the second query representation 303 is initialized. The feature transformation model 330 can generate semantically relevant features 525 for the sample image 521 based on the sample image features of the sample image 531 (e.g., the output of the image encoder 320), the second predetermined prompt 304, and the second query representation 303. The feature transformation model 330 and the initialized second query representation 303 are updated by denoising and adding noise to the sample image 530, conditioned on the semantically relevant features 525 of the sample image 521 and text 529 indicating the style of the sample image 530. When the predetermined conditions are met, updating of the feature transformation model 330 and the second query representation 303 ceases. In some embodiments, if the model 240 includes a trainable transformation unit (e.g., as described with reference to FIG. 4 ), such a transformation unit is also updated during training.

[0078] The SERE task described above is a non-reconstruction task. This training task helps decouple the style and semantics of the reference image, while also ensuring that the image information (in this case, semantic information) is not too strong during training, thereby drowning out the input text.

[0079] In some embodiments, in order to avoid image information omission caused by non-reconstruction tasks, a reconstruction task may be additionally performed. Figure 5C shows a schematic diagram of a training task 500C for image reconstruction according to some embodiments of the present disclosure.

[0080] 5C , the upper half branch is the style branch. During training, the feature transformation model 230 can generate style-related features 533 of the sample image 531 based on the sample image features of the sample image 531 (e.g., the output of the image encoder 220 ), the first predetermined hint 204 , and the first query representation 203 .

[0081] In the semantic branch, semantically relevant features 532 of the sample image 531 may be obtained. For example, the semantically relevant features 532 may be obtained using the feature transformation model 330. As shown, the feature transformation model 330 may generate the semantically relevant features 532 of the sample image 531 based on the sample image features of the sample image 531 (e.g., the output of the image encoder 320), the second predetermined hint 304, and the second query representation 303.

[0082] In this way, the feature transformation model 230 and the first query representation 203 can be updated by reconstructing the sample image 531, taking the style-related features 533 and the semantic-related features 532 as conditions. Accordingly, the feature transformation model 330 and the second query representation 303 are also updated. In some embodiments, if the model 240 includes a trainable transformation unit (e.g., as described with reference to FIG. 4 ), such a transformation unit is also updated as part of the training based on the reconstruction task.

[0083] As described above, during the execution of training tasks 500A, 500B, and 500C, feature transformation model 230, feature transformation model 330, first query representation 203, and second query representation 303 are trainable, model 240 is partially trainable (as described in reference FIG. 4 ), and the remaining models are frozen.

[0084] In some embodiments, in order to support the above-mentioned training tasks, corresponding sample sets can be created. For example, a group of style words indicating different styles and a group of subject words indicating different contents can be determined. Through the combination of style words and subject words, multiple prompt information can be obtained, for example, multiple prompt words. Each prompt information indicates the style and content to be generated. For example, the same prompt information can generate multiple sample images. In some embodiments, any pair of sample images among the multiple sample images generated by the same prompt information can be used for the training task of style feature extraction. For example, sample image 511 and sample image 520 in Figure 5A are generated by the same prompt information. Compared with image pairs using the same style words but different subject words, image pairs generated using the same prompt information can obtain better stylization effects.

[0085] In some embodiments, sample image pairs with the same subject word and different style words can be used for semantic feature extraction training tasks. For example, sample image 521 and sample image 530 in FIG5B can be generated based on the same subject word and different style words.

[0086] Example procedures, devices, and equipment

[0087] FIG6 shows a flowchart of a process 600 of image processing according to some embodiments of the present disclosure. The process 600 may be implemented at the electronic device 110.

[0088] In block 610 , the electronic device 110 acquires a first reference image feature of a first reference image.

[0089] In block 610 , the electronic device 110 determines style-related features of the first reference image based on the first reference image features, a first predetermined cue for style feature extraction, and a first query representation.

[0090] At block 610 , the electronic device 110 generates a first target image based on the style-related features of the first reference image and first input text indicating target image content. The first target image matches the image style of the first reference image and the target image content.

[0091] In some embodiments, generating a first target image includes: obtaining a first text feature of a first input text; generating a first target image feature using a first model based on style-related features and the first text feature, wherein the first model includes multiple processing layers, the multiple processing layers are respectively used to process features of corresponding sizes, and the style-related features are provided to a first number of processing layers having a size greater than a first threshold size; and generating a first target image based on the first target image feature.

[0092] In some embodiments, generating a first target image feature using a first model includes: for a given processing layer among the first number of processing layers: converting the input image feature of the given processing layer into a query feature; generating a key feature and a value feature by converting the first text feature and the style-related feature; and generating an output image feature of the given processing layer based on the query feature, the key feature, and the value feature.

[0093] In some embodiments, the input image feature is converted into a query feature using a first conversion unit, and generating a key feature and a value feature includes: converting the first text feature into a first intermediate text feature using a second conversion unit; converting the first text feature into a second intermediate text feature using a third conversion unit; converting the style-related feature into a first intermediate style feature using a fourth conversion unit; converting the style-related feature into a second intermediate style feature using a fifth conversion unit; combining the first intermediate text feature and the first intermediate style feature into a key feature; and combining the second intermediate text feature and the second intermediate style feature into a value feature.

[0094] In some embodiments, the first transformation unit, the second transformation unit, and the third transformation unit are obtained through a first training pattern, and the fourth transformation unit and the fifth transformation unit are obtained through a second training pattern different from the first training pattern.

[0095] In some embodiments, process 600 also includes: obtaining second reference image features of the second reference image; determining semantically relevant features of the second reference image based on the second reference image features, a second predetermined prompt for semantic feature extraction, and a second query representation; and generating a second target image based on the semantically relevant features and a second input text indicating the style of the target image, the second target image matching the image content and the target image style of the second reference image.

[0096] In some embodiments, generating a second target image includes: obtaining second text features of a second input text; generating second target image features according to a first model based on semantically related features and the second text features, wherein the first model includes multiple processing layers, the multiple processing layers are respectively used to process features of corresponding sizes, and the semantically related features are provided to a second number of processing layers having a size less than a second threshold size; and generating a second target image based on the second target image features.

[0097] In some embodiments, style-related features are determined using a feature transformation model, the first query representation is obtained through training of the feature transformation model, and the training of the feature transformation model includes: obtaining a first sample image and a second sample image having the same image style; determining the style-related features of the first sample image using the feature transformation model based on the first sample image features, a first predetermined prompt, and a first query representation; and updating the feature transformation model and the first query representation by denoising and de-noising the second sample image, using the style-related features of the first sample image and text indicating the image content of the second sample image as conditions.

[0098] In some embodiments, the first sample image and the second sample image are generated based on the same prompt information, where the prompt information indicates the style and content of the image to be generated.

[0099] In some embodiments, the training of the feature transformation model also includes: determining the style-related features of the third sample image using the feature transformation model based on the third sample image features of the third sample image, the first predetermined prompt and the first query representation; obtaining the semantic-related features of the third sample image; and updating the feature transformation model and the first query representation by reconstructing the third sample image with the style-related features and the semantic-related features as conditions.

[0100] 7 shows a schematic structural block diagram of an apparatus 700 for image processing according to certain embodiments of the present disclosure. Apparatus 700 may be implemented as or included in electronic device 110. Each module / component in apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.

[0101] As shown, apparatus 700 includes a first image feature extraction module 710 configured to obtain first reference image features of a first reference image. Apparatus 700 also includes a style-related feature module 720 configured to determine style-related features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query representation. Apparatus 700 also includes a first target image generation module 730 configured to generate a first target image based on the style-related features of the first reference image and first input text indicating target image content, wherein the first target image matches the image style and target image content of the first reference image.

[0102] In some embodiments, the first target image generation module 730 is further configured to: obtain a first text feature of a first input text; generate a first target image feature using a first model based on the style-related feature and the first text feature, wherein the first model includes multiple processing layers, the multiple processing layers are respectively used to process features of corresponding sizes, and the style-related features are provided to a first number of processing layers having a size greater than a first threshold size; and generate a first target image based on the first target image feature.

[0103] In some embodiments, the first target image generation module 730 is further configured to: for a given processing layer in the first number of processing layers: convert the input image features of the given processing layer into query features; generate key features and value features by converting the first text features and style-related features; and generate output image features of the given processing layer based on the query features, key features, and value features.

[0104] In some embodiments, the input image feature is converted into a query feature using a first conversion unit, and the first target image generation module 730 is further configured to: convert the first text feature into a first intermediate text feature using a second conversion unit; convert the first text feature into a second intermediate text feature using a third conversion unit; convert the style-related feature into a first intermediate style feature using a fourth conversion unit; convert the style-related feature into a second intermediate style feature using a fifth conversion unit; combine the first intermediate text feature and the first intermediate style feature into a key feature; and combine the second intermediate text feature and the second intermediate style feature into a value feature.

[0105] In some embodiments, the first transformation unit, the second transformation unit, and the third transformation unit are obtained through a first training pattern, and the fourth transformation unit and the fifth transformation unit are obtained through a second training pattern different from the first training pattern.

[0106] In some embodiments, the device 700 also includes: a second image feature extraction module, configured to obtain second reference image features of the second reference image; a semantic related feature module, configured to determine the semantic related features of the second reference image based on the second reference image features, a second predetermined prompt for semantic feature extraction and a second query representation; and a second target image generation module, configured to generate a second target image based on the semantic related features and a second input text indicating the style of the target image, the second target image matching the image content and the target image style of the second reference image.

[0107] In some embodiments, the second target image generation module is further configured to: obtain second text features of the second input text; generate second target image features according to the first model based on the semantically related features and the second text features, wherein the first model includes multiple processing layers, the multiple processing layers are respectively used to process features of corresponding sizes, and the semantically related features are provided to a second number of processing layers whose sizes are smaller than a second threshold size; and generate a second target image based on the second target image features.

[0108] In some embodiments, style-related features are determined using a feature transformation model, the first query representation is obtained through training of the feature transformation model, and the training of the feature transformation model includes: obtaining a first sample image and a second sample image having the same image style; determining the style-related features of the first sample image using the feature transformation model based on the first sample image features, a first predetermined prompt, and a first query representation; and updating the feature transformation model and the first query representation by denoising and de-noising the second sample image, using the style-related features of the first sample image and text indicating the image content of the second sample image as conditions.

[0109] In some embodiments, the first sample image and the second sample image are generated based on the same prompt information, where the prompt information indicates the style and content of the image to be generated.

[0110] In some embodiments, the training of the feature transformation model also includes: determining the style-related features of the third sample image using the feature transformation model based on the third sample image features of the third sample image, the first predetermined prompt and the first query representation; obtaining the semantic-related features of the third sample image; and updating the feature transformation model and the first query representation by reconstructing the third sample image with the style-related features and the semantic-related features as conditions.

[0111] FIG8 shows a block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 800 shown in FIG8 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 800 shown in FIG8 can be used to implement the electronic device 110 of FIG1 .

[0112] As shown in FIG8 , electronic device 800 is a general-purpose electronic device. Components of electronic device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit 810 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 800.

[0113] The electronic device 800 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 800.

[0114] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG8 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0115] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 800 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 800 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0116] The input device 850 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 860 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 800 may also communicate with one or more external devices (not shown) via the communication unit 840 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 800, or with any device that allows the electronic device 800 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0117] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0118] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0119] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0120] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0121] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0122] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. An image processing method, comprising: Obtaining first reference image features of a first reference image; Determining style-related features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query representation; And Generating a first target image based on the style-related features of the first reference image and a first input text indicating target image content, the first target image matching the image style of the first reference image and the target image content.

2. The method according to claim 1, wherein generating the first target image comprises: Obtaining first text features of the first input text; Generating first target image features based on the style-related features and the first text features by using a first model, wherein the first model comprises a plurality of processing layers, the plurality of processing layers are respectively used for processing features of corresponding sizes, and the style-related features are provided to a first number of processing layers having a size larger than a first threshold size; And Generating the first target image based on the first target image features.

3. The method according to claim 2, wherein generating the first target image features by using the first model comprises: For a given processing layer among the first number of processing layers: Converting input image features of the given processing layer into query features; Generating key features and value features by transforming the first text features and the style-related features; Generating output image features of the given processing layer based on the query features, the key features, and the value features.

4. The method according to claim 3, wherein the input image features are converted into the query features by using a first conversion unit, and generating the key features and the value features comprises: Converting the first text features into first intermediate text features by using a second conversion unit; Converting the first text features into second intermediate text features by using a third conversion unit; Converting the style-related features into first intermediate style features by using a fourth conversion unit; Converting the style-related features into second intermediate style features by using a fifth conversion unit; Combining the first intermediate text features and the first intermediate style features into the key features; and Combining the second intermediate text features and the second intermediate style features into the value features.

5. The method according to claim 4, wherein the first conversion unit, the second conversion unit, and the third conversion unit are obtained through a first training mode, and the fourth conversion unit and the fifth conversion unit are obtained through a second training mode different from the first training mode.

6. The method according to claim 1, further comprising: Obtaining second reference image features of a second reference image; Determining semantic-related features of the second reference image based on the second reference image features, a second predetermined prompt for semantic feature extraction, and a second query representation; And Generate a second target image based on the semantic correlation features and a second input text indicating the target image style, where the second target image matches the image content of the second reference image and the target image style.

7. The method according to claim 6, wherein generating the second target image comprises: Obtain second text features of the second input text; Based on the semantic correlation features and the second text features, generate second target image features according to a first model, where the first model includes a plurality of processing layers, the plurality of processing layers are respectively configured to process features of corresponding sizes, and the semantic correlation features are provided to a second number of processing layers with sizes smaller than a second threshold size; And Generate the second target image based on the second target image features.

8. The method according to claim 1, wherein the style correlation features are determined using a feature transformation model, the first query representation is obtained through training of the feature transformation model, and the training of the feature transformation model includes: Obtain a first sample image and a second sample image with the same image style; Based on the first sample image features of the first sample image, the first predetermined prompt, and the first query representation, determine the style correlation features of the first sample image using the feature transformation model; Using the style correlation features of the first sample image and the text indicating the image content of the second sample image as conditions, update the feature transformation model and the first query representation by adding noise to and denoising the second sample image.

9. The method according to claim 8, wherein the first sample image and the second sample image are generated based on the same prompt information, and the prompt information indicates the style and content of the image to be generated.

10. The method according to claim 8, wherein the training of the feature transformation model further includes: Based on the third sample image features of a third sample image, the first predetermined prompt, and the first query representation, determine the style correlation features of the third sample image using the feature transformation model; Obtain the semantic correlation features of the third sample image; and Using the style correlation features and the semantic correlation features as conditions, update the feature transformation model and the first query representation by reconstructing the third sample image.

11. An apparatus for image processing, comprising: A first image feature extraction module configured to obtain first reference image features of a first reference image; A style correlation feature module configured to determine the style correlation features of the first reference image based on the first reference image features, a first predetermined prompt for style feature extraction, and a first query representation; And A first target image generation module configured to generate a first target image based on the style correlation features of the first reference image and a first input text indicating the target image content, where the first target image matches the image style of the first reference image and the target image content.

12. An electronic device, comprising: At least one processing unit; And at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.

14. A computer program product tangibly stored in a computer storage medium and including computer-executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 - 10.

Citation Information

Patent Citations

  • Image processing method and device and electronic equipment

    CN113393371A

  • Stylized image generation method and device, computer equipment and storage medium

    CN116012488A

  • Video generation method and server

    CN116233491A

  • Image generation method and device, electronic equipment, storage medium and program product

    CN116958323A

  • Style migration image processing method and device, electronic equipment and storage medium

    CN118447262A