Image processing method and device, equipment and storage medium

By using the prompt information obtained by training in the diffusion model, a feature representation suitable for visual perception tasks is generated, which solves the problems of large amount of computing and general-purpose insufficient in the prior art, and achieves more efficient and flexible image processing.

CN120198911APending Publication Date: 2025-06-24FACE CUTE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311794444.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When adapted to general visual perception tasks, the existing diffusion model relies on additional text models, which increases the computational volume, and is insufficient in general style, making it difficult to generate appropriate prompt features through category labels or explanatory text.

Method used

The prompt information corresponding to the target visual task is obtained by training the diffusion model, and the first feature representation of the input image is used to generate the second feature representation, thereby realizing the result generation of the target visual task.

Benefits of technology

This reduces the computational cost of image processing, improves the adaptability of the diffusion model to various visual tasks, and avoids dependence on additional text models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198911A_ABST
    Figure CN120198911A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring a first feature representation of a to-be-processed input image; based on the first feature representation and prompt information corresponding to the target visual task, using a diffusion model to generate a second feature representation of the input image, the prompt information being obtained through training of the diffusion model; and generating a result of the input image about the target visual task based on the second feature representation. In this way, on one hand, the image processing cost is reduced, and on the other hand, the diffusion model can better adapt to various visual tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for image processing. Background Art

[0002] In the field of computer vision (CV), various image processing technologies based on machine learning have been significantly developed and have wide applications. Computer vision can be applied to a variety of different image processing tasks, for example, image generation tasks and visual perception tasks. In visual perception tasks, it is necessary to process a target image to obtain desired perception results, such as classification results, segmentation results, and so on. Summary of the Invention

[0003] In a first aspect of the present disclosure, there is provided an image processing method. The method includes: obtaining a first feature representation of an input image to be processed; generating a second feature representation of the input image by using a diffusion model based on the first feature representation and hint information corresponding to a target visual task, where the hint information is obtained through training of the diffusion model; and generating a result of the input image regarding the target visual task based on the second feature representation.

[0004] In a second aspect of the present disclosure, there is provided an apparatus for image processing. The apparatus includes: a first representation obtaining module configured to obtain a first feature representation of an input image to be processed; a second representation obtaining module configured to generate a second feature representation of the input image by using a diffusion model based on the first feature representation and hint information corresponding to a target visual task, where the hint information is obtained through training of the diffusion model; and a result generating module configured to generate a result of the input image regarding the target visual task based on the second feature representation.

[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device includes at least one processing unit; and at least one memory, where at least one memory is coupled to at least one processing unit and stores instructions for execution by at least one processing unit. The instructions, when executed by at least one processing unit, cause the device to execute the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0007] It should be understood that the content described in this content part is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings

[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0009] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0010] Figure 2 A schematic diagram showing an image processing architecture for a target visual task according to some embodiments of the present disclosure;

[0011] Figure 3 A schematic diagram showing an image processing flow for a target visual task according to some embodiments of the present disclosure;

[0012] Figure 4 A schematic diagram showing the conversion of multi-scale features according to some embodiments of the present disclosure;

[0013] Figure 5 A flowchart showing the process of image processing according to some embodiments of the present disclosure;

[0014] Figure 6 A block diagram showing a device for image processing according to some embodiments of the present disclosure; and

[0015] Figure 7 A block diagram showing a device capable of implementing multiple embodiments of the present disclosure. Detailed Embodiments

[0016] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0017] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0018] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0019] It is understood that the above notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manners of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manners of the present disclosure.

[0020] It is understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0021] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0022] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. In addition, the embodiments described in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different sections / subsections in any manner.

[0023] In this document, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.

[0024] In the description of the embodiments of the present disclosure, the term "including" and its like terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter. The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0025] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, for a given input, a corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In this document, "model" can also be referred to as "machine learning model", "machine learning network", or "network", and these terms are used interchangeably in this document. A model can also include different types of processing units or networks.

[0026] Example environment

[0027] Figure 1 FIG. shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In environment 100, an image processing system 120, also simply referred to as system 120, is deployed in an electronic device 110. The image processing system 120 is configured to perform a target vision task on an input image 101 to generate a corresponding task result 105.

[0028] The target vision task can include any suitable type of visual perception task that performs visual perception on the input image to obtain a corresponding result. The visual perception task is a non-generative task and can include, but is not limited to, image segmentation, image classification, object detection, key point detection, depth estimation, etc. As an example, when the image processing system 120 is used for image classification, the task result 105 can be the classification of the object in the input image 101. As another example, when the image processing system 120 is used for depth estimation, the task result 105 can be a depth map corresponding to the input image 101. It should be understood that the above-listed visual perception tasks are only exemplary and are not intended to limit the scope of the present disclosure. The image processing system 120 can be applied to any type of perception task.

[0029] In environment 100, the electronic device 110 can be any type of device with computing capabilities, including a terminal device or a server device. The terminal device can be any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server device can include, for example, a computing system / server, such as a mainframe, an edge computing node, an electronic device in a cloud environment, and so on.

[0030] It should be understood that the structure and functions of the environment 100 are described only for exemplary purposes, without implying any limitation to the scope of the present disclosure.

[0031] As briefly mentioned above, CV technology has been applied to various image processing tasks, including visual perception tasks. Generative pre-training schemes are one of the technical routes to achieve general visual perception. Currently, diffusion models have been applied to image generation tasks. By training on a large amount of image-text pair data, diffusion models have achieved high-quality text-to-image generation. Therefore, a scheme has been proposed to adapt diffusion models to general visual perception tasks.

[0032] However, existing adaptation schemes rely on additional text information related to the processed images, such as class labels of objects in the images or captions of the images, etc. Such additional text information needs to cooperate with a pre-trained text model to generate prompt features. The generated prompt features will be used in the diffusion model.

[0033] Such adaptation schemes have some problems. For example, relying on additional text models, etc., to generate prompt features increases the computational cost. Also, the generality of such adaptation schemes is insufficient. Visual tasks such as depth estimation cannot generate appropriate prompt features through class labels or captions.

[0034] To this end, embodiments of the present disclosure propose an improved scheme for image processing. According to various embodiments of the present disclosure, during the process of using a diffusion model to perform a target visual task, prompt information corresponding to the target visual task is utilized. Such prompt information is obtained through the training of the diffusion model. Specifically, based on a first feature representation of an input image and the prompt information, the diffusion model is used to generate a second feature representation of the input image. Then, based on the second feature representation, a result of the input image regarding the target visual task is generated.

[0035] According to embodiments of the present disclosure, when using a diffusion model to perform a visual task, it is not necessary to rely on an additional text model, but rather use the prompt information obtained through the training of the diffusion model. In this way, on the one hand, the image processing cost is reduced, and on the other hand, the diffusion model can better adapt to various visual tasks.

[0036] Some example embodiments of the present disclosure will be further described below with reference to the accompanying drawings.

[0037] Example architecture for target vision tasks

[0038] Figure 2 A schematic diagram of an image processing architecture 200 for target visual tasks according to some embodiments of the present disclosure is shown. The architecture 200 can be implemented by an image processing system 120. AsFigure 2 As shown, the image processing system 120 can obtain a first feature representation 210 of the input image 101 to be processed. The first feature representation 210 can be regarded as the features of the input image 101 in the latent space. Compared with the input image 101, the first feature representation 210 can be dimension-reduced or compressed. Any suitable method or machine learning model can be used to generate the first feature representation 210. Alternatively, the first feature representation 210 can be generated by an external system and input into the image processing system 120. Embodiments of the present disclosure are not limited in this regard.

[0039] The image processing system 120 can also include or utilize hint information 230 corresponding to the target visual task. Specifically, the image processing system 120 can use the diffusion model 201 to generate a second feature representation 220 of the input image 101 based on the hint information 230 and the first feature representation 210. The hint information 230 is obtained through the training of the diffusion model 201. Specifically, the hint information 230 can be determined by training the diffusion model 201 for the target visual task and can be regarded as a description of the target visual task. The diffusion model 201 can be pre-trained and fine-tuned for the target visual task.

[0040] The hint information 230 can be implemented in any suitable manner. In some embodiments, the hint information 230 can include multiple hint representations of the same dimension. For example, the hint information 230 can include multiple embedding representations, also known as meta-hints. These meta-hints are incorporated into the diffusion model 201. Exemplarily, the hint information 230 can be represented as where N represents the number of meta-hints, and D represents the dimension of the meta-hints. In some embodiments, the number of hint representations (e.g., the value of N) can be associated with the target visual task. That is, the number of hint representations can be set according to the specific visual task.

[0041] Such hint information 230 can be used to mimic text embeddings. With the hint information 230, external text prompts (such as classification labels, image captions, etc.) are no longer necessary for the diffusion model 201, and the use of pre-trained text encoders is also avoided.

[0042] Such hint information 230 is learnable and learns and adapts to the specific target visual task as the diffusion model 201 is trained. For example, the diffusion model 201 and the hint information 230 can be trained end-to-end according to the target visual task and the corresponding dataset, so as to construct hint information 230 customized for the diffusion model 201 and the target visual task.

[0043] In some embodiments, during the training phase, an initial prompt representation for the prompt information can be generated. For example, the prompt information can be randomly initialized. Based on the initial prompt representation, using a diffusion model, the training loss for the target visual task can be determined, and the diffusion model and the initial prompt representation can be updated based on the training loss until a predetermined condition is met. For example, the diffusion model 201 and the prompt representation can be updated by minimizing the training loss until the training loss is less than a predetermined threshold or a predetermined number of training rounds is reached. In this way, the updated initial prompt representation can be determined as the prompt information 230.

[0044] At initialization, the prompt information (such as having random values) may not possess any meaningful information about the target visual task. As the training process progresses, the prompt information undergoes a transformation process. Through iterative updates during training, the prompt information evolves from meaningless noise into valuable semantic information related to the target visual task. Thus, such prompt information 230 obtained by training the diffusion model 201 for the target visual task can be regarded as containing rich and task-specific semantic information.

[0045] Depending on the specific structure of the diffusion model 201, the diffusion model 201 can combine the first feature representation 210 with the prompt information 230 in any suitable manner. In some embodiments, the diffusion model 201 can be based on a cross-attention mechanism, such as a denoising U-Net. In such an embodiment, the cross-attention can be applied to the first feature representation 210 (or the transformed first feature representation 210) and the prompt information 230 by the diffusion model 201 to obtain an attention map. Then, the diffusion model 201 can determine the second feature representation 220 based on the attention. The diffusion model 201 can be configured with any suitable network structure to apply cross-attention, and the embodiments of the present disclosure are not limited in this regard.

[0046] In such an embodiment, the first feature representation 210 is an image feature, and the prompt information 230, as a semantic feature, can achieve cross-modal fusion with the image feature through the cross-attention mechanism. As a replacement for text embedding, such prompt information 230 can eliminate the difference between the text-to-image diffusion model and the visual perception task.

[0047] Continuing with the architecture 200. The image processing system 120 can further generate a task result 105 for the target visual task based on the second feature representation 220 of the input image 101. For example, the second feature representation 220 or the transformed second feature representation 220 can be fed into a decoder for the target visual task to obtain the task result 105.

[0048] The above describes an example image processing architecture for a target visual task. By using the prompt information determined through the training of the diffusion model, the gap between the generative model and the non-generative visual perception task can be bridged without additional text information. Thus, the diffusion model can be adapted to various visual perception tasks. In this way, the embodiments of the present disclosure implement a general solution for visual perception.

[0049] Multi-scale feature reconstruction

[0050] Figure 3 FIG. 300 is a schematic diagram showing an image processing flow for a target visual task according to some embodiments of the present disclosure. The image processing flow 300 can be regarded as an example flow implemented by the image processing system 120.

[0051] As Figure 3 shown, the encoder 301 can generate a first feature representation 210 of the input image 101 based on the input image 101. For example, the encoder 301 can perform image compression on the input image 101 to reduce the resolution of the input image 101, thereby obtaining a feature representation in the latent space. The encoder 301 can be implemented with any suitable network structure. As an example, a variational autoencoder (VAE), such as a vector quantization variational autoencoder (VQVAE), can be used. However, it should be understood that this is only exemplary and is not intended to impose any limitation. The encoder 302 can have any suitable network structure.

[0052] In some embodiments, the encoder 301 can be fixed. In other words, the encoder 301 can not be trained together with the diffusion model 201. Therefore, any general-purpose encoding model or image feature extraction model can be used to implement the encoder 301.

[0053] Next, the generation of the second feature representation 220 can be based on the first feature representation 210 and the prompt information 230. In some embodiments, the generation of the second feature representation 220 can include multiple steps, such as Figure 3 shown, t steps, where t is a positive integer. Exemplarily, the generation of the second feature representation 220 can include 3 steps. Note that compared with the use of the diffusion model for image generation, the number of steps required in the case of visual perception tasks is greatly reduced.

[0054] In each step, the prompt information 230 is fed into the diffusion model 201. For example, in each step, the image processing system 120 can determine the input feature representation of the step based on the first feature representation 201 (for the first step) or the input of the previous step of the step (for steps after the first step). Then, the image processing system 120 can use the prompt information 230 as at least part of the condition of the diffusion model 201 to generate the output feature representation of the step from the input feature representation of the step. If the step is the last step, the generated output feature representation can be used as the second feature representation 220. If the step is not the last step, the generated output feature representation can be used as the input feature representation of the diffusion model 201 in the next step.

[0055] Through a process such as the one described above, the second feature representation 220 can be generated. In some embodiments, the valuable semantic information contained in the prompt information 230 can be further utilized. As Figure 3 shown, the image processing system 120 can use the prompt information 230 to transform the second feature representation 220 to obtain the transformed feature representation 330. This transformation of the second feature representation can be regarded as feature rearrangement to make the transformed feature representation more relevant to the target visual task.

[0056] In some embodiments, depending on the specific structure of the diffusion model 201, the second feature representation 220 can include features at multiple scales. For example, the features close to the output layer mainly focus on finer low-level details. In some cases, this low-level detail is not sufficient for visual perception tasks that emphasize texture and granularity, and some visual perception tasks may require an understanding of both low-level details and high-level semantics.

[0057] In view of this, features at multiple scales can be combined to generate the transformed feature representation 330. Exemplarily, the image processing system 120 can use the prompt information 230 to transform the features at these scales respectively to obtain transformed features at multiple scales, and can combine the transformed features at multiple scales into the transformed feature representation 330.

[0058] Refer to Figure 4 to describe an example. Features at multiple scales can be represented by where H and W represent the height and width of the feature map respectively. This example includes features at 4 scales, but this is only exemplary and is not intended to impose any limitation. The prompt information 230 can be represented by as described in reference Figure 2 . Accordingly, the transformed feature at each scale can be obtained by the following formula:

[0059]

[0060] Among them, through the dot product operation between the meta-prompt and the feature maps at various scales, the task-adaptive characteristics of the multi-scale features and the prompt information are combined.

[0061] As described above, the prompt information is task-adaptive and embeds the context knowledge of the dataset specific to the target visual task. This context awareness enables the prompt information to be used as a filter to guide the above-mentioned feature reconstruction process to find the features most relevant to the target visual task from the features generated by the diffusion model.

[0062] Continue to refer to Figure 3 After that, the image processing system 120 can use the decoder 302 to determine the task result 105 based on the transformed feature representation 330. For example, for the depth estimation task, the task result 105 can include the estimated depth map. For another example, for the image segmentation task, the task result 105 can include the segmentation mask. The decoder 302 is specific to the target visual task. For example, the decoder 302 can be trained together with the diffusion model 201 for the target visual task. The decoder 302 can have any suitable structure, and the embodiments of the present disclosure are not limited in this regard.

[0063] Multi-step modulation

[0064] As described above, in some embodiments, the generation of the second feature representation 220 can include multiple steps. The output of the diffusion model 201 in one step can be used as the input of the diffusion model 201 in the next step. In some cases, this may pose challenges. As these steps progress, the distribution of the input features may experience a shift. However, the parameters of the diffusion model 201 remain unchanged in different steps.

[0065] For this reason, in some embodiments, in each of these steps, learnable model modulation information can be used. In Figure 2 the example, the model modulation information 310-1, 310-2,..., 310-t (represented by and respectively) for these t steps are shown, which are also collectively or individually referred to as the model modulation information 310. Exemplarily, the model modulation information 310 can be used as the time step encoding of the diffusion model 201 in the corresponding step. In Figure 2 the example, in each step, taking the prompt information 230 as a condition and the model modulation information of this step as the time step encoding, the output feature representation is generated from the input feature representation.

[0066] The model modulation information 310 for different steps can be different, depending on the training of the diffusion model 201. The model modulation information 310 can be a time step embedding for adjusting the parameters of the diffusion model 201.

[0067] The model modulation information is learnable and adapts to changes in the input features as the diffusion model 201 is trained. For example, the diffusion model 201, the prompt information 230, and the model modulation information 310 can be trained end-to-end according to the target visual task and the corresponding dataset.

[0068] In some embodiments, during the training phase, an initial modulation representation for the model modulation information can be generated. For example, the model modulation information for each step can be randomly initialized. Based on the initial prompt representation and the initial modulation representation, using the diffusion model, the training loss for the target visual task can be determined, and the diffusion model, the initial prompt representation, and the initial modulation representation can be updated based on the training loss until a predetermined condition is met. For example, the diffusion model 201, the initial prompt representation, and the initial modulation representation can be updated by minimizing the training loss until the training loss is less than a predetermined threshold or a predetermined number of training rounds is reached. In this way, the updated initial modulation representation can be determined as the model modulation information 310.

[0069] Through the model modulation information, it can be ensured that the diffusion model 201 remains adaptive and responsive to the changing nature of the input features at different steps, thereby optimizing the feature extraction process and enhancing the performance of the diffusion model in visual perception tasks.

[0070] Example process

[0071] Figure 5 A flowchart of a process 500 for image processing according to some embodiments of the present disclosure is shown. The process 500 can be implemented at the electronic device 110 and can be implemented by, for example, the image processing system 120.

[0072] In block 510, the image processing system 120 obtains a first feature representation of the input image to be processed.

[0073] In block 520, the image processing system 120 generates a second feature representation of the input image using the diffusion model based on the first feature representation and the prompt information corresponding to the target visual task. The prompt information is obtained through the training of the diffusion model.

[0074] In block 530, the image processing system 120 generates a result of the input image regarding the target visual task based on the second feature representation.

[0075] In some embodiments, generating the result of the input image for the target visual task includes: obtaining a transformed second feature representation by transforming the second feature representation by using the prompt information; and determining the result based on the transformed second feature representation by using a decoder for the target visual task.

[0076] In some embodiments, the second feature representation includes features of multiple scales, and obtaining the transformed second feature representation includes: respectively transforming the features of multiple scales by using the prompt information to obtain transformed features of multiple scales; and combining the transformed features of multiple scales into the transformed second feature representation.

[0077] In some embodiments, the decoder is trained for the target visual task together with the diffusion model.

[0078] In some embodiments, generating the second feature representation of the input image includes multiple steps, and a given step in the multiple steps includes: determining the input feature representation of the given step based on the first feature representation or the output of the previous step of the given step; and using the diffusion model to generate the output feature representation of the given step from the input feature representation conditioned on the prompt information.

[0079] In some embodiments, generating the output feature representation of the given step includes: using the diffusion model to generate the output feature representation from the input feature representation conditioned on the prompt information and with the model modulation information for the given step encoded as the time step, where the model modulation information is obtained through the training of the diffusion model.

[0080] In some embodiments, the prompt information is obtained by the following method: generating an initial prompt representation for the prompt information; determining the training loss for the target visual task based on the initial prompt representation by using the diffusion model; updating the diffusion model and the initial prompt representation based on the training loss until a predetermined condition is satisfied; and determining the updated initial prompt representation as the prompt information.

[0081] In some embodiments, generating the second feature representation of the input image includes multiple steps, and in a given step of the multiple steps, the model modulation information is used as the time step encoding of the diffusion model, and the model modulation information is obtained by the following method: generating an initial modulation representation for the model modulation information; determining the training loss based on the initial prompt representation and the initial modulation representation by using the diffusion model; updating the diffusion model, the initial encoding representation, and the initial modulation representation based on the training loss until a predetermined condition is satisfied; and determining the updated initial modulation representation as the model modulation information.

[0082] In some embodiments, generating a second feature representation of an input image includes: applying cross-attention to the first feature representation and the prompt information to obtain an attention map; and determining the second feature representation based on the attention map.

[0083] In some embodiments, the prompt information includes a plurality of prompt representations with the same dimension, and the number of the plurality of prompt representations is associated with the target visual task.

[0084] Example devices and equipment

[0085] Figure 6 FIG. shows a schematic structural block diagram of an apparatus 600 for image processing according to certain embodiments of the present disclosure. The apparatus 600 may be implemented as or included in an electronic device 110, such as an image processing system 120. Each module / component in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.

[0086] As shown in the figure, the apparatus 600 includes a first representation acquisition module 610 configured to acquire a first feature representation of an input image to be processed. The apparatus 600 further includes a second representation acquisition module 620 configured to generate a second feature representation of the input image based on the first feature representation and prompt information corresponding to a target visual task, using a diffusion model, where the prompt information is obtained through training of the diffusion model. The apparatus 600 further includes a result generation module 630 configured to generate a result of the input image regarding the target visual task based on the second feature representation.

[0087] In some embodiments, the result generation module 630 includes: a feature transformation module configured to transform the second feature representation by using the prompt information to obtain a transformed second feature representation; and a decoding module configured to determine a result based on the transformed second feature representation by using a decoder for the target visual task.

[0088] In some embodiments, the second feature representation includes features of multiple scales, and the feature transformation module is further configured to: respectively transform the features of multiple scales by using the prompt information to obtain transformed features of multiple scales; and combine the transformed features of multiple scales into a transformed second feature representation.

[0089] In some embodiments, the decoder is trained for the target visual task together with the diffusion model.

[0090] In some embodiments, generating the second feature representation of the input image includes multiple steps, and the second representation acquisition module 620 is further configured to: in a given step of the multiple steps, determine the input feature representation of the given step based on the first feature representation or the output of the previous step of the given step; and use a diffusion model to generate the output feature representation of the given step from the input feature representation conditioned on the prompt information.

[0091] In some embodiments, the second representation acquisition module 620 is further configured to: use a diffusion model to generate the output feature representation from the input feature representation conditioned on the prompt information and with the model modulation information for the given step encoded as the time step, where the model modulation information is obtained through the training of the diffusion model.

[0092] In some embodiments, the apparatus 600 further includes a prompt information acquisition module configured to acquire the prompt information by: generating an initial prompt representation for the prompt information; based on the initial prompt representation, using a diffusion model to determine the training loss for the target visual task; updating the diffusion model and the initial prompt representation based on the training loss until a predetermined condition is met; and determining the updated initial prompt representation as the prompt information.

[0093] In some embodiments, generating the second feature representation of the input image includes multiple steps, and in a given step of the multiple steps, the model modulation information is used as the time step encoding of the diffusion model, and the apparatus 600 further includes a modulation information acquisition module configured to acquire the model modulation information by: generating an initial modulation representation for the model modulation information; based on the initial prompt representation and the initial modulation representation, using a diffusion model to determine the training loss; updating the diffusion model, the initial encoding representation, and the initial modulation representation based on the training loss until a predetermined condition is met; and determining the updated initial modulation representation as the model modulation information.

[0094] In some embodiments, the second representation acquisition module 620 is further configured to: apply cross-attention to the first feature representation and the prompt information to obtain an attention map; and determine the second feature representation based on the attention map.

[0095] In some embodiments, the prompt information includes multiple prompt representations with the same dimension, and the number of the multiple prompt representations is associated with the target visual task.

[0096] Figure 7 A block diagram of an electronic device 700 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 7 The illustrated electronic device 700 is merely exemplary and should not impose any limitation on the functions and scopes of the embodiments described herein. Figure 7The illustrated electronic device 700 can be used to implement Figure 1 the electronic device 110.

[0097] As Figure 7 shown, the electronic device 700 is in the form of a general-purpose electronic device. The components of the electronic device 700 can include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 700.

[0098] The electronic device 700 generally includes multiple computer storage media. Such media can be any accessible media that can be obtained by the electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 700.

[0099] The electronic device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 7 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 720 can include a computer program product 725 that has one or more program modules, and these program modules are configured to execute various methods or actions of various embodiments of the present disclosure.

[0100] The communication unit 740 enables communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 700 can be implemented by a single computing cluster or multiple computer machines that can communicate through a communication connection. Therefore, the electronic device 700 can operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.

[0101] The input device 750 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 760 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 700 can also communicate with one or more external devices (not shown) as needed through the communication unit 740. The external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 700, or communicate with any device that enables the electronic device 700 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0102] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.

[0103] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0104] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0105] Computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to generate a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0106] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart, and combinations of blocks in the block diagrams and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or act, or by a combination of dedicated hardware and computer instructions.

[0107] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvement of technologies in the market, or to enable other ordinary skilled artisans in the art to understand the various implementations disclosed herein.

Claims

1. An image processing method, comprising: Obtaining a first feature representation of an input image to be processed; Based on the first feature representation and hint information corresponding to a target visual task, using a diffusion model to generate a second feature representation of the input image, where the hint information is obtained through training of the diffusion model; And Based on the second feature representation, generating a result of the input image regarding the target visual task.

2. The method according to claim 1, wherein generating the result of the input image regarding the target visual task includes: By using the hint information to transform the second feature representation, obtaining a transformed second feature representation; And Based on the transformed second feature representation, using a decoder for the target visual task to determine the result.

3. The method according to claim 2, wherein the second feature representation includes features of multiple scales, and obtaining the transformed second feature representation includes: Using the hint information to transform the features of the multiple scales respectively to obtain transformed features of the multiple scales; And Combining the transformed features of the multiple scales into the transformed second feature representation.

4. The method according to claim 2, wherein the decoder is trained for the target visual task together with the diffusion model.

5. The method according to claim 1, wherein generating the second feature representation of the input image includes multiple steps, and a given step in the multiple steps includes: Based on the first feature representation or the output of the previous step of the given step, determining an input feature representation of the given step; And Using the diffusion model, conditional on the hint information, to generate an output feature representation of the given step from the input feature representation.

6. The method according to claim 5, wherein generating the output feature representation of the given step includes: Using the diffusion model, conditional on the hint information and with model modulation information for the given step encoded as the time step, to generate the output feature representation from the input feature representation, where the model modulation information is obtained through training of the diffusion model.

7. The method according to claim 1, wherein the hint information is obtained by: Generating an initial hint representation for the hint information; Based on the initial hint representation, using the diffusion model to determine a training loss for the target visual task; Updating the diffusion model and the initial hint representation based on the training loss until a predetermined condition is met; And Determining the updated initial hint representation as the hint information.

8. The method according to claim 7, wherein generating the second feature representation of the input image includes multiple steps, and in a given step of the multiple steps, model modulation information is used as the time step encoding of the diffusion model, and the model modulation information is obtained by: Generating an initial modulation representation for the model modulation information; Based on the initial prompt representation and the initial modulation representation, use the diffusion model to determine the training loss; Update the diffusion model, the initial encoding representation, and the initial modulation representation based on the training loss until the predetermined condition is satisfied; and Determine the updated initial modulation representation as the model modulation information.

9. The method according to claim 1, wherein generating the second feature representation of the input image comprises: Applying cross-attention to the first feature representation and the prompt information to obtain an attention map; And Determining the second feature representation based on the attention map.

10. The method according to claim 1, wherein the prompt information comprises a plurality of prompt representations having the same dimension, and the number of the plurality of prompt representations is associated with the target visual task.

11. An apparatus for image processing, comprising: A first representation acquisition module configured to acquire a first feature representation of an input image to be processed; A second representation acquisition module configured to generate a second feature representation of the input image based on the first feature representation and prompt information corresponding to a target visual task, the prompt information being obtained through training of the diffusion model; And A result generation module configured to generate a result of the input image regarding the target visual task based on the second feature representation.

12. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to execute the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.