Image processing apparatus, image processing method, and computer-readable storage medium
By using the pre-trained diffusion model and the Schrödinger bridge diffusion model of the image to the image in the image generation unit, and adjusting the estimated image in the true image correspondence relationship, the problem of insufficient image generation accuracy in the prior art is solved, and image processing with higher accuracy is achieved.
Patent Information
- Application Number
- CN202410132340.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-08-01
AI Technical Summary
The existing image processing technology is difficult to improve the accuracy of generated images without retraining the pre-trained diffusion model.
The image generation unit uses the pre-trained diffusion model to perform multi-step processing, including adjustment and update of the estimated image and predicted image, using the pre-trained Schrödinger bridge diffusion model based on the image to image to perform image generation, combining the correspondence between the first image and the true image to perform image adjustment, and optimizing the generation process.
The accuracy of generated images is significantly improved, the difference between generated images and true images is reduced, and the accuracy of image recognition and processing is improved.
Smart Images

Figure CN120411268A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and particularly to an image processing apparatus, an image processing method, and a computer-readable storage medium. Background Art
[0002] With the continuous development of technology, image-based technologies (such as image recognition) are widely used. Summary of the Invention
[0003] A brief overview of the present disclosure is given below to provide a basic understanding of certain aspects of the present disclosure. However, it should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify the key or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is merely to present certain concepts of the present disclosure in a simplified form as a prelude to the more detailed description given later.
[0004] The object of the present disclosure is to provide an improved image processing apparatus, an image processing method, and a computer-readable storage medium, for example, which can improve the accuracy of the generated image without retraining a pre-trained diffusion model.
[0005] According to an aspect of the present disclosure, there is provided an image processing apparatus, including: an image generation unit configured to perform processing including m steps on a first image by using a pre-trained diffusion model to generate an m-th predicted image as a second image. Wherein, the i-th step includes: generating an estimated image of the second image based on the (i - 1)-th predicted image generated in the (i - 1)-th step, and generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image. Wherein, generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image includes: adjusting the estimated image based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image to obtain an adjusted estimated image, and generating the i-th predicted image based on the adjusted estimated image and the (i - 1)-th predicted image; and / or generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image under the guidance of the first image. Wherein, m is a natural number greater than 1, and i is a natural number, 1 ≤ i ≤ m.
[0006] According to another aspect of the present disclosure, there is provided an image processing apparatus, including: an image generation unit configured to perform processing including m steps on a first image by using a pre-trained diffusion model to generate an m-th predicted image as a second image. Wherein, the i-th step includes: generating an estimated image of the second image based on the (i - 1)-th predicted image generated in the (i - 1)-th step, and generating an i-th predicted image based on the estimated image and the (i - 1)-th predicted image. Wherein, the image generation unit is further configured to: perform at least two rounds of the B-th step to the A-th step, and before performing the B-th step in the second round or subsequent rounds, update the (B - 1)-th predicted image based on the first image and the A-th predicted image generated in the A-th step of the previous round, and perform at least two rounds of the B-th step to the A-th step. Wherein, in the second round or subsequent rounds, the B-th step includes: generating a new B-th predicted image based on the first image and the A-th predicted image generated in the A-th step of the previous round. Wherein, i, m, A, and B are natural numbers, 1 ≤ i ≤ m, 1 < B < A ≤ m.
[0007] According to still another aspect of the present disclosure, there is provided an image processing method, including: performing processing including m steps on a first image by using a pre-trained diffusion model to generate an m-th predicted image as a second image. Wherein, the i-th step includes: generating an estimated image of the second image based on the (i - 1)-th predicted image generated in the (i - 1)-th step, and generating an i-th predicted image based on the estimated image and the (i - 1)-th predicted image. Wherein, generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image includes: adjusting the estimated image based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image to obtain an adjusted estimated image, and generating the i-th predicted image based on the adjusted estimated image and the (i - 1)-th predicted image; and / or generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image under the guidance of the first image. Wherein, m is a natural number greater than 1, and i is a natural number, 1 ≤ i ≤ m.
[0008] According to other aspects of the present disclosure, there are also provided computer program codes and computer program products for implementing the methods according to the present disclosure as described above, and computer-readable storage media having recorded thereon the computer program codes for implementing the methods according to the present disclosure as described above.
[0009] In the following specification part, other aspects of the embodiments of the present disclosure are given, wherein, the preferred embodiments for fully disclosing the embodiments of the present disclosure are described in detail without imposing limitations thereon. Description of the Drawings
[0010] The present disclosure can be better understood by referring to the detailed description given below in conjunction with the accompanying drawings, in which the same or similar reference numerals are used throughout the drawings to denote the same or similar components. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of the specification, and are used to further illustrate the preferred embodiments of the present disclosure and to explain the principles and advantages of the present disclosure. Among them:
[0011] Figure 1 is a block diagram showing an example of the functional configuration of an image processing apparatus according to an embodiment of the present disclosure;
[0012] Figure 2A is a schematic diagram showing an example of a process of generating a second image using other methods;
[0013] Figures 2B to 2D is a schematic diagram showing an example of a process of an image processing apparatus generating a second image according to an embodiment of the present disclosure;
[0014] Figure 3 is a block diagram showing an example of the functional configuration of an image processing apparatus according to an embodiment of the present disclosure;
[0015] Figure 4 is a block diagram showing an example of the functional configuration of an image processing apparatus according to an embodiment of the present disclosure;
[0016] Figure 5 is a block diagram showing an example of the functional configuration of an image processing apparatus according to an embodiment of the present disclosure;
[0017] Figure 6 and Figure 7 is a graph showing a performance comparison between an image processing apparatus according to an embodiment of the present disclosure and other methods;
[0018] Figure 8 and Figure 9 is a graph showing an example of a second image generated using an image processing apparatus according to an embodiment of the present disclosure and other methods;
[0019] Figure 10 is a block diagram showing an example of the functional configuration of an image processing apparatus according to an embodiment of the present disclosure;
[0020] Figure 11 is a flowchart showing an example of a process of an image processing method according to an embodiment of the present disclosure;
[0021] Figure 12 is a flowchart showing an example of a process of an image processing method according to an embodiment of the present disclosure; and
[0022] Figure 13It is a block diagram showing an example structure of a personal computer that can be adopted in an embodiment of the present disclosure. Detailed implementation
[0023] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the accompanying drawings. For clarity and conciseness, not all features of the actual implementation are described in the specification. However, it should be understood that many implementation-specific decisions must be made in the process of developing any such actual implementation in order to achieve the specific goals of the developer, for example, to comply with those system- and business-related constraints, and these constraints may vary depending on the implementation. In addition, it should also be understood that although the development work may be very complex and time-consuming, for those skilled in the art who benefit from the present disclosure, such development work is merely a routine task.
[0024] Here, it should also be noted that in order to avoid obscuring the present disclosure with unnecessary details, only the device structures and / or processing steps closely related to the solution according to the present disclosure are shown in the drawings, while other details less related to the present disclosure are omitted.
[0025] The embodiments of each aspect according to the present disclosure will be described in detail below with reference to the accompanying drawings.
[0026] Embodiments of the first aspect
[0027] Figure 1 It is a block diagram showing an example of the functional configuration of an image processing apparatus 100 according to an embodiment of the present disclosure. As Figure 1 shown, the image processing apparatus 100 according to an embodiment of the present disclosure may include an image generation unit 102.
[0028] The image generation unit 102 may be configured to use a pre-trained diffusion model to perform a process including m steps on the first image x m to generate the m-th predicted image x0 as the second image. m is a natural number greater than 1.
[0029] For example, the i-th (i = 1, 2,... m) step may include: based on the (i - 1)-th predicted image x m-i+1 generated in the (i - 1)-th step, generating an estimated image of the second image and based on the estimated image (that is, the estimated image generated in the i-th step ) and the (i - 1)-th predicted image x m-i+1 generating the i-th predicted image x m-i .
[0030] In the case of i = 1, since there is no 0-th step, in the first step, the first image x mis used as the 0th predicted image. That is to say, the first step may include: based on the first image x m generating an estimated image of the second image and based on the estimated image (i.e., the estimated image generated in the first step ) and the first image x m generating the 1st predicted image x m-1 .
[0031] For example, generating the ith predicted image x based on the estimated image m-i+1 and the (i - 1)th predicted image x m-i may include a first operation: based on the first image x m and the correspondence between the first image x m and the ground truth image corresponding to the first image to adjust the estimated image to obtain an adjusted estimated image and generating the ith predicted image based on the adjusted estimated image and the (i - 1)th predicted image.
[0032] The purpose of adjusting the estimated image is to make the difference between the adjusted estimated image and the ground truth image smaller than the difference between the estimated image before adjustment and the ground truth image so as to improve the accuracy of the finally generated second image. For example, the accuracy of the second image can be characterized by the difference between the second image and the ground truth image.
[0033] As an example, a set of predicted images of the ground truth image can be generated based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image, and the estimated image can be adjusted to one of the predicted images in the set of predicted images whose difference from the estimated image is less than or equal to a predetermined threshold, thereby further improving the accuracy of the finally generated second image. For example, the above - mentioned predetermined threshold can be set according to actual needs or through a limited number of experiments.
[0034] As another example, the estimated image can be adjusted to the predicted image in the set of predicted images with the smallest difference from the estimated image, as Figure 2B shown, thereby further improving the accuracy of the finally generated second image. In Figure 2B , the circle including the ground truth image represents the set of predicted images.
[0035] For example, comparing Figure 2B and Figure 2ASchematic diagrams showing the process of generating a second image by the image processing apparatus 100 and the process of generating a second image without including the above adjustment operation on the estimated image. It can be seen that through the above adjustment operation, the adjusted estimated image obtained and the ground truth image the difference between them is reduced, and the difference between the finally generated second image x0 and the ground truth image is reduced.
[0036] For example, the first image may be an image of the first version, the ground truth image corresponding to the first image may be the ground truth image of the second version corresponding to the image of the first version, the second image may be a predicted image of the second version, and the estimated image may be an estimated image of the second version. For example, the image of the first version may be an image with low resolution, and the image of the second version may be an image with high resolution. For example, the image of the first version may be a blurred or occluded image, and the image of the second version may be a clearer or less occluded (even unoccluded) image.
[0037] Embodiments of the second aspect
[0038] Figure 3 is a block diagram showing a functional configuration example of an image processing apparatus 300 according to an embodiment of the present disclosure. As Figure 3 shown, the image processing apparatus 300 according to an embodiment of the present disclosure may include an image generation unit 302.
[0039] The image generation unit 302 may be configured to perform a process including m steps on the first image x m using a pre-trained diffusion model to generate the m-th predicted image x0 as the second image. m is a natural number greater than 1.
[0040] For example, the i-th (i = 1, 2,... m) step may include: generating an estimated image of the second image based on the (i - 1)-th predicted image x m-i+1 generated in the (i - 1)-th step, and based on the estimated image and the (i - 1)-th predicted image x m-i+1 generating the i-th predicted image x m-i .
[0041] In the case of i = 1, since there is no 0-th step, in the 1-st step, the first image x m can be used as the 0-th predicted image.
[0042] For example, the i-th predicted image may be generated under the guidance of the first image, so that the accuracy of the finally generated second image can be improved.
[0043] In the case of generating the i-th predicted image under the guidance of the first image, based on the estimated image and the (i - 1)-th predicted image x m-i+1 generate the i-th predicted image x m-i may include a second operation: generating an initial i-th predicted image based on the estimated image and the (i - 1)-th predicted image; and adjusting the initial i-th predicted image based on the first image to obtain an adjusted i-th predicted image as the i-th predicted image.
[0044] Embodiments of the third aspect
[0045] Figure 4 is a block diagram showing an example of the functional configuration of an image processing apparatus 400 according to an embodiment of the present disclosure. As Figure 4 shown, the image processing apparatus 400 according to an embodiment of the present disclosure may include an image generation unit 402.
[0046] The image generation unit 402 may be configured to use a pre-trained diffusion model to perform a process including m steps on the first image x m to generate the m-th predicted image x0 as the second image. m is a natural number greater than 1.
[0047] For example, the i-th (i = 1, 2,..., m) step may include: based on the (i - 1)-th predicted image x generated in the (i - 1)-th step m-i+1 generate an estimated image of the second image and based on the estimated image and the (i - 1)-th predicted image x m-i+1 generate the i-th predicted image x m-i .
[0048] In the case of i = 1, since there is no 0-th step, in the first step, the first image x m can be used as the 0-th predicted image.
[0049] For example, some of the above m steps may be selectively executed two or more rounds to update the predicted images generated in the corresponding steps, thereby improving the accuracy of the finally generated second image.
[0050] For example, the steps from step B to step A can be executed for at least two rounds, where A and B are natural numbers, 1 < B < A ≤ m. In this case, after executing step A in the previous round, the next round of processing is started, that is, the steps from step B to step A in the next round are started. In addition, after executing step A in the last round, step A + 1 in the first round is started. For example, when the steps from step B to step A are executed for two rounds, after executing step A in the first round, the second round of processing is started, that is, the steps from step B to step A in the second round are started. In addition, after executing step A in the second round, step A + 1 in the first round is started.
[0051] For steps B + 1 to A, the processing executed in each round can be the same. That is, in each round, step j (j = B + 1, B + 2,... A) can include: generating an estimated image of the second image based on the (j - 1)-th predicted image generated in this round, and generating a new j-th predicted image based on the estimated image and the (j - 1)-th predicted image, that is, updating the j-th predicted image generated in the previous round.
[0052] For step B, the processing executed in the second round or subsequent rounds can be similar to the processing executed in the first round. For example, in the first round, step B can include: generating an estimated image of the second image based on the (B - 1)-th predicted image generated in the first round, and generating the B-th predicted image based on the estimated image and the (B - 1)-th predicted image.
[0053] In the second round or subsequent rounds, step B - 1 is not executed. If the (B - 1)-th image generated in the first round is used to generate a new B-th predicted image, the finally obtained second image is the same as the second image generated when all m steps are executed for only one round. The inventors of the present application have found through a large number of experiments that by updating the (B - 1)-th image generated in the previous round based on the first image and the A-th predicted image generated in the previous round of step A in the second round or subsequent rounds and generating the B-th predicted image based on the updated (B - 1)-th image, the accuracy of the finally generated second image can be further improved.
[0054] For example, in the second round or subsequent rounds, the (B - 1)-th image generated in the previous round can be updated based on the first image and the A-th predicted image generated in the previous round of step A to obtain an updated (B - 1)-th image. In addition, in the second round or subsequent rounds, step B can include: generating an estimated image of the second image based on the updated (B - 1)-th predicted image, and generating the B-th predicted image based on the estimated image and the updated (B - 1)-th predicted image.
[0055] Embodiments of the fourth aspect
[0056] Figure 5 is a block diagram showing an example of a functional configuration of an image processing apparatus 500 according to an embodiment of the present disclosure. As Figure 5 shown, the image processing apparatus 500 according to an embodiment of the present disclosure may include an image generation unit 502.
[0057] The image processing apparatus 500 may be obtained by combining any two or three of the embodiments of the above first aspect to third aspect.
[0058] For example, the image processing apparatus 500 may be obtained by combining the embodiments of the above first aspect and second aspect. In this case, the image generation unit 502 included in the image processing apparatus 500 may generate the i-th predicted image by executing a third operation obtained by combining the first operation described above for the image generation unit 102 and the second operation described above for the image generation unit 302, so that the accuracy of the finally generated second image may be further improved.
[0059] The third operation may include: adjusting the estimated image based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image to obtain an adjusted estimated image; and generating the i-th predicted image based on the adjusted estimated image and the (i - 1)-th predicted image under the guidance of the first image.
[0060] For example, generating the i-th predicted image based on the adjusted estimated image and the (i - 1)-th predicted image under the guidance of the first image may include: generating an initial i-th predicted image based on the adjusted estimated image and the (i - 1)-th predicted image; and adjusting the i-th predicted image based on the first image to obtain an adjusted i-th predicted image as the i-th predicted image, as Figure 2C shown. From Figure 2C it can be seen that by adjusting the i-th predicted image based on the first image, the difference between the finally generated second image x0 and the ground truth image can be further reduced.
[0061] For example, the image processing apparatus 500 may be obtained by combining the embodiments of the above first aspect and third aspect. In this case, the image generation unit 502 included in the image processing apparatus 500 may selectively execute two or more rounds of some of the m steps described for the image generation unit 102 to update the predicted images generated in the corresponding steps, so that the accuracy of the finally generated second image may be further improved.
[0062] For example, the image processing apparatus 500 may be obtained by combining the embodiments of the second and third aspects described above. In this case, the image generation unit 502 included in the image processing apparatus 500 may selectively perform two or more rounds of some of the m steps described for the image generation unit 302 to update the predicted images generated in the corresponding steps, so that the accuracy of the finally generated second image can be further improved.
[0063] For example, the image processing apparatus 500 may be obtained by combining the embodiments of the first to third aspects described above. In this case, as Figure 2D shown, the image generation unit 502 included in the image processing apparatus 500 may selectively perform two or more rounds of some of the m steps including the third operation described above to update the predicted images generated in the corresponding steps, so that the accuracy of the finally generated second image can be further improved.
[0064] For example, the pre-trained diffusion model may be a pre-trained diffusion model based on the Image-to-Image Schrödinger Bridge (I SB) (e.g., see Reference 1: Guan-Horng Liu, Arash Vahdat, De-An Huang, Evangelos A Theodorou, Weili Nie, and Anima Anandkumar, “I 2 sb: Image-to-image 2 bridge,” arXiv preprint arXiv:2302.05872, 2023), but the present disclosure is not limited thereto. Instead, those skilled in the art can select a suitable pre-trained diffusion model according to actual needs.
[0065]
[0066] 2 SB diffusion model, an example of the image processing apparatus 500 obtained by combining the embodiments of the first to third aspects described above will be further described in detail below. Note that the following description can be similarly applied to other embodiments.
[0066] The image generation unit 502 may be configured to use the pre-trained I 2 SB diffusion model to perform a process including m steps on the first image x m to generate the m-th predicted image x0 as the second image.
[0067] For example, the i-th (i = 1, 2, …, m) step may include: based on the (i - 1)-th predicted image x generated in the (i - 1)-th step m-i+1 generating an estimated image of the second image and based on the estimated image (i.e., the estimated image generated in the i-th step ) and the (i - 1)-th predicted image x m-i+1 generating the i-th predicted image x m-i .
[0068] The following will describe in detail the processing example of the i-th step by taking the example of i = (m - n + 1) (1 ≤ n ≤ m). It can be understood that similar processing can be applied to other similar steps.
[0069] In the (m - n + 1)-th step, the following first to sixth processes can be executed.
[0070] In the first process, an estimated image of the second image can be generated based on the (m - n)-th predicted image x generated in the (m - n)-th step n generating an estimated image of the second image For example, the first process can be represented by the following formula (1):
[0071]
[0072] In formula (1), t n represents the discrete time point corresponding to the (m - n)-th predicted image x n . Additionally, ∈ θ (x n , t n ) represents a neural network that can be used to estimate the noise difference between x n and x0. is calculated by the formula , where βτ is the diffusion coefficient equation. For the meanings of the respective parameters in the above formula (1), reference can also be made to the above-mentioned reference 1.
[0073] In the second process, the estimated image m can be adjusted based on the first image x m and the correspondence between the first image x and the ground truth image corresponding to the first image to obtain an adjusted estimated image For example, the correspondence between the first image x m (which can also be referred to as the observation signal y) and the ground truth image corresponding to the first image can be expressed as the following formula (2):
[0074]
[0075] For example, a set of predicted images of the ground truth image can be generated based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image And the estimated image can be adjusted to the predicted image in the set of predicted images that has the smallest difference from the estimated image. For example, in this case, the second process can be represented by the following formula (3):
[0076]
[0077] If in formula (3) is represented in matrix form as H, then the above formula (3) can be converted to formula (4):
[0078]
[0079] In formula (4), represents the Moore-Penrose pseudo-inverse of H, and I represents the identity matrix.
[0080] In the third process, the adjusted estimated image can be calculated respectively by the following formulas (5) to (8) The (m - n)th predicted image x n and the coefficients μ0, μ of the Wiener noise n and μ z .
[0081]
[0082]
[0083]
[0084]
[0085] For the meanings of the respective parameters in the above formulas (5) to (8), reference can be made to the above-mentioned reference 1.
[0086] In the fourth process, an initial (m - n + 1)th predicted image x can be generated based on the adjusted estimated image n and the (m - n)th predicted image x n-1 . For example, the fourth process can be represented by the following formula (9):
[0087]
[0088] In the fifth process, based on the first image x mAdjust the initial (m - n + 1)-th predicted image x n-1 to obtain the adjusted (m - n + 1)-th predicted image as the (m - n + 1)-th predicted image. Since there is no direct analyzable closed-form relationship between the first image x m (which can also be referred to as the observed signal y) and the (m - n)-th predicted image x n-1 used to generate the initial (m - n + 1)-th predicted image x n , but the first image x m is indirectly linked to the (m - n)-th predicted image x n through the ground truth image . For example, the relationship between the (m - n)-th predicted image x n and the ground truth image can be represented by the following equation (10), and the relationship between the observed signal y and the ground truth image can be represented by the above equation (2).
[0089]
[0090] In equation (10), the random time which can be converted into a set of discrete values 0 = t0 < … t n < … t m = 1. g0 represents the ground truth image, g1 represents the first image x m , and
[0091] In addition, in equation (10), where β τ represents the diffusion coefficient function. The meanings of the respective parameters in equation (10) can be referred to the above reference 1.
[0092] For example, the conditional fraction in the case of the observed signal y can be estimated as the sum of the posterior score of y and the unconditional fraction. The (m - n)-th predicted image x n already includes the unconditional fraction, so the conditional fraction can be obtained by adding the posterior score of the observed signal y to it. For example, the posterior score can be obtained by the following equation (11):
[0093]
[0094] In equation (11), which can be approximated as the adjusted estimated image For the meanings of the various parameters in Equation (11), reference can be made to, for example, Reference 2 (Hyungjin Chung, Jeongsol Kim, Michael T Mccann, Marc L Klasky, and Jong Chul Ye, “Diffusion posterior sampling for general noisy inverse problems,” arXiv preprint arXiv:2209.14687, 2022). Thus, the fifth process can be represented by the following Equation (12):
[0095]
[0096] In Equation (12), η represents the sliding step size, which can be adjusted according to the effect of image generation. Generally, it is a very small value, such as 0.01. Through the above fifth process, the (m - n + 1)-th predicted image x n-1 can slide along the feasible gradient direction on the image manifold and gradually approach the true value image.
[0097] In the sixth process, Wiener noise can be added to the (m - n + 1)-th predicted image x n-1 . For example, the sixth process can be represented by the following Equation (13):
[0098]
[0099] In Equation (13), the Wiener noise
[0100] Note that in the m-th step, the sixth process is not executed.
[0101] For example, in the case of updating the predicted images generated during the time period [t B , t A , that is, in the case of performing two or more rounds for the B-th step to the A-th step, for the (B + 1)-th step to the A-th step, the processing performed in each round can be the same. That is to say, in each round, the (B + 1)-th step to the A-th step can include the above first process to the sixth process. Note that in the case of A = m, the A-th step does not include the above sixth process.
[0102] For the B step, the processing performed in the second round or subsequent rounds can be similar to the processing performed in the first round. For example, in the first round, the B step may include the first to sixth processes described above. In the second round or subsequent rounds, the B step includes processes that can be similar to the first to sixth processes described above, with the only difference being that the B step uses the updated B-1 image based on the first image and the A prediction image generated in the A step of the previous round. For example, when performing the B step in the first round, replace x in the above equations (1) and (9) with the B-1 image x generated in the B-1 step of the first round, while in the second round or subsequent rounds, replace x in the above equations (1) and (9) with the updated B-1 image. Of course, in the second round or subsequent rounds, the updated B-1 images are different from each other. n with the B-1 image x generated in the B-1 step of the first round m-(B-1) instead, and in the second round or subsequent rounds, replace x in the above equations (1) and (9) with the n updated B-1 image instead. Of course, in the second round or subsequent rounds, the updated B-1 images are different from each other.
[0103] Based on the first image x m and the A prediction image generated in the A step of the previous round to update the B-1 image The process can be represented by the following equation (14):
[0104]
[0105] Figure 6 , Figure 7 , Figure 8 and Figure 9 are diagrams showing a comparison between the image processing apparatus according to an embodiment of the present disclosure and other technologies. In Figures 6 to 8 I 2SB represents a first comparison method for generating a second image without including the above adjustment operation on the estimated image (see the above reference 1), DDRM represents a second comparison method using a denoising diffusion restoration model (see reference 3: Bahjat Kawar, Michael Elad, Stefano Ermon, and Jiaming Song, “Denoising diffusion restoration models,” NeurIPS, vol. 35, pp. 23593 - 23606, 2022), DPS represents a third comparison method using denoised posterior sampling (see the above reference 2), +P. represents the image processing apparatus 100 according to an embodiment of the first aspect, +S. represents the image processing apparatus 300 according to an embodiment of the second aspect, +R. represents the image processing apparatus 400 according to an embodiment of the third aspect, +P&S. represents an example of the image processing apparatus 500 obtained by combining the embodiments according to the first and second aspects, and PSR. represents an example of the image processing apparatus 500 obtained by combining the embodiments according to the first to third aspects. Additionally, in Figure 6 and Figure 7 the first test image set represents an image set obtained by blurring the original image set (which may also be referred to as the “ground truth image set”) using the Gauss method, the second test image set represents an image set obtained by blurring the original image set using the uniform method, the third test image set represents an image set obtained by compressing the original image into the JPEG format, the fourth test image set represents an image set obtained by processing the original image set using a bicubic-based super-resolution technique, and the fifth test image set represents an image set obtained by processing the original image set using a pooling-based super-resolution technique. Additionally, in Figure 6 and Figure 7 FID represents the distance between the generated second image and the ground truth image corresponding to the first image, SSIM represents the similarity between the generated second image and the ground truth image corresponding to the first image, and PSNR represents the signal-to-noise ratio of the generated second image.
[0106] From Figure 6 and Figure 7 it can be seen that in the case of different test image sets, compared with the first to third comparison methods, the FID of the second image generated by the image processing apparatuses according to the embodiments of the first to third aspects of the present disclosure decreases, and the SSIM and PSNR increase, which indicates that the accuracy of the second image generated by the image processing apparatus according to the present disclosure is improved. Additionally, from Figure 6 and Figure 7It can be seen that the accuracy of the second image generated by the image processing apparatus obtained by combining the embodiments of the first to third aspects of the present disclosure is further improved.
[0107] In addition, from Figure 8 and Figure 9 it can also be seen that, compared with the first comparison method and the second comparison method, the difference between the second image generated by the image processing apparatus according to the present disclosure and the ground truth image is reduced, which also indicates that the accuracy of the second image generated by the image processing apparatus according to the present disclosure is improved.
[0108] Embodiments of the fifth aspect
[0109] Figure 10 is a block diagram showing a functional configuration example of an image processing apparatus 1000 according to an embodiment of the present disclosure. As Figure 10 shown, the image processing apparatus 1000 according to an embodiment of the present disclosure may include an image generation unit 1002 and an image recognition unit 1004.
[0110] The image generation unit 1002 may be configured to use a pre-trained diffusion model to perform a process including m steps on the first image x m to generate the m-th predicted image x0 as the second image. For example, the image generation unit 1002 may have a function and configuration similar to any one of the image generation units 102, 302, 402, and 502, and thus will not be described in detail below.
[0111] The image recognition unit 1004 may be configured to recognize an object (such as a face, an object, an animal, etc.) involved in the second image generated by the image generation unit 1002.
[0112] For example, the image processing apparatus 1000 may be implemented as an electronic device (such as a mobile phone, a self-checkout machine, etc.) or may be integrated into an electronic device. For example, the image generation unit 1002 may process the first image acquired by the electronic device (such as through an image acquisition device integrated in the electronic device) to generate the second image. For example, in the case where the first image acquired by the electronic device is blurred or blocked, the image generation unit 1002 may generate a clear or less blocked (even unblocked) second image corresponding to the first image, and provide the generated second image to the image recognition unit 1004, thereby improving the image recognition accuracy.
[0113] For example, the image processing apparatus 1000 can be implemented as a cash register or integrated into a cash register. For instance, the cash register may further include an image acquisition device for acquiring an image of an item. The image generation unit 1002 can process the item image acquired by the image acquisition device to generate a second image, and the image recognition unit 1004 can recognize the second image to obtain the category of the item for the cash register to determine the price of the item. In this way, even when the item is occluded or the item image is unclear, the category of the item can be accurately recognized.
[0114] For example, the image processing apparatus 1000 can be implemented as an identity recognition device or integrated into an identity recognition device. For instance, the identity recognition device may further include an image acquisition device for acquiring an image of a person or a part of the person (such as the face). The image generation unit 1002 can process the image acquired by the image acquisition device to generate a second image, and the image recognition unit 1004 can recognize the second image to obtain the identity of the person. In this way, even when the acquired image is unclear, the identity of the person can be accurately recognized.
[0115] Embodiments of the sixth aspect
[0116] Figure 11 is a flowchart showing a flow example of an image processing method according to an embodiment of the present disclosure. As Figure 11 shown, the image processing method 1100 according to an embodiment of the present disclosure can start at S1102 and end at S1106.
[0117] As Figure 11 shown, the image processing method 1100 may include an image generation step S1104. In the image generation step S1104, a diffusion model trained in advance can be used to perform processing including m steps on the first image to generate the m-th predicted image as the second image. For example, the image generation step S1104 can be executed by any one of the image generation units 102, 302, 402, and 502 described above, so it will only be briefly described below.
[0118] For example, the i-th (i = 1, 2,... m) step may include: based on the (i - 1)-th predicted image x m-i+1 generated in the (i - 1)-th step, generating an estimated image of the second image and based on the estimated image (i.e., the estimated image generated in the i-th step ) and the (i - 1)-th predicted image x m-i+1 generating the i-th predicted image x m-i .
[0119] For example, based on the estimated image and the (i-1)-th predicted image x m-i+1 generate the i-th predicted image x m-i may include a first operation: based on the first image x m and the first image x m and the ground-truth image corresponding to the first image to adjust the estimated image to obtain an adjusted estimated image and generate the i-th predicted image based on the adjusted estimated image and the (i-1)-th predicted image, thereby improving the accuracy of the finally generated second image.
[0120] As an example, a set of predicted images of the ground-truth image may be generated based on the first image and the correspondence between the first image and the ground-truth image corresponding to the first image, and the estimated image may be adjusted to one of the predicted images in the set of predicted images whose difference from the estimated image is less than or equal to a predetermined threshold, thereby further improving the accuracy of the finally generated second image. For example, the above-mentioned predetermined threshold may be set according to actual needs or through a limited number of experiments.
[0121] As another example, the estimated image may be adjusted to the predicted image in the set of predicted images that has the smallest difference from the estimated image, thereby further improving the accuracy of the finally generated second image.
[0122] For example, the i-th predicted image may be generated under the guidance of the first image, thereby improving the accuracy of the finally generated second image.
[0123] For example, some of the above m steps may be selectively executed two or more rounds to update the predicted images generated in the corresponding steps, thereby improving the accuracy of the finally generated second image.
[0124] Embodiments of the seventh aspect
[0125] Figure 12 is a flowchart showing a flow example of an image processing method according to an embodiment of the present disclosure. As Figure 12 shown, the image processing method 1200 according to an embodiment of the present disclosure may start at S1202 and end at S1208.
[0126] As Figure 12 shown, the image processing method 1200 according to an embodiment of the present disclosure may include an image generation step S1204 and an image recognition step S1206.
[0127] In the image generation step S1204, using a pre-trained diffusion model, the first image is processed in m steps to generate the m-th predicted image as the second image. For example, the image generation step S1204 can be performed by the above-mentioned image generation unit 1002, so it will not be described below.
[0128] In the image recognition step S1206, the object (such as a face, an object, an animal, etc.) involved in the second image generated in the image generation step S1204 can be recognized. For example, the image recognition step S1206 can be performed by the above-mentioned image recognition unit 1004, so it will not be described below.
[0129] It should be noted that although the functional configurations and operations of the image processing apparatus and method according to the embodiments of the present disclosure are described above, these are merely examples and not limitations, and those skilled in the art can modify the above embodiments according to the principles of the present disclosure. For example, functions modules in each embodiment can be added, deleted, or combined, etc., and such modifications all fall within the scope of the present disclosure.
[0130] In addition, it should also be noted that the method embodiments here correspond to the above-mentioned apparatus embodiments. Therefore, for the content not described in detail in the method embodiments, reference can be made to the corresponding parts in the apparatus embodiments, and it will not be repeated here.
[0131] It should be understood that the machine-executable instructions in the storage medium and program product according to the embodiments of the present disclosure can also be configured to execute the above-mentioned image processing method. Therefore, for the content not described in detail here, reference can be made to the previous corresponding parts, and it will not be repeated here.
[0132] Correspondingly, the storage medium for carrying the above-mentioned program product including machine-executable instructions is also included in the disclosure of the present invention. The storage medium includes but is not limited to floppy disks, optical discs, magneto-optical discs, memory cards, memory sticks, and the like.
[0133] In addition, it should also be noted that the above-mentioned series of processes and systems can also be implemented by software and / or firmware. In the case of implementation by software and / or firmware, a program constituting the software is installed from a storage medium or a network into a computer having a dedicated hardware structure, such as Figure 13 the general personal computer 1300 shown. When various programs are installed on this computer, it can execute various functions, etc.
[0134] In Figure 13Among them, the central processing unit (CPU) 1301 executes various processes according to the programs stored in the read-only memory (ROM) 1302 or the programs loaded from the storage device 1308 into the random access memory (RAM) 1303. In the RAM 1303, data required when the CPU 1301 executes various processes, etc. is also stored as needed.
[0135] The CPU 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. The input / output interface 1305 is also connected to the bus 1304.
[0136] The following components are connected to the input / output interface 1305: an input device 1306 including a keyboard, a mouse, etc.; an output device 1307 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage device 1308 including a hard disk, etc.; and a communication device 1309 including a network interface card such as a LAN card, a modem, etc. The communication device 1309 executes communication processing via a network such as the Internet.
[0137] As needed, a drive 1310 is also connected to the input / output interface 1305. A removable medium 1311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1310 as needed, so that the computer program read therefrom is installed in the storage device 1308 as needed.
[0138] In the case where the above series of processes are implemented by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 1311.
[0139] Those skilled in the art should understand that such a storage medium is not limited to Figure 13 the removable medium 1311 shown in which the program is stored and distributed separately from the device to provide the program to the user. Examples of the removable medium 1311 include a magnetic disk (including a floppy disk (registered trademark)), an optical disk (including a compact disc read-only memory (CD-ROM) and a digital versatile disc (DVD)), a magneto-optical disk (including a mini disc (MD) (registered trademark)), and a semiconductor memory. Alternatively, the storage medium may be the ROM 1302, a hard disk included in the storage device 1308, etc., in which the program is stored and distributed to the user together with the device containing them.
[0140] The preferred embodiments of the present disclosure have been described above with reference to the accompanying drawings, but the present disclosure is of course not limited to the above examples. Those skilled in the art can obtain various changes and modifications within the scope of the appended claims, and it should be understood that these changes and modifications will naturally fall within the technical scope of the present disclosure.
[0141] For example, in the above embodiments, multiple functions included in one unit can be implemented by separate devices. Alternatively, multiple functions implemented by multiple units in the above embodiments can be separately implemented by separate devices. Additionally, one of the above functions can be implemented by multiple units. Needless to say, such configurations are included within the technical scope of the present disclosure.
[0142] In this specification, the steps described in the flowcharts include not only processes executed in chronological order in the stated sequence, but also processes executed in parallel or individually rather than necessarily in chronological order. Moreover, even in steps of chronological processing, needless to say, the order can be appropriately changed.
[0143] The various techniques described in this specification can be executed independently of each other unless there is a contradiction. Of course, any multiple techniques can be executed in combination. In one example, a part or all of the techniques described in any embodiment can be executed in combination with a part or all of the techniques described in another embodiment. Additionally, any part or all of the techniques described above can be executed in combination with another technique not described above.
[0144] Additionally, the technique according to the present disclosure can also be configured as follows.
[0145] Supplementary Note 1. An image processing apparatus, comprising:
[0146] An image generation unit configured to perform a process including m steps on a first image using a pre-trained diffusion model to generate an mth predicted image as a second image,
[0147] wherein the ith step includes: generating an estimated image of the second image based on the (i - 1)th predicted image generated in the (i - 1)th step, and generating the ith predicted image based on the estimated image and the (i - 1)th predicted image,
[0148] wherein generating the ith predicted image based on the estimated image and the (i - 1)th predicted image includes:
[0149] adjusting the estimated image based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image to obtain an adjusted estimated image, and generating the ith predicted image based on the adjusted estimated image and the (i - 1)th predicted image; and / or
[0150] generating the ith predicted image based on the estimated image and the (i - 1)th predicted image under the guidance of the first image,
[0151] wherein m is a natural number greater than 1, and i is a natural number, 1 ≤ i ≤ m.
[0152] Supplementary Note 2. The image processing apparatus according to Supplementary Note 1, wherein adjusting the estimated image based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image includes:
[0153] generating a set of predicted images of the ground truth image based on the first image and the correspondence; and
[0154] adjusting the estimated image to one of the predicted images in the set of predicted images whose difference from the estimated image is less than or equal to a predetermined threshold, or the predicted image in the set of predicted images having the smallest difference from the estimated image.
[0155] Supplementary Note 3. The image processing apparatus according to Supplementary Note 1, wherein generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image under the guidance of the first image includes:
[0156] generating an initial i-th predicted image based on the estimated image and the (i - 1)-th predicted image; and adjusting the initial i-th predicted image based on the first image to obtain an adjusted i-th predicted image as the i-th predicted image.
[0157] Supplementary Note 4. The image processing apparatus according to Supplementary Note 1, wherein the image generation unit is further configured to:
[0158] perform at least two rounds of the B-th step to the A-th step, and
[0159] before performing the B-th step in the second round or subsequent rounds, update the (B - 1)-th predicted image based on the first image and the A-th predicted image generated in the A-th step of the previous round,
[0160] where A and B are natural numbers, 1 < B < A ≤ m.
[0161] Supplementary Note 5. An image processing apparatus, comprising:
[0162] an image generation unit configured to perform a process including m steps on a first image using a pre-trained diffusion model to generate an m-th predicted image as a second image, wherein the i-th step includes: generating an estimated image of the second image based on the (i - 1)-th predicted image generated in the (i - 1)-th step, and generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image,
[0163] wherein the image generation unit is further configured to:
[0164] perform at least two rounds of the B-th step to the A-th step, and
[0165] Before performing the step B in the second round or subsequent rounds, update the B-1 prediction image based on the first image and the A prediction image generated in step A of the previous round.
[0166] Wherein, i, m, A, and B are natural numbers, 1 ≤ i ≤ m, 1 < B < A ≤ m.
[0167] Supplementary Note 6. The image processing apparatus according to any one of Supplementary Notes 1 to 5, wherein the pre-trained diffusion model includes a diffusion model based on an image-to-image Schrödinger bridge.
[0168] Supplementary Note 7. The image processing apparatus according to any one of Supplementary Notes 1 to 5, further comprising:
[0169] An image recognition unit configured to recognize an object involved in the second image generated by the image generation unit.
[0170] Supplementary Note 8. An image processing method, comprising:
[0171] Using a pre-trained diffusion model, perform processing including m steps on a first image to generate an m-th prediction image as a second image.
[0172] Wherein, the i-th step includes: generating an estimated image of the second image based on the (i - 1)-th prediction image generated in the (i - 1)-th step, and generating an i-th prediction image based on the estimated image and the (i - 1)-th prediction image.
[0173] Wherein, generating the i-th prediction image based on the estimated image and the (i - 1)-th prediction image includes:
[0174] Adjusting the estimated image based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image to obtain an adjusted estimated image, and generating the i-th prediction image based on the adjusted estimated image and the (i - 1)-th prediction image; and / or
[0175] Generating the i-th prediction image based on the estimated image and the (i - 1)-th prediction image under the guidance of the first image.
[0176] Wherein, m is a natural number greater than 1, and i is a natural number, 1 ≤ i ≤ m.
[0177] Supplementary Note 9. The image processing method according to Supplementary Note 8, wherein adjusting the estimated image based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image includes:
[0178] Generating a set of predicted images of the ground truth image based on the first image and the corresponding relationship; and
[0179] Adjusting the estimated image to one of the predicted images in the set of predicted images whose difference from the estimated image is less than or equal to a predetermined threshold, or the predicted image in the set of predicted images with the smallest difference from the estimated image.
[0180] Supplement 10. The image processing method according to Supplement 8, wherein generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image under the guidance of the first image includes:
[0181] Generating an initial i-th predicted image based on the estimated image and the (i - 1)-th predicted image; and adjusting the initial i-th predicted image based on the first image to obtain an adjusted i-th predicted image as the i-th predicted image.
[0182] Supplement 11. The image processing method according to Supplement 8, wherein at least two rounds of steps from step B to step A are performed, and
[0183] The image processing method further includes: before performing step B in the second round or subsequent rounds, updating the (B - 1)-th predicted image based on the first image and the A-th predicted image generated in step A of the previous round,
[0184] where A and B are natural numbers, 1 < B < A ≤ m.
[0185] Supplement 12. The image processing method according to any one of Supplements 8 to 11, wherein the pre-trained diffusion model includes a diffusion model based on an image-to-image Schrödinger bridge.
[0186] Supplement 13. The image processing method according to any one of Supplements 8 to 11 further includes:
[0187] Identifying the object involved in the generated second image.
[0188] Supplement 14. A computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to execute the image processing method according to any one of Supplements 8 to 13.
Claims
1. An image processing apparatus, comprising: An image generation unit configured to perform processing including m steps on a first image by using a pre-trained diffusion model to generate an m-th predicted image as a second image, wherein the i-th step includes: generating an estimated image of the second image based on the (i - 1)-th predicted image generated in the (i - 1)-th step, and generating an i-th predicted image based on the estimated image and the (i - 1)-th predicted image, wherein generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image includes: Adjusting the estimated image based on the first image and the correspondence between the first image and a ground truth image corresponding to the first image to obtain an adjusted estimated image, and generating the i-th predicted image based on the adjusted estimated image and the (i - 1)-th predicted image; and / or Generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image under the guidance of the first image, wherein m is a natural number greater than 1, i is a natural number, and 1 ≤ i ≤ m.
2. The image processing apparatus according to claim 1, wherein, Adjusting the estimated image based on the first image and the correspondence between the first image and a ground truth image corresponding to the first image includes: Generating a set of predicted images of the ground truth image based on the first image and the correspondence; and Adjusting the estimated image to: one of the predicted images in the set of predicted images whose difference from the estimated image is less than or equal to a predetermined threshold, or the predicted image in the set of predicted images with the smallest difference from the estimated image.
3. The image processing apparatus according to claim 1, wherein, Generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image under the guidance of the first image includes: Generating an initial i-th predicted image based on the estimated image and the (i - 1)-th predicted image; and adjusting the initial i-th predicted image based on the first image to obtain an adjusted i-th predicted image as the i-th predicted image.
4. The image processing apparatus according to claim 1, wherein, The image generation unit is further configured to: Execute at least two rounds of the B-th step to the A-th step, and Before executing the B-th step in the second round or subsequent rounds, update the (B - 1)-th predicted image based on the first image and the A-th predicted image generated in the A-th step of the previous round, wherein A and B are natural numbers, 1 < B < A ≤ m.
5. An image processing apparatus, comprising: An image generation unit configured to perform processing including m steps on a first image by using a pre-trained diffusion model to generate an m-th predicted image as a second image, wherein the i-th step includes: generating an estimated image of the second image based on the (i - 1)-th predicted image generated in the (i - 1)-th step, and generating an i-th predicted image based on the estimated image and the (i - 1)-th predicted image, wherein the image generation unit is further configured to: Execute at least two rounds of the B-th step to the A-th step, and Before executing the B-th step in the second round or subsequent rounds, update the (B - 1)-th predicted image based on the first image and the A-th predicted image generated in the A-th step of the previous round, Among them, i, m, A, and B are natural numbers, where 1 ≤ i ≤ m, and 1 < B < A ≤ m.
6. The image processing apparatus according to any one of claims 1 to 5, wherein, The pre-trained diffusion model includes a diffusion model based on an image-to-image Schrödinger bridge.
7. The image processing apparatus according to any one of claims 1 to 5 further includes: An image recognition unit configured to recognize an object involved in the second image generated by the image generation unit.
8. An image processing method includes: Using a pre-trained diffusion model, performing processing including m steps on a first image to generate an m-th predicted image as the second image, where the i-th step includes: generating an estimated image of the second image based on the (i - 1)-th predicted image generated in the (i - 1)-th step, and generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image, where generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image includes: Adjusting the estimated image based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image to obtain an adjusted estimated image, and generating the i-th predicted image based on the adjusted estimated image and the (i - 1)-th predicted image; and / or Generating the i-th predicted image based on the estimated image and the (i - 1)-th predicted image under the guidance of the first image, where m is a natural number greater than 1, i is a natural number, and 1 ≤ i ≤ m.
9. The image processing method according to claim 8, wherein, Adjusting the estimated image based on the first image and the correspondence between the first image and the ground truth image corresponding to the first image includes: Generating a set of predicted images of the ground truth image based on the first image and the correspondence; and Adjusting the estimated image to one of the predicted images in the set of predicted images whose difference from the estimated image is less than or equal to a predetermined threshold, or the predicted image in the set of predicted images with the smallest difference from the estimated image.
10. A computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to execute the image processing method according to claim 8 or 9.