Image processing method, image processing device, and program

WO2026203510A1PCT designated stage Publication Date: 2026-10-01CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/039595
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2025-11-12
Publication Date
2026-10-01

Smart Images

  • Figure JP2025039595_01102026_PF_FP_ABST
    Figure JP2025039595_01102026_PF_FP_ABST
Patent Text Reader

Abstract

[Problem] To provide an image processing method that enables improvement of the quality of an output image. [Solution] This image processing method comprises: a step of generating a first map that is the result of segmentation corresponding to a captured image and includes a plurality of labels; and a step of generating an output image corresponding to the captured image using a first machine learning model on the basis of the captured image, the first map, and text relating to the meaning of each of the plurality of labels.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method, image processing apparatus, and program

[0001] The present disclosure relates to an image processing method that uses a generative model including images in both input and output.

[0002] Non-Patent Document 1 discloses a method for generating an upscaled image having fine structures and less unnaturalness by inputting an image and text including semantic information of a subject to a generative model.

[0003] K. V. Gandikota and P. Chandramouli, “Text-guided Explorable Image Super-resolution”, Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp.25900-25911

[0004] However, the method of Non-Patent Document 1 has a problem that the quality of the generated image is degraded. For example, in the method of Non-Patent Document 1, the text “The Santa ornament hangs from a Christmas tree branch amongst the colorful bright lights.” is used. This text does not describe in which region of the image the ornament is located, and thus has spatial ambiguity. Therefore, in the generated image, the positional relationship between subjects may deviate from that in the input image of the generative model, or an unnatural texture that does not correspond to the original subject may be generated.

[0005] An image processing method as one aspect of the present invention comprises: a step of generating a first map that is a segmentation result corresponding to a captured image and includes a plurality of labels; and a step of generating an output image corresponding to the captured image using a first machine learning model based on the captured image, the first map, and texts relating to the respective meanings of the plurality of labels.

[0006] According to the present disclosure, an image processing method capable of improving the quality of an output image can be provided.

[0007] This is a block diagram of the image processing system in Example 1. This is an external view of the image processing system in Example 1. This is an explanatory diagram of the diffusion models in Examples 1 and 2. This is a flowchart of the training of the first machine learning model in Example 1. This is a diagram showing the flow of output image generation in Example 1. This is an explanatory diagram of the first map in Examples 1 and 2. This is a diagram showing the structure of the first machine learning model in Example 1. This is a flowchart of estimation by the first machine learning model in Example 1. This is a block diagram of the image processing system in Example 2. This is an external view of the image processing system in Example 2. This is a flowchart of the training of the third machine learning model in Example 2. This is a flowchart of the training of the first machine learning model in Example 2. This is a diagram showing the flow of output image generation in Example 2. This is a diagram showing the structure of the first machine learning model in Example 2. This is a flowchart of estimation by the first and third machine learning models in Example 2.

[0008] The embodiments of the present invention will be described in detail below with reference to the drawings. In each figure, the same reference numeral is used for identical components, and redundant explanations are omitted.

[0009] Before describing this embodiment in detail, the gist of the present invention will be explained. First, a first map is generated, which is the result of segmentation corresponding to the captured image and includes multiple labels. Based on the captured image, the first map, and text describing the meaning of each of the multiple labels, an output image is generated using a first machine learning model. The text contains semantic information of the subjects present in the captured image. The first map also contains spatial information indicating which region of the captured image each subject represented in the text is located in. In other words, the first map and text clarify the meaning and spatial location of the subjects present in the captured image. This suppresses deviations in the positional relationships between subjects from the captured image and the generation of unnatural textures that do not correspond to the original subjects, and enables the generation of an output image that corresponds to the captured image.

[0010] The first type of machine learning model is a generative model (Image to Image) that includes images in both its input and output. Image to Image tasks include upscaling, blur sharpening, denoising, dehazing, and other contrast enhancements, as well as dynamic range expansion, gradation enhancement, refocusing, relighting, and perspective changes. Non-generative models output the average solution for the input image; for example, in upscaling, edges remaining in the low-resolution image are sharpened, but fine structures such as textures that have been lost due to blurring cannot be generated. In contrast, generative models output one of several solutions expected from the low-resolution image, so they can generate an upscaled image with finer structures. In other words, in upscaling, the output image has a finer sample count than the captured image, and it is expected that the generative model will generate finer structures. In blur sharpening, the output image has stronger sharpness than the captured image, and it is expected that the generative model will generate textures that have been lost due to blurring. In denoising, the output image has less noise than the captured image, and it is expected that the generative model will generate fine structures that were removed along with the noise. In contrast enhancement, the output image has stronger contrast than the captured image, and it is expected that the generative model will generate textures that were crushed by haze, etc. In dynamic range expansion, the output image has a wider dynamic range than the captured image, and it is expected that the generative model will generate structures that were crushed in dark or bright areas. In gradation enhancement, the output image has finer gradation than the captured image, and it is expected that the generative model will generate textures that were crushed due to insufficient gradation. In refocusing, the output image is focused at a different distance than the captured image, and it is expected that the generative model will generate fine structures and textures in the focused area of ​​the output image (the area that was out of focus in the captured image). In relighting, the output image is under different lighting conditions than the captured image, and it is expected that the generative model will generate structures that were crushed by shadows and highlights in the captured image. In perspective change, the output image has a different perspective than the captured image.When capturing the same subject at the same size on the image sensor, a wide-angle lens will make distant subjects in the background appear smaller compared to a telephoto lens. This is the difference in perspective. To change perspective, it is necessary to change the scaling ratio depending on the distance to the subject, and it is also necessary to capture subjects in areas where the structure is not captured due to occlusion by the foreground or the lens's field of view. It is expected that a generative model will generate natural structures in these enlarged areas and areas where the structure was not captured.

[0011] Figure 1 is a block diagram of the image processing system 100 in this embodiment. Figure 2 is an external view of the image processing system 100. The image processing system 100 includes a training device 101 and an imaging device 102. The training device 101 includes a storage unit 111, a communication unit 112, an acquisition unit 113, a calculation unit 114, a determination unit 115, and an update unit 116. The imaging device 102 includes an optical system 121, an image sensor 122, a storage unit 123, a communication unit 124, an acquisition unit 125, a calculation unit 126, and a display unit 127.

[0012] The optical system 121 collects light incident from the subject space and forms a subject image. The image sensor 122 receives the subject image and generates an image by photoelectric conversion. The image is processed by the calculation unit 126, which performs predetermined image processing such as removing defective pixels, and stored in the storage unit 123. In this embodiment, upscaling is performed on the stored image using the first machine learning model to generate an output image with finer sampling than the image. However, this can also be applied to tasks other than upscaling, such as converting an image to another image (Image to Image). The image to be upscaled may be an undeveloped RAW image or a developed image. Upscaling may be performed automatically during imaging, or at any timing specified by the user after imaging. The model parameters used in the first machine learning model are trained in the training device 101, are acquired in advance by the imaging device 102 via the communication units 112 and 124, and stored in the storage unit 123. Details regarding the training of the first machine learning model and estimation (upscaling) using the trained first machine learning model will be described later. The output image after upscaling is displayed on the display unit 127 and stored in the storage unit 123.

[0013] The first machine learning model is a generative model. A generative model is a model that learns and acquires a probability distribution for generating various desired data (images in the case of Image to Image), and then generates data according to that probability distribution. Examples of generative models include GAN (Generative Adversarial Network) and VAE (Variational Autoencoder). There are also diffusion models and flow-based generative models. In this example, a diffusion model is used as the first machine learning model. However, other generative models may also be used. The overview of the diffusion model will be explained using Figure 3. Figure 3 is an explanatory diagram of the diffusion model. In the diffusion model, we consider a diffusion process in which noise is gradually added to an image with desired properties to obtain a completely noisy image, and a dediffusion process in which noise is gradually removed from a noisy image to obtain an image with desired properties. Here, the noise is Gaussian noise. Each step is represented by time t, where t=0 is the state of the image with desired properties, and t=T is the state of the completely noisy image. Note that the diffusion model is inspired by thermodynamics, so the expression "time" is used for convenience, but time does not actually change, and t can also be expressed as the calculation step or the noise intensity level. The diffusion process can be easily performed because it adds noise to the image, but in the dediffusion process, simply subtracting an appropriate amount of Gaussian noise will not, of course, yield an image with desired properties. Therefore, in the diffusion model, a neural network is used for noise removal at each time step of the dediffusion process. By repeatedly performing noise removal using a neural network from the completely noisy image at t=T, we obtain the noisy image at t=0 (the image with desired properties). In the Image to Image diffusion model, in addition to the noise image, a constraint image is also input to the neural network. If the task is upscaling, the constraint image is the low-resolution image before upscaling. The noise image at t=0 (the image with the desired properties) is the high-resolution image after upscaling.By learning the dediffusion process for images of various subjects, the diffusion model can acquire probability distributions for generating various images with desired properties.

[0014] The training of the first machine learning model (determination of model parameters) performed by the training device 101 will be described below with reference to Figure 4. Figure 4 is a flowchart showing the training of the first machine learning model in this embodiment. Model parameters refer to the parameters of the machine learning model determined by training. Examples include the weight coefficients of the fully connected layer, the filter coefficients of the convolutional layer, and the bias.

[0015] In step S101, the acquisition unit 113 acquires one or more training input images and ground truth images from the storage unit 111. Since the task performed in this embodiment is upscaling, the training input images are low-resolution images before upscaling, and the ground truth images are high-resolution images. By downscaling the high-resolution images, low-resolution images containing the same subject can be generated. The high-resolution images may be real photographs or computer graphics (CG) may be used.

[0016] In step S102, the calculation unit 114 generates a first map based on the training input image using a second machine learning model. The second machine learning model is a pre-trained machine learning model that performs segmentation on the input image. Furthermore, segmentation on the training input image does not necessarily have to be performed using a machine learning model; a rule-based method may also be used.

[0017] Figure 5 is a diagram illustrating the flow of output image generation in this embodiment. The training input image 200 is input to the second machine learning model 212, and the first map 201 is generated. The first map 201 is the result of segmentation corresponding to the training input image 200 and contains multiple labels. Each of the multiple labels represents the meaning of a different subject. Figure 6 is an explanatory diagram of the first map 201. Figure 6(A) is an example of the training input image 200, and Figure 6(B) is the first map 201 corresponding to the training input image 200 in Figure 6(A). The labels of the first map 201 may be represented by pixel values ​​or by channels. For example, if the pixel value of the first map 201 is 0, it may be represented as label 1, if it is 1, as label 2, if it is 2, as label 3, and so on. Alternatively, the labels may be represented by a combination of the values ​​of multiple channels. Furthermore, the probabilities of the first channel of the first map 201 corresponding to label 1, the probabilities of the second channel corresponding to label 2, and the probabilities of the third channel corresponding to label 3 may also be expressed in this way.

[0018] In this embodiment, the multiple labels included in the first map 201 include labels representing the material of the subject. In Figure 6(B), skin, hair, and cloth are examples of materials. Other examples of materials include sky, paper, wood (trunk and branches), leaves, water, fire, stone, soil, gravel, leather, metal, asphalt, concrete, and brick. The first map 201 and the text 202 describing the meaning of each of the multiple labels included in the first map 201 are used for upscaling by the first machine learning model 211. Since the first machine learning model 211, which is a generative model, aims to produce an output image with a texture that is appropriate for the subject and has minimal inconsistencies, it is desirable that the labels represent materials linked to the structure of the texture. However, labels representing the meaning of individual people, cars, dogs, etc., may also be used. Furthermore, labels that further classify materials may be used. For example, skin may be divided into races such as East Asian, Black, and White. Cloth may also be further divided into shirts, ties, etc.

[0019] Furthermore, in this embodiment, the multiple labels included in the first map 201 include a label representing a defocused subject. In Figure 6(B), the defocused subject corresponds to this label. If the first map 201 were to assign labels representing materials, etc., to subjects that are outside the depth of field of the training input image 200, the first machine learning model 211 might generate high-resolution textures even for subjects outside the depth of field. In this case, the depth of field would change before and after upscaling. By including a label representing a defocused subject in the first map 201, it is possible to suppress the generation of high-resolution textures for subjects outside the depth of field. Since the purpose is to suppress texture generation, the meaning of the labels may be "blurred subject," "subject without texture," "gradient," "flat subject," etc.

[0020] In step S103, the calculation unit 114 generates the ground truth noise and noise image at time t. Here, time t is set to any value from 1 to T, excluding 0. The ground truth noise is all the noise added during the diffusion process between time 0 and t. The noise image at time t is the image with the noise added to the ground truth image. The model parameters of the first machine learning model 211 are updated using the ground truth noise and noise image. If multiple ground truth images are obtained in step S101, ground truth noise and noise image are generated for each of them. Different values ​​for time t may be set for each ground truth image.

[0021] In step S104, the calculation unit 114 uses the first machine learning model 211 to estimate the noise (estimated noise) added up to time t, based on the training input image 200, the first map 201, the text 202, and the time t set in step S103. The text 202 relates to the meaning of each of the multiple labels contained in the first map 201. The text 202 may include the relationship between the meaning of each of the multiple labels and the value (or channel) of the first map 201 corresponding to that meaning. For example, if the first map 201 is a one-channel map, the text 202 may say, "Value 0.0 is skin. Value 0.1 is hair. ... Value 1.0 is a defocused subject." If the first map 201 represents the probability of each of the multiple channels corresponding to a label, the text 202 may say, "Channel 1 is the probability of skin. Channel 2 is the probability of hair. ..." The text 202 may also include the image caption. A caption, for example, in Figure 6(B), is a descriptive text such as, "There is one man behind two men standing side by side, and the background is out of focus and blurred."

[0022] Figure 7 shows the structure of the first machine learning model 211. However, the configuration of the neural network is not limited to this. In this embodiment, the first machine learning model 211 uses a foundation model, which is pre-trained on a large amount of training data including a large number of images and their descriptions, as the initial model during training. Examples of foundation models that link language and images include Stable Diffusion and Midjourney. By further training the foundation model, it becomes possible to easily generate high-quality images. However, it is also possible to prepare a large number of pairs of images and their descriptions and train the first machine learning model 211 from the beginning.

[0023] The first machine learning model 211 includes an image generation model 231, a time encoder 232, and a text encoder 233. The noise image 220 at time 204, the training input image 200, and the first map 201 are concatenated in the channel direction and input to the image generation model 231. Here, the number of vertical and horizontal pixels in the noise image 220 is the same as that of the ground truth image (after upscaling), and therefore greater than that of the training input image 200 (before upscaling). Therefore, the vertical and horizontal pixels of the noise image 220 are rearranged in the channel direction to match the number of pixels. For example, if the number of vertical and horizontal pixels in the noise image 220 is 2H × 2W and the number of pixels in the training input image 200 is H × W, the noise image 220 is rearranged in the channel direction to H × W × 4. However, instead of rearranging the noise image 220, the training input image 200 and the first map 201 may be enlarged by bilinear interpolation or the like to match the number of vertical and horizontal pixels. Furthermore, the image and map input to the image generation model 231 may be vectorized by rearranging the pixels either vertically or horizontally. For example, if the number of pixels in the vertical, horizontal, and channel is H × W × C, they may be rearranged to HW × 1 × C. Note that concatenation in the channel direction before inputting to the image generation model 231 is not mandatory. For example, the noise image 220, the training input image 200, and the first map 201 may each be input to different convolutional layers within the image generation model 231 to be converted into features, and these features may be concatenated in the channel direction. Note that the features are vectors or maps of two or more dimensions.

[0024] Time 204 is input to the time encoder 232, converted into time features, and then input to the image generation model 231. In Figure 7, the arrow is shown with a dashed line for easier distinction from other arrows. The time encoder 232 has one or more full connection layers. The time features are processed in the residual block (Res Block in Figure 7) within the image generation model 231 by performing at least one of element-wise summation or multiplication with the features. In the first residual block, at least one of element-wise summation or multiplication is performed with the transformed features of the noise image 220, the training input image 200, and the first map 201, and the time features. Alternatively, the features may be concatenated in the channel direction instead of summation or multiplication. The residual block has one or more full connection layers or convolutional layers, activation functions, and skip connections.

[0025] Text 202 is input to the text encoder 233, converted into text features, and then input to the image generation model 231. In the attention block (Attn Block in Figure 7) within the image generation model 231, cross-attention is performed with the features output from the residual block or the features converted within the attention block. In cross-attention, a query is generated from the features output from the residual block or the features converted within the attention block, and a queue and value are generated from the text features. The attention block has cross-attention and may also have self-attention, a fully connected layer, a convolutional layer, an activation function, etc. Note that the number and order of residual blocks and attention blocks are not limited to the configuration shown in Figure 7.

[0026] In Figure 7, Downscale refers to the process of reducing the number of pixels in at least one of the vertical or horizontal directions of the feature. Pixels may be rearranged in the channel direction, or pooling or bilinear interpolation may be used. Upscale refers to the process of increasing the number of pixels in at least one of the vertical or horizontal directions of the feature. Note that the presence or absence of Downscale and Upscale, and their number, are not limited to the configuration shown in Figure 7. Concat. represents concatenation in the channel direction. Instead of Concat., element-wise summation may be taken. Subtract is the element-wise difference, and subtracting the estimated noise 203 from the noise image 220 results in the output image 221. The output image 221 is an upscaled image corresponding to the training input image 200. Note that in Figure 7, the estimated noise 203 and the output image 221 are in a state where the vertical and horizontal pixels have been rearranged in the channel direction.

[0027] Note that the noise image at time t=0 may be used as the residual image of the ground truth image and the training input image 200, rather than the ground truth image. If the task is upscaling, the training input image 200 is enlarged using bilinear interpolation or the like before the residuals are taken. In this case, the result of subtracting the estimated noise 203 from the noise image 220 is the residual image, which is then added to the enlarged training input image to obtain the output image 221.

[0028] In step S105, the update unit 116 updates the model parameters of the first machine learning model 211 based on the estimated noise 203 and the ground truth noise. In this embodiment, the mean square error of the estimated noise 203 and the ground truth noise is used, but the mean absolute error or the like may also be used. When calculating the error, the number of vertical and horizontal pixels is matched by rearranging the pixels in the channel direction of the estimated noise 203 vertically and horizontally, or by rearranging the pixels in the vertical and horizontal directions of the ground truth noise vertically and horizontally in the channel direction. Alternatively, the error may be calculated from the output image 221 and the ground truth image. In this embodiment, only the model parameters in the image generation model 231 are updated, but the model parameters of the time encoder 232 and the text encoder 233 may also be updated. Backpropagation or the like may be used for the update.

[0029] In step S106, the determination unit 115 determines whether the training of the model parameters of the first machine learning model 211 is complete. Specifically, the determination unit 115 determines the completion of training based on the number of updates, the magnitude of the model parameter updates, etc. If it is determined that training is not complete, one or more new sets of training input images and correct images are acquired in step S101. If it is determined that training is complete, the storage unit 111 stores the model parameters.

[0030] Through the above process, it is possible to train a generative model in which images are included in both the input and output, and in which the quality of the output image is improved.

[0031] The following describes the upscaling of captured images using a first trained machine learning model, with reference to Figure 8. Figure 8 is a flowchart showing the estimation by the first machine learning model.

[0032] In step S201, the acquisition unit 125 acquires the captured image. The captured image may be the entire area of ​​the image, or a partial area of ​​the entire area.

[0033] In step S202, the calculation unit 126 uses a second machine learning model to generate a first map which is the result of segmenting the captured image and includes multiple labels. The process to obtain the output image obtained by upscaling the captured image is the same as in Figure 5 during training, except that the training input image 200 is replaced with the captured image. As described in step S102, a rule-based method may also be used to generate the first map. The generated first map includes labels indicating the material of the subject and labels indicating the defocused subject, as in training. The effect is as explained in step S102.

[0034] In step S203, the calculation unit 126 generates a noise image. Here, the noise image is a complete Gaussian noise image corresponding to time t = T. Depending on the noise distribution in the noise image, subtle changes occur in the texture and other elements generated in the output image. The noise image may be generated randomly from a fixed seed value, or the date and time when this step is executed may be used as the seed value to generate a noise image with a different distribution each time. Alternatively, a noise image that has been generated in advance and stored in the storage unit 123 may be retrieved.

[0035] In step S204, the calculation unit 126 generates an output image based on the captured image, the first map, and text describing the meaning of each of the multiple labels, using a first machine learning model. The output image is an upscaled image corresponding to the captured image. The text is the same as described in step S104. The first map and text clarify the meaning and spatial position of the subjects present in the captured image. This makes it possible to generate an upscaled image while suppressing deviations in the positional relationships between subjects from the captured image and the generation of unnatural textures that do not correspond to the original subjects.

[0036] Similar to Figure 7 during training, the noise image at t=T is used at the position of noise image 220, and the captured image is used at the position of training input image 200. In this case, the output image at t=T (corresponding to output image 221 in Figure 7) is estimated, but the estimation accuracy tends to be low. Therefore, a more accurate output image may be generated using the iterative calculation shown below. The noise image at t=T is generated by multiplying the estimated noise at t=T (corresponding to estimated noise 203 in Figure 7) by a coefficient and subtracting it from the noise image at t=T. A small noise component may be introduced at this time. The noise image 220 and time 204 in Figure 7 are changed to the generated noise image at t=T-1 and time t=T-1, and the estimated noise at t=T-1 is estimated. By repeating this, an output image corresponding to t=0 is generated. The output image generated by this method has high estimation accuracy, but the calculation time until generation is long, so the range of time advanced at one time may be widened. For example, the time may be reduced from t=T to T-10, T-20, etc.

[0037] The generated output image (upscaled image) is generated based on information regarding the focus of the captured image. As described in steps S102 and S202, the first map includes labels representing defocused subjects. In other words, labels other than these indicate focused subjects. By using information indicating whether or not a subject is in focus, it is possible to obtain an upscaled image in which the depth of field does not deviate from that of the captured image. Furthermore, when generating the first map, it is advisable to add labels representing defocused subjects to areas where labels could not be determined or where the determination accuracy was low. Since no texture is generated in areas labeled as defocused subjects, it is possible to suppress the generation of unnatural structures that do not match the subject.

[0038] As described above, according to the configuration of this embodiment, the quality of the output image can be improved in a generation model in which images are included in the input and output.

[0039] Figure 9 is a block diagram of the image processing system 300 in this embodiment. Figure 10 is an external view of the image processing system 300. The image processing system 300 includes a training device 301, a sharpening device 302, an imaging device 303, and a lens device 304. The training device 301 includes a storage unit 311, a communication unit 312, an acquisition unit 313, a calculation unit 314, a determination unit 315, and an update unit 316. The sharpening device 302 includes a storage unit 321, a communication unit 322, an acquisition unit 323, a calculation unit 324, and a display unit 325. The imaging device 303 includes an image sensor 331, a storage unit 332, a communication unit 333, a calculation unit 334, and a display unit 305. The lens device 304 includes an optical system 341, a storage unit 342, and a communication unit 343.

[0040] The imaging device 303 is an interchangeable-lens camera and is configured to be connectable to multiple types of lens devices 304. The subject image formed by the lens device 304 is photoelectrically converted by the image sensor 331. Each pixel of the image sensor 331 has two photoelectric conversion units arranged horizontally. The two photoelectric conversion units divide the pupil of the optical system 341 horizontally, and light beams from the divided partial pupils are incident on each photoelectric conversion unit. As a result, two parallax images with parallax are acquired. The calculation unit 334 generates a parallax map of the subject space from the two parallax images. The amount of parallax shift at each pixel of the parallax map represents the amount of defocus. The calculation unit 334 also generates an image by combining the two parallax images. The image contains the effects of blur due to aberrations and diffraction generated in the optical system 341. In this embodiment, the blur due to aberrations and diffraction is sharpened by a first machine learning model. However, the same method can be applied to sharpening other types of blur, such as defocus blur and image smudge. Furthermore, similar effects can be obtained for image transformation tasks other than blur sharpening. The captured image and parallax map are stored in the memory unit 332. Information regarding the type of lens device 304 and the state of the optical system 341 (focal length, F-number, and focus distance) at the time the image was captured is acquired via the communication units 333 and 343 and written to the metadata of the captured image. Information regarding the type of lens device 304, etc., is stored in the memory unit 342.

[0041] The sharpening device 302 uses a first machine learning model to sharpen the blur in the captured image and generate a sharpened image (output image). In this embodiment, a diffusion model is used for the first machine learning model. However, other generative models may also be used. In this embodiment, blur sharpening is performed in advance by a third machine learning model, which is a non-generative model, and then processing is performed by the first machine learning model, which is a generative model. By performing blur sharpening in advance, the burden on the first machine learning model can be reduced. The model parameters of the first and third machine learning models are trained by the training device 301, and are acquired in advance by the sharpening device 302 via the communication units 312 and 322 and stored in the storage unit 321. The sharpened image undergoes other necessary processing by the calculation unit 324 and is stored in the storage unit 321 or storage unit 332, or displayed in the display unit 325 or display unit 335.

[0042] The training of the third machine learning model performed on the training device 301 will now be described with reference to Figure 11. Figure 11 is a flowchart showing the training of the third machine learning model in this embodiment. Since multiple types of lens devices 304 can be connected to the imaging device 303, different training is performed for each type of lens device 304, and different model parameters are determined for each.

[0043] In step S301, the acquisition unit 313 acquires one or more sets of training input images and ground truth images from the storage unit 311. The training input image is an image blurred by aberration occurring in the optical system 341 or blur caused by diffraction. The ground truth image is an image with smaller blur than that of the training input image, or an image with no blur. The pair of the training input image and the ground truth image may be prepared by actual shooting, or may be prepared by imaging simulation. In the case of actual shooting, the ground truth image can be prepared, for example, by imaging the same subject with an optical system having higher resolution performance than the optical system 341. In the case of imaging simulation, a training input image is generated by preparing an original image as a subject, and adding blur caused by aberration occurring in the optical system 341 and diffraction to the original image. A ground truth image can be prepared by adding blur smaller than the blur caused by aberration and diffraction occurring in the optical system 341, or using the original image as it is. The plurality of training input images stored in the storage unit 311 include various types of blur occurring in the optical system 341. Blur caused by aberration and diffraction changes depending on the focal length, F-number, and focus distance of the optical system 341, as well as image plane coordinates, the optical low-pass filter of the image sensor 331, and the pixel pitch. Blur obtained when these parameters take various values is included in the plurality of training input images.

[0044] In step S302, the arithmetic unit 314 generates a second intermediate image based on the training input image using a third machine learning model. Specifically, the training input image is input to the third machine learning model to generate the second intermediate image. The network structure of the third machine learning model is, for example, ResNet, U-net, or the like. Information specifying the blur acting on the training input image may be input to the third machine learning model together with the training input image. The information specifying the blur is, for example, the focal length, F-number and focus distance of the optical system 341, the image plane coordinates, and the optical low-pass filter and pixel pitch of the image sensor 331. The effect of blur sharpening can be improved by inputting the information specifying the blur. Note that the second intermediate image may be a residual image that becomes an image with sharpened blur by being added to the training input image.

[0045] In step S303, the updating unit 316 updates the model parameters of the third machine learning model based on the second intermediate image and the ground truth image. In this embodiment, the mean square error between the second intermediate image and the ground truth image is used as the loss function, but mean absolute error or the like may also be used. When the second intermediate image is a residual image, an error between the second intermediate image and the ground truth image is calculated after adding the second intermediate image to the training input image. Other descriptions that are the same as those in step S105 are omitted here.

[0046] In step S304, the determining unit 315 determines whether the training of the third machine learning model is completed. The model parameters obtained after the training is completed are stored in the storage unit 311. Other descriptions that are the same as those in step S106 are omitted here.

[0047] Hereinafter, the training of the first machine learning model will be described with reference to FIG. 12. FIG. 12 is a flowchart showing the training of the first machine learning model according to the present embodiment.

[0048] In step S401, the acquiring unit 313 acquires one or more sets of training input images, a second map, and a ground truth image from the storage unit 311. The second map is a map representing a focused region of the training input image. Other descriptions that are the same as those in step S301 are omitted here.

[0049] In step S402, the arithmetic unit 314 uses the third machine learning model to generate a second intermediate image based on the training input image. FIG. 13 is a diagram showing a flow of generating an output image according to the present embodiment. The trained third machine learning model 413 is used, and the model parameters of the third machine learning model 413 are not updated thereafter. Since the second intermediate image 402 is an image sharpened by a non-generative model, edges are sharpened with respect to the training input image 400, but the sharpening effect for textures and the like may be insufficient.

[0050] In step S403, the calculation unit 314 generates a first map 401 based on the training input image 400 using the second machine learning model 412. The training input image 400 may be input to the second machine learning model 412, but in this embodiment, a second intermediate image 402 (an image based on the training input image 400) is input to the second machine learning model 412 to generate the first map 401. If the degradation of the training input image 400 (blurring in this embodiment) is large, the accuracy of segmentation may decrease, so it is desirable to generate the first map 401 from a second intermediate image 402 in which the degradation has been suppressed. In this embodiment, the first map 401 does not include labels indicating whether or not there is focus. Instead, a second map 404 representing the focused region of the training input image 400 is input to the first machine learning model 411 to suppress the generation of textures in the defocused region.

[0051] In step S404, the calculation unit 314 generates the correct noise and noise image at time t. The explanation is the same as in step S103 and will be omitted.

[0052] In step S405, the calculation unit 314 uses the first machine learning model 411 to estimate the noise (estimated noise) added up to time t. In this embodiment, the network configuration shown in Figure 14 is used as the first machine learning model 411, but is not limited to this. The first machine learning model 411 has an image generation model 431, a time encoder 432, and a text encoder 433, and the image generation model 431 generates the estimated noise 406. In this embodiment, the noise image 420, the second intermediate image 402, and the second map 404 are linked in the channel direction and input to the image generation model 431. If the second intermediate image 402 is a residual image added with the training input image 400, it is desirable to also input the training input image 400. By inputting the second map 404 to the image generation model 431, it is possible to identify the areas where texture should be generated (focus areas) and areas where it should not be generated (defocus areas), thereby suppressing deviations in depth of field from the training input image 400. Time 405 is converted into a time feature by the time encoder 432, similar to Example 1, and input to the residual block of the image generation model 431. Text 403 is converted into a text feature by the text encoder 433, similar to Example 1, and input to the attention block of the image generation model 431. The first map 401 is also input to the attention block. In the attention block, the feature output from the residual block, or the feature converted within the attention block, and the first map 401 are concatenated in the channel direction, and then cross-attention is taken with the text feature. If there is vertical or horizontal scaling (Upscale or Downscale in Figure 7) within the image generation model 431, the first map 401 should also be scaled to match the number of pixels in the feature for which cross-attention is taken. Furthermore, the method of input to the attention block is not limited to this. For example, the first map 401 may be converted into a map feature using a map encoder, and then cross-attention may be taken. Furthermore, the second map 404 may be input in the same way as the first map 401. In this embodiment, since the task is to sharpen blur caused by aberrations and diffraction, the depth of field of the training input image 400 and the ground truth image are the same.If the task is to sharpen blur caused by defocusing (expand depth of field) or refocus, the depth of field of the training input image 400 and the ground truth image will not match. If you want to perform these tasks, the second map 404 will be a map representing the desired focus region of the output image 421, rather than the training input image 400. Alternatively, as in step S302, information identifying the blur of the training input image 400 may be input to the first machine learning model 411.

[0053] In step S406, the update unit 316 updates the model parameters of the first machine learning model 411 based on the estimated noise 406 and the ground truth noise. The same explanation as in step S105 is omitted.

[0054] In step S407, the determination unit 315 determines whether the training of the first machine learning model 411 is complete. If training is complete, the model parameters are stored in the storage unit 311. Further explanations similar to those for step S106 are omitted.

[0055] The following describes the sharpening of blurred images performed by the sharpening device 302, with reference to Figure 15. Figure 15 is a flowchart showing the estimation by the first and third machine learning models.

[0056] In step S501, the acquisition unit 323 acquires the captured image and the disparity map.

[0057] In step S502, the calculation unit 324 generates a second map showing the focused region of the captured image based on the parallax map. As mentioned above, the amount of parallax shift in the parallax map is linked to the amount of defocus. The second map is generated with the region where the amount of defocus falls within the depth of field as the focused region. In this embodiment, since the image sensor 331 can acquire a parallax image, the second map was generated based on the parallax map, but the second map may be generated by other methods. For example, the imaging device 303 may have an optical system and an image sensor capable of acquiring a depth map of the subject space, and the second map may be generated from the depth map and the focus distance when the captured image was taken. Alternatively, the depth map and focus map may be estimated from the captured image using a machine learning model.

[0058] If the task is to increase the depth of field or refocus, the second map represents the desired focused region of the output image. In this case, if a map representing the amount of defocus (including information about the focus of the captured image) is known, the second map can be generated from that map. If the depth map is known, the second map can be generated by combining it with the focus distance of the captured image. Therefore, even when the second map represents the focused region of the output image, the second map is based on information about the focus of the captured image.

[0059] In step S503, the calculation unit 324 generates a second intermediate image based on the captured image using a trained third machine learning model. The second intermediate image is an image in which the blur caused by aberrations and diffraction generated in the optical system 341 has been sharpened using a non-generating model. The type of lens device 304 used when the captured image was taken is obtained from the metadata of the captured image, and model parameters corresponding to the type of lens device 304 are used in the third machine learning model. The training input image 400 in Figure 13 is replaced with the captured image, and the same process is performed. If information to identify blur is input to the third machine learning model during training, the same information is input during estimation. Information regarding the state of the optical system 341 (focal length, F-number, and focus distance) when the captured image was taken is obtained from the metadata of the captured image.

[0060] In step S504, the calculation unit 324 generates a first map based on a second intermediate image (an image based on the captured image) using a second machine learning model. By performing segmentation on an image that has been blurred and sharpened using a non-generative model, improved accuracy can be expected compared to directly segmenting the captured image.

[0061] In step S505, the calculation unit 324 generates a noise image. The same explanation as in step S203 is omitted.

[0062] In step S506, the calculation unit 324 generates an output image using the trained first machine learning model. Similar to step S503, model parameters corresponding to the type of lens device 304 used to capture the image are used in the first machine learning model. The first machine learning model is input with the same information as during training: a noise image, a second intermediate image (an image based on the captured image), a first map, text representing the meaning of each of the multiple labels, and a second map. By inputting a second map representing the focused region of the captured image (or the output image in the case of depth of field expansion or refocusing), it is possible to generate an output image that suppresses the unnatural texture of the defocused region.

[0063] Alternatively, instead of inputting the second map into the first machine learning model, the first intermediate image generated by the first machine learning model and the second intermediate image (or captured image) may be combined based on the second map to generate the output image. Since the second map is not input into the first machine learning model, textures may be generated even in the defocused areas of the first intermediate image. By reducing the weight of the first intermediate image in the defocused areas in the second map and combining it with the second intermediate image (or captured image), an output image with suppressed unnatural textures can be generated. Furthermore, the intensity of textures generated for a specific subject may be adjusted using the first map. For example, skin areas can be identified from the first map, and the texture of skin areas can be suppressed by reducing the weight of the first intermediate image in the skin areas and combining it with the second intermediate image (or captured image).

[0064] As described above, according to the configuration of this embodiment, the quality of the output image can be improved in a generation model in which images are included in the input and output. [Other Embodiments] The present invention can also be realized by supplying a program that implements one or more of the functions of the above embodiment to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (for example, an ASIC) that implements one or more functions.

[0065] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its gist.

Claims

1. An image processing method characterized by comprising the steps of: generating a first map which is the result of segmentation corresponding to an captured image and includes a plurality of labels; and generating an output image corresponding to the captured image using a first machine learning model based on the captured image, the first map, and text relating to the meaning of each of the plurality of labels.

2. The image processing method according to claim 1, characterized in that the output image is generated based on information regarding the focus of the captured image.

3. The image processing method according to claim 1 or 2, characterized in that a second map representing the focused region of the captured image or the output image is input to the first machine learning model to generate the output image.

4. The image processing method according to any one of claims 1 to 3, characterized in that the text includes the relationship between the meaning of each of the plurality of labels and the value or channel of the first map corresponding to the meaning.

5. The image processing method according to any one of claims 1 to 4, characterized in that the plurality of labels include a label representing a defocused subject.

6. The image processing method according to any one of claims 1 to 5, characterized in that the plurality of labels include labels representing the material of the subject.

7. The image processing method according to any one of claims 1 to 6, characterized in that the output image is generated by combining the captured image and a first intermediate image generated by the first machine learning model using a second map representing the focused region of the captured image or the output image.

8. The image processing method according to any one of claims 1 to 7, characterized in that the first map is generated based on the captured image using a second machine learning model.

9. The image processing method according to any one of claims 1 to 8, further comprising the step of generating a second intermediate image using a third machine learning model based on the captured image, wherein the output image is generated based on the second intermediate image.

10. The image processing method according to claim 9, characterized in that the output image is generated by combining at least one of the captured image and the second intermediate image with a first intermediate image generated by the first machine learning model, using a second map representing the focused region of the captured image or the output image.

11. The image processing method according to claim 9 or 10, wherein the first map is a feature generated based on the second intermediate image.

12. The image processing method according to any one of claims 1 to 11, wherein the first machine learning model comprises an image generation model and a text encoder, and the output image is generated using the image generation model based on the captured image, the first map, and text features obtained by converting the text by the text encoder.

13. The image processing method according to any one of claims 1 to 12, characterized in that the output image is an image that differs from the captured image in at least one of the following: number of pixels, sharpness, contrast, noise, dynamic range, gradation, focus distance, lighting, and perspective.

14. An image processing apparatus characterized by comprising: means for generating a first map which is the result of segmentation corresponding to an captured image and includes a plurality of labels; and means for generating an output image corresponding to the captured image using a first machine learning model based on the captured image, the first map, and text relating to the meaning of each of the plurality of labels.

15. A program characterized by causing a computer to execute the image processing method described in any one of claims 1 to 13.