Image processing method, image processing device, program and image processing system

JP2024021485A5Pending Publication Date: 2025-08-08CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022124335
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-08-03
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing image processing methods using machine learning models for blurred images suffer from accuracy decreases due to brightness saturation, leading to artifacts and incorrect task performance in recognition or regression tasks.

Method used

An image processing method that generates a first map representing the relationship between the expanded area of brightness saturation and signal value magnitude in a blurred image, followed by a second map correction based on saturated region information, using two machine learning models to improve task accuracy.

Benefits of technology

The method effectively suppresses accuracy decreases caused by brightness saturation, enhancing the performance of recognition and regression tasks by accurately distinguishing between saturated and unsaturated blurred areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an image processing method that suppresses task accuracy deterioration due to luminance saturation in a recognition or regression task that uses machine learning on captured images in which blurring has occurred.SOLUTION: An image processing method includes: first step S201 of acquiring a captured image obtained by image-capturing; second step S202 of generating a first map based on the captured image using a first machine learning model; and third step S203 of generating a second map by modifying the first map based on information regarding the position of a luminance-saturated area of the captured image. The first map is a map representing a relationship between extent in which a subject in the luminance-saturated area has expanded due to blurring caused by the image-capturing and a magnitude of a signal value corresponding to the area.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image processing method for performing a recognition or regression task using a machine learning model on a captured image in which blurring has occurred. [Background technology]

[0002] Non-Patent Document 1 discloses a method for sharpening blur in a captured image using a convolutional neural network (CNN), which is one of the machine learning models. A training data set is generated by blurring an image having a signal value equal to or greater than the luminance saturation value of the captured image, and a CNN is trained on the training data set, thereby suppressing adverse effects even around the luminance saturation region and performing blur sharpening. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Li Xu,et al.,Deep Convolutional Neural Network for Image Deconvolution,Advances in Neural Information Processing Systems 27,NIPS2014 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the method disclosed in Non-Patent Document 1, there is a possibility that artifacts (false structures) may occur in the subject at positions unrelated to the brightness saturation. Specifically, the artifacts are localized decreases or increases in signal values ​​that are different from the actual structure of the subject. Details of the artifacts and the reasons for their occurrence will be described later. In tasks for images in which blur other than blur sharpening occurs, the accuracy of the task similarly decreases due to the influence of brightness saturation.

[0005] Therefore, an object of the present invention is to provide an image processing method capable of suppressing a decrease in task accuracy due to brightness saturation in a recognition or regression task using machine learning for a blurred captured image. [Means for solving the problem]

[0006] An image processing method as one aspect of the present invention includes a first step of acquiring an image obtained by imaging, a second step of generating a first map based on the image using a first machine learning model, and a third step of generating a second map by modifying the first map based on information regarding the position of a brightness saturation region in the image, wherein the first map is a map that represents the relationship between the range of an area in which a subject in the brightness saturation region has expanded due to blurring caused by the imaging, and the magnitude of the signal value corresponding to the area.

[0007] Other objects and features of the present invention are illustrated in the following examples. Effect of the Invention

[0008] According to the present invention, it is possible to provide an image processing method capable of suppressing a decrease in task accuracy due to brightness saturation in a recognition or regression task using machine learning for a blurred captured image. [Brief description of the drawings]

[0009] [Figure 1] FIG. 13 is a diagram showing a generation process of a model output in the first embodiment. [Diagram 2] 4 is an explanatory diagram of the relationship between a subject and a captured image and a first map in the first to third embodiments. FIG. [Diagram 3] 4A to 4C are explanatory diagrams of a captured image, a first map, and a model output in the first embodiment. [Figure 4] 1 is a block diagram of an image processing system according to a first embodiment. [Diagram 5]1 is an external view of an image processing system according to a first embodiment. [Figure 6] FIG. 1 is an explanatory diagram of an artifact in the first embodiment. [Figure 7] 1 is a flowchart of training the first and second machine learning models in Examples 1 to 3. [Figure 8] FIG. 1 is a diagram showing the training process of the first and second machine learning models in the first embodiment. [Figure 9] FIG. 2 is an explanatory diagram of a training data set in the first embodiment. [Figure 10] 1 is a flowchart of estimation of a first and a second machine learning model in the first or second embodiment. [Figure 11] FIG. 11 is a block diagram of an image processing system according to a second embodiment. [Figure 12] FIG. 11 is an external view of an image processing system according to a second embodiment. [Figure 13] FIG. 11 is an explanatory diagram of a correction to a first map in the second embodiment. [Figure 14] FIG. 11 is a block diagram of an image processing system according to a third embodiment. [Figure 15] FIG. 11 is an external view of an image processing system according to a third embodiment. [Figure 16] 13 is a flowchart of estimation of the first and second machine learning models in the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the drawings, the same components are designated by the same reference numerals, and duplicated descriptions will be omitted.

[0011] Before describing each embodiment in detail, the problem to be solved by the present invention will be described. In a theory-based method for image recognition or regression tasks, the accuracy of the task may be reduced due to elements ignored by assumptions or approximations. In contrast, in a method using machine learning, a machine learning model is trained using training data that includes those elements, so that estimation based on training data without assumptions or approximations can be realized, thereby improving the accuracy of the task. In other words, in a task for image recognition or regression, a method using a machine learning model can achieve higher accuracy than a theory-based method.

[0012] For example, in a technology for sharpening a blurred captured image, the above-mentioned factor is brightness saturation of the captured image (also called blown-out highlights). In theory-based methods such as the Wiener filter, it is assumed that there is no brightness saturation, so the blur is not properly sharpened around a brightness-saturated region (brightness-saturated region), resulting in problems such as ringing. In contrast, a machine learning method can sharpen the blur even if brightness saturation exists, as in Non-Patent Document 1. However, the method in Non-Patent Document 1 has a problem in that artifacts occur in the corrected image.

[0013] The object of the present invention is to suppress the deterioration of task accuracy (the above-mentioned artifacts) caused by brightness saturation when a recognition or regression task is performed on a blurred captured image using a machine learning model. Here, the blur refers to any of the following: aberration, diffraction, or defocus of the optical system used to capture the captured image, blur due to an optical low-pass filter, blur due to the pixel aperture of the image sensor, blur due to camera shake or subject motion during capture, etc. Or, it refers to a combination of a plurality of these. Also, a recognition task is a task for determining a class corresponding to an input image. For example, examples of recognition tasks include tasks for recognizing properties and meanings within an image, such as a task for classifying a subject in an image into people, dogs, cars, etc., and a task for classifying facial expressions from a face image into smiling faces, crying faces, etc. A class is generally a discrete variable. Also, a class is a signal sequence in which recognition labels, which are scalar values, or recognition labels such as segmentation maps are spatially arranged. In contrast, a regression task refers to a task for determining a signal sequence in which continuous variables corresponding to an input image are spatially arranged. For example, regression tasks include estimating a blur-sharpened image from a blurred image, or estimating a depth map of object space from an image.

[0014] Using FIG. 2(A), the difference in the properties between the surrounding area including the brightness saturation area and the other areas in a captured image in which blur has occurred will be described. FIG. 2(A) is a diagram showing the relationship between the object and the brightness distribution of the captured image. In FIG. 2(A), the horizontal axis is the spatial coordinate, and the vertical axis is the brightness. The solid line is the captured image without blur, and the dashed line is the actual captured image in which blur has occurred. The dashed line represents the brightness distribution before clipping due to brightness saturation. Even if the object 221 is blurred in the imaging process, it only has a brightness equal to or less than the brightness saturation value. Therefore, clipping due to the brightness saturation value does not occur, and the object 221 becomes a non-saturated blurred image 231. On the other hand, the object 222 becomes blurred in the imaging process, and has a brightness equal to or more than the brightness saturation value, so clipping due to the brightness saturation value occurs, and the object 222 becomes a saturated blurred image 232. In the non-saturated blurred image 231, the information of the object is attenuated due to blurring. In contrast, in the saturated blur image 232, information about the subject is attenuated not only by blur but also by clipping of the signal value due to brightness saturation. Therefore, the manner in which information about the subject is attenuated differs depending on whether or not there is brightness saturation. This is the first reason why the characteristics of the surrounding area including the brightness saturation area differ from those of the other areas.

[0015] Next, the second factor that differs in nature will be described. That is, at the edge of the brightness saturation region, false edges that do not actually exist in the subject are generated due to clipping of the signal value. The saturated blur image 232 is originally a smooth distribution of brightness in the region above the brightness saturation value, which is represented by a dashed line, but discontinuous edges are formed due to clipping of the brightness saturation value.

[0016] Furthermore, in the captured image, signal values ​​leak from the subject 222 in the brightness saturation region to its surroundings due to blur. The magnitude and range of this leaked signal value increases as the brightness of the subject 222 in the brightness saturation region increases, but since the signal value is clipped by brightness saturation, the magnitude and range of the leaked signal value are not easily known. Therefore, the third element of the difference in nature is that in the surrounding region including the brightness saturation region, the signal value of the subject and the signal value leaked due to blur cannot be separated (even if the shape of the blur is known).

[0017] These three factors result in different properties between the surrounding area including the brightness saturated area and the other areas, so unless different processing is performed for each area, high-precision tasks cannot be achieved.

[0018] The machine learning model can execute different processes depending on the characteristics of an input image, instead of a uniform process. Therefore, for example, when considering an example of sharpening the blur of a captured image, the machine learning model internally determines whether the region of interest is a blurred image including luminance saturation (saturated blurred image) or a blurred image other than that (non-saturated blurred image), and executes different sharpening processes. In this way, both blurred images can be sharpened. However, the machine learning model may not make a correct judgment. For example, in the saturated blurred image 232 of FIG. 2(A), if the region of interest is near the luminance saturation region, the machine learning model can determine that the region of interest is a region affected by luminance saturation because there is a luminance saturation region near the region of interest. However, when the region of interest is a position 233 away from the luminance saturation region, it is not easy to determine whether the position 233 is affected by luminance saturation, and ambiguity increases. As a result, the machine learning model may make an erroneous judgment at the position 233 away from the luminance saturation region. As a result, when the task is to sharpen the blur, the sharpening process corresponding to the saturated blur image is executed on the non-saturated blur image. At this time, artifacts occur in the image where the blur is sharpened, and the accuracy of the task is reduced. The artifacts will be described in detail in the first embodiment.

[0019] The same is true for tasks other than blur sharpening, where the machine learning model erroneously determines which areas are affected by brightness saturation and which are not, resulting in a decrease in task accuracy. For example, in a recognition task, if a non-saturated blurred image is erroneously determined to be a saturated blurred image, it will be determined that the blurred image is in a state in which signal values ​​leaking out from the brightness saturated area due to blurring have been added, and a feature value different from that of the actual non-saturated blurred image will be extracted, resulting in a decrease in task accuracy.

[0020] Next, the gist of the present invention for solving this problem will be described. In the present invention, a first map is generated from a captured image in which blurring occurs during the imaging process by using a first machine learning model. The first map is a map that represents the relationship between the range of an area in which a subject in a luminance saturated area of ​​the captured image spreads due to blurring caused by imaging, and the magnitude of a signal value corresponding to the area. The first map can also be said to be a map (a signal sequence arranged spatially) that represents the magnitude and range of signal values ​​in an area in which a subject in a luminance saturated area of ​​the captured image spreads due to blurring caused during the imaging process of the captured image. In other words, the first map is a map that represents the spread of luminance values ​​in a high luminance area including the luminance saturated area of ​​the captured image (a map that represents a distribution in which a subject that is so high in luminance that it is saturated in luminance spreads due to blurring caused during the imaging process).

[0021] As an example, the first map for the captured image in Fig. 2(A) is shown by a dashed line in Fig. 2(B). By explicitly having the first machine learning model generate the first map, in a task to be executed thereafter (such as blur sharpening), it is possible to execute the processing to be executed on the area affected by brightness saturation and the processing to be executed on the other areas, respectively, on the appropriate areas. Therefore, by having the first machine learning model generate the first map, the accuracy of the task is improved.

[0022] However, there is a possibility that an erroneous estimation occurs in the generated first map. This will be described with reference to FIG. 3. FIG. 3 is an explanatory diagram of a captured image, a first map, and a model output. The model output in FIG. 3 is a blur-sharpened image in which the blur of the captured image is sharpened. For example, the first map shown by the dashed line in FIG. 3(B) may be estimated for the captured image shown by the dashed line in FIG. 3(A). Since FIG. 3(A) shows a subject with a flat signal distribution less than the brightness saturation value, it is correct that the first map has the same value (first signal value) that represents a non-saturated blur image. However, as shown in FIG. 3(B), a region 241 having a value that represents the influence of brightness saturation may be generated. This is an erroneous estimation that occurs due to the learning method of the machine learning model. The principle of this erroneous estimation will be described in detail in the explanation of the first embodiment. Because region 241 exists in the first map, when blur sharpening is performed based on the first map, an artifact region 242 that does not exist in the actual subject appears in the blur sharpened image (model output), as represented by the solid line in Figure 3(A).

[0023] Therefore, in the present invention, the first map is further modified based on information regarding the position of the brightness saturation region of the captured image to generate a second map. The first map should not have values ​​that represent the influence of brightness saturation at positions unrelated to the brightness saturation region of the captured image. Therefore, the second map can be generated by correcting the erroneous estimation of the first map based on the information regarding the position of the saturated region. This can further improve the accuracy of the task.

[0024] In the following, the stage of determining the weights of a machine learning model based on a training dataset is called training, and the stage of executing a recognition or regression task from a captured image using the machine learning model with the trained weights is called estimation. Machine learning models include, for example, neural networks, genetic programming, and Bayesian networks. Neural networks include CNNs (Convolutional Neural Networks), GANs (Generative Adversarial Networks), RNNs (Recurrent Neural Networks), Transformers, and the like. Example 1 An image processing system 100 according to a first embodiment of the present invention will be described. In this embodiment, the task executed after the second map is generated is to sharpen the blur in a captured image including luminance saturation. The blur to be sharpened is the blur caused by aberration and diffraction occurring in the optical system and an optical low-pass filter. However, the effect of the invention can be obtained in the same way even when the blur caused by pixel aperture, defocus, or shaking is sharpened. Moreover, the invention can be implemented in the same way to obtain the effect for tasks other than blur sharpening.

[0025] FIG. 4 is a block diagram of the image processing system 100 in this embodiment. FIG. 5 is an external view of the image processing system 100. The image processing system 100 has a training device 101 and an image processing device 103 connected by a wired or wireless network. The training device 101 has a storage unit 101a, an acquisition unit 101b, a calculation unit 101c, and an update unit 101d. The image processing device 103 has a storage unit 103a, an acquisition unit 103b, and a calculation unit 103c. The image processing device 103 is connected to an imaging device 102, a display device 104, a recording medium 105, and an output device 106 by wire or wirelessly. The imaging device 102 has an optical system 102a and an image sensor 102b.

[0026] A captured image of a subject space captured using the imaging device 102 is input to the image processing device 103. The captured image is blurred due to aberration and diffraction caused by the optical system 102a in the imaging device 102 and the optical low-pass filter of the imaging element 102b, and information about the subject is attenuated. The image processing device 103 estimates a first map from the captured image using a first machine learning model. Furthermore, the image processing device 103 generates a second map by correcting the first map based on information about the position of the saturated region of the captured image, and generates a blur-sharpened image (model output) from the captured image and the second map using the second machine learning model. The first and second machine learning models are trained by the training device 101, and the image processing device 103 acquires information about the first and second machine learning models from the training device 101 in advance and stores it in the storage unit 103a. Furthermore, the image processing device 103 has a function of adjusting the intensity of blur sharpening. The training and estimation of the first and second machine learning models, and the adjustment of the strength of blur sharpening will be described in detail later.

[0027] A user can adjust the intensity of blur sharpening while checking the blur sharpened image displayed on the display device 104. The blur sharpened image subjected to the intensity adjustment is stored in the storage unit 103a or the recording medium 105, and is output to an output device 106 such as a printer as necessary. It is preferable that each of the training device 101 and the image processing device 103 has a processing means suitable for parallel calculation such as a GPU (Graphics Processing Unit) that can process the machine learning model at high speed. The captured image may be grayscale or may have multiple color components. The captured image may be an undeveloped RAW image or a developed image.

[0028] Next, with reference to Figs. 6(A) to (C), artifacts that occur when blur sharpening is performed using a machine learning model will be described. An artifact is a local decrease or increase in signal value that is different from the actual structure of the subject. Figs. 6(A) to (C) are explanatory diagrams of artifacts, with the horizontal axis indicating spatial coordinates and the vertical axis indicating signal values. Figs. 6(A) to (C) show spatial changes in image signal values, which correspond to the color components R, G, and B (Red, Green, and Blue), respectively. Here, the image is an 8-bit image, and the brightness saturation value is 255.

[0029] In Fig. 6(A) to (C), the dashed line indicates the captured image (blurred image), and the thin solid line indicates the correct image without blur. Since none of the pixels have reached the brightness saturation value, there is no effect of brightness saturation. The dotted line indicates a blur-sharpened image in which the blur of the captured image is sharpened using a conventional machine learning model to which this embodiment is not applied. In the blur-sharpened image represented by the dotted line, the blur of the edge is sharpened, but a decrease in signal value occurs near the center that is not present in the correct image. This decrease occurs at a position away from the edge, not adjacent to it, and the occurrence area is wide, so it is a different problem from undershoot. This is an artifact that occurs when the blur is sharpened.

[0030] Also, as can be seen from a comparison of Figs. 6(A) to (C), the degree of signal value reduction differs depending on the color component. In Figs. 6(A) to (C), the degree of signal value reduction increases in the order of G, R, and B. This tendency is also seen in undeveloped RAW images. Therefore, even though the correct image has flat areas, dark areas colored green occur as artifacts in the conventional blur-sharpened image represented by the dotted line. Note that Figs. 6(A) to (C) show examples in which the signal is reduced from the correct image, but conversely, there are also cases in which the signal value increases.

[0031] The reason why this artifact occurs is that, as mentioned above, the machine learning model misjudged the area affected by brightness saturation from the other area, and misapplied the blur sharpening that should be applied to the saturated blur image to the non-saturated blur image. As can be seen from FIG. 2(A), the greater the brightness of the subject, the larger the absolute value of the blur sharpening residual component (the difference between the captured image and the unblurred captured image). Therefore, if the blur sharpening that should be applied to the saturated blur image is applied to the non-saturated blur image, the signal value will be changed excessively. As a result, as shown by the dotted lines in FIG. 6(A) to (C), an area occurs where the signal value is smaller than that of the correct image (solid line).

[0032] In addition, optical systems for visible light are generally designed to provide the best performance for G among RGB. In other words, because the blur (PSF: point spread function) of R and B is larger than that of G, the edges of a saturated blurred image of a highly luminous subject tend to be colored in R or B (purple fringing corresponds to this). When correcting this saturated blurred image, the residual components of the blur sharpening in R and B are larger than that in G. Therefore, if a non-saturated blurred image is erroneously determined to be a saturated blurred image, the reduction in the signal values ​​of R and B is larger than that in G, and artifacts occur as dark areas colored green, as shown in Figures 6(A) to (C).

[0033] In contrast, the dashed lines shown in Fig. 6(A) to (C) are the results of performing blur sharpening using this embodiment. It can be seen that the occurrence of artifacts is suppressed and the blur is sharpened. This is because the first map and the second map that corrects the erroneous estimation make it difficult for the second machine learning model that performs blur sharpening to erroneously determine areas affected by brightness saturation from other areas. From Fig. 6(A) to (C), it can be seen that the deterioration of task accuracy is suppressed by this embodiment.

[0034] Next, the training of the first and second machine learning models executed by the training device 101 will be described with reference to Fig. 7. Fig. 7 is a flowchart of the training of the first and second machine learning models. Each step in Fig. 7 is executed by the storage unit 101a, the acquisition unit 101b, the calculation unit 101c, or the update unit 101d of the training device 101.

[0035] In step S101, the acquisition unit 101b acquires one or more original images from the storage unit 101a. The original image is an image having a signal value greater than a second signal value. Here, the second signal value is a signal value corresponding to the luminance saturation value of the captured image. However, since the signal value may be normalized when inputting to the first and second machine learning models, the second signal value does not necessarily have to match the luminance saturation value of the captured image. Since the first and second machine learning models are trained based on the original image, it is desirable that the original image is an image having various frequency components (edges, gradations, flat parts, etc. with different orientations and intensities). The original image may be a real-life image or CG (Computer Graphics).

[0036] In step S102, the calculation unit 101c adds blur to the original image to generate a blurred image. The blurred image is an image input to the first and second machine learning models during training, and corresponds to the captured image during estimation. The blur to be added is the blur to be sharpened. In this embodiment, blur caused by the aberration and diffraction of the optical system 102a and the optical low-pass filter of the image sensor 102b is added. The shape of the blur caused by the aberration and diffraction of the optical system 102a varies depending on the image plane coordinates (image height and azimuth). It also varies depending on the magnification, aperture, and focus state of the optical system 102a. When it is desired to train the second machine learning model that sharpens all of these blurs at once, it is preferable to generate multiple blurred images using multiple blurs generated by the optical system 102a. In addition, in the blurred image, a signal value exceeding the second signal value is clipped. This is performed to reproduce the luminance saturation that occurs in the capture process of the captured image. If necessary, noise generated in the image sensor 102b may be added to the blurred image.

[0037] In step S103, the calculation unit 101c sets a first region based on an image based on the original image and a signal value threshold. In the first embodiment, a blurred image is used as an image based on the original image, but the original image itself may be used. The first region is set by comparing the signal value of the blurred image with the signal value threshold. More specifically, the first region is a region where the signal value of the blurred image is equal to or greater than the signal value threshold. In this embodiment, the signal value threshold is the second signal value. Therefore, the first region represents a brightness saturation region of the blurred image. However, the signal value threshold and the second signal value do not necessarily have to match. The signal value threshold may be set to a value slightly smaller than the second signal value (for example, 0.9 times).

[0038] In step S104, the calculation unit 101c generates a first image having a signal value of the original image in a first region. The first image has a signal value different from that of the original image in a region other than the first region. More preferably, the first image has a first signal value in a region other than the first region. In this embodiment, the first signal value is 0, but is not limited thereto. That is, in this embodiment, the first image has the signal value of the original image only in the brightness saturation region of the blurred image, and the signal value of the other region is 0.

[0039] In step S105, the calculation unit 101c applies blur to the first image to generate a first correct answer map. The applied blur is the same as the blur applied to the blurred image. This generates a first correct answer map, which is a map (a signal sequence arranged spatially) that represents the magnitude and range of signal values ​​that have leaked out from a subject in a brightness saturation region of the blurred image to the periphery due to blur. In this embodiment, the first correct answer map is clipped with the second signal value, as with the blurred image, but clipping is not necessarily required.

[0040] In step S106, the acquisition unit 101b acquires a correct model output. In this embodiment, since the task is blur sharpening, the correct model output is an image with less blur than the blurred image. In this embodiment, the correct model output is generated by clipping the original image with the second signal value. If the original image lacks high-frequency components, an image obtained by reducing the original image may be used as the correct model output. In this case, reduction is also performed when generating a blurred image in step S102. Also, step S106 may be executed at any time after step S101 and before step S107. By step S106, training data (training data set in the case of multiple blurred images) used for training the first and second machine learning models are prepared.

[0041] In step S107, the calculation unit 101c generates a first map and a model output based on the blurred image using the first and second machine learning models. FIG. 8 is a diagram showing the training process of the first and second machine learning models. In this embodiment, the configuration shown in FIG. 8 is used in the training of the first and second machine learning models, but is not limited to this. In FIG. 8, a blurred image 251 and a brightness saturation map 252 are input to the first machine learning model 211. The blurred image 251 and the brightness saturation map 252 have a spatial two-dimensional signal distribution, but in FIG. 8, for ease of understanding, they are depicted as a one-dimensional signal distribution in a certain cross section. The brightness saturation map 252 is a map showing a brightness saturated region (where the signal value is equal to or greater than the second signal value) of the blurred image 251. For example, the brightness saturation map 252 can be generated by binarizing the blurred image 251 with the second signal value. In FIG. 8, the blurred image 251 is normalized by the second signal value, and binarized with 1 as a threshold value to generate the brightness saturation map 252. However, the method of generating the brightness saturation map 252 is not limited to this. In addition, the brightness saturation map 252 is not necessarily required. The blurred image 251 and the brightness saturation map 252 are linked in the channel direction and input to the first machine learning model 211. However, this embodiment is not limited to this. For example, the blurred image 251 and the brightness saturation map 252 may be converted into feature maps, and the feature maps may be linked in the channel direction. In addition, information other than the brightness saturation map 252 may be added to the input.

[0042] The first machine learning model 211 and the second machine learning model 212 have a plurality of layers, and a linear sum of the input of the layer and the weight is taken in each layer. The initial value of the weight can be determined by a random number or the like. In this embodiment, a CNN using a convolution of the input and the filter as a linear sum (the value of each element of the filter corresponds to the weight, and may include a sum with a bias) is used as the first and second machine learning models 211 and 212, but is not limited thereto. In addition, in each layer, a nonlinear transformation is performed by an activation function such as a ReLU (Rectified Linear Unit) or a sigmoid function as necessary. Furthermore, the first and second machine learning models 211 and 212 may have a residual block or a Skip Connection (also called a Shortcut Connection) as necessary.

[0043] In the first machine learning model 211, a first map 253 is generated. Next, the first correct answer map 254 and the blurred image 251 are concatenated in the channel direction and input to the second machine learning model 212 to generate a model output 255. Instead of the first correct answer map 254, the first map 253 or a second map obtained by correcting the first map 253 based on information about the position of the saturated region of the blurred image 251 may be input to the second machine learning model 212. Note that the training of the first and second machine learning models 211 and 212 does not need to be performed simultaneously, and may be performed separately.

[0044] Returning to FIG. 7, in step S108, the update unit 101d updates the weights of the first machine learning model 211 and the second machine learning model 212 based on the loss function. In this embodiment, the loss function of the first machine learning model 211 is based on the first map 253 and the first correct answer map 254. The loss function of the second machine learning model 212 is based on the model output 255 and the correct answer model output. As the loss function, MSE (Mean Squared Error) is used, but is not limited to this. Backpropagation or the like can be used to update the weights.

[0045] In step S109, the update unit 101d determines whether the training of the first machine learning model 211 and the second machine learning model 212 is completed. The completion of the training can be determined by whether the number of iterations of the weight update reaches a predetermined number, whether the loss function at the time of update or the amount of change of the weight at the time of update is smaller than a predetermined value, or the like. If it is determined in step S109 that the training is not completed, the process returns to step S101, and the acquisition unit 101b acquires one or more new original images. On the other hand, if it is determined that the training is completed, the update unit 101d ends the training, and stores the configurations and weight information of the first and second machine learning models 211 and 212 in the storage unit 101a.

[0046] By the above-mentioned training method, the first machine learning model 211 can estimate the first map 253 that represents the magnitude and range of signal values ​​spread by blurring of a subject in a brightness-saturated region of the blurred image 251 (captured image at the time of estimation). However, as shown in FIG. 8, an erroneous estimation region 260 due to the learning method may occur in the first map 253.

[0047] The principle of the occurrence of this erroneous estimation will be explained with reference to Figs. 9(A) and (B). Figs. 9(A) and (B) are explanatory diagrams of the training data set in this embodiment, and in Fig. 9(A), the horizontal axis indicates spatial coordinates, and the vertical axis indicates luminance. In Fig. 9(A), the dashed line indicates the blurred image 251, and the solid line indicates the correct model output 256. As in Fig. 2(A), the dashed line indicates the signal value before being clipped by the luminance saturation value. The left side of the blurred image 251 is a non-saturated blurred image because there is no clipping by the luminance saturation value, and the right side of the blurred image 251 is a saturated blurred image. In Fig. 9(B), the dashed line is the first correct map 254 corresponding to Fig. 9(A). For example, assume that the first machine learning model 211 is trained using the area 261 shown in Figs. 9(A) and (B). At this time, the first machine learning model 211 needs to estimate the first correct map 254 from the input blurred image 251. However, since there is no brightness saturation region in the region 261, it is impossible to determine that the left side of the blurred image 251 is a non-saturated blurred image and the right side of the blurred image 251 is a saturated blurred image. Therefore, the trained first machine learning model 211 cannot estimate the first correct map 254 from the blurred image 251, and estimates a solution that minimizes the loss function, for example, the first map 253 as shown by the solid line in FIG. 9(B). This first map 253 has an erroneous estimation region at a position corresponding to the non-saturated blurred image of the blurred image 251. Due to such a principle, the erroneous estimation region 260 shown in FIG. 8 occurs.

[0048] As a learning method for suppressing the occurrence of the erroneous estimation region 260, for example, the following method is conceivable. In this method, the blurred image 251 (the region 261 of the dashed line in FIG. 9(A)) is input to the first machine learning model 211, and only the region 262 of the estimated first map 253 excluding the periphery is used. At this time, the weight of the first machine learning model 211 is updated using a loss function of the first map 253 and the first correct answer map 254 in the region 262. The occurrence of the erroneous estimation region 260 can be suppressed by training the first machine learning model 211 excluding the range where the saturated region outside the region 261 is influenced by blurring. However, since the information used for training the first machine learning model 211 is reduced, it is necessary to expand the region 261 to maintain the accuracy of the training, and there is a problem that the calculation load of the training becomes very large. Therefore, in this embodiment, the erroneous estimation region is suppressed by correcting the first map at the time of estimation after training.

[0049] Next, blur sharpening (estimation) of a captured image using the trained first and second machine learning models executed by the image processing device 103 will be described with reference to Figs. 1 and 10. Fig. 1 is a diagram showing a process of generating a model output. Fig. 10 is a flowchart of estimation of the first and second machine learning models. Each step of Fig. 10 is executed by the memory unit 103a, acquisition unit 103b, or calculation unit 103c of the image processing device 103.

[0050] In step S201, the acquisition unit (acquisition means) 103b acquires the captured image 201, the first machine learning model 211, and the second machine learning model 212. Information on the configurations and weights of the first and second machine learning models 211 and 212 is acquired from the storage unit 103a.

[0051] In step S202, the calculation unit (generation means) 103c generates a first map 203 from the captured image 201 and the brightness saturation map 202 corresponding to the captured image 201 using the first machine learning model 211. The configuration of the first machine learning model 211 is the same as that during training. The first map 203 is a map that represents the magnitude and range of the signal value of the area in which the subject in the brightness saturation area of ​​the captured image 201 spreads due to blurring generated during the imaging process of the captured image 201. However, the first map 203 may have an erroneous estimation area 220 in a position unrelated to the saturated blur image. In general, the saturation signal value of each pixel of the image sensor 102b does not become a constant value due to manufacturing variations. Therefore, when generating the brightness saturation map 202, a value obtained by multiplying the design value of the brightness saturation in the image sensor 102b by a coefficient of 1 or less (such as 0.9. The value may be determined according to the magnitude of manufacturing variations) may be used as the brightness saturation value in all pixels of the captured image 201.

[0052] In step S203, the calculation unit 103c generates a second map 205 by correcting the erroneously estimated region 220 of the first map 203 based on information about the position of the brightness saturation region of the captured image 201. In this embodiment, the first map 203 is corrected by the method shown in FIG. 1, but the method is not limited to this. A MAX filter (maximum value filter) 213 is convolved with the brightness saturation map 202 (fourth map) to generate a third map 204, which is a map for distinguishing between a peripheral region of the brightness saturation region including the brightness saturation region and a region other than the peripheral region. The third map 204 represents a region within a predetermined range from each saturated pixel of the captured image 201 and the other region. In this embodiment, the size of the predetermined range is determined by the filter size of the MAX filter 213. The filter size of the MAX filter 213 may be determined based on the spread of blur generated in the captured image 201. The convolution filter may be other than the MAX filter 213. For example, the third map 204 may be generated by convolving a filter whose elements are all 1 and binarizing the result as zero or non-zero. The third map 204 is a map that distinguishes between the peripheral region of the brightness saturation region including the brightness saturation region and the other region in the captured image 201, and has a value of 1 in the peripheral region including the brightness saturation region and a value of 0 in the other region in this embodiment. The second map 205 is generated by using a multiplication operation 214 for each element of the first map 203 and the third map 204. The multiplication operation 214 with the third map 204 can suppress an erroneous estimation region 220 that exists outside the peripheral region of the brightness saturation region including the brightness saturation region. The correction method of the first map 203 shown in this embodiment is composed of convolution and multiplication, and can be easily executed by a parallel calculation means such as a GPU. Therefore, when the estimation of the first and second machine learning models 211, 212 is executed by a parallel computing means, steps S202 to S204 can be executed continuously by the same parallel computing means, enabling high-speed processing. Note that the third map 204 may be generated before step S203.

[0053] Also, threshold processing may be performed on the first map 203 or the second map 205. For example, threshold processing is effective when a very weak erroneous estimation component exists over a wide area of ​​the first map 203 or the second map 205. Furthermore, the threshold processing is preferably soft thresholding processing so that the first map 203 or the second map 205 does not have discontinuity in values ​​at the threshold boundary. After the soft thresholding processing, the first map 203 or the second map 205 may be rescaled by multiplying it by a coefficient so that the maximum value of the first map 203 or the second map 205 does not change. Note that, for simplicity, FIG. 1 illustrates a case where the captured image 201 has a single color component, but when the captured image 201 has multiple color components, step S203 is performed for each color component.

[0054] In step S204, the calculation unit 103c uses the second machine learning model 212 to generate a model output 206, which is an image in which the blur of the captured image 201 has been sharpened, from the captured image 201 and the second map 205. By using the second map 205 in which the erroneous estimation region 220 is suppressed, the second machine learning model 212 can distinguish between a non-saturated blur image and a saturated blur image with high accuracy. Therefore, the second machine learning model 212 can perform the sharpening of the blur while suppressing the occurrence of artifacts. Note that a method other than machine learning (such as a Wiener filter or a Richardson-Lucy method) may be used for the sharpening of the blur. Since the second map 205 allows the regions of a non-saturated blur image and a saturated blur image to be distinguished with high accuracy, it is preferable to sharpen them by a method suitable for each. For example, the non-saturated blurred image region may be sharpened using a Wiener filter, and only the saturated blurred image region may be sharpened using the second machine learning model 212, and the results of both may be combined.

[0055] In step S205, the calculation unit 103c synthesizes the captured image 201 and the model output 206, which is an image corresponding to the captured image 201, based on the second map 205. The surrounding area including the luminance saturated area of ​​the captured image 201 has attenuation of the information of the object due to luminance saturation compared to other areas, so it is difficult to sharpen the blur (estimate the attenuated object information). Therefore, in the surrounding area including the luminance saturated area, problems associated with the sharpening of the blur (ringing, undershoot, etc.) are likely to occur. In order to suppress this problem, the model output 206 and the captured image 201 are synthesized. At this time, by synthesizing based on the second map 205, it is possible to increase the weight of the captured image 201 only in the surrounding area including the luminance saturated area where problems are likely to occur, while suppressing the deterioration of the blur sharpening effect of the non-saturated blur image. In this embodiment, the synthesis is performed by the following method. The second map 205 is normalized by the second signal value, and this is used as a weight map of the captured image 201 and weighted averaged with the model output 206. At this time, for the model output 206, a weight map obtained by subtracting the weight map of the captured image 201 from a map of all 1 is used. It is also possible to adjust the balance between the blur sharpening effect and the adverse effects by changing the signal value for normalizing the second map 205. Alternatively, a synthesis method may be used in which the model output 206 is replaced with the captured image 201 only in an area where the second map 205 has a value equal to or greater than a predetermined signal value.

[0056] With the above configuration, it is possible to provide an image processing system capable of suppressing accuracy degradation due to brightness saturation in blur sharpening using a machine learning model.

[0057] Next, a description will be given of desirable conditions for obtaining the effect of this embodiment. In step S107, it is desirable to input the first correct answer map 254 to the second machine learning model 212. If the generated first map 253 is input instead of the first correct answer map 254 to train the second machine learning model 212, artifacts may occur in the model output 255. This principle will be described with reference to FIGS. 9(A) and 9(B). The second machine learning model 212 must estimate the correct answer model output 256 in the region 261 from the blurred image 251 in the region 261. The blurred image 251 has a similar signal distribution on the left and right sides, but the signal distribution of the correct answer model output 256 is significantly different on the left and right sides. When the first correct answer map 254 is input to the second machine learning model 212, the difference in the value of the first correct answer map 254 corresponds to the correct answer model output 256. Therefore, the second machine learning model 212 can estimate a model output 255 close to the correct model output 256 by changing the sharpening performed on the blurred image 251 based on the value of the first correct map 254. In contrast, the value of the first map 253 in FIG. 9(B) does not correspond to the difference in the correct model output 256. Therefore, when the first map 253 is input to the second machine learning model 212 instead of the first correct map 254, the second machine learning model 212 may not be able to distinguish between a non-saturated blur image and a saturated blur image in the blurred image 251, resulting in artifacts. Example 2 An image processing system in the second embodiment will be described. In this embodiment, the task executed after the second map is generated is the conversion of the blur of the captured image including luminance saturation. The conversion of the blur is a task of converting the blur caused by defocus acting on the captured image into a blur of a different shape from the defocus blur. For example, when a double line blur or vignetting occurs in the defocus blur, it is converted into a blur represented by a circular disk (a shape with flat intensity) or Gaussian. In the conversion of the blur, the defocus blur is made larger, and the sharpening of the blur (estimation of attenuated subject information) is not performed. The method described in this embodiment can be similarly effective for tasks other than the conversion of the blur.

[0058] FIG. 11 is a block diagram of the image processing system 300 in this embodiment. FIG. 12 is an external view of the image processing system 300. The image processing system 300 includes a training device 301, an imaging device 302, and an image processing device 303. The training device 301 includes a storage unit 311, an acquisition unit 312, a calculation unit 313, and an update unit 314. The image processing device 303 includes a storage unit 331, a communication unit 332, an acquisition unit 333, and a calculation unit 334. The imaging device 302 includes an optical system 312, an image sensor 322, a storage unit 323, a communication unit 324, and a display unit 325. The training device 301 and the image processing device 303, and the image processing device 303 and the imaging device 302 are each connected by a wired or wireless network. The captured image captured by the imaging device 302 is subjected to defocus blur having a shape corresponding to the optical system 321. The captured image is transmitted to the image processing device 303 via the communication unit 324. The image processing device 303 receives the captured image via the communication unit 332, and converts the blur using information on the configuration and weights of the first and second machine learning models stored in the storage unit 331. The information on the configuration and weights of the first and second machine learning models is trained by the training device 301, and is acquired in advance from the training device 301 and stored in the storage unit 331. A blur-converted image (model output) in which the blur of the captured image is converted is transmitted to the imaging device 302, stored in the storage unit 323, and displayed on the display unit 325.

[0059] Next, the training of the first and second machine learning models executed by the training device 301 will be described with reference to the flowchart of Fig. 7, but parts similar to those in the first embodiment will be omitted. Each step in Fig. 7 is executed by the storage unit 311, the acquisition unit 312, the calculation unit 313, or the update unit 314 of the training device 301.

[0060] In step S 101 , the acquisition unit 312 acquires one or more original images from the storage unit 311 .

[0061] In step S102, the calculation unit 313 sets a defocus amount for the original image, and generates a blurred image by adding a defocus blur corresponding to the defocus amount to the original image. The shape of the defocus blur changes depending on the magnification and aperture of the optical system 321. The defocus blur also changes depending on the focus distance of the optical system 321 and the defocus amount of the subject at that time. Furthermore, the defocus blur also changes depending on the image height and azimuth. If it is desired to train a second machine learning model capable of converting all of these defocus blurs at once, it is preferable to generate multiple blurred images using multiple defocus blurs generated by the optical system 321. In addition, in the conversion of the blur, it is desirable that the focus subject that is not defocused remains unchanged before and after the conversion. Therefore, it is necessary to train the second machine learning model so that the focus subject does not change. For this reason, a blurred image when the defocus amount is 0 is also generated. A blurred image with a defocus amount of 0 may not be blurred, or may be blurred due to aberration or diffraction on the focus plane of the optical system 321.

[0062] In step S103, the calculation unit 313 sets a first region based on the blurred image and a threshold value of the signal value.

[0063] In step S104, the calculation unit 313 generates a first image having the signal values ​​of the original image in the first region.

[0064] In step S105, the calculation unit 313 imparts the same defocus blur as the blurred image to the first image, and generates a first correct answer map.

[0065] In step S106, the acquisition unit 312 acquires the correct model output. In this embodiment, the second machine learning model is trained so that the defocus blur is converted into a disk blur (a blur that is circular and has a flat intensity distribution). Therefore, the disk blur is added to the original image to generate the correct model output. However, the shape of the blur to be added is not limited to this. A disk blur having a spread corresponding to the defocus amount of the blurred image is added. The added disk blur is larger than the defocus blur added in the generation of the blurred image. In other words, the disk blur has a lower MTF (modulation transfer function) than the defocus blur added in the generation of the blurred image. Also, when the defocus amount is 0, it is the same as the generation of the blurred image.

[0066] In step S107, the calculation unit 313 generates a first map from the blurred image using the first machine learning model, and generates a model output from the blurred image and the first correct answer map using the second machine learning model.

[0067] In step S108, the update unit 314 updates the weights of the first and second machine learning models from the loss function.

[0068] In step S109, the update unit 314 determines whether the training of the first and second machine learning models is completed. Information on the configurations and weights of the trained first and second machine learning models is stored in the storage unit 311.

[0069] Next, the conversion of blur of a captured image using the trained first and second machine learning models executed by the image processing device 303 will be described with reference to the flowchart of Fig. 10, but the same parts as those in the first embodiment will be omitted. Each step in Fig. 10 is executed by the storage unit 331, the communication unit 332, the acquisition unit 333, or the calculation unit 334 of the image processing device 303.

[0070] In step S201, the acquisition unit 333 acquires a captured image, a first machine learning model, and a second machine learning model.

[0071] In step S202, the calculation unit 334 generates a first map from the captured image using the first machine learning model.

[0072] In step S203, the calculation unit 334 generates a second map by correcting the first map based on information on the position of the brightness saturation area of ​​the captured image. In this embodiment, the first map is corrected based on whether or not the closed space that satisfies a predetermined condition of the first map includes the position of the brightness saturation area of ​​the captured image, thereby generating the second map. This will be described in detail with reference to Figs. 13(A) and 13(B). Fig. 13(A) shows a map obtained by binarizing the first map. The binarization is performed using values ​​that represent a non-saturated blur image, and the shaded area represents a non-saturated blur image, and the white area represents a saturated blur image affected by brightness saturation, or an area of ​​erroneous estimation. For example, when 0 in the first map represents a non-saturated blur image, the area of ​​0 becomes a shaded area, and the non-zero area becomes a white area. Fig. 13(B) shows a brightness saturation map corresponding to the captured image. The white area represents a saturated area of ​​the captured image, and the shaded area represents a non-saturated area. In Fig. 13(A), there are closed spaces 401 and 402 that satisfy a predetermined condition (estimated not to be a non-saturated blurred image) in the first map. If the closed spaces do not include a saturated region of the captured image, it is immediately clear that the closed spaces are erroneously estimated regions. Since the closed space 402 does not include a saturated region of the captured image, the first map is corrected with the closed space 402 as an erroneously estimated region, and a second map is generated.

[0073] In step S204, the calculation unit 334 generates a model output from the captured image and the second map using the second machine learning model. The model output is a blur-transformed image in which the defocus blur of the captured image is converted into a blur with a different shape.

[0074] In step S205, the calculation unit 334 combines the captured image and the model output based on the second map.

[0075] With the above configuration, it is possible to provide an image processing system that can suppress accuracy degradation due to brightness saturation in blur conversion using a machine learning model. Example 3 An image processing system in the third embodiment will be described. In this embodiment, the task executed after generating the second map is to estimate a depth map for a captured image. Since the shape of blur in an optical system changes depending on the amount of defocus, the shape of blur can be associated with depth (amount of defocus). The machine learning model can generate a depth map of the subject space by estimating (explicitly or implicitly) the shape of blur in each region of the input captured image within the model. Note that the method described in this embodiment can be similarly effective for tasks other than estimating a depth map.

[0076] FIG. 14 is a block diagram of an image processing system 500 in this embodiment. FIG. 15 is an external view of the image processing system 500. The image processing system 500 has a training device 501 and an imaging device 502 connected by wire or wirelessly. The training device 501 has a storage unit 511, an acquisition unit 512, a calculation unit 513, and an update unit 514. The imaging device has an optical system 512, an imaging element 522, an image processing unit 523, a storage unit 524, a communication unit 525, a display unit 526, and a system controller 527. The image processing unit 523 has an acquisition unit 523a, a calculation unit 523b, and a blurring unit 523c. In FIG. 15, both the front and back surfaces of the imaging device 502 are drawn. The imaging device 502 forms an image of the subject space via the optical system 521, and the image is acquired as a captured image by the imaging element 522. The captured image is blurred due to aberration and defocus of the optical system 521. The image processing unit 523 generates a depth map of the subject space from the captured image using the first and second machine learning models. The first and second machine learning models are trained by the training device 501, and information on their configuration and weights is acquired in advance from the training device 501 via the communication unit 525 and stored in the storage unit 524. The captured image and the estimated depth map are stored in the storage unit 524 and displayed on the display unit 526 as necessary. The depth map is used to add blur to the captured image and to cut out the subject. A series of controls are performed by the system controller 527.

[0077] Next, the training of the first and second machine learning models executed by the training device 501 will be described with reference to the flowchart of Fig. 7, but the same parts as those in the first embodiment will be omitted. Each step in Fig. 7 is executed by the storage unit 511, the acquisition unit 512, the calculation unit 513, or the update unit 514 of the training device 501.

[0078] In step S101, the acquisition unit 512 acquires one or more original images.

[0079] In step S102, the calculation unit 513 applies blur to the original image to generate a blurred image. The calculation unit 513 sets a depth map (or a defocus map) corresponding to the original image and the focus distance of the optical system 521, and applies blur corresponding to the focus distance of the optical system 521 and the defocus amount therefrom. When the aperture value is fixed, the larger the absolute value of the defocus amount, the larger the blur caused by defocus. Furthermore, due to the influence of spherical aberration, the shape of the blur changes before and after the focus plane. When the spherical aberration is in the negative direction, in the subject space, in the direction away from the focus plane to the optical system 521 (object side), it becomes a two-line blur, and in the direction approaching (image side), it becomes a blur with a shape having a peak at the center. When the spherical aberration is positive, the relationship is reversed. Furthermore, outside the optical axis, the shape of the blur changes further according to the defocus amount due to the influence of astigmatism and the like.

[0080] In step S103, the calculation unit 513 sets a first region based on the blurred image and the signal threshold value.

[0081] In step S104, the calculation unit 513 generates a first image having the signal values ​​of the original image in the first region.

[0082] In step S105, the calculation unit 513 blurs the first image and generates a first correct answer map. In this embodiment, the first correct answer map is not clipped by the second signal value. This allows the first machine learning model to be trained to estimate the luminance of the luminance saturated region before clipping when generating the first map.

[0083] In step S106, the acquisition unit 512 acquires the correct model output, which is the depth map set in step S102.

[0084] In step S107, the calculation unit 513 generates a first map from the blurred image using the first machine learning model, and generates a model output from the blurred image and the first correct answer map using the second machine learning model.

[0085] In step S108, the update unit 514 updates the weights of the first and second machine learning models using the loss function.

[0086] In step S109, the update unit 514 determines whether the training of the first and second machine learning models is completed.

[0087] Next, the estimation of a depth map of a captured image using the first and second machine learning models and the addition of blur to the captured image, which are executed by the image processing unit 523, will be described with reference to the flowchart of FIG. 16, but the same parts as those in the first embodiment will be omitted. FIG. 16 is a flowchart of the estimation of the first and second machine learning models. Each step of FIG. 16 is executed by the acquisition unit 523a, the calculation unit 523b, or the blurring unit 523c of the image processing unit 523.

[0088] In step S401, the acquisition unit 523a acquires a captured image, a first machine learning model, and a second machine learning model. From the storage unit 524, information on the configurations and weights of the first and second machine learning models is acquired.

[0089] In step S402, the calculation unit 523b generates a first map from the captured image using the first machine learning model.

[0090] In step S403, the calculation unit 523b generates a second map by correcting the first map based on information about the position of the brightness saturation region of the captured image. The correction is performed in the same manner as in the first embodiment.

[0091] In step S404, the calculation unit 523b generates a model output from the captured image and the second map using the second machine learning model. The model output is a depth map corresponding to the captured image.

[0092] In step S405, the blurring unit 523c blurs the captured image based on the model output and the second map to generate an image with a blurred effect (a shallow depth of field). From the depth map, which is the model output, a blur is set for each region of the captured image according to the defocus amount. No blur is added to the focus region, and a larger blur is added to the region with a larger defocus amount. In addition, the second map estimates the luminance of the luminance saturated region of the captured image before clipping. The signal value of the luminance saturated region of the captured image is replaced with this luminance, and then blur is added. This makes it possible to generate an image with a natural blur, without the addition of blur to sunlight filtering through the trees, reflected light from the water surface, or lights in a night view, etc.

[0093] With the above configuration, it is possible to provide an image processing system capable of suppressing accuracy degradation due to brightness saturation in estimating a depth map using a machine learning model. (Other Examples) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-mentioned embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0094] According to each embodiment, an image processing method, an image processing device, and an image processing program can be provided that are capable of suppressing accuracy degradation due to brightness saturation in recognition or regression tasks using a machine learning model for a blurred captured image.

[0095] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.

[0096] For example, an image processing system may be configured such that an imaging device (first device) of each embodiment and a device (second device) on a cloud are capable of communicating with each other, and the second device executes the process of FIG. 10 or FIG. 16 based on a request from the first device. In this case, the first device has a transmitting means for transmitting a captured image and a request for execution of processing to the second device. The second device has a receiving means for receiving the captured image and the request from the first device, and a generating means for generating a first map based on the captured image using a first machine learning model in response to the received request. Then, the generating means generates a second map by modifying the first map based on information regarding the position of the brightness saturation region of the captured image. Furthermore, the generating means generates a model output based on the captured image and the second map using a second machine learning model.

[0097] The disclosure of each of the above embodiments includes the following methods and configurations.

[0098] (Method 1) A first step of acquiring a captured image obtained by imaging; a second step of generating a first map based on the captured images using a first machine learning model; and a third step of generating a second map by modifying the first map based on information about a position of a brightness saturation region of the captured image, an image processing method characterized in that the first map is a map that represents the relationship between the range of an area in which a subject in the brightness saturation area has expanded due to blurring caused by the imaging, and the magnitude of the signal value corresponding to that area. (Method 2) and a fourth step of generating a model output based on the captured images and the second map using a second machine learning model; The image processing method according to Method 1, wherein the model output is a recognition label or a spatially arranged signal sequence corresponding to the captured image. (Method 3) The image processing method described in Method 1 or 2, characterized in that the information is a third map that distinguishes between a peripheral area of ​​the brightness saturation area including the brightness saturation area and an area other than the peripheral area in the captured image. (Method 4) The image processing method according to method 3, wherein the third step generates the second map using a multiplication operation based on the first map and the third map. (Method 5) The image processing method according to Method 3 or 4, wherein the third map is generated based on a convolution operation between a fourth map representing the brightness saturation region of the captured image and a filter. (Method 6) 6. The image processing method according to method 5, wherein the filter is a MAX filter. (Method 7) 6. The image processing method according to method 5, wherein the filter is an all-ones filter. (Method 8) The image processing method according to method 2, wherein the second step, the third step, and the fourth step are executed by a same processing means capable of performing parallel computation. (Method 9) the model output is an image corresponding to the captured image; The image processing method according to method 2 or 8, further comprising a fifth step of generating an image by combining the captured image and the model output based on the second map. (Method 10) The image processing method described in any one of Methods 2, 8, and 9, characterized in that the model output includes an image in which the blur of the captured image has been sharpened, an image in which the blur of the captured image has been converted into a blur of a different shape, or a depth map of an object space corresponding to the captured image. (Method 11) 11. The image processing method according to any one of methods 1 to 10, wherein the captured image has a plurality of color components, and the third step is performed for each of the color components. (Method 12) The image processing method according to any one of methods 1 to 11, characterized in that the third step generates the second map based on whether or not a closed space satisfying a predetermined condition of the first map includes the position of the brightness saturation region. (Configuration 1) An acquisition means for acquiring a captured image obtained by imaging; generating a first map based on the captured image using a first machine learning model; a generating means for generating a second map by modifying the first map based on information about a position of a brightness saturation region of the captured image, The image processing device according to claim 1, wherein the first map is a map that represents a relationship between the range of an area in which a subject in the brightness saturation area has expanded due to blurring caused by the imaging, and the magnitude of the signal value corresponding to the area. (Configuration 2) The generating means generates a model output based on the captured image and the second map using a second machine learning model; 2. The image processing device according to configuration 1, wherein the model output is a recognition label corresponding to the captured image or a spatially arranged signal sequence. (Configuration 3) A program for causing a computer to execute the image processing method according to any one of Methods 1 to 12. (Configuration 4) An image processing system having a first device and a second device capable of communicating with each other, the first device has a transmission means for transmitting a captured image obtained by imaging and a request for execution of processing to the second device; The second device comprises: a receiving means for receiving the captured image and the request from the first device; a generating means for generating a first map based on the captured image by using a first machine learning model in response to the request, and for generating a second map by modifying the first map based on information regarding a position of a brightness saturation region of the captured image, The image processing system according to claim 1, wherein the first map is a map that represents a relationship between the range of an area in which a subject in the brightness saturation area has expanded due to blurring caused by the imaging, and the magnitude of the signal value corresponding to the area. (Configuration 5) The generating means generates a model output based on the captured image and the second map using a second machine learning model; 5. The image processing system according to configuration 4, wherein the model output is a recognition label or a spatially arranged signal sequence corresponding to the captured image. [Explanation of symbols]

[0099] 103 Image processing device 103b Acquisition unit (acquisition means) 103c Arithmetic unit (generation means)

Claims

1. a first step of acquiring a captured image obtained by imaging; a second step of generating a first map based on the captured image using a first machine learning model; a third step of generating a second map by correcting the first map based on information about a position of a brightness saturation region in the captured image, 10. An image processing method, wherein the first map is a map that represents a blurred area around a brightness saturated area of the captured image.

2. a fourth step of generating a model output based on the captured image and the second map using a second machine learning model; 2. The image processing method according to claim 1, wherein the model output is a recognition label or a spatially arranged signal sequence corresponding to the captured image.

3. 2. The image processing method according to claim 1, wherein the information is a third map that distinguishes between a peripheral area of the brightness saturated area including the brightness saturated area and an area other than the peripheral area in the captured image.

4. 4. The image processing method according to claim 3, wherein the third step generates the second map by using a multiplication operation based on the first map and the third map.

5. 4. The image processing method according to claim 3, wherein the third map is generated based on a convolution operation between a fourth map representing the brightness saturation region of the captured image and a filter.

6. 6. The image processing method according to claim 5, wherein the filter is a MAX filter.

7. 6. The image processing method according to claim 5, wherein the filter is a filter in which all elements are 1.

8. 3. The image processing method according to claim 2, wherein the second step, the third step, and the fourth step are executed by the same processing means capable of parallel computation.

9. the model output is an image corresponding to the captured image; 3. The image processing method according to claim 2, further comprising a fifth step of generating an image by combining the captured image and the model output based on the second map.

10. 3. The image processing method according to claim 2, wherein the model output includes an image in which the blur of the captured image is sharpened, an image in which the blur of the captured image is converted into a blur of a different shape, or a depth map of a subject space corresponding to the captured image.

11. 2. The image processing method according to claim 1, wherein the captured image has a plurality of color components, and the third step is executed for each of the color components.

12. 2. The image processing method according to claim 1, wherein the third step generates the second map based on whether or not a closed space that satisfies a predetermined condition of the first map includes the position of the brightness saturated region.

13. An image processing method described in any one of claims 1 to 12, characterized in that the first map represents a blurred area in which the subject in the brightness saturation area has expanded due to blurring caused by the imaging, and the magnitude of the signal value in the blurred area.

14. an acquisition means for acquiring a captured image obtained by imaging; generating a first map based on the captured image using a first machine learning model; a generating means for generating a second map by correcting the first map based on information about the position of a brightness saturation region of the captured image, The image processing device according to claim 1, wherein the first map is a map that represents a blurred area around a brightness saturated area of the captured image.

15. the generating means generates a model output based on the captured image and the second map using a second machine learning model; 15. The image processing apparatus according to claim 14, wherein the model output is a recognition label corresponding to the captured image or a spatially arranged signal sequence.

16. A program causing a computer to execute the image processing method according to any one of claims 1 to 13.

17. An image processing system comprising the image processing device according to claim 14 or 15 and a control device capable of communicating with the image processing device, the control device has a transmission means for transmitting a request for execution of processing on the captured image to the image processing device, The image processing system is characterized in that the image processing device has a means for executing processing on the captured image in response to the request.

18. the generating means generates a model output based on the captured image and the second map using a second machine learning model; 18. The image processing system according to claim 17, wherein the model output is a recognition label or a spatially arranged signal sequence corresponding to the captured image.