Image processing method, image processing apparatus, storage medium, image processing system, and learned model manufacturing method
By using input images and maps beyond the dynamic range as input data in neural networks, the problem of reduced estimation accuracy caused by luminance saturated or obstructed areas in the image is solved, and higher image processing accuracy is achieved.
Patent Information
- Application Number
- CN202510176171.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-29
- Filing Date
- 2020-03-24
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to maintain the estimation accuracy of the neural network when the luminance is saturated or obscured in the image, resulting in inaccurate eigenvalue extraction.
By using the input image and its corresponding map beyond the dynamic range as input data for the neural network, the region where the luminance is saturated or obscured is specified, thereby suppressing the reduction in estimation accuracy.
The reduction in estimation accuracy caused by luminance saturation or obscured shadowed areas is effectively suppressed, and the accuracy of image processing is improved.
Smart Images

Figure CN120107674A_ABST
Abstract
Description
[0001] This application is a divisional application based on a patent application with application number 202010210572.8, application date March 24, 2020, and invention name “Image processing method, image processing device, storage medium, image processing system and learned model manufacturing method”. Technical Field
[0002] The present invention relates to an image processing method capable of suppressing a decrease in estimation accuracy of a neural network. Background Art
[0003] Japanese Patent Publication No. (“JP”) 2016-110232 discloses a method of determining the position of a recognition target in an image with high accuracy using a neural network.
[0004] However, the method disclosed in JP 2016-110232 reduces the determination accuracy when the image has a luminance-saturated area or an obscured shadow area. Depending on the dynamic range of the image sensor and the exposure during imaging, the luminance-saturated area or the obscured shadow area may appear in the image. In the luminance-saturated area or the obscured shadow area, there is a possibility that information related to the structure in the object space cannot be obtained, and pseudo edges that do not exist originally appear at the boundaries between these areas. This leads to the extraction of feature values different from the original values of the object, reducing the estimation accuracy. Summary of the invention
[0005] The present invention provides an image processing method, an image processing device, a storage medium, an image processing system, and a learned model manufacturing method, each of which can suppress the reduction in estimation accuracy of a neural network even when brightness saturation or obscured shadows occur.
[0006] As an aspect of the present invention, an image processing method includes the following steps: obtaining a first mapping map representing an area beyond a dynamic range of the input image based on a signal value in an input image and a threshold value of the signal value, and inputting input data including the input image and the first mapping map and performing a recognition task or a regression task.
[0007] An image processing apparatus configured to execute the above-mentioned image processing method and a storage medium storing a computer program for enabling a computer to execute the above-mentioned image processing method also constitute another aspect of the present invention.
[0008] An image processing system as an aspect of the present invention includes a first device and a second device that can communicate with the first device. The first device includes a transmitter configured to transmit a request for the second device to perform processing on a captured image. The second device includes: a receiver configured to receive the request sent by the transmitter; an obtainer configured to obtain a first map representing an area of the captured image that exceeds the dynamic range based on a signal value in the captured image and a threshold value of the signal value; a processor configured to input data including the captured image and the first map into a neural network and perform a recognition task or a regression task; and a transmitter configured to send the result of the task.
[0009] As an aspect of the present invention, an image processing method includes the following steps: obtaining a training image, a first mapping map representing an area of the training image that exceeds a dynamic range based on a signal value in the training image and a threshold value of the signal value, and ground truth data, and using input data including the training image, the first mapping map, and the ground truth data to enable a neural network to learn to perform a recognition task or a regression task.
[0010] A storage medium storing a computer program for enabling a computer to execute the above-mentioned image processing method also constitutes another aspect of the present invention.
[0011] As one aspect of the present invention, a method for manufacturing a learned model includes the following steps: obtaining a training image, a first mapping map representing an area of the training image that is out of a dynamic range based on a signal value in the training image and a threshold value of the signal value, and ground truth data, and using input data including the training image, the first mapping map, and the ground truth data to enable a neural network to learn to perform a recognition task or a regression task.
[0012] An image processing apparatus configured to execute the above-mentioned image processing method also constitutes another aspect of the present invention.
[0013] Further features of the present invention will become apparent from the following description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a diagram illustrating the configuration of a neural network according to the first embodiment.
[0015] Figure 2 is a block diagram of an image processing system according to the first embodiment.
[0016] Figure 3 is an external view of the image processing system according to the first embodiment.
[0017] Figure 4 is a flowchart related to weight learning according to the first embodiment.
[0018] Figure 5A and Figure 5B is a diagram illustrating an example of a training image and a ground truth class map according to the first embodiment.
[0019] Fig. 6A and Figure 6B : is a diagram illustrating an example of a luminance-saturated area and a map exceeding the dynamic range of a training image according to the first embodiment.
[0020] Figure 7 is a flowchart related to the generation of the estimated class map according to the first embodiment.
[0021] Figure 8 is a block diagram of an image processing system according to a second embodiment.
[0022] Fig. 9 is an external view of the image processing system according to the second embodiment.
[0023] Fig.10 is a flowchart related to weight learning according to the second embodiment.
[0024] Fig.11A and Fig. 11B : is a diagram illustrating an example of a map of a luminance-saturated area and an obscured shadow area and an out-of-dynamic range in a training image according to the second embodiment.
[0025] Fig. 12A and Fig. 12B is a diagram illustrating four-channel transformation on a training image according to the second embodiment.
[0026] Fig.13 is a diagram illustrating the configuration of a neural network according to the second embodiment.
[0027] Fig.14 is a flowchart related to the generation of a weighted average image according to the second embodiment.
[0028] Fig.15 is a block diagram of an image processing system according to a third embodiment.
[0029] Fig.16 is a flowchart related to generation of an output image according to the third embodiment. DETAILED DESCRIPTION
[0030] Now, with reference to the accompanying drawings, a detailed description of embodiments according to the present invention will be given. Corresponding elements in the respective drawings will be denoted by the same reference numerals, and their repeated description will be omitted.
[0031] First, before describing the embodiments in detail, the gist of the present invention will be given. The present invention suppresses the reduction of estimation accuracy caused by luminance saturation or occlusion shadows in an image during a recognition task or regression task using a neural network. Here, the input data input to the neural network is x (d-dimensional vector, d is a natural number). Recognition is a task for finding a category y corresponding to the vector x. For example, there are tasks for identifying the characteristics and importance of an object, such as tasks for classifying objects in an image as people, dogs, or cars, and tasks for identifying expressions such as smiling faces and crying faces from facial images. Category y is generally a discrete variable and can be a vector in the generation of a segmentation map, etc. On the other hand, regression is a task for finding a continuous variable y corresponding to the vector x. For example, there is a task for estimating a noise-free image from a noisy image, and a task for estimating a high-resolution image before downsampling from a downsampled image.
[0032] As described above, an area with luminance saturation or obscured shadow (hereinafter referred to as a luminance saturation area or an obscured shadow area) has lost information about the structure in the object space, and a pseudo edge may appear at the boundary between each area. Therefore, it is difficult to correctly extract the feature values of the object. Therefore, the estimation accuracy of the neural network is reduced. In order to suppress this reduction, the present invention uses an input image and an out-of-dynamic range map corresponding to the input image as input data of the neural network. The out-of-dynamic range map (first map) is a map representing a luminance saturation area or an obscured shadow area in the input image. Using the out-of-dynamic range input map, the neural network can specify problematic areas as described above so as to suppress the reduction in estimation accuracy.
[0033] In the following description, the step for learning the weights of the neural network will be referred to as a learning phase, and the step for performing recognition or regression using the learned weights will be referred to as an estimation phase.
[0034] First embodiment
[0035] A description will now be given of an image processing system according to a first embodiment of the present invention. In the first embodiment, a neural network performs a recognition task for detecting a region of a person in an image (whether it is a segmentation of a person). However, the present invention is not limited to this embodiment and is similarly applicable to other recognition tasks and regression tasks.
[0036] Figure 2 is a block diagram illustrating the image processing system 100 in this embodiment. Figure 3 is an external view of the image processing system 100 . Figure 3The front and back of the imaging device (image processing device) 102 are illustrated. The image processing system 100 includes a learning device (image processing device) 101, an imaging device 102, and a network 103. The learning device 101 includes a memory 111, an acquirer (acquisition unit) 112, a detector (learning unit) 113, and an updater (learning unit) 114, and is configured to learn the weights of a neural network for detecting a region of a person. The details of this learning will be described later. The memory 111 stores the weight information learned by the learning device 101. The imaging device 102 performs acquisition of a captured image and detection of a region of a person using a neural network.
[0037] The imaging device 102 includes an optical system 121 and an image sensor 122. The optical system 121 collects light entering the imaging device 102 from the object space. The image sensor 122 receives (photoelectrically converts) an optical image (object image) formed via the optical system 121 and obtains a captured image. The image sensor 122 is, for example, a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal Oxide Semiconductor) sensor.
[0038] The image processor 123 includes an acquirer (acquisition unit) 123a and a detector (processing unit) 123b, and is configured to use at least a portion of the captured image as an input image to detect the area of the person using the weight information stored in the memory 124. The weight information is read in advance from the learning device 101 via the wired or wireless network 103 and stored in the memory 124. The stored weight information may be the weight value itself or the encoding format. A detailed description related to the detection process for the area of the person will be given later. The image processor 123 performs processing based on the detected area of the person and generates an output image. For example, the image processor 123 adjusts the brightness in the captured image so that the area of the person has a suitable brightness. The recording medium 125 stores the output image. Alternatively, the captured image may be stored as it is on the recording medium 125, after which the image processor 123 may read the captured image from the recording medium 125 and detect the area of the person. The display 126 displays the output image stored in the recording medium 125 according to the user's instruction. The system controller 127 controls this series of operations.
[0039] Reference now Figure 4 , a description will be given of weight learning (making a learned model) performed by the learning device 101 in this embodiment. Figure 4 is a flowchart related to weight learning. Mainly, the acquirer 112, the detector 113 or the updater 114 in the learning device 101 executes Figure 4 Each step in .
[0040] First, in step S101, the obtainer 112 obtains one or more sets of training images and ground truth class maps (also called ground truth segmentation maps or ground truth data) and maps beyond the dynamic range. The training image is an input image during the learning phase of the neural network. The ground truth class map is a ground truth segmentation map corresponding to the training image.
[0041] Figure 5A and Figure 5B Examples of training images and ground truth category maps are shown. Fig. 6A and Figure 6B is a diagram illustrating an example of a map of a luminance-saturated area and an area outside the dynamic range for a training image. Figure 5A shows examples of training images, and Figure 5B The corresponding ground truth category map is shown. Figure 5B The white areas in the figure are the categories representing the human areas, while the black areas are the categories representing other areas. Figure 5A The training images in have areas of saturated brightness.
[0042] Fig. 6A An image is illustrated in which the wavy lines represent luminance saturation areas. In this embodiment, the out-of-dynamic range map (first map) is a map that indicates whether luminance saturation has occurred for each pixel in the training image. However, the present invention is not limited to this embodiment, and the map may be a map that represents occluded shadows. The signal value at each pixel in the training image is compared with the luminance saturation value as a threshold. When the signal value is equal to or greater than the luminance saturation value, an out-of-dynamic range map is generated, wherein the map indicates that the signal value exceeds the dynamic range. Alternatively, a map outside the dynamic range can be pre-generated according to the above-mentioned method for training images, and can be obtained by reading it out.
[0043] In this embodiment, the out-of-dynamic-range map is a binary map of 1 or 0 (information indicating whether luminance saturation occurs), such as Figure 6BAs shown in . The numerical significant bits can be reversed. The binary map has the advantage of reducing data capacity. However, the present invention is not limited to this embodiment. The mapping map beyond the dynamic range can be a mapping map with intermediate values to indicate how close the signal value is to the luminance saturation value. The learning stage uses multiple training images of various imaging scenes so that the estimation stage stably detects the area of the person even in the image of the unknown imaging scene. Multiple training images can be obtained by changing the brightness in the same imaging scene. The training image has the same format as the input image in the estimation stage. If the input image in the estimation stage is an undeveloped RAW image, then the training image is also an undeveloped RAW image. If the input image in the estimation stage is a developed image, then this also applies to the training image. When the training image is a RAW image, a mapping map beyond the dynamic range can be generated after applying white balance. The input image and the training image in the estimation stage do not have to have the same number of pixels.
[0044] Next, in Figure 4 In step S102, the detector 113 inputs the training image and the map beyond the dynamic range into the neural network and generates an estimated class map. In this embodiment, the neural network uses Figure 1 The U-Net shown in (for a detailed description, see O. Ronneberger, P. Fischer, and T. Brox, "U-net: Convolutional networks for biomedical image segmentation", MICACI, 2015), but the present invention is not limited to this embodiment. Input data 201 is data obtained by cascading a training image with a map out of the dynamic range in the channel direction. The order of the cascade is not limited, and other data can be inserted between them. The training image can have multiple channels of RGB (red, green, blue). The map out of the dynamic range can have only one channel or the same number of channels as the training image. When the map out of the dynamic range has one channel, the map is, for example, a map that expresses the presence or absence of luminance saturation for luminance components other than color differences. The number of pixels (number of elements) of each channel is the same between the training image and the map out of the dynamic range. Even if training images including various scenes with luminance saturation are input to the neural network, by including a map exceeding the dynamic range in the input data, the neural network can recognize luminance saturation areas in the training images and can suppress a decrease in estimation accuracy.
[0045] If necessary, the input data may be normalized. When the training image is a RAW image, the black level may be different depending on the image sensor or ISO sensitivity. Therefore, after the black level is subtracted from the signal value in the training image, the training image is input to the neural network. Normalization may be performed after subtracting the black level. Figure 1 The convolution in represents one or more convolution layers, the max pooling represents the maximum value pooling, the up-convolution represents one or more convolution layers including upsampling, and the cascade represents the cascade in the channel direction. At the first learning, a random number is used to determine the weight of the filter in each convolution layer. An estimated category map 202 is calculated as the output of the U-Net corresponding to the training image.
[0046] Only one of the input image and the out-of-dynamic-range map may be input to the first layer in the neural network, and at least the feature map output from the first layer may be cascaded with the input image and the other out-of-dynamic-range map that has not been input to the first layer in the channel direction, and may be input to a subsequent layer. Alternatively, the input portion of the neural network may be branched, the input image and the out-of-dynamic-range map may be converted into feature maps in different layers, and the feature maps may be cascaded with each other and may be input to a subsequent layer.
[0047] Later, in Figure 4 In step S103 of FIG. 1 , the updater 114 updates the weights of the neural network based on the estimated class map and the ground truth class map. The first embodiment uses the cross entropy of the estimated class map and the ground truth class map as a loss function, but the present invention is not limited to this implementation. The weights are updated according to the values calculated from the loss function by back propagation or the like.
[0048] Later in Figure 4 In step S104, the updater 114 determines whether the weight learning has been completed. The completion can be determined based on whether the number of iterations of learning (weight update) has reached a predetermined value or whether the weight change during the update is less than a predetermined value. If the weight learning is determined to be incomplete, the process returns to step S101 to newly obtain one or more sets of training images, maps that exceed the dynamic range, and ground truth category maps. On the other hand, if the weight learning is determined to be completed, the learning is terminated and the memory 111 stores the weight information.
[0049] Reference now Figure 7 , a description will be given of detection of a region of a person in an input image (generation of an estimated class map, estimation phase) performed by the image processor 123 in this embodiment. Figure 7 1 is a flowchart related to the generation of an estimated class map. Mainly, the acquirer 123a or the detector 123b in the image processor 123 performs Figure 7 Each step in .
[0050] First, in step S201, the obtainer 123a obtains an input image and a threshold value (in this embodiment, a luminance saturation value) corresponding to the input image. The input image is at least a portion of a captured image captured by the image sensor 122. The memory 124 has stored the luminance saturation value of the image sensor 122, and the value is read and obtained. Then in step S202, the obtainer 123a generates an out-of-dynamic range map based on a comparison between the signal value at each pixel in the input image and the threshold value. Then in step S203, the detector 123b inputs the input image and the out-of-dynamic range map as input data to the neural network, and generates an estimated category map. At this time, the image is detected using the image sensor 122. Figure 1 The neural network and the weights obtained during the learning phase.
[0051] This embodiment can provide an image processing system that can generate a highly accurate segmentation map even when luminance saturation occurs.
[0052] Second embodiment
[0053] A description will now be given of an image processing system in a second embodiment of the present invention. In this embodiment, a neural network is configured to perform a regression task for deblurring a captured image having blur caused by aberration and diffraction. However, the present invention is not limited to this embodiment and may be applied to another recognition task or regression task.
[0054] Figure 8 is a block diagram of the image processing system 300 in this embodiment. Fig. 9 300 is an external view of the image processing system 300. The image processing system 300 includes a learning device (image processing device) 301, an imaging device 302, an image estimation device (image processing device) 303, a display device 304, a recording medium 305, an output device 306, and a network 307.
[0055] The learning device 301 includes a memory 301a, an acquirer (acquisition unit) 301b, a generator (learning unit) 301c and an updater (learning unit) 301d. The imaging device 302 includes an optical system 302a and an image sensor 302b. The captured image captured by the image sensor 302b includes blur caused by aberration and diffraction of the optical system 302a, and shadows blocked by luminance saturation due to the dynamic range of the image sensor 302b. The image estimation device 303 includes a memory 303a, an acquirer 303b and a generator 303c, and is configured to generate an estimated image obtained by deblurring an input image that is at least a part of the captured image, and generate a weighted average image based on the input image and the estimated image. The input image and the estimated image are RAW images. A neural network is used for deblurring, and its weight information is read from the memory 303a. The learning device 301 has learned the weights, and the image estimation device 303 has read out the weight information from the memory 301a in advance via the network 307, and the memory 303a has stored the weight information. A detailed description of the weight learning and the deblurring process using the weights will be given later. The image estimation device 303 performs a development process on the weighted average image and generates an output image. The output image is output to at least one of the display device 304, the recording medium 305, and the output device 306. The display device 304 is, for example, a liquid crystal display or a projector. Via the display device 304, the user can perform editing work, etc. while checking the image being processed. The recording medium 305 is, for example, a semiconductor memory, a hard disk drive, or a server on the network. The output device 306 is a printer, etc.
[0056] Reference now Fig.10 , a description will be given of the weight learning (learning phase) performed by the learning device 301. Fig.10 is a flowchart related to weight learning. Mainly, the acquirer 301b, the generator 301c or the updater 301d in the learning device 301 executes Fig.10 Each step in .
[0057] First, in step S301, the acquirer 301b obtains one or more sets of source images and imaging conditions. A pair of blurred images (hereinafter referred to as the first training image) and a non-blurred image (hereinafter referred to as the ground truth image) are required for deblurring learning of aberration and diffraction. This embodiment generates the pair of images from the source image by imaging simulation. However, the present invention is not limited to this embodiment, and the pair of images can be prepared by imaging the same object using a lens that may cause blurring due to aberration and diffraction and a lens with higher performance.
[0058] This embodiment uses RAW images for learning and deblurring. However, the present invention is not limited to this embodiment, and the image can be used after development. The source image is a RAW image, and the imaging condition is a parameter for using the source image as an imaging simulation of the object. The parameters include the optical system used for imaging, the state of the optical system (zoom, aperture stop and focus distance), image height, the presence or absence of an optical low-pass filter and the type, the noise characteristics of the image sensor, pixel spacing, ISO sensitivity, color filter array, dynamic range, black level, etc. This embodiment learns the weights to be used in the deblurring for each optical system. This embodiment sets multiple combinations of state, image height, pixel spacing, ISO sensitivity, etc. for a specific optical system, and generates a pair of first training images and ground truth images (ground truth data) under different imaging conditions. The source image can be an image with a wider dynamic range than the training image. When the dynamic range between the source image and the training image is the same, the blurring process deletes a smaller luminance saturation area or a smaller shaded shadow area in the source image, making it difficult to perform learning. A source image having a wide dynamic range may be prepared by capturing an image using an image sensor having a wide dynamic range or by capturing and combining images of the same object under different exposure conditions.
[0059] Then in step S302, the generator 301c generates a first training image, a second training image, and a ground truth image from the source image based on the imaging conditions. The first training image and the ground truth image are respectively an image obtained by adding blur caused by aberration and diffraction of the optical system to the source image, and an image of the source image without adding blur. If necessary, noise can be added to the first training image and the ground truth image. When no noise is added to the first training image, the neural network amplifies the noise and performs deblurring in the estimation stage. When noise is added to the first training image, and no noise is added to the ground truth image, or noise that has no correlation with the noise in the first training image is added to the ground truth image, the neural network learns deblurring and denoising. On the other hand, when noise that has correlation with the noise in the first training image is added to the ground truth image, the neural network learns deblurring, wherein noise variation is suppressed.
[0060] This embodiment adds correlated noise to the first training image and the ground truth image. If the dynamic range in the source image is greater than the dynamic range of the first training image, the signal value is clipped so that the dynamic range in the first training image and the ground truth image is brought into the original ground truth range. This embodiment uses a Wiener filter for the first training image and generates a second training image (hereinafter referred to as an intermediate deblurred image in the learning phase) in which blur has been corrected to some extent. The Wiener filter is a filter calculated from the blur assigned to the first training image. However, the correction method is not limited to the Wiener filter, and another inverse filter-based method or Richardson-Lucy method can be used. By using the second training image, it is possible to improve the robustness of deblurring for blur changes in the neural network. If necessary, the source image can be reduced during the imaging simulation. When the source image is prepared not by CG (computer graphics) but by actual imaging, the source image is an image captured by a certain optical system. Therefore, the source image already includes blur caused by aberrations and diffraction. However, this reduction can reduce the impact of blur and generate a ground truth image including high frequencies.
[0061] Then in step S303, the generator 301c generates a map out of the dynamic range based on a comparison between the signal value in the first training image (the input image of the learning phase) and the threshold value of the signal value. However, the map out of the dynamic range can be generated from the signal value in the second training image.
[0062] In this embodiment, the signal threshold is based on the luminance saturation value and the black level of the image sensor 302b. Fig.11A and Fig. 11B Shown in. Fig.11A is the first training image, in which the wavy lines represent areas having signal values equal to or greater than the luminance saturation value (hereinafter referred to as the first threshold). The vertical lines represent areas having signal values equal to or less than a value obtained by adding a constant to the black level (hereinafter referred to as the second threshold). At this time, the out-of-dynamic-range mapping corresponding to the first training image is as follows: Fig. 11B As shown in . The area with a signal value equal to or greater than the first threshold is set to 1, the area with a signal value equal to or less than the second threshold is set to 0, and the other areas are set to 0.5. However, the present invention is not limited to this embodiment. For example, the area with a signal value greater than the second threshold and less than the first threshold can be set to 0, and the other areas with a structure of luminance saturation or blocked shadows can be set to 1.
[0063] Next is the reason for adding a constant to the black level in the second threshold. Since noise is added to the first training image, even if the true signal value is the black level, the signal value may exceed the black level due to noise. Therefore, taking into account the increase in the signal value due to noise, a constant is added to the second threshold. The constant may be a value reflecting the amount of noise. For example, the constant may be set to n times (n is a positive real number) the standard deviation of the noise. Maps that exceed the dynamic range are input to the neural network in both the learning phase and the estimation phase. Because the learning phase adds noise during simulation, the standard deviation of the noise in the input image is known, but in the estimation phase, the standard deviation of the noise in the input image is unknown. Therefore, the estimation phase may measure the noise characteristics of the image sensor 302b in advance, and may determine the constant to be added to the second threshold according to the ISO sensitivity during imaging. If the noise is small enough, the constant may be zero.
[0064] Then in step S304, the generator 301c inputs the first training image and the second training image and the map outside the dynamic range into the neural network and generates an estimated image (i.e., a deblurred image). This embodiment converts the first training image and the second training image and the map outside the dynamic range into four-channel formats respectively and inputs them into the neural network. Fig. 12A and Fig. 12B This conversion will be described. Fig. 12A The color filter array in the first training image is shown. G1 and G2 represent the two green components. The first training image is converted to a four-channel format during input to the neural network, as shown in Fig. 12B The dashed lines represent each channel component at the same position. However, the color order in the array is not limited to Fig. 12A and Fig. 12B Similarly, the second training image and the out-of-dynamic-range map are converted to a four-channel format. It is not always necessary to perform the conversion to a four-channel format. If necessary, the first training image and the second training image may be normalized and the black level may be subtracted.
[0065] This example uses Fig.13The neural network shown in , but the present invention is not limited to this embodiment, and for example, a GAN (generative adversarial network) may be used. The input data 511 is data obtained by cascading the first training image 501 converted into a four-channel format, the second training image, and the mapping map that exceeds the dynamic range in the channel direction. There is no restriction on the order of cascading in the channel direction. Convolution represents one or more convolution layers, and deconvolution represents one or more deconvolution layers. The second skip connection to the fourth skip connection 522 to 524 takes the sum of each element in the two feature maps, or these elements may be cascaded in the channel direction. The first skip connection 521 obtains an estimated image 512 by taking the sum of the first training image 501 (or the second training image) and the residual image output from the final layer. However, the number of skip connections is not limited to Fig.13 The number in. Fig. 12B As shown in , the estimated image 512 is also a four-channel image.
[0066] In image resolution enhancement and contrast enhancement such as deblurring, problems occur near areas where object information is lost due to luminance saturation or blocked shadows. In addition, deblurring can reduce areas with object information loss. Unlike other areas, the neural network needs to perform repair processing in areas with object information loss. Using an input map that exceeds the dynamic range as input, the neural network can specify these areas and can perform highly accurate deblurring.
[0067] Subsequently, in step S305, the updater 301d updates the weights of the neural network based on the estimated image and the ground truth image. This embodiment defines the Euclidean norm of the difference in signal values between the estimated image and the ground truth image as a loss function. However, the loss function is not limited to this. Before taking the difference, the ground truth image is also converted into a four-channel format based on the estimated image. The second embodiment removes areas with luminance saturation or occluded shadows from the loss. Since this area loses information about the object space, the restoration task is required as described above in order to make the estimated image similar to the ground truth image. Since restoration may cause erroneous construction, the second embodiment excludes this area from the estimation and also replaces this area with the input image in the estimation stage. As Fig. 12BAs shown in , the first training image includes multiple color components. Thus, even if a certain color component has luminance saturation or an obscured shadow, the structure of the object can be obtained through other color components. In this case, since information about an area with luminance saturation or an obscured shadow can be estimated based on pixels present at very close positions, erroneous structure rarely occurs. Therefore, in a map that exceeds the dynamic range, a loss weight map is generated in which all pixels of a channel that exceeds the dynamic range are set to 0, while other pixels are set to 1, and the loss is calculated by taking the product of each component related to the difference between the estimated image and the ground truth image. Thus, it is possible to exclude only areas that may have erroneous structures. It is not always necessary to exclude areas that may have erroneous structures from the loss.
[0068] Then in step S306, the updater 301d determines whether the learning has been completed. If the learning has not been completed, the process returns to step S301 to newly obtain one or more sets of source images and imaging conditions. On the other hand, if the learning is completed, the memory 301a stores the weight information.
[0069] Reference now Fig.14 , a description will be given of the deblurring of aberrations and diffraction in the input image performed by the image estimation device 303 (generation of a weighted average image, estimation stage). Fig.14 is a flowchart related to the generation of a weighted average image. Mainly, the acquirer 303b and the generator 303c in the image estimation device 303 perform Fig.14 Each step in .
[0070] First, in step S401, the obtainer 303b obtains an input image and a threshold value corresponding to the input image from a captured image. The first threshold value is a luminance saturation value of the image sensor 302b, and the second threshold value is a value obtained by adding a constant to the black level of the image sensor 302b. The constant is determined according to the ISO sensitivity when the captured image is captured using the noise characteristics of the image sensor 302b.
[0071] Then in step S402, the generator 303c generates an out-of-dynamic range map based on the comparison of the signal value of the input image with the first threshold and the second threshold. The out-of-dynamic range map is generated by a method similar to that in step S303 in the learning phase.
[0072] Then in step S403, the generator 303c generates an intermediate deblurred image from the input image. The generator 303c generates the intermediate deblurred image by reading out information about the Wiener filter that corrects blur caused by aberration and diffraction of the optical system 302a from the memory 303a and applying the information to the input image. Since the input image has different blur for each image height, shift variable correction is performed. Either step S402 or step S403 may be performed first.
[0073] Then in step S404, the generator 303c inputs the input image, the intermediate deblurred image and the out-of-dynamic-range map into the neural network and generates an estimated image. Fig.13 The configuration shown in , and input data obtained by cascading an input image (corresponding to the first training image), an intermediate deblurred image (corresponding to the second training image), and a map out of the dynamic range in the same order as in the learning in the channel direction. An estimated image is generated by reading out weight information corresponding to the optical system 302a from the memory 303a. If normalization or black level subtraction has been performed in step S404 when input to the neural network, scaling for restoring the signal value and processing for adding the black level are performed on the estimated image.
[0074] Subsequently, in step S405, the generator 303c calculates a weight map based on a comparison between the signal value in the input image and the first threshold and the second threshold. That is, the generator 303c obtains a weight map based on the signal value in the input image and the threshold of the signal value. Similar to the calculation of the loss weight map in the learning phase, this embodiment uses a map that exceeds the dynamic range to calculate the weight map. For example, when luminance saturation or an obscured shadow occurs in a target pixel having a certain color component, if luminance saturation or an obscured shadow occurs in all the closest pixels having other colors, the weight is set to 0; in other cases, the weight is set to 1.
[0075] As described above, in this embodiment, the input image includes a plurality of color components. If luminance saturation or blocked shadows appear in the target pixel of the input image and in all pixels having color components different from the color components in the target pixel within a predetermined area (e.g., the closest area), a weight map is generated so that the weight at the target pixel position in the input image is greater than the output from the neural network. On the other hand, if luminance saturation or blocked shadows do not appear in the target pixel of the input image and / or in any of the pixels having color components different from the color components in the target pixel within the predetermined area, a weight map is generated so that the weight at the target pixel position in the input image is less than the output from the neural network.
[0076] Blurring may be performed to reduce discontinuities in the calculated weight map, or the weight map may be generated by another method.The weight map may be generated at any time between step S401 and step S406.
[0077] Subsequently, in step S406, the generator 303c weights and averages the input image and the estimated image based on the weight map, and generates a weighted average image. That is, the generator 303c generates a weighted average image based on the output from the neural network (estimated image or residual image), the input image, and the weight map. The weighted average image is generated by taking the product of the weight map and each element in the estimated image and the sum of the products of the map obtained by subtracting the weight map from the map of all elements in the input image and each element. Instead of step S406, by using the weight map, when the skip connection 521 takes the sum of the input image and the residual image in step S404, an estimated image can be generated, wherein the input image replaces the area in the estimated image that may have an erroneous structure. In this case, the pixels that may have an erroneous structure indicated by the weight map are set to the input image, and the other pixels are set to the sum of the input image and the residual image. By performing the same process, in step S305, the learning stage can also exclude areas that may have erroneous structures from the loss function.
[0078] This embodiment can provide an image processing system that can perform deblurring with high accuracy even when luminance saturation or blocked shadows occur.
[0079] Therefore, in the first and second embodiments, the obtaining unit (obtainer 123a; obtainer 303b and generator 303c) obtains the out-of-dynamic-range map of the input image based on the signal value in the input image and the threshold value of the signal value. The processing unit (detector 123b; generator 303c) inputs the input data including the input image and the out-of-dynamic-range map to the neural network and performs a recognition task or a regression task.
[0080] Third embodiment
[0081] A description will now be given of an image processing system in a third embodiment of the present invention. The image processing system in this embodiment differs from the first and second embodiments in that the image processing system includes a processing device (computer) configured to send a captured image to be processed (input image) to an image estimation device and to receive a processed output image from the image estimation device.
[0082] Fig.15601 is a block diagram of an image processing system 600 in this embodiment. The image processing system 600 includes a learning device 601, an imaging device 602, an image estimation device 603, and a processing device (computer) 604. The learning device 601 and the image estimation device 603 are, for example, servers. The computer 604 is, for example, a user terminal (personal computer or smart phone). A network 605 connects the computer 604 and the image estimation device 603. A network 606 connects the image estimation device 603 and the learning device 601. That is, the computer 604 and the image estimation device 603 are configured to be communicable, and the image estimation device 603 and the learning device 601 are configured to be communicable. The computer 604 corresponds to the first device, and the image estimation device 603 corresponds to the second device. The configuration of the learning device 601 is the same as that of the learning device 301 in the second embodiment, so its description will be omitted. The configuration of the imaging device 602 is the same as that of the imaging device 302 in the second embodiment, so its description will be omitted.
[0083] The image estimation device 603 includes a memory 603a, an acquirer (acquisition unit) 603b, a generator (processing unit) 603c, and a communicator (receiving unit and sending unit) 603d. The memory 603a, the acquirer 603b, and the generator 603c are respectively the same as the memory 103a, the acquirer 103b, and the generator 103c in the image estimation device 303 in the second embodiment. The communicator 603d has a function of receiving a request sent from the computer 604, and a function of sending an output image generated by the image estimation device 603 to the computer 604.
[0084] The computer 604 includes a communicator (transmitting unit) 604a, a display 604b, an image processor 604c, and a recorder 604d. The communicator 604a has a function of transmitting a request for the image estimation device 603 to perform processing on a captured image to the image estimation device 603, and a function of receiving an output image processed by the image estimation device 603. The display 604b has a function of displaying various information. The information displayed by the display 604b includes, for example, a captured image to be transmitted to the image estimation device 603 and an output image received from the image estimation device 603. The image processor 604c has a function of performing further image processing on the output image received from the image estimation device 603. The recorder 604d records a captured image obtained from the imaging device 602, an output image received from the image estimation device 603, and the like.
[0085] Reference now Fig.16 , a description of the image processing in this embodiment will be given. The image processing in this embodiment is equivalent to the deblurring processing described in the second embodiment ( Fig.14 ).
[0086] Fig.16 604 is a flowchart related to the generation of output images. When the user issues an instruction to start image processing via computer 604, Fig.16 The image processing shown in is started. First, the operation in the computer 604 will be described.
[0087] In step S701, the computer 604 sends a request for processing a captured image to the image estimation device 603. It does not matter how the captured image to be processed is sent to the image estimation device 603. For example, the captured image may be uploaded from the computer 604 to the image estimation device 603 at the same time as step S701, or may be uploaded to the image estimation device 603 before step S701. Instead of an image recorded on the computer 604, the captured image may be an image stored on a server different from the image estimation device 603. In step S701, the computer 604 may send ID information for authenticating a user, etc., and a request for processing the captured image. In step S702, the computer 604 receives an output image generated in the image estimation device 603. Similar to the second embodiment, the output image is an estimated image obtained by deblurring the captured image.
[0088] A description will now be given of the operation of the image estimation device 603. In step S801, the image estimation device 603 receives a request for processing a captured image sent from the computer 604. The image estimation device 603 determines that processing (deblurring processing) of the captured image has been instructed, and performs processing after step S802. Steps S802 to S807 are the same as steps S401 to S406 in the second embodiment. In step S808, the image estimation device 603 transmits the estimated image (weighted average image) as a result of the regression task to the computer 604 as an output image.
[0089] Although this embodiment has been described as performing the deblurring process similarly to the second embodiment, this embodiment can be similarly applied to the detection of the area of a person in the first embodiment ( Figure 7 ). This embodiment has described that the image estimation device 603 performs all the processing corresponding to steps S401 to S406 in the second embodiment, but the present invention is not limited to this embodiment. For example, the computer 604 may perform one or more of steps S401 to S406 in the second embodiment (corresponding to steps S802 to S807 in this embodiment), and may send the result to the image estimation device 603.
[0090] As described in this embodiment, the image estimating device 603 can be controlled using the computer 604 communicably connected to the image estimating device 603 .
[0091] For example, the regression task in each embodiment is to form a defocus blur in a captured image. Forming a defocus blur is a task for converting double-line blur, vignetting, a ring pattern caused by an aspherical lens mold, a ring-shaped defocus blur of a mirror lens, etc. into a blur with an arbitrary distribution. At this time, a problem occurs in an area where information loss occurs due to luminance saturation or an obscured shadow. However, by inputting a map that exceeds the dynamic range into a neural network, it is possible to perform the formation of a defocus blur while suppressing side effects.
[0092] Other embodiments
[0093] Embodiments of the present invention may also be implemented by a computer of a system or device and a method performed by a computer of the system or device, the computer reading and executing computer executable instructions (e.g., one or more programs) recorded on a storage medium (which may also be more completely referred to as a "non-transitory computer-readable storage medium") to perform the functions of one or more of the embodiments described above, and / or including one or more circuits (e.g., application-specific integrated circuits (ASICs)) for performing the functions of one or more of the embodiments described above, the method by, for example, reading and executing computer executable instructions from the storage medium to perform the functions of one or more of the embodiments described above, and / or controlling the one or more circuits to perform the functions of one or more of the embodiments described above. The computer may include one or more processors (e.g., a central processing unit (CPU), a microprocessing unit (MPU)) and may include a network of independent computers or independent processors to read and execute the computer executable instructions. The computer executable instructions may be provided to the computer from, for example, a network or the storage medium. The storage medium may include, for example, one or more of a hard disk, a random access memory (RAM), a read-only memory (ROM), a storage device of a distributed computing system, an optical disk (such as a compact disk (CD), a digital versatile disk (DVD) or a Blu-ray disk (BDTM)), a flash memory device, a memory card, etc.
[0094] Other embodiments
[0095] The embodiments of the present invention may also be implemented by providing software (program) for performing the functions of the above-described embodiments to a system or device via a network or various storage media, and a computer or a central processing unit (CPU) or a microprocessing unit (MPU) of the system or device reads and executes the program.
[0096] The above-described embodiments can provide an image processing method, an image processing apparatus, a program, an image processing system, and a learned model manufacturing method, each of which can suppress a decrease in estimation accuracy of a neural network even when luminance saturation or an obscured shadow occurs.
[0097] While the present invention has been described with reference to exemplary embodiments, it is to be understood that the invention is not limited to the disclosed exemplary embodiments.The scope of the following claims is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures and functions.
Claims
1. An image processing method, The following steps are involved: Obtaining a first map representing the area based on a signal value in the input image and a threshold value of the signal value; as well as performing a recognition task or a regression task by inputting input data including the input image and the first map, The first mapping map is a mapping map representing at least one of a luminance-saturated area and an obscured shadow area in the input image.
2. The image processing method according to claim 1, Features The threshold is set based on at least one of a luminance saturation value and a black level in the input image.
3. The image processing method according to claim 1, Features The input data is obtained by concatenating the input image and the first map.
4. The image processing method according to claim 1, Features The execution step converts the input image or the first map into a feature map by inputting only one of the input image and the first map into a first layer of a neural network, cascading the feature map with the other one of the input image and the first map that has not been input to the first layer in a channel direction, and inputting the cascaded data into subsequent layers of the neural network.
5. The image processing method according to claim 1, Features The step of converting the input image and the first map into feature maps in different layers by inputting the input image and the first map into different layers, cascading the feature maps in a channel direction, and thereafter inputting the feature maps into subsequent layers.
6. The image processing method according to claim 1, Features The number of pixels in each channel between the input image and the first map is equal to each other.
7. The image processing method according to claim 1, Features The task includes deblurring the input image.
8. The image processing method according to claim 1, further comprising: The following steps are involved: obtaining a weight map based on the signal value and a threshold value of the signal value; as well as A weighted average image is generated based on the output from the neural network, the input image, and the weight map.
9. The image processing method according to claim 8, Features The input image includes a plurality of color components, and Wherein, in the input image, when luminance saturation or occluded shadow appears in a target pixel and all pixels in a predetermined area having a color component different from the color component of the target pixel, the weight map is generated so that the weight at the position of the target pixel in the input image is greater than the weight of the output.
10. The image processing method according to claim 8 or 9, Features The input image has multiple color components, and Wherein, in the input image, when neither luminance saturation nor occluded shadows occur in a target pixel and all pixels in a predetermined area having a color component different from the color component of the target pixel, the weight map is generated so that the weight at the position of the target pixel in the input image is less than the weight of the output.
11. An image processing device, include: an obtaining unit configured to obtain a first map representing a region based on a signal value in an input image and a threshold value of the signal value; as well as a processing unit configured to perform a recognition task or a regression task by inputting data including the input image and the first map into a neural network, The first mapping map is a mapping map representing at least one of a luminance-saturated area and an obscured shadow area in the input image.
12. The image processing apparatus according to claim 11, further comprising a memory configured to store information about the neural network.
13. An image capturing device, include: A capture unit for obtaining an input image by capturing an object; as well as An image processing device according to claim 11 or 12. 14 . A non-transitory computer-readable storage medium storing a computer program for causing a computer to execute the image processing method according to claim 1 .
15. An image processing system, include: The image processing device according to claim 11 or 12; as well as a control device capable of communicating with the image processing device, The control device comprises a transmitter configured to send a request for causing the image processing device to perform processing on the captured image; The image processing device comprises: A processor is configured to perform processing on the captured image in response to the request.
16. A method for generating a learned model, The following steps are involved: obtaining a training image, a first map representing a region, and ground truth data, the first map being based on signal values in the training image and thresholds for the signal values; as well as Using input data including the training images and the first map and the ground truth data, a neural network is caused to learn to perform a recognition task or a regression task, The first mapping map is a mapping map representing at least one of a luminance-saturated area and an obscured shadow area in the input image. 17 . A non-transitory computer-readable storage medium storing a computer program for causing a computer to execute the generating method according to claim 16 .
Citation Information
Patent Citations
Object recognition device, object recognition method, and program
JP2016110232A