Image processing methods, image processing devices, programs

The image processing method employs a machine learning model to address image quality issues from geometric transformations by dynamically adjusting corrections based on deformation information, improving image clarity.

JP7851206B2Active Publication Date: 2026-04-24CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CANON KK
Filing Date
2022-07-20
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing methods for correcting image quality degradation due to geometric transformations apply a fixed amount of deformation, leading to risks of undercorrection or overcorrection, particularly in images captured with fisheye lenses.

Method used

An image processing method using a machine learning model that inputs geometric transformation-applied images and deformation information to accurately correct image quality degradation by adjusting deformation-specific corrections.

Benefits of technology

The method effectively corrects image quality degradation caused by geometric transformations with high accuracy, reducing undercorrection and overcorrection, thereby enhancing image clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007851206000001
    Figure 0007851206000001
  • Figure 0007851206000002
    Figure 0007851206000002
  • Figure 0007851206000003
    Figure 0007851206000003
Patent Text Reader

Abstract

To provide an image processing method capable of accurately correcting deterioration of image quality due to geometric transformation using a machine learning model, an image processing apparatus, an image processing system, a program, a storage medium, a learning apparatus, and a method of manufacturing a trained model.SOLUTION: An image processing method includes the steps of: acquiring information 21 about an optical system, and a second image 23 obtained by applying geometric transformation to a first image 22; acquiring information 24 about a deformation amount of the first image in the geometric transformation; and generating a third image 25 by inputting the second image 23 and the information 24 about the deformation amount to a machine learning model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for correcting a reduction in image quality caused by performing a geometric transformation on an image.

Background Art

[0002] By photographing a subject using a fisheye lens, a clear image over a wide range can be obtained. However, the image obtained using a fisheye lens is distorted more greatly toward the edges. Therefore, an image obtained using a fisheye lens needs to be corrected for distortion by geometric transformation. The greater the amount of deformation (correction amount) of the image in geometric transformation, the greater the reduction in the image quality of the image subjected to the geometric transformation.

[0003] Non-Patent Document 1 discloses a method for correcting a reduction in image quality caused by performing a geometric transformation on an image using a machine learning model.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the method disclosed in Non-Patent Document 1, regardless of the geometric transformation applied to the image, the reduction in image quality is corrected with a fixed amount of deformation for each pixel. Therefore, depending on the geometric transformation applied to the image, there is a risk of undercorrection or overcorrection.

[0006] Therefore, the present invention aims to correct the degradation of image quality caused by applying geometric transformations to images with high accuracy using a machine learning model. [Means for solving the problem]

[0007] One aspect of the present invention is an image processing method comprising the steps of: obtaining a second image obtained by applying a geometric transformation to a first image; and obtaining information regarding the amount of deformation of the first image in the geometric transformation. Furthermore, it comprises the step of inputting the second image and the information regarding the amount of deformation into a machine learning model to generate a third image. [Effects of the Invention]

[0008] According to the present invention, it is possible to correct the degradation of image quality caused by geometric transformations with high accuracy using a machine learning model. [Brief explanation of the drawing]

[0009] [Figure 1] This diagram shows the process for generating the estimated image in Example 1. [Figure 2] This is a block diagram of the image processing system in Example 1. [Figure 3] This is an external view of the image processing system in Example 1. [Figure 4] This diagram shows the flow of weight updates in Example 1. [Figure 5] This is a flowchart illustrating the weight update process in Example 1. [Figure 6] This is a flowchart showing the image processing method in Example 1. [Figure 7] This is an explanatory diagram regarding the amount of deformation in Example 1. [Figure 8] This is a block diagram of the image processing system in Example 2. [Figure 9] This is an external view of the image processing system in Example 2. [Figure 10] This is a flowchart showing the image processing method in Example 2. [Figure 11] This is an explanatory diagram regarding the amount of deformation in Example 2. [Figure 12] This is a block diagram of the image processing system in Example 3. [Figure 13] This is a flowchart of the image processing method in Example 3. [Modes for carrying out the invention]

[0010] Embodiments of the present invention will be described in detail below with reference to the drawings. In each figure, the same reference numerals are used for the same components, and redundant descriptions are omitted.

[0011] First, before giving a detailed description of the embodiment, the gist of this embodiment will be explained. In this embodiment, an image (second image) generated by applying a geometric transformation to the original image (first image) and information regarding the amount of deformation of the first image in the geometric transformation are input to a machine learning model to generate an estimated image (third image).

[0012] The geometric transformation according to this embodiment is performed, for example, to reduce distortion and chromatic aberration caused by the characteristics of the optical system in the imaging device used to acquire the first image. Furthermore, the geometric transformation may employ a projection method different from the central projection method and be performed to display an image acquired using an optical system that distorts the subject and images a wide area (e.g., a fisheye lens), resulting in an image represented by a projection method or display method different from that of the original image. The projection methods of the image obtained by the geometric transformation include equidistant projection, equisolid angle projection, orthogonal projection, stereographic projection, central projection, etc. The display methods of the image obtained by the geometric transformation include azimuthal projection, cylindrical projection, conic projection, etc.

[0013] The degradation of image quality due to geometric transformation in this embodiment is caused by a decrease in resolution or the occurrence of aliasing noise. The decrease in resolution occurs due to the relative shift of frequency components to the lower frequency side with respect to the Nyquist frequency. On the other hand, aliasing noise is the occurrence of a false structure in the image that does not exist in the original subject due to the folding back of relatively high frequency components to the lower frequency side with respect to the Nyquist frequency (aliasing). Since the frequency components in the image change depending on the amount of deformation of the image in the geometric transformation, the aliasing noise can obtain a theoretical value (calculated value) based on the amount of deformation.

[0014] In this embodiment, the information regarding the amount of deformation of the image is represented by the ratio (expansion rate or reduction rate) of the corresponding shapes (line segments or areas) of the images before and after the geometric transformation. However, the information regarding the amount of deformation is not limited to this. For example, it may be represented by the amount of movement from a point in the first image to a point in the second image corresponding to the point in the first image. The amount of deformation of the image may vary depending on the position within the image depending on the method of geometric transformation. At this time, the degradation of image quality due to geometric transformation also varies depending on the position within the image. Note that a pixel in an image is the region of the image corresponding to one pixel of the imaging element in the imaging device used to acquire the image. Furthermore, the information regarding the amount of deformation may have the amount of deformation for a plurality of different regions or pixels in the first image.

[0015] By inputting the information regarding the amount of deformation of the first image in the geometric transformation applied to the first image and the second image into a machine learning model, it is possible to cause the machine learning model to perform correction processing corresponding to the degradation of image quality for each pixel of the second image. Therefore, since it is possible to reduce undercorrection or overcorrection due to geometric transformation, it becomes possible to accurately correct the degradation of image quality due to geometric transformation in the second image.

[0016] In this embodiment, the machine learning model is generated by performing learning using a neural network. The machine learning model may be learned by genetic programming, a Bayesian network, or the like. As the neural network, CNN (Convolutional Neural Network), GAN (Generative Adversarial Network), RNN (Recurrent Neural Network), or the like can be adopted. In the neural network, a filter for convolution on an image, a bias to be added, and an activation function for performing a non-linear transformation are used. The filter and the bias are called weights and are updated (learned) using training images and correct images. In this embodiment, this process is called the learning phase. Further, the image processing method in this embodiment inputs an image generated by geometric transformation and information regarding the above-described deformation amount into the machine learning model, and performs a process of outputting an estimated image in which deterioration in image quality (degradation in resolution) caused by performing geometric transformation on the image is corrected. In this embodiment, this process is called the estimation phase. Note that the above image processing method is an example, and the present invention is not limited thereto. Details of other image processing methods and the like are described in the following examples.

[0017] [Example 1] Referring to FIGS. 2 and 3, the image processing system 100 according to Example 1 will be described. In this example, a process of correcting deterioration in image quality due to geometric transformation is learned and executed in the machine learning model. FIG. 2 is a block diagram of the image processing system 100 in this example. FIG. 3 is an external view of the image processing system 100. The image processing system 100 includes a learning device 101 and an imaging device 102, and the learning device 101 and the imaging device 102 are connected to each other by a wired or wireless network 103.

[0018] The learning device 101 includes a storage unit 111, an acquisition unit 112, a generation unit 113, and an update unit 114, and determines the weights of the machine learning model.

[0019] The imaging device 102 includes an optical system 121, an image sensor 122, an image estimation unit 123, a storage unit 124, a recording medium 125, a display unit 126, and a system controller 127. The optical system 121 collects light incident from the subject space and generates a subject image. The optical system 121 has functions such as zoom, aperture adjustment, and autofocus as needed. In this embodiment, it is assumed that the optical system 121 has distortion aberration. The image sensor 122 converts the subject image generated by the optical system 121 into an electrical signal and generates the original image. The image sensor 122 is, for example, a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal-Oxide Semiconductor) sensor.

[0020] The image estimation unit 123 comprises an acquisition unit 123a, a calculation unit 123b, and an estimation unit 123c. The image estimation unit 123 acquires the original image and generates an input image by geometric transformation. Furthermore, it generates an estimated image using a machine learning model. The degradation of image quality due to geometric transformation is corrected using a multi-layer neural network. The weight information in the multi-layer neural network is generated by the learning device 101, and the imaging device 102 reads the weight information from the storage unit 111 via the network 103 in advance and stores it in the storage unit 124. The weight information to be stored may be the numerical weights themselves or in an encoded format. Details regarding weight updates and the generation of estimated images using the weights will be described later. The image estimation unit 123 has the function of generating an output image by performing development processing and other image processing as needed. The estimated image may also be used as the output image. The image estimation unit 123 can be a processor within the imaging device 102, an external device, or another storage medium.

[0021] The recording medium 125 records the output image. The display unit 126 displays the output image when the user gives an instruction to output the output image. The above operations are controlled by the system controller 127.

[0022] Next, with reference to Figures 4 and 5, the method for updating weights (weight information) performed by the learning device 101 in this embodiment (method for manufacturing a trained model) will be described. Figure 4 is a diagram showing the flow of the learning phase. Figure 5 is a flowchart related to weight updating. Each step in Figure 5 is mainly performed by the acquisition unit 112, the generation unit 113, and the update unit 114.

[0023] First, in step S101, the ground truth patch, training patch, and deformation patch are obtained. The ground truth patch, training patch, and deformation patch are generated by the generation unit 113. A patch refers to an image with a predetermined number of pixels (for example, 64 x 64 pixels). The generation of the ground truth patch, training patch, and deformation patch will be described later.

[0024] Next, in step S102, the generation unit 113 inputs the training patch and deformation patch into a multilayer neural network to generate an estimated patch. The estimated patch is an image obtained from the training patch by a machine learning model, and ideally matches the ground truth patch. The convolutional layer CN and the deconvolutional layer DC calculate the sum of the input, the convolution of the filter, and the bias, and process the result with an activation function. The initial values ​​of each component of the filter and the bias are arbitrary and are determined by random numbers in this embodiment. For example, the activation function can be ReLU (Rectified Linear Unit) or the sigmoid function. The output of each layer except the final layer is called a feature map. Skip connections 32 and 33 synthesize the feature maps output from non-contiguous layers. The feature maps may be synthesized by taking element-wise sums or by concatenation in the channel direction. In this embodiment, element-wise sums are used. Skip connection 31 takes the estimated residuals of the training patch and the ground truth patch and the training patch and sums them to generate an estimated patch. In this embodiment, the neural network configuration shown in Figure 4 is used, but the present invention is not limited to this.

[0025] Next, in step S103, the update unit 114 updates the neural network weights based on the error between the estimated patch and the ground truth patch. In this embodiment, the weights include the filter components and biases of each layer. Backpropagation is used to update the weights, but the present invention is not limited thereto. In the case of mini-batch learning, the errors of multiple ground truth patches and the multiple estimated patches corresponding to them are calculated, and the weights are updated. For the loss function, for example, the L2 norm or the L1 norm may be used. However, the present invention is not limited thereto, and online learning or batch learning may also be used.

[0026] Next, in step S104, the update unit 114 determines whether the weight update is complete. The completion of the update can be determined by whether the number of weight update iterations has reached a predetermined number, or whether the amount of change in the weight during the update is less than a predetermined value. If it is determined that the weight update is not complete, the process returns to step S101, and the acquisition unit 112 acquires one or more sets of correct answer patches, training patches, and deformation amount patches. On the other hand, if it is determined that the weight update is complete, the update unit 114 terminates the learning process and stores the weight information in the storage unit 111.

[0027] Next, the method for generating training data will be explained. The training data consists of correct answer patches, training patches, and deformation amount patches, and is mainly generated by the generation unit 113.

[0028] First, the generation unit 113 acquires the correct image 10, the first training image 12, and information 11 related to the optical system corresponding to the first training image 12 from the storage unit 111.

[0029] The ground truth image 10 consists of multiple images, and may be images acquired by the imaging device 102 or computer graphics (CG) images. The ground truth image 10 may be represented in grayscale or may have multiple channel components. Furthermore, by using images obtained by imaging various subjects as the ground truth image 10, the robustness of the machine learning model to diverse subjects can be improved. For example, the image may have edges, textures, gradients, and flat areas with varying intensities and directions. If necessary, information regarding the optical system corresponding to the ground truth image 10 may be stored in the memory unit 111.

[0030] In this embodiment, the optical system information 11 is information about the distortion aberration of the optical system used to acquire the first training image 12, and is stored in the storage unit 124 as a lookup table representing the relationship between the ideal image height and the actual image height of the optical system. The ideal image height is the image height at which an image is formed in the case of no aberration, and the actual image height is the image height at which an image is actually formed when distortion aberration is taken into account. In addition, a lookup table is generated for each imaging condition. The imaging conditions are, for example, focal length, F-number, and subject distance. The distortion aberration D[%] is expressed by the following equation (1) using the ideal image height r and the actual image height r'. D = (r'-r) / r·100…(1) However, the information 11 regarding the optical system is not limited to a lookup table representing the relationship between the ideal image height and the actual image height of the optical system, but may also be stored as the amount of distortion of the optical system. For example, it may be a lookup table representing the relationship between the ideal image height and the amount of distortion, or the relationship between the actual image height and the amount of distortion.

[0031] In this embodiment, the first training image 12 is an image obtained by capturing the same subject as the ground truth image 10, and has distortion aberration originating from the optical system. Alternatively, the first training image 12 may be an image generated by applying a geometric transformation based on the ground truth image 10 and information about the optical system corresponding to the ground truth image 10.

[0032] Furthermore, before applying geometric transformations to the ground truth image 10, processing may be performed to reduce aliasing noise occurring in the first training image 12 (anti-aliasing). By applying anti-aliasing to the ground truth image 10 according to the amount of deformation, aliasing noise occurring in the first training image 12 can be reduced to a desired level.

[0033] Next, information 14 regarding the deformation amount (first deformation amount) of the second training image 13 and the first training image 12 in the geometric transformation is generated. The second training image 13 and the deformation amount information 14 are calculated from the optical system information 11 and the first training image 12.

[0034] The second training image 13 is obtained by applying a geometric transformation to the first training image 12 based on the optical system information 11. The second training image 13 may also be interpolated as needed. Known interpolation methods such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation can be used. Furthermore, if the optical axis does not coincide with the center of each image before and after the transformation, the amount of shift from the optical axis to the center of each image must be considered. The second training image 13 may also be an undeveloped RAW image. When training is performed using RAW images and developed images for the second training image 13 and the ground truth image 10, the generated machine learning model can perform development processing in addition to correcting the degradation of image quality due to the geometric transformation. Development processing is the process of converting the RAW image into an image file such as JPEG (Joint Photographic Experts Group) or TIFF (Tag Image File Format).

[0035] The deformation amount information 14 is information represented as a scalar value or a two-dimensional map (feature map), and indicates the deformation amount of corresponding shapes in the first training image 12 and the second training image 13. Note that multiple deformation amounts of corresponding shapes in the first training image 12 and the second training image 13 may be obtained for each position. The shape is, for example, the distance between two corresponding points (line segments) in the first training image 12 and the second training image 13, or the area of ​​a corresponding region. When the deformation amount is expressed as a ratio, it can be expressed as, for example, a magnification ratio where the value increases as the image is magnified and decreases as it is reduced, or a reduction ratio where the value decreases as the image is magnified and increases as it is reduced. Furthermore, the deformation amount of an image may be expressed using the difference (change amount) of corresponding shapes in the image before and after the geometric transformation. In addition, the deformation amount of an image may be expressed using the amount of movement from one point in the first image 22 to one point in the second image corresponding to one point in the first image 22.

[0036] In this embodiment, the deformation amount information 14 consists of two or more two-dimensional maps, each showing the deformation amount in a different direction. The deformation amount information 14 in this embodiment is represented by two types of two-dimensional maps, each corresponding to the horizontal and vertical directions, which are the pixel arrangement directions. The deformation amount in the horizontal direction is a value calculated using the distance between any two points in the horizontal direction within the second training image 13 and the distance between two points in the first training image 12 that correspond to the distance between any two points in the horizontal direction within the second training image 13. The horizontal two-dimensional map is generated by determining multiple deformation amounts based on the distances between any two different points in the first training image 12 and the second training image 13. The vertical two-dimensional map can be generated in a similar manner. By using two-dimensional maps that include many deformation amounts at different positions in the first training image 12 and the second training image 13, it is possible to generate a neural network that can correct image quality degradation with high accuracy.

[0037] Although examples of two types of two-dimensional maps showing deformation amounts in the horizontal and vertical directions have been shown, deformation amounts in multiple directions that are different from each other are acceptable. For example, it could be two directions tilted at 45 degrees and 135 degrees from the horizontal, or two directions in the concentric circle direction and the radial direction. The deformation amount information 14 may be calculated from the deformation amount for a portion of the image, or it may be calculated from the deformation amount of that portion of the image to calculate the deformation amount for all corresponding pixels in the second training image 13 by interpolation or other means. Furthermore, the deformation amount information 14 may be subjected to normalization processing.

[0038] Furthermore, multiple sets of second training images 13 and deformation amount information 14 may be extracted from the first training image 12 and optical system information 11. Depending on the deformation amount indicated by the deformation amount information 14, there may be a bias in the number of patches extracted. For example, by extracting more patches from areas with large deformation amounts, it is possible to update the weights that have a high effect in correcting image quality degradation.

[0039] Regarding the relationship between the second training image 13 and the ground truth image 10, the sampling pitch of the second training image 13 and the sampling pitch of the ground truth image 10 may be different, as long as the second training image 13 and the ground truth image 10 contain the same subject in their respective regions. For example, by using the ground truth image 10 in combination with the second training image 13, which has a smaller sampling pitch than the ground truth image 10, as training data, it is possible to generate a machine learning model that can perform upscaling in addition to correcting the degradation of image quality in geometric transformations. Upscaling is a process in the estimation phase that makes the sampling pitch of the output image smaller than the sampling pitch of the input image.

[0040] Finally, ground truth patches, training patches, and deformation patches are generated. The ground truth patches, training patches, and deformation patches are generated by extracting images with a predetermined number of pixels from regions representing the same subject from the ground truth image 10, the second training image 13, and the deformation information 14, respectively. Alternatively, the ground truth image 10, the second training image 13, and the deformation information 14 may be used as the ground truth patch, training patch, and deformation patch, respectively. Furthermore, while the deformation patches in this embodiment have different pixel values ​​depending on their position within the patch, the pixel values ​​within the patch may be the same. For example, a patch may be used in which all pixels have the average value of the pixel values ​​within the deformation patch or the pixel value at the center position. Alternatively, instead of deformation patches, the average value of the pixel values ​​within a patch or the pixel value at the center position may be used as scalar values ​​for training.

[0041] Furthermore, images acquired by the imaging device 102 may be used to generate training data. In this case, a second training image 13 can be generated by using the acquired image as the first training image 12. The ground truth image 10 is obtained by imaging the same subject as the first training image 12 using an optical system with less distortion aberration than the optical system 121.

[0042] Next, with reference to Figures 1 and 6, we will describe in detail the image processing method (estimation phase) using a trained machine learning model. Figure 1 is a diagram showing the flow of the estimation phase. Figure 6 is a flowchart of the estimation phase in this embodiment. Each step in Figure 6 is performed in the acquisition unit 123a, the calculation unit 123b, or the estimation unit 123c of the image estimation unit 123.

[0043] First, in step S201, the acquisition unit 123a acquires information about the optical system 21, a first image 22, and weight information. The information about the optical system 21 is pre-stored in the storage unit 124, and the acquisition unit 123a acquires the information about the optical system 21 corresponding to the imaging conditions. The weight information is read out in advance from the storage unit 111 and stored in the storage unit 124. The information about the optical system 21 corresponds to the information about the optical system 11 in the learning phase. The first image 22 corresponds to the first training image 12 in the learning phase.

[0044] Next, in step S202, the calculation unit 123b generates a second image 23 from the optical system information 21 and the first image 22. In this embodiment, the second image 23 is an image generated by applying a geometric transformation to the first image 22 in order to reduce the distortion aberration that occurred in the first image 22 due to the optical system 121.

[0045] The second image 23 corresponds to the second training image in the learning phase and is obtained by applying a geometric transformation to the first image 22 based on the optical system information 21. The second image 23 may also be interpolated as needed.

[0046] Next, in step S203, the calculation unit 123b uses the optical system information 21 and the first image 22 to generate information 24 regarding the deformation amount (second deformation amount) of the first image 22 in the geometric transformation. The deformation amount information 24 represents the deformation amount when the second image 23 is generated in step S202. In this embodiment, the deformation amount information 24 consists of two types of two-dimensional maps, each showing the deformation amount in the horizontal or vertical direction. The deformation amount information 24 in this embodiment will now be explained with reference to Figure 7. Figure 7(A) is an example of the first image 22, and Figure 7(B) is an example of the second image 23. Figure 7(C) is a two-dimensional map showing the horizontal deformation amount when the second image 23 is generated from the first image 22. Figure 7(D) is a two-dimensional map showing the vertical deformation amount when the second image 23 is generated from the first image 22. In this embodiment, the two types of two-dimensional maps shown in Figures 7(C) and 7(D) constitute the deformation amount information 24. The method for generating the deformation amount information 24 is the same as for the deformation amount information 14. Note that steps S202 and S203 in this embodiment may be processed simultaneously.

[0047] Furthermore, information regarding deformation amounts 24 can be obtained when multiple second images 23 are generated in step S202 using multiple first images 22 and information regarding multiple optical systems 21 corresponding to the multiple first images 22. At this time, each of the multiple first images 22 is corrected for distortion aberration by geometric transformation.

[0048] Furthermore, the image estimation unit 123 may be included in an image processing device different from the imaging device 102. In that case, the image acquired by the acquisition unit 123a may be an image corresponding to the second image 23 instead of the first image 22. In other words, step S202 may be performed in advance using an image processing device different from the image estimation unit 123 to generate the second image 23 from the optical system information 21 and the first image 22.

[0049] Next, in step S204, the estimation unit 123c inputs the second image 23 and information 24 regarding the amount of deformation into a machine learning model to generate an estimated image (third image) 25. The third image 25 is an image obtained by correcting the degradation of image quality due to geometric transformation from the second image 23.

[0050] As described above, according to this embodiment, it is possible to provide an image processing system that can accurately correct the degradation of image quality caused by geometric transformation in the second image 23, in which distortion aberration has been reduced by geometric transformation, using a machine learning model.

[0051] [Example 2] Next, with reference to Figures 8 and 9, the image processing system 200 according to Embodiment 2 will be described. In this embodiment, a machine learning model is trained and executed to correct the degradation of image quality due to geometric transformation. The image processing system 200 in this embodiment differs from Embodiment 1 in that it acquires the original image from the imaging device 202 and the image estimation device 203 performs image processing. Figure 8 is a block diagram of the image processing system 200 in this embodiment. Figure 9 is an external view of the image processing system 200. The image processing system 200 includes a learning device 201, an imaging device 202, an image estimation device 203, a display device 204, a storage medium 205, an output device 206, and a network 207.

[0052] The learning device 201 includes a storage unit 201a, an acquisition unit 201b, a generation unit 201c, and an update unit 201d, and determines the weights of the machine learning model.

[0053] The imaging device 202 has an optical system 202a and an image sensor 202b, and acquires a first image 22. The optical system 202a collects light incident from the subject space and generates a subject image. The image sensor 202b converts the subject image generated by the optical system 202a into an electrical signal and generates the first image 22. In this embodiment, the optical system 202a has a fisheye lens employing an equisolid angle projection method, and the subject in the first image 22 has distortion corresponding to the equisolid angle projection method. However, the optical system 202a is not limited to this, and an optical system employing any projection method may be used.

[0054] The image estimation device 203 includes a storage unit 203a, an acquisition unit 203b, a generation unit 203c, and an estimation unit 203d. The image estimation device 203 generates an estimated image using a machine learning model. In this embodiment, the geometric transformation is a transformation from a first image 22 represented by an equisolid angle projection method (first projection method) to a second image 23 represented by a central projection method (second projection method). However, this embodiment is not limited to this, and images represented by any projection method or representation method may be used. The degradation of image quality due to the geometric transformation is corrected using a machine learning model, and the weight information of the machine learning model is generated by the learning device 201. The image estimation device 203 reads the weight information from the storage unit 201a via the network 207 and stores it in the storage unit 203a. Note that the weight update performed by the learning device 201 is the same as that of the learning device 101 in Embodiment 1, so the explanation is omitted. Further details regarding the method of generating training data and image processing using weights will be described later. The image estimation device 203 may also have a function to generate an output image by performing development processing or other image processing as needed.

[0055] The output image generated by the image estimation device 203 is output to at least one of the display device 204, storage medium 205, or output device 206. The display device 204 is, for example, a liquid crystal display or a projector. The display device 204 may be used to allow the user to perform editing work while checking the image in progress. The storage medium 205 is, for example, a semiconductor memory, a hard disk, or a server on a network, and stores the output image. The output device 206 is, for example, a printer.

[0056] The recording medium 125 records the output image. The display unit 126 displays the output image when the user gives an instruction to output the output image. The above operations are controlled by the system controller 127.

[0057] Next, we will explain how to generate the training data. The training data consists of ground truth patches, training patches, and deformation patches, and is mainly generated by the generation unit 201c.

[0058] First, the acquisition unit 201b acquires the ground truth image 10 and information 11 related to the optical system corresponding to the ground truth image 10 from the storage unit 201a. In this embodiment, the ground truth image 10 is an image acquired by an optical system employing a central projection method.

[0059] In this embodiment, the optical system information 11 includes information about the projection method employed by the optical system used to acquire each image. The projection method is a method by which an optical system with focal length f represents an object located at an angle θ from the optical axis on a two-dimensional plane, and is expressed using the image height r of the optical system.

[0060] The equisolid angle projection method is a projection method characterized by the fact that the solid angle of the subject is proportional to its area on a two-dimensional plane. An optical system employing the equisolid angle projection method represents the subject on a two-dimensional plane according to the following equation (2). r = 2·f·sin(θ / 2) ... (2)

[0061] Furthermore, an optical system employing the central projection method represents the subject on a two-dimensional plane according to the following equation (3). r = f·tanθ …(3)

[0062] Furthermore, the information regarding the optical system 11 is not limited to the relationship between the angle from the optical axis of the subject and the image height of the optical system; it is sufficient if it can associate the position of the subject with its position on the two-dimensional plane in which it is represented.

[0063] Next, the first training image 12 is generated. In this embodiment, the first training image 12 is an image obtained by imaging the same subject as the ground truth image 10, and is an image acquired by an optical system employing an equisolid angle projection method. However, the projection method for the first training image 12 is not limited to this.

[0064] Next, a second training image 13 and deformation amount information 14 are generated. The second training image 13 and deformation amount information 14 are calculated from the optical system information 11 and the first training image 12. The second training image 13 is an image generated by applying a geometric transformation to the first training image 12, which is represented using an equisolid angle projection method, and is represented using a central projection method. The second training image 13 may also be interpolated as needed. However, the second training image 13 is not limited to this, and should be represented using at least the same projection method as the ground truth image 10.

[0065] The deformation amount information 14 is generated in the same manner as in Example 1. The correct answer patch, training patch, and deformation amount patch are also generated in the same manner as in Example 1.

[0066] Next, with reference to Figures 1 and 10, the image processing method using the trained machine learning model will be described in detail. Figure 10 is a flowchart of the estimation phase in this embodiment. Each step in Figure 10 is performed by the acquisition unit 203b, the generation unit 203c, and the estimation unit 203d.

[0067] First, in step S301, the acquisition unit 203b acquires information about the optical system 21, the first image 22, and weight information. In this embodiment, the information about the optical system 21 includes information about the projection method employed by the optical system used to acquire the first image 22. The weight information is read in advance from the storage unit 201a and stored in the storage unit 203a.

[0068] Next, in step S302, the generation unit 203c generates (calculates) a second image 23 using the optical system information 21 and the first image 22. The second image 23 is an image generated by applying a geometric transformation to the first image 22, which is represented using the equisolid angle projection method, and is represented using the central projection method. The second image 23 may also be interpolated as needed.

[0069] Next, in step S303, the generation unit 203c generates deformation amount information 24 using optical system information 21 and the first image 22. In this embodiment, deformation amount information 24 consists of two types of two-dimensional maps showing the horizontal and vertical deformation amounts associated with the transformation (geometric transformation) from equisolid angle projection to central projection. Here, the deformation amount information 24 will be explained with reference to Figure 11. Figure 11(A) is an example of the first image 22 represented using equisolid angle projection. Figure 11(B) is an example of the second image 23 represented using central projection. Figure 11(C) is a two-dimensional map showing the horizontal deformation amount when the second image 23 is generated from the first image 22. Figure 11(D) is a two-dimensional map showing the vertical deformation amount when the second image 23 is generated from the first image 22. In this embodiment, the two types of two-dimensional maps shown in Figures 11(C) and 11(D) constitute the deformation amount information 24. The method for generating the deformation amount information 24 is the same as for the deformation amount information 14. Note that steps S302 and S303 in this embodiment may be processed simultaneously.

[0070] Next, in step S304, the estimation unit 203d inputs the second image 23 and the deformation amount information 24 into the machine learning model to generate a third image 25. The third image 25 is an image in which the degradation of image quality due to geometric transformation in the second image 23 has been corrected.

[0071] As described above, this embodiment provides an image processing system that can accurately correct the degradation of image quality in the second image 23, which has had its projection scheme transformed by geometric transformation, using a machine learning model.

[0072] [Example 3] Next, with reference to Figures 12 and 13, the image processing system 300 according to Embodiment 3 will be described. In this embodiment, a machine learning model is trained and executed to correct the degradation of image quality due to geometric transformation.

[0073] The image processing system 300 of this embodiment differs from Embodiment 1 in that it has a control device 304 that acquires optical system information 21 and a first image 22 from an imaging device 302 and makes requests to an image estimation device (image processing device) 303 regarding image processing of the first image 22.

[0074] Figure 12 is a block diagram of the image processing system 300 in this embodiment. The image processing system 300 includes a learning device 301, an imaging device 302, an image estimation device 303, and a control device 304. In this embodiment, the learning device 301 and the image estimation device 303 may be servers. The control device 304 is a user terminal such as a personal computer or a smartphone. The control device 304 is connected to the image estimation device 303 via a network 305. The image estimation device 303 is connected to the learning device 301 via a network 306. In other words, the control device 304 and the image estimation device 303, as well as the image estimation device 303 and the learning device 301, are configured to communicate with each other.

[0075] The learning device 301 and imaging device 302 in the image processing system 300 have the same configuration as the learning device 201 and imaging device 202, respectively, so their descriptions are omitted.

[0076] The image estimation device 303 includes a storage unit 303a, an acquisition unit (acquisition means) 303b, a generation unit (generation means) 303c, an estimation unit (estimation means) 303d, and a communication unit (receiving means) 303e. The storage unit 303a, acquisition unit 303b, generation unit 303c, and estimation unit 303d in the image estimation device 303 are the same as the storage unit 203a, acquisition unit 203b, generation unit 203c, and estimation unit 203d, respectively.

[0077] The control device 304 includes a communication unit (transmission means) 304a, a display unit (display means) 304b, an input unit (input means) 304c, a processing unit (processing means) 304d, and a recording unit 304e. The communication unit 304a can transmit a request to the image estimation device 303 to perform processing on the first image 22. It can also receive the output image processed by the image estimation device 303. The communication unit 304a may also communicate with the imaging device 302. The display unit 304b displays various information. The various information displayed by the display unit 304b includes, for example, the first image 22, the second image 23, or the output image received from the image estimation device 303. The input unit 304c can receive instructions from the user to start image processing. The processing unit 304d can perform arbitrary image processing on the output image received from the image estimation device 303. The recording unit 304e stores information 21 about the optical system acquired from the imaging device 302, the first image 22, and the output image received from the image estimation device 303.

[0078] The method by which the first image 22 to be processed is transmitted to the image estimation device 303 is not limited; for example, the first image 22 may be uploaded to the image estimation device 303 at the same time as S401, or it may be uploaded to the image estimation device 303 before S401. Also, the first image 22 may be an image stored on a server different from the image estimation device 303.

[0079] Next, the generation of the output image (estimated image) in this embodiment will be explained. Figure 13 is a flowchart of the estimation phase in this embodiment.

[0080] The operation of the control device 304 will now be described. In this embodiment, image processing is initiated by the user via the control device 304 when an instruction to start image processing is given.

[0081] First, in step S401 (the first transmission step), the communication unit 304a transmits a request for processing of the first image 22 to the image estimation device 303. In step S401, the control device 304 may also transmit, along with the request for processing of the first image 22, an ID for user authentication and shooting conditions corresponding to the first image 22.

[0082] Next, in step S402 (first receiving step), the communication unit 304a receives the third image 25 generated by the estimation device 303.

[0083] Next, the operation of the image estimation device 303 will be described. First, in step S501 (second receiving step), the communication unit 303e receives a request for processing the first image 22 transmitted from the communication unit 304a. Upon receiving the instruction to process the first image 22, the image estimation device 303 executes the processing from step S502 onward.

[0084] Next, in step S502, the acquisition unit 303b acquires information about the optical system 21 and the first image 22. In this embodiment, the information about the optical system 21 and the first image 22 are transmitted from the control device 304. Note that the processing in steps S501 and S502 may be performed simultaneously. Also, steps S503 to S505 are the same as steps S202 to S204, so their explanation is omitted.

[0085] Next, in step S506 (the second transmission step), the image estimation device 303 transmits the third image 25 to the control device 304.

[0086] As described above, this embodiment provides an image processing system that can accurately correct the degradation of image quality due to geometric transformation in the second image 23 using a machine learning model. In this embodiment, the control device 304 only requests processing for a specific image. The actual image processing is performed by the image estimation device 303. Therefore, if the control device 304 is used as a user terminal, the processing load on the user terminal can be reduced. Consequently, the user can obtain the output image with a low processing load.

[0087] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions. The image processing apparatus in the present invention can be any device having the image processing function of the present invention, and can be implemented in the form of an imaging device or a PC.

[0088] According to each embodiment, it is possible to provide an image processing method, an image processing system, and a program that can accurately correct the degradation of image quality caused by geometric transformation in an image that has undergone geometric transformation using a machine learning model.

[0089] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its essence.

[0090] [Method 1] The steps include obtaining a second image by applying a geometric transformation to the first image, A step of obtaining information regarding the amount of deformation of the first image in the geometric transformation, An image processing method characterized by comprising the steps of inputting the second image and the information regarding the amount of deformation into a machine learning model and generating a third image.

[0091] [Method 2] The image processing method according to Method 1, wherein the information regarding the amount of deformation includes the ratio of the distance between two points in the first image to the distance between two points corresponding to the two points in the second image.

[0092] [Method 3] The image processing method according to either method 1 or 2, wherein the information relating to the amount of deformation includes the ratio of the area of ​​the region in the first image to the area of ​​the region corresponding to the region in the second image.

[0093] [Method 4] The image processing method according to any one of methods 1 to 3, characterized in that the information relating to the amount of deformation includes the amount of movement from one point in the first image to a point in the second image corresponding to the aforementioned point.

[0094] [Method 5] The image processing method according to any one of methods 1 to 4, characterized in that the information relating to the amount of deformation includes the value of the amount of deformation for each pixel position in the first image.

[0095] [Method 6] The image processing method according to any one of methods 1 to 5, characterized in that the geometric transformation is a transformation in which the amount of deformation differs for each pixel position in the first image.

[0096] [Method 7] The image processing method according to any one of methods 1 to 6, characterized in that the geometric transformation is a transformation from a first projection scheme of the first image to a second projection scheme of the second image.

[0097] [Program 8] A program characterized by causing a computer to execute the image processing method described in any one of Methods 1 to 7.

[0098] [Composition 9] A storage medium characterized by storing the program described in Program 8.

[0099] [Configuration 10] A means for obtaining a second image obtained by applying a geometric transformation to the first image, Means for obtaining information regarding the amount of deformation of the first image in the geometric transformation, An image processing apparatus characterized by having means for inputting the second image and the information regarding the amount of deformation into a machine learning model and generating a third image.

[0100] [Composition 11] A means for acquiring a first training image obtained by imaging using an optical system and an image sensor, information about the optical system, and a ground truth image. A means for generating a second training image by applying a geometric transformation to the first training image based on information about the optical system, Means for obtaining information regarding the amount of deformation of the first training image in the geometric transformation, A means for inputting the second training image and the information regarding the amount of deformation into a machine learning model and generating an estimated image, A learning device characterized by having means for updating the weights of a neural network based on the aforementioned correct image and the aforementioned estimated image.

[0101] [Method 12] A step of acquiring a first training image obtained by taking images using an optical system and an image sensor, information about the optical system, and a ground truth image. The steps include generating a second training image by applying a geometric transformation to the first training image based on the information regarding the optical system, A step of obtaining information regarding the amount of deformation of the first training image in the geometric transformation, The steps include inputting the second training image and the information regarding the amount of deformation into a machine learning model to generate an estimated image, A method for manufacturing a trained model, characterized by having means for updating the weights of a neural network based on the ground truth image and the estimated image.

[0102] [Program 13] A program characterized by causing a computer to execute the method for manufacturing a trained model described in Method 12.

[0103] [Composition 14] An image processing system including an imaging device and a learning device capable of communicating with the imaging device, The learning device is A means for acquiring a first training image obtained by imaging using an optical system and an image sensor, information about the optical system, and a ground truth image. A means for generating a second training image by applying a geometric transformation to the first training image based on information about the optical system, Means for obtaining information regarding the first deformation amount of the first training image in the geometric transformation of the first training image, A means for inputting the second training image and the information regarding the first deformation amount into a machine learning model and generating an estimated image, The system includes means for updating the weights of the neural network based on the aforementioned correct image and the aforementioned estimated image. The imaging device comprises an optical system, an image sensor, and an image estimation unit. The image estimation unit is, Means for acquiring a first image acquired using the imaging device and information regarding the optical system of the imaging device, A means for generating a second image by applying a geometric transformation to the first image based on information about the optical system, Means for obtaining information regarding the second deformation amount of the first image in the geometric transformation of the first image, An image processing system characterized by comprising means for inputting the second image and information relating to the second amount of deformation into a machine learning model and generating a third image.

[0104] [Composition 15] An image processing system including a control device and an image processing device capable of communicating with the control device, The control device has means for transmitting a request to the image processing device to perform processing on the first image acquired by imaging using the optical system and image sensor. The aforementioned image processing device is Means for receiving the aforementioned request, An acquisition means for acquiring the first image and the optical system information, Means for acquiring a second image obtained by applying a geometric transformation to the first image based on the information of the optical system, Means for obtaining information regarding the amount of deformation of the first image in the geometric transformation, An image processing system characterized by having means for inputting the second image and the information regarding the amount of deformation into a machine learning model and generating a third image. [Explanation of Symbols]

[0105] 22 Image 1 23. Image 2 24. Information regarding deformation amount

Claims

1. The steps include obtaining a second image by applying a geometric transformation to the first image, A step of obtaining information regarding the amount of deformation of the first image in the geometric transformation, An image processing method characterized by comprising the steps of inputting the second image and the information regarding the amount of deformation into a machine learning model to generate a third image.

2. The image processing method according to claim 1, wherein the information relating to the amount of deformation includes the ratio of the distance between two points in the first image to the distance between two points corresponding to the two points in the second image.

3. The image processing method according to claim 1, wherein the information relating to the amount of deformation includes the ratio of the area of ​​the region in the first image to the area of ​​the region corresponding to the region in the second image.

4. The image processing method according to claim 1, characterized in that the information relating to the amount of deformation includes the amount of movement from one point in the first image to a point in the second image corresponding to the one point.

5. The image processing method according to claim 1, characterized in that the information relating to the amount of deformation includes the value of the amount of deformation for each pixel position in the first image.

6. The image processing method according to claim 1, characterized in that the geometric transformation is a transformation in which the amount of deformation differs for each pixel position in the first image.

7. The image processing method according to claim 1, characterized in that the geometric transformation is a transformation from a first projection scheme of the first image to a second projection scheme of the second image.

8. A program characterized by causing a computer to execute the image processing method described in any one of claims 1 to 7.

9. A storage medium characterized by storing the program described in claim 8.

10. A means for obtaining a second image obtained by applying a geometric transformation to the first image, Means for obtaining information regarding the amount of deformation of the first image in the geometric transformation, An image processing apparatus characterized by having means for inputting the second image and the information regarding the amount of deformation into a machine learning model and generating a third image.

11. A means for acquiring a first training image obtained by imaging using an optical system and an image sensor, information about the optical system, and a ground truth image. A means for generating a second training image by applying a geometric transformation to the first training image based on the information about the optical system, Means for obtaining information regarding the amount of deformation of the first training image in the geometric transformation, A means for inputting the second training image and the information regarding the amount of deformation into a machine learning model and generating an estimated image, A learning device characterized by having means for updating the weights of a neural network based on the aforementioned correct image and the aforementioned estimated image.

12. A step of acquiring a first training image obtained by taking images using an optical system and an image sensor, information about the optical system, and a ground truth image. The steps include generating a second training image by applying a geometric transformation to the first training image based on the information regarding the optical system, A step of obtaining information regarding the amount of deformation of the first training image in the geometric transformation, The steps include inputting the second training image and the information regarding the amount of deformation into a machine learning model to generate an estimated image, A method for manufacturing a trained model, characterized by having means for updating the weights of a neural network based on the ground truth image and the estimated image.

13. A program characterized by causing a computer to execute the method for manufacturing a trained model described in claim 12.

14. An image processing system including an imaging device and a learning device capable of communicating with the imaging device, The learning device is A means for acquiring a first training image obtained by imaging using an optical system and an image sensor, information about the optical system, and a ground truth image. A means for generating a second training image by applying a geometric transformation to the first training image based on the information about the optical system, Means for obtaining information regarding the first deformation amount of the first training image in the geometric transformation of the first training image, A means for inputting the second training image and the information regarding the first deformation amount into a machine learning model to generate an estimated image, The system includes means for updating the weights of the neural network based on the aforementioned correct image and the aforementioned estimated image. The imaging device comprises an optical system, an image sensor, and an image estimation unit. The image estimation unit is, Means for acquiring a first image acquired using the imaging device and information regarding the optical system of the imaging device, A means for generating a second image by applying a geometric transformation to the first image based on information about the optical system, Means for obtaining information regarding the second deformation amount of the first image in the geometric transformation of the first image, An image processing system characterized by comprising means for inputting the second image and information regarding the second amount of deformation into a machine learning model and generating a third image.

15. An image processing system including a control device and an image processing device capable of communicating with the control device, The control device has means for transmitting a request to the image processing device to perform processing on the first image acquired by imaging using the optical system and image sensor. The aforementioned image processing device is Means for receiving the aforementioned request, An acquisition means for acquiring the first image and the optical system information, Means for acquiring a second image obtained by applying a geometric transformation to the first image based on the information of the optical system, Means for obtaining information regarding the amount of deformation of the first image in the geometric transformation, An image processing system characterized by having means for inputting the second image and the information regarding the amount of deformation into a machine learning model and generating a third image.

Citation Information

Patent Citations

  • Image Distortion Correction Method and Device and Electronic Device

    US20210125310A1