Image processing system, image processing method, and computer program
The image processing system enhances super-resolution model training by using both high-resolution and reduced-resolution images, minimizing errors and improving feature detection accuracy through pixel exclusion and degradation process modeling.
Patent Information
- Application Number
- JP2024159473
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2044-09-13
AI Technical Summary
Existing super-resolution models trained with high-resolution and low-resolution images or reduced-resolution images separately may not achieve sufficient accuracy due to positional, shape, and color differences, leading to inappropriate training models.
An image processing system that trains a learning model using both high-resolution and reduced-resolution images, minimizing errors by converting image resolutions and excluding pixels with significant loss, and using a degradation process reproduction model to generate accurate super-resolved images.
The system improves the accuracy of super-resolved images by considering both high-resolution and reduced-resolution image relationships, enabling precise feature detection and interpretation.
Smart Images

Figure 0007802881000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing system, an image processing method, and a computer program. [Background technology]
[0002] Conventionally, images of geographical features, such as aerial photographs or satellite images, have been used to perform processes such as interpreting building changes, creating maps, and creating city models. In order to perform processes such as interpreting building changes, creating maps, and creating city models with high accuracy, it is preferable to use images with higher resolution.
[0003] Patent Document 1 discloses a learning device that generates a learning model for inferring an estimated image corresponding to a target image with improved image quality from a target image obtained by remote sensing of a target satellite. This learning device receives, as learning data, a pair of a reference image with higher image quality than the target image and a degraded image that is a reference image with lower image quality. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2023-150145 Summary of the Invention [Problem to be solved by the invention]
[0005] In recent years, super-resolution models that generate high-resolution images from low-resolution images have been developed. If a super-resolution model is trained using a high-resolution image and a low-resolution image captured separately from the high-resolution image, the position, shape, and / or color of features in each image may differ slightly, which may result in an inappropriate training model. On the other hand, if a super-resolution model is trained using a high-resolution image and a reduced-resolution version of the high-resolution image, the appearance (appearance) of objects contained in each image may differ, which may result in an inappropriate training model. Therefore, even when using a training model trained only with pairs of high-resolution and low-resolution images, as described above, or a training model trained only with pairs of high-resolution and reduced-resolution images, it may not be possible to obtain images with a resolution high enough to sufficiently improve the accuracy of feature interpretation, detection, and the like.
[0006] An object of the present invention is to provide an image processing system, an image processing method, and a computer program that are capable of obtaining an output image by super-resolving an input image. [Means for solving the problem]
[0007] The image processing system of the present invention comprises an acquisition means for acquiring an input image in which a feature is captured at a first resolution, and an output means for outputting information regarding an output image which includes the feature and has a second resolution higher than the first resolution, which is output from the learning model when the input image is input to the learning model, and the learning model is trained to output a first learning output image which includes the specified area and has the second resolution when a first learning input image in which a specified area is captured at the first resolution is input, and to output a second learning output image which includes the specified area and has the second resolution when a second learning input image in which a specified area is captured at the second resolution is converted so that the resolution becomes the first resolution.
[0008] In addition, in the image processing system of the present invention, it is preferable that the learning model is trained so as to minimize the error between each of the first learning output image and the second learning output image and the same correct image in which a specified area is captured at the second resolution.
[0009] In the image processing system according to the present invention, it is preferable that the first learning input image undergoes hue conversion before being input to the learning model.
[0010] In addition, in the image processing system according to the present invention, it is preferable that the color of the first learning input image is converted by performing standardization or by converting the gradation value of each pixel by histogram matching.
[0011] In the image processing system according to the present invention, it is preferable that, of the pixels included in the first learning output image, pixels whose loss relative to the corresponding pixel in the correct image satisfies a predetermined condition are excluded from the learning target.
[0012] In the image processing system according to the present invention, the pixels that satisfy the predetermined condition are preferably pixels whose loss is equal to or greater than a predetermined value, or pixels that are included at a predetermined ratio in descending order of loss.
[0013] In the image processing system according to the present invention, the second learning input image is preferably generated from an image in which a predetermined area is captured at the second resolution, using a model that reproduces the deterioration process.
[0014] In addition, the image processing method of the present invention includes acquiring an input image in which a feature is captured at a first resolution, and outputting information regarding an output image which includes the feature and has a second resolution higher than the first resolution, which is output from the learning model when the input image is input to the learning model, and the learning model is trained to output a first learning output image which includes the specified area and has the second resolution when a first learning input image in which a specified area is captured at the first resolution is input, and to output a second learning output image which includes the specified area and has the second resolution when a second learning input image in which a specified area is captured at the second resolution is converted so that the resolution becomes the first resolution.
[0015] In addition, the computer program of the present invention causes a computer to acquire an input image in which a feature is captured at a first resolution, and output information regarding an output image which includes the feature and has a second resolution higher than the first resolution, which is output from the learning model when the input image is input into the learning model, and the learning model is trained to output a first learning output image which includes the specified area and has the second resolution when a first learning input image in which a specified area is captured at the first resolution is input, and to output a second learning output image which includes the specified area and has the second resolution when a second learning input image in which a specified area is captured at the second resolution is converted so that the resolution becomes the first resolution. [Effects of the Invention]
[0016] The image processing system, image processing method, and computer program according to the present invention are capable of obtaining an output image by super-resolving an input image. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a diagram illustrating an example of a configuration of an image processing system 1. FIG. [Figure 2] 10 is a flowchart illustrating an example of the flow of a learning process. [Figure 3]FIG. 10A is a schematic diagram for explaining changes in features, and FIG. 10B is a schematic diagram for explaining changes in imaging angle. [Figure 4] 10 is a flowchart illustrating an example of the flow of an output process. DETAILED DESCRIPTION OF THE INVENTION
[0018] Various embodiments of the present invention will be described below with reference to the drawings. It should be noted that the technical scope of the present invention is not limited to these embodiments, but encompasses the inventions set forth in the claims and their equivalents.
[0019] FIG. 1 is a diagram showing an example of the configuration of an image processing system 1 according to the present invention.
[0020] The image processing system 1 includes an information processing device 100 and a server device 200. The information processing device 100 and the server device 200 are examples of image processing devices, and are communicatively connected to each other via a network N. The network N is an intranet, the Internet, or the like.
[0021] The information processing device 100 is a personal computer, a notebook personal computer, a server, etc. The information processing device 100 includes an operation device 101, a display device 102, a communication device 103, a storage device 110, a processing circuit 120, etc.
[0022] The operation device 101 has input devices such as a keyboard and a mouse, and an interface circuit that acquires signals from the input devices, accepts operations by a user, and outputs to the processing circuit 120 a signal according to the user's input.
[0023] The display device 102 is an example of an output means. The display device 102 has a display configured with a liquid crystal display, an organic electroluminescence display, or the like, and an interface circuit that outputs image data to the display, and displays image data on the display in accordance with instructions from the processing circuit 120.
[0024] The communication device 103 is an example of an output means. The communication device 103 includes a wired or wireless communication interface circuit and connects the information processing device 100 to a communication network. The communication device 103 performs wired communication according to a communication protocol such as TCP / IP (Transmission Control Protocol / Internet Protocol). The communication device 103 may also perform wireless communication using a wireless communication method conforming to the IEEE (Institute of Electrical and Electronics Engineers) 802.11 standard. The communication device 103 transmits information supplied from the processing circuit 120 to an external device. The communication device 103 also supplies information received from an external device to the processing circuit 120.
[0025] The storage device 110 includes, for example, semiconductor memory such as RAM (Random Access Memory) and ROM (Read Only Memory), a fixed disk device such as a hard disk, or a portable storage device such as an optical disk. The storage device 110 stores computer programs, data, and the like used for processing by the processing circuit 120. The computer programs are installed in the storage device 110 from a server (not shown) via the communication device 103. Note that the computer programs may be installed in the storage device 110 from a computer-readable portable recording medium using a known setup program or the like. The portable recording medium is, for example, a CD-ROM, a DVD-ROM, or the like. The computer programs may be distributed from a server or the like and installed in the storage device 110.
[0026] Furthermore, the storage device 110 stores a learning model 111. When an image is input, the learning model 111 outputs an image that corresponds to the input image and has a higher resolution than the input image.
[0027] The processing circuit 120 is, for example, a CPU (Central Processing Unit). The processing circuit 120 may be an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or the like. The processing circuit 120 is connected to the operation device 101, the display device 102, the communication device 103, the storage device 110, and the like, and controls each of these components. The processing circuit 120 reads a program stored in the storage device 110 and operates in accordance with the read program, thereby functioning as a learning means 121, an acquisition means 122, and an output control means 123. The processing circuit 120 generates a learning model 111 and, using the learning model 111, generates an output image that includes features and has a higher resolution than the input image from an input image in which the features are captured.
[0028] FIG. 2 is a flowchart showing an example of the flow of the learning process executed by the information processing device 100.
[0029] An example of the operation of the learning process of the information processing device 100 will be described below with reference to the flowchart shown in Fig. 2. The flow of the operation described below is executed mainly by the processing circuitry 120 in cooperation with each element of the information processing device 100 based on a program stored in advance in the storage device 110.
[0030] First, the learning means 121 acquires a training dataset (step S101). The training dataset includes a plurality of combinations of a first training input image in which a predetermined region is captured at a first resolution and a reference image in which the predetermined region is captured at a second resolution. The first training input image and the reference image to be combined are images of the same region (a small positional error is allowed). On the other hand, the reference images may include images of the same region, but basically, different regions are captured. The learning means 121 acquires the training dataset by receiving it from the server device 200 via the communication device 103. Note that the training dataset may be pre-stored in the storage device 110, and the learning means 121 may acquire the training dataset by reading it from the storage device 110.
[0031] The predetermined area is an aerial photograph or satellite image (or an orthoimage based on the aerial photograph or satellite image) of the ground surface, which is the target area for processes such as interpreting building changes, creating maps, and creating city models, and includes features captured from the sky. Features include buildings, trees, rocks, etc. The first training input image and the correct answer image are different types of images (aerial photographs, satellite images, etc.), images captured by different imaging devices (imaging sensors), images captured at different times, images captured from different positions at different angles, or images captured under two or more of these conditions. For example, a satellite image can be used as the first training input image, and an aerial photograph with a higher resolution can be used as the correct answer image. Therefore, although the first training input image and the correct answer image include the same area, the position, shape, and / or color of features in each image are likely to be slightly different.
[0032] Next, the learning means 121 generates second learning input images by converting each correct image included in the learning dataset so that its resolution becomes the first resolution (step S102). The learning means 121 generates second learning input images from each correct image using a degradation process reproduction learning model that reproduces the degradation process. The degradation process reproduction learning model is trained to output an image that is naturally degraded when an image is input. The degradation process reproduction learning model reduces the resolution of the input image and outputs an image to which noise, defocus (image blur) and / or imaging blur have been added or which has a changed color. As the degradation process reproduction learning model, for example, BSRGAN (Blind Super Resolution Generative Adversarial Network) or the like is used. For more information on BSRGAN, see, for example, "Designing a Practical Degradation Model for Deep Blind Image Super-Resolution" by Kai Zhang, Jingyun Liang, Luc Van Gool, and Radu Timofte (https: / / arxiv.org / abs / 2103.14006).
[0033] The learning means 121 may generate the second learning input image by thinning and / or interpolating pixels of each ground truth image. In this case, the learning means 121 may add noise to each ground truth image by randomly changing the pixel values of specific pixels, apply a smoothing filter to a specific region to add defocus or imaging blur, or change the color tone by gamma correction or the like.
[0034] The image processing system 1 can generate various patterns of learning data by using images obtained by degrading each correct answer image as the second learning input image, thereby improving the learning accuracy of the learning model 111. Note that the learning means 121 may generate the second learning input image without adding noise, defocus (image blur), or imaging shake to each correct answer image and without changing the color. Furthermore, since the output of the learning model is enlarged by a constant magnification, if the resolution of the second learning input image differs from the resolution of the first learning input image, the loss with the high-resolution image may not be calculated correctly. The image processing system 1 can correctly calculate the loss with the high-resolution image by making the resolution of the second learning input image the same as the resolution of the first learning input image.
[0035] Next, the learning means 121 converts the color tone of each first learning input image included in the learning dataset (step S103). For example, the learning means 121 performs standardization on the first learning input image, converting the color tone of the first learning input image so that the gradation values (R value, G value, and B value) of each pixel included in the first learning input image form a standard normal distribution. In this case, the learning means 121 calculates the average value and variance of the gradation values of all pixels included in the first learning input image for each color (red, green, and blue), and converts the gradation value of each pixel based on the calculated average value and variance according to the following equation (1). I out =(I in -I mean ) / σ (1) where I out is the converted tone value, and I in is the tone value before conversion, and I mean is the average value of the gradation values, and σ is the variance of the gradation values. As a result, the average value of the gradation values of all pixels of the converted first learning input image becomes 0, the variance becomes 1, and the gradation values of the converted first learning input image form a standard normal distribution.
[0036] In this way, before inputting the first learning input image to the learning model 111, the learning means 121 converts the hue of the first learning input image so that the gradation values of each pixel included in the first learning input image form a standard normal distribution. This allows the image processing system 1 to prevent the first learning input image from including pixels with outliers. Furthermore, the image processing system 1 can prevent the learning from being affected by differences in hue between the first learning input image and the correct image (overlearning of hue conversion). Therefore, the image processing system 1 can improve the learning accuracy of the learning model 111.
[0037] The learning means 121 may convert the color tone of the first learning input image by histogram matching. In this case, the information processing device 100 generates a histogram of the gradation values of each pixel included in a predetermined sample image for each color (red, green, and blue) and stores it in the storage device 110. The sample image is set to a satellite image, aerial photograph, or the like that does not contain any outlier pixels and has an average color tone. The learning means 121 converts the gradation value of each pixel included in the first learning input image for each color (red, green, and blue) so that the histogram of the gradation values of each pixel included in the first learning input image matches the histogram of the gradation values of each pixel included in the sample image.
[0038] In this way, the learning means 121 converts the tone values of each pixel included in the first learning input image by histogram matching before inputting the first learning input image to the learning model 111, thereby converting the hue of the first learning input image. This allows the image processing system 1 to prevent the first learning input image from including pixels with outliers. Furthermore, the image processing system 1 can prevent the learning from being affected by differences in hue between the first learning input image and the correct image (overlearning of hue conversion). Therefore, the image processing system 1 can improve the learning accuracy of the learning model 111.
[0039] The process of step S103 may be omitted.
[0040] Next, the learning means 121 generates a learning model 111 using the first learning input image whose hue has been converted in step S103, the second learning input image generated in step S102, and the correct image acquired in step S101 (step S104). The learning model 111 is, for example, a model using a neural network, SwinIR (Image Restoration Using Swin Transformer), or the like. The learning model 111 may also be another model using RAISR (Rapid and Accurate Image Super-Resolution), or the like.
[0041] The learning means 121 trains the learning model 111 so that, when a first learning input image is input, the learning model 111 outputs a first learning output image that includes a predetermined area captured in the first learning input image and has a second resolution. The learning means 121 trains the learning model 111 so that the first learning output image approximates a correct answer image that includes the predetermined area included in the first learning input image. For example, the learning means 121 trains the learning model 111 so that the error between the first learning output image and a correct answer image that includes the predetermined area included in the first learning input image, i.e., the loss function, is minimized. The loss function is, for example, the sum of squared errors of the gradation values (R value, G value, and B value) of corresponding pixels, or the sum of absolute values of the differences. The learning means 121 changes the values of parameters such as weights in the learning model 111 using backpropagation or the like based on the error between the first learning output image and the correct answer image.
[0042] Furthermore, the learning means 121 trains the learning model 111 so that, when a second learning input image is input, the learning model 111 outputs a second learning output image that includes a predetermined region captured in a gold standard image that is the basis for the second learning input image and has a second resolution. The learning means 121 trains the learning model 111 so that the second learning output image approximates the gold standard image that is the basis for the second learning input image. For example, the learning means 121 trains the learning model 111 so that the error between the second learning output image and the gold standard image that is the basis for the second learning input image is minimized. The learning means 121 changes the values of parameters such as weights in the learning model 111 by error backpropagation or the like based on the error between the second learning output image and the gold standard image.
[0043] The learning means 121 executes optimization of parameters in the learning model 111 at the point in time when the learning means 121 inputs a first learning input image to the learning model 111 to obtain a first learning output image from the learning model 111 and inputs a second learning input image to the learning model 111 to obtain a second learning output image from the learning model 111. This allows the image processing system 1 to optimize the parameters of the learning model 111 in consideration of both the relationship between the first learning input image and the correct image and the relationship between the second learning input image and the correct image.
[0044] In this way, the learning model 111 is trained to output a first learning output image having a second resolution and including a predetermined area when a first learning input image in which the predetermined area is captured at a first resolution is input, and to output a second learning output image having a second resolution and including the predetermined area when a second learning input image in which the reference image in which the predetermined area is captured at the second resolution is input, resulting from converting the resolution of the reference image to the first resolution. As described above, the first learning input image and the reference image differ from each other in terms of type, imaging device, imaging timing, position, or imaging angle of the imaging device. Therefore, although the first learning input image and the reference image include the same area, the position, shape, and / or color of features in each image are likely to be slightly different. Therefore, the learning accuracy of a learning model trained using only the first learning input image and the reference image may be low.
[0045] On the other hand, the appearance (appearance) of objects included in a second learning input image generated by reducing the resolution of a reference image captured at the second resolution to the first resolution differs from that of an image captured at the first resolution from the beginning. Therefore, the learning accuracy of a learning model trained using only the second learning input image and the reference image may be low. However, the type, imaging device, imaging timing, and imaging device position or imaging angle of the second learning input image and the reference image are identical. Therefore, the position, shape, and / or color of features are identical in the second learning input image and the reference image. Therefore, the image processing system 1 can improve the learning accuracy of the learning model 111 by training the learning model 111 using both the first learning input image and the second learning input image.
[0046] Furthermore, the image processing system 1 trains the learning model 111 so as to minimize the error between each of the first learning output image and the second learning output image and the same (single) correct image in which a predetermined area included in the first learning output image and the second learning output image is captured at the second resolution. This allows the image processing system 1 to more efficiently improve the learning accuracy of the learning model 111. Note that the image processing system 1 may use different images as the correct image for the first learning input image and the correct image for the second learning input image.
[0047] The learning means 121 may exclude from the learning model 111, pixels included in the first learning output image whose loss relative to the corresponding pixel in the corresponding reference image satisfies a predetermined condition. The predetermined condition is that the loss is sufficiently large, i.e., that the difference in gradation value from the corresponding pixel in the reference image is sufficiently large. For example, a pixel satisfying the predetermined condition is a pixel whose loss relative to the corresponding pixel in the reference image is equal to or greater than a predetermined value. The predetermined value is set in advance to, for example, a value corresponding to a difference in gradation value (e.g., 20) that allows a person to visually distinguish the difference in the image. The predetermined value may be dynamically set to a predetermined rank (e.g., the top 10%) of losses relative to the corresponding pixel in the corresponding reference image in a predetermined number (e.g., 10) of first learning output images most recently learned, in descending order of the largest loss.
[0048] The pixels that satisfy the predetermined condition may be pixels included in a predetermined percentage of the first training output image in descending order of loss relative to the corresponding pixels in the corresponding ground truth image. The predetermined percentage is set, for example, to a percentage (e.g., the top 1%) that corresponds to pixels that can be considered outliers.
[0049] FIG. 3A is a schematic diagram illustrating feature changes. In FIG. 3A, image P11 represents a first learning input image in which a predetermined area is captured at a predetermined timing, and image P12 represents a correct image in which the predetermined area is captured at a later timing. In the example shown in FIG. 3A, a tree T is present in an area S that was vacant when image P11 was captured, when image P12 was captured. Similarly, a building may be newly constructed, destroyed, expanded, or renovated within the area included in each image between the time when the first learning input image was captured and the time when the correct image was captured. As such, if the first learning input image and the correct image are captured at different times, features included in each image may change. The image processing system 1 can exclude pixels in the first learning output image that have a large loss compared to the corresponding pixel in the correct image from the learning target, thereby excluding pixels in which the captured features have changed. Therefore, the image processing system 1 can improve the learning accuracy of the learning model 111.
[0050] FIG. 3B is a schematic diagram illustrating changes in the imaging angle. In FIG. 3B, image P21 represents a first learning input image in which a predetermined area is captured at a predetermined angle, and image P22 represents a correct image in which the predetermined area is captured at an angle different from the predetermined angle. In the example shown in FIG. 3B, image P21 includes building B captured from the upper left, while image P22 includes building B captured from the lower right. As shown in FIG. 3B, when the imaging angles of the first learning input image and the correct image differ, misalignment due to so-called tilting may occur, and one image may include a side of a building that is not captured in the other image. The image processing system 1 excludes pixels included in the first learning output image that have a large loss compared to the corresponding pixel in the correct image from the learning target, thereby excluding pixels corresponding to portions included only in one image from the learning target. Therefore, the image processing system 1 can improve the learning accuracy of the learning model 111.
[0051] Next, the learning means 121 stores the generated learning model 111 in the storage device 110 (step S105), and ends the learning process.
[0052] FIG. 4 is a flowchart showing an example of the flow of the output process executed by the information processing device 100.
[0053] An example of the operation of the output process of the information processing device 100 will be described below with reference to the flowchart shown in Fig. 4. The flow of the operation described below is executed mainly by the processing circuitry 120 in cooperation with each element of the information processing device 100 based on a program stored in advance in the storage device 110.
[0054] First, the acquisition means 122 acquires an input image in which a feature is captured at a first resolution (step S201). The acquisition means 122 acquires the input image by receiving it from the server device 200 via the communication device 103. The input image is an aerial photograph, satellite image (or an orthoimage based on the aerial photograph or satellite image) or the like of the earth's surface, which is the target area for processing such as interpreting building changes, creating maps, and creating city models, and includes a feature captured from the sky.
[0055] Next, the output control means 123 inputs the input image acquired by the acquisition means 122 to the learning model 111, and acquires an output image output from the learning model 111 when the input image is input to the learning model 111 (step S202). The output image is an image that includes the features included in the input image and has a second resolution.
[0056] Next, the output control means 123 outputs information about the acquired output image by transmitting it to an external device via the communication device 103 or by displaying it on the display device 102 (step S203), and ends the output process. The information about the output image is, for example, the output image itself. The information about the output image may also be an image obtained by performing a predetermined correction process on the output image. The information about the output image may also be information indicating features detected from the output image, objects attached to the features (for example, solar panels attached to the roof of a house), or the positions of the features or the objects. In this case, the output control means 123 uses known image processing techniques to detect features included in the output image, objects attached to the features, or the positions of the features or the objects from the output image.
[0057] As described above, the image processing system 1 trains the learning model 111 using a first learning input image in which a predetermined area is captured at a first resolution and a second learning input image in which a correct answer image in which a predetermined area is captured at a second resolution is converted to have the first resolution. Then, the image processing system 1 generates an output image having the second resolution from the input image captured at the first resolution using the learning model 111. This enables the image processing system 1 to obtain an output image in which the input image has been super-resolved.
[0058] Furthermore, the image processing system 1 can generate the learning model 111 by using different types of images as the first learning input images and the correct answer images. Satellite images are captured more frequently and at lower cost than aerial photographs. For example, the image processing system 1 can generate the learning model 111 efficiently and at low cost by using a combination of satellite images and aerial photographs as the first learning input images and the correct answer images.
[0059] Furthermore, by obtaining an output image from an input image having a higher resolution than the resolution of the input image, the image processing system 1 can more accurately detect features included in the input image and / or objects attached to those features.
[0060] With a typical camera or subject, the difference between a low-resolution image and a reduced-size image is not significant. However, with aerial photographs and satellite images, the difference between the low-resolution image and the reduced-size image is greater than with a typical camera due to the influence of disturbances such as the shooting equipment and timing, as well as regional characteristics and atmospheric conditions. When a super-resolution model is trained using only pairs of reduced and high-resolution images, if there is a large difference in image distribution between the pairs of images used for training, repeated training will result in large loss and an appropriate trained model will not be generated. Therefore, when training a super-resolution model, it is necessary for the image distribution between the pairs of images used for training to be similar. Furthermore, when a super-resolution model is trained using only pairs of low-resolution and high-resolution images, the image distribution between the pairs of images used for training will differ depending on disturbances (such as shooting timing and shooting angle), so loss will not be sufficiently reduced and an appropriate trained model will not be generated. The image processing system 1 can solve the above problem by generating a trained model using both pairs of low-resolution and high-resolution images and pairs of reduced and high-resolution images. The image processing system 1 can improve the interpretation accuracy by using the low-resolution image that is actually to be enlarged during learning, while suppressing an increase in loss by using reduced images and high-resolution images in areas where the image distributions of the low-resolution image and the high-resolution image are far apart. Therefore, the image processing system 1 can appropriately generate a learning model.
[0061] Although preferred embodiments have been described above, the embodiments are not limited to these. For example, the correct answer image of the first learning input image used to train the learning model 111 and the correct answer image of the second learning input image corresponding to the first learning input image may not be the same image, but may be mutually different images.
[0062] Alternatively, instead of the information processing device 100, the server device 200 may have the learning means 121 and execute the learning process of FIG.
[0063] 4, the output control means 123 transmits an input image to the server device 200 via the communication device 103. The server device 200 inputs the input image received from the information processing device 100 to the learning model 111, and transmits an output image output from the learning model 111 to the information processing device 100. The output control means 123 acquires the output image by receiving it from the server device 200 via the communication device 103.
[0064] It should be understood by those skilled in the art that various changes, substitutions, and alterations can be made to the present invention without departing from the spirit and scope of the present invention. For example, the above-described embodiments and modifications may be implemented in appropriate combinations within the scope of the present invention. [Explanation of symbols]
[0065] 1 Image processing system, 100 Information processing device, 102 Display device, 103 Communication device, 111 Learning model, 121 Learning means, 122 Acquisition means, 123 Output control means
Claims
1. an acquisition means for acquiring an input image of a feature captured at a first resolution; an output means for outputting information about an output image that includes the feature and has a second resolution higher than the first resolution, the output image being output from the learning model when the input image is input to the learning model; the learning model is trained to output a first learning output image that includes the predetermined area and has the second resolution when a first learning input image in which a predetermined area is captured at the first resolution is input, and to output a second learning output image that includes the predetermined area and has the second resolution when a second learning input image in which the image in which the predetermined area is captured at the second resolution is converted so that its resolution becomes the first resolution, the learning model is trained so that the first learning output image output when the first learning input image is input is approximate to a correct image that includes the predetermined region and has the second resolution, the first learning input image and the correct answer image are images of different types, images captured by different imaging devices, or images captured from different positions at different angles, the learning model is trained so that the second learning output image output when the second learning input image is input is approximate to a correct image that includes the predetermined region and has the second resolution, The second learning input image is an image to which defocus or imaging blur is added to the correct answer image. An image processing system comprising:
2. 2. The image processing system of claim 1, wherein the learning model is trained so as to minimize an error between each of the first learning output image and the second learning output image and an identical correct image in which the specified area is captured at the second resolution.
3. The image processing system according to claim 1 , wherein the first training input image is subjected to a color conversion before being input to the training model.
4. 4. The image processing system according to claim 3, wherein the first learning input image is subjected to hue conversion so that the gradation values of each pixel form a standard normal distribution, or by converting the gradation values of each pixel by histogram matching.
5. The image processing system according to claim 1 , wherein pixels included in the first learning output image, whose loss relative to a corresponding pixel in a correct answer image satisfies a predetermined condition, are excluded from the learning target.
6. 6. The image processing system according to claim 5, wherein the pixels that satisfy the predetermined condition are pixels whose loss is equal to or greater than a predetermined value or pixels that are included in a predetermined proportion in descending order of the loss.
7. The image processing system according to claim 1 , wherein the second learning input image is generated from an image of the predetermined area captured at the second resolution using a model that reproduces a deterioration process.
8. acquiring an input image of a feature at a first resolution; outputting information about an output image that includes the feature and has a second resolution higher than the first resolution, the output image being output from the learning model when the input image is input to the learning model; the learning model is trained to output a first learning output image that includes the predetermined area and has the second resolution when a first learning input image in which a predetermined area is captured at the first resolution is input, and to output a second learning output image that includes the predetermined area and has the second resolution when a second learning input image in which the image in which the predetermined area is captured at the second resolution is converted so that its resolution becomes the first resolution, the learning model is trained so that the first learning output image output when the first learning input image is input is approximate to a correct image that includes the predetermined region and has the second resolution, the first learning input image and the correct answer image are images of different types, images captured by different imaging devices, or images captured from different positions at different angles, the learning model is trained so that the second learning output image output when the second learning input image is input is approximate to a correct image that includes the predetermined region and has the second resolution, The second learning input image is an image to which defocus or imaging blur is added to the correct answer image. An image processing method comprising:
9. acquiring an input image of a feature at a first resolution; causing a computer to output information about an output image output from the learning model when the input image is input to the learning model, the output image including the feature and having a second resolution higher than the first resolution; the learning model is trained to output a first learning output image that includes the predetermined area and has the second resolution when a first learning input image in which a predetermined area is captured at the first resolution is input, and to output a second learning output image that includes the predetermined area and has the second resolution when a second learning input image in which the image in which the predetermined area is captured at the second resolution is converted so that its resolution becomes the first resolution, the learning model is trained so that the first learning output image output when the first learning input image is input is approximate to a correct image that includes the predetermined region and has the second resolution, the first learning input image and the correct answer image are images of different types, images captured by different imaging devices, or images captured from different positions at different angles, the learning model is trained so that the second learning output image output when the second learning input image is input is approximate to a correct image that includes the predetermined region and has the second resolution, The second learning input image is an image to which defocus or imaging blur is added to the correct answer image. A computer program characterized by:
Citation Information
Patent Citations
Binocular deep learning method based on adaptive single-peak stereo matching cost filtering
CN111709977A
Transformers-based low-resolution image super-resolution method and system
CN112862690A
Breast tumor classification method and device based on ultrasonic image two-stage deep learning
CN113033667A
MRI image and CT image conversion method and terminal based on deep learning
CN114266929A
Number-of-targets estimation device, number-of-targets estimation method, and program
JP2018072938A