Image processing to improve the resolution of images taken of specimens
The image processing device enhances image resolution using a super-resolution network trained on opposing cross-sections of specimens, addressing the limitations of optical microscopes and electron microscopes, enabling detailed observation and diagnosis with improved image quality.
Patent Information
- Application Number
- JP2024561545
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-11-30
- Filing Date
- 2023-11-29
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2043-11-29
AI Technical Summary
Existing optical microscopes have a resolution limit due to the diffraction of light, making it difficult to observe specimens in detail, while electron microscopes require more effort, time, and cost to prepare and operate, and specialized equipment like dermatoscopes are challenging to use due to environment, cost, and effort, necessitating a method to improve the resolution of input images to exceed the diffraction limit.
An image processing device that uses a trained super-resolution network to estimate an output image with higher resolution by inputting an image captured with a first wave and estimating it as if photographed with a second wave, utilizing training data from opposing cross-sections of specimens captured with different waves, and aligning and correcting images to accommodate deformation.
Enables the estimation of high-resolution output images from lower-resolution input images, effectively overcoming the diffraction limit and facilitating detailed observation and diagnosis without the need for electron microscopes or specialized equipment.
Smart Images

Figure 0007722757000001 
Figure 0007722757000002 
Figure 0007722757000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing system, an image processing method, a program, and an information recording medium that obtain an output image by performing estimation to improve the resolution of an input image of a sample. [Background technology]
[0002] A super-resolution technology has been proposed that estimates an output image with improved resolution from an input image. Super-resolution estimation is based on the characteristics of the object captured in the input image, and a neural network trained by deep learning or the like can be used. Training a neural network often involves preparing a photograph of the subject expected to be processed, applying noise, mosaicking, blurring, or reducing the resolution to the photograph, which is then used as an input image for training. This photograph is then used as the output (correct) image for training (see, for example, Patent Document 1 and Non-Patent Document 2).
[0003] Meanwhile, various proposals have been made for image registration (image alignment) techniques for matching feature points of two images of the same subject and obtaining a mapping representing the transformation from one image to the other, or the inverse transformation (see, for example, Non-Patent Document 1).
[0004] Now, in the field of pathological photography using optical microscopes, for example, The obtained specimen is immersed in formalin water or similar to fix the tissue, The water in the tissue is replaced with paraffin and embedded. Thinly slice the tissue using a microtome or other instrument to obtain sections. Each section was attached to a glass plate. The paraffin is deparaffinized by melting it back into water. Staining with various dyes A tissue specimen is prepared according to the procedure.
[0005] By observing and photographing this tissue specimen under an optical microscope, a digital pathological image can be obtained.
[0006] However, optical microscopes have a resolution limit due to the diffraction of light. This limit, called the diffraction limit, is smaller than a biological cell (1 μm to 100 μm) and larger than viruses (100 nm), proteins (10 nm), and less complex molecules (1 nm) for typical optical systems.
[0007] In order to obtain a resolution exceeding the diffraction limit, it is conceivable to use an electron microscope. In the field of electron microscopy, for example, The obtained specimen is immersed in an aldehyde-based fixative to pre-fix the tissue. After washing with rinse solution, Post-fixation was performed with osmium solution. Dehydrate using ethanol, Embedding in epoxy resin, Ultra-thin slices were obtained using an ultramicrotome, diamond knife, etc. Perform electron staining A tissue specimen is prepared according to the procedure.
[0008] By observing and photographing this tissue specimen under an electron microscope, a digital electron micrograph is obtained.
[0009] Because electron microscopes use electron beams rather than light for observation, their diffraction limit exceeds that of optical microscopes and is generally considered to be on the order of a few nanometers.
[0010] In Patent Document 2, machine learning is performed based on images captured using an optical microscope (first microscope) and an electron microscope (second microscope). The image captured by the first microscope is a cell nucleus A1 in an iPS cell (first region), and the image captured by the electron microscope is a cell nucleus A2 in a brain tissue cell (second region), and different specimens are observed for different regions.
[0011] In this way, when the waves used for imaging are different between visible light and electron beams, the wavelengths of the waves also differ, and generally, the shorter the wavelength, the higher the spatial resolution.
[0012] In addition, in the examination of skin tumors or cancers, dermatoscopic diagnosis is performed, in which the affected area is observed and photographed visually or through a magnifying glass using normal light, and then, if necessary, observed and photographed using polarized light. An observation instrument called a dermatoscope is capable of observation and photography using both normal light and polarized light. Furthermore, for dermatoscopes, a technology has been proposed that switches the wavelength of the irradiated light to a specific wavelength band, as disclosed in Patent Document 3, for example.
[0013] Here, normal light (unpolarized waves) makes it possible to observe and photograph the surface layer of the skin. Polarized light (polarized waves) makes it possible to observe and photograph the structure of layers deeper than the surface of the skin. Furthermore, by using light of different wavelength bands, it becomes possible to observe and photograph tissues with specific structures that correspond to the wavelength bands in an emphasized manner.
[0014] Therefore, by using different waves, it is possible to improve the resolution for areas that are far from the surface of the object of observation in the depth direction, and the resolution based on the structure of the tissue. [Prior art documents] [Patent documents]
[0015] [Patent Document 1] Japanese Patent Publication No. 2022-056769 [Patent Document 2] Patent Publication No. 2021-18582 [Patent Document 3] Special Publication No. 2022-527642 [Non-patent literature]
[0016] [Non-Patent Document 1] Jeremy Joslove and Emna Kamoun, "Image Registration: From SIFT to Deep Learning", Sicara's blog, [online], https: / / medium.com / sicara / image-registration-sift-deep-learning-3c794d794b7a (overview and introduction), https: / / www.sicara.fr / blog-technique / 2019-07-16-image-registration-deep-learning (main text), July 16, 2019 [Non-patent document 2] Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J. Fleet, and Mohammad Norouzi, "Image Super-Resolution via Iterative Refinement", Google Research, Brain Team [online], https: / / doi.org / 10.48550 / arXiv.2104.07636, April 15, 2021 Summary of the Invention [Problem to be solved by the invention]
[0017] From the viewpoint of the diffraction limit, i.e., spatial resolution, the performance of electron microscopes exceeds that of optical microscopes. Therefore, there is a demand for observing a pathological photograph taken with an optical microscope, and then observing the site photographed in the pathological photograph in more detail using an electron microscope photograph, thereby making various diagnoses and judgments.
[0018] However, specimens for electron microscopes often require more effort, time, and cost to prepare than specimens for light microscopes. The use of electron microscopes themselves can also be problematic in terms of effort, time, and cost compared to light microscopes.
[0019] Furthermore, because the procedures for preparing specimens differ between optical and electron microscopes, it is difficult to obtain appropriate electron microscope photographs even when a specimen intended for an optical microscope is observed and photographed using an electron microscope.
[0020] In addition, while it may be easy to photograph the surface of the affected area using a digital camera in today's working environment, photographing using specialized equipment such as a dermatoscope is often difficult due to the working environment, cost, and effort involved.
[0021] Therefore, there is a strong demand for an input image (for example, a pathology image based on a photograph of a pathology specimen taken of a specimen for an optical microscope, or an epidermal image of the skin surface of an affected area taken with a digital camera) to be estimated to improve the resolution of the input image, preferably to an extent that exceeds the diffraction limit at the time of taking the input image, thereby obtaining an output image (for example, an estimated image that represents the results of estimating elements of how the tissue corresponding to the specimen appears as if it were photographed with an electron microscope, or how the affected area appears as if it were photographed with polarized light or light in a specific wavelength band in dermatoscopy), and applying this to various diagnoses and judgments.
[0022] The present invention is devised to solve the above-mentioned problems, and aims to provide an image processing device, an image processing system, an image processing method, a program, and an information recording medium that perform estimation to improve the resolution of an input image of a specimen and obtain an output image. [Means for solving the problem]
[0023] The image processing device according to the present invention comprises: When an input image of a target specimen is inputted by a first wave having a first resolution, By providing the input image to a trained super-resolution network, an output image that should be obtained by photographing the target specimen with a second wave having a second resolution is estimated. Here, the second resolution is higher than the first resolution.
[0024] In addition, in the image processing device of the present invention, The input image is a cross-section of the target specimen. It can be configured as follows.
[0025] In addition, in the image processing device of the present invention, The super-resolution network comprises: a training input based on a first training image obtained by photographing a first training cross section that appears by cutting a training specimen with the first wave; a training output based on a second training image obtained by photographing a second training cross section that appears opposite to the first training cross section with the second wave; It is learned from training data including It can be configured as follows. [Effects of the Invention]
[0026] According to the present invention, it is possible to provide an image processing device, an image processing system, an image processing method, a program, and an information recording medium that perform estimation to improve the resolution of an input image of a specimen and obtain an output image. [Brief explanation of the drawings]
[0027] [Figure 1] 1 is an explanatory diagram showing a schematic configuration of an image processing apparatus according to an embodiment of the present invention; [Figure 2] FIG. 1 is an explanatory diagram illustrating an example of a super-resolution network used by an image processing device according to an embodiment of the present invention. [Figure 3] 10 is a flowchart showing the control flow of a training process for training a super-resolution network executed in an image processing device according to an embodiment of the present invention. [Figure 4] 4 is a flowchart showing a control flow of image processing executed by the image processing apparatus according to the embodiment of the present invention. [Figure 5] 1 is an explanatory diagram showing a schematic configuration of an image processing system according to an embodiment of the present invention; [Figure 6] FIG. 2 is an explanatory diagram showing an example of a display on a screen of a terminal computer according to an embodiment of the present invention. [Figure 7]FIG. 2 is an explanatory diagram showing an example of a display on a screen of a terminal computer according to an embodiment of the present invention. [Figure 8A] This is a photograph, which is a substitute for a drawing, showing a photographed pathological image, an electron microscope image estimated from the pathological image, and the photographed electron microscope image arranged vertically in grayscale. [Figure 8B] This is a photograph in place of a drawing, showing a photographed pathological image, an electron microscope image estimated from the pathological image, and the photographed electron microscope image arranged vertically in monochrome binary. [Figure 9A] This is a photograph, which is a substitute for a drawing, showing a photographed pathological image, an electron microscope image estimated from the pathological image, and the photographed electron microscope image arranged vertically in grayscale. [Figure 9B] This is a photograph in place of a drawing, showing a photographed pathological image, an electron microscope image estimated from the pathological image, and the photographed electron microscope image arranged vertically in monochrome binary. [Figure 10A] This is a photograph, which is a substitute for a drawing, showing a photographed pathological image, an electron microscope image estimated from the pathological image, and the photographed electron microscope image arranged vertically in grayscale. [Figure 10B] This is a photograph in place of a drawing, showing a photographed pathological image, an electron microscope image estimated from the pathological image, and the photographed electron microscope image arranged vertically in monochrome binary. [Figure 11] 1 is an explanatory diagram showing a schematic configuration of an image processing apparatus according to an embodiment of the present invention; [Figure 12] FIG. 1 is an explanatory diagram showing the configuration of an Encoder in a VAE, which is an example of a generation network in an image processing device according to an embodiment of the present invention. [Figure 13] FIG. 2 is an explanatory diagram showing the configuration of a decoder in a VAE, which is an example of a generation network in an image processing device according to an embodiment of the present invention. [Figure 14] FIG. 10 is an explanatory diagram showing the configuration of an Encoder in a VAE, which is another example of a generation network in an image processing device according to an embodiment of the present invention. [Figure 15] FIG. 10 is an explanatory diagram showing the configuration of a decoder in a VAE, which is another example of a generation network in an image processing device according to an embodiment of the present invention. [Figure 16]This is a photograph in place of a drawing, which shows a captured pathological image, an electron microscope image estimated from the pathological image, and a generated reference image side by side in grayscale. [Figure 17] This is a photograph in place of a drawing, showing the captured pathological image, the electron microscope image estimated from the pathological image, and the generated reference image side by side in monochrome binary. DETAILED DESCRIPTION OF THE INVENTION
[0028] The following describes embodiments of the present invention. Note that these embodiments are for illustrative purposes only and do not limit the scope of the present invention. Therefore, those skilled in the art can adopt embodiments in which each or all of the elements of the present embodiments are replaced with equivalents. Furthermore, elements described in each example can be omitted as appropriate depending on the application. In this way, all embodiments constructed in accordance with the principles of the present invention are included in the scope of the present invention.
[0029] (composition) The image processing device according to this embodiment is typically realized by a computer executing a program. The computer is connected to various output devices and input devices, and transmits and receives information to and from these devices.
[0030] A program executed by a computer can be distributed or sold by a server to which the computer is connected for communication, or it can be recorded on a non-transitory information recording medium such as a CD-ROM (Compact Disk Read Only Memory), flash memory, or EEPROM (Electrically Erasable Programmable ROM), and then the information recording medium can be distributed, sold, etc.
[0031] The program is installed on a non-transitory information recording medium such as a hard disk, solid-state drive, flash memory, EEPROM, etc., of the computer. The image processing device of this embodiment is then realized by the computer. Generally, the computer's central processing unit (CPU) reads the program from the information recording medium into random access memory (RAM) under the control of the computer's operating system (OS), and then interprets and executes the code contained in the program. However, in an architecture in which the information recording medium can be mapped within a memory space accessible by the CPU, explicit loading of the program into RAM may not be necessary. Various pieces of information required during program execution can be temporarily stored in RAM.
[0032] Furthermore, as mentioned above, it is desirable for computers to be equipped with a GPU (Graphics Processing Unit) to perform various image processing calculations at high speed. By using a GPU and libraries such as TensorFlow, it becomes possible to utilize the learning and classification functions in various artificial intelligence processes under the control of the CPU.
[0033] It is also possible to configure the image processing device of this embodiment using a dedicated electronic circuit, rather than using a general-purpose computer. In such an embodiment, an electronic circuit that satisfies the specifications defined in the program is configured using an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the electronic circuit functions as a dedicated device that performs the functions defined in the program, thereby realizing the image processing device of this embodiment.
[0034] (Image processing device) For ease of understanding, the following description will be given assuming that the image processing device is realized by a computer executing a program. Fig. 1 is an explanatory diagram showing the general configuration of an image processing device according to an embodiment of the present invention.
[0035] As shown in the figure, an image processing device 101 according to this embodiment includes an input unit 102 and a super-resolution unit 103.
[0036] Here, an input image of a target specimen captured with a first wave having a first resolution is input to the input unit 102.
[0037] On the other hand, the super-resolution unit 103 provides an input image to the trained super-resolution network 104, causing it to estimate an output image that should be obtained by photographing the target specimen with a second wave having a second resolution.
[0038] Any neural network can be used as the super-resolution network 104. For example, a neural network such as that shown in Fig. 2 may be used, or a complex neural network such as that shown in Non-Patent Document 2 may also be used.
[0039] Here, the second waveband of the second wave can be configured to be shorter than the first waveband of the first wave.
[0040] The first wave may be light and the second wave may be an electron beam.
[0041] In these configurations, an output image to be captured by an electron microscope is estimated from an input image based on a photograph captured by an optical microscope.
[0042] The first wave may be unpolarized light and the second wave may be polarized light.
[0043] In these configurations, the output image to be captured by the dermatoscope is estimated from an input image based on a photograph taken under normal light.
[0044] The super-resolution network 104 is trained using multiple sets of training data.
[0045] Here, each training data includes a training input (input data) and a training output (correct answer data). The training data is prepared by preparing images of the same object photographed in the past using two types of waves.
[0046] For example, when making inferences about dermatoscopes, the training inputs are photographs taken under normal light, and the training outputs are photographs taken under polarized light.
[0047] As mentioned above, specimens and samples for observation with an optical microscope and those for observation with an electron microscope are treated differently, so it can be difficult to photograph the same object using two different types of waves.
[0048] However, when observing a specimen, the state of its cross section is often examined. Therefore, in this embodiment, it is possible to utilize the fact that when a specimen is cut, two opposing cross sections are obtained.
[0049] In this embodiment, input data based on a first training image obtained by photographing a first training cross section that appears by cutting a training specimen with a first wave can be used as training input, and correct answer data based on a second training image obtained by photographing a second training cross section that appears opposite the first training cross section with a second wave can be used as training output.
[0050] Note that when the training specimen is cut into two samples, the cross section itself may be significantly deformed due to the cutting. Therefore, after fixing the specimen, several thin slices may be cut from the initial cross section, and the inner layer (the layer farthest from the initial cross section) may be photographed and used as the training input and output.
[0051] In this case, the training input and training output images are images of sections spaced apart from the cross section of the specimen. Therefore, the physical locations of the regions imaged in the training input and training output images within the specimen are not the opposing cross sections themselves, but rather sections located nearby. However, if the distance between them is sufficiently small, these sections can be treated as essentially opposing cross sections.
[0052] (Basic operation of image processing device) The following describes the basic operation of the image processing device 101 that uses the trained super-resolution network 104. Fig. 4 is a flowchart showing the flow of control of image processing executed by the image processing device according to the embodiment of the present invention.
[0053] First, in the image processing device 101, the input unit 102 receives an input of an input image of a target object captured with a first wave having a first resolution (step S401).
[0054] Then, the super-resolution unit 103 provides the received input image to the super-resolution network 104 (step S402).
[0055] The super-resolution network 104 estimates, based on the input image, an image that should be obtained by imaging the target object with the second wave having the second resolution (step S403).
[0056] Then, the image processing device 101 outputs the estimated image as an output image (step S404), and ends this processing.
[0057] If the basic operation of the image processing apparatus 101 is provided as a service by a server computer, after step S404, control returns to step S401 and the above processing is repeated.
[0058] (Training a super-resolution network) As described above, the training input and training output used to train the super-resolution network are data obtained by photographing opposing cross sections of a single specimen using different waves, and when the specimen is cut and sampled, deformation of the specimen may occur.
[0059] Therefore, when training the super-resolution network 104, correction and alignment of the training input and training output are required to accommodate deformation of the sample.
[0060] In this embodiment, the image processing device 101 trains the super-resolution network 104 through the following processing. Fig. 3 is a flowchart showing the control flow of the training processing for training the super-resolution network executed by the image processing device according to the embodiment of the present invention. The following description will be made with reference to this figure.
[0061] For ease of understanding, the following describes a mode in which the image processing device 101 trains the super-resolution network 104. However, this training process may be executed by the image processing device 101, or may be executed by a device other than the image processing device 101, and the super-resolution network 104 learned by the training process may then be made available to the image processing device 101.
[0062] First, the image processing device 101 receives a first training image obtained by photographing a first training section using a first wave motion and a second training image obtained by photographing a second training section opposite the first training section using a second wave motion (step S301).
[0063] Then, the image processing device 101 obtains a mapping for aligning the first training image with the second training image (step S302).
[0064] The first method for obtaining the mapping is as follows. That is, since the first training image and the second training image have different resolutions, a scale mapping that enlarges or reduces the resolution of both images to match is prepared and applied to the first training image. The scaling ratio of the scale mapping can be determined based on the shooting conditions when the first training image and the second training image were captured, i.e., the "number of pixels in the width and height of the image" determined by the lens magnification and image resolution, and the "actual width and height of the subject corresponding to one pixel in the image."
[0065] Then, a transform map is calculated to align the first and second training images after scaling, i.e., the two images with the same resolution.
[0066] Finally, the desired mapping is obtained by combining the scale and transform mappings.
[0067] As described above, the scale mapping can be uniquely determined by determining the enlargement / reduction ratio based on the conditions under which the first training image and the second training image were captured.
[0068] On the other hand, the transform map can be obtained by applying the technique disclosed in Non-Patent Document 1 as well as various other techniques for aligning two images with the same resolution.
[0069] The second method for determining the mapping is as follows. That is, the image registration techniques disclosed in Non-Patent Document 1, such as the combination of feature extraction and convolutional neural networks, and reinforcement learning and deep learning that determine the direct sum of multiple homography mappings or diffeomorphism mapping, can be applied even when the resolutions of the two images do not match. In such cases, the desired mapping can be determined by combining the scale mapping and the transform mapping. Therefore, in such cases, it is not necessary to determine the mapping in two separate steps: the scale mapping and the transform mapping.
[0070] Next, the image processing device 101 corrects the first training image using the obtained mapping (step S303).
[0071] Then, the image processing device 101 divides the corrected first training image into a plurality of regions (step S304), and repeats the following process for each of the plurality of regions (step S305). Note that, among the plurality of regions, the region currently being processed in the repetition will be referred to as the "first region" hereinafter.
[0072] First, the image processing device 101 cuts out the first region of the corrected first training image to obtain a first partial image (step S306).
[0073] Then, the image processing device 101 determines candidate positions of an area (hereinafter referred to as a "second area") where the first area is projected in the second training image by the above mapping (step S307).
[0074] Next, the image processing device 101 scans a predetermined scanning range around the candidate position using a window of the same shape as the second region, and obtains an image within the window at each position (step S308). For example, if the predetermined scanning range is a range centered on the candidate position and shifted up, down, left, and right by n pixels, the number of images within the window obtained by scanning, including the center, is (2×n+1) 2 If not included, it becomes (2×n) 2 Becomes an individual.
[0075] Furthermore, the image processing device 101 calculates the similarity between each of the images within the multiple windows obtained by scanning with the windows and the first partial image, and selects the image with the greatest similarity as the second partial image (step S309).
[0076] Then, the position of the second region is identified as the position of the window into which the second partial image is cut out (step S310).
[0077] Next, the image processing device 101 generates training data to be given to the super-resolution network 104, in which the first partial image is used as a training input and the second partial image is used as a training output (step S311).
[0078] After performing this repetitive process for each of the divided regions (step S312), the image processing device 101 removes training data whose similarity is an outlier based on the distribution of similarities between training inputs and training outputs in the generated training data (step S313), and then provides the remaining training data to the super-resolution network 104 to proceed with learning (step S314), thereby terminating this process.
[0079] If multiple pairs of first training images and second training images are prepared, the above process may be repeated for each pair, or training data may be generated for all pairs and then provided to the super-resolution network 104 all at once to train the super-resolution network 104.
[0080] In addition, in step S307, a candidate position for the second region relative to the first region is determined. At the beginning of the repetition, the value of n may be set to a value that covers a wide range of the second training image or the entire image, and after the repetition has progressed to a certain extent, the value of n may be reduced after determining a candidate position for the current first region based on the correspondence between the positions of the first region and second region identified in the past.
[0081] (Example of the first method for finding a mapping) Below, we will explain in more detail an example of a first technique for finding a mapping that aligns a first training image with a second training image through two steps: a step for finding a scale mapping and a step for finding a transform mapping.
[0082] In this method, as described above, the resolutions of the first training images and the second training images are matched by scale mapping, and then the image processing device 101 performs feature point detection and feature point matching. Any method can be applied to feature point detection and feature point matching, from classical methods such as SIFT and AKAZE to various methods such as those disclosed in Non-Patent Document 1.
[0083] Then, (a) a first feature point detected in a first training image; (b) second feature points detected in the second training image, the second feature points corresponding to the first feature points; A plurality of pairs consisting of the above are extracted.
[0084] The target mapping is one that minimizes the difference between the projections of the first feature points of each pair obtained here and the second feature points of each pair. Therefore, the target mapping can be obtained by solving the minimization problem.
[0085] Therefore, the image processing device divides the first training image into a plurality of polygons whose vertices are the feature points detected by the feature point detection. In the simplest case, triangles can be used as polygons, and Delaunay division can be used as division.
[0086] Next, consider a plurality of candidate mappings each having a domain of a plurality of polygons.
[0087] Each candidate mapping of the plurality of candidate mappings is (a) projecting a first polygon that is the domain of each candidate mapping onto a second polygon in a second training image; (b) A mapping that projects the vertices of a first polygon onto the vertices of a second polygon that correspond to the vertices of the first polygon through feature point matching, and can be defined most simply by inversion, translation, rotation, scaling, shearing, trapezoidal transformation, or a combination of these.
[0088] Here, inversion, translation, rotation, scaling, shear, and combinations thereof can be expressed as affine transformations, and trapezoidal transformations and combinations thereof can be expressed as projective transformations (homography transformations). Any of these transformations can be specified by a transformation matrix.
[0089] In an ideal case where the samples for the first training image and the second training image are both undeformed, the transformation matrices representing each candidate mapping obtained here should be identical; however, in reality, deformations and other influences occur when the samples are created.
[0090] Therefore, clustering is performed to remove inappropriate candidates from among the multiple candidate mappings. As mentioned above, each candidate mapping is expressed by a transformation matrix, so the difference between the transformation matrices is used as the distance to perform clustering and detect outliers. For example, using k-nearest neighbor (k-NN) or Local Outlier Factor (LOF), clusters are distinguished into clusters to which outliers belong, i.e., minority clusters, and clusters to which major values belong, i.e., majority clusters.
[0091] Candidate images belonging to the minority cluster, that is, minority maps, are considered to have an incorrect matching between the domain area and the range area, or to have significant deformations in the domain area or range area.
[0092] Therefore, the majority mappings belonging to the majority cluster are identified, and the direct sum of the majority mappings is adopted as the target mapping.
[0093] After obtaining the mapping in this way, if the first partial image is not included in the domain of the mapping, or if the second partial image is not included in the range of the mapping, then this set will not be adopted as training data.
[0094] By performing such processing, it is possible to obtain appropriate training data that has been freed from the influence of incorrect feature point matching, partial deformation of the sample, and the like.
[0095] (An example of the second method for finding the mapping) As mentioned above, the second method for finding a mapping uses a combination of feature extraction and a convolutional neural network, and image registration techniques such as reinforcement learning and deep learning that find multiple homography mappings and diffeomorphisms. These techniques can be applied even when the resolutions of the two images do not match, but the desired mapping can also be obtained through a two-step process in which the resolutions are matched using the scale mapping described above, and then registration is performed to find a transform mapping.
[0096] In the case of medical image alignment as in this embodiment, the desired mapping cannot be simply described by a single homography matrix, so it is necessary to find the direct sum of multiple homography mappings or a diffeomorphism represented by a displacement vector field.
[0097] Such techniques include a robust registration technique using agent-based action learning by Julian Krebs et al., the DIRNET technique by Bob D. de Vos et al., and the Quicksilver technique by Xiao Yang et al., and these techniques can be applied to this embodiment.
[0098] (Image Processing System) The training of the super-resolution network 104 and the estimation of images by the super-resolution network 104 can be performed on a server computer in which various medical images based on electronic medical records, etc. are collected. In other words, the server computer can function as the image processing device 101 described above.
[0099] In such an embodiment, an image processing system can be used in which a terminal consisting of a computer provided in a hospital or the like serves as an input / output interface for actual diagnosis. FIG. 5 is an explanatory diagram showing a schematic configuration of an image processing system according to an embodiment of the present invention. FIG. 6 is an explanatory diagram showing an example of a display on a screen of a terminal computer according to an embodiment of the present invention. FIG. 7 is an explanatory diagram showing an example of a display on a screen of a terminal computer according to an embodiment of the present invention. The following description will be made with reference to these figures.
[0100] The image processing system 201 includes a server computer 202 that implements the image processing device 101, a terminal computer 203, and a computer communication network 204 that connects the two so that they can communicate with each other.
[0101] Here, the terminal computer 203 comprises a first display unit 211 , a reception unit 212 , a transmission / reception unit 213 , and a second display unit 214 .
[0102] The first display unit 211 displays, in a first area 411 on the screen 401, a group of images of a plurality of cross sections that appear when the target specimen is cut into layers.
[0103] In this drawing, each image 412 in the image group is parallel projected from an oblique angle and drawn side by side. The original shape of the edges of each image 412 is rectangular or square, and when parallel projected from an oblique angle, it becomes a parallelogram, but the edges are not shown in this drawing.
[0104] That is, the first display unit 211 arranges the multiple images included in the image group on multiple parallel planes set in a virtual three-dimensional space in the order in which the multiple cross sections for the multiple images are arranged in the target specimen, and depicts the appearance of the virtual three-dimensional space in the first area 411 by parallel (oblique) projection or perspective projection, thereby creating a pseudo-three-dimensional feeling.
[0105] On the other hand, the receiving unit 212 receives an image selection instruction to select one of the images from the displayed image group.
[0106] An image selection instruction is given by selecting a desired image with a mouse, keyboard, etc. In this figure, the edge of image 412 (which, as mentioned above, is originally a square or rectangle, but has been transformed into a parallelogram due to perspective viewing) is highlighted with dotted line 413 to indicate that the image has been selected.
[0107] In first area 411, by scrolling using scroll bar 415, it is possible to sequentially view obliquely a large number of images that do not fit in first area 411. In this case, it may be configured so that an image located in the center of first area 411 is selected simply by scrolling.
[0108] Furthermore, the selected image is displayed in its original form as viewed normally, not obliquely, in second area 422. That is, in Fig. 6, the part enclosed by dotted line 413 in first area 411 corresponds to an oblique projection of the image displayed in second area 422.
[0109] For example, when an observer such as a doctor selects images in order in the first area 411, a three-dimensional view of the target specimen is displayed as an animation in the second area 422.
[0110] In addition, the image displayed in the second area 422 can be enlarged or reduced, and a portion of the image selected in the first area 411 can be displayed in the second area 422. When only a portion of the image selected by zooming in or the like is displayed in the second area 422, the area corresponding to that portion can be shown to the observer by being surrounded by a thick line 414 in the first area 411, as shown in Fig. 7. When the entire image is displayed in the second area 422, the thick line 414 is not drawn, as shown in Fig. 6.
[0111] The viewer can zoom in or out of the image using the zoom-in button 425 and zoom-out button 426 located near the second area 422, and can move the position displayed within the second area 422 using the scroll bars 427 and 428 located on the edges of the second area 422.
[0112] Furthermore, an area to be observed in detail can be specified in the second area 422. That is, the receiving unit 212 can receive an area selection instruction to select an area from the selected image based on an operation of the mouse, keyboard, or the like.
[0113] 6 and 7, the area within the second area 422 surrounded by a thick frame 423 is the selected area. Of the selected images, the image within this area becomes the input image to be processed. For example, the position of the thick frame 423 can be changed by specifying and dragging the area within the thick frame 423. Furthermore, the size of the thick frame 423 may be changed by dragging a vertex or edge of the thick frame 423.
[0114] The area selection instruction can also be cancelled, in which case the entire selected image becomes the input image.
[0115] The selected input image is transmitted from the terminal computer 203 to the server computer 202 via the computer communication network 204 upon explicit instruction from the observer or upon the passage of a certain period of time after image selection and area selection, and is input to the image processing device 101 as an input image.
[0116] The output image estimated by the image processing device 101 is then transmitted from the server computer 202 to the terminal computer 203 via the computer communication network 204 .
[0117] Here, the transmission and reception of the input image and the output image is carried out by the transmission and reception unit 213 in the terminal computer.
[0118] The second display unit 214 displays the output image estimated by the image processing device 101 in a third area 433 on the screen 401 .
[0119] Let us consider a case where the image group to be estimated is color photographic images of stained specimens taken by an optical microscope, and the estimation results are the results of photographing by an electron microscope, which generally are grayscale images.
[0120] Therefore, in order to easily grasp the degree to which the magnification ratio of the output image displayed in the third area 433 differs from the magnification ratio of the input image displayed in the second area 422, i.e., the degree to which the output image displayed in the third area 433 is estimated from the input image that was actually captured, the output image displayed in the third area 433 can be colored.
[0121] For example, in the second region 422, the brightness, hue, and saturation of the pixel value displayed for each pixel are respectively set as follows: the brightness of the pixel in the output image corresponding to each pixel; the hue of the pixel in the input image corresponding to each pixel; Saturation determined according to the magnification ratio of the output image for the cross section of the target specimen By setting it to , it becomes possible to easily grasp the magnification ratio of the output image.
[0122] In this way, when the image processing system 201 is applied to use in medical institutions and research institutes, medical images (photographic images taken with optical microscopes or electron microscopes) provided by each institution are stored in the server computer 202, and medical images relating to the same target specimen (opposite cross sections) are used as training data for the super-resolution network 104, enabling better estimation.
[0123] In addition, each institution can select a medical image of the target specimen taken with an optical microscope on the terminal computer 203, and select a desired area within the image as needed, thereby obtaining an estimate of how the target specimen will appear when observed with an electron microscope, which can be useful for diagnosis and research.
[0124] (Experimental example) An experimental example in which the above embodiment is applied to converting a pathological image captured by an optical microscope into an electron microscope image captured by an electron microscope will be described below.
[0125] First, the pathology images and electron microscope images used as training data are scale-mapped to match their resolution, and then AKAZE features are detected for the entire image. A transform map is then calculated to match each position on the entire electron microscope image, and the entire pathology image is corrected.
[0126] Furthermore, the entire corrected pathological image is slid by 32 pixels to extract a 256x256 pixel pathological tile image.
[0127] On the other hand, for electron microscope images, a scanning range of 1024 x 1024 pixels ((1024-256) / 2 = 384 pixels up, down, left and right) is used, and the image with the highest similarity is selected.
[0128] As the super-resolution network, SR3, a technology disclosed in Non-Patent Document 2, is used.
[0129] FIG. 8A is a drawing-substitute photograph showing a photographed pathology image, an electron microscope image estimated from the pathology image, and the photographed electron microscope image arranged vertically in grayscale. FIG. 8B is a drawing-substitute photograph showing a photographed pathology image, an electron microscope image estimated from the pathology image, and the photographed electron microscope image arranged vertically in monochrome binary. FIG. 9A is a drawing-substitute photograph showing a photographed pathology image, an electron microscope image estimated from the pathology image, and the photographed electron microscope image arranged vertically in grayscale. FIG. 9B is a drawing-substitute photograph showing a photographed pathology image, an electron microscope image estimated from the pathology image, and the photographed electron microscope image arranged vertically in monochrome binary. FIG. 10A is a drawing-substitute photograph showing a photographed pathology image, an electron microscope image estimated from the pathology image, and the photographed electron microscope image arranged vertically in grayscale. FIG. 10B is a drawing-substitute photograph showing a photographed pathology image, an electron microscope image estimated from the pathology image, and the photographed electron microscope image arranged vertically in monochrome binary. In these figures, the representative photographs of drawings with drawing numbers ending in B are monochrome binarized versions of the representative photographs of drawings with the same drawing numbers and ending in A. For example, the drawing substitute photograph of Figure 8B is the drawing substitute photograph of Figure 8A that has been monochrome binarized.
[0130] In these figures, A pathological image of one cross section of the target specimen taken with an optical microscope at 400x magnification, an electron microscope image estimated from the pathological image according to this embodiment; and a power image of the other cross section of the target specimen taken with an electron microscope at a magnification of 1000 times, the power image being at a position corresponding to the pathological image; are shown side by side, enlarged and reduced for comparison.
[0131] As can be seen from these figures, the electron microscope images estimated from pathology images have higher resolution than the pathology images, while there is little difference between them and the actual electron microscope images taken, indicating that the estimation is of a quality that can be useful for doctors' diagnoses and researchers' research.
[0132] In this manner, in this embodiment, a pathological image is used as an input image, an electron microscope image is estimated from the input image, and the electron microscope image is output as an output image. The output image is merely reference information and is intended to be used for diagnostic support. Researchers are expected to be the primary users of the output image, but clinicians can also use it.
[0133] For example, kidney disease is difficult to treat, and once dialysis is required, the condition continues for the rest of a patient's life, so early treatment is necessary. In this embodiment, by generating electron microscope images from pathological photographs of glomeruli taken by the Renal Biomedical Laboratory, reference information for early diagnosis and early treatment can be provided. The Renal Biomedical Laboratory itself is often handled by an internist, but may also be handled by a pathologist.
[0134] Furthermore, in intestinal abnormalities caused by cancer, gene mutations cause changes in the structure of the intestinal tract, making early diagnosis using electron microscope images necessary, and the electron microscope images generated by this embodiment can be used as reference material for such diagnosis.
[0135] In the case of the Cardiac Biology Institute, by generating electron microscope images from pathological images, the state of the myocardium and fibers can be viewed as reference information. In this case, pathologists are often in charge.
[0136] In addition, amyloidosis, which affects people all over the world, can be found in any organ other than the intestine, and is classified as a primary disease of unknown cause that is endemic, and as a secondary disease caused by decreased renal function that also affects dialysis patients. The electron microscope images obtained in this embodiment can be used as reference information when observing the state of the affected area.
[0137] In general, photographing using an electron microscope is expensive because the methods of fixing specimens differ between pathological photographs and electron microscope photographs, and the number of photographers skilled in electron microscope photography is decreasing.
[0138] The electron microscope images output by this embodiment are generated from pathological images, so their cost is low. Therefore, by additionally using these electron microscope images as diagnostic support and reference information, they can be useful for low-cost early diagnosis and early treatment.
[0139] (Quantitative evaluation of output images) In the above embodiment, an electron microscope image is estimated from a pathological image and used as an output image. In this embodiment, the plausibility and reliability of this output image are quantitatively evaluated, and the evaluation value is provided as reference information to users of the output image.
[0140] 11 is an explanatory diagram showing the schematic configuration of an image processing device according to an embodiment of the present invention. The image processing device 101 according to this diagram is obtained by adding a generation unit 502 and an evaluation unit 503 to the image processing device 101 according to the above embodiment. The following description will be made with reference to this diagram.
[0141] First, the generation unit 502 provides the output image output from the super-resolution unit 103 to the trained generation network 505, causing it to generate a reference image.
[0142] The generative network 505 is a network that obtains an output that matches the original input as closely as possible by removing information from the input, such as by reducing the dimension of the input or adding noise to the input.
[0143] The simplest example of the generative network 505 is an autoencoder. Various types of autoencoders can be used, such as a stacked autoencoder, a convolutional autoencoder (CAE), a variational autoencoder (VAE), and a conditional variational autoencoder (CVAE).
[0144] Fig. 12 is an explanatory diagram showing the configuration of an encoder in a VAE, which is an example of a generative network in an image processing device according to an embodiment of the present invention. Fig. 13 is an explanatory diagram showing the configuration of a decoder in a VAE, which is an example of a generative network in an image processing device according to an embodiment of the present invention. A relatively simple VAE network as shown in this figure can be used as generative network 505. Fig. 14 is an explanatory diagram showing the configuration of an encoder in a VAE, which is another example of a generative network in an image processing device according to an embodiment of the present invention. Fig. 15 is an explanatory diagram showing the configuration of a decoder in a VAE, which is another example of a generative network in an image processing device according to an embodiment of the present invention. A network using a VAE with another configuration as shown in this figure can also be used as generative network 505. Also, Neural network based on the diffusion model, Generative Adversarial Network (GAN), A neural network based on a flow-based generative model (Flow-based Generative Network), A neural network that implements dimensionality reduction and restoration based on the Transformer Various networks such as the following may also be used as the generating network 505.
[0145] It is desirable that the generative network 505 proceeds with learning using a training sample different from the training sample used for learning the super-resolution network 104. The generative network 505 proceeds with learning using, as training input and training output, other training images obtained by photographing, with the second wave, other training cross sections that appear by cutting the other training sample.
[0146] To explain this in accordance with the above application example, in a mode in which an electron microscope image is estimated from a pathological image taken with an optical microscope and used as the output image, other electron microscope images of the same affected area are prepared as training images, and learning of the generative network 505 is carried out based on these electron microscope images.
[0147] Then, when an image is given as input, the generative network 505 is expected to output an image that closely represents the features specific to electron microscope images in that image.
[0148] Therefore, if the output image output by the super-resolution unit 103 is a plausible image that closely resembles an electron microscope image, the difference between it and the reference image generated by the generation unit 502 will be small, and if the output image output by the super-resolution unit 103 is different from an electron microscope image (for example, an image that contains abnormal information), the difference between the output image and the reference image will be large.
[0149] Therefore, the evaluation unit 503 quantitatively evaluates the output image numerically based on the difference between the output image and the reference image, and outputs the evaluation result.
[0150] The parameter values used for quantitative evaluation may include the following: The number of pixels that differ between the output image and the reference image, and the ratio of the number of pixels that differ to the total number of pixels. The sum of the different pixel values between the output image and the reference image, or the average of the different pixel values. The Mahalanobis distance between the output image and the reference image. Cosine similarity between the output image and the reference image. The similarity between the output image and the reference image calculated using machine learning and deep learning technologies such as AugNet. The deviation in the distribution of any of the above parameter values, a variance used in autoregressive models and One Class SVM outlier detection.
[0151] Figure 16 is a drawing substitute photograph showing a captured pathology image, an electron microscope image estimated from the pathology image, and a generated reference image arranged in grayscale. Figure 17 is a drawing substitute photograph showing a captured pathology image, an electron microscope image estimated from the pathology image, and a generated reference image arranged in monochrome binary. In these figures, for example (a) and example (b), the electron microscope image in the center is estimated from the pathology image on the left, and a reference image is generated from the estimated electron microscope image using the VAE disclosed in Figures 12 and 13. In example (a), a plausible electron microscope image is estimated, but in example (b), abnormal shapes (arrows, squares, and circles) are drawn, indicating that the generation of the electron microscope image failed.
[0152] When the above VAE is used to generate a reference image, in both example (a) and example (b), the generated image appears different from the estimated electron microscope image.
[0153] However, when the difference between the electron microscope image and the reference image is calculated by the sum of the pixel value differences, the difference is 2,397,066 in example (a) and 3,871,511 in example (b), so the difference is smaller in example (a).
[0154] Furthermore, the similarity between the electron microscope image and the reference image calculated by AugNet was 28.441 for example (a) and 34.211 for example (b). Example (a) has a smaller value, and according to AugNet's definition of similarity, the smaller the value, the more similar the images are, so example (a) has a smaller difference.
[0155] According to the inventor's experiment, for 199 electron microscope images, The percentage of pixels with a sum of pixel value differences of 1.5 million or less is 100%. 90% are under 2 million yen; 56% have incomes of 2.5 million yen or less; 27% are below 3 million yen; 7% for those under 3.5 million yen; 0% for items under 4 million yen This is what happened.
[0156] Similarly, The percentage of similarities below 25 is 100%. 77% were below 27.5; 57% are under 30; 38% had a score of 32.5 or less; 36% are under 35; 18% had a score of 37.5 or less; 0% for those under 40; This is what happened.
[0157] In addition, when a reference image was generated for a similar electron microscope image using the VAE disclosed in Figures 14 and 15, the difference and similarity in example (a) were 2763465 and 28.517, and the difference and similarity in example (b) were 3342139 and 29.369. In addition, for 199 electron microscope images, the sum of the pixel value differences was The percentage of items below 1.9 million yen is 100%. 68% have incomes of 2.3 million yen or less; 43% have incomes of 2.7 million yen or less; 41% have incomes of 3.1 million yen or less; 18% are below 3.5 million yen; 0% for items under 4 million yen and the similarity is The percentage of those under 16 is 92%. 68% were under 20; 50% for those under 24; 43% were 28 or younger; 42% are under 32; 0% for those under 36 This is what happened.
[0158] The percentages in these experimental examples indicate the percentage of low-quality electron microscope images that were inappropriately estimated by the super-resolution unit 103. Therefore, in general, if you want to consider only high-quality output images, you can reduce the threshold for the sum of pixel value differences or the AugNet similarity and select images with parameter values smaller than the threshold.
[0159] Furthermore, by calculating the proportion of appropriate images when the value parameter of the difference for the estimated electron microscope image is used as a threshold, a numerical value representing the likelihood of the estimated electron microscope image can be obtained.
[0160] In this way, based on the difference between the estimated electron microscope image and the reference image generated from it, it is possible to obtain a quantitative evaluation result that indicates the likelihood and quality of the electron microscope image.
[0161] Therefore, by providing a pathological image as an input image to the image processing device 101, researchers, doctors, etc. can obtain an output image from the image processing device 101 and a quantitative evaluation indicating how likely the output image is as an electron microscope image, and can decide whether or not to use the output image as reference information for diagnostic support.
[0162] The above explanation has been given of an embodiment in which an output image in which the resolution of an electron microscope and the electron beam are used as a second resolution and a second wave is estimated from an input image captured using a first resolution, the resolution of an optical microscope, and light as a first resolution and a first wave. However, the resolution and type of wave are not limited to these and can be applied to various embodiments, and these embodiments are also included in the scope of the present invention.
[0163] Furthermore, in a method in which an estimated output image is given to a generative network trained using images captured at a second resolution and a second wave to obtain a reference image, and the output image is quantitatively evaluated based on the difference between the output image and the reference image, the specific generative network and evaluation method are not limited to the above-mentioned aspects, and various outlier detection and anomaly detection technologies can be applied, and these aspects are also included in the scope of the present invention.
[0164] (summary) As described above, the image processing apparatus according to this embodiment: an input unit to which an input image of a target object captured with a first wave having a first resolution is input; a super-resolution unit that estimates an output image that should be obtained by photographing the target specimen with a second wave having a second resolution by providing the input image to a trained super-resolution network; Equipped with The second resolution is higher than the first resolution. Configure it as follows.
[0165] In addition, in the image processing device according to this embodiment, The second waveband of the second vibration is shorter than the first waveband of the first vibration. It can be configured as follows.
[0166] In addition, in the image processing device according to this embodiment, the first wave is light, The second wave is an electron beam It can be configured as follows.
[0167] In addition, in the image processing device according to this embodiment, the first wave is unpolarized light; The second wave is polarized light It can be configured as follows.
[0168] In addition, in the image processing device according to this embodiment, The image of the cross section of the target specimen is taken as the input image. It can be configured as follows.
[0169] In addition, in the image processing device according to this embodiment, The super-resolution network comprises: a training input based on a first training image obtained by photographing a first training cross section that appears by cutting a training specimen with the first wave; a training output based on a second training image obtained by photographing a second training cross section that appears opposite to the first training cross section with the second wave; It is learned from training data including It can be configured as follows.
[0170] In addition, in the image processing device according to this embodiment, The image processing device includes: correcting the first training images with a mapping that aligns the first training images with the second training images; a first region of the corrected first training image is cut out as a first partial image; Scanning a predetermined scanning range from the second region using a window having the same shape as the second region projected onto the second training image by the mapping of the first region, and selecting an image within the scanned window that has the highest similarity to the first partial image as a second partial image; The super-resolution network uses the first partial image as the training input and the second partial image as the training output. It is learned by It can be configured as follows.
[0171] In addition, in the image processing device according to this embodiment, In the alignment, By matching the resolution of the first training image and the second training image and then performing feature point detection and feature point matching, a plurality of pairs are obtained, each of the plurality of pairs being: first feature points detected in the first training image; second feature points detected in the second training image, the second feature points corresponding to the first feature points; Extract a number of pairs consisting of The mapping is a projection of each of the pairs of first feature points by the mapping; the plurality of pairs of second feature points; is a mapping that minimizes the difference between It can be configured as follows.
[0172] In addition, in the image processing device according to this embodiment, In the alignment, Dividing the first training image into a plurality of polygons having vertices that correspond to the feature points detected in the first training image by the feature point detection; A plurality of candidate mappings each having a domain of the plurality of polygons, wherein each candidate mapping of the plurality of candidate mappings comprises: projecting a first polygon that is the domain of each candidate mapping onto a second polygon in the second training image; The vertices of the first polygon are projected onto the vertices of the second polygon that are respectively associated with the vertices of the first polygon by the feature point matching. Find multiple candidate mappings, clustering the plurality of candidate mappings to identify minority mappings belonging to a minority cluster and majority mappings other than the minority cluster; Let the direct sum of the majority maps be the map, the first sub-image is included in the domain of the mapping; The second partial image is included in the range of the mapping. It can be configured as follows.
[0173] In addition, in the image processing device according to this embodiment, the polygon is a triangle, the partitioning is a Delaunay partitioning, The plurality of candidate mappings may be inversion, translation, rotation, scaling, shear, trapezoidal transformation, or a combination thereof. It can be configured as follows.
[0174] In addition, in the image processing device according to this embodiment, The mapping is a direct sum or diffeomorphism of multiple homography mappings learned by reinforcement learning or deep learning. It can be configured as follows.
[0175] In addition, in the image processing device according to this embodiment, a generating unit that generates a reference image by providing the output image to a generating network that has been trained by using other training images, which are obtained by cutting other training cross sections that appear by cutting other training specimens and are photographed with the second wave, as training inputs and training outputs; an evaluation unit that quantitatively evaluates the output image based on a difference between the output image and the reference image; The device may be configured to further include:
[0176] In addition, in the image processing device of this embodiment, The generator network comprises: Autoencoders, including stacked autoencoders, convolutional autoencoders (CAEs), variational autoencoders (VAEs), and conditional variational autoencoders (CVAEs). Neural network based on the diffusion model, Generative Adversarial Network (GAN), A neural network based on a flow-based generative model (Flow-based Generative Network), Transformer-based neural networks The configuration can be any one of the above.
[0177] In addition, in the image processing device of this embodiment, The evaluation unit outlier detection based on the number of different pixels or the distribution of different pixel values between the output image and the reference image (including outlier detection based on an autoregressive model); Outlier detection based on One Class SVM of the difference; outlier detection based on a Mahalanobis distance between the output image and the reference image; the cosine similarity between the output image and the reference image; AugNet-based similarity between the output image and the reference image The output image is quantitatively evaluated by any one of the following methods. It can be configured as follows.
[0178] The image processing system according to this embodiment includes a terminal and the image processing device described above, The terminal a first display unit that displays a plurality of images in a first area within the screen; a reception unit that receives an image selection instruction for selecting one of the displayed images; a transmitting / receiving unit that transmits the selected image to the image processing device as the input image and receives the output image estimated by the image processing device from the image processing device; a second display unit that displays the received output image in a second area within the screen; The present invention can be configured to include the following.
[0179] In addition, in the image processing system according to this embodiment, The plurality of images are a group of images obtained by photographing a plurality of cross sections that appear when the target specimen is cut into layers. It can be configured as follows.
[0180] In addition, in the image processing system according to this embodiment, When the receiving unit receives an area selection instruction to select any area from the selected image, the second display unit sets an image in the selected area of the selected image as the input image, and displays the output image obtained for the input image in the second area on the screen. It can be configured as follows.
[0181] In addition, in the image processing system according to this embodiment, the set of images are color images, the output image is a grayscale image; The lightness, hue, and saturation of the pixel value displayed in each pixel of the second region are respectively: the brightness of the pixel of the output image corresponding to each pixel; the hue of the pixel of the input image corresponding to each pixel; a saturation determined according to a magnification ratio of the output image for the cross section of the target specimen; is set to It can be configured as follows.
[0182] In addition, in the image processing system according to this embodiment, The first display unit is arranging the images included in the image group on a plurality of planes parallel to each other set in a virtual three-dimensional space in the order in which a plurality of cross sections corresponding to the images are arranged in the target specimen; The state of the virtual three-dimensional space is drawn in the first area. and displaying the group of images in the first area. It can be configured as follows.
[0183] The image processing method according to this embodiment includes: an input step of inputting an input image of a target specimen captured with a first wave having a first resolution; a super-resolution step of estimating an output image that should be obtained by photographing the target specimen with a second wave having a second resolution by providing the input image to a trained super-resolution network; Equipped with The second resolution is higher than the first resolution. Configure it as follows.
[0184] The program according to this embodiment executes the following steps: an input unit to which an input image of a target object captured with a first wave having a first resolution is input; a super-resolution unit that estimates an output image that should be obtained by photographing the target specimen with a second wave having a second resolution by providing the input image to a trained super-resolution network; It functions as The second resolution is higher than the first resolution. Configure it as follows.
[0185] The program according to this embodiment can be recorded on a non-transitory computer-readable information recording medium and distributed or sold, or can be distributed or sold via a temporary transmission medium such as a computer communication network.
[0186] The present invention allows various embodiments and modifications without departing from the broad spirit and scope of the present invention. Furthermore, the above-described embodiments are intended to explain the present invention and do not limit the scope of the present invention. That is, the scope of the present invention is defined by the claims, not the embodiments. Various modifications made within the scope of the claims and the meaning of the invention equivalent thereto are considered to be within the scope of the present invention. This application claims priority based on patent application No. 2022-192062, filed in Japan on Wednesday, November 30, 2022, and the contents of that basic application are incorporated into this application to the extent permitted by the laws and regulations of the designated countries. [Industrial Applicability]
[0187] According to the present invention, it is possible to provide an image processing device, an image processing system, an image processing method, a program, and an information recording medium that perform estimation to improve the resolution of an input image of a specimen and obtain an output image. [Explanation of symbols]
[0188] 101 Image processing device 102 Input section 103 Super-resolution unit 104 Super-resolution Network 201 Image Processing System 202 Server Computer 203 Terminal Computer 204 Computer Communication Network 211 1st display section 212 Reception Department 213 Transmitter / Receiver 214 2nd display section 401 screen 411 First area Each image in the 412 image set 413 Dotted line representing selected image 414 Bold lines representing the part displayed in the second area 415 Scrollbar 422 Second area 423 Bold border representing selected area 425 Zoom in button 426 Zoom out button 427 Scrollbar 428 Scrollbar 433 Third area 502 Generation part 503 Evaluation Department 505 Generative Network
Claims
1. an input unit to which an input image of a cross section of a target specimen captured with a first wave having a first resolution is input; a super-resolution unit that estimates an output image that should be obtained by photographing the cross section of the target specimen with a second wave having a second resolution by providing the input image to a trained super-resolution network; Equipped with the second resolution is higher than the first resolution; The second waveband of the second vibration is shorter than the first waveband of the first vibration.
1. An image processing device comprising:
2. the first wave is light, The second wave is an electron beam 2. The image processing device according to claim 1, wherein:
3. the first wave is unpolarized light; The second wave is polarized light 2. The image processing device according to claim 1, wherein:
4. The super-resolution network comprises: a training input based on a first training image obtained by photographing a first training cross section that appears by cutting a training specimen with the first wave; a training output based on a second training image obtained by photographing a second training cross section that appears opposite to the first training cross section with the second wave; It is learned from training data including 2. The image processing device according to claim 1, wherein:
5. The image processing device includes: correcting the first training images with a mapping that aligns the first training images with the second training images; a first region of the corrected first training image is cut out as a first partial image; Scanning a predetermined scanning range from the second region using a window having the same shape as the second region projected onto the second training image by the mapping of the first region, and selecting an image within the scanned window that has the highest similarity to the first partial image as a second partial image; The super-resolution network uses the first partial image as the training input and the second partial image as the training output. It is learned by 5. The image processing device according to claim 4.
6. The mapping is a direct sum or diffeomorphism of multiple homography mappings learned by reinforcement learning or deep learning.
6. The image processing device according to claim 5,
7. a generating unit that generates a reference image by providing the output image to a generating network that has been trained by using other training images, which are obtained by cutting other training cross sections that appear by cutting other training specimens and are photographed with the second wave, as training inputs and training outputs; an evaluation unit that quantitatively evaluates the output image based on a difference between the output image and the reference image; 2. The image processing device according to claim 1, further comprising:
8. The generator network comprises: Autoencoders, including stacked autoencoders, convolutional autoencoders (CAEs), variational autoencoders (VAEs), and conditional variational autoencoders (CVAEs). Neural network based on the diffusion model, Generative Adversarial Network (GAN), A neural network based on a flow-based generative model (Flow-based Generative Network), Transformer-based neural networks 8. The image processing device according to claim 7, wherein the image processing device is one of the above.
9. The evaluation unit outlier detection based on the number of different pixels or the distribution of different pixel values between the output image and the reference image (including outlier detection based on an autoregressive model); Outlier detection based on One Class SVM of the difference; outlier detection based on a Mahalanobis distance between the output image and the reference image; the cosine similarity between the output image and the reference image; AugNet-based similarity between the output image and the reference image The output image is quantitatively evaluated by any one of the following methods.
8. The image processing device according to claim 7,
10. An image processing system including a terminal and an image processing device, The image processing device includes: an input unit to which an input image of a target object captured with a first wave having a first resolution is input; a super-resolution unit that estimates an output image that should be obtained by photographing the target specimen with a second wave having a second resolution by providing the input image to a trained super-resolution network; Equipped with the second resolution is higher than the first resolution; a second waveband of the second vibration is shorter than a first waveband of the first vibration; The terminal a first display unit that displays a plurality of images in a first area within the screen; a reception unit that receives an image selection instruction for selecting one of the displayed images; a transmitting / receiving unit that transmits the selected image to the image processing device as the input image and receives the output image estimated by the image processing device from the image processing device; a second display unit that displays the received output image in a second area within the screen; An image processing system comprising:
11. The image of the cross section of the target specimen is taken as the input image.
11. The image processing system according to claim 10.
12. The plurality of images are a group of images obtained by photographing a plurality of cross sections that appear when the target specimen is cut into layers.
12. The image processing system according to claim 11.
13. When the receiving unit receives an area selection instruction to select any area from the selected image, the second display unit sets an image in the selected area of the selected image as the input image, and displays the output image obtained for the input image in the second area on the screen.
13. The image processing system according to claim 12.
14. the set of images are color images, the output image is a grayscale image; The lightness, hue, and saturation of the pixel value displayed in each pixel of the second region are respectively: the brightness of the pixel of the output image corresponding to each pixel; the hue of the pixel of the input image corresponding to each pixel; a saturation determined according to a magnification ratio of the output image for the cross section of the target specimen; is set to 14. The image processing system according to claim 13.
15. The first display unit is arranging the images included in the image group on a plurality of planes parallel to each other set in a virtual three-dimensional space in the order in which a plurality of cross sections corresponding to the images are arranged in the target specimen; The state of the virtual three-dimensional space is drawn in the first area. and displaying the group of images in the first area.
13. The image processing system according to claim 12.
16. an input step of inputting an input image of a cross section of a target specimen captured with a first wave having a first resolution; a super-resolution step of estimating an output image that should be obtained by photographing the cross section of the target specimen with a second wave having a second resolution by providing the input image to a trained super-resolution network; Equipped with the second resolution is higher than the first resolution; The second waveband of the second vibration is shorter than the first waveband of the first vibration. An image processing method comprising:
17. Computer, an input unit to which an input image of a cross section of a target specimen captured with a first wave having a first resolution is input; a super-resolution unit that estimates an output image that should be obtained by photographing the cross section of the target specimen with a second wave having a second resolution by providing the input image to a trained super-resolution network; It functions as the second resolution is higher than the first resolution; The second waveband of the second vibration is shorter than the first waveband of the first vibration. A program characterized by:
18. A non-transitory computer-readable information recording medium on which the program according to claim 17 is recorded.
Citation Information
Patent Citations
Learning device, determination device, microscope, trained model, and program
JP2021018582A
Image processing method, program, image processing device, trained model manufacturing method, learning method, learning device, and image processing system
JP2022056769A
Systems and methods for converting holographic microscope images into microscope images of various modalities
JP2022507259A
Medical devices using narrowband light observation
JP2022527642A