Method and apparatus for generating high-depth field images using stereo images, and apparatus for training a high-depth field image generation model using stereo images.
The use of stereo images and deep learning models in microscopes and scanners addresses the challenge of capturing high-depth images by segmenting and estimating depth ranges, resulting in efficient and clear all-in-focus videos.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- VIEWORKS CO LTD
- Filing Date
- 2022-07-01
- Publication Date
- 2026-05-15
AI Technical Summary
Existing microscope and slide scanner technologies face challenges in capturing high-depth-of-field images due to the need for optical structures that change focal planes and lengthy shooting times, especially when imaging objects with varying heights or three-dimensional shapes.
A method and apparatus that generate high-depth images using stereo images, employing a deep learning model to segment and estimate depth ranges, eliminating the need for optical structures and reducing shooting time by generating images with a wider range of depths in focus.
This approach allows for the generation of all-in-focus high-depth videos without requiring multiple focal plane adjustments, significantly shortening shooting time and enhancing image clarity for objects at various positions.
Smart Images

Figure 0007859893000002 
Figure 0007859893000003 
Figure 0007859893000004
Abstract
Description
Technical Field
[0001] The present invention relates to a high-depth video generation method and apparatus, and a high-depth video generation model learning apparatus.
Background Art
[0002] A microscope is a mechanism for magnifying and observing minute objects and microorganisms that are difficult to observe with the human eye. A slide scanner used in conjunction with a microscope is a device capable of automatically scanning one or more slides to save and observe and analyze images. Generally, since a microscope uses a high-magnification lens for photographing tissues and cells, etc., the depth of field is low, and it is difficult to simultaneously photograph cells distributed at various heights. For example, when using a specimen of a 4-μm-thick tissue in a pathological examination, the depth of focus of a 40-fold objective lens is at most about 1 μm. If two or more cells in a single photographing area are distributed with a height difference of 1 μm or more, it is difficult to photograph all of these cells in focus in one image. Also, since an object having a three-dimensional shape rather than a planar shape must be photographed, it is necessary to focus on an uneven surface. Generally, since many cells in an image exist at different positions from each other, it is difficult to obtain an overall focused image.
[0003] Therefore, in order to obtain a high-depth-of-field video in a microscope or a slide scanner, a z-stack (focus stack) technique is used, in which a plurality of videos are photographed while changing the focal plane of the z-axis from a position where the x and y axes are fixed, and then they are synthesized.
[0004] However, such z-stack techniques have various problems, including the need for an optical structure that allows for multiple images at different depths to change the focal plane of the z axis, and the need to repeatedly determine, change, and capture a very large number of focal planes before combining them, resulting in long shooting times. Furthermore, methods that use technologies such as lasers to determine the distance (depth) to focus are also unsuitable for imaging subjects where the focus of the image base needs to be determined. [Overview of the project] [Problems that the invention aims to solve]
[0005] The technical problem that the present invention aims to solve is to provide a method and apparatus for generating depth-of-depth images that does not require optical structures for multiple images at different depths and can generate depth-of-depth images from captured images, as well as a learning apparatus for a depth-of-depth image generation model therefor. [Means for solving the problem]
[0006] To solve the aforementioned technical problems, the high-depth image generation apparatus according to the present invention generates stereo images of tissue or cells. The depth ranges are different from each other. For each of the left and right images Located in each of the aforementioned depth ranges The region of the tissue or cells is divided, and the left image The regions of the tissue or cells that are divided by, and 、 The aforementioned right video Each of the regions of the tissue or cells that have been divided by the divisions is to include the respective region of the tissue or cells. The system includes: a region division unit that generates region data that separates the tissue or cell region from other regions in the stereo image; a depth estimation unit that estimates the depth for each tissue or cell region in the stereo image and generates depth data for the tissue or cell region; and a depth image generation unit that generates a depth image with a wider range of depths in focus than each image constituting the stereo image from the stereo image, the region data, and the depth data. The depth image generation unit generates the left image The regions of the tissue or cells that are divided by, and 、 The aforementioned right video Divided by The high-depth image is generated, which includes the region of the tissue or cells. The high-depth image generation unit may generate the high-depth image using a trained deep learning model.
[0007] The region division unit or the depth estimation unit may generate the region data or the depth data using a trained deep learning model.
[0010] The trained deep learning model may be implemented to mimic blind deconvolution using a point-spread function.
[0011] To solve the aforementioned technical problems, the high-depth image generation method according to the present invention provides stereo images of tissue or cells The depth ranges are different from each other. For each of the left and right images Located in each of the aforementioned depth ranges The region of the tissue or cells is divided, and the left image The regions of the tissue or cells that are divided by, and 、 The aforementioned right video Each of the regions of the tissue or cells that have been divided by the divisions is to include the respective region of the tissue or cells. The process includes: a region segmentation process that generates region data that separates the tissue or cell region from other regions in the stereo image; a depth estimation process that estimates the depth of each tissue or cell region in the stereo image and generates depth data for the tissue or cell region; and a depth-of-field image generation process that generates a depth-of-field image with a wider range of depths in focus than each image constituting the stereo image from the stereo image, the region data, and the depth data, wherein the depth-of-field image generation process includes the left image The regions of the tissue or cells that are divided by, and 、 The aforementioned right video Divided by The high-depth image may be generated that includes the region of the tissue or cells.
[0012] The aforementioned high-depth image generation process may utilize a trained deep learning model.
[0013] The domain segmentation process or the depth estimation process may utilize a trained deep learning model.
[0016] The trained deep learning model may be implemented in a way that mimics blind deconvolution using a point-spread function.
[0017] To solve the aforementioned technical problems, the high-depth image generation model learning device according to the present invention is a learning model that is implemented to output a high-depth image from a stereo image of an input tissue or cell, wherein the stereo image The depth ranges are different from each other. For each of the left and right images Located in each of the aforementioned depth ranges The region of the tissue or cells is divided, and the left image The regions of the tissue or cells that are divided by, and 、 The aforementioned right video Each of the regions of the tissue or cells that have been divided by the divisions is to include the respective region of the tissue or cells. A learning model including: a region division unit that generates region data that separates the tissue or cell region from other regions in the stereo image; a depth estimation unit that estimates the depth for each tissue or cell region in the stereo image and generates depth data for the tissue or cell region; and a high-depth image generation unit that generates a high-depth image with a wider range of depths in focus than each image constituting the stereo image from the stereo image, the region data, and the depth data; and a learning unit that trains the learning model from learning data including the stereo image and a corresponding reference high-depth image, wherein the high-depth image generation unit generates the left image The regions of the tissue or cells that are divided by, and 、 The aforementioned right video Divided by The high-depth image is generated, which includes the region of the tissue or cells.
[0018] The learning unit may calculate a cost function from the high-depth video output from the learning model and the reference high-depth video, and use the cost function to train the learning model.
[0019] The high-depth image generation unit may generate the high-depth image using a deep learning model.
[0020] The area division unit or the depth estimation unit may generate the area data or the depth data using a deep learning model.
[0023] The deep learning model may be implemented to simulate blind deconvolution using a point-spread function.
[0024] The high-depth video generation model learning device may further include a video preprocessing unit that preprocesses the stereo video and inputs it to the learning model.
Advantages of the Invention
[0025] According to the present invention, an optical structure for multiple shootings at different depths is not required, and a high-depth video can be generated from a stereo video. Therefore, the shooting time is significantly shortened, and a high-depth video can be effectively obtained from the stereo video.
[0026] The basic information required to change the focal plane is the depth (z-axis) at which the object is located. The present invention utilizes a stereo technology like the human eye to grasp the depth at which the cell is located. In addition, in order to eliminate the time-consuming repeated shooting process, a deconvolution algorithm via a stereo video can be performed to obtain the effect of increasing the depth. Further, based on the information obtained by distinguishing an object from a stereo video and estimating the depth of each object, a video adjusted to be in focus for each region of the object can be generated through deconvolution restoration.
[0027] According to the present invention, it is possible to generate an all-in-focus high-depth video for objects (such as cells) existing at various positions even on a slide thicker than before.
Brief Description of the Drawings
[0028] [Figure 1]This document shows a high-depth image generation model learning device according to one embodiment of the present invention. [Figure 2] This is a diagram illustrating an example of stereoscopic imagery. [Figure 3a] This diagram illustrates examples of determining the camera angle in cytopathology and histopathology, respectively. [Figure 3b] This diagram illustrates examples of determining the camera angle in cytopathology and histopathology, respectively. [Figure 4] This is a diagram illustrating an example of a standard high-depth image. [Figure 5] This is a diagram illustrating the process by which region data is generated. [Figure 6] This diagram illustrates the process by which depth is estimated through cell size, central position, and observation position. [Figure 7] An example of depth data obtained for the stereo image in Figure 2 is shown. [Figure 8] This shows a high-depth image generation device according to one embodiment of the present invention. [Figure 9] An example of cell region segmentation images, region data, and depth data obtained from stereo video by one embodiment of the present invention is shown. [Figure 10] An example of a high-depth image obtained from stereo image, region data, and depth data according to an embodiment of the present invention is shown. [Modes for carrying out the invention]
[0029] Preferred embodiments of the present invention will be described in detail below with reference to the drawings. Substantially identical components in the following description and accompanying drawings will be denoted by the same reference numerals, and redundant explanations will be omitted. In describing the present invention, if it is determined that a specific description of a related known function or configuration would unnecessarily obscure the gist of the invention, such detailed description will be omitted.
[0030] The inventors of this application focused on the present invention because stereo images captured through a stereo camera contain depth information, and since depth information is closely related to focus, high-depth images can be generated from stereo images through a deep learning model. In embodiments of the present invention, stereo images are described using images of cells as an example, but of course, the objects of photography can be diverse, including not only biological tissues such as cells, but also tissues, materials, products, etc. In the following description, high-depth images refer to images with a wider range of depths in focus than single-image images, meaning, for example, images with a wider range of depths in focus than each image that makes up the stereo image.
[0031] Figure 1 shows a high-depth image generation model learning device according to one embodiment of the present invention. The high-depth image generation model learning device according to this embodiment may include a video preprocessing unit 110 for preprocessing stereo video, a learning model 120 that is implemented to output high-depth video from stereo video, and a learning unit 130 for training the learning model 120 from training data.
[0032] The training data consists of a dataset of stereo images and their corresponding reference depth-of-depth images. Figure 2 illustrates an example of stereo images. The stereo images may include left and right images captured through the left and right cameras, respectively. Referring to Figure 2, the left and right cameras capture different focal planes (depth ranges) that are tilted at a certain angle relative to each other. Therefore, for example, if there are three cells C1, C2, and C3, the left image may capture only cells C2 and C3, while the right image may capture only cells C1 and C2. Alternatively, the same cell may appear in focus and sharp in one image, but out of focus and blurred in the other image.
[0033] JPEG0007859893000001.jpg98168
[0034] On the other hand, stereo images can be acquired using two or more cameras having different optical paths, or they can be acquired using a structure in which one optical path is split into two or more optical paths via optical splitting means (e.g., beam splitter, prism, mirror, etc.), and multiple cameras obtain images with focal planes tilted at a certain angle.
[0035] Figure 4 illustrates an example of a reference depth-of-field image. The reference depth-of-field image is a depth-of-field image acquired of the same object as the stereo image, and may be acquired using existing z-stack techniques. That is, the reference depth-of-field image may be acquired by combining multiple images taken while changing the focal plane of the z axis at the same (x,y) offset position as the stereo image. Referring to Figure 4, any of the three cells C1, C2, and C3 can appear in the reference depth-of-field image.
[0036] The video preprocessing unit 110 can perform video augmentation and video normalization as preprocessing for stereo video.
[0037] By increasing the amount of training data through image augmentation, robustness against noise can be ensured, and the model can learn under various shooting conditions. The image preprocessing unit 110 can increase the amount of training data by arbitrarily adjusting the brightness, contrast, RGB values, etc., of the image. For example, the image preprocessing unit 110 can adjust the brightness, contrast, RGB values or their distribution according to the mean and standard deviation, or adjust the intensity of staining by arbitrarily adjusting the absorption coefficients obtained according to the Beer-Lambert Law of Light Absorption.
[0038] Video normalization can improve the performance of many learning models and increase the learning convergence speed. Stereo video training data generally consists of images of various target cells taken at multiple time points through multiple devices, meaning that the shooting conditions may differ. Furthermore, even for the same target cell, the images may be obtained differently due to shooting conditions and various environmental variables replicated through augmentation. Therefore, video normalization may be performed to minimize various variations and align the color space of the images. For example, staining normalization may utilize the light absorption coefficients of H&E staining according to the Vier-Lambert rule.
[0039] Depending on the embodiment, both video enhancement and video normalization may be performed, or only one of them may be performed, or both may be omitted, depending on the constraints of the learning resources.
[0040] Referring again to Figure 1, the learning model 120 according to an embodiment of the present invention is constructed to include a region division unit 121 that divides cell regions from an input stereo image to generate region data, a depth estimation unit 122 that estimates depth from an input stereo image to generate depth data, and a high-depth image generation unit 123 that generates a high-depth image from the input stereo image, the region data output from the region division unit 121, and the depth data output from the depth estimation unit 122.
[0041] The region division unit 121 can divide the cell region for both the left and right images, and then combine the left and right images with the divided cell regions to generate region data. When dividing the cell region, the region division unit 121 can separate the cell region from other regions into foreground and background, while also separating each cell as a separate individual.
[0042] The process of segmenting cell regions may be carried out via deep learning models such as DeepLab v3 or U-Net. Using deep learning models offers advantages such as fast segmentation speed, robustness against various shooting conditions like image mutations, high accuracy for complex images, and the ability to fine-tune according to the purpose. If resources for training are limited, the process of segmenting cell regions may be carried out using existing region segmentation algorithms such as Otsu Thresholding, Region Growing, Watershed algorithm, Graph cut, Active contour model, or Active shape model.
[0043] The left and right images that make up a stereo image have low depth, and not all cells may be observed equally from both images. Therefore, the left and right images may be joined so that cells located at different depths appear. In the process of joining the left and right images, the positions of each cell can be joined through a deep learning model, resulting in natural stitching. Figure 5 is a diagram illustrating the process of generating region data. If the left and right images differ at the same position even considering stereo parallax, the foreground color that is farther from the background color may be selected from the left and right images, taking into account the background color estimated from the overall image. The background color may be defined as the color with the most prevalent distribution in the edge regions of the image. As illustrated in Figure 5, the region data may have values that distinguish the background and each cell region. For example, as shown in Figure 5, the background may be given 0, and each cell region may be given different pixel values such as 1, 2, 3, etc.
[0044] Generally, stereo images aim to generate three-dimensional images by estimating the distance or depth of an object using the visual difference between two images. In contrast, embodiments of the present invention use stereo images to perform region segmentation. In particular, high-magnification images have a very low depth of field, and often very different objects are captured in stereo images that simultaneously capture a single imaging area. For example, when analyzing stereo images, if a particular cell is captured in only one of the two images, when processed using a general stereo depth estimation method, this cell may appear very blurry or be excluded from the image. In embodiments of the present invention, detail regions are segmented based on the shape and form of the object captured in each of the two stereo images. In this case, if the segmentation positions of the two stereo images do not coincide, the object segmented from the final image can be clearly represented using depth information for the relevant region. Through this, it is possible to generate high-depth images that exceed the physical depth of field of the objective lens, and the depth in this case is determined by the parallax angle of the stereo image.
[0045] Generally, in high-magnification imaging, image analysis is performed to determine the optimal lens position for focus. However, in the embodiments of the present invention, in addition to image analysis, it is also possible to measure the height (position) of the tissue sample slide via a laser sensor or the like. The method of estimating the slide height using a laser sensor or the like has lower accuracy compared to the lens depth of field, and even if the slide height is known, it is not possible to know at what height (position) of the sample thickness the cells are located on the slide. Therefore, this method is not used in general high-resolution optical imaging systems. However, in the embodiments of the present invention, the depth of field can be increased to a thickness similar to that of the sample. Therefore, by knowing the slide height, an optimized focused image can be secured, which has the advantage of making laser sensors, which were previously difficult to use biologically, applicable to the lens height adjustment method. When using a laser sensor, the position of focus can be determined more quickly compared to the method of determining the focal height through image analysis, and the shooting speed can be increased. Furthermore, in the embodiments of the present invention, since the depth of field of the image is longer than the depth of field of the lens, it is not necessary to analyze the optimal position of the image focus while adjusting the lens height in units similar to or short increments to the lens depth of field. Therefore, a nanometer-level ultra-precise lens position adjustment mechanism is not required to adjust the height of the lens focal point.
[0046] The depth estimation unit 122 extracts feature maps from both the left and right images, and can estimate the depth from the feature maps through the cell size and central position or observation position. While the central part of the image has many areas where objects (cells) are similarly observed, the outer edges of the image have relatively smaller areas where objects (cells) are similarly observed due to differences in depth (z). This phenomenon becomes more pronounced as the angle of the camera sensor increases due to high depth, which can affect the performance of image segmentation and depth estimation. Therefore, morphology, color, central position, and observation position may be considered to distinguish objects (cells). The process of extracting feature maps from the left and right images can be performed via a Convolutional Neural Network (CNN) model such as VGG, ResNet, or Inception. If resources for training are limited, feature maps may be extracted using existing stereo matching algorithms.
[0047] Figure 6 illustrates the process by which depth is estimated through cell size, central position, and observation position. When the same cell is observed in both the left and right images due to disparity, the depth can be estimated by comparing the cell's range (size) and central position. For example, cell C2 is observed to the right of the x-axis center, and its size is smaller and its center is more off-center in the right image than in the left image. Therefore, it can be estimated to be located above the reference focal point (the point where the two focal planes intersect, see Figure 6) (i.e., at a shallow depth). When a cell is observed in only one of the left or right images, the depth can be estimated by considering the angle of the camera that captured that image. For example, cell C1 is observed in the right image, and its center is to the left of the x-axis center. Therefore, it can be estimated to be located above the reference focal point (i.e., at a shallow depth), while cell C3 is observed in the left image, and its center is to the left of the x-axis center. Therefore, it can be estimated to be located below the reference focal point (i.e., at a deep depth). In this case, the depth value can be estimated by considering the size of the cell and the distance between the center of the cell and the center of the x-axis. However, the depth value may also be obtained from the feature maps of the left and right images via a CNN model such as DeepFocus. Figure 7 shows an example of depth data obtained for the stereo image of Figure 2. For example, the depth data may have a depth value of 0 for the background and depth values of 1 to 255 for cell regions depending on their depth.
[0048] The results from the region division unit 121 and the depth estimation unit 122 influence each other to generate the final high-depth image, thereby improving the image quality.
[0049] The depth estimation unit 122 may, in estimating the depth of the object being photographed, independently estimate the depth for each divided region of the left and right images, which have been divided via the region division unit 121, and then generate the final depth data. Furthermore, the depth estimation results for each region may affect the region division results of the final high-depth image. As illustrated in Figure 6, in the case of cell C2 captured simultaneously from the left and right images, the left and right images are divided into different regions relative to cell C2, and the determination of which region better reflects the actual size of the cell is as follows: The depth estimation unit 122 determines that the image with a relatively large size and clear contrast among the regions of the left and right images was captured from a focal height closer to the actual value, estimates the depth of the object, and can extract information about the divided region from that image. Therefore, the region range of cell C2 in the final high-depth image can be generated to be similar to the size that appears in the left image.
[0050] Referring again to Figure 1, the high-depth image generation unit 123 may receive stereo video, region data, and depth data as input. If the stereo video is RGB video, the high-depth image generation unit 123 may receive 8 channels of data, including 3 RGB channels for the left video and 3 channels of region data for the right video, 1 channel of region data, and 1 channel of depth data for each video. The high-depth image generation unit 123 can generate 3 RGB channel high-depth video from such 8 channels of input data through a deep learning model.
[0051] The high-depth image generation unit 123 can perform deconvolution on the stereo image, considering the depth for each region, through a deep learning model based on region data, depth data, and the stereo image. The deep learning model may be implemented to determine the degree of focus of a region or sub-region of the image, i.e., the in-focus or out-focus level, from the input data and apply the learned point diffusion function. The point diffusion function is a function that represents the shape of light scattering when a point light source is photographed, and by applying the point diffusion function inversely, a clear image can be obtained from a blurred image. The deep learning model may be a CNN model and may be implemented to mimic blind deconvolution, which estimates and applies the inverse function of the point diffusion function. To improve the learning performance of the deconvolution model, the input data may be preprocessed using algorithms such as the Jansson-Van Cittert algorithm, Agard's Modified algorithm, Regularized least squares minimization method, MLE (Maximum Likelihood Estimation), and EM (Expectation Maximization).
[0052] The learning unit 130 trains the learning model 120 from training data that includes stereo video and corresponding reference high-depth video. At this time, the learning unit 130 may train the learning model 120 using end-to-end learning.
[0053] The learning unit 130 can calculate a cost function from the high-depth video output from the high-depth video generation unit 123 and the reference high-depth video, and update the parameters (weights) of the learning model 120 using the cost function. If the domain division unit 121, depth estimation unit 122, and high-depth video generation unit 123 that constitute the learning model 120 are all implemented as deep learning models, all of these parameters can be updated throughout the learning process. If only some of these are implemented as deep learning models, the parameters of the corresponding deep learning model can be updated. The cost function may consist of a sum of loss functions or weights such as Residual (difference between the output high-depth video and the reference high-depth video), PSNR (Peak Signal-to-noise ratio), MSE (Mean Squared Error), SSIM (Structural Similarity), and Perceptual Loss. Residual, PSNR, and MSE can be used to reduce the absolute error between the output high-depth video and the reference high-depth video. SSIM may be used to improve learning performance by reflecting structural features such as brightness and contrast. Perceptual Loss can be used to improve learning performance for fine details and human-perceived characteristics. Segmentation Loss may be further used as a loss function to improve the performance of the region segmentation unit 121. Segmentation Loss may use a Dice Coefficient formula to compare the region data output from the region segmentation unit 121 with the region segmentation labels of the reference high-depth image. The learning unit 130 can update the parameters of the learning model 120 using the error back propagation method. The backpropagated values may be adjusted by an optimization algorithm. For example, the search direction, learning rate, decay, momentum, etc., may be adjusted based on the previous state (backpropagated values and direction, etc.). This allows for optimization to make the learning direction more robust to noise and faster. As an optimization algorithm, Adam Optimizer, SGD (Stochastic Gradient Descent), AdaGrad, RMSProp, etc. may be used. Furthermore, batch normalization may be used to improve learning speed and robustness.
[0054] The learning unit 130 may train the learning model 120 until the value of the cost function decreases to a certain level or below during the learning process, or until a set epoch is reached.
[0055] Figure 8 shows a high-depth image generation device according to one embodiment of the present invention. The high-depth image generation device according to this embodiment may include a video preprocessing unit 110' that preprocesses stereo video and a learning model 120 that has been learned to output high-depth video from stereo video via the aforementioned high-depth image generation model learning device.
[0056] The video preprocessing unit 110' can perform video normalization as a preprocessing step for stereo video. The video normalization may be the same as the video normalization performed by the video preprocessing unit 110 in Figure 1.
[0057] The learning model 120 may include a region division unit 121 that divides cell regions from the input stereo image to generate region data, a depth estimation unit 122 that estimates depth from the input stereo image to generate depth data, and a high-depth image generation unit 123 that generates high-depth images from the input stereo image, the region data output from the region division unit 121, and the depth data output from the depth estimation unit 122. The high-depth image generation unit 123 can be implemented by a deep learning model learned through the high-depth image generation model learning device described above. The region division unit 121 or the depth estimation unit 122 may also be implemented by a deep learning model learned through the high-depth image generation model learning device described above.
[0058] Figure 9 shows an example of cell region segmentation video, region data, and depth data obtained from stereo video according to an embodiment of the present invention. Referring to Figure 9, region data c obtained by combining video b1, in which cell regions are segmented from left video a1, and video b2, in which cell regions are segmented from right video a2, is shown, as well as depth data d generated from left video a1 and right video a2.
[0059] Figure 10 shows an example of a high-depth image obtained from stereo video, region data, and depth data according to an embodiment of the present invention. Referring to Figure 10, when the left video (3 channels), right video (3 channels), region data (1 channel), and depth data (1 channel) are input to the high-depth image generation unit 123, the high-depth image generation unit 123 outputs a high-depth image (3 channels) via a deep learning model.
[0060] The apparatus according to an embodiment of the present invention may include a processor, memory for storing and executing program data, permanent storage such as a disk drive, a communication port for communicating with external devices, and user interface devices such as a touch panel, keys, and buttons. The method embodied in the software module or algorithm may be stored on a computer-readable recording medium as computer-readable code or program instructions that can be executed on the processor. Here, computer-readable recording media include magnetic storage media (e.g., ROM (read-only memory), RAM (random-access memory), floppy disks, hard disks, etc.) and optical reading media (e.g., CD-ROM, DVD (Digital Versatile Disc), etc.). The computer-readable recording media may be distributed across a network of computer systems, and computer-readable code may be stored and executed in a distributed manner. The medium may be read by a computer, stored in memory, and executed by the processor.
[0061] Embodiments of the present invention can be represented as functional block configurations and various processing stages. Such functional blocks may be embodied in various numbers of hardware and / or software configurations that perform specific functions. For example, embodiments may employ integrated circuit configurations such as memory, processing, logic, and look-up tables, which can perform various functions under the control of one or more microprocessors or other control devices. Similar to how the components of the present invention can be executed by software programming or software elements, embodiments may include various algorithms embodied in combinations of data structures, processes, routines, or other programming configurations, which may be embodied in programming or scripting languages such as C, C++, Java, and assembler. Functional aspects may be embodied in algorithms executed by one or more processors. Embodiments may also employ prior art for electronic environment setup, signal processing, and / or data processing. Terms such as “mechanism,” “element,” “means,” and “configuration” can be used broadly and are not limited to mechanical and physical configurations. The aforementioned term may also include the meaning of a series of software processes (routines) that work in conjunction with a processor or other components.
[0062] The specific executions described in the embodiments are one embodiment and do not in any way limit the scope of the embodiments. For the sake of brevity of the specification, descriptions of conventional electronic configurations, control systems, software, and other functional aspects of said systems may be omitted. Furthermore, the lines or connecting members between components shown in the drawings illustrate functional connections and / or physical or circuit connections and may be substituted or shown as various additional functional connections, physical connections, or circuit connections in actual devices. Also, components may not be necessary for applying the invention unless specifically mentioned as "essential" or "important."
[0063] The present invention has been described above, focusing on preferred embodiments. Those with ordinary skill in the art to which the present invention pertains will understand that the present invention may be embodied in modified forms that do not depart from its essential characteristics. Therefore, the disclosed embodiments should be considered in an explanatory rather than restrictive manner. The scope of the present invention is defined in the claims rather than in the above description, and all differences within an equivalent scope should be construed as being included within the scope of the present invention. [Explanation of Symbols]
[0064] 110...Video pre-processing unit 110'...Video pre-processing unit 120...Learning Model 121...Area division part 122... Depth estimation unit 123...High depth image generation section 130…Learning Department
Claims
1. A region division unit that divides the region of the tissue or cell located within the respective depth ranges of the left and right stereo images of which images of tissue or cells have different depth ranges, and generates region data that separates the region of the tissue or cell from other regions in the stereo image, such that the respective regions of the tissue or cell are included at the positions of the regions of the tissue or cell divided in the left image and the regions of the tissue or cell divided in the right image; A depth estimation unit that estimates the depth of each tissue or cell region in relation to the stereo image and generates depth data for each tissue or cell region; and A high-depth image generation unit generates a high-depth image with a wider depth range of focus than each individual image constituting the stereo image, from the stereo image, the region data, and the depth data. Includes, The depth-of-depth image generation unit is a depth-of-depth image generation device that generates the depth-of-depth image including the region of tissue or cells divided in the left image and the region of tissue or cells divided in the right image.
2. The high-depth image generation device according to claim 1, wherein the high-depth image generation unit generates the high-depth image using a trained deep learning model.
3. The high-depth video generation apparatus according to claim 2, wherein the region division unit or the depth estimation unit generates the region data or the depth data using a trained deep learning model.
4. The high-depth image generation apparatus according to claim 2, wherein the learned deep learning model is implemented to mimic blind deconvolution using a point-spread function.
5. A region segmentation process that divides the region of the tissue or cell located within the respective depth ranges of the left and right stereo images of which images of tissue or cells have different depth ranges, and generates region data that separates the region of the tissue or cell from other regions in the stereo image so that the respective regions of the tissue or cell are included at the locations of the regions of the tissue or cell divided in the left image and the regions of the tissue or cell divided in the right image; A depth estimation process that estimates the depth of each tissue or cell region in relation to the stereo image and generates depth data for each tissue or cell region; A high-depth image generation process that generates a high-depth image with a wider depth range of focus than each individual image constituting the stereo image, from the stereo image, the region data, and the depth data. Includes, The depth-of-depth image generation process is a method for generating a depth-of-depth image that includes the region of tissue or cells divided in the left image and the region of tissue or cells divided in the right image.
6. The method for generating high-depth images according to claim 5, wherein the high-depth image generation process utilizes a trained deep learning model.
7. The method for generating high-depth images according to claim 6, wherein the region segmentation process or the depth estimation process utilizes a trained deep learning model.
8. The method for generating high-depth images according to claim 6, wherein the trained deep learning model is implemented to mimic blind deconvolution using a point-spread function.
9. A learning model embodied to output a depth-of-depth image from a stereo image of an input tissue or cell, comprising: a region division unit that divides the regions of the tissue or cell located within the respective depth ranges of the left and right images of the stereo image, each having different depth ranges, and generates region data that separates the regions of the tissue or cell from other regions in the stereo image, such that each region of the tissue or cell is included at the respective positions of the regions of the tissue or cell divided in the left image and the regions of the tissue or cell divided in the right image; a depth estimation unit that estimates the depth for each region of the tissue or cell in the stereo image and generates depth data for the regions of the tissue or cell; and a depth-of-depth image generation unit that generates a depth-of-depth image with a wider depth range in focus than each image constituting the stereo image from the stereo image, the region data, and the depth data; and A learning unit that trains the learning model using training data including stereo video and corresponding reference high-depth video. Includes, The depth-of-depth image generation unit is a depth-of-depth image generation model learning device that generates the depth-of-depth image including the tissue or cell region divided in the left image and the tissue or cell region divided in the right image.
10. The learning unit calculates a cost function from the high-depth video output from the learning model and the reference high-depth video, and uses the cost function to train the learning model, as described in claim 9.
11. The high-depth image generation model learning device according to claim 9, wherein the high-depth image generation unit generates the high-depth image using a deep learning model.
12. The high-depth image generation model learning apparatus according to claim 11, wherein the region division unit or the depth estimation unit generates the region data or the depth data using a deep learning model.
13. The deep learning model is implemented to mimic blind deconvolution using a point-spread function, as described in claim 11, for the high-depth image generation model learning device.
14. The high-depth video generation model learning apparatus according to claim 9, further comprising a video preprocessing unit that preprocesses the stereo video and inputs it to the learning model.