Focal position estimation method, focal position estimation program, focal position estimation system, model generation method, model generation program, model generation system, and focal position estimation model
Patent Information
- Application Number
- JP2023008567
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-01-24
- Publication Date
- 2026-08-27
- Estimated Expiration
- 2043-01-24
AI Technical Summary
【0031】 本発明によれば、画像に基づく焦点位置の推定を、画像の位置に応じて行うことができる。
Smart Images

Figure 0007911974000005 
Figure 0007911974000006 
Figure 0007911974000007
Abstract
Description
[Technical Field]
[0001] The present invention relates to a focal position estimation method, a focal position estimation program, and a focal position system for estimating the focal position at the time of focus corresponding to an image to be estimated; a model generation method, a model generation program, and a model generation system for generating a focal position estimation model used for estimating the focal position at the time of focus; and a focal position estimation model. [Background technology]
[0002] Conventionally, virtual slide scanners have been used, which use images obtained by scanning a glass slide as virtual microscope images. In such devices, it is necessary to perform imaging with the focal point aligned with the sample. In contrast, it has been proposed to estimate the appropriate focal point based on the image of the sample. For example, Patent Document 1 shows estimation using a machine learning algorithm. [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2013-50713 [Overview of the project] [Problems that the invention aims to solve]
[0004] In the conventional method described in Patent Document 1, one focal position is estimated for the entire image. However, if the object, such as a sample, in the image has irregularities or is tilted, the focal position (Z position) in which it is in focus will differ depending on its position (XY position) in the image. In such cases, the single focal position for the entire image estimated by the conventional method described above is not necessarily appropriate.
[0005] The present invention has been made in view of the above, and aims to provide a focal position estimation method, a focal position estimation program, a focal position estimation system, a model generation method, a model generation program, a model generation system, and a focal position estimation model that can estimate the focal position based on an image according to the position of the image. [Means for solving the problem]
[0006] To achieve the above objective, the focal position estimation method according to the present invention is a focal position estimation method for estimating the focal position at the time of focus corresponding to an image to be estimated, and includes an image acquisition step for acquiring an image to be estimated, and a focal position estimation step for estimating the focal position at the time of focus corresponding to an image to be estimated, from an image to be estimated acquired in the image acquisition step, using a focal position estimation model that is generated by training machine learning and outputs information indicating the focal position at the time of focus according to the position in the image as input to information based on the image.
[0007] In the focal position estimation method according to the present invention, a focal position estimation model is used to estimate the focal position at the time of focus according to the position in the target image. Thus, according to the focal position estimation method according to the present invention, the focal position based on an image can be estimated according to the position in the image.
[0008] In a focal position estimation method, the focal position estimation model may be generated by a learning image acquisition step, which involves acquiring multiple training images of the same object with different focal positions, each associated with a focal position, and focal position information indicating the focal position when the multiple training images are in focus; a learning focus information generation step, which involves inputting the information based on each of the multiple training images acquired in the learning image acquisition step into a focal position estimation model under training, performing calculations according to the focal position estimation model to acquire information indicating the focal position when the multiple training images are in focus according to their position, and generating learning focus information indicating the focal position when the multiple training images are in focus according to their position in the image used for machine learning training, from the acquired information and the focal position information; and a learning step, which involves training machine learning to generate a focal position estimation model using the information based on each of the multiple training images acquired in the learning image acquisition step and the learning focus information corresponding to each of the multiple training images generated in the learning focus information generation step. In this configuration, training focus information is generated, and a machine learning model is trained to create a focal position estimation model. This model is then used to estimate the focal position at the time of focus, corresponding to the position in the target image. This allows for accurate and reliable estimation of the focal position based on the image.
[0009] The focus position estimation model may be generated in the training focus information generation step by calculating a single common focus position for all training images, corresponding to the position in each of the training images, from the focal position at the time of focus, which corresponds to the position in each of the training images, based on the information obtained using the focus position estimation model during training. Then, training focus information is generated for each of the training images from this single focal position. This configuration allows for more appropriate and reliable estimation of the focus position based on images.
[0010] In the position estimation step, a feature output model is used that takes image-based information as input and outputs the feature quantities of the image to be input to the focus position estimation model. The feature quantities of the target image are obtained from the target image acquired in the estimation target image acquisition step, and the focus position at the time of focus corresponding to the target image and its position is estimated from these feature quantities. The feature output model may be generated in the training step by creating two different feature training images corresponding to multiple training images, each with its own corresponding focus position, based on information indicating the focus position at the time of focus corresponding to the position of each training image, obtained using the focus position estimation model during training. The combination of these two feature training images is treated as a single unit, and the feature quantities of the two feature training images are compared according to the focus position associated with the two feature training images. Machine learning training is then performed based on the comparison results. With this configuration, the feature output model is used to obtain the feature quantities of the target image, and the focus position at the time of focus corresponding to the position of the target image is estimated from these feature quantities. By using feature-based estimation, the estimation of focal point based on images can be made more accurate and reliable.
[0011] The feature output model may be generated by training the machine learning model in the training step such that, when two feature training images relate to the same focal position, the difference in features between the two feature training images is small, and when the two feature training images relate to different focal positions, the difference in features between the two feature training images is large. This configuration allows for more appropriate and reliable estimation of focal positions based on images.
[0012] In the focal position estimation step, it may be possible to estimate the inclination of the imaging object captured in the estimated target image from the focal position at the time of focusing according to the position in the estimated target image. According to this configuration, the inclination of the imaging object captured in the estimated target image can be appropriately estimated.
[0013] In the focal position estimation step, it may be possible to control the focal position at the time of imaging of the imaging object captured in the estimated target image based on the focal position at the time of focusing according to the position in the estimated target image. According to this configuration, the imaging object can be appropriately imaged. For example, an image with focus at all positions can be obtained.
[0014] In the focal position estimation step, it may be possible to output information indicating the focus state according to the position in the estimated target image based on the focal position at the time of focusing according to the position in the estimated target image. According to this configuration, the focus state according to the position in the estimated target image can be grasped. For example, the positions with focus and the positions without focus in the estimated target image can be grasped.
[0015] In the estimated target image acquisition step, a plurality of estimated target images with different focal positions related to the same imaging object are acquired, and in the focal position estimation step, from at least one of the plurality of estimated target images acquired in the estimated target image acquisition step, the focal position at the time of focusing according to the position in the estimated target image is estimated, and based on the estimated focal position, it may be possible to generate one image from the plurality of estimated target images. According to this configuration, an appropriate image can be obtained. For example, an image with focus at all positions can be obtained.
[0016] By the way, in addition to being described as an invention of the focal position estimation method as described above, the present invention can also be described as an invention of a focal position estimation program and a focal position estimation system as follows. These are only different in category and are substantially the same invention, and exhibit the same actions and effects.
[0017] The focus position estimation program according to the present invention is a focus position estimation program that causes a computer to function as a focus position estimation system for estimating a focus position at the time of focusing corresponding to an estimation target image, and causes the computer to function as an estimation target image acquisition means for acquiring an estimation target image, and uses a focus position estimation model that is generated by training machine learning and inputs information based on an image and outputs information indicating a focus position at the time of focusing corresponding to the position in the image, to estimate from the estimation target image acquired by the estimation target image acquisition means a focus position at the time of focusing corresponding to the estimation target image and corresponding to the position in the estimation target image as focus position estimation means.
[0018] In the focus position estimation program, the focus position estimation model includes a learning image acquisition step of acquiring a plurality of learning images with different focus positions related to the same imaging object, each associated with a focus position, and focus position information indicating the focus position at the time of focusing for the plurality of learning images; an information based on each of the plurality of learning images acquired in the learning image acquisition step is input into the focus position estimation model during training, and an operation corresponding to the focus position estimation model is performed to obtain information indicating the focus position at the time of focusing corresponding to the position in each of the plurality of learning images, and from the obtained information and the focus position information, for each of the plurality of learning images, a learning focus information indicating the focus position at the time of focusing corresponding to the position in the image used for training machine learning is generated as a learning focus information generation step; and a learning step of performing machine learning training to generate a focus position estimation model using the information based on each of the plurality of learning images acquired in the learning image acquisition step and the learning focus information corresponding to each of the plurality of learning images generated in the learning focus information generation step. It may be generated by the above.
[0019] The focal position estimation system according to the present invention is a focal position estimation system that estimates the focal position at the time of focus corresponding to an image to be estimated, comprising: an image acquisition means for acquiring an image to be estimated; and a focal position estimation means that uses a focal position estimation model, which is generated by training machine learning and outputs information indicating the focal position at the time of focus according to the position in the image as input to information based on the image, to estimate the focal position at the time of focus corresponding to the image to be estimated and according to the position in the image, from the image to be estimated acquired by the image acquisition means.
[0020] In a focus position estimation system, the focus position estimation model may be generated by a learning image acquisition step, which involves acquiring multiple training images of the same object with different focal positions, each associated with a focal position, and focus position information indicating the focal position when the multiple training images are in focus; a learning focus information generation step, which involves inputting the information based on each of the multiple training images acquired in the learning image acquisition step into a focus position estimation model under training, performing calculations according to the focus position estimation model to acquire information indicating the focal position when the multiple training images are in focus according to their position, and generating learning focus information indicating the focal position when the multiple training images are in focus according to their position in the image used for machine learning training, from the acquired information and focus position information; and a learning step, which involves training machine learning to generate a focus position estimation model using the information based on each of the multiple training images acquired in the learning image acquisition step and the learning focus information corresponding to each of the multiple training images generated in the learning focus information generation step.
[0021] To achieve the above objective, the model generation method according to the present invention is a model generation method for generating a focal position estimation model that takes image-based information as input and outputs information indicating the focal position at the time of focus according to the position in the image, and includes: a learning image acquisition step of acquiring a plurality of learning images of different focal positions relating to the same image target object, each with an associated focal position, and focus position information indicating the focal position at the time of focus for the plurality of learning images; a learning focus information generation step of inputting the information based on each of the plurality of learning images acquired in the learning image acquisition step into a focus position estimation model under training and performing calculations according to the focus position estimation model to acquire information indicating the focal position at the time of focus according to the position in each of the plurality of learning images, and generating learning focus information indicating the focal position at the time of focus according to the position in the image used for training machine learning for each of the plurality of learning images from the acquired information and focus position information; and a learning step of performing machine learning training to generate a focal position estimation model using the information based on each of the plurality of learning images acquired in the learning image acquisition step and the learning focus information corresponding to each of the plurality of learning images generated in the learning focus information generation step.
[0022] According to the model generation method of the present invention, learning focus information is generated, machine learning is trained, and a focus position estimation model is generated. The generated focus position estimation model is used to estimate the focal position at the time of focus according to the position in the target image. Thus, according to the model generation method of the present invention, the focal position estimation based on an image can be performed according to the position in the image.
[0023] In the learning focus information generation step, the learning focus information may be generated for each of the multiple training images by calculating a single focal position common to all of the multiple training images, corresponding to the position in each training image, from the focal position at the time of focus, which corresponds to the position in each training image, as indicated by the information obtained using the focal position estimation model during training. This configuration generates a focal position estimation model that can estimate the focal position based on images more appropriately and reliably.
[0024] In the learning step, information based on an image is input to generate a feature output model that outputs the features of the image to be input to the focus position estimation model. Alternatively, in the learning step, based on information indicating the focal position at the time of focus according to the position in each of the multiple training images obtained using the focus position estimation model during training, two different feature training images are generated corresponding to the multiple training images, each with its own associated focal position. The combination of these two feature training images is treated as a single unit, and the features of these two feature training images are compared according to the focal position associated with the two feature training images. Machine learning is then performed based on the comparison results to generate a feature output model. With this configuration, a feature output model is generated that outputs features used to estimate the focal position at the time of focus according to the position in the target image. By performing estimation using these features, the estimation of the focal position based on the image can be made more appropriate and reliable.
[0025] In the learning step, if the two feature training images relate to the same focal position, the machine learning model may be trained so that the difference in features between the two feature training images is small, and if the two feature training images relate to different focal positions, the difference in features between the two feature training images is large. This configuration allows for more appropriate and reliable estimation of focal positions based on images.
[0026] Incidentally, in addition to being described as an invention of a model generation method as described above, the present invention can also be described as an invention of a model generation program and a model generation system as follows. These are substantially the same invention, differing only in category, and produce similar functions and effects.
[0027] The model generation program according to the present invention is a model generation program that causes a computer to function as a model generation system that takes image-based information as input and generates a focal position estimation model that outputs information indicating the focal position at the time of focus according to the position in the image, wherein the computer functions as: a learning image acquisition means that acquires a plurality of training images of the same image target object with different focal positions, each with a corresponding focal position, and focal position information indicating the focal position at the time of focus for the plurality of training images; a learning focus information generation means that inputs the information based on each of the plurality of training images acquired by the learning image acquisition means into a focal position estimation model in the process of training, performs calculations according to the focal position estimation model to acquire information indicating the focal position at the time of focus according to the position in each of the plurality of training images, and generates learning focus information indicating the focal position at the time of focus according to the position in the image used for machine learning training for each of the plurality of training images from the acquired information and the focal position information; and a learning means that performs machine learning training to generate a focal position estimation model using the information based on each of the plurality of training images acquired by the learning image acquisition means and the learning focus information corresponding to each of the plurality of training images generated by the learning focus information generation means.
[0028] The model generation system according to the present invention is a model generation system that takes image-based information as input and generates a focal position estimation model that outputs information indicating the focal position at the time of focus according to the position in the image, and includes: a learning image acquisition means that acquires a plurality of training images of the same image target object with different focal positions, each of which has an associated focal position, and focus position information indicating the focal position at the time of focus for the plurality of training images; a learning focus information generation means that inputs the information based on each of the plurality of training images acquired by the learning image acquisition means into a focal position estimation model in the process of training, performs calculations according to the focal position estimation model, acquires information indicating the focal position at the time of focus according to the position in each of the plurality of training images, and generates learning focus information indicating the focal position at the time of focus according to the position in the image used for machine learning training for each of the plurality of training images from the acquired information and focus position information; and a learning means that performs machine learning training to generate a focal position estimation model using the information based on each of the plurality of training images acquired by the learning image acquisition means and the learning focus information corresponding to each of the plurality of training images generated by the learning focus information generation means.
[0029] Furthermore, the focal position estimation model according to the present invention itself is an invention with a novel configuration. That is, the focal position estimation model according to the present invention is generated by machine learning training and is a focal position estimation model that functions a computer to take image-based information as input and output information indicating the focal position at the time of focus according to the position in the image.
[0030] The focal position estimation model may be generated by a learning image acquisition step, which involves acquiring multiple training images of the same object with different focal positions, each associated with a specific focal position, and focal position information indicating the focal position when the multiple training images are in focus; a learning focus information generation step, which involves inputting the information based on each of the multiple training images acquired in the learning image acquisition step into a focal position estimation model in the process of training, performing calculations according to the focal position estimation model to acquire information indicating the focal position when the multiple training images are in focus according to their position, and generating learning focus information indicating the focal position when the multiple training images are in focus according to their position in the image used for machine learning training, from the acquired information and the focal position information; and a learning step, which involves training machine learning to generate a focal position estimation model using the information based on each of the multiple training images acquired in the learning image acquisition step and the learning focus information corresponding to each of the multiple training images generated in the learning focus information generation step. [Effects of the Invention]
[0031] According to the present invention, the focal position can be estimated based on an image, depending on the position in the image. [Brief explanation of the drawing]
[0032] [Figure 1] This figure shows the configuration of a focal position estimation system and a model generation system according to an embodiment of the present invention. [Figure 2] This figure shows an example of estimation in a focus position estimation system. [Figure 3] This figure shows an example of a specific method of estimation in a focus position estimation system. [Figure 4] This diagram schematically shows a focal position map for an image. [Figure 5] This is a focal position map used to explain the estimation of the tilt of the object being imaged. [Figure 6]This is a diagram to explain the estimation of the tilt of the object being imaged. [Figure 7] This is a diagram to explain the estimation of the tilt of the object being imaged. [Figure 8] This is a diagram to explain the estimation of the tilt of the object being imaged. [Figure 9] This is a diagram illustrating imaging using estimation results. [Figure 10] This is a diagram illustrating imaging using estimation results. [Figure 11] This is an example of an image showing a focused state. [Figure 12] This figure shows an example of generating a single image from multiple target images. [Figure 13] This figure shows the focal position map of the training images and the estimation results. [Figure 14] This is a diagram showing the extraction of images used for training. [Figure 15] This figure shows the generation of a collection map from the focal position map of the estimation results. [Figure 16] This figure shows the generation of a focal position map for training from a collection map. [Figure 17] This figure shows the iterative training of encoding and decoding. [Figure 18] This diagram illustrates the generation of an encoder through machine learning training. [Figure 19] This diagram illustrates the generation of an encoder through machine learning training. [Figure 20] This diagram illustrates the generation of decoders through machine learning training. [Figure 21] This flowchart shows a focal position estimation method, which is a process performed in a focal position estimation system according to an embodiment of the present invention. [Figure 22] This flowchart shows a model generation method, which is a process performed by the model generation system according to an embodiment of the present invention. [Figure 23] This figure shows examples of the focal point position estimated for each pixel position when in focus. [Figure 24] This graph shows the focal position of the estimated target image and the focal position at the time of focus estimated from that target image. [Figure 25] This figure shows an example of a two-dimensional display of the feature quantities at the output of each block of the encoder. [Figure 26] This figure shows a specific example of the estimation performed according to this embodiment. [Figure 27] This figure shows a specific example of the estimation performed according to this embodiment. [Figure 28] This figure shows a specific example of the estimation performed according to this embodiment. [Figure 29] This figure shows a specific example of the estimation performed according to this embodiment. [Figure 30] This figure shows the configuration of the focus position estimation program and the model generation program according to an embodiment of the present invention, along with the recording medium. [Modes for carrying out the invention]
[0033] Hereinafter, embodiments of the focal position estimation method, focal position estimation program, focal position estimation system, model generation method, model generation program, model generation system, and focal position estimation model according to the present invention will be described in detail with reference to the drawings. In the description of the drawings, the same elements are denoted by the same reference numerals, and redundant explanations are omitted.
[0034] Figure 1 shows a computer 10, which is a focal position estimation system and model generation program according to this embodiment. The computer 10 functionally includes a focal position estimation system 20 and a model generation system 30. The focal position estimation system 20 is a system (device) that estimates the focal position at the time of focus corresponding to the target image. The model generation system 30 is a system (device) that takes image-based information as input and generates a focal position estimation model that outputs information indicating the focal position at the time of focus according to the position in the image. The focal position estimation model is used for estimation by the focal position estimation system 20.
[0035] As shown in Figure 1, the computer 10 is connected to an imaging device 40. The imaging device 40 is a device that obtains images by taking images, etc. The computer 10 acquires the images obtained by the imaging device 40. The focal position estimation system 20 and the model generation system 30 perform processing using the images acquired from the imaging device 40.
[0036] The imaging device 40 may be included in an inspection device for inspecting devices such as semiconductor devices. Alternatively, the imaging device 40 may be an observation device that images a biological sample placed on a slide glass and observes the image of the imaged biological sample. In this case, the image obtained by imaging with the imaging device 40 becomes, for example, an image for realizing a virtual microscope. Furthermore, the sample is not limited to devices such as semiconductor devices or biological samples placed on slide glass; the imaging device 40 may be a microscope device used for other purposes. Conventional imaging devices 40 can be used as the imaging device 40 itself. In addition, the imaging device 40 may have a function that allows it to be controlled from the focus position estimation system 20, as will be described later.
[0037] According to the functions of the computer 10, the focal position at the time of focusing corresponding to the image obtained by imaging by the imaging device 40 is estimated. For example, this estimation allows the imaging device 40 to take another image with the object being imaged in focus. Alternatively, it is possible to determine whether the image was taken with the object being imaged in focus.
[0038] Furthermore, the computer 10 only needs to be able to acquire images to be processed by the focal position estimation system 20 and the model generation system 30, and does not need to acquire images directly from the imaging device 40.
[0039] Computer 10 is a conventional computer that includes hardware such as a processor (CPU, Central Processing Unit), memory, and communication modules. Computer 10 may also be a computer system comprising multiple computers. Furthermore, Computer 10 may be configured using cloud computing. The various functions of Computer 10, described later, are performed by the operation of these components through programs, etc. Computer 10 and the imaging device 40 are connected to each other to enable the transmission and reception of information. In Figure 1, the focus position estimation system 20 and the model generation system 30 are implemented by the same computer 10, but they may be implemented by separate computers 10.
[0040] Next, the functions of the focus position estimation system 20 and the model generation system 30 included in the computer 10 according to this embodiment will be described. As shown in Figure 1, the focus position estimation system 20 is configured to include an estimation target image acquisition unit 21 and a focus position estimation unit 22.
[0041] The estimated target image acquisition unit 21 is an estimated target image acquisition means for acquiring the estimated target image. For example, the estimated target image acquisition unit 21 receives and acquires an image captured by the imaging device 40 from the imaging device 40. The estimated target image acquisition unit 21 divides the image acquired from the imaging device 40 into sizes that can be estimated by the focal position estimation unit 22, and each of the divided images is used as the estimated target image. The sizes that can be estimated by the focal position estimation unit 22 will be described later. Note that the acquisition of the estimated target image does not have to be done by the method described above, and may be done by any other method. The estimated target image acquisition unit 21 outputs the acquired estimated target image to the focal position estimation unit 22.
[0042] The focus position estimation unit 22 is a focus position estimation means that uses a focus position estimation model to estimate the focal position at the time of focus corresponding to the target image and its position in the target image, from the target image acquired by the target image acquisition unit 21. The focus position estimation model is generated by machine learning training and is a model that takes image-based information as input and outputs information indicating the focal position at the time of focus corresponding to its position in the image. The focus position estimation unit 22 may also use a feature output model to acquire feature quantities of the target image acquired by the target image acquisition unit 21, and use the focus position estimation model to estimate the focal position at the time of focus corresponding to the target image and its position in the target image from these feature quantities. The feature output model is a model that takes image-based information as input and outputs the feature quantities of the image that are input to the focus position estimation model.
[0043] For example, the focal position estimation unit 22 corresponds to the target image and estimates the focal position at the time of focus according to the position in the target image, as follows. In this embodiment, the focal position estimation unit 22 estimates the focal position at the time of focus for each pixel of the target image as the focal position at the time of focus according to the position in the target image. Figure 2 shows an example of the target image 50 and a focal position map 60 that shows the focal position at the time of focus estimated from the target image 50. The focal position map 60 is data that has the focal position at the time of focus for each pixel of the target image 50. Since the focal position map 60 has information corresponding to each pixel of the target image 50, it can be an image of the same size as the target image 50.
[0044] The focal position map 60 has a value (information) for each pixel that indicates the focal position at the time of focus at the position of that pixel (for example, the XY position of that pixel). For example, the value of each pixel in the focal position map 60 indicates the direction (direction in the depth direction or towards the viewer) and distance from the focal position when the estimated target image 50 was captured to the focal position when it was in focus. This value is, for example, the value obtained by subtracting the distance corresponding to the position of the object being captured when the estimated target image 50 was captured (the distance from the position of the lens, such as the objective lens, to the position of the object being captured when the estimated target image 50 was captured) from the distance corresponding to the position of the object being captured when it was in focus (for example, the distance from the position of the lens, such as the objective lens, to the position of the object being captured when it was in focus, which corresponds to the focal length of the lens).
[0045] In other words, in this case, the value represents the focal position at the time of focus in a coordinate system where the focal position at the time the estimated target image 50 was captured is set to 0. If the distance to the position of the object being photographed when the estimated target image 50 was captured is longer than the distance to the position of the object being photographed at the time of focus, the value will be negative. If the distance to the position of the object being photographed when the estimated target image 50 was captured is shorter than the distance to the position of the object being photographed at the time of focus, the value will be positive. The focal position at the time of focus is the position where the object being photographed in the estimated target image 50 is in focus. The distance to the object being photographed at the time of focus refers to the distance from the position of the lens, such as the objective lens, to the position of the object being photographed when the object being photographed in the estimated target image 50 is in focus, and generally corresponds to the focal length of the lens. By changing the position of the lens, etc., by the above difference from the position of the lens, etc., when the estimated target image 50 was captured, an image in focus on the object being photographed can be captured. The focal position map 60 shown in Figure 2 represents the numerical value of each pixel in terms of color intensity.
[0046] The values of the focal position map 60, i.e., the imaging direction values, may be values in a predetermined unit. For example, a unit length may be predetermined (e.g., 50 μm), and the value may be set to a unit length of 1. In the following examples as well, imaging direction values such as the value indicating the focal position will be shown using this numerical value.
[0047] In conventional methods, as described above, information regarding one focal position is estimated for the entire image. In contrast, in this embodiment, as described above, the focal position at the time of focus (e.g., Z position) is estimated from a single target image 50, corresponding to the position (e.g., XY position) in the target image 50. This allows for the estimation of an appropriate focal position at the time of focus, corresponding to the position in the target image 50, even if the object being imaged, such as a sample, has irregularities or inclination. Note that the information indicating the estimated focal position at the time of focus does not necessarily have to be the focal position map 60 described above.
[0048] The focus position estimation unit 22 performs the above estimation using an encoder (feature extraction layer) 70, which is a feature output model, and a decoder (discrimination layer) 71, which is a focus position estimation model. The encoder 70 and decoder 71 are trained models generated by the model generation system 30. The focus position estimation unit 22 pre-inputs and stores the encoder 70 and decoder 71 generated by the model generation system 30 and uses them for estimation.
[0049] The encoder 70 is a model that takes image-based information as input and outputs feature quantities of the image that are input to the decoder 71. The feature quantities output from the decoder 71 are information that indicates the features of the input image. In this embodiment, these features reflect the focal position when the image was captured. That is, the encoder 70 is an optical model relating to optical features. These feature quantities are, for example, vectors with a predetermined number of dimensions (for example, 1024 dimensions).
[0050] The encoder 70 is configured, for example, to include a neural network. The neural network may be multi-layered. That is, the encoder 70 and decoder 71 may be generated by deep learning. The neural network may also be a convolutional neural network (CNN).
[0051] The input layer of the encoder 70 is provided with neurons for inputting image-based information. For example, the information input to the encoder 70 is the pixel value of each pixel in the image. In this case, the input layer is provided with a number of neurons equal to the number of pixels in the image, and the pixel value of the corresponding pixel is input to each neuron. The image related to the information input to the encoder 70 is an image of a predetermined size. This image size is the size of the image that can be estimated at one time by the focus position estimation unit 22.
[0052] The information input to the encoder 70 does not have to be the pixel values of each pixel, as long as it is based on an image. For example, this information may be feature quantities for input to the encoder 70, obtained by performing preprocessing such as conventional image processing on the image to reduce the influence of the imaging environment. By performing such preprocessing, it is possible to improve the efficiency of machine learning and the accuracy of the generated encoder 70.
[0053] The output layer of the encoder 70 is provided with neurons for outputting feature vectors. For example, neurons with a number of dimensions equal to the feature vector are provided.
[0054] The decoder 71 is a model that takes the feature quantities of the image output from the encoder 70 as input and outputs information indicating the focal position at the time of focus, corresponding to the position in the image. For example, the decoder 71 outputs a focal position map 60 as the estimated focal position at the time of focus. Alternatively, the decoder 71 may output information indicating the focal position itself at the time of focus for each pixel corresponding to the position in the target image 50 (for example, the distance corresponding to the focal position at the time of focus). Furthermore, candidate values for the above may be set in advance for each pixel corresponding to the position in the target image 50, and the decoder 71 may output a value indicating the degree to which each of these candidates is valid. For example, if the candidate values for the above are +1, 0, -1, -2, ..., the decoder 71 outputs a value indicating the degree to which each candidate is valid. For example, the candidate with the highest value may be set as the above value.
[0055] The decoder 71 is composed of, for example, a neural network. The neural network may be multi-layered. That is, the decoder 71 may be generated by deep learning. The neural network may also be a convolutional neural network (CNN).
[0056] The input layer of the decoder 71 is provided with neurons for inputting feature quantities. For example, the input layer is provided with neurons corresponding to the neurons provided in the output layer of the encoder 70. That is, the input layer is provided with the same number of neurons as the output layer of the encoder 70. The output layer of the decoder 71 is provided with neurons for outputting the estimation result of the focal position at the time of focus as described above. For example, the output layer is provided with the number of neurons equal to the number of pixels in the target image 50 for outputting the focal position map 60.
[0057] A more detailed example of the encoder 70 and decoder 71 will be described. The encoder 70 has 15 layers of neurons. In these layers, the encoder 70 halves the resolution and doubles the number of channels by max pooling every three blocks. Each block consists of Conv2d, batch normalization, and ReLU. The number of channels in each layer, from the input side, is 3, 64, 128, 256, 512, and 1024, respectively.
[0058] The decoder 71 has three layers of neurons in a fully connected layer. The decoder 71 receives the feature quantities output from the encoder 70 as input by connecting the output of the 1024-channel block (final layer).
[0059] Alternatively, decoder 71 has 12 layers + 1 layer of neurons. In these layers, decoder 71 doubles the resolution and halves the number of channels by upsampling every 3 blocks. Each block consists of Conv2d, batch normalization, and ReLU. The last layer is a 64-to-1 convolutional layer. Decoder 71 receives the feature vectors output from encoder 70 as input by connecting the block outputs of 1024, 512, and 256 channels.
[0060] Unlike this embodiment, if the focus position at a single point of focus is to be estimated from the image, rather than for each position in the image, for example, each pixel, then the encoder may consist of convolutional layers and the decoder may consist of fully connected layers. In that case, the encoded feature quantities will be the average value of the image region (which becomes the average space for each channel due to the global average pooling layer). Therefore, if the object being photographed in the image has irregularities or inclination, the feature quantities will be a mixture of features from different imaging directions (Z position) and will not show the correct features.
[0061] In this embodiment, the encoder 70 may not have a global average pooling layer and may output features without spatial averaging. This results in less feature mixing. For example, mixing is limited to the local region level of the convolution.
[0062] Note that the encoder 70 and decoder 71 may be composed of something other than a neural network.
[0063] The encoder 70 and decoder 71 may be for specific types of images. For example, the specific type of image may be an image in which radiation from an object is detected (an image used for light emission / heating analysis), an image in which light from an object is detected when light is shone on the object (an image used for pattern analysis), or an image in which the electrical characteristics of an object are detected when light is shone on the object (an image used for laser analysis). Alternatively, the type of image may be the type of object captured in the image.
[0064] In this case, when generating images through machine learning training of the encoder 70 and decoder 71, and when estimating the focal position at the time of focus using the encoder 70 and decoder 71, the specific type of image is used. Similarly, the encoder 70 and decoder 71 may be for a specific type of imaging device 40. In this way, by making the encoder 70 and decoder 71 for a specific type of image or a specific type of imaging device 40, more appropriate estimation can be performed according to the type of image or the type of imaging device 40. However, the encoder 70 and decoder 71 may be common to multiple types of images or types of imaging devices 40.
[0065] The encoder 70 and decoder 71 are intended to be used as program modules that are part of artificial intelligence software. The encoder 70 and decoder 71 are used, for example, in a computer equipped with a processor and memory, and the computer's processor operates according to instructions from a model stored in memory. For example, the computer's processor operates according to the instructions to input information to the model, perform calculations according to the model, and output a result from the model. Specifically, the computer's processor operates according to the instructions to input information to the input layer of a neural network, perform calculations based on parameters such as learning weight coefficients in the neural network, and output a result from the output layer of the neural network.
[0066] The focal position estimation unit 22 receives the target image 50 from the target image acquisition unit 21. The focal position estimation unit 22 inputs information based on the target image 50 to the encoder 70, performs calculations according to the encoder 70, and obtains feature quantities of the target image 50, which are the output from the encoder 70. The focal position estimation unit 22 inputs the obtained feature quantities to the decoder 71, performs calculations according to the decoder 71, and obtains a focal position map 60, which is the output from the decoder 71, as the estimated focal position at the time of focus according to the position in the target image 50.
[0067] As shown in Figure 3, the target image 50 may be multiple images (image patches) obtained by dividing the image 51 acquired from the imaging device 40. The focal position estimation unit 22 acquires a focal position map 60 from the multiple target images 50 using an encoder 70 and a decoder 71. The focal position estimation unit 22 may also generate a focal position map 61 for the image 51 acquired from the imaging device 40 by stitching together (tiling) the focal position maps 60 acquired from each target image 50.
[0068] Figure 4(a) schematically shows a focal position map 61 for an image 51 acquired from the imaging device 40 in this embodiment. Figure 4(b) schematically shows a focal position map in a comparative example of this embodiment, for example, when one focal position at the time of focus is estimated for each of the estimated target images 50 divided from the image 51 acquired from the imaging device 40. Figure 4 shows the focal position at the time of focus for each position of the object to be imaged 52 on a plane (XZ plane) viewed from the side (Y axis) in the imaging direction (Z axis direction).
[0069] In this embodiment, the focal position at the time of focus is estimated for each pixel of the image 51, so the tilt or distortion of the object being imaged 52 (sample) can be correctly grasped (measured). In this embodiment, the focal position at the time of focus in the imaging direction shown by the focal position map 60 is assumed to be the position of the object being imaged. For example, in the example shown in Figure 4(a), the warping of the object being imaged 52 increases as it moves away from the center, but this can be correctly grasped. On the other hand, in the example shown in Figure 4(b), the interval between the positions in the X direction where the focal position at the time of focus is estimated is large, so the tilt or distortion of the object being imaged 52 cannot be correctly grasped (measured). For example, in the example shown in Figure 4(b), if linear interpolation is performed from the estimated focal position at the time of focus, the resulting line will be off from the object being imaged 52.
[0070] The focal position estimation unit 22 may output the acquired focal position map 60, or a focal position map 61 for the image 51 acquired from the imaging device 40. For example, the focal position estimation unit 22 may output these in a format that is recognizable to the computer 10 user (e.g., display). Alternatively, the focal position estimation unit 22 may transmit these to another device or module.
[0071] Furthermore, the focal position estimation unit 22 may use the acquired focal position map 60, or the focal position map 61 for the image 51 acquired from the imaging device 40, for example, as follows.
[0072] The focal position estimation unit 22 may estimate the inclination of the object being imaged in the estimated target image 50 from the focal position at the time of focus, corresponding to the position in the estimated target image 50. For example, the focal position estimation unit 22 estimates the inclination of the object being imaged in the estimated target image 50 as follows: The focal position estimation unit 22 estimates the inclination of the object being imaged for each of the two coordinate axes, the X axis and the Y axis, which are parallel to each side of the focal position map 60 as shown in Figure 5. For example, the focal position estimation unit 22 estimates the angle θ1 of the inclination of the object being imaged in the X axis with respect to a plane 62 perpendicular to the imaging direction (Z axis direction), as shown in Figure 6, and the angle θ2 of the inclination of the object being imaged in the Y axis with respect to a plane 62 perpendicular to the imaging direction (Z axis direction), as shown in Figure 7.
[0073] Figure 6 shows information indicating the focal position at the time of focus for each pixel shown by the focal position map 60 (specifically, information regarding the focal position at the time of focus relative to the focal position when the image was captured) (these are values such as -2, 0, and 1 in the matrix). Figure 6 shows an example where the object being imaged is tilted in the X-axis direction, specifically, the right side is lower. Figure 7 shows an example where the object being imaged is tilted in the Y-axis direction, specifically, the front side is lower.
[0074] The focal position estimation unit 22 calculates angles θ1 and θ2 using the following formulas.
number
[0075] x and y are determined based on the positions of the pixels of the focal position map 60 as shown in FIG. 5. For example, x is the difference in the X-axis direction between two preset pixels P of the focal position map 60 that are separated from each other in the X-axis direction used for estimating the inclination. a , P b of the positions. That is, x = |P b - P a | (only the X-axis component). y is the difference in the Y-axis direction between two preset pixels P of the focal position map 60 that are separated from each other in the Y-axis direction used for estimating the inclination. a , P c of the positions. That is, y = |P c - P a | (only the X-axis component). The pixels P of the focal position map 60 for which the inclination is to be estimated a , P b , P c may be set arbitrarily.
[0076] z1 and z2 are calculated from the focal positions at the time of focusing corresponding to each pixel shown by the focal position map 60 as shown in FIG. 6. For example, z1 is the difference between the focal positions Z a , Z b at the time of focusing corresponding to the above two pixels P. That is, z1 = |Z a - Z b |. z1 is the difference between the focal positions Z b - Z a at the time of focusing corresponding to the above two pixels P a , P c . That is, z2 = |Z a - Z c |. c - Z a |.
[0077] The focal position estimation unit 22 may estimate the inclination for a plurality of positions of the imaging object (that is, the focal position map 60) and calculate a statistical value such as an average value or a median value. Thereby, the accuracy of the inclination of the imaging object to be estimated can be improved.
[0078] Alternatively, the focal position estimation unit 22 may estimate the inclination of the object being imaged as follows. The focal position estimation unit 22 calculates the least squares plane of the focal position at the time of focus from the focal position at the time of focus and the position of each pixel shown by the focal position map 60. The coordinate system used in this case is, for example, a coordinate system in which the plane of the image of the focal position map 60 is the XY plane and the direction perpendicular to that plane is the Z axis. The focal position estimation unit 22 calculates the normal vector n1=(a,b,c) of the calculated least squares plane shown in Figure 8. The focal position estimation unit 22 calculates the angle θ of the normal vector n1 of the least squares plane with respect to the normal vector n2=(0,0,1) of the XY plane as the inclination of the object being imaged using the following formula which is stored in advance.
number
[0079] If the focal position estimation unit 22 estimates the inclination of the object being imaged in the estimation target image 50 from the focal position map 60, it may estimate the inclination using a method other than those described above. Furthermore, the focal position estimation unit 22 may estimate an angle other than those described above as the inclination of the object being imaged.
[0080] The focal position estimation unit 22 may output information indicating the tilt of the estimated image target. For example, the focal position estimation unit 22 may output this information in a format that is recognizable to the user of the computer 10 (e.g., a display). Alternatively, the focal position estimation unit 22 may transmit this information to another device or module.
[0081] Furthermore, the focal position estimation unit 22 may use the estimated tilt of the object to be imaged to control the imaging device 40. In this case, the imaging device 40 is configured to control the tilt (orientation) of the object to be imaged, for example, as follows. The imaging device 40 has a mounting section, which is a member on which the object to be imaged is placed when imaging. The mounting section is configured so that the tilt of the mounting surface on which the object to be imaged is placed is variable with respect to the imaging direction. That is, the imaging device 40 can perform tilt correction of the object to be imaged. The imaging device 40 can be a conventional one that can control the tilt of the object to be imaged.
[0082] The focal position estimation unit 22 controls the imaging device 40 so that the estimated tilt of the object to be imaged is eliminated during imaging. Specifically, the focal position estimation unit 22 controls the imaging device 40 to tilt the object to be imaged in the opposite direction to the estimated tilt of the object to be imaged. The controlled imaging device 40 adjusts the tilt of the object to be imaged during imaging, for example by operating the mounting unit. In this way, the focal position estimation unit 22 controls the tilt correction in the imaging device 40. After tilt correction is performed, the imaging device 40 takes an image of the object to be imaged, so that an image of the object to be imaged with an appropriate tilt can be obtained. Alternatively, the user of the computer 10 may manually control the tilt correction by checking the information indicating the tilt of the object to be imaged output from the focal position estimation unit 22.
[0083] The focal position estimation unit 22 may control the focal position of the object being imaged in the estimated target image 50 when it is captured, based on the focal position at the time of focusing corresponding to the position in the estimated target image 50. This control is for when the object being imaged is captured again after the estimated target image 50 has been obtained. For example, the focal position estimation unit 22 controls the focal position of the object being imaged in the estimated target image 50 when it is captured, as follows.
[0084] As shown in Figure 9, the imaging device 40 images the object 52 on the slide glass 42, for example, using the objective lens 41. The parts indicated by the Z and Y axes in Figure 9 are the objective lens 41, slide glass 42, and object 52 as viewed from the side (X axis) in the imaging direction (Z axis direction), while the parts indicated by the X and Y axes are the slide glass 42 and object 52 as viewed from above in the imaging direction (Z axis direction).
[0085] The imaging device 40 is configured to control the focal position relative to the object to be imaged 52, for example, as follows: The imaging device 40 is configured so that the position of the objective lens 41 in the imaging direction, i.e., the height of the objective lens 41, is variable. The imaging device 40 is also configured so that the position of the plane (XY plane) perpendicular to the imaging direction (Z axis direction) is variable. Specifically, the objective lens 41 may be movable, or the mounting part on which the slide glass 42 is placed may be movable, or both may be present.
[0086] The focus position estimation unit 22 calculates the inclination of the object 52 as described above, from the focal position at the time of focus of each pixel position 52a of the object 52 corresponding to the position of each pixel in the estimated focal position map 60 (estimated target image 50) (for example, +2, which is the difference between the focal position at the time of focus and the focal position 0 when the estimated target image 50 was captured, as shown in Figure 9). The focus position estimation unit 22 calculates the focus plane 52b of the object 52 from the calculated inclination of the object 52 and the focal positions at the time of focus of each position 52a of the object 52. Note that the focus plane 52b does not take inclination. For example, the average of the focal positions at the time of focus of each position 52a of the object 52 may be used as the focus plane 52b.
[0087] The focus position estimation unit 22 controls the imaging position (imaging direction and position in the XY plane) of the imaging device 40 so that the object to be imaged 52 is imaged along the focus plane 52b. The controlled imaging device 40 scans the object to be imaged 52 and performs imaging to acquire a high-magnification image.
[0088] Alternatively, instead of calculating the focus plane 52b for the entire object 52 as described above, the focal position may be controlled by setting a partial region 52c (for example, the rectangular region shown in Figure 10) that divides the area of the object 52. In this case, the focus position estimation unit 22 calculates the focus plane 52d for each partial region 52c. The calculation of the focus plane 52d for each partial region 52c may be performed in the same manner as described above, or by other methods. For example, the focus plane 52d for the partial region 52c may be a plane (XY plane) perpendicular to the imaging direction (Z axis direction) that passes through the focal position when the center of the partial region 52c is in focus.
[0089] The focus position estimation unit 22 controls the imaging position (imaging direction and position in the XY plane) of the imaging device 40 so that imaging of the object to be imaged 52 is performed along the focus plane 52d for each sub-region 52c. The controlled imaging device 40 scans the object to be imaged 52 and performs imaging to acquire a high-magnification image.
[0090] The acquired images can be used in the same way as before. By using the focal position map 60 estimated when the target object 52 is imaged again in this way, an image in focus on the target object 52 can be acquired easily and quickly.
[0091] The focus position estimation unit 22 may output information indicating the focus state corresponding to the position in the estimated target image 50, based on the focal position at the time of focus corresponding to the position in the estimated target image 50. For example, the focus position estimation unit 22 outputs the following information.
[0092] As shown in Figure 11, multiple positions 50a indicating the in-focus state are set in advance on the target image 50. The positions 50a to be set are, for example, the grid positions shown in Figure 10. However, the positions 50a to be set can be any position. Also, the positions 50a may be an area having a certain range (focus determination area). The focal position estimation unit 22 refers to the focal position at the time of focus of the position 50a indicated by the focal position map 60 and determines the in-focus state for each set position 50a. The determined in-focus state is, for example, the degree to which the focal position at the time the target image 50 was captured matches the estimated focal position at the time of focus, that is, the degree to which the target image 50 is in focus at the set positions 50a.
[0093] For example, the focus position estimation unit 22 calculates a focus score indicating the degree of the above for each position 50a of the target image 50 (for example, the higher the focus score, the higher the degree of the above). The focus score for the in-focus state may be determined according to the proportion of the number of positions 50a where the focal position at the time of focus matches the focal position when the target image 50 was captured (for example, the number of pixels in the focal position map 60 whose value is 0). For example, the higher the proportion, the higher the focus score. Alternatively, it may be determined according to a weighted sum of the pixel values in the focal position map 60.
[0094] The focus position estimation unit 22 determines whether the image is in focus at each position 50a of the target image 50 based on the calculated focus score. The focus position estimation unit 22 determines that the image is in focus at position 50a if the focus score is equal to or greater than a preset threshold. The focus position estimation unit 22 determines that the image is out of focus at position 50a if the focus score is less than a preset threshold.
[0095] The focal position estimation unit 22 generates an image in which information indicating the above judgment result is superimposed on each position 50a of the target image 50. For example, as shown in Figure 11, the target image 50 is generated in which a green rectangle is superimposed on the in-focus position 50b and a red rectangle is superimposed on the out-of-focus position 50c. The focal position estimation unit 22 outputs the generated target image 50 as information indicating the focus state. The output can be performed in the same way as the output of the focal position map 60 described above.
[0096] By referring to the focus state corresponding to the position in the estimated target image 50 shown in this way, it is easy to determine whether the image has been acquired properly. Furthermore, if the position 50c is out of focus, imaging may be performed again. Note that the information indicating the focus state corresponding to the position in the estimated target image 50 may be other than that described above. Also, this information does not have to be superimposed on the estimated target image 50 as described above, and may be in any format.
[0097] As shown in Figure 12, a single image 53 may be generated from multiple estimated target images 50 of the same object at different focal positions. The generated image 53 is an image that is in focus (or nearly in focus) at each position within the image 53.
[0098] In this case, the estimated target image acquisition unit 21 acquires multiple estimated target images 50 of the same object at different focal positions, as shown in Figure 12. For example, multiple estimated target images 50 are obtained by the imaging device 40 fixing the position (XY) during imaging other than the imaging direction (Z axis direction) for the same object, and performing multiple consecutive imagings with different focal positions. In this case, as shown in Figure 12, the focal positions are set to differ at regular intervals (steps). The interval of the focal positions may be, for example, one unit interval in the pre-set units described above (the example shown in Figure 12 also uses this interval). Note that the interval of the focal positions for the multiple estimated target images 50 does not necessarily have to be a regular interval (step).
[0099] For example, as shown in Figure 12, multiple target images 50 with focal positions of +2, +1, 0, -1, and -2 are acquired. The focal position of 0 is a pre-set reference focal position, and the other focal positions indicate the deviation from the focal position of 0 in the pre-set units mentioned above. The target image acquisition unit 21 acquires information indicating the focal position when each target image 50 was captured (for example, the information regarding +2, +1, 0, -1, and -2 mentioned above) along with the multiple target images 50, and outputs it to the focal position estimation unit 22.
[0100] The focus position estimation unit 22 estimates the focal position at the time of focus according to the position in the estimated target image 50 from at least one of the multiple estimated target images 50 acquired by the estimation target image acquisition unit 21, and generates one image 53 from the multiple estimated target images 50 based on the estimated focal position.
[0101] For example, the focal position estimation unit 22 generates one image 53 from multiple target images 50 as follows: The focal position estimation unit 22 receives multiple target images 50 and information indicating the focal position when each target image 50 was captured from the target image acquisition unit 21. The focal position estimation unit 22 generates a focal position map 60 from the input multiple target images 50. For example, the focal position estimation unit 22 generates a focal position map 60 from one target image 50, for example, a target image 50 with a focal position of 0.
[0102] In theory, with multiple target images 50 that differ only in their focal positions, the same focal position map 60 will be generated regardless of which target image 50 is used. However, in actual calculations, it is possible that the focal position map 60 will differ slightly for each of the multiple target images 50. Therefore, it is also possible to generate a focal position map 60 from each of the multiple target images 50, and then take the average of these pixel-wise values to generate the focal position map 60 to be used in subsequent processing.
[0103] Next, the focal position estimation unit 22 extracts a pixel from the estimated target image 50 that corresponds to the focal position closest to the focal position at the time of focus indicated by the focal position map 60, for each pixel in the focal position map 60. The focal position estimation unit 22 extracts a pixel from one of the multiple estimated target images 50 as described above for all pixels. The focal position estimation unit 22 combines the pixels extracted from one of the multiple estimated target images 50 while maintaining the pixel positions to generate a single image 53.
[0104] The images 53 generated in this way, which are in focus (or nearly in focus) at each position, are used in the same way as before. For example, when using image 53 as a virtual microscope image, since image 53 is clear throughout, it becomes unnecessary to switch to an image that is in focus at each position, as was done in the past. This reduces the amount of data required to realize a virtual microscope. The above describes the functions of the focus position estimation system 20.
[0105] As shown in Figure 1, the model generation system 30 is configured to include a learning image acquisition unit 31, a learning focus information generation unit 32, and a learning unit 33.
[0106] The learning image acquisition unit 31 is a learning image acquisition means that acquires multiple learning images of the same object being imaged, each with a corresponding focal position, but at different focal positions, and focal position information indicating the focal position when the multiple learning images are in focus. The learning images and focal position information are used to generate the encoder 70 and decoder 71. For example, the learning image acquisition unit 31 acquires multiple learning images and focal position information as follows.
[0107] The learning image acquisition unit 31 acquires multiple learning images 80 of the same object being imaged, each with a different focal position, as shown in Figure 13. For example, the learning image acquisition unit 31 acquires images captured by the imaging device 40. The acquired images show the object being imaged for the learning images 80. The object being imaged for the learning images 80 may be something captured by the imaging device 40, or it may be something else. The learning images 80 are generated from images captured by the imaging device 40.
[0108] The images that form the basis for the multiple training images 80 are obtained by the imaging device 40 fixing the position (XY) during imaging other than the imaging direction (Z-axis direction) of the same object, and taking multiple consecutive images with different focal positions. In this process, the focal positions are made to differ at regular intervals (steps), as shown in the multiple training images 80 in Figure 13. The interval of the focal positions may be, for example, one unit interval in the pre-set units mentioned above (the example in Figure 13 also uses this interval). Note that the interval of the focal positions for the images that form the basis for the multiple training images 80 does not necessarily have to be a regular interval (step).
[0109] Each training image 80 corresponds to an estimated image 50 used as input to the encoder 70. As described above, the estimated image 50 input to the encoder 70 is not the entire image captured by the imaging device 40, but a part of that image. Therefore, as shown in Figure 14, the training image acquisition unit 31 extracts training images 80, which are partial images (image patches) of a preset size used as input to the encoder 70, from the image 81 captured by the imaging device 40. The extraction from the image 81 captured by the imaging device 40 is performed on the same location (XY) region of each image 81 with different focal positions. As shown in Figure 13, the multiple training images 80 extracted from the same location (XY) are considered a set for machine learning training, as described below. In this embodiment, multiple images captured from the same location (XY) with different focal positions are called a Z-stack. The training image acquisition unit 31 acquires a number of training images 80, which are Z-stacks, sufficient to appropriately generate the encoder 70 and decoder 71.
[0110] The position from which the training image 80 is extracted in the image 81 captured by the imaging device 40 is the portion in which the object to be imaged is visible. However, the training image 80 may include training images 80 in which the object to be imaged is not visible. The position from which the training image 80 is extracted in the image 81 captured by the imaging device 40 may be predetermined. Alternatively, the position from which the training image 80 is extracted may be determined by performing image recognition on the image 81 captured by the imaging device 40 and estimating that the object to be imaged is visible. When generating a single image 53 from the above-mentioned multiple estimated target images 50, the multiple estimated target images 50 may also be generated by extracting them from the image in the same way as multiple training images 80.
[0111] The learning image acquisition unit 31 acquires information indicating the focal position and in-focus position information for each learning image 80 in the Z stack. The information indicating the focal position and in-focus position information are values (values from +5 to -5) indicating the focal position in the preset units described above, as shown in Figure 13. These values are associated with each learning image 80 in the Z stack. Of these values, ±0 indicates the focal position at the time of focus and is associated with the in-focus learning image 80 in the Z stack. The other values indicate information regarding the direction and distance of the focal position in the in-focus learning image 80 relative to the focal position at the time of imaging and are associated with the learning image 80 corresponding to that information. This value is, for example, the value obtained by subtracting the distance corresponding to the position of the object being imaged when the estimated target image 50 was imaged from the distance corresponding to the position of the object being imaged at the time of focus.
[0112] In other words, in this case, the value represents the focal position in a coordinate system where the focal position when the in-focus training image 80 was captured is set to 0. If the distance to the object being photographed related to this value is longer than the distance to the object being photographed when the in-focus training image 80 was captured, the value will be negative. If the distance to the object being photographed related to this value is shorter than the distance to the object being photographed when the in-focus training image 80 was captured, the value will be positive. The distance to the object being photographed when the in-focus training image 80 was captured refers to the distance from the position of the objective lens or other lens to the position of the object being photographed when the object being photographed in the training image 80 is in focus, and generally corresponds to the focal length of the lens. The focal position when in focus is the position where the object being photographed in the training image 80 is in focus.
[0113] To perform proper machine learning training, each training image 80 in the Z-stack should include as many in-focus training images 80 (training images 80 with a focal position of ±0) as possible in the center. That is, the number of training images 80 corresponding to positive focal positions in each training image 80 of the Z-stack should be roughly equal to the number of training images 80 corresponding to negative focal positions. Note that each training image 80 in the Z-stack does not necessarily have to include a training image 80 with a focal position of ±0.
[0114] As described above, this embodiment takes into account that the focal position at the time of focus may differ for each position in the image. However, the focus position information associated with the Z-stack is in units of 80 training images, not in units of the position (pixel) of the training images 80. It is difficult to determine the appropriate focal position at the time of focus in advance in units of the position (pixel) of the training images 80. As will be described later, in this embodiment, an encoder 70 and a decoder 71 are generated that can estimate the focal position at the time of focus for each position in the image from the focus position information of 80 training images.
[0115] The training images 80 associated with ±0 in the Z-stack can be identified in advance by conventional methods such as identifying in-focus images or measuring the focal position at the time of focus. For example, a contrast evaluation value (e.g., the sum of the absolute values of the differences in peripheral pixel values) can be calculated for each training image 80 in the Z-stack, and the training image 80 with the highest evaluation value can be designated as the training image 80 associated with ±0.
[0116] The learning image acquisition unit 31 may acquire information indicating the focal position and a value indicating the focal position in a preset unit as focus position information by receiving input from the user or from another device. Alternatively, the learning image acquisition unit 31 may store the interval of the focal positions between the learning images 80 of the Z stack in advance and calculate and acquire the value itself from the stored interval and the learning images 80 themselves. The learning image acquisition unit 31 outputs the acquired Z stack and information indicating the focal position corresponding to the Z stack to the learning focus information generation unit 32 and the learning unit 33.
[0117] The learning image acquisition unit 31 may acquire focus position information indicating the focal position when in focus for the Z-stack and multiple learning images corresponding to the Z-stack, by methods other than those described above. Furthermore, the acquired multiple learning images and focus position information may be other information, as long as they are multiple learning images with different focal positions relating to the same image target, each with a corresponding focal position, and focus position information indicating the focal position when in focus for those multiple learning images. The method of acquiring the information is also not limited to those described above.
[0118] The learning focus information generation unit 32 is a learning focus information generation means that inputs information based on each of the multiple training images 80 acquired by the learning image acquisition unit 31 into a focus position estimation model in the middle of training, performs calculations according to the focus position estimation model, acquires information indicating the focal position at the time of focus according to the position in each of the multiple training images 80, and generates learning focus information for each of the multiple training images from the acquired information and the focus position information, indicating the focal position at the time of focus according to the position in the image used for machine learning training. The learning focus information generation unit 32 may also calculate a single focal position common to the multiple training images, corresponding to the position in each of the multiple training images, from the focal positions at the time of focus according to the position in each of the multiple training images 80, as indicated by the information acquired using the focus position estimation model in the middle of training, and generate learning focus information for each of the multiple training images from the single focal position at the time of focus.
[0119] For example, the learning focus information generation unit 32 generates learning focus information as follows. The learning focus information is a learning focal position map (training image data) corresponding to each learning image 80. The learning focal position map is data that has the focal position at the time of focus for each pixel of the learning image 80. The learning focal position map is data in the same format as the focal position map 60 generated from the target image 50. However, during machine learning training, the focal position at the time of focus related to the learning focal position map does not necessarily need to be highly accurate. As shown below, the learning focal position map is repeatedly generated when the encoder 70 and decoder 71 are generated, but its accuracy increases as machine learning training progresses.
[0120] The learning focus information generation unit 32 receives information from the learning image acquisition unit 31 indicating the Z-stack and the corresponding focal position. The learning focus information generation unit 32 also receives the encoder 70 and decoder 71, which are in the process of being trained, from the learning unit 33. The learning focus information generation unit 32 inputs information based on each learning image 80 of the Z-stack into the encoder 70 in the process of being trained, performs calculations according to the encoder 70 in the process of being trained, and obtains the feature quantities of the learning image 80, which are the output from the encoder 70 in the process of being trained. The learning focus information generation unit 32 inputs the obtained feature quantities into the decoder 71 in the process of being trained, performs calculations according to the decoder 71 in the process of being trained, and obtains the estimated focal position map 90, which is the output from the decoder 71 in the process of being trained, as the estimated result of the focal position at the time of focus according to the position in the learning image 80. Figure 13 shows an example of each focal position map 90 of the estimated results obtained from each learning image 80 of the Z-stack.
[0121] As described above, each training image 80 in the Z stack is an image that differs only in its focal position. Therefore, if the encoder 70 and decoder 71 are accurate enough, the focal position at the time of focus shown in each acquired focal position map 90 should be the same. However, because the encoder 70 and decoder 71 are not accurate enough during training, the focal position at the time of focus shown in each acquired focal position map 90 is usually not the same. The training focus information generation unit 32 generates a training focal position map from each acquired focal position map 90 in which the focal position at the time of focus is the same.
[0122] The learning focus information generation unit 32 generates a label map (logical coordinate map) 82, as shown in Figure 15, for each learning image 80 in the Z stack, where all pixel values in images of the same size are values (values from +5 to -5) indicating the focal position associated with the learning image 80. For each learning image 80, the learning focus information generation unit 32 generates a label collection map 91 by adding the pixel values of the label map 82 and the estimated focal position map 90 pixel by pixel. The estimated focal position map 90 is data for the coordinate system of each learning image 80 (a coordinate system based on the focal position when the learning image 80 was captured), but the label collection map 91, which is combined with the label map 82, can be in the same coordinate system (a coordinate system based on the focal position of the learning image 80 at ±0) between them.
[0123] The learning focus information generation unit 32 calculates the average of the pixel values for each pixel in the label collection map 91 and generates a collection map 92 in which the pixel values of each pixel are set to that average. When generating the collection map 92, outliers in the pixel values of each pixel in the label collection map 91 may be excluded. Also, when generating the collection map 92, a weighted average may be used instead of a simple average. The weights of the weighted average are set in advance. For example, among the multiple training images 80, the weight of the label collection map 91 corresponding to images whose corresponding focal position is close to the focal position at the time of focus (those whose focal position is set to ±0) is increased, and the weight of the label collection map 91 corresponding to images that are farther away is decreased. The collection map 92 is data that shows the focal position at the time of focus in the same coordinate system estimated from the encoder 70 and decoder 71 during training, that is, data that shows a single focal position at the time of focus common to multiple training images 80 in the Z stack.
[0124] As shown in Figure 16, the learning focus information generation unit 32 generates a learning focal position map 93 for each label map 82, i.e., for each learning image 80, from the generated collection map 92 and each label map 82. The learning focal position map 93 is generated by subtracting the pixel value of the label map 82 from the pixel value of the collection map 92 for each pixel.
[0125] The generation of the focal position map 93 for training follows the reverse process of the generation of the collection map 92. The collection map 92 is calculated in an ensemble from label collection maps 91 estimated from each training image 80, and therefore has high accuracy due to the ensemble effect. By applying this collection map 92 to each label map 82 as described above, it is possible to generate a highly accurate focal position map 93 for training, that is, one that can be used to generate a more appropriate encoder 70 and decoder 71. Furthermore, by removing outliers and performing a weighted average during the generation of the collection map 92 as described above, it is possible to generate a highly accurate focal position map 93 for training.
[0126] The learning focus information generation unit 32 outputs the generated collection map 92 and the learning focal position map 93 to the learning unit 33. The learning focus information generated by the learning focus information generation unit 32 may be any other learning focus information other than the learning focal position map 93 described above, as long as it indicates the focal position at the time of focus according to the position in the image used for machine learning training, generated from the calculation results and focus position information corresponding to the learning model during training. Furthermore, the generation of learning focus information may be performed by methods other than those described above.
[0127] The learning unit 33 is a learning means that trains machine learning to generate a focus position estimation model using information based on each of the multiple training images 80 acquired by the training image acquisition unit 31, and training focus information corresponding to each of the multiple training images generated by the training focus information generation unit 32. The learning unit 33 may also generate a feature output model that takes image-based information as input and outputs the feature quantities of the image to be input to the focus position estimation model. Based on information indicating the focal position at the time of focus according to the position in each of the multiple training images acquired using the focus position estimation model during training, the learning unit 33 generates two different feature learning images corresponding to the multiple training images, each with a corresponding focal position, and treats the combination of these two feature learning images as a single unit. The learning unit compares the feature quantities of the two feature learning images according to the focal position associated with the two feature learning images, and generates a feature output model by training machine learning based on the comparison result. The learning unit 33 may perform machine learning training such that, when the two feature learning images relate to the same focal position, the difference between the features of the two feature learning images becomes small, and when the two feature learning images relate to different focal positions, the difference between the features of the two feature learning images becomes large.
[0128] The learning unit 33 generates, for example, a decoder 71, which is a focal position estimation model, and an encoder 70, which is a feature output model, as shown below. The training of the encoder 70 and the decoder 71 by the learning unit 33, as well as the updating (generation) of the collection map 92 and the updating (generation) of the focal position map 93 for training by the learning unit 33, are repeated in order as shown in Figure 17. As shown in Figure 17, the updating of the collection map 92 and the updating of the focal position map 93 for training are performed by the learning unit 33 before the training of the encoder 70 and the decoder 71 is performed.
[0129] The learning unit 33 receives information from the learning image acquisition unit 31 indicating the Z-stack and the corresponding focal position. The learning unit 33 receives a collection map 92 and a learning focal position map 93 from the learning focus information generation unit 32.
[0130] As shown in Figure 18, the learning unit 33 trains the encoder 70 using the input information. For training the encoder 70, the learning unit 33 generates a feature learning image 83 from each training image 80 in the Z stack. As described above, in the training image 80, the position within the image, that is, the difference between the focal position when the training image 80 was captured and the focal position when it is in focus, may differ for each pixel. The feature learning image 83 is an image in which these differences do not differ for each pixel (or so it is presumed). In other words, the feature learning image 83 is an image in which the unevenness and tilt of the training image 80 have been corrected (or so it is thought).
[0131] The learning unit 33 generates feature learning images 83 by referring to the collection map 92. For example, the learning unit 33 refers to the collection map 92 and obtains pixels from any of the multiple learning images 80 in the Z stack where the difference between the focal position when the learning image 80 was captured and the focal position at the time of focus is the same, and combines (re-combines) them to generate feature learning images 83. In training the encoder 70, the above difference is used as one focal position at the time of capture for the feature learning images 83. As shown in Figure 18, the learning unit 33 generates multiple feature learning images 83 with different focal positions at the time of capture from a single Z stack. Note that feature learning images 83 may be generated by methods other than those described above.
[0132] Furthermore, the learning unit 33 generates feature learning images 83 from multiple Z stacks. The generated feature learning images 83 include multiple feature learning images 83 with the same focal position (the above-mentioned shift) and multiple feature learning images 83 with different focal positions (the above-mentioned shift). Figure 19 shows an image of the generated feature learning images 83. The vertical direction of the part of Figure 19 showing the feature learning images 83 is the imaging direction. The feature learning images 83 can be considered as partial images (image patches) cut out from a plane 84 where the shift between the focal position when the image was captured and the focal position at the time of focus is the same at all positions in the image. There are multiple planes 84, and the focal positions at the time of capture differ at regular intervals (steps) (ΔZ).
[0133] The focal positions at the time of acquisition of feature learning images 83 that are thought to have been extracted from the same plane 84 are the same, while the focal positions at the time of acquisition of feature learning images 83 that are thought to have been extracted from different planes 84 are different from each other.
[0134] The learning unit 33 uses two feature learning images 83 selected from the multiple feature learning images 83 generated as a set to train the encoder 70's machine learning. The set used for training the machine learning includes both a set of feature learning images 83 relating to the same focal position and a set of feature learning images 83 relating to different focal positions. The selection of the set of feature learning images 83 can be done using a pre-configured method that satisfies the above conditions.
[0135] The learning unit 33 uses information based on the selected set of feature learning images 83 as input to the encoder 70 to perform machine learning training. As shown in Figure 19, when each of the feature learning images 83 in a set is input to the encoder 70, a feature is obtained as output for each of the feature learning images 83. In Figure 19, the values of each element of the feature vector are shown as a bar graph. In this case, the encoder 70 that inputs one set of feature learning images 83 is used as the training target, and the encoder 70 that inputs the other set of feature learning images 83 is used as the comparison target. However, these encoders 70 are the same ones during the training process.
[0136] The learning unit 33 compares two output features according to the focal position of the feature training image 83 and performs machine learning training based on the comparison result. If the focal positions of the two feature training images 83 are the same (i.e., they are on the same plane), the learning unit 33 performs machine learning so that the difference between the features of the two feature training images 83 is small. If the focal positions of the two feature training images 83 are different (i.e., their Z positions are different), the learning unit 33 performs machine learning so that the difference between the features of the two feature training images 83 is large. Note that if the two feature training images 83 are (considered to be) cut out from the same plane 84, the focal positions of the two feature training images 83 will be the same. Also, if the focal positions of the two feature training images 83 are close enough to be considered identical, the focal positions of the two feature training images 83 may be considered identical.
[0137] In other words, the feature quantities of sub-images extracted from an image at the same focal plane are designed to have a high correlation regardless of the extraction position. On the other hand, the feature quantities of sub-images extracted from images at different focal planes are designed to have a low correlation. Through this machine learning training, the feature quantities output from the encoder 70 reflect features corresponding to the focal position.
[0138] Specifically, if the focal positions of the two feature learning images 83 are the same, the learning unit 33 performs machine learning using the following loss_xy as the loss function.
number
[0139] If the focal positions of the two feature learning images 83 are different, the learning unit 33 performs machine learning using the following loss_z as the loss function.
number
[0140] The learning unit 33 generates the encoder 70 by repeatedly selecting a set of feature learning images 83 and training the machine learning model. For example, the learning unit 33 generates the encoder 70 by repeating the above process until the generation of the encoder 70 converges based on pre-set conditions, or for a predetermined number of pre-set times, similar to conventional machine learning training.
[0141] The learning unit 33 may generate the encoder 70 using an existing trained model generated by machine learning. The existing trained model is a model that takes image-based information as input, similar to the encoder 70 in this embodiment. That is, an existing trained model that has the same input as the encoder 70 in this embodiment may be used. The existing trained model is, for example, a model for image recognition, specifically ResNet, VGG, MobileNet, etc. A part of the existing trained model is used to generate the encoder 70. The output layer of the existing trained model is deleted, and the part of the existing trained model up to the intermediate layer is used to generate the encoder 70. The existing trained model used to generate the encoder 70 may include all of the intermediate layers, or it may include only a part of the intermediate layers.
[0142] The learning unit 33 takes a portion of the existing trained model as input and uses it as the encoder 70 at the start of machine learning. That is, the learning unit 33 uses a portion of the existing trained model as the initial parameter of the encoder 70 to perform fine tuning. Alternatively, the encoder 70 at the start of machine learning may be obtained by adding a new output layer to the output side of the portion of the trained model. Furthermore, if a new output layer is added, the encoder 70 at the start of machine learning may be obtained by adding a new hidden layer between the output side of the portion of the trained model and the new output layer.
[0143] The learning unit 33 may generate the encoder 70 without using an existing trained model. For example, a model with random values as initial parameters, similar to conventional machine learning, may be used as the encoder 70 at the start of machine learning training.
[0144] Using an existing pre-trained model to generate the encoder 70 offers the following advantages: Training time can be significantly reduced. A highly accurate encoder 70, i.e., an encoder 70 that can output more appropriate features, can be generated even with a small number of feature training images 83. The aforementioned existing pre-trained model has already acquired the ability to separate features at a low level of abstraction. Therefore, it is only necessary to train it on features at a high level of abstraction using new feature training images 83.
[0145] As shown in Figure 20, the learning unit 33 trains the decoder 71 using the input information. The learning unit 33 inputs the information based on the training image 80 into the encoder 70 after the machine learning training described above, performs calculations according to the encoder 70, and obtains the feature quantities of the training image 80, which are the output from the encoder 70. The learning unit 33 uses the obtained feature quantities as input to the decoder 71 and trains the machine learning model using the information based on the training focal position map 93 corresponding to the training image 80 input to the encoder 70 as the output of the focal position estimation model. The information based on the training focal position map 93 is considered to be the information corresponding to the output from the decoder 71. Furthermore, as described above, the training focal position map 93 is generated by subtracting the label map 82 (Z logical coordinates of image patches) from the collection map 92.
[0146] For example, the input to the encoder 70 during machine learning training is the pixel value of each pixel in the training image 80. If the decoder 71 outputs a focal position map 60, the output from the decoder 71 during machine learning training is the pixel value of each pixel in the training focal position map 93. If the decoder 71 outputs the candidate values described above, the information based on the training focal position map 93 is, for example, a candidate value (one-hot vector for each pixel) where the candidate value corresponding to the pixel value of the training focal position map 93 is set to 1 for each pixel, and the value of the candidate that does not correspond is set to 0. The learning unit 33 generates information based on the training focal position map 93 corresponding to the output from the decoder 71 before performing machine learning training, as needed.
[0147] The machine learning training itself, that is, the updating of the decoder 71 parameters, can be carried out in the same way as before. For example, as shown in Figure 20, the learning unit 33 inputs the feature data into the decoder 71, performs calculations according to the decoder 71, and obtains the output 85 from the decoder 71. The learning unit 33 compares the obtained output 85 with information based on the training focal position map 93 and updates the decoder 71 parameters by backpropagation. Note that since this machine learning training is only for the decoder 71, the encoder 70 parameters are not updated during this machine learning training.
[0148] The learning unit 33 generates the decoder 71 by repeating the machine learning training process for a predetermined number of times, or until the generation of the decoder 71 converges based on predetermined conditions, similar to conventional machine learning training.
[0149] The learning unit 33 determines whether to terminate the training of the encoder 70 and decoder 71 after training the encoder 70 and decoder 71. For example, if the training of the encoder 70 and decoder 71 by the learning unit 33, as shown in Figure 17, and the updating (generation) of the collection map 92 and the updating (generation) of the focal position map 93 for learning by the learning unit 33 have been performed a predetermined number of times, the learning unit 33 determines that the training of the encoder 70 and decoder 71 is terminated. Alternatively, the learning unit 33 may determine to terminate the training of the encoder 70 and decoder 71 based on criteria other than those described above.
[0150] When the learning unit 33 determines that it has finished training the encoder 70 and decoder 71, it sets the encoder 70 and decoder 71 at that point in time as the final encoder 70 and decoder 71 generated by the model generation system 30. The learning unit 33 outputs the encoder 70 and decoder 71 generated by training to the focus position estimation system 20. The generated encoder 70 and decoder 71 may be used for purposes other than those described in this embodiment. In that case, for example, the learning unit 33 transmits or outputs the encoder 70 and decoder 71 to another device or module that uses the encoder 70 and decoder 71. Alternatively, the learning unit 33 may store the generated encoder 70 and decoder 71 in the computer 10 or other device so that they can be used by other devices or modules that use the encoder 70 and decoder 71.
[0151] If the learning unit 33 determines that it has not finished training the encoder 70 and decoder 71, it outputs the partially trained encoder 70 and decoder 71 to the learning focus information generation unit 32. The learning focus information generation unit 32 takes the partially trained encoder 70 and decoder 71 as input. Using the input partially trained encoder 70 and decoder 71, the learning focus information generation unit 32 generates the collection map 92 and the learning focal position map 93 again and outputs them to the learning unit 33.
[0152] The learning unit 33 again receives the collection map 92 and the focal position map 93 for learning from the learning focus information generation unit 32. Using these, the learning unit 33 trains the encoder 70 and decoder 71 again in the same manner as above, and then decides whether to terminate the training of the encoder 70 and decoder 71. The processing after this decision is the same as above. The above describes the functions of the model generation system 30.
[0153] Next, the processes performed by the computer 10 according to this embodiment (the methods performed by the computer 10) will be explained using the flowcharts in Figures 21 and 22. First, the process performed when estimating the focal position at the time of focus corresponding to the target image 50, that is, the focal position estimation method which is performed by the focal position estimation system 20 according to this embodiment, will be explained using the flowchart in Figure 21.
[0154] In this process, first, the estimated target image acquisition unit 21 acquires the estimated target image 50 (S01, estimated target image acquisition step). The estimated target image 50 is based, for example, on an image obtained by imaging by the imaging device 40. Next, the focus position estimation unit 22 uses the encoder 70 and decoder 71 to estimate the focal position at the time of focus from the estimated target image 50, corresponding to the estimated target image 50 and its position in the estimated target image 50 (S02, focus position estimation step). For example, the focal position at the time of focus is estimated for each pixel of the estimated target image 50.
[0155] The estimated focal position at the time of focus is used by the focal position estimation unit 22 (S03, focal position estimation step). For example, as described above, it is used to estimate the inclination of the object being imaged in the estimation target image 50, to control the focal position of the object being imaged in the estimation target image 50 at the time of imaging, to output information indicating the focus state according to the position in the estimation target image 50, or to generate a single image from multiple estimation target images 50. In addition, information indicating the estimated focal position at the time of focus may be output from the focal position estimation unit 22. The above is the focal position estimation method, which is the process executed in the focal position estimation system 20 according to this embodiment.
[0156] Next, using the flowchart in Figure 22, we will explain the process performed when generating the encoder 70 and decoder 71, that is, the model generation method which is performed by the model generation system 30 according to this embodiment.
[0157] In this process, first, the training image acquisition unit 31 acquires multiple training images 80 of the same object being imaged, each with a corresponding focal position, but at different focal positions, and focal position information indicating the focal position when the multiple training images 80 are in focus (S11, training image acquisition step). Subsequently, the training focus information generation unit 32 uses the trained encoder 70 and decoder 71 to generate a collection map 92 and a training focal position map 93 (S12, training focus information generation step).
[0158] Next, the learning unit 33 trains the encoder 70 based on the training image 80 and the collection map 92 (S13, learning step). Next, the learning unit 33 trains the decoder 71 based on the training image 80 and the training focal position map 93 (S14, learning step). Next, the learning unit 33 determines when the training of the encoder 70 and decoder 71 is complete (S15, learning step).
[0159] If it is determined that training of the encoder 70 and decoder 71 is complete (YES in S15), the generated encoder 70 and decoder 71 are output from the learning unit 33 (S16). The generated encoder 70 and decoder 71 are used, for example, in the focus position estimation system 20.
[0160] If it is determined that training of the encoder 70 and decoder 71 is not to be completed (NO in S15), the encoder 70 and decoder 71 that are in the middle of training at that point are used, and the above processing (S12 to S15) is performed again by the learning focus information generation unit 32 and the learning unit 33. The above is the model generation method which is the processing performed by the model generation system 30 according to this embodiment.
[0161] In this embodiment, a decoder 71 is used that outputs information indicating the focal position at the time of focus according to the position in the image, and the focal position at the time of focus according to the position in the target image 50 is estimated. Therefore, according to this embodiment, the focal position based on the image can be estimated according to the position in the image.
[0162] Furthermore, as in this embodiment, a decoder 71 generated by the model generation system 30 may be used to estimate the focal position at the time of focus. This allows for appropriate and reliable estimation of the focal position based on the image. However, the focal position estimation model used for estimation does not necessarily have to be a decoder 71 generated by the model generation system 30; it just needs to be generated by machine learning training and take image-based information as input to output information indicating the focal position at the time of focus according to the position in the image.
[0163] Furthermore, as in this embodiment, in addition to the decoder 71, an encoder 70 generated by the model generation system 30 may be used to estimate the focal position at the time of focus. With this configuration, the encoder 70 is used to acquire feature quantities of the target image 50, and the focal position at the time of focus corresponding to the position in the target image 50 is estimated from the feature quantities. By performing estimation using feature quantities, the estimation of the focal position based on the image can be made more appropriate and reliable. However, the encoder 70 is not required to be used to estimate the focal position at the time of focus. In that case, the decoder 71 is input information based on the target image 50, not the feature quantities output from the encoder 70.
[0164] Furthermore, as in this embodiment, the information indicating the estimated focal position at the time of focus may be used to estimate the inclination of the object being photographed in the estimated target image 50. With this configuration, the inclination of the object being photographed in the estimated target image 50 can be appropriately estimated.
[0165] Furthermore, as in this embodiment, the information indicating the estimated focal position at the time of focus may be used to control the focal position of the object to be imaged in the estimated target image 50 during imaging. With this configuration, the object to be imaged can be properly captured. For example, an image in focus at all positions can be obtained.
[0166] Furthermore, as in this embodiment, the information indicating the estimated focal position at the time of focus may be used to output information indicating the focus state according to the position in the estimated target image 50. With this configuration, it is possible to grasp the focus state according to the position in the estimated target image 50. For example, it is possible to grasp the positions in the estimated target image 50 that were in focus and the positions that were out of focus.
[0167] Furthermore, as in this embodiment, the information indicating the estimated focal position at the time of focus may be used to generate a single image from multiple estimated target images 50. With this configuration, an appropriate image can be obtained. For example, an image in which all positions are in focus can be obtained. However, the information indicating the estimated focal position at the time of focus does not have to be used in any of the above-described ways and may be used arbitrarily for any purpose. Also, the information indicating the estimated focal position at the time of focus does not have to be used in the focal position estimation system 20 and may be used in other devices or modules.
[0168] In this embodiment, a training focal position map 93, which is training focus information, is generated, machine learning is trained, and a decoder 71 is generated. The decoder 71 thus generated can estimate the focal position based on an image, according to the position in the image. Furthermore, in the generation of the decoder 71 according to this embodiment, the decoder 71 can be generated without needing to know the focal position at the time of focus for each position in the training image 80, for example, for each pixel of the training image 80. Therefore, the decoder 71 can be generated easily.
[0169] Furthermore, as in this embodiment, a single collection map 92 common to multiple training images 80 may be calculated from the focal position maps 90 of the estimation results for each of the multiple training images 80 obtained using the decoder 71 during training, and a training focal position map 93 may be generated for each of the multiple training images 80 from the collection map 92. With this configuration, the training focal position map 93 used for training can be made more appropriate, and a decoder 71 that can estimate the focal position more appropriately and reliably is generated. However, the generation of the training focal position map 93 may be performed based on information acquired by the training image acquisition unit 31, and may be performed by methods other than those described above.
[0170] Furthermore, as in this embodiment, an encoder 70 may also be generated in addition to the decoder 71. With this configuration, estimation is performed using the features output from the encoder 70, making it possible to estimate the focal position based on the image more appropriately and reliably. Also, as in this embodiment, training is performed using the training image 80 and the feature training image 83 generated based on the collection map 92, making the encoder 70 capable of outputting more appropriate features.
[0171] Furthermore, by generating feature learning images 83 and using them for training, it is possible to minimize the mixing of features mentioned above. Reducing the mixing of features enables highly accurate machine learning training.
[0172] Furthermore, as in this embodiment, if the two feature training images 83 relate to the same focal position, the machine learning training may be performed so that the difference between the features of the two feature training images 83 is small, and if the two feature training images 83 relate to different focal positions, the difference between the features of the two feature training images 83 is large. With this configuration, the features output from the encoder 70 can be made more appropriate, and as a result, the estimation of the focal position based on the image can be performed more appropriately and reliably. However, the machine learning training does not necessarily have to be performed as described above, and can be performed based on the comparison result of the features of the two feature training images 83. Also, the generation of the encoder 70 itself is not necessarily required. In that case, the decoder 71 to be generated can be generated to take information based on the target image 50, which is not the features output from the encoder 70, as input.
[0173] In this embodiment, the computer 10 includes a focal position estimation system 20 and a model generation system 30, but the focal position estimation system 20 and the model generation system 30 may be implemented independently.
[0174] Next, a specific example of estimating the focal position at the time of focus according to this embodiment will be described. Figure 23 shows the focal position at the time of focus for each pixel position when the focal position at the time of focus is estimated for the target image 50 of the Z stack. In Figure 23, the horizontal axis shows the pixel position of the target image 50, and the vertical axis shows the focal position of the Z stack. The thick dashed line in the figure shows the estimated focal position at the time of focus.
[0175] In the example shown in Figure 23, there are 9 Z-stacks. If Z-stack number 4 is taken as the focal point when in focus, the range of the Z-position of the Z-stack relative to it is -3 to +5 (similarly for pixel positions 2, 3, 5, 6, and 7). However, the focal point of pixel 1 when in focus is at Z-stack number 8, and the range of the Z-position of the Z-stack relative to it is -7 to +1. Similarly, the focal point of pixel 4 when in focus is at Z-stack number 3, and the range of the Z-position of the Z-stack relative to it is -2 to +6. Looking at the range of distances furthest from the focal point when in focus, the range of the Z-position is extended to -7 to +6, showing that it can handle a wider range of information than the original Z-stack's Z-position range of -3 to +5.
[0176] Figure 24 shows a graph of the focal position (Z-stack position) of the target image 50 and the focal position estimated from that target image 50 when in focus. The horizontal axis represents the focal position of the target image 50, and the vertical axis represents the focal position estimated from the target image 50 when in focus. Each point in the graph corresponds to one pixel of the target image 50. Furthermore, all points in a single graph correspond to pixels of the same target image 50. The points in the graph correspond to 128x128 positions (pixels) extracted from a 10x10 grid of positions in each target image 50.
[0177] The graph in Figure 24(a) is a graph of the estimation performed using a decoder that was trained with a uniform value for all pixels in the focal position map for learning, without using the collection map 92, unlike in this embodiment. The graph in Figure 24(b) is a graph of the estimation performed in this embodiment.
[0178] An ideal graph is one where the points on the graph form a band of constant width. The graph in Figure 24(b) according to this embodiment is closer to the ideal than the graph in Figure 24(a). In particular, the difference is large at the edges of the Z-stack.
[0179] This embodiment explains why some shortcut paths for high resolution are not used, unlike in typical SegNet and U-Net. Figure 25 shows an example of a 2D display of the feature quantities at the output of each block of the encoder 70 using the dimensionality reduction display method UMAP (Uniform Manifold Approximation and Projection). Note that the XY in the figure is 2D for UMAP and cannot be compared due to data dependency. Figure 25(a) is a display of the 6th layer, 128 channels of the encoder 70. Figure 25(b) is a display of the 9th layer, 256 channels of the encoder 70. Figure 25(c) is a display of the 12th layer, 512 channels of the encoder 70. Figure 25(d) is a display of the 15th layer, 1024 channels of the encoder 70.
[0180] The Z-stack consists of 51 images. The displayed density corresponds to Z-stack numbers 1 through 51. For each image, a 128x128 array of features is estimated from a 5x5 grid grid. That is, there are 25 points with the same density, resulting in a 51-level gradient.
[0181] In the sixth layer, structural information indicating the structure of the object captured in the image remains, resulting in many areas where features represented by the same intensity are far apart. In layers 9, 12, and 15, the structural information decreases, and the information becomes more in the Z-axis direction (imaging direction). As a result, features at different locations within the same image, represented by the same intensity, are closer together and aligned along the Z-axis.
[0182] In other words, by connecting the features from the 9th layer onward using shortcut paths, it becomes possible to achieve higher spatial accuracy in the Z-position information obtained through 2D decoding. Note that the number of layers at which structural information disappears will vary depending on the magnification of the optical system and the resolution of the image.
[0183] Figure 26 shows an example where an image is extracted from a different position in the training image 80 and used as the target image 50 for estimation. Figure 26(a) is the target image 50 in this example. Figure 26(b) is the focal position map 60 estimated from the target image 50.
[0184] The graphs in Figures 26(c) and (e) differ from this embodiment in that they estimate the focal point position at a single point of focus from a single image, rather than estimating the focal point position for each position in the image, for example, for each pixel. The graphs in Figures 26(d) and (f) are graphs when the estimation is performed according to this embodiment. The points in the graphs in Figures 26(c) and (d) correspond to images extracted from a 10x10 grid pattern in the target image 50. The points in the graphs in Figures 26(e) and (f) correspond to three locations (each at the same position) near the center of the target image 50. The graphs in Figures 26(c) to (f) are similar to the graph in Figure 24.
[0185] Figure 27 shows an example where an image is extracted from a different position in the training image 80 and used as the target image 50 for estimation. Figure 27(a) is the target image 50 in this example. Figure 27(b) is the focal position map 60 estimated from the target image 50.
[0186] The graphs in Figures 27(c) and (e) differ from this embodiment in that they estimate the focal point position at a single point of focus from a single image, rather than estimating the focal point position for each position in the image, for example, for each pixel. The graphs in Figures 27(d) and (f) are graphs when the estimation is performed according to this embodiment. The points in the graphs in Figures 27(c) and (d) correspond to images extracted from a 10x10 grid pattern in the target image 50. The points in the graphs in Figures 27(e) and (f) correspond to three locations (each at the same position) near the center of the target image 50. The graphs in Figures 27(c) to (f) are similar to the graph in Figure 24.
[0187] Figure 28 shows an example where an image is extracted from a different position in the training image 80 and used as the target image 50 for estimation. Figure 28(a) is the target image 50 in this example. Figure 28(b) is the focal position map 60 estimated from the target image 50.
[0188] The graphs in Figures 28(c) and (e) differ from this embodiment in that they estimate the focal point position at a single point of focus from a single image, rather than estimating the focal point position for each position in the image, for example, for each pixel. The graphs in Figures 28(d) and (f) are graphs when the estimation is performed according to this embodiment. The points in the graphs in Figures 28(c) and (d) correspond to images extracted from a 10x10 grid pattern in the target image 50. The points in the graphs in Figures 28(e) and (f) correspond to three locations (each at the same position) near the center of the target image 50. The graphs in Figures 28(c) to (f) are similar to the graph in Figure 24.
[0189] Figure 29 shows an example of estimation performed using an estimation target image 50, which was captured separately from the training image 80. Figure 29(a) is the estimation target image 50 in this example. Figure 29(b) is the focal position map 60 estimated from the estimation target image 50.
[0190] The graphs in Figures 29(c) and (e) differ from this embodiment in that they estimate the focal point position at a single point of focus from a single image, rather than estimating the focal point position for each position in the image, for example, for each pixel. The graphs in Figures 29(d) and (f) are graphs when the estimation is performed according to this embodiment. The points in the graphs in Figures 29(c) and (d) correspond to images extracted from a 10x10 grid pattern in the target image 50. The points in the graphs in Figures 29(e) and (f) correspond to three locations (each at the same position) near the center of the target image 50. The graphs in Figures 29(c) to (f) are similar to the graph in Figure 24.
[0191] As shown in Figures 26 to 29, the estimation method according to this embodiment allows for appropriate estimation based on the position in the image.
[0192] Next, a focus position estimation program and a model generation program for executing the processing performed by the series of focus position estimation systems 20 and model generation systems 30 described above will be explained. As shown in Figure 30(a), the focus position estimation program 200 is stored in a program storage area 211 formed on a computer-readable recording medium 210 that is inserted into and accessed by a computer, or is provided by a computer. The recording medium 210 may be a non-temporary recording medium.
[0193] The focus position estimation program 200 comprises an image acquisition module 201 and a focus position estimation module 202. The functions realized by executing the image acquisition module 201 and the focus position estimation module 202 are the same as the functions of the image acquisition unit 21 and the focus position estimation unit 22 of the focus position estimation system 20 described above.
[0194] As shown in Figure 30(b), the model generation program 300 is stored in a program storage area 311 formed on a computer-readable recording medium 310 that is inserted into and accessed by the computer, or is provided by the computer. The recording medium 310 may be a non-temporary recording medium. If the focus position estimation program 200 and the model generation program 300 are executed on the same computer, the recording medium 310 may be the same as the recording medium 210.
[0195] The model generation program 300 comprises a training image acquisition module 301, a training focus information generation module 302, and a training module 303. The functions realized by executing the training image acquisition module 301, the training focus information generation module 302, and the training module 303 are the same as the functions of the training image acquisition unit 31, the training focus information generation unit 32, and the training unit 33 of the model generation system 30 described above.
[0196] Furthermore, the focal position estimation program 200 and the model generation program 300 may be configured such that part or all of them are transmitted via a transmission medium such as a communication line, received by other equipment, and recorded (including installation). Also, each module of the focal position estimation program 200 and the model generation program 300 may be installed on any of multiple computers, not just one. In that case, the series of processes described above will be performed by a computer system consisting of these multiple computers.
[0197] The focal position estimation method, focal position estimation program, focal position estimation system, model generation method, model generation program, model generation system, and focal position estimation model of this disclosure have the following configurations. [1] A method for estimating the focal position at the time of focus corresponding to the target image, The steps include: acquiring the target image for estimation, and A focal position estimation step is performed using a focal position estimation model that is generated by machine learning training and outputs information indicating the focal position at the time of focus according to the position in the image, based on the estimated target image acquired in the estimation target image acquisition step, in which the focal position at the time of focus corresponding to the estimated target image and the position in the estimated target image is estimated. A method for estimating the focal position, including the method itself. [2] The focal position estimation model is: A learning image acquisition step involves acquiring multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and acquiring focus position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation step involves inputting information based on each of the multiple training images acquired in the training image acquisition step into the focus position estimation model during training, performing calculations according to the focus position estimation model to obtain information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generating learning focus information for each of the multiple training images, indicating the focal position at the time of focus according to the position in the image used for machine learning training, from the acquired information and the focus position information. A learning step in which machine learning is trained to generate the focus position estimation model using information based on each of the multiple learning images acquired in the learning image acquisition step, and learning focus information corresponding to each of the multiple learning images generated in the learning focus information generation step, The focal position estimation method described in [1] is generated by [1]. [3] The focal position estimation model is The focus position estimation method described in [2], wherein in the step of generating learning focus information, a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, is calculated from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focus position estimation model during training, and learning focus information is generated for each of the multiple learning images from the single focal position at the time of focus. [4] In the focal position estimation step, using a feature output model that takes image-based information as input and outputs feature quantities of the image to be input to the focal position estimation model, feature quantities of the target image are obtained from the target image acquired in the estimation target image acquisition step, and the focal position estimation model is used to estimate the focal position at the time of focus corresponding to the target image and its position in the target image from these feature quantities. The aforementioned feature output model is, The method for estimating a focal position as described in [2] or [3], wherein in the learning step, based on information indicating the focal position at the time of focus according to the position in each of the plurality of learning images, obtained using the focal position estimation model during training, the focal position is associated with each of the plurality of learning images, and two different feature learning images corresponding to the plurality of learning images are generated, and the combination of the two feature learning images is treated as a single unit, and the features of the two feature learning images are compared according to the focal position associated with the two feature learning images, and machine learning is trained based on the comparison result. [5] The feature output model is The focal position estimation method described in [4] is generated by training machine learning in the learning step such that, if the two feature learning images relate to the same focal position, the difference between the features of the two feature learning images becomes small, and if the two feature learning images relate to different focal positions, the difference between the features of the two feature learning images becomes large. [6] A method for estimating a focal position according to any one of [1] to [5], wherein in the focal position estimation step, the tilt of the object being photographed in the estimated target image is estimated from the focal position at the time of focus corresponding to the position in the estimated target image. [7] A focal position estimation method according to any one of [1] to [6], wherein in the focal position estimation step, the focal position of the object to be imaged in the estimated target image is controlled based on the focal position at the time of focusing corresponding to the position in the estimated target image. [8] A method for estimating a focal position according to any one of [1] to [7], wherein in the focal position estimation step, information indicating the focus state corresponding to the position in the estimated target image is output based on the focal position at the time of focus corresponding to the position in the estimated target image. [9] In the estimation target image acquisition step, multiple estimation target images with different focal positions relating to the same object are acquired, A method for estimating a focal position according to any one of [1] to [8], wherein in the focal position estimation step, the method estimates the focal position at the time of focus according to the position in the estimated target image from at least one of the plurality of estimated target images acquired in the estimation target image acquisition step, and generates one image from the plurality of estimated target images based on the estimated focal position.
[10] A focal position estimation program that causes a computer to function as a focal position estimation system for estimating the focal position at the time of focus corresponding to an image to be estimated, The computer in question, An estimated target image acquisition means for acquiring an estimated target image, A focal position estimation means estimates the focal position at the time of focus corresponding to the estimated target image and its position in the estimated target image, using a focal position estimation model that is generated by machine learning training and outputs information indicating the focal position at the time of focus according to the position in the image as input to information based on the image, from the estimated target image acquired by the estimated target image acquisition means. A program for estimating the focal position that functions as such.
[11] The focal position estimation model is A learning image acquisition step involves acquiring multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and acquiring focus position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation step involves inputting information based on each of the multiple training images acquired in the training image acquisition step into the focus position estimation model during training, performing calculations according to the focus position estimation model to obtain information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generating learning focus information for each of the multiple training images, indicating the focal position at the time of focus according to the position in the image used for machine learning training, from the acquired information and the focus position information. A learning step in which machine learning is trained to generate the focus position estimation model using information based on each of the multiple learning images acquired in the learning image acquisition step, and learning focus information corresponding to each of the multiple learning images generated in the learning focus information generation step, The focal position estimation program described in
[10] was generated by
[10] . [11-2] The aforementioned focal position estimation model is The focus position estimation program described in
[11] is generated by, in the step of generating learning focus information, calculating a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focus position estimation model during training, and generating the learning focus information for each of the multiple learning images from the single focal position at the time of focus. [11-3] The focal position estimation means uses a feature output model that takes image-based information as input and outputs feature quantities of the image to be input to the focal position estimation model to obtain feature quantities of the target image from the target image acquired by the target image acquisition means, and uses the focal position estimation model to estimate the focal position at the time of focus corresponding to the target image and its position in the target image from these feature quantities. The aforementioned feature output model is, The focal position estimation program described in
[11] or [11-2] is generated in the learning step by associating the focal position with each of the multiple learning images and generating two different feature learning images corresponding to the multiple learning images, based on information indicating the focal position at the time of focus according to the position in each of the multiple learning images obtained using the focal position estimation model during training, treating the combination of the two feature learning images as one unit, comparing the features of the two feature learning images according to the focal position associated with the two feature learning images, and training machine learning based on the comparison result. [11-4] The feature output model described above is The focal position estimation program described in [11-3] is generated by training machine learning in the learning step such that, when the two feature learning images relate to the same focal position, the difference between the features of the two feature learning images becomes small, and when the two feature learning images relate to different focal positions, the difference between the features of the two feature learning images becomes large. [11-5] The focal position estimation means is a focal position estimation program according to any one of
[10] to [11-4] which estimates the inclination of an object captured in an estimated target image from the focal position at the time of focus corresponding to the position in the estimated target image. [11-6] The focal position estimation means controls the focal position of an object captured in the estimated target image at the time of imaging, based on the focal position at the time of focusing corresponding to the position in the estimated target image, according to any of
[10] to [11-5]. [11-7] The focal position estimation means outputs information indicating the focus state corresponding to the position in the estimated target image based on the focal position at the time of focus corresponding to the position in the estimated target image, according to any of
[10] to [11-6]. [11-8] The focal position estimation means acquires multiple estimated target images of the same object at different focal positions, The focal position estimation means estimates the focal position at the time of focus according to the position in the estimated target image from at least one of the plurality of estimated target images acquired by the estimation target image acquisition means, and generates one image from the plurality of estimated target images based on the estimated focal position, as described in any of
[10] to [11-7].
[12] A focal position estimation system for estimating the focal position at the time of focus corresponding to the target image, An estimated target image acquisition means for acquiring an estimated target image, A focal position estimation means estimates the focal position at the time of focus corresponding to the estimated target image and its position in the estimated target image, using a focal position estimation model that is generated by machine learning training and outputs information indicating the focal position at the time of focus according to the position in the image as input to information based on the image, from the estimated target image acquired by the estimated target image acquisition means. A focal position estimation system equipped with the following features.
[13] The focal position estimation model is A learning image acquisition step involves acquiring multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and acquiring focus position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation step involves inputting information based on each of the multiple training images acquired in the training image acquisition step into the focus position estimation model during training, performing calculations according to the focus position estimation model to obtain information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generating learning focus information for each of the multiple training images, indicating the focal position at the time of focus according to the position in the image used for machine learning training, from the acquired information and the focus position information. A learning step in which machine learning is trained to generate the focus position estimation model using information based on each of the multiple learning images acquired in the learning image acquisition step, and learning focus information corresponding to each of the multiple learning images generated in the learning focus information generation step, The focal position estimation system described in
[12] was generated by
[12] . [13-2] The aforementioned focal position estimation model is: The focus position estimation system described in
[13] is generated by, in the step of generating learning focus information, calculating a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focus position estimation model during training, and generating the learning focus information for each of the multiple learning images from the single focal position at the time of focus. [13-3] The focal position estimation means uses a feature output model that takes image-based information as input and outputs feature quantities of the image to be input to the focal position estimation model to obtain feature quantities of the target image from the target image acquired by the target image acquisition means, and uses the focal position estimation model to estimate the focal position at the time of focus corresponding to the target image and its position in the target image from these feature quantities. The aforementioned feature output model is, The focal position estimation system described in
[13] or [13-2] is generated in the learning step by associating the focal position with each of the multiple learning images based on information indicating the focal position at the time of focus according to the position in each of the multiple learning images, obtained using the focal position estimation model during training, and generating two different feature learning images corresponding to the multiple learning images, treating the combination of the two feature learning images as one unit, comparing the features of the two feature learning images according to the focal position associated with the two feature learning images, and training machine learning based on the comparison result. [13-4] The feature output model described above is The focal position estimation system described in [13-3] is generated by training machine learning in the learning step such that, when the two feature learning images relate to the same focal position, the difference between the features of the two feature learning images becomes small, and when the two feature learning images relate to different focal positions, the difference between the features of the two feature learning images becomes large. [13-5] The focal position estimation means estimates the inclination of an object captured in the estimated target image from the focal position at the time of focus corresponding to the position in the estimated target image, according to any one of
[12] to [13-4]. [13-6] A focal position estimation system according to any one of
[12] to [13-5], wherein the focal position estimation means controls the focal position at the time of imaging of an object captured in the estimated target image based on the focal position at the time of focusing corresponding to the position in the estimated target image. [13-7] A focal position estimation system according to any one of
[12] to [13-6], wherein the focal position estimation means outputs information indicating the focus state corresponding to the position in the estimated target image based on the focal position at the time of focus corresponding to the position in the estimated target image. [13-8] The focal position estimation means acquires multiple estimated target images of the same object at different focal positions, A focal position estimation system according to any one of
[12] to [13-7], wherein the focal position estimation means estimates the focal position at the time of focus according to the position in the estimated target image from at least one of the plurality of estimated target images acquired by the estimation target image acquisition means, and generates one image from the plurality of estimated target images based on the estimated focal position.
[14] A model generation method for generating a focal position estimation model that takes image-based information as input and outputs information indicating the focal position at the time of focus according to the position in the image, A learning image acquisition step involves acquiring multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and acquiring focus position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation step involves inputting information based on each of the multiple training images acquired in the training image acquisition step into the focus position estimation model during training, performing calculations according to the focus position estimation model to obtain information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generating learning focus information for each of the multiple training images, indicating the focal position at the time of focus according to the position in the image used for machine learning training, from the acquired information and the focus position information. A learning step in which machine learning is trained to generate the focus position estimation model using information based on each of the multiple learning images acquired in the learning image acquisition step, and learning focus information corresponding to each of the multiple learning images generated in the learning focus information generation step, A model generation method that includes this.
[15] The model generation method according to
[14] , wherein in the step of generating learning focus information, the model generation method is characterized in that, in the step of generating learning focus information, the model generates learning focus information for each of the multiple learning images by calculating a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focal position estimation model during training, and generating learning focus information for each of the multiple learning images from the single focal position at the time of focus.
[16] In the learning step, a feature output model is generated that takes image-based information as input and outputs the feature quantities of the image to be input to the focus position estimation model, A model generation method according to
[14] or
[15] , wherein in the learning step, based on information indicating the focal position at the time of focus according to the position in each of the plurality of learning images, obtained using the focal position estimation model during training, two different feature learning images are generated corresponding to the plurality of learning images, each with a corresponding focal position, and the combination of the two feature learning images is treated as a single unit, and the features of the two feature learning images are compared according to the focal position associated with the two feature learning images, and machine learning training is performed based on the comparison result to generate the feature output model.
[17] The model generation method according to
[16] , wherein in the learning step, if the two feature learning images relate to the same focal position, the machine learning is trained so that the difference between the features of the two feature learning images becomes small, and if the two feature learning images relate to different focal positions, the difference between the features of the two feature learning images becomes large.
[18] A model generation program that causes a computer to function as a model generation system that takes image-based information as input and generates a focal position estimation model that outputs information indicating the focal position at the time of focus according to the position in the image, The computer in question, A learning image acquisition means that acquires multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and focal position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation means inputs information based on each of the multiple training images acquired by the learning image acquisition means into the focus position estimation model during training, performs calculations according to the focus position estimation model, obtains information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generates learning focus information indicating the focal position at the time of focus according to the position in the image used for machine learning training for each of the multiple training images from the acquired information and the focus position information, A learning means for training machine learning to generate the focus position estimation model using information based on each of the multiple training images acquired by the training image acquisition means, and training focus information corresponding to each of the multiple training images generated by the training focus information generation means. A model generation program that functions as such. [18-2] The model generation program described in
[18] , wherein the learning focus information generation means calculates a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focal position estimation model during training, and generates the learning focus information for each of the multiple learning images from the single focal position at the time of focus. [18-3] The learning means inputs information based on an image and generates a feature output model that outputs the feature quantities of the image to be input to the focus position estimation model, The learning means generates two different feature learning images corresponding to the multiple learning images, each with a corresponding focal position, based on information indicating the focal position at the time of focus according to the position in each of the multiple learning images, obtained using the focal position estimation model during training, and treating the combination of the two feature learning images as a single unit, compares the features of the two feature learning images according to the focal position associated with the two feature learning images, and performs machine learning training based on the comparison result to generate the feature output model as described in
[18] or [18-2]. [18-4] A model generation program as described in [18-3], wherein in the learning step, if the two feature learning images relate to the same focal position, the machine learning is trained so that the difference between the features of the two feature learning images becomes small, and if the two feature learning images relate to different focal positions, the difference between the features of the two feature learning images becomes large.
[19] A model generation system that generates a focal position estimation model that takes image-based information as input and outputs information indicating the focal position at the time of focus according to the position in the image, A learning image acquisition means that acquires multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and focal position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation means inputs information based on each of the multiple training images acquired by the learning image acquisition means into the focus position estimation model during training, performs calculations according to the focus position estimation model, obtains information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generates learning focus information indicating the focal position at the time of focus according to the position in the image used for machine learning training for each of the multiple training images from the acquired information and the focus position information, A learning means for training machine learning to generate the focus position estimation model using information based on each of the multiple training images acquired by the training image acquisition means, and training focus information corresponding to each of the multiple training images generated by the training focus information generation means. A model generation system that includes this. [19-2] The model generation system according to
[19] , wherein the learning focus information generation means calculates a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focal position estimation model during training, and generates the learning focus information for each of the multiple learning images from the single focal position at the time of focus. [19-3] The learning means takes information based on an image as input and generates a feature output model that outputs the feature quantities of the image to be input to the focus position estimation model, The learning means generates two different feature learning images corresponding to the multiple learning images, each with a corresponding focal position, based on information indicating the focal position at the time of focus according to the position in each of the multiple learning images, obtained using the focal position estimation model during training, and treating the combination of the two feature learning images as a single unit, compares the features of the two feature learning images according to the focal position associated with the two feature learning images, and trains the machine learning model based on the comparison result to generate the feature output model as described in
[19] or [19-2]. [19-4] The model generation system according to [19-3], wherein in the learning step, if the two feature learning images relate to the same focal position, the machine learning is trained so that the difference between the features of the two feature learning images becomes small, and if the two feature learning images relate to different focal positions, the difference between the features of the two feature learning images becomes large.
[20] A focal position estimation model that is generated by machine learning training and which enables a computer to take image-based information as input and output information indicating the focal position at the time of focus, corresponding to the position in the image.
[21] A learning image acquisition step that acquires multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and focal position information indicating the focal position when the multiple learning images are in focus, A learning focus information generation step involves inputting information based on each of the multiple training images acquired in the training image acquisition step into the focus position estimation model during training, performing calculations according to the focus position estimation model to obtain information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generating learning focus information for each of the multiple training images, indicating the focal position at the time of focus according to the position in the image used for machine learning training, from the acquired information and the focus position information. A learning step in which machine learning is trained to generate the focus position estimation model using information based on each of the multiple learning images acquired in the learning image acquisition step, and learning focus information corresponding to each of the multiple learning images generated in the learning focus information generation step, The focal position estimation model described in
[20] was generated by
[20] . [21-2] The focus position estimation model according to
[21] , wherein in the learning focus information generation step, a single focus position common to the multiple learning images, corresponding to the position in each of the multiple learning images, is calculated from the focus positions at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focus position estimation model during training, and the learning focus information is generated for each of the multiple learning images from the single focus position at the time of focus. [Explanation of Symbols]
[0198] 10...Computer, 20...Focus position estimation system, 21...Estimation target image acquisition unit, 22...Focus position estimation unit, 30...Model generation system, 31...Training image acquisition unit, 32...Training focus information generation unit, 33...Learning unit, 40...Imaging device, 200...Focus position estimation program, 201...Estimation target image acquisition module, 202...Focus position estimation module, 210...Recording medium, 211...Program storage area, 300...Model generation program, 301...Training image acquisition module, 302...Training focus information generation module, 303...Learning module, 310...Recording medium, 311...Program storage area.
Claims
1. A method for estimating the focal position at the time of focus corresponding to the target image, The steps include: acquiring the target image for estimation, and A focal position estimation step is performed using a focal position estimation model that is generated by machine learning training and outputs information indicating the focal position at the time of focus according to the position in the image, based on the estimated target image acquired in the estimation target image acquisition step, in which the focal position at the time of focus corresponding to the estimated target image and the position in the estimated target image is estimated. Includes, The aforementioned focal position estimation model is, A learning image acquisition step involves acquiring multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and acquiring focus position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation step involves inputting information based on each of the multiple training images acquired in the training image acquisition step into the focus position estimation model during training, performing calculations according to the focus position estimation model to obtain information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generating learning focus information for each of the multiple training images, indicating the focal position at the time of focus according to the position in the image used for machine learning training, from the acquired information and the focus position information. A learning step in which machine learning is trained to generate the focus position estimation model using information based on each of the multiple learning images acquired in the learning image acquisition step, and learning focus information corresponding to each of the multiple learning images generated in the learning focus information generation step, It was generated by, The aforementioned focal position estimation model is, A focus position estimation method that, in the step of generating learning focus information, calculates a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focus position estimation model during training, and generates the learning focus information for each of the multiple learning images from the single focal position at the time of focus.
2. In the focal position estimation step, using a feature output model that takes image-based information as input and outputs feature quantities of the image to be input to the focal position estimation model, feature quantities of the target image are obtained from the target image acquired in the estimation target image acquisition step, and the focal position estimation model is used to estimate the focal position at the time of focus corresponding to the target image and its position in the target image from these feature quantities. The aforementioned feature output model is, The method for estimating a focal position according to claim 1, wherein in the learning step, based on information indicating the focal position at the time of focus according to the position in each of the plurality of learning images, obtained using the focal position estimation model during training, the focal position is associated with each of the plurality of learning images, and two different feature learning images corresponding to the plurality of learning images are generated, and the combination of the two feature learning images is treated as a single unit, and the features of the two feature learning images are compared according to the focal position associated with the two feature learning images, and machine learning is trained based on the comparison result.
3. The aforementioned feature output model is, The method for estimating a focal position according to claim 2, wherein in the learning step, if the two feature learning images relate to the same focal position, the machine learning training is performed such that the difference between the features of the two feature learning images becomes small, and if the two feature learning images relate to different focal positions, the difference between the features of the two feature learning images becomes large.
4. The method for estimating a focal position according to claim 1 or 2, wherein in the focal position estimation step, the tilt of an object captured in the estimated target image is estimated from the focal position at the time of focus corresponding to the position in the estimated target image.
5. The method for estimating a focal position according to claim 1 or 2, wherein in the focal position estimation step, the focal position of an object captured in the estimated target image is controlled based on the focal position at the time of focusing, corresponding to the position in the estimated target image.
6. The method for estimating a focal position according to claim 1 or 2, wherein in the focal position estimation step, information indicating the focus state corresponding to the position in the estimated target image is output based on the focal position at the time of focus corresponding to the position in the estimated target image.
7. In the above estimation target image acquisition step, multiple estimation target images with different focal positions relating to the same object to be imaged are acquired. The method for estimating a focal position according to claim 1 or 2, wherein in the focal position estimation step, the method involves estimating the focal position at the time of focus according to the position in the estimated target image, from at least one of the plurality of estimated target images acquired in the estimation target image acquisition step, and generating one image from the plurality of estimated target images based on the estimated focal position.
8. A focal position estimation program that causes a computer to function as a focal position estimation system that estimates the focal position at the time of focus corresponding to the target image, The computer in question, An estimated target image acquisition means for acquiring an estimated target image, A focal position estimation means estimates the focal position at the time of focus corresponding to the estimated target image and its position in the estimated target image, using a focal position estimation model that is generated by machine learning training and outputs information indicating the focal position at the time of focus according to the position in the image as input to information based on the image, from the estimated target image acquired by the estimated target image acquisition means. To make it function as, The aforementioned focal position estimation model is, A learning image acquisition step involves acquiring multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and acquiring focus position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation step involves inputting information based on each of the multiple training images acquired in the training image acquisition step into the focus position estimation model during training, performing calculations according to the focus position estimation model to obtain information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generating learning focus information for each of the multiple training images, indicating the focal position at the time of focus according to the position in the image used for machine learning training, from the acquired information and the focus position information. A learning step in which machine learning is trained to generate the focus position estimation model using information based on each of the multiple learning images acquired in the learning image acquisition step, and learning focus information corresponding to each of the multiple learning images generated in the learning focus information generation step, It was generated by, The aforementioned focal position estimation model is, A focus position estimation program generated by, in the learning focus information generation step, calculating a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focus position estimation model during training, and generating the learning focus information for each of the multiple learning images from the single focal position at the time of focus.
9. A focal position estimation system that estimates the focal position at the time of focus corresponding to the target image, An estimated target image acquisition means for acquiring an estimated target image, A focal position estimation means estimates the focal position at the time of focus corresponding to the estimated target image and its position in the estimated target image, using a focal position estimation model that is generated by machine learning training and outputs information indicating the focal position at the time of focus according to the position in the image as input to information based on the image, from the estimated target image acquired by the estimated target image acquisition means. Equipped with, The aforementioned focal position estimation model is, A learning image acquisition step involves acquiring multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and acquiring focus position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation step involves inputting information based on each of the multiple training images acquired in the training image acquisition step into the focus position estimation model during training, performing calculations according to the focus position estimation model to obtain information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generating learning focus information for each of the multiple training images, indicating the focal position at the time of focus according to the position in the image used for machine learning training, from the acquired information and the focus position information. A learning step in which machine learning is trained to generate the focus position estimation model using information based on each of the multiple learning images acquired in the learning image acquisition step, and learning focus information corresponding to each of the multiple learning images generated in the learning focus information generation step, It was generated by, The aforementioned focal position estimation model is, A focus position estimation system generated by, in the learning focus information generation step, calculating a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focus position estimation model during training, and generating the learning focus information for each of the multiple learning images from the single focal position at the time of focus.
10. A model generation method for generating a focal position estimation model that takes image-based information as input and outputs information indicating the focal position at the time of focus according to the position in the image, A learning image acquisition step involves acquiring multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and acquiring focus position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation step involves inputting information based on each of the multiple training images acquired in the training image acquisition step into the focus position estimation model during training, performing calculations according to the focus position estimation model to obtain information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generating learning focus information for each of the multiple training images, indicating the focal position at the time of focus according to the position in the image used for machine learning training, from the acquired information and the focus position information. A learning step in which machine learning is trained to generate the focus position estimation model using information based on each of the multiple learning images acquired in the learning image acquisition step, and learning focus information corresponding to each of the multiple learning images generated in the learning focus information generation step, Includes, A model generation method comprising the steps for generating learning focus information, wherein in the learning focus information generation step, a single focal position common to the multiple learning images is calculated from the focal position at the time of focus
11. In the learning step described above, a feature output model is generated that takes image-based information as input and outputs the feature quantities of the image to be input to the focus position estimation model. The model generation method according to claim 10, wherein in the learning step, based on information indicating the focal position at the time of focus according to the position in each of the plurality of learning images, obtained using the focal position estimation model during training, two different feature learning images are generated corresponding to the plurality of learning images, each with a corresponding focal position, and the combination of the two feature learning images is treated as a single unit, and the features of the two feature learning images are compared according to the focal position associated with the two feature learning images, and machine learning training is performed based on the comparison result to generate the feature output model.
12. The model generation method according to claim 11, wherein in the learning step, if the two feature learning images relate to the same focal position, the machine learning training is performed such that the difference in features between the two feature learning images becomes small, and if the two feature learning images relate to different focal positions, the difference in features between the two feature learning images becomes large.
13. A model generation program that causes a computer to function as a model generation system that takes image-based information as input and generates a focal position estimation model that outputs information indicating the focal position at the time of focus according to the position in the image, The computer in question, A learning image acquisition means that acquires multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and focal position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation means inputs information based on each of the multiple training images acquired by the learning image acquisition means into the focus position estimation model during training, performs calculations according to the focus position estimation model, obtains information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generates learning focus information indicating the focal position at the time of focus according to the position in the image used for machine learning training for each of the multiple training images from the acquired information and the focus position information, A learning means for training machine learning to generate the focus position estimation model using information based on each of the multiple training images acquired by the training image acquisition means, and training focus information corresponding to each of the multiple training images generated by the training focus information generation means. To make it function as, The learning focus information generation means is a model generation program that calculates a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focal position estimation model during training, and generates the learning focus information for each of the multiple learning images from the single focal position at the time of focus.
14. A model generation system that generates a focal position estimation model by inputting information based on an image and outputting information indicating the focal position at the time of focus according to the position in the image, A learning image acquisition means that acquires multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and focal position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation means inputs information based on each of the multiple training images acquired by the learning image acquisition means into the focus position estimation model during training, performs calculations according to the focus position estimation model, obtains information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generates learning focus information indicating the focal position at the time of focus according to the position in the image used for machine learning training for each of the multiple training images from the acquired information and the focus position information, A learning means for training machine learning to generate the focus position estimation model using information based on each of the multiple training images acquired by the training image acquisition means, and training focus information corresponding to each of the multiple training images generated by the training focus information generation means. Equipped with, The aforementioned training focus information generation means is a model generation system that calculates a single focal position common to the multiple training images, corresponding to the position in each of the multiple training images, from the focal position at the time of focus corresponding to the position in each of the multiple training images, as indicated by the information obtained using the focal position estimation model during training, and generates the training focus information for each of the multiple training images from the single focal position at the time of focus.
15. A focal position estimation model that is generated by machine learning training and is used to enable a computer to take image-based information as input and output information indicating the focal position at the time of focus according to the position in the image, A learning image acquisition step involves acquiring multiple learning images of the same object being imaged, each with a corresponding focal position, but with different focal positions, and acquiring focus position information indicating the focal position when the multiple learning images are in focus. A learning focus information generation step involves inputting information based on each of the multiple training images acquired in the training image acquisition step into the focus position estimation model during training, performing calculations according to the focus position estimation model to obtain information indicating the focal position at the time of focus according to the position in each of the multiple training images, and generating learning focus information for each of the multiple training images, indicating the focal position at the time of focus according to the position in the image used for machine learning training, from the acquired information and the focus position information. A learning step in which machine learning is trained to generate the focus position estimation model using information based on each of the multiple learning images acquired in the learning image acquisition step, and learning focus information corresponding to each of the multiple learning images generated in the learning focus information generation step, It was generated by, In the learning focus information generation step, a focus position estimation model is generated by calculating a single focal position common to the multiple learning images, corresponding to the position in each of the multiple learning images, from the focal position at the time of focus corresponding to the position in each of the multiple learning images, as shown by the information obtained using the focus position estimation model during training, and then generating the learning focus information for each of the multiple learning images from the single focal position at the time of focus.
Citation Information
Patent Citations
Fully focusing image pickup device
JP2000316120A
Method for determining in-focus position and vision inspection system
JP2013050713A
JPP7205013B
Step measuring method and apparatus, and exposure method and apparatus
WO2005088686A1