Image analysis apparatus, image analysis method, and program
Patent Information
- Application Number
- JP2023003395
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-01-12
- Publication Date
- 2026-01-16
AI Technical Summary
Three-dimensional imaging using confocal microscopes is unsuitable for observing high-speed events due to the time required for scanning in the depth direction, and the large data size of volumetric images poses challenges for transmission, especially with wireless communication.
An image analysis device and method that utilizes a group of elemental images from shifted viewpoints to track the three-dimensional spatial position of an object by reconstructing images at arbitrary depth positions and analyzing parallax information in light field images, using a microlens array and image sensor to generate multi-view image data.
Enables high-speed tracking of three-dimensional spatial positions of objects from high frame rate moving image data, reducing data size and facilitating storage and transmission, particularly suitable for observing biomolecules and intracellular behavior.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an image analysis device, an image analysis method, and a program, and in particular to an image analysis technique suitable for tracking the three-dimensional spatial position of an object in an image from continuous data of a group of elemental images obtained from a light field image that records a light space (light field). [Background technology]
[0002] High-speed, large-scale three-dimensional imaging is required to observe biomolecules and intracellular behaviors in a delicate and dynamic manner. In general, such three-dimensional images are generated by using a confocal microscope to continuously change the focal position of the sample in the depth direction (Z direction) at a specified pitch and capture the image of the sample at each focal position (see, for example, Patent Documents 1 and 2). In addition, a spinning disk confocal microscope is known as a microscope that enables three-dimensional imaging. In the spinning disk confocal method, a laser beam is irradiated onto a rotating disk with a number of pinholes to create multiple parallel beams of light, which are then used to scan the sample at high speed to form a confocal image. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2011-044016 A [Patent Document 2] JP 2014-157158 A Summary of the Invention [Problem to be solved by the invention]
[0004] Three-dimensional imaging using a confocal microscope requires scanning in the depth direction, which takes time for measurement, and is not suitable for observing high-speed events. For example, when imaging a three-dimensional space by scanning 30 times in the Z direction at 30 fps, the imaging speed of the entire three-dimensional space is 30 fps / 30 = 1 vps, and it takes one second to capture the three-dimensional space. Since neural activity is a phenomenon on the order of milliseconds, three-dimensional imaging using a confocal microscope is insufficient in terms of speed for observing biomolecules and intracellular behavior.
[0005] In addition, because the data size per unit volume of 3D imaging is very large, when such volumetric images are converted into time-series video data, the data size can reach the scale of several terabytes. Such large-scale data would take a very long time to transmit even using wired communication, which has a high transmission speed, and there is a risk that it would be virtually impossible to transmit using wireless communication, which has an even slower transmission speed.
[0006] Another option for 3D imaging is light field microscopy (LFM). A light field is a collection of light rays in a 3D space. LFM uses a microlens array, which is a two-dimensional arrangement of many microlenses, placed on an intermediate image plane, and records the light field by obtaining the position of the light rays passing on the aperture stop of the objective lens and the position of the light rays being acquired on the image sensor. From the light field image recorded by LFM, a subaperture image equivalent to a small-diameter lens can be obtained. A subaperture image is a partial image of a subject captured by shifting the viewpoint, and each subaperture image has a deep depth of field due to the pinhole effect, and at the same time, it has parallax between other subaperture images. Therefore, by shifting and overlapping the subaperture images, it is possible to reconstruct an image refocused at any depth position.
[0007] In this way, LFM is suitable for observing high-speed events because it can capture 3D images in a single shot at high speed without scanning in the depth direction. In terms of data size, the light field images themselves are 2D images, so the data size is small, and the data size of video data that is a series of these images is also small.
[0008] Therefore, an object of the present invention is to analyze video data in which two-dimensional images containing parallax information, such as light field images, are successive in time, and to track the three-dimensional spatial position of an object in the images. [Means for solving the problem]
[0009] According to one aspect of the present invention, there is provided an image analysis device for analyzing moving image data in which a group of elemental images, in which a plurality of elemental images obtained by capturing the same subject from a plurality of mutually shifted viewpoints are arranged, are successive in time, and for tracking a three-dimensional spatial position of an object in the image, the image analysis device including: an image reconstruction unit for reconstructing an image refocused on a depth position of the object from a first group of elemental images; a template selection unit for selecting an elemental image including the largest number of bright points from the first group of elemental images as a template; a bright point representative point identification unit for determining a representative point of each bright point in the second and subsequent groups of elemental images; a bright point tracking unit for determining a positional change of the representative point of each bright point in the second and subsequent groups of elemental images; The image analysis device includes a bright point labeling unit that assigns a label for identifying an object to a representative point of each bright point consisting of the bright points, and assigns the same label to a representative point of each bright point in the template in each element image in the second and subsequent element image groups, which corresponds to a representative point of each bright point in the template; and a three-dimensional coordinate updating unit that determines the three-dimensional spatial coordinates of the object in the reconstructed image and sets the coordinates as the initial coordinates of the object identified by the label, updates the XY coordinates of the object identified by the label based on a position change of the representative point of the bright point with the same label in the second and subsequent element image groups, and updates the Z coordinate of the object based on the distance between the representative points between adjacent element images.
[0010] According to another aspect of the present invention, there is provided an image analysis method for analyzing moving image data in which a group of elemental images, each of which is an array of elemental images obtained by capturing an identical subject from a plurality of mutually shifted viewpoints, is continuous in time, and tracking a three-dimensional spatial position of an object in the image, the method comprising the steps of: reconstructing an image refocused on a depth position of the object from a first group of elemental images; attaching a label for identifying the object to a representative point of each bright point, which is a group of pixels corresponding to the object in the reconstructed image in the first group of elemental images; determining the three-dimensional spatial coordinates of the object in the reconstructed image and setting the coordinates to initial coordinates of the object identified by the label; and selecting an elemental image including the most bright points from the first group of elemental images. the step of selecting a representative point of each bright point in the second or subsequent element image groups as a template; a step of determining a representative point of each bright point in the second or subsequent element image groups; a step of labeling a representative point of each bright point in each element image of the second or subsequent element image groups that corresponds to a representative point of each bright point in the template with the same label as the representative point of each bright point in the template; and a step of updating the XY coordinates of an object identified by the label based on the position change of the representative point of the bright point with the same label in the second or subsequent element image groups, and updating the Z coordinate of the object based on the distance between the representative points between adjacent element images.
[0011] According to yet another aspect of the present invention, there is provided a program for causing a computer to analyze video data of a group of elemental images, each of which is an array of elemental images of the same subject captured from a plurality of mutually shifted viewpoints, and for tracking a three-dimensional spatial position of an object in the image, the program including: a three-dimensional image reconstructing means for reconstructing an image refocused on a depth position of the object from a first group of elemental images; a template selecting means for selecting an elemental image including the largest number of bright points from the first group of elemental images as a template; a bright point representative point identifying means for determining a representative point of each bright point in the second and subsequent groups of elemental images; a bright point tracking means for determining a positional change of the representative point of each bright point in the second and subsequent groups of elemental images; a group of pixels corresponding to the object in the reconstructed image in the first group of elemental images; The present invention provides a program that causes a computer to function as a bright point labeling means that assigns a label for identifying an object to a representative point of each bright point consisting of a cluster, and assigns the same label to a representative point of each bright point in the template in each element image in the second or subsequent element image groups that corresponds to the representative point of each bright point in the template, and a three-dimensional coordinate updating means that determines the three-dimensional spatial coordinates of the object in the reconstructed image and sets the coordinates as the initial coordinates of the object identified by the label, updates the XY coordinates of the object identified by the label based on a positional change of the representative point of the bright point with the same label in the second or subsequent element image groups, and updates the Z coordinate of the object based on the distance between the representative points between adjacent element images. Effect of the Invention
[0012] According to the present invention, it is possible to track the three-dimensional spatial position of an object in an image from high frame rate video data such as continuous data of elemental image groups, so that it is possible to observe high-speed events such as biomolecules and intracellular behavior. In addition, video data of continuous elemental image groups, which are two-dimensional images, requires an overwhelmingly smaller data size than continuous data of volumetric image data obtained by a confocal microscope, so it does not put a strain on storage capacity for storage, and can be transmitted by wireless communication, etc. [Brief description of the drawings]
[0013] [Figure 1]1 is a schematic diagram of an optical system including an image analysis device according to an embodiment of the present invention; [Diagram 2] 1A is a diagram showing an example of multi-viewpoint image data, (b) a group of elemental images, and (c) a reconstructed image. [Diagram 3] FIG. 2 is a diagram illustrating a geometric relationship between parameters of an optical system according to an example. [Figure 4] 1A and 1B are diagrams for explaining the relationship between a PQ array and superimposition of elemental images. [Figure 5A] 1 is a first half of a flowchart of an image analysis method according to one embodiment of the present invention. [Figure 5B] 11 is a flowchart showing the second half of the image analysis method according to one embodiment of the present invention. [Figure 6] 11A and 11B are diagrams illustrating an example of three-dimensionally reconstructing an object from an initial group of elemental images, performing bright point labeling, setting initial coordinates of the object, and selecting a template. [Figure 7] 11A and 11B are diagrams for explaining how to specify a representative point of each bright point in a group of elemental images. [Figure 8] 11A and 11B are diagrams for explaining the tracking of representative points of each bright point in a group of elemental images. [Figure 9] FIG. 13 is a diagram illustrating bright point labeling in an elemental image group. [Figure 10] 13A to 13C are diagrams illustrating template selection when one element image contains multiple bright spots and bright spot labeling in the second and subsequent element image groups. [Figure 11] 11 is a diagram illustrating the interval between representative points of bright points between adjacent elemental images in an elemental image group. FIG. [Figure 12] 13 is a graph showing the tracking performance in the Z direction. [Figure 13] FIG. 2 is a diagram showing a state in which an object is tracked by an image analyzing device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] Hereinafter, the embodiments of the present invention will be described in detail with reference to the drawings as appropriate. However, more detailed explanations than necessary may be omitted. For example, detailed explanations of already well-known matters or duplicate explanations of substantially identical configurations may be omitted. This is to avoid the following explanation becoming unnecessarily redundant and to facilitate understanding by those skilled in the art. Note that the inventor provides the accompanying drawings and the following explanation so that those skilled in the art can fully understand the present invention, and does not intend to limit the subject matter described in the claims by them. In addition, the dimensions of each member depicted in the drawings, detailed shapes of details, etc. may differ from the actual ones.
[0015] <Embodiment> 1 is a schematic diagram of an optical system including an image analysis device according to an embodiment of the present invention. The optical system 100 includes an optical system 10, a storage device 20, and an image analysis device 30. The storage device 20 may be located on a cloud. In addition, all or a part of the components of the image analysis device 30 may also be located on a cloud.
[0016] The optical system 10 includes a main lens 1, a microlens array 2, and an image sensor 3. The main lens 1 may include an objective lens and an imaging lens, but for convenience, the lens including these is referred to as the main lens 1. The microlens array 2 is configured by arranging a plurality of microlenses (lenslets) 2a two-dimensionally, and is disposed between the main lens 1 and the image sensor 3. The image sensor 3 is a CCD sensor or a CMOS sensor configured by arranging a plurality of light receiving elements (not shown) two-dimensionally, which perform photoelectric conversion of input light and output an electrical signal. The electrical signal output from each light receiving element is converted into digital data by an A / D converter (not shown), and a light field image is output from the optical system 10 as multi-viewpoint image data.
[0017] The storage device 20 is a collection of storage devices such as RAM, ROM, SSD, and HDD. The RAM is mainly used as a working memory when the image analysis device 30 performs image analysis. The ROM, SSD, and HDD mainly store computer programs for operating the image analysis device 30, moving image data to be analyzed by the image analysis device 30, reconstructed images, PQ arrays, three-dimensional images, label tables, etc., which are generated and used in the image analysis described below, temporarily or for a long period of time.
[0018] The image analysis device 30 analyzes the video data stored in the storage device 20 and tracks the three-dimensional spatial position of the object in the image. Specifically, the image analysis device 30 includes an image reconstruction unit 31, a template selection unit 32, a bright point representative point identification unit 33, a bright point tracking unit 34, a bright point labeling unit 35, and a three-dimensional coordinate update unit 36. Each of these components can be realized as hardware such as ASIC or FPGA, or as software that executes a computer program stored in the storage device 20 by a processor (not shown) such as a CPU or GPU. It can also be realized by appropriately combining hardware and software. The image analysis device 30 appropriately accesses various storage devices constituting the storage device 20 during image analysis to write and read data.
[0019] The moving image data to be analyzed by the image analysis device 30 is image data in which a group of elemental images, each of which is an array of elemental images of the same subject captured from a plurality of mutually shifted viewpoints, is a sequence of such frame images in time. From such a group of elemental images, an image refocused at an arbitrary depth position can be reconstructed by utilizing the parallax of each elemental image contained therein. The moving image data can be generated, for example, from multi-viewpoint image data (light field image) continuously output from the optical system 10.
[0020] FIG. 2 is a diagram showing an example of multi-viewpoint image data, an element image, and a reconstructed image. (a) Multi-viewpoint image data is obtained by capturing an image of a sheet with the characters " / 5" written on it, placed at a position shifted from the focal plane of the optical system 10. The multi-viewpoint image data is a group of microlens images in which a large number of microlens images MI captured by each microlens 2a of the microlens array 2 of the optical system 10 are arranged in a square lattice pattern. As shown by the dashed line, the shape of each microlens image MI is approximately circular, reflecting the circular shape of each microlens 2a. In this way, the multi-viewpoint image data is a collection of a plurality of sub-images (microlens images MI in the example of FIG. 2) captured of the same subject from a plurality of viewpoints shifted from each other.
[0021] The area near the boundary of the microlens image MI is dark due to the small amount of light, and the central part of the microlens image MI is used because the area is significantly distorted due to the light rays that pass through the edge of the microlens 2a. That is, from each of the roughly circular microlens images MI, for example, a square area inscribed in the circle is cut out as an element image EI, and the element images EI are arranged in a square lattice to form the (b) element image group.
[0022] As described above, since the sheet on which the characters " / 5" are written is located at a position shifted from the focal plane of the optical system 10, (a) in the multi-viewpoint image data, microlens images are recorded in which the characters " / 5" are partially captured by shifting the viewpoint using each microlens 2a, that is, microlens images with parallax are recorded. For example, when focusing on the upper left corner of the character "5," the light beam emitted from that part of the subject is recorded with a distance Δ between adjacent microlens images. As shown in (b) in the elemental image group, the distance Δ between the microlens images is Δ between the elemental images. EI Here, the pitch of the microlenses 2a is MLP, and the length of one side of the element image is EI L Then, Δ EI is expressed by the following equation (1). Δ EI =Δ-(MLP-EI L ) …(1)
[0023] The reconstructed image (c) is obtained by superimposing adjacent elemental images in the elemental image group while shifting them by an amount corresponding to the parallax. In the example of FIG. 2, adjacent elemental images are shifted by Δ EI By shifting the elemental images by an appropriate amount and superimposing them, a reconstructed image (c) is obtained, refocused at the position of the sheet with the letters " / 5" written on it. In this way, by superimposing the elemental images while appropriately changing the shift amount, it is possible to reconstruct an image refocused at any depth position based on ray tracing in the reverse direction from the acquired image.
[0024] Here, the depth of the object to be refocused, which is the deviation from the object-side focal plane NOP (Native Object Plane) of the objective lens in the main lens 1, is d obj Then, the distance Δ between adjacent microlens images of light rays originating from the same point and the depth d of that point are obj The relationship between the parameters of the optical system 10 and the depth d can be expressed by using the parameters of the optical system 10. FIG. 3 is a diagram showing a schematic diagram of the geometric relationship between the parameters of the optical system 10 according to an example. In the diagram, O represents the depth d obj A is the image point of the object O formed by the main lens 1, B and C are the centers of the adjacent microlenses 2a, and B' and C' are the positions of the light rays from the object O captured by the image sensor 3 via B and C. When we look at the triangles ABC and AB'C' in the figure, from the similarity of the triangles, BC:B'C'=CA:C'A This can be expressed as parameters as follows: MLP: Δ = (a + d obj M 2 ):(a+d obj M 2 -b) Solving this for Δ gives us the following equation (2). Δ=MLP(1-b / (d obj M 2 -a)) …(2) where MLP is the pitch of the microlenses 2a, M is the magnification of the objective lens of the main lens 1, a is the distance from the image-side focal plane T-NIP (Native Image Plane) of the imaging lens of the main lens 1 to the center of the microlens 2a, and b is the distance from the center of the microlens 2a to the image sensor 3.
[0025] In addition, since there is no anisotropy in two mutually perpendicular directions (X direction, Y direction) in a plane perpendicular to the depth direction (Z direction), the same Δ can be applied to both the X direction and the Y direction.
[0026] From equations (1) and (2), obj and Δ EI Therefore, at any depth position (depth d obj ), the elemental images are shifted by the amount of shift Δ EI The elemental images can be easily overlapped by referring to the PQ array described below.
[0027] FIG. 4 is a diagram for explaining the relationship between the PQ array and the overlapping of element images. Overlapping pixels are shown in a shaded manner, the left column shows the overlapping state of the element images, and the right column shows the overlapping state of the element images. For convenience, the element image group is composed of four element images D, E, F, and G, arranged in two rows and two columns, and each element image is five pixels in size. The element images are arranged such that element image E is to the right of element image D, element image F is below element image D, and element image G is below element image E, i.e., to the right of element image F. In order to refer to each pixel in the element image group, consecutive numbers starting from 1 are assigned to the rows and columns. For example, pixels belonging to element image D are referred to in the range from the first row to the fifth row and the first column to the fifth column, and pixels belonging to element image G are referred to in the range from the sixth row to the tenth row and the sixth column to the tenth column. In the following, pixel [m, n] refers to the pixel located in the mth row and the nth column in the element image group.
[0028] Z in the figure is d objFor convenience, d obj and the shift amount of the element image Δ EI are expressed by the same Z value. That is, when Z=n, d obj = n, and the shift amount of the element image Δ EI represents n pixels.
[0029] When Z=0, that is, at depth d obj If =0, the elemental images are not superimposed on each other, and the elemental images become the reconstructed image as they are.
[0030] When Z=1, that is, at depth d obj To reconstruct an image refocused to Z=1, adjacent elemental images are shifted by one pixel and then superimposed. Which pixels in the elemental image group should be superimposed can be easily determined by referring to the PQ array, which shows how the pixels are superimposed. The PQ array is a representation of the row or column numbers of pixels in the elemental image group folded over on an elemental image basis, and as shown in the figure, for example, the PQ array corresponding to Z=1 is expressed as follows: 1 2 3 4 5 6 7 8 9 10 This PQ array means that the pixel in the fifth row (fifth column) and the pixel in the sixth row (sixth column) overlap. By preparing such a PQ array for each Z value in advance and storing it in the storage device 20, it is possible to perform pixel overlap by referring to the PQ array corresponding to the Z value.
[0031] When Z=2, that is, at depth d obj To reconstruct an image refocused to Z = 2, adjacent elemental images are shifted by two pixels and overlapped. As shown in the figure, for example, the PQ array corresponding to Z = 2 is expressed as follows: 1 2 3 4 5 6 7 8 9 10 This PQ array means that the pixels in the 4th and 5th rows (4th and 5th columns) overlap with the pixels in the 6th and 7th rows (6th and 7th columns), respectively.
[0032] Note that there may be cases where three or more element images overlap in the column and row directions, in which case the PQ array will be composed of three or more rows. Also, when the matrix of element images is composed of the same number of pixels as in the above example, a common PQ array can be used for overlapping element images in the row direction and the column direction, but when the matrix of element images has a different number of pixels, a PQ array in the row direction and a PQ array in the column direction must be prepared.
[0033] Next, image analysis by the image analysis device 30 will be described. Fig. 5A is the first half of a flowchart of an image analysis method according to one embodiment of the present invention, showing the image analysis procedure for the first group of element images (corresponding to time t=1). Fig. 5B is the second half of a flowchart of an image analysis method according to one embodiment of the present invention, showing the image analysis procedure for the second and subsequent groups of element images (corresponding to time t=2 and after). Each step is executed by the above-mentioned components of the image analysis device 30.
[0034] When the first group of element images in the video data is read from the storage device 20 to the image analysis device 30, the image reconstruction unit 31 reconstructs an image refocused on the depth position of the object from the group of element images (S1). When the reconstructed image is generated, the bright spot labeling unit 35 attaches a label for identifying the object to a representative point of each bright spot consisting of a group of pixels corresponding to the object in the reconstructed image (S2). The bright spot here refers to a group of one or more pixels having a certain luminance or pixel value, that is, a point or a small image area having a certain brightness. After the labeling is completed, the three-dimensional coordinate update unit 36 obtains the three-dimensional space coordinates of the object in the reconstructed image and sets the coordinates as the initial coordinates of the object identified by the label (S3). In parallel or asynchronously with these steps S1 to S3, the template selection unit 32 selects the element image containing the most bright spots from the group of element images as a template (S4).
[0035] 6 is a diagram for explaining an example of three-dimensionally reconstructing an object from an initial group of elemental images, labeling bright spots, setting initial coordinates of the object, and selecting a template. For example, the image reconstruction unit 31 appropriately refers to the PQ array stored in the storage device 20, shifts the elemental images in the group of elemental images by a shift amount according to the depth position, and overlaps them to reconstruct an image refocused at a plurality of depth positions. Depth d obj As the size of the element image increases, the shift amount Δ EI becomes larger, and the size of the reconstructed image becomes smaller accordingly. The image reconstruction unit 31 resizes the reconstructed images, the sizes of which vary depending on the depth position, according to the depth position and stacks them in the Z direction (depth direction) to reconstruct a three-dimensional image of the target object. Note that, as a method for reconstructing an image refocused at an arbitrary depth position from a group of elemental images, a high-resolution image can be reconstructed by using the optical sectioning technology disclosed in the application of the present inventors (Patent Application No. 2022-202566).
[0036] In order to identify the pixel corresponding to the three-dimensional image in the element image group, for example, the bright spot labeling unit 35 multiplies the XY coordinates of each voxel of the three-dimensional image by the resize ratio of the image at each depth position to convert it into the pixel coordinates of the first element image group. For example, if the size of an image of a certain Z coordinate in the three-dimensional image is m and the size of the image before resizing is n, the XY coordinates of each voxel of the Z coordinate can be multiplied by the resize ratio n / m to convert it into the corresponding pixel coordinates in the first element image group. However, since the coordinates of other pixels overlapping with the converted pixel coordinates are not known, the bright spot labeling unit 35 refers to the PQ array to identify the coordinates of other pixels overlapping with the pixel of the converted pixel coordinates when reconstructing an image refocused at that depth position, and attaches a label that identifies the object to the representative point of each bright spot consisting of the pixel of the converted pixel coordinates and a collection of pixels overlapping with the pixel when reconstructing an image refocused at that depth position. For example, each pixel constituting a bright spot is binarized, its center of gravity is found, and the center of gravity can be set as the representative point of the bright spot. The same label is attached to representative points of bright spots corresponding to the same three-dimensional image.
[0037] The correspondence between the labels and the representative points of each bright point is tabulated and stored as a label table in the storage device 20. The labels are also associated with the three-dimensional coordinates of the object. Specifically, the three-dimensional coordinate update unit 36 obtains the coordinates of each voxel of the three-dimensional image, and obtains, for example, the center of gravity of the three-dimensional image as the coordinates of the three-dimensional space, and sets the center of gravity as the initial coordinates of the object identified by the label.
[0038] In parallel or asynchronous with the above image reconstruction, bright spot labeling, and initial coordinate setting of the object, the template selection unit 32 selects the element image containing the most bright spots as the template. For convenience, since the object is one in the example of FIG. 6, the number of bright spots in each element image is one. In this way, when there are multiple element images that are candidates for the template, the element image closest to the origin in the element image group, that is, the element image at the top left can be selected as the template. The template is referred to in the bright spot labeling process of the second element image group. This completes the analysis process for the first element image group.
[0039] Returning to FIG. 5B, when the next group of elemental images in the video image data is read from the storage device 20 to the image analysis device 30, the bright spot representative point identification unit 33 determines the representative point of each bright spot in the group of elemental images (S5). FIG. 7 is a diagram for explaining the determination of the representative point of each bright spot in the group of elemental images. Since the group of elemental images is a grayscale image, the bright spot representative point identification unit 33 smoothes the group of elemental images using an averaging filter as necessary and then performs binarization processing on the group of elemental images. The bright spot representative point identification unit 33 then determines the center of gravity of each binarized bright spot and sets the center of gravity as the representative point of that bright spot. Thereafter, labeling and tracking are performed on the representative point of each bright spot.
[0040] The threshold for binarization may be an appropriate fixed value, or may be determined according to the characteristics of the elemental image group. In the latter case, first, the approximate ratio of the number of pixels in which the object appears is calculated based on the range of the field of view and the size of the object. For example, if the field of view is 422.4 μm × 422.4 μm and the object is a circle with a diameter of 4 μm, the ratio of the pixels in which the object appears in the field of view (in the elemental image group) is approximately 0.0070%. Next, the pixel value and its frequency percentage of each pixel in the elemental image group are calculated, and the frequencies are accumulated in descending order until the ratio of the number of pixels in which the object appears is reached, and the pixel value corresponding to the last added frequency is set as the threshold. For example, in the above example, the pixel values are accumulated from the highest frequency, and the pixel value when the accumulated value exceeds 99.993% is set as the threshold. This method of determining the threshold is effective when the contrast is high, such as in a microscope image, the background and the object are clearly separated, and the ratio of the object to the entire field of view is very small. If the object can be observed in advance, the pixel values of the background and object can be obtained in advance using appropriate image analysis software, and the obtained values can be used as the threshold value. In addition, Otsu's binarization, which is a general automatic binarization method, is difficult to apply when the background and the object are clearly separated and the proportion of the object in the entire field of view is very small, but it can be applied depending on the object.
[0041] Returning to FIG. 5B, when the representative point of each bright point in the element image is determined, the bright point tracking unit 34 obtains the position change of the representative point of each bright point in the element image group (S6). FIG. 8 is a diagram for explaining the tracking of the representative point of each bright point in the element image group. For convenience, in FIG. 8, two element image groups that are successive in time are superimposed. The bright point tracking unit 34 finds the destination of the representative point of each bright point by nearest neighbor search between the element image groups that are successive in time. If the frame rate of the moving image data is sufficiently high with respect to the moving speed of the target object, as shown in FIG. 8, the deviation of the coordinates of the representative point of each bright point between the element image groups that are successive in time is very small, and the destination of the representative point of each bright point can be obtained by nearest neighbor search. For example, as shown in the enlarged view, the destinations of the representative points P1 to P5 of the five bright points are found by nearest neighbor search between the element image groups that are successive in time. When the destination of the representative point of each bright point is found, the bright point tracking unit 34 obtains the position change of each bright point from the positional relationship before and after the movement of the representative point of each bright point. For example, for the representative point P1 of the bright point, the XY coordinates at time t-1 are P1 tー1 =(xP1 tー1 ,yP1 t-1 ) and the XY coordinates at time t are P1 t =(xP1 t ,yP1 t ), then the position change of P1 is P1 t -P1 tー1 In the enlarged image, the representative point below P1 and to the right of P3 appears for the first time at time t, so it cannot be tracked from time t-1 to time t. If this representative point continues to appear after time t, its position will be tracked.
[0042] Returning to FIG. 5B, in parallel with tracking the representative points of each bright point, the bright point labeling unit 35 assigns the same label as the representative point of each bright point in the template to the representative point of each bright point in each element image of the element image group that corresponds to the representative point of each bright point in the template (S7). FIG. 9 is a diagram for explaining bright point labeling in element image groups. The left diagram shows the kth element image group, the center diagram shows the k+1th element image group in which the representative point of the bright point is nearest neighbor searched using the template, and the right diagram shows the k+1th element image group after labeling processing. For convenience, it is assumed that there is one representative point of the bright point in each element image, and the same label is assigned to the representative point of each bright point in the kth element image group in the left diagram. At this time, as shown in the center diagram, the bright point labeling unit 35 finds the representative point of the bright point corresponding to the representative point P0 of each bright point in the template in each element image in the k+1th element image group by nearest neighbor search. For convenience, the element image group in the center diagram shows a template superimposed on each element image in the k+1th element image group. In this example, each element image has only one representative point of a bright spot, so each representative point is found as a representative point corresponding to the representative point of the bright spot in the template. When the representative point of the corresponding bright spot is found in each element image, the bright spot labeling unit 35 assigns the same label as the representative point P0 of the corresponding bright spot in the template to the found representative point. As a result, as shown on the right side of the diagram, a label is assigned to the representative point of each bright spot in the k+1th element image group. In this way, a label is assigned to the representative point of each bright spot in the new element image group.
[0043] When there are multiple objects, the element image may contain multiple bright spots derived from the objects. In such a case, the bright spot labeling unit 35 may label element images containing a predetermined number or more of bright spots. For example, if the number of bright spots in the template is N, element images containing N-1 or more bright spots are labeled. This reduces the number of element images to be labeled, thereby reducing the load related to the labeling process. FIG. 10 is a diagram for explaining template selection when one element image contains multiple bright spots and bright spot labeling in the second and subsequent element image groups. The left diagram shows the first element image group, the center diagram shows the m-th element image group, and the right diagram shows the n (n>m)-th element image group. The template selection unit 32 selects the element image containing the most bright spots as the template from the first element image group. In this example, an element image containing four bright spots is selected as the template, but there are multiple such element images. In such a case, the template selection unit 32 selects the element image with the largest degree of separation between the bright spots as the template. If the degree of separation is expressed by SEP, the SEP can be defined, for example, by the following formula. SEP=X max -X min +Y max -Y min However, X max is the maximum X coordinate among the XY coordinates of the representative points of multiple bright points included in the element image of the template candidate, and X min is the minimum X coordinate, Y max is the maximum Y coordinate, Y min is the minimum Y coordinate.
[0044] As described above, when cutting out an elemental image from a microlens image, a part of the microlens image is cut, but when the target object moves, the bright spots included in the cut area move into the elemental image, and the number of bright spots in the elemental image may increase or decrease over time. For example, when focusing on the elemental image enlarged in FIG. 10, the number of bright spots in the elemental image is 1 in the first elemental image group (left), increases to 2 in the m-th elemental image group (center), and increases to 4 in the n-th elemental image group (right). In the m-th elemental image group, the number of bright spots in the elemental image increases to 2, but the condition of N-1 or more is not satisfied, so the elemental image is not subject to labeling. When the number of bright spots in the elemental image increases to 4 in the n-th elemental image group, the condition of N-1 or more is satisfied, so the elemental image is subject to labeling. For the elemental images to be labeled, the bright point labeling unit 35 assigns the same label as that of each bright point in the template to the bright point corresponding to the bright point in the template.
[0045] Returning to FIG. 5B, after tracking and labeling of the representative points of each bright spot, the three-dimensional coordinate update unit 36 updates the XY coordinates of the object identified by the label based on the position change of the representative points of the bright spots with the same label, and updates the Z coordinate of the object based on the interval between the representative points between adjacent element images (S8). For example, the XY coordinates of the object at time t can be obtained by adding the average of the position change of the representative points of the bright spots with the same label from time t-1 to time t to the XY coordinates at time t-1. This can be expressed by the following formula. (x t ,y t )=(x t-1 ,y t-1 )+〈P t -P t-1 〉 However, (x t ,y t ) is the XY coordinate of the object at time t, (x t-1 ,y t-1 ) are the XY coordinates of the object at time t-1, P tare the XY coordinates of the representative point of the bright point with the same label at time t, P t-1 are the XY coordinates of the representative point of the bright points with the same label at time t-1, and the operator 〈〉 represents the average value of each XY coordinate.
[0046] The Z coordinate of the target object at time t can be directly obtained based on the interval between representative points of bright points with the same label in the element image group at time t between adjacent element images. Fig. 11 is a diagram for explaining the interval between representative points of bright points between adjacent element images in the element image group. For example, as shown in the enlarged view, the coordinates of representative points P1 to P5 of five bright points in the element image group are known in advance, so the interval Δ EI1 , Δ EI2 , Δ EI3 , Δ EI4 is easily obtained. For the representative points of the same labeled bright points at time t, the interval Δ EI1 , Δ EI2 , Δ EI3 , Δ EI4 , …, Δ EIn The average of  ̄ Δ EIt Then, from equations (1) and (2), the Z coordinate of the object identified by the label at time t is expressed by the following equation: z t =d obj =((  ̄ Δ EIt -EI L )a-MLP*b) / (  ̄ Δ EIt -EI L )M 2
[0047] FIG. 12 is a graph showing the tracking performance in the Z direction. The tracking performance in the Z direction was confirmed by capturing images of a fluorescent particle with a radius of 2 μm being moved in the Z direction from 0 μm to 40 μm in 1 μm increments on an XYZ motorized stage, and analyzing the moving image data with an image analyzer 30. (a) shows the input to the motorized stage carrying the target (fluorescent particle), and (b) shows the result of tracking the position of the target (fluorescent particle) by the image analyzer 30. The reason why position tracking is not possible near Z=0 (NOP) in (b) is that refocusing is difficult at that part due to the characteristics of the optical system 10. It can be said that the movement of the target in the Z direction is well tracked for other depth positions where refocusing is possible.
[0048] Effect
[0049] The image analysis device 30 according to this embodiment performs three-dimensional reconstruction of the object from the elemental image group at the beginning (time t=1) to obtain the initial coordinates, but thereafter (time t=2 and after), the three-dimensional spatial position of the object is tracked only by analyzing the elemental image group of the two-dimensional image without three-dimensionally reconstructing the object from the elemental image group. Fig. 13 is a diagram showing the state of object tracking by the image analysis device 30. For example, after determining the initial coordinates (x, y, z) = (553, 362, 44) of particle A, which is the object, from the first elemental image group, the three-dimensional spatial position of particle A is tracked by analyzing the elemental image group of the two-dimensional image without three-dimensionally reconstructing particle A.
[0050] In this way, the image analyzing device 30 according to the present embodiment can track the three-dimensional spatial position of an object in an image from high frame rate video data such as continuous data of a group of elemental images, making it possible to observe high-speed events such as biomolecules and intracellular behavior. As an example, the three-dimensional spatial position of an object can be tracked by analyzing video data of 100 fps.
[0051] In addition, the data size of the moving image data, which is a series of elemental images that are two-dimensional images, is overwhelmingly smaller than the continuous data of volumetric image data obtained by a confocal microscope, so that the data size does not put a strain on storage capacity for storage, and transmission by wireless communication or the like is possible. As an example, the data size of the above-mentioned 100 fps moving image data is 1.27 MB, while the total data size of the volumetric image reconstructed three-dimensionally from each elemental image group is 1.53 GB. In other words, the image analysis device 30 according to this embodiment can track the three-dimensional spatial position of the target object from the continuous data of the elemental image group that is a two-dimensional image without requiring temporally continuous volumetric images, and therefore can reduce the data size related to storage and transmission to approximately 1 / 1000.
[0052] <<Variations>> Instead of reconstructing a three-dimensional image of the object and determining the initial coordinates of the object from the three-dimensional image as explained in Fig. 6, the three-dimensional coordinates of the object may be directly obtained from an image reconstructed from the initial elemental image group. Since the Z coordinate of the reconstructed image is known in advance, the initial coordinates of the object in three-dimensional space can be determined by determining the center of gravity coordinates of the object in the reconstructed image, i.e., the XY coordinates.
[0053] The template selection unit 32 may update the template by appropriately selecting a template not only from the first element image group but also from the second and subsequent element image groups. In particular, when there are multiple objects that rotate or change direction, it is preferable to reselect a template from the second and subsequent element image groups.
[0054] The above explanation is based on the premise that the element images are arranged in a square lattice in the element image group, but when a microlens array in which microlenses are arranged in a honeycomb pattern is used, the element images may be arranged in a honeycomb pattern. Even in such a case, the above-mentioned identification of representative points of bright spots in the element image group, bright spot labeling, template selection, bright spot tracking, and three-dimensional coordinate updating are applicable.
[0055] The multi-viewpoint image data is not limited to a light field image, but may be a collection of multiple sub-images of the same subject captured from multiple viewpoints that are offset from one another. For example, it may be captured by a multi-camera system in which multiple cameras are arranged. In addition, the images captured by each camera in the multi-camera system do not need to be compiled into a single image data set, and the images captured by each camera may be stored in separate files. In this case, the set of image files is the multi-viewpoint image data. From such multi-viewpoint image data, the image data to be analyzed by the image analysis device 30 can be generated.
[0056] As described above, the embodiment has been described as an example of the technology of the present invention. For this purpose, the attached drawings and detailed description have been provided. Therefore, among the components described in the attached drawings and detailed description, not only components essential for solving the problem but also components that are not essential for solving the problem in order to illustrate the above technology may be included. Therefore, just because those non-essential components are described in the attached drawings or detailed description, it should not be immediately recognized that those non-essential components are essential. In addition, since the above-mentioned embodiment is intended to illustrate the technology of the present invention, various modifications, substitutions, additions, omissions, etc. may be made within the scope of the claims or their equivalents. [Industrial Applicability]
[0057] The image analysis device of the present invention can track the three-dimensional spatial position of an object in an image from high frame rate video data, such as continuous data of a group of elemental images, and is therefore useful for observing high-speed events such as biomolecules and intracellular behavior. [Explanation of symbols]
[0058] 30 Image analysis device 31 Image reconstruction unit 32 Template Selection Section 33 Bright spot representative point identification part 34 Bright Spot Tracking Unit 35 Bright spot labeling section 36 3D coordinate update section
Claims
1. An image analysis device that analyzes moving image data in which a group of elemental images, each of which is an array of elemental images obtained by capturing an identical subject from a plurality of mutually shifted viewpoints, is continuous in time, and tracks a three-dimensional spatial position of an object in the image, comprising: an image reconstruction unit that reconstructs an image refocused at a depth position of the object from the initial elemental image group; a template selection unit that selects an element image including the most bright points from the initial element image group as a template; a bright point representative point specifying unit for determining a representative point of each bright point in the second and subsequent element image groups; a bright point tracking unit for determining a change in position of a representative point of each bright point in the second and subsequent elemental image groups; a bright point labeling unit that assigns a label for identifying the object to a representative point of each bright point consisting of a group of pixels corresponding to the object in the reconstructed image in the first element image group, and assigns the same label as that of each bright point in the template to a representative point of each bright point corresponding to the representative point of each bright point in the template in each element image in the second and subsequent element image groups; a three-dimensional coordinate updating unit which obtains three-dimensional spatial coordinates of an object in the reconstructed image and sets the coordinates as initial coordinates of the object identified by the label, updates the XY coordinates of the object identified by the label based on positional changes of representative points of bright points with the same label in second and subsequent elemental image groups, and updates the Z coordinate of the object based on the intervals between the representative points of adjacent elemental images; An image analysis device comprising:
2. the image reconstruction unit shifts each element image in the initial element image group by a shift amount according to the depth position, superimposes the element images, and reconstructs images refocused at a plurality of depth positions, resizes the reconstructed images according to the depth positions, and stacks them in the Z direction to reconstruct a three-dimensional image of the object; The bright point labeling unit converts the XY coordinates of each voxel of the three-dimensional image into pixel coordinates of the initial elemental image group by multiplying the XY coordinates of each voxel of the three-dimensional image by the resizing ratio of the image at each depth position, and attaches a label for identifying an object to a representative point of each bright point consisting of a pixel of the converted pixel coordinates and a group of pixels that are superimposed on the pixel when reconstructing an image refocused at the depth position.
2. The image analysis device according to claim 1 .
3. The template selection unit selects, as a template, an element image having the largest degree of separation between the bright points among element images including a plurality of bright points.
2. The image analysis device according to claim 1 .
4. The template selection unit selects a template from the second or subsequent element image group and updates the template.
2. The image analysis device according to claim 1 .
5. The bright point labeling unit finds a representative point of a bright point corresponding to a representative point of each bright point in the template by nearest neighbor search in each element image in the second and subsequent element image groups.
2. The image analysis device according to claim 1 .
6. The bright point labeling unit labels element images including a predetermined number or more of bright points in the second and subsequent element image groups.
2. The image analysis device according to claim 1 .
7. The bright point representative point specifying unit performs binarization processing on the elemental image group, obtains the center of gravity of each binarized bright point, and sets the center of gravity as the representative point of the bright point.
2. The image analysis device according to claim 1 .
8. The bright point tracking unit finds the movement destination of the representative point of each bright point by nearest neighbor search between the element image groups that are successive in time, and obtains the position change of the representative point of each bright point from the positional relationship before and after the movement of the representative point of each bright point.
2. The image analysis device according to claim 1 .
9. 1. An image analysis method for tracking a three-dimensional spatial position of an object in an image by analyzing moving image data in which a group of elemental images, in which a plurality of elemental images of the same object are arranged and captured from a plurality of mutually shifted viewpoints, are successive in time, the method comprising: Reconstructing an image refocused at a depth position of the object from the initial elemental images; A step of attaching a label for identifying an object to a representative point of each bright point, which is a group of pixels corresponding to an object in a reconstructed image, in the first elemental image group; determining three-dimensional spatial coordinates of the object in the reconstructed image and setting the coordinates as initial coordinates of the object identified by the label; selecting an element image including the most bright points from the initial element image group as a template; determining a representative point of each bright point in the second and subsequent elemental image groups; A step of determining a position change of a representative point of each bright point in the second and subsequent elemental image groups; labeling representative points of bright points corresponding to representative points of bright points in the template in each element image of the second and subsequent element image groups with the same label as that of the representative points of bright points in the template; updating the XY coordinates of the object identified by the label based on a change in position of a representative point of a bright point with the same label in the second and subsequent element image groups, and updating the Z coordinate of the object based on an interval between the representative points of the same label in adjacent element images; The image analysis method includes:
10. A program for causing a computer to analyze video data in which a group of elemental images, each of which is an array of elemental images obtained by capturing an identical subject from a plurality of mutually shifted viewpoints, is continuous in time, and to track a three-dimensional spatial position of an object in the image, an image reconstruction means for reconstructing an image refocused at a depth position of the object from the initial elemental image group; a template selection means for selecting an element image including the largest number of bright points from the initial element image group as a template; A bright point representative point specifying means for determining a representative point of each bright point in the second and subsequent element image groups; A bright point tracking means for determining a change in position of a representative point of each bright point in the second and subsequent elemental image groups; a bright point labeling means for labeling a representative point of each bright point consisting of a group of pixels corresponding to an object in a reconstructed image in a first elemental image group, the representative point being a group of pixels corresponding to the object in the reconstructed image, and labeling a representative point of each bright point corresponding to a representative point of each bright point in the template in each elemental image in a second or subsequent elemental image group with the same label as that of the representative point of each bright point in the template; a three-dimensional coordinate updating means for determining three-dimensional spatial coordinates of an object in the reconstructed image and setting the coordinates as initial coordinates of the object identified by the label, updating the XY coordinates of the object identified by the label based on positional changes of representative points of bright points with the same label in second and subsequent elemental image groups, and updating the Z coordinate of the object based on the intervals between the representative points between adjacent elemental images; A program that makes a computer function as a