Information processing device and information processing method
By extracting flat areas from microscope images using an information processing device and employing multi-source reflected light and machine learning methods, the problem of microscopes being unable to accurately measure minute shapes has been solved, achieving high-precision 3D image reconstruction.
Patent Information
- Application Number
- CN202080077511.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-15
- Filing Date
- 2020-11-09
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-11-09
AI Technical Summary
Existing technologies struggle to capture 3D images of minute shapes with high accuracy using microscopes, and non-contact 3D measuring instruments are costly, while the Time-of-Flight (ToF) method lacks sufficient accuracy.
The system acquires target images from sensors using an information processing device, extracts flat areas using reflected light emitted from multiple light sources, calculates the shape information of the target surface based on brightness values and sensor information, and generates normal and depth images using machine learning methods.
It improves the accuracy of shape measurement, suppresses the decrease in the accuracy of depth calculation, and achieves high-precision 3D image reconstruction.
Smart Images

Figure CN114667431B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to information processing apparatus and information processing methods. Background Technology
[0002] Microscopes, which are relatively inexpensive and can be easily used to perform measurements, have been widely used as devices for observing objects in minute forms.
[0003] Techniques for analyzing the color and dirt on the skin surface by using differences in the incident angle of the illumination unit (e.g., Patent Document 1) are known to be microscope-related techniques. Furthermore, techniques for reducing defocus and distortion when capturing images of the skin surface by arranging transparent glass at a predetermined distance from the distal dome of the microscope are known (e.g., Patent Document 2).
[0004] Citation List
[0005] Patent documents
[0006] Patent Document 1: JP H10-333057A
[0007] Patent Document 2: JP 2008-253498 A Summary of the Invention
[0008] Technical issues
[0009] According to conventional techniques, the quality of images captured by a microscope can be improved.
[0010] However, conventional techniques only improve the quality of planar images and struggle to obtain 3D images that reproduce the minute shapes (bumps and ridges) of an object. Note that non-contact 3D measuring instruments and 3D scanners are used as devices for measuring the minute shapes of objects, but their introduction comes at the cost of relatively high prices. Also note that distance measuring devices using the Time-of-Flight (ToF) method are relatively inexpensive, but their accuracy may be insufficient.
[0011] Therefore, this disclosure provides an information processing apparatus and an information processing method that can improve the accuracy of shape measurement.
[0012] Solution to the problem
[0013] According to this disclosure, an information processing apparatus is provided. The information processing apparatus includes a control unit. The control unit acquires a captured image of a target imaged by a sensor. The captured image is an image obtained based on reflected light emitted from multiple light sources respectively arranged at different locations and directed to the target. The control unit extracts a flat region from the captured image based on its brightness values. The control unit calculates shape information about the surface shape of the target based on the flat region of the captured image and information about the sensor. Attached Figure Description
[0014] Figure 1 This is a diagram used to describe a method for calculating depth based on captured images.
[0015] Figure 2 It is a table used to describe the approximate shape of the surface of a target.
[0016] Figure 3 This is a diagram that serves as an overview of a first embodiment of the present disclosure.
[0017] Figure 4 This is a block diagram illustrating an example configuration of an information processing system according to a first embodiment of the present disclosure.
[0018] Figure 5 This is a diagram illustrating an example configuration of a microscope according to a first embodiment of the present disclosure.
[0019] Figure 6 This is a block diagram illustrating an example configuration of an information processing apparatus according to a first embodiment of the present disclosure.
[0020] Figure 7 This is a diagram showing an example of the brightness values of a captured image.
[0021] Figure 8 This is a diagram showing an example of a mask image generated by a region acquisition unit.
[0022] Figure 9 This is a diagram used to describe an example of obtaining a flat region by a region acquisition unit.
[0023] Figure 10 This is an example diagram used to describe normal information.
[0024] Figure 11 This is a diagram used to describe a learning machine according to an embodiment of the present disclosure.
[0025] Figure 12 It is a graph used to describe the length of each pixel in a captured image.
[0026] Figure 13 This is a diagram showing an example of an image displayed on a display unit by a display control unit.
[0027] Figure 14 This is a flowchart illustrating an example of depth computing processing according to a first embodiment of the present disclosure.
[0028] Figure 15 This is a diagram used to describe a captured image acquired by an acquisition unit according to a second embodiment of the present disclosure.
[0029] Figure 16 This is a diagram used to describe a captured image acquired by a region acquisition unit according to a second embodiment of the present disclosure.
[0030] Figure 17 This is a diagram used to describe the smoothing process performed by the region acquisition unit according to the third embodiment of this disclosure.
[0031] Figure 18 This is a diagram used to describe the acquisition of a flat region by a region acquisition unit according to the fourth embodiment of the present disclosure.
[0032] Figure 19 This is a diagram used to describe the acquisition of a flat region by a region acquisition unit according to the fourth embodiment of the present disclosure.
[0033] Figure 20 It is a graph used to describe the depth of field of the sensor.
[0034] Figure 21 This is a diagram used to describe the acquisition of a flat region by a region acquisition unit according to the fifth embodiment of the present disclosure.
[0035] Figure 22 This is a diagram used to describe the acquisition of a flat region by a region acquisition unit according to the fifth embodiment of the present disclosure.
[0036] Figure 23 This is a diagram used to describe a plurality of captured images acquired by an acquisition unit according to a sixth embodiment of the present disclosure.
[0037] Figure 24 This is a diagram used to describe the acquisition of a flat region by a region acquisition unit according to the sixth embodiment of this disclosure.
[0038] Figure 25 This is a diagram illustrating an example configuration of an information processing system according to a seventh embodiment of the present disclosure.
[0039] Figure 26 This is a diagram used to describe frequency separation performed by the normal frequency separation unit according to the seventh embodiment of this disclosure.
[0040] Figure 27 This is a diagram illustrating an example configuration of an information processing apparatus according to an eighth embodiment of the present disclosure.
[0041] Figure 28 This is a diagram used to describe frequency separation performed by the normal frequency separation unit according to the eighth embodiment of this disclosure.
[0042] Figure 29 This is a diagram used to describe the learning method of the learning machine according to the ninth embodiment.
[0043] Figure 30 This is a diagram illustrating an example configuration of the control unit of the information processing apparatus according to the tenth embodiment of this disclosure.
[0044] Figure 31 This is a block diagram illustrating an example of a schematic configuration of a patient in vivo information acquisition system using a capsule endoscope to which the technology (the technology) according to this disclosure can be applied.
[0045] Figure 32 This is a diagram illustrating an example of a schematic configuration of an endoscopic surgical system to which the technology (the technology) according to this disclosure can be applied.
[0046] Figure 33 It is shown Figure 32 A block diagram illustrating an example of the functional configuration of the camera head and camera control unit (CCU). Detailed Implementation
[0047] In the following, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Note that in this specification and the drawings, components having substantially the same functional configuration are indicated by the same reference numerals, thus omitting repeated descriptions of these components.
[0048] Note that the descriptions will be provided in the following order.
[0049] 1. Background
[0050] 1.1. Surface Shape Calculation Method
[0051] 1.2. Problems with Calculation Methods
[0052] 2. First Implementation Method
[0053] 2.1. Overview of the First Embodiment
[0054] 2.2. Example of system configuration
[0055] 2.3. Example of microscope configuration
[0056] 2.4. Example of Information Processing Device Configuration
[0057] 2.5. Depth Computation Processing
[0058] 3. Second Implementation Method
[0059] 4. Third Implementation Method
[0060] 5. Fourth Implementation Method
[0061] 6. Fifth Implementation Method
[0062] 7. Sixth Implementation Method
[0063] 8. Seventh Implementation Method
[0064] 9. Eighth Implementation Method
[0065] 10. Ninth Implementation Method
[0066] 11. Tenth Implementation Method
[0067] 12. Other implementation methods
[0068] 13. Applicable Examples
[0069] 14. Supplementary Description
[0070] <1. Background>
[0071] <1.1. Surface Shape Calculation Method>
[0072] First, before describing the details of the implementation of this disclosure, the background of the implementation of this disclosure by the inventors will be described.
[0073] As a method for obtaining a 3D image in which the shape (convexity / concavity) of an object is reproduced based on a captured image (RGB image) captured by an imaging device, methods for directly calculating the surface shape (depth to the surface) based on an RGB image using machine learning such as a convolutional neural network (CNN) are known. However, in the case of directly calculating the depth based on an RGB image using a CNN, there are uncertainties and it is difficult to improve the accuracy.
[0074] As another method for calculating the depth to an object's surface based on captured images, there exists a method that uses machine learning, such as CNNs, to calculate the surface normal information of the object based on the captured images and then transforms the normal information into depth using an expression. (Refer to...) Figure 1 Describe such a method. Figure 1 This is a diagram used to describe a calculation method for calculating depth based on a captured image. For example, it is used by the information processing device 200 with a reference... Figure 1 The method described.
[0075] In this method, the information processing device 200 first acquires a captured image obtained by the microscope (step S1). At this time, the information processing device 200 also acquires, for example, information about the microscope. Based on the captured image, the information processing device 200 calculates the normals (normal information) of the target's surface (step S2). For example, the information processing device 200 obtains the normal information as output data by inputting the captured image into a learning machine (model) trained using a CNN or similar method.
[0076] The information processing device 200 calculates the distance (depth information) to the target surface based on the normal information (step S3). When the distance between the microscope sensor and the target is known, which is a piece of information about the microscope, the information processing device 200 can measure the distance to the target surface based on the normal information. Here, when the user performs imaging of the target using the microscope, the microscope head mount is brought into contact with the target to perform imaging. Therefore, the distance between the microscope sensor and the target is a known value corresponding to the length of the head mount.
[0077] Specifically, the information processing device 200 calculates the depth to the target surface based on the normal information by minimizing W in the following expression (1).
[0078] [Mathematical Expression 1]
[0079]
[0080] Note that the parameters of expression (1) are as follows.
[0081] p: x-direction of the calculated normal
[0082] q: The y-direction of the calculated normal.
[0083] Z x The partial derivative of the depth in the x-direction (x-direction of the normal) to be obtained.
[0084] Z y The partial derivative of the depth in the y-direction (the y-direction of the normal) to be obtained.
[0085] Z xx The desired depth is the second partial derivative in the x-direction (the partial derivative of the normal in the x-direction).
[0086] Z yy The desired depth is the second partial derivative in the y-direction (the partial derivative of the normal in the y-direction).
[0087] Z xy The second-order partial derivatives of the depth in the x and y directions to be obtained.
[0088] Note that x and y indicate coordinates on the captured image, and the x-direction is, for example, the horizontal direction of the captured image, and the y-direction is the vertical direction of the captured image.
[0089] When the Fourier transform (frequency transform) of expression (1) is performed, expression (2) is obtained. The depth to the target surface is calculated by minimizing it. Note that Z in expression (2) F It is obtained by Fourier transform of the desired depth, and u and v represent coordinates in frequency space.
[0090] [Mathematical Expression 2]
[0091]
[0092] The following expression (3) is obtained by expanding expression (2), and the depth to the target surface is calculated by performing an inverse Fourier transform on expression (3).
[0093] [Mathematical Expression 3]
[0094]
[0095] The information processing device 200 calculates the distance (depth) to the target surface by using the above transformations of expressions (1) to (3).
[0096] <1.2. Problems with Calculation Methods>
[0097] In the above calculation method, the second and third rows on the right side of expression (1) are the cost terms in expression (1). The second row represents the sum of the absolute values of the normals, and the third row represents the sum of the absolute values of the derivatives of the normals.
[0098] Here, in the above calculation method, the depth is calculated under the assumption that the weights (λ, μ) of the cost term are small—that is, the sum of the absolute values of the normals is small and the sum of the absolute values of the derivatives of the normals is small. The assumption that the sum of the absolute values of the normals is small (hereinafter also referred to as Assumption 1) implies that the surface of the target to which the calculated depth is located is flat. The assumption that the sum of the absolute values of the derivatives of the normals is small (hereinafter also referred to as Assumption 2) implies that the curvature of the surface of the target to which the calculated depth is located is small. That is, in the above calculation method, the depth to the target is calculated under the assumption that the approximate surface shape of the target to which the calculated depth is located is a flat surface shape.
[0099] Here, we will refer to Figure 2 To describe the approximate shape of the target surface. Figure 2 This is a table used to describe the approximate shape of the target's surface. Note that in Figure 2 In the diagram, the actual shape of the target surface is indicated by dashed lines, and the approximate shape of the target surface is indicated by straight lines. Furthermore, the normals to the target surface are indicated by arrows. Figure 2 In this context, the horizontal direction of the target is defined as the x-direction, and the vertical direction is defined as the z-direction (depth).
[0100] The approximate shape of the target surface is the shape of the target surface in the entire captured image, and the approximate shape being a flat surface shape means that there are small variations when viewing the shape of the target surface in the entire captured image. Note that because the local unevenness of the target surface is calculated, it is not included in the approximate shape. When the cost term of the optimization formula shown in the above expression (1) is satisfied, that is, the shape of the target surface in the entire captured image is closer to a flat surface shape, the information processing device 200 can calculate the depth to the target surface with high accuracy.
[0101] For example, Figure 2 The target shown in Table (1) has a shape that includes minute irregularities in a straight surface. Therefore, since the target shown in (1) satisfies both of the above assumptions 1 and 2 (indicated by “○” in the table), the information processing device 200 can calculate the depth of the target with high accuracy.
[0102] Figure 2 The target shown in Table (2) has a shape including minute irregularities in a slightly upwardly curved surface. Therefore, the target shown in (2) slightly satisfies Assumption 1 (indicated by “Δ” in the table) and Assumption 2. Therefore, the information processing device 200 can calculate the depth of the target shown in (2) with high accuracy, although the accuracy is lower than that of (1).
[0103] Figure 2 The target shown in Table (3) has a shape including minute irregularities in a surface that curves sharply upwards on the left. Therefore, the target shown in (3) does not satisfy both assumptions 1 and 2 (indicated by “×” in the table). Therefore, the information processing device 200 cannot accurately calculate the depth of the target shown in (3).
[0104] Figure 2 The target shown in table (4) has a stepped surface. As mentioned above, since the target shown in (4) has the following shape, which includes minute irregularities in the surface with multiple straight surfaces having different heights and orientations, the target shown in (4) does not satisfy both assumptions 1 and 2 above. Therefore, the information processing device 200 cannot accurately calculate the depth of the target shown in (4).
[0105] As mentioned above, in the above calculation method, if the assumption that the surface of the depth calculation target is roughly flat is not met, there is a problem of reduced accuracy in calculating the depth.
[0106] Therefore, in view of this situation, this disclosure creates each embodiment of the present disclosure relating to an information processing apparatus 200 capable of improving the accuracy of shape measurement by improving the accuracy of depth calculation. Therefore, details of each embodiment according to the present disclosure will be described sequentially below.
[0107] <2. First Implementation Method>
[0108] <2.1. Overview of the First Embodiment>
[0109] Figure 3 This is a diagram outlining a first embodiment of the present disclosure. According to the first embodiment, the information processing apparatus 200A (not shown) extracts regions (flat regions) that satisfy the above assumptions 1 and 2 from a captured image, thereby suppressing the degradation of the accuracy of shape information (depth information) related to the shape of the target surface and improving the accuracy of shape measurement.
[0110] First, the information processing device 200A acquires a captured image M11 obtained by imaging the target S using the sensor 150 of the microscope 100. The captured image M11 is an image obtained based on the reflected light from light IA and light IB emitted to the target S from multiple light sources 160A and 160B respectively arranged at different positions.
[0111] Here, microscope 100 will be briefly described. For example... Figure 3 As shown on the left, the microscope 100 includes a sensor 150, a head mount 10 as a cylindrical mechanism mounted between the sensor 150 and the imaging target, and multiple light sources 160A and 160B. Note that the sensor 150 can be interpreted as a lens, camera device, etc. The microscope 100 is an imaging device used by a user who holds the microscope in his / her hand and brings the head mount 10 into contact with the target S to guide the sensor 150 toward the target S.
[0112] Microscope 100 exposes the target S to reflect the light IA and IB emitted simultaneously from light sources 160A and 160B, thereby creating an image of the target S. At this time, for example, if the surface of the target S is not a generally flat surface, occlusions (areas not illuminated by light) appear on the surface of the target S.
[0113] For example, light IA emitted from light source 160A illuminates regions SA and SAB on the surface of target S, but not region SB. On the other hand, light IB emitted from light source 160B illuminates regions SB and SAB on the surface of target S, but not region SA. Region SAB, illuminated by both light IA and light IB from multiple light sources 160A and 160B, is a flat region without obstruction.
[0114] like Figure 3 As shown on the right, the captured image M11 of the microscope 100 is a dark image in which the brightness values of regions SA and SB, which are illuminated only by one of the light sources IA and IB of 160A and 160B, are lower than the brightness value of region SAB, which is illuminated by both light IA and IB.
[0115] Therefore, the information processing apparatus 200A extracts the flat region SAB from the captured image M11 based on the brightness value of the captured image M11. For example, the information processing apparatus 200A extracts the region SAB of the captured image M11 whose brightness value is equal to or greater than a threshold as the flat region.
[0116] The information processing device 200A calculates shape information (depth information) about the surface of the target S based on the flat region SAB of the captured image M11 and information about the sensor 150. For example, the information processing device 200A obtains the normal information of the flat region SAB by inputting the flat region SAB of the captured image M11 into a learning machine trained using a CNN. The information processing device 200A calculates the depth information about the depth to the surface of the target S by performing a transformation on the obtained normal information using the above expressions (1) to (3).
[0117] In this way, the information processing device 200A can improve the accuracy of shape measurement by extracting the regions that satisfy the above assumptions 1 and 2 of expression (1) from the captured image M11.
[0118] The details of the information processing system 1, which includes the aforementioned information processing device 200A, will be described below.
[0119] <2.2. Example of system configuration>
[0120] Figure 4 This is a block diagram illustrating an example configuration of an information processing system 1 according to a first embodiment of the present disclosure. (See diagram for example.) Figure 4 As shown, the information processing system 1 includes a microscope 100 and an information processing device 200A.
[0121] Microscope 100 is an imaging device used by a user who holds the microscope in his / her hand and guides sensor 150 toward the imaging target.
[0122] The information processing device 200A calculates shape information about the surface of the imaging target S based on the captured image M11 captured by the microscope 100. (See below for further details.) Figure 6 To describe the details of the information processing device 200A.
[0123] The microscope 100 and the information processing device 200A are connected via, for example, a cable. Alternatively, the microscope 100 and the information processing device 200A can be directly connected via wireless communication such as Bluetooth (registered trademark) or Near Field Communication (NFC). The microscope 100 and the information processing device 200A can be connected wired or wirelessly via, for example, a network (not shown). Alternatively, the microscope 100 and the information processing device 200A can send and receive captured images via an externally mounted storage medium such as a hard disk, magnetic disk, magneto-optical disk, optical disk, USB memory, or memory card. Furthermore, the microscope 100 and the information processing device 200A can be configured as a single unit. Specifically, for example, the information processing device 200A can be arranged inside the main body of the microscope 100.
[0124] <2.3. Example of microscope configuration>
[0125] Figure 5 This is a diagram illustrating an example configuration of a microscope 100 according to a first embodiment of the present disclosure. The microscope 100 includes a sensor 150 and a head mount 10 as a cylindrical mechanism mounted between the sensor 150 and the imaging target. Furthermore, the microscope 100 includes point light sources 160A and 160B (not shown) arranged at different positions.
[0126] The head mount 10 is a mechanism mounted on the distal end of the microscope 100. The head mount 10 is also referred to as a tip head or lens tube, for example. For example, a mirror may be disposed inside the head mount 10, and light emitted from point light sources 160A and 160B may be totally reflected by the side surface of the head mount 10.
[0127] The user brings the head mount 10 into contact with the target S to image the target S. As a result, the distance between the sensor 150 and the target S is fixed, and the focus (focal length) can be prevented from shifting during imaging.
[0128] The microscope 100 includes point light sources 160A and 160B disposed inside the head mounting portion 10. Therefore, the microscope 100 exposes the target S by reflecting light IA and IB emitted from the point light sources 160A and 160B to the target S, thereby imaging the target S. Note that the term "point light source" in this specification ideally refers to a point-based light source; however, point-based light sources do not actually exist, and therefore, point light sources include light sources with extremely small dimensions (within a few millimeters or smaller).
[0129] Note that this specification describes an example using a point light source, but the light source is not limited to a point light source. Furthermore, the number of light sources is not limited to two. Multiple light sources can be set, and three or more light sources can be set, as long as the light sources are arranged in different locations.
[0130] <2.4. Example of Information Processing Device Configuration>
[0131] Figure 6 This is a block diagram illustrating an example configuration of an information processing apparatus 200A according to a first embodiment of the present disclosure. The information processing apparatus 200A includes a control unit 220 and a storage unit 230.
[0132] (Control unit)
[0133] The control unit 220 controls the operation of the information processing device 200A. The control unit 220 includes an acquisition unit 221, a region acquisition unit 225, a normal calculation unit 222, a depth calculation unit 223, and a display control unit 224. Each functional unit, including the acquisition unit 221, region acquisition unit 225, normal calculation unit 222, depth calculation unit 223, and display control unit 224, is implemented by the control unit 220, for example, by executing a program stored internally in the control unit 220 using a random access memory (RAM) or similar memory as its working area. Note that the internal structure of the control unit 220 is not limited to... Figure 6 The configuration shown is for control unit 220, and control unit 220 may have other configurations, as long as it performs the information processing described later. Furthermore, the connection relationships between the various processing units included in control unit 220 are not limited to... Figure 6 The connection relationship shown can be any other connection relationship.
[0134] (Acquisition Unit)
[0135] The acquisition unit 221 acquires the captured image M11 captured by the microscope 100 and information about the microscope 100. For example, the information about the microscope 100 includes information about the structure of the microscope 100, such as the length d and focal length f of the head mount 10.
[0136] For example, the acquisition unit 221 can control the point light sources 160A and 160B of the microscope 100 and the sensor 150. In this case, the acquisition unit 221 controls the point light sources 160A and 160B so that light IA and light IB are emitted simultaneously from the point light sources 160A and 160B. Furthermore, the acquisition unit 221 controls the sensor 150 so that the sensor 150 images the target S when light IA and light IB are emitted simultaneously from the point light sources 160A and 160B. In this way, the information processing device 200A can control the microscope 100.
[0137] Alternatively, the acquisition unit 221 may acquire information about the imaging conditions other than the captured image M11. This information about the imaging conditions could include, for example, information indicating that the captured image M11 was captured when point light sources 160A and 160B simultaneously emitted light IA and light IB. In this case, the region acquisition unit 225, described later, could extract a flat region SAB from the captured image M11, for example, based on the imaging conditions.
[0138] (Region Acquisition Unit)
[0139] The region acquisition unit 225 extracts the flat region SAB from the captured image M11. For example, the region acquisition unit 225 compares the brightness value L(x,y) of each pixel in the captured image M11 with a threshold th. Figure 7 As shown, the region acquisition unit 225 sets the region containing pixels whose brightness value L(x,y) is equal to or greater than the threshold th as the processing region, and sets the region containing pixels whose brightness value L(x,y) is less than the threshold th as the exclusion region. Note that Figure 7 This is a diagram showing an example of the brightness values of the captured image M11, and Figure 7 The relationship between the brightness value L(x, β) and the pixel (x, β) is shown, where the value of the y-axis is a predetermined value "β".
[0140] Note that when the captured image M11 is an RGB image, the luminance value L(x,y) is calculated using the expression L(x,y) = (R(x,y) + 2G(x,y) + B(x,y)) / 4. R(x,y) is the red (R) component of the pixel value of pixel (x,y), G(x,y) is the green (G) component of the pixel value of pixel (x,y), and B(x,y) is the blue (B) component of the pixel value of pixel (x,y).
[0141] The region acquisition unit 225 compares the brightness value L(x,y) with the threshold th for all pixels to obtain the region determined as the processing region as the extraction region (flat region SAB) to be extracted from the captured image M11.
[0142] For example, the region acquisition unit 225 generates a pixel by setting it to white when the pixel's brightness value L(x,y) is equal to or greater than a threshold th, and setting the pixel to black when the brightness value L(x,y) is less than the threshold th. Figure 8 The mask image shown is M12. Note that... Figure 8 This is a diagram illustrating an example of a mask image M12 generated by the region acquisition unit 225.
[0143] like Figure 9As shown, the region acquisition unit 225 compares the generated mask image M12 with the captured image M11, and extracts the white pixels in the same coordinates from the captured image M11, thereby acquiring the flat region SAB. Note that Figure 9 This is a diagram used to describe an example of obtaining a flat region SAB by the region acquisition unit 225.
[0144] Alternatively, the region acquisition unit 255 can generate a mask image M12 by setting the brightness value of a pixel to "1" when the brightness value L(x,y) of the pixel is equal to or greater than a threshold th, and setting the brightness value of the pixel to "0" when the brightness value L is less than the threshold th. In this case, the region acquisition unit 255 obtains the flat region SAB by multiplying the captured image M11 and the mask image M12.
[0145] The region acquisition unit 255 outputs the acquired flat region SAB to the normal calculation unit 222. Note that the mask image M12 generated by the region acquisition unit 255 can be output to the normal calculation unit 222, and the normal calculation unit 222 can acquire the flat region SAB from the captured image M11 using the mask image M12.
[0146] (Normal calculation unit)
[0147] The normal calculation unit 222 calculates the normal information of the surface of the target S as a normal image based on the flat region SAB of the captured image M11. For example, the normal calculation unit 222 generates the normal image using a CNN. Alternatively, the normal calculation unit 222 generates the normal image using a learning machine pre-trained with a CNN. The learning machine 300 (see...) Figure 11 The number of input channels for the learning machine 300 is three corresponding to the RGB values of the input image (in this case, the flat SAB region) or one corresponding to the grayscale values of the input image. Furthermore, the number of output channels for the learning machine 300 is three corresponding to the RGB values of the output image (in this case, the normal image). The resolution of the output image is assumed to be equal to the resolution of the input image.
[0148] Here, we will refer to Figure 10 Describe the relationship between normal information and the RGB values of the normal image. Figure 10 This is an example diagram used to describe normal information. Assume... Figure 10 The normal N shown is the normal vector in a predetermined pixel of the captured image M11 obtained by imaging the target S. In this case, the relationship between the normal N(θ, φ) in the polar coordinate system and the normal N(x, y, z) in the orthogonal coordinate system is established by the following expression (4), where the zenith angle of the normal N is θ and the azimuth angle is φ. Note that the normal N is assumed to be a unit vector.
[0149] [Mathematical Expression 4]
[0150]
[0151] For example, the normal information is the normal N(x, y, z) in the orthogonal coordinate system described above, and the normal information is calculated for each pixel of the flat region SAB of the captured image M11. Furthermore, the normal image is an image obtained by replacing the normal information of each pixel with RGB values. That is, the normal image is an image obtained by replacing x with R (red), y with G (green), and z with B (blue) of the normal N.
[0152] Description Return to Figure 6 The normal calculation unit 222 generates a normal image as the output image by inputting the flat region SAB as the input image to the learning machine 300. It is assumed that the learning machine 300 is pre-generated using learned ground truth values as weights in the normal learning CNN.
[0153] Here, we will refer to Figure 11 Describe the method for generating the learning machine 300. Figure 11 This is a diagram used to describe the learning machine 300 according to an embodiment of the present disclosure. Here, for the sake of simplicity, it is assumed that the control unit 220 of the information processing device 200A generates the learning machine 300, but the learning machine 300 may be generated by, for example, another information processing device (not shown).
[0154] exist Figure 11 In the example shown, the information processing device 200A calculates the similarity between the output image M15 and the normal image M16, which serves as the ground truth for learning, when the input image M14 is input to the learning machine 300. Here, the information processing device 200A calculates the least squares error (L2) between the output image M15 and the normal image M16, and performs learning processing using L2 as the loss. The information processing device 200A updates the weights of the learning machine 300, for example, by backpropagating the calculated loss. As a result, the information processing device 200A generates a learning machine 300 that outputs a normal image when the input captured image is received.
[0155] Note that here, as an example of machine learning, the learning machine 300 is generated using a CNN, but the invention is not limited to this. As a machine learning method, various methods such as recurrent neural networks (RNNs) can be used to generate the learning machine 300 in addition to CNNs. Furthermore, in the example above, the weights of the learning machine 300 are updated by backpropagating the calculated loss, but the invention is not limited to this. In addition to backpropagation, any learning method such as stochastic gradient descent can be used to update the weights of the learning machine 300. In the example above, the loss is the least squares error, but it is not limited to this. The loss can be the minimum average error.
[0156] Description Return to Figure 6 The normal calculation unit 222 generates a normal image using the generated learning machine 300. The normal calculation unit 222 outputs the calculated normal image to the depth calculation unit 223.
[0157] (Depth Calculation Unit)
[0158] The depth calculation unit 223 takes each pixel value of the normal image as input and performs a transformation using the above expressions (1) to (3) to calculate the distance (depth) from the sensor 150 to the target S. Specifically, the depth calculation unit 223 calculates Z' by performing an inverse Fourier transform on the above expression (3) as follows, and calculates the depth Z to be obtained as Z = Z' × P.
[0159] [Mathematical Expression 5]
[0160]
[0161] Note that the parameters of the above expression (3), etc., are as follows.
[0162] P(u,v): Fourier transform of the normal in the x-direction
[0163] Q(u,v): Fourier transform of the normal in the y-direction
[0164] u, v: Coordinates of each pixel in frequency space
[0165] P: Length of each pixel (one pixel) in the captured image xy [μm]
[0166] Z F The depth of the Fourier transform to be obtained
[0167] Z': The depth per pixel to obtain.
[0168] Z: Depth to be obtained [μm]
[0169] P (the length of each pixel (one pixel) of the captured image xy [μm]) is a parameter whose value is determined based on the configuration of the sensor 150 of the imaging device (microscope 100).
[0170] For example, such as Figure 12 As shown, assume the focal length of sensor 150 is f [mm] and the length of the housing (head mounting part 10) is d [mm]. Furthermore, based on P:d = p:f, the actual length P of each pixel in the captured image is P = (d × p) / f [μm], where the pixel pitch of sensor 150 is p [μm]. Note that... Figure 12 It is a graph used to describe the length of each pixel in a captured image.
[0171] The depth calculation unit 223 calculates the depth of the target S for each pixel of the captured image Mll, for example, based on the normal image and information about the sensor 150. The information about the sensor 150 includes, for example, information about the structure of the sensor 150, and specifically, information about the focal length of the sensor 150 and the distance between the sensor 150 and the target S. In the case where the imaging device is a microscope 100, the distance between the sensor 150 and the target S is the length d of the head mount 10.
[0172] (Display control unit)
[0173] The display control unit 224 displays various images on a display unit (not shown), such as a liquid crystal display. Figure 13 This is a diagram illustrating an example of an image M17 displayed on a display unit by the display control unit 224. (See diagram for example.) Figure 13 As shown, the display control unit 224 causes the display unit to display, for example, an image M17 including a captured image M17a, a normal image M17b, a depth image M17c, an image M17d showing a depth curve, and a 3D image M17e of the target S.
[0174] The captured image M17a is, for example, an image obtained by cropping the flat region SAB from the captured image M11 captured by the microscope 100. The normal image M17b is an image in which the normal information of the captured image M17a is displayed in RGB. The depth image M17c is an image indicating the depth in each pixel of the captured image M17a, and shows, for example, that the lighter the color, the greater the depth (the farther the distance from the sensor 150 to the target S). The image M17d, showing a depth curve, is an image in which the depth along the straight line shown in the normal image M17b, depth image M17c, and 3D image M17e of the target S is plotted. The 3D image M17e of the target S is an image that displays the target S in three dimensions based on the depth image M17c.
[0175] In this way, the display control unit 224 can display depth curves or three-dimensional images on the display unit in addition to the captured image, the generated normal image, and the depth image.
[0176] (Storage unit)
[0177] The storage unit 230 is implemented by a read-only memory (ROM) and a random access memory (RAM). The ROM stores programs, operating parameters, etc., for the processing to be performed by the control unit 220, and the RAM temporarily stores parameters that are changed as appropriate.
[0178] <2.5. Depth Computation Processing>
[0179] Figure 14 This is a flowchart illustrating an example of depth computing processing according to a first embodiment of the present disclosure. Figure 14 The depth calculation processing shown is implemented by the control unit 220 of the information processing device 200A executing a program. Figure 14 The depth calculation processing shown is performed after the microscope 100 captures the captured image M11. Alternatively, it can be performed in response to instructions from the user. Figure 14 The depth calculation process is shown.
[0180] like Figure 14 As shown, the control unit 220 acquires the captured image M11 from the microscope 100, for example (step S101). Alternatively, the control unit 220 may acquire the captured image M11 from another device, such as a network (not shown). Examples of other devices include another information processing device and a cloud server.
[0181] The control unit 220 acquires information about the sensor 150 (step S102). The control unit 220 can acquire information about the sensor 150 from the microscope 100. Alternatively, if the sensor information about the sensor 150 is stored in the storage unit 230, the control unit 220 can acquire the sensor information from the storage unit 230. Furthermore, the control unit 220 can acquire sensor information, for example, through input from a user.
[0182] The control unit 220 obtains the flat region SAB from the acquired captured image M11 (step S103). Specifically, the control unit 220 extracts the region where the brightness value of each pixel in the captured image M11 is equal to or greater than the threshold th as the flat region SAB.
[0183] The control unit 220 calculates the normal information of the acquired flat region SAB (step S104). Specifically, the control unit 220 inputs the flat region SAB as an input image to the learning machine 300 to generate a normal image including normal information as an output image.
[0184] The control unit 220 calculates depth information based on the normal information and information about the sensor 150 (step S105). Specifically, the control unit 220 calculates depth information by performing a transformation on the normal information based on the above expression (3), etc.
[0185] As described above, the information processing apparatus 200A according to the first embodiment includes a control unit 220. The control unit 220 acquires a captured image M11 obtained by the sensor 150 capturing a target S. The captured image M11 is an image obtained from reflected light IA and IB emitted to the target S from a plurality of point light sources 160A and 160B respectively arranged at different positions. The control unit 220 extracts a flat region SAB from the captured image M11 based on the brightness value of the captured image M11. The control unit 220 calculates shape information (depth information) regarding the shape of the surface of the target S based on the flat region SAB of the captured image M11 and information about the sensor 150.
[0186] Specifically, the information processing apparatus 200A according to the first embodiment acquires a flat region SAB from a captured image M11, which is obtained from reflected light IA and IB emitted simultaneously from multiple point light sources 160A and 160B to the target S.
[0187] As described above, when the information processing device 200A acquires the flat region SAB, assumptions 1 and 2 can be satisfied in the transformation performed by the depth calculation unit 223, and the decrease in depth calculation accuracy can be suppressed. As a result, the information processing device 200A can improve the accuracy of shape measurement.
[0188] <3. Second Implementation Method>
[0189] In the first embodiment, it has been described that the information processing apparatus 200A calculates depth based on a captured image M11 obtained by the microscope 100 when multiple point light sources 160A and 160B simultaneously emit light IA and light IB. In addition to the above example, the information processing apparatus 200A can also calculate depth based on a captured image obtained by the microscope 100 when light is emitted from a single point light source. Therefore, in the second embodiment, an example will be described whereby the information processing apparatus 200A calculates depth based on a captured image obtained for each of the point light sources 160A and 160B, obtained from reflected light emitted from one of the multiple point light sources 160A and 160B towards the target S. Note that, except for the operation of the acquisition unit 221 and the area acquisition unit 225, the information processing apparatus 200A according to the second embodiment has the same configuration and operates in the same manner as the information processing apparatus 200A according to the first embodiment, and therefore, the information processing apparatus 200A according to the second embodiment is indicated by the same reference numerals, and a portion of the description of the information processing apparatus 200A according to the second embodiment will be omitted.
[0190] Figure 15 This is a diagram used to describe the captured images M21 and M22 acquired by the acquisition unit 221 according to the second embodiment of the present disclosure.
[0191] like Figure 15 As shown, the microscope 100 first emits light IA from the point light source 160A to image the target S. At this time, for example, if the surface shape of the target S is not a flat surface, occlusions (areas not illuminated by light) appear on the surface of the target S.
[0192] For example, light IA emitted from light source 160A illuminates region SA1 on the surface of target S, but not region SA2. In this case, for example, as... Figure 15 As shown, the acquisition unit 221 acquires a captured image M21 in which region SA2 is darker than region SA1.
[0193] Subsequently, the microscope 100 emits light IB from the point light source 160B to image the target S. At this time, similar to the case where light IA is emitted from the point light source 160A, occlusion (areas not illuminated by light) appears on the surface of the target S.
[0194] For example, light IB emitted from light source 160B illuminates region SB1 on the surface of target S, but not region SB2. In this case, for example, as... Figure 15 As shown, the acquisition unit 221 acquires the captured image M22 in which region SB2 is darker than region SB1.
[0195] In this way, the acquisition unit 221 acquires the captured images M21 and M22 by sequentially turning on two point light sources 160A and 160B through the microscope 100.
[0196] The region acquisition unit 225 acquires the flat region SAB from the captured images M21 and M22 acquired by the acquisition unit 221. As described above, the flat region SAB is illuminated by the light IA and light IB from both point light sources 160A and 160B. In other words, the occluded area that is not illuminated by at least one of the light IA or light IB is a non-flat region. As described above, in the captured images M21 and M22, the light IA and light IB from point light sources 160A and 160B illuminate the flat region SAB, such that the brightness values of the pixels in the flat region SAB in the captured images M21 and M22 are substantially the same.
[0197] Therefore, the region acquisition unit 225 acquires the flat region SAB based on the brightness values of the captured images M21 and M22. The region acquisition unit 225 extracts the region in each corresponding pixel of the captured images M21 and M22 where the brightness value does not change or the change in brightness value is small (less than the threshold th2) as the flat region SAB.
[0198] Reference Figure 16 Describes a specific method for obtaining the flat region SAB by the region acquisition unit 225. Figure 16 This is a diagram used to describe the acquisition of a flat region by the region acquisition unit 225 according to the second embodiment of the present disclosure.
[0199] The region acquisition unit 225 calculates, for example, the difference in brightness values of corresponding pixels (x, y) between captured images M21 and M22, D(x, y) = L1(x, y) - L2(x, y). Note that L1(x, y) is the brightness value of pixel (x, y) in captured image M21, and L2(x, y) is the brightness value of pixel (x, y) in captured image M22. Figure 16 This is a graph showing the relationship between the difference D and x when the y-value of a pixel is fixed to a predetermined value. Note that, as... Figure 15 As shown, the horizontal direction of captured images M21 and M22 is the x-direction, and the vertical direction of captured images M21 and M22 is the y-direction.
[0200] For example, in Figure 16 In region SB2 shown, the brightness value of captured image M21 is high, and the brightness value of captured image M22 is low (see [reference]). Figure 15 Therefore, the value of difference D in region SB2 increases in the positive direction. On the other hand, in Figure 16 In region SA2 shown, the brightness value of captured image M21 is low, while the brightness value of captured image M22 is high (see [reference]). Figure 15Therefore, the difference D in region SA2 increases in the negative direction. Furthermore, in regions outside SB2 and SA2, since captured images M21 and M22 both have substantially the same high brightness value, the difference D becomes close to 0.
[0201] Therefore, the region acquisition unit 225 sets the region SB2, whose difference D is equal to or greater than the threshold th2, as an exclusion region for which depth calculation is not performed. Furthermore, the region acquisition unit 225 sets the region SA2, whose difference D is equal to or less than the threshold -th2, as an exclusion region for which depth calculation is not performed. The region acquisition unit 225 sets the region whose difference D is within the range of the threshold ± th2 as the processing region for which depth calculation is to be performed. The region acquisition unit 225 acquires the regions determined as processing regions as the extraction regions (flat regions SAB) to be extracted from the captured images M21 and M22.
[0202] More specifically, the region acquisition unit 225 performs threshold determination on the absolute value abs(D) of the difference D, setting regions containing pixels where the absolute value abs(D) of the difference D is equal to or greater than the threshold th2 as black regions, and regions containing pixels where the absolute value abs(D) of the difference D is less than the threshold th2 as white regions, thereby generating, for example... Figure 8 The mask image M12 is shown. For example, the region acquisition unit 225 compares the generated mask image M12 with the captured image M21, and extracts pixels from the captured image M21 that are white in the mask image M12 at the same coordinates, thereby obtaining the flat region SAB. Alternatively, the region acquisition unit 225 may also compare the mask image M12 with the captured image M22 to obtain the flat region SAB. Note that the processing after generating the mask image M12 is the same as the processing in the first embodiment.
[0203] In the second embodiment, one of the point light sources 160A and 160B emits light to perform imaging. In this case, compared to the simultaneous light emission in the first embodiment described above, the amount of light illuminating the target S can be suppressed to a low level, thereby suppressing the overall brightness of the captured images M21 and M22 while maintaining a high contrast. Therefore, the difference between the brightness values of the illuminated areas and the unilluminated occluded areas in the captured images M21 and M22 can be increased. As a result, the region acquisition unit 225 can more easily acquire the flat region SAB from the captured images M21 and M22.
[0204] Meanwhile, in the information processing apparatus 200A according to the first embodiment, since only one capture image M11 needs to be acquired, the image acquisition time can be reduced to half the time of the second embodiment. Furthermore, when the target S or sensor 150 moves while capture images M21 and M22 are captured sequentially, the region acquisition unit 225 needs to perform processes such as alignment of capture images M21 and M22 before acquiring the flat region SAB. Therefore, in the case of performing depth calculations in real time, it is desirable to use the information processing apparatus 200A according to the first embodiment, which has a short image acquisition time and does not require processes such as alignment, to perform depth calculations.
[0205] As described above, the information processing apparatus 200A according to the second embodiment acquires captured images M21 and M22 for each of the point light sources 160A and 160B, based on the reflected light of light IA and light IB emitted from one of the multiple point light sources 160A and 160B to the target S. The information processing apparatus 200A extracts a region where the absolute value of the difference D between the multiple captured images M21 and M22 acquired for each of the point light sources 160A and 160B is less than a predetermined threshold th2 as a flat region SAB.
[0206] As a result, the information processing device 200A can increase the contrast of the captured images M21 and M22, improve the accuracy of extracting the flat region SAB, and improve the accuracy of shape measurement.
[0207] Note that the case where microscope 100 includes two point light sources 160A and 160B has been described here as an example, but the number of point light sources is not limited to two. For example, microscope 100 may include three or more point light sources.
[0208] In this case, the microscope 100 sequentially turns on three or more point light sources to image the target S. The information processing device 200A acquires multiple captured images captured by the microscope 100. The information processing device 200A selects regions with small changes in brightness values among the acquired multiple captured images (in other words, regions where the absolute value of the difference in brightness values of corresponding pixels between multiple captured images is less than a predetermined threshold th2) as flat regions SAB.
[0209] Note that even when the microscope 100 has three or more point light sources, it is not necessary to sequentially activate all point light sources to image the target S. For example, when the microscope 100 has three point light sources, imaging can be performed by sequentially activating two of the three point light sources. In this case, for example, the microscope 100 can select and activate the farthest point light source among the multiple point light sources.
[0210] <4. Third Implementation Method>
[0211] In the first and second embodiments, the case of obtaining the flat region SAB from the captured image acquired by the acquisition unit 221 has been described. Besides the above examples, the information processing apparatus 200A can obtain the flat region SAB after performing smoothing processing on the captured image. Therefore, in the third embodiment, an example of the information processing apparatus 200A performing smoothing processing before obtaining the flat region SAB from the captured image will be described. Note that, except for the operation of the region acquisition unit 225, the information processing apparatus 200A according to the third embodiment has the same configuration and operates in the same manner as the information processing apparatus 200A according to the first embodiment, and therefore, the information processing apparatus 200A according to the third embodiment is indicated by the same reference numerals, and a portion of the description of the information processing apparatus 200A according to the third embodiment will be omitted.
[0212] Figure 17 This is a diagram used to describe the smoothing process performed by the region acquisition unit 225 according to the third embodiment of this disclosure.
[0213] like Figure 17 As shown, the captured image M31 acquired by the acquisition unit 221 includes the local bumps and concavities of the target S and noise. Note that the local bumps and concavities of the target S are the target for depth information calculation by the depth calculation unit 223.
[0214] In the region acquisition unit 225, an attempt is made to obtain data in pixels. Figure 17 When the threshold is determined in the captured image M31 shown, the accuracy of acquiring the flat region SAB is degraded due to the aforementioned unevenness and noise.
[0215] Therefore, the region acquisition unit 225 performs smoothing processing on the captured image M31 before performing threshold determination to obtain the flat region SAB, and generates, as shown in the image. Figure 17 The smoothed image M32 is shown below. Specifically, the region acquisition unit 225 generates the smoothed image M32 by applying a low-pass filter to the captured image M31. The region acquisition unit 225 performs threshold determination on the smoothed image M32 to generate... Figure 8 The mask image M12 is shown.
[0216] For example, the region acquisition unit 225 compares the generated mask image M12 with the captured image M31, and extracts the pixels in the captured image M31 that are white in the mask image M12 at the same coordinates, thereby obtaining the flat region SAB in the captured image M31.
[0217] As described above, the region acquisition unit 225 generates a mask image M12 based on the smoothed image M32, enabling the generation of the mask image M12 while reducing the influence of minor bumps and noise in the captured image M31. Furthermore, by using the mask image M12 to acquire the flat region SAB of the captured image M31, the region acquisition unit 225 can extract the flat region SAB, including local bumps, and accurately calculate the depth of the target S.
[0218] Note that, as described in the second embodiment, when the acquisition unit 221 acquires multiple captured images, the region acquisition unit 225 performs smoothing processing on all acquired captured images. As a result, the information processing apparatus 200A can suppress the decrease in accuracy when acquiring the flat region SAB and can improve the accuracy of shape measurement.
[0219] <5. Fourth Implementation Method>
[0220] In the first and second embodiments, it has been described that the region acquisition unit 225 acquires the flat region SAB by performing a threshold determination on the brightness value of the pixels. In addition to the above examples, the information processing apparatus 200A can acquire a flat region by dividing the captured image into multiple blocks. Therefore, in the fourth embodiment, an example of the information processing apparatus 200A dividing the captured image into multiple blocks will be described. Note that, except for the operation of the region acquisition unit 225, the information processing apparatus 200A according to the fourth embodiment has the same configuration and operates in the same manner as the information processing apparatus 200A according to the first embodiment, and therefore, the information processing apparatus 200A according to the fourth embodiment is indicated by the same reference numerals, and a portion of the description of the information processing apparatus 200A according to the fourth embodiment will be omitted.
[0221] Figure 18 This is a diagram used to describe the flat region acquired by the region acquisition unit 225 according to the fourth embodiment of the present disclosure.
[0222] like Figure 18 As shown, the region acquisition unit 225, for example, divides the captured image M41 acquired by the acquisition unit 221 into multiple blocks, and acquires one block M42 as a flat region. Note that the region acquisition unit 225 determines the size of the block based on information about the microscope 100, and divides the captured image M41 into blocks of the determined size.
[0223] like Figure 19 As shown, even when the shape of the target S varies greatly throughout the captured image M41, the local region SC can be considered substantially flat. Note that... Figure 19This is a diagram used to describe the flat region acquired by the region acquisition unit 225 according to the fourth embodiment of the present disclosure.
[0224] For example, when the target S is human skin, and the area of the target S is empirically 5mm × 5mm or smaller, the region can be a flat surface that satisfies assumptions 1 and 2 above. Therefore, for example, the region acquisition unit 225 divides the captured image M41 into blocks of 5mm × 5mm or smaller. As described above, the actual length P [μm] of each pixel of the captured image is calculated based on the focal length f [mm] of the sensor 150, the length d [mm] of the housing (head mounting part 10), and the pixel pitch p [μm] of the sensor 150.
[0225] For example, as described above, when the captured image M41 is divided into blocks of 5mm × 5mm or smaller, the width w [pixel] of the block is w = 5000 / P = 5000f / (d × p). Specifically, when the focal length f of the sensor 150 is 16 [mm], the length d of the head mounting portion 10 is 100 [mm], and the pixel pitch p of the sensor 150 is 5 [μm], the length P of each pixel is P = 31.25 [μm]. Therefore, when the captured image M41 is divided into blocks of 5mm × 5mm or smaller, the block size is 160 × 160 [pixel].
[0226] Note that normal calculation unit 222 and depth calculation unit 223 calculate normal and depth information for all partitioned blocks. Except for the number of blocks (flat regions) to be calculated, the depth calculation process is similar to... Figure 13 The calculation process shown is the same.
[0227] As described above, in the fourth embodiment, the information processing apparatus 200A divides the captured image M41 into multiple blocks and calculates the depth. As a result, the information processing apparatus 200A can obtain a flat region that satisfies assumptions 1 and 2 in the above expressions (1) to (3), and can accurately calculate the depth. As described above, the information processing apparatus 200A according to the fourth embodiment can improve the accuracy of shape measurement.
[0228] Here, the size of the block is 5mm × 5mm, but it is not limited to this. For example, the imaging target of the microscope 100 may not be human skin. Since the size of the block, which can be considered flat, varies depending on the imaging target, the size of the block can differ depending on the imaging target.
[0229] For example, when the microscope 100 images multiple types of targets, a table relating the types of imagesd targets to appropriate block sizes is stored in the storage unit 230. In this case, the information processing device 200A selects the block size based on the type of imaged target. Alternatively, the user can specify the block size. As mentioned above, the block size is not limited to the above example, and various sizes can be selected.
[0230] <6. Fifth Implementation Method>
[0231] In the fourth embodiment, the case where the region acquisition unit 225 acquires a flat region by dividing the captured image into blocks has been described. Besides the above example, the information processing apparatus 200A can acquire a flat region based on the contrast value of the divided blocks. Therefore, in the fifth embodiment, an example of the information processing apparatus 200A acquiring a flat region based on the contrast value of the divided blocks will be described. Note that, except for the operation of the region acquisition unit 225, the information processing apparatus 200A according to the fifth embodiment has the same configuration and operates in the same manner as the information processing apparatus 200A according to the fourth embodiment, and therefore, the information processing apparatus 200A according to the fifth embodiment is denoted by the same reference numerals, and a portion of the description of the information processing apparatus 200A according to the fifth embodiment is omitted.
[0232] Figure 20 This is a diagram used to describe the depth of field of sensor 150. Sensor 150 (image-taking device) has a depth of field (the range of depth at which a subject can be imaged without blur) corresponding to the aperture of the lens. In the optical system of the image-taking device, there exists a focal length f that allows for sharp imaging of the subject at the highest resolution through focusing. For example, Figure 20 The plane M52 shown is the focusing plane, that is, the plane that is in focus and can image the subject clearly at the highest resolution. The range from plane M53 to plane M51 is the depth range (depth of field) that can image the subject without blur.
[0233] When imaging is performed by bringing the head mount 10 into contact with the target S, as in microscope 100, the distance to the target S becomes shorter, and therefore the amount of light incident on the sensor 150 decreases. Therefore, in order to maximize the amount of light incident on the sensor 150, imaging is performed in microscope 100 with the aperture value F reduced—that is, the aperture is opened. Therefore, as... Figure 21 As shown, the depth of field P1 decreases. Note that... Figure 21 This is a diagram used to describe the acquisition of a flat region by the region acquisition unit 225 according to the fifth embodiment of the present disclosure.
[0234] like Figure 21As shown, because the depth of field P1 of the microscope 100 is small, it is possible to image the target S included in the depth of field P1—that is, the surface SD1 of the target S, which is approximately the same distance from the sensor 150 in terms of approximate shape—without blurring. On the other hand, it is possible to image the target S not included in the depth of field P1—that is, the surfaces SD2 and SD3, which are at different distances from the sensor 150 than the surface SD1—without blurring.
[0235] Therefore, the region acquisition unit 225 acquires the imaging region included in the depth of field P1 as a flat region. The region acquisition unit 225 acquires the flat region by acquiring the in-focus region and the non-blurred region from the captured image. (Refer to...) Figure 22 This will be described. Figure 22 This is a diagram used to describe the acquisition of a flat region by the region acquisition unit 225 according to the fifth embodiment of the present disclosure.
[0236] The region acquisition unit 225 divides the captured image M54 acquired by the acquisition unit 221 into multiple blocks. Note that the size of the blocks to be divided here may be the same as or different from the size of the blocks in the fourth embodiment described above.
[0237] Subsequently, the region acquisition unit 225 calculates the contrast value for each segmented block. The contrast value is obtained through the expression "contrast = (Lmax - Lmin) / (Lmax + Lmin)". Note that Lmax is the maximum value of the luminance value L in the block, and Lmax = max(L(x,y)). Furthermore, Lmin is the minimum value of the luminance value L in the block, and Lmin = min(L(x,y)).
[0238] A focused and unblurred imaging region has a high contrast value `contrast`, while an unfocused and blurred imaging region has a low contrast value `contrast`. Therefore, the region acquisition unit 225 acquires flat regions based on the calculated contrast value `contrast`. Specifically, the region acquisition unit 225 compares the calculated contrast value `contrast` with a threshold `th3` for each block. Figure 22 As shown, the region acquisition unit 225 extracts blocks with a contrast value equal to or greater than the threshold th3 as flat regions for which depth calculation is to be performed.
[0239] Normal calculation unit 222 and depth calculation unit 223 calculate the normal information and depth information of the flat region acquired by region acquisition unit 225. The calculation method is the same as that in the first embodiment. Note that normal calculation unit 222 and depth calculation unit 223 can calculate normal information and depth information on a block-by-block basis. Alternatively, normal calculation unit 222 and depth calculation unit 223 can jointly calculate normal information and depth information for the region extracted as a flat region by region acquisition unit 225—that is, multiple blocks with a contrast value contrast equal to or greater than threshold th3.
[0240] As described above, in the fifth embodiment, the information processing apparatus 200A divides the captured image M54 into multiple blocks. The information processing apparatus 200A obtains flat regions based on the contrast value contrast calculated for each block.
[0241] As described above, the information processing apparatus 200A obtains the flat region based on the contrast value of the captured image M54, and then calculates the depth in the flat region that satisfies assumptions 1 and 2 in the above expressions (1) to (3). As a result, the information processing apparatus 200A according to the fifth embodiment can improve the accuracy of shape measurement.
[0242] <7. Sixth Implementation Method>
[0243] In the fifth embodiment, it has been described that the region acquisition unit 225 divides a captured image into multiple blocks and acquires a flat region based on the contrast value of each block. In addition to the above example, the information processing apparatus 200A can divide each of multiple captured images with different subject depths into multiple blocks and acquire a flat region based on the contrast value of each block. Therefore, in the sixth embodiment, an example of the information processing apparatus 200A acquiring a flat region based on the contrast value of each block in each of the multiple captured images will be described. Note that, except for the operation of the acquisition unit 221 and the region acquisition unit 225, the information processing apparatus 200A according to the sixth embodiment has the same configuration and operates in the same manner as the information processing apparatus 200A according to the fifth embodiment, and therefore, the information processing apparatus 200A according to the sixth embodiment is denoted by the same reference numerals, and a portion of the description of the information processing apparatus 200A according to the sixth embodiment will be omitted.
[0244] Figure 23 This is a diagram illustrating a plurality of captured images acquired by the acquisition unit 221 according to a sixth embodiment of the present disclosure. The acquisition unit 221 acquires a plurality of captured images with different subject depths. For example... Figure 23As shown, the microscope 100 acquires multiple captured images with different subject depths (focal planes) by moving the sensor 150 of the microscope 100. Figure 23 In the example, microscope 100 captures three capture images by simultaneously moving sensor 150 up and down three times to capture the target S. For example, during the first imaging, microscope 100 first captures the image where the subject depth is closest to sensor 150, and during the third imaging, microscope 100 captures the image where the subject depth is furthest from sensor 150. For example, during the second imaging, microscope 100 acquires a capture image where the subject depth is between the depth of the subject in the first imaging and the depth of the subject in the second imaging.
[0245] Notice, Figure 23 The subject depth and number of captured images shown are examples, and the invention is not limited thereto. For example, microscope 100 may capture two, four, or more captured images. Furthermore, the range of subject depth at each imaging session may be continuous or partially overlapping.
[0246] The acquisition unit 221 controls, for example, a microscope 100 to acquire captured images M61 to M63 with different subject depths.
[0247] Figure 24 This is a diagram illustrating the acquisition of a flat region by the region acquisition unit 225 according to the sixth embodiment of this disclosure. Here, the acquisition unit 221 causes the microscope 100 to capture a region having, for example, a flat region. Figure 23 The captured images at different subject depths are shown, thus obtaining images such as... Figure 24 The captured images shown are M61 to M63.
[0248] The region acquisition unit 225 divides each of the captured images M61 to M63 into blocks. For each block, the region acquisition unit 225 calculates a contrast value (contrast). The region acquisition unit 225 compares the calculated contrast value (contrast) of the block with a threshold (th3) and extracts blocks whose contrast value (contrast) is equal to or greater than the threshold (th3) as flat regions.
[0249] For example, in Figure 24In the captured image M61, region acquisition unit 225 extracts region M61A from the captured image M61. Region M61A includes blocks in the captured image M61 whose contrast value (contrast) is equal to or greater than the threshold th3. Similarly, region acquisition unit 225 extracts region M62A from the captured image M62. Region M62A includes blocks in the captured image M62 whose contrast value (contrast) is equal to or greater than the threshold th3. Similarly, region acquisition unit 225 extracts region M63A from the captured image M63. Region M62A includes blocks in the captured image M63 whose contrast value (contrast) is equal to or greater than the threshold th3.
[0250] Finally, the region acquisition unit 225 combines the extracted regions M61A to M63A to generate an image M64 that includes flat regions. In other words, the region acquisition unit 225 generates image M64 by combining the focused and unblurred regions of the captured images M61 to M63.
[0251] Note that the processing performed by the normal calculation unit 222 and the depth calculation unit 223 on the image M64, which is a flat region, is the same as the processing in the fifth embodiment.
[0252] As described above, the information processing apparatus 200A according to the sixth embodiment acquires multiple captured images with different subject depths. The information processing apparatus 200A divides each of the acquired multiple captured images into multiple blocks. The information processing apparatus 200A obtains a flat region based on the contrast value (contrast) of each divided block.
[0253] As a result, the information processing device 200A obtains a flat region based on the contrast value contrast, and then calculates the depth in the flat region that satisfies assumptions 1 and 2 in the above expressions (1) to (3). Furthermore, the information processing device 200A can obtain flat regions from multiple captured images and expand the flat regions based on the contrast value contrast.
[0254] Note that here, the region acquisition unit 225 combines regions M61A to M63A, but the invention is not limited thereto. For example, the normal calculation unit 222 and the depth calculation unit 223 can calculate normal information and depth information for each of regions M61A to M63A respectively, and for example, when the display control unit 224 causes the display unit to display the results, the result of combining regions M61A to M63A can be displayed.
[0255] <8. Seventh Implementation Method>
[0256] In the first to sixth embodiments described above, the case where the region acquisition unit 225 acquires a flat region from the captured image has been described. Besides the examples above, high-frequency components can also be extracted from the normal information calculated by the normal calculation unit 222. Therefore, in the seventh embodiment, an example of extracting high-frequency components from the normal information and calculating depth information will be described.
[0257] Figure 25 This is a diagram illustrating an example configuration of the information processing apparatus 200B according to the seventh embodiment of this disclosure. Except that the control unit 220B does not include the region acquisition unit 225 and includes the normal frequency separation unit 226B, Figure 25 The information processing device 200B shown has the same as Figure 6 The components are the same as those in the information processing device 200A shown.
[0258] The normal calculation unit 222 calculates the normal information in each pixel of the captured image acquired by the acquisition unit 221.
[0259] Normal frequency separation unit 226 separates high-frequency components (hereinafter also referred to as high-frequency normals) from normal information. Here, as described above, the depth information to be calculated by depth calculation unit 223 is information about the local convexity and concavity of target S. In other words, the depth information calculated by depth calculation unit 223 is the high-frequency component of the shape of the surface of target S, and does not include low-frequency components. Normal frequency separation unit 226 removes the low-frequency components of normal information as approximate shape from normal information to generate high-frequency normals. Since high-frequency normals do not include low-frequency components (normal information of approximate shape), high-frequency normals satisfy assumptions 1 and 2 of the above expressions (1) to (3).
[0260] Figure 26 This is a diagram used to describe the frequency separation performed by the normal frequency separation unit 226B according to the seventh embodiment of this disclosure.
[0261] For example, suppose it will be as follows Figure 26 The normal information X shown on the upper side is input to the normal frequency separation unit 226B. Note that in Figure 26 In the diagram, the actual surface of target S is indicated by dashed lines, the approximate shape of the surface of target S is indicated by solid lines, and the normal to the surface of target S is indicated by arrows.
[0262] In this case, the normal frequency separation unit 226B separates the high-frequency and low-frequency components of the normal information by using, for example, a convolution filter F(X), and extracts the high-frequency and low-frequency components of the normal information. Figure 26 The high-frequency normal Y is shown on the middle side. Note that the convolution filter F(X) is a pre-designed filter used to extract the high-frequency component when input normal information is received.
[0263] Depth calculation unit 223 can perform a transformation on the high frequency of the normal separated by normal frequency separation unit 226B using the above expressions (1) to (3), and calculate depth information Z (see below) under the condition of satisfying assumptions 1 and 2. Figure 26 (the lower side).
[0264] In this way, the information processing device 200B extracts high-frequency components from the normal information and calculates depth information based on the extracted high-frequency components.
[0265] As a result, the information processing device 200B can perform the transformation using the above expressions (1) to (3) while satisfying assumptions 1 and 2, and thus, the reduction in accuracy when calculating depth information can be suppressed. Therefore, compared with the case where frequency separation is not performed, the information processing device 200B can improve the accuracy of shape measurement.
[0266] <9. Eighth Implementation Method>
[0267] In the seventh embodiment, the case of calculating depth information based on the high-frequency components of the normal information has been described. Besides the example above, depth information can also be calculated using low-frequency components other than the high-frequency components of the normal information. Therefore, in the eighth embodiment, an example of calculating depth information by separating the high-frequency and low-frequency components from the normal information will be described.
[0268] Figure 27 This is a diagram illustrating an example configuration of the information processing apparatus 200C according to the eighth embodiment of this disclosure. Besides the normal frequency separation unit 226C and the depth calculation unit 223C of the control unit 220C, Figure 27 The information processing device 200C shown includes and Figure 25 The components are the same as those in the information processing device 200B shown.
[0269] The normal frequency separation unit 226C separates high-frequency components and low-frequency components (hereinafter also referred to as low-frequency normals) from the normal information and outputs each of the high-frequency and low-frequency components to the depth calculation unit 223C. As mentioned above, the local unevenness of the target S contributes to the high-frequency components of the normal information. Furthermore, the general shape of the target S contributes to the low-frequency components of the normal information. (Refer to...) Figure 28 This will be described.
[0270] Figure 28 This is a diagram illustrating frequency separation performed by the normal frequency separation unit 226C according to the eighth embodiment of this disclosure.
[0271] For example, suppose that Figure 28(a) The normal information X shown is input to the normal frequency separation unit 226C. Note that in Figure 28 In the diagram, the actual surface of target S is indicated by dashed lines, the approximate shape of the surface of target S is indicated by solid lines, and the normal to the surface of target S is indicated by arrows.
[0272] In this case, the normal frequency separation unit 226C extracts the high-frequency components of the normal information by using, for example, a convolution filter F(X), and extracts the high frequency of the normal YH = F(X). Figure 28 As shown on the upper side of (b), the surface shape represented by the high frequency of the normal YH is a shape with the approximate shape (low frequency of the normal YL) removed and with local bumps and depressions on a flat surface. Note that the convolution filter F(X) is a pre-designed filter used to extract the high frequency component when the input normal information is received.
[0273] Furthermore, the normal frequency separation unit 226C extracts the low normal frequency YL by subtracting the high normal frequency YH from the normal information X. That is, the normal frequency separation unit 226C extracts the low normal frequency YL by calculating the low normal frequency YL = X - YH = XF(X). Figure 28 As shown on the lower side of (b), the surface shape represented by the low frequency of the normal YL is the approximate shape with local bumps removed.
[0274] The depth calculation unit 223C performs a transformation on each of the high-frequency normal YH and the low-frequency normal YL using an expression, and calculates as follows: Figure 28 (c) shows the high-frequency component ZH (hereinafter also referred to as high-frequency depth) and low-frequency component ZL (hereinafter also referred to as low-frequency depth) of the depth information.
[0275] Here, since the normal high frequency YH satisfies assumptions 1 and 2 of the transformation using the above expression, the depth calculation unit 223C can accurately calculate the depth high frequency ZH.
[0276] Furthermore, the depth calculation unit 223C calculates the low-frequency depth ZL by performing a transformation on the low-frequency normal YL with each of λ and μ in the above expressions (1) to (3) set to 0 (λ = 0 and μ = 0). As described above, by setting the coefficient of the cost term to 0, the low-frequency depth ZL can be calculated without excluding assumptions 1 and 2. Note that the low-frequency normal YL is the low-frequency component of the normal information X after removing the high-frequency normal YH, and is normal information indicating the state in which local unevenness has been removed from the shape of the surface of the target S. Therefore, even when performing a transformation on the low-frequency normal YL by using the expression without excluding assumptions 1 and 2, the depth calculation unit 223C can accurately calculate the low-frequency depth ZL.
[0277] The depth computing unit 223C calculates depth by combining the high-frequency depth ZH and the low-frequency depth ZL. Figure 28 (d) shows the depth information Z.
[0278] In this manner, the information processing device 200C performs frequency separation on the normal information calculated based on the captured image. The information processing device 200C calculates depth information for each separated frequency, combines the calculated depth information, and calculates the depth information of the target S.
[0279] As a result, in addition to the depth of local unevenness, the information processing device 200C can also calculate the depth of the approximate shape. Furthermore, the information processing device 200C can combine these depths to calculate the depth information of the target S more accurately.
[0280] Note that here, the depth calculation unit 223C performs a transformation on the normal low frequency YL with each of λ and μ in the above expressions (1) to (3) set to 0 (λ = 0 and μ = 0), but the invention is not limited thereto. The depth calculation unit 223C can calculate the depth low frequency ZL under the condition that the influence of assumptions 1 and 2 is small, and can set the values of λ and μ to sufficiently small values such that assumptions 1 and 2 do not affect the calculation of the depth low frequency ZL.
[0281] <10. Ninth Implementation Method>
[0282] In the first embodiment, the case of training the learning machine 300 using the image M16 as the truth value has been described. Besides the example above, the information processing apparatus 200A can train the learning machine 300 using an image obtained by shifting the image M16 as correct answer data. Therefore, in the ninth embodiment, reference will be made to... Figure 29 The description information processing device 200A is an example of training the learning machine 300 by using an image obtained by shifting the image M16 as the correct answer data. Figure 29 This is a diagram illustrating the learning method of the learning machine 300 according to the ninth embodiment. Note that, except for the operation of the normal calculation unit 222, the information processing apparatus 200A according to the ninth embodiment has the same configuration and operates in the same manner as the information processing apparatus 200A according to the first embodiment, and therefore, the information processing apparatus 200A according to the ninth embodiment is indicated by the same reference numerals, and a portion of the description of the information processing apparatus 200A according to the ninth embodiment is omitted.
[0283] As described above, the normal calculation unit 222 performs learning by updating the weights of the learning machine 300 using the captured image M14 as input data and the normal image M16 as correct answer data (true value). Here, the captured image M14 is captured by a device such as a microscope 100 that captures RGB images. On the other hand, the normal image M16 is generated based on shape information measured by a device such as a non-contact 3D measuring instrument that directly measures the surface shape of the target S.
[0284] As described above, when the optical systems of the device acquiring the input image and the device acquiring the correct answer image are different, it becomes difficult to align the input image and the correct answer image. Therefore, the output image and the correct answer image obtained by inputting the input image to the learning machine 300 can be shifted by several pixels. As described above, when the learning machine 300 is trained in the state of shifting the output image and the correct answer image, the learning machine 300 generates a blurred output image for the input image.
[0285] Therefore, as Figure 29 As shown, during the learning process to determine the filter coefficients (weights) of the learning machine (CNN) 300, the normal calculation unit 222 shifts the normal image M16 in the x and y directions and calculates the similarity between the shifted normal image M16 and the output image M15. Note that the normal calculation unit 222 shifts the normal image M16 in the x direction within the range of 0 to ±x and in the y direction within the range of 0 to ±y.
[0286] For example, the normal calculation unit 222 calculates the least squares error between the shifted normal image M16 and the output image M15 for each shift amount. The normal calculation unit 222 updates the weights of the learning machine 300 using the minimum value of the calculated least squares error as the final loss. Alternatively, instead of the least squares error, the minimum average error between the shifted normal image M16 and the output image M15 is calculated for each shift amount, and the minimum value of the calculated minimum average error can be used as the final loss.
[0287] In this way, the learning machine 300 is generated by updating weights based on the similarity between a shifted image (shifted normal image) obtained by shifting the ground truth image (normal image) and the output image M15. More specifically, the learning machine 300 is generated by updating weights based on the similarity between each of a plurality of shifted images (also called shifted normal images or correct answer candidate images) and the output image M15, said plurality of shifted images being obtained by shifting the ground truth image (normal image) by different amounts of shift (different numbers of pixels).
[0288] As a result, even without aligning the output image M15 of the learning machine 300 with the normal image M16, the weights of the learning machine 300 can be updated while the two images are aligned. Consequently, the accuracy of training the learning machine 300 can be further improved. Furthermore, since the normal calculation unit 222 calculates normal information using the learning machine 300, the normal information can be calculated with high accuracy, and the accuracy of shape measurement by the information processing device 200A can be further improved.
[0289] <11. Tenth Implementation Method>
[0290] In the first embodiment described above, the case of obtaining a flat region from a captured image has been described. Besides the example above, the information processing apparatus 200A can divide the captured image into multiple regions and obtain a flat region for each of the divided regions. Therefore, in the tenth embodiment, reference will be made to... Figure 30 The description information processing device 200A divides a captured image into multiple regions and obtains an example of a flat region for each of the divided regions. Figure 30 This is an example showing the configuration of the control unit 220D of the information processing apparatus 200A according to the tenth embodiment of the present disclosure.
[0291] like Figure 30 As shown, in addition to including a region acquisition unit 225D instead of a region acquisition unit 225, and also including a depth combination unit 227, the control unit 220D of the information processing device 200A and Figure 6 The control unit 220 shown is the same.
[0292] like Figure 30 As shown, the region acquisition unit 225D divides the captured image M71 acquired by the acquisition unit 221 into three regions M71A to M71C. Note that the number of regions divided by the region acquisition unit 225D is not limited to three, and can be two, four or more.
[0293] Region acquisition unit 225D acquires flat regions M72A to M72C for each of the divided regions M71A to M71C. Normal calculation unit 222 calculates normal region images M73A to M73C based on the flat regions M72A to M72C, and depth calculation unit 223 calculates depth region images M74A to M74C based on the normal region images M73A to M73C.
[0294] The depth combination unit 227 combines depth region images M74A to M74C to generate a depth image M75, and outputs the depth image M75 to the display control unit 224.
[0295] Note that, in addition to depth region images M74A to M74, depth combination unit 227 can also combine the divided flat regions M72A to M72C or combine normal region images M73A to M73C. Furthermore, the processing in each unit can be performed sequentially for each region, or the processing in each unit can be performed in parallel for each region.
[0296] In this manner, the information processing device 200A divides the captured image into multiple segmented regions M71A to M71C, and for each of the segmented regions M71A to M71C, it acquires a segmented flat region M72A to M72C. Furthermore, the information processing device 200A calculates normal information and depth information for each of the segmented flat regions M72A to M72C.
[0297] As a result, the size of the captured image (region) to be processed by each unit can be reduced, the processing load can be reduced, and the processing speed can be increased.
[0298] <12. Other Implementation Methods>
[0299] The processing according to each of the above embodiments can be performed in various different modes other than those in each of the above embodiments.
[0300] In the above embodiments, an example has been described where the head mount 10 is a cylindrical portion mounted on the distal end of the microscope 100. However, the head mount 10 need not have a cylindrical shape as long as it is a structure used to keep the distance between the target S and the sensor 150 of the microscope 100 constant.
[0301] <13. Applicable Examples>
[0302] Reference Figure 31 Examples of applicable information processing apparatus 200A to 200C according to the first to tenth embodiments are described. Figure 31 This is a block diagram illustrating an example of a schematic configuration of a patient in vivo information acquisition system using a capsule endoscope to which the technology (the technology) according to this disclosure can be applied.
[0303] The in vivo information acquisition system 10001 includes a capsule endoscope 10100 and an external control device 10200.
[0304] The capsule endoscope 10100 is swallowed by the patient during the examination. The capsule endoscope 10100 has imaging and wireless communication functions. As it moves within an organ such as the stomach or intestine through peristaltic movements until it is naturally expelled from the patient's body, it sequentially captures images of the inside of the organ (hereinafter also referred to as in vivo images) at predetermined intervals, and sequentially wirelessly transmits information about the in vivo images to an external control device 10200 located outside the body.
[0305] The external control device 10200 controls the operation of the in vivo information acquisition system 10001 as a whole. In addition, the external control device 10200 receives information related to in vivo images transmitted from the capsule endoscope 10100, and generates image data for displaying the in vivo images on a display device (not shown) based on the received information about the in vivo images.
[0306] In this way, the in vivo information acquisition system 10001 can obtain in vivo images of the patient's internal state at any time from swallowing the capsule endoscope 10100 to expelling the capsule endoscope 10100.
[0307] The configuration and functions of the capsule endoscope 10100 and the external control device 10200 will be described in more detail.
[0308] The capsule endoscope 10100 includes a capsule-shaped housing 10101, and a light source unit 10111, an imaging unit 10112, an image processing unit 10113, a wireless communication unit 10114, a power supply unit 10115, a power supply unit 10116, and a control unit 10117 are housed in the housing 10101.
[0309] The light source unit 10111 is implemented, for example, by a light source such as a light-emitting diode (LED), and uses light to illuminate the imaging field of view of the imaging unit 10112.
[0310] Imaging unit 10112 includes an imaging element and an optical system, the optical system including a plurality of lenses disposed upstream of the imaging element. Reflected light (hereinafter referred to as observation light) illuminating the body tissue to be observed is collected by the optical system and incident on the imaging element. In imaging unit 10112, the observation light incident on the imaging element undergoes photoelectric conversion, and an image signal corresponding to the observation light is generated. The image signal generated by imaging unit 10112 is provided to image processing unit 10113.
[0311] The image processing unit 10113 is implemented by a processor such as a central processing unit (CPU) or a graphics processing unit (GPU), and performs various types of signal processing on the image signal generated by the imaging unit 10112. The image processing unit 10113 provides the processed image signal as raw data to the wireless communication unit 10114.
[0312] The wireless communication unit 10114 performs predetermined processing, such as modulation processing, on the image signal that has undergone signal processing by the image processing unit 10113, and transmits the image signal to the external control device 10200 via the antenna 10114A. Furthermore, the wireless communication unit 10114 receives control signals related to the drive control of the capsule endoscope 10100 from the external control device 10200 via the antenna 10114A. The wireless communication unit 10114 provides the control signals received from the external control device 10200 to the control unit 10117.
[0313] The power supply unit 10115 includes a power receiving antenna coil, a power regeneration circuit that regenerates power from the current generated in the antenna coil, a boost circuit, etc. In the power supply unit 10115, power is generated using a so-called contactless charging principle.
[0314] The power supply unit 10116 is implemented using a secondary battery and stores the power generated by the power feeder unit 10115. Figure 31 In order to avoid the complexity of the figure, arrows indicating the destination of the power supplied from the power supply unit 10116 are not shown. However, the power stored in the power supply unit 10116 is supplied to the light source unit 10111, the imaging unit 10112, the image processing unit 10113, the wireless communication unit 10114, and the control unit 10117, and can be used to drive these units.
[0315] The control unit 10117 is implemented by a processor such as a CPU, and appropriately controls the driving of the light source unit 10111, the imaging unit 10112, the image processing unit 10113, the wireless communication unit 10114, and the power supply unit 10115 according to the control signals sent from the external control device 10200.
[0316] The external control device 10200 is implemented by a processor such as a CPU or GPU, a microcomputer that integrates a processor and a storage element such as memory, a control board, etc. The external control device 10200 controls the operation of the capsule endoscope 10100 by sending control signals to the control unit 10117 of the capsule endoscope 10100 via antenna 10200A. In the capsule endoscope 10100, for example, the light illumination conditions of the observed target in the light source unit 10111 can be changed according to the control signals from the external control device 10200. Furthermore, imaging conditions (e.g., frame rate, exposure value, etc. in the imaging unit 10112) can be changed according to the control signals from the external control device 10200. Additionally, the processing content in the image processing unit 10113 and the conditions for transmitting image signals in the wireless communication unit 10114 (e.g., transmission interval, number of images to be transmitted, etc.) can be changed according to the control signals from the external control device 10200.
[0317] Furthermore, the external control device 10200 performs various types of image processing on the image signals transmitted from the capsule endoscope 10100 and generates image data for displaying the captured in vivo image on a display device. For example, various types of signal processing, such as enhancement processing (de-mosaic processing), high image quality processing (bandwidth enhancement processing, super-resolution processing, noise reduction processing, image stabilization processing, etc.), and magnification processing (electronic zoom processing), can be performed individually or in combination as image processing. The external control device 10200 controls the drive of the display device to display the captured in vivo image based on the generated image data. Alternatively, the external control device 10200 can cause a recording device (not shown) to record the generated image data or cause a printing device (not shown) to print out the generated image data.
[0318] Examples of in vivo information acquisition systems to which the technology according to this disclosure can be applied have been described above. The technology according to this disclosure can be applied to, for example, the external control device 10200 in the configuration described above. By applying the technology according to this disclosure to the external control device 10200, the shape of surfaces within the body can be measured based on in vivo images captured by the capsule endoscope 10100.
[0319] (Examples of the application of endoscopic surgical systems)
[0320] The technology based on this disclosure can also be applied to endoscopic surgical systems. Figure 32 This is a diagram illustrating an example of a schematic configuration of an endoscopic surgical system to which the technology (the technology) according to this disclosure can be applied.
[0321] Figure 32The illustration shows an operator (physician) 11131 performing surgery on a patient 11132 on a hospital bed 11133 using an endoscopic surgical system 11000. As shown, the endoscopic surgical system 11000 includes an endoscope 11100, other surgical instruments 11110 such as a pneumoperitoneum tube 11111 and an energy therapy tool 11112, a support arm assembly 11120 supporting the endoscope 11100, and a trolley 11200 equipped with various devices for endoscopic surgery.
[0322] Endoscope 11100 includes a lens barrel 11101 and a camera head 11102, in which a region corresponding to a predetermined length from the distal end is inserted into the body cavity of patient 11132, and the camera head 11102 is connected to the proximal end of the lens barrel 11101. In the example shown, an endoscope 11100 is shown configured as a so-called rigid endoscope including a rigid lens barrel 5117; however, endoscope 11100 can be configured as a so-called flexible endoscope including a flexible lens barrel 5117.
[0323] An opening for mounting to the objective lens is provided at the distal end of the lens barrel 11101. A light source device 11203 is connected to the endoscope 11100, and light generated by the light source device 11203 is guided to the distal end of the lens barrel via a light guide extending inside the lens barrel 11101, and emitted via the objective lens toward the target for observation within the body cavity of the patient 11132. Note that the endoscope 11100 can be a forward-looking endoscope, an oblique-looking endoscope, or a lateral-looking endoscope.
[0324] An optical system and imaging element are housed inside the camera head 11102, and reflected light (observation light) from the observed target is converged onto the imaging element via the optical system. The imaging element performs photoelectric conversion on the observation light and generates an electrical signal corresponding to the observation light—in other words, an image signal corresponding to the observed image. The image signal is sent as raw data to the camera control unit (CCU) 11201.
[0325] The CCU 11201 includes a CPU, GPU, etc., and controls the operation of the endoscope 11100 and the display device 11202 as a whole. In addition, the CCU 11201 receives image signals from the camera head 11102 and performs various types of image processing based on the image signals—for example, image signal rendering (de-mosaic processing)—for displaying the image.
[0326] The display device 11202 displays an image based on an image signal that has undergone image processing by the CCU 11201 under the control of the CCU 11201.
[0327] The light source device 11203 is implemented, for example, by a light source such as a light-emitting diode (LED), and supplies illumination light to the endoscope 11100 for capturing images of surgical sites, etc.
[0328] Input device 11204 is an input interface for endoscopic surgical system 11000. Users can input various types of information or commands into endoscopic surgical system 11000 via input device 11204. For example, users can input commands to change the imaging conditions (type of illumination light, magnification, focal length, etc.) of endoscope 11100, etc.
[0329] Treatment tool control device 11205 controls the drive of energy therapy tool 11112 used for tissue cauterization and incision, vascular closure, etc. Pneumoperitoneum device 11206 feeds gas into the patient's body cavity 11132 via pneumoperitoneum tube 11111 to inflate the body cavity, ensuring a clear view for endoscope 11100 and ensuring the operator's working space. Recorder 11207 is a device capable of recording various types of information about the surgery. Printer 11208 is a device capable of printing various types of information about the surgery in various formats such as text, images, or graphics.
[0330] Note that the light source device 11203 supplying illumination light to the endoscope 11100 when capturing images of the surgical site may include, for example, a white light source implemented by an LED, a laser light source, or a combination of a white light source and a laser light source. In the case of implementing a white light source using a combination of RGB laser light sources, the output intensity and timing of each color (each wavelength) can be controlled with high accuracy, and therefore, white balance adjustment of the captured image can be performed within the light source device 11203. Furthermore, in this case, the observation target is illuminated with lasers from each of the RGB laser light sources in a time-division manner, and the driving of the imaging element of the camera head 11102 is controlled synchronously with the illumination timing, so that images corresponding to each of the RGB light sources can also be captured in a time-division manner. Using this method, color images can be obtained without setting a color filter in the imaging element.
[0331] Furthermore, the drive of the light source device 11203 can be controlled to change the intensity of the light to be output at each predetermined time. The drive of the imaging element of the camera head 11102 is controlled in time synchronization with the change in light intensity so as to acquire images in a time-divided manner and combine the images, thereby generating high dynamic range images without the so-called underexposure and overexposure.
[0332] Furthermore, the light source device 11203 can be configured to supply light in a predetermined wavelength band corresponding to a specific light observation. In specific light observation, for example, so-called narrowband imaging is performed, in which images of predetermined tissues, such as blood vessels in the mucosal epithelium, are captured with high contrast by irradiating light with a narrower band than the irradiating light (i.e., white light) during normal observation, by utilizing the wavelength dependence of light absorption in body tissue. Alternatively, in specific light observation, fluorescence observation can be performed to obtain an image by irradiating the generated fluorescence with excitation light. In fluorescence observation, for example, fluorescence from body tissue can be observed by irradiating the body tissue with excitation light (autofluorescence observation), or a fluorescence image can be obtained by locally injecting a reagent such as indocyanine green (ICG) into the body tissue and irradiating the body tissue with excitation light corresponding to the fluorescence wavelength of the reagent. The light source device 11203 can be configured to supply narrowband light and / or excitation light corresponding to such specific light observation.
[0333] Figure 33 It is shown Figure 32 A block diagram illustrating an example of the functional configuration of the camera head 11102 and CCU 11201.
[0334] The camera head 11102 includes a lens unit 11401, an imaging unit 11402, a driving unit 11403, a communication unit 11404, and a camera head control unit 11405. The CCU 11201 includes a communication unit 11411, an image processing unit 11412, and a control unit 11413. The camera head 11102 and the CCU 11201 are communicatively connected to each other via a transmission cable 11400.
[0335] Lens unit 11401 is an optical system disposed at the portion where the camera head 11102 connects to the lens barrel 11101. Observation light incident from the distal end of the lens barrel 11101 is guided to the camera head 11102 and incident on lens unit 11401. Lens unit 11401 is implemented by combining multiple lenses, including zoom lenses and focusing lenses.
[0336] Imaging unit 11402 includes imaging elements. The number of imaging elements included in imaging unit 11402 can be one (so-called single-plate type) or multiple (so-called multi-plate type). When imaging unit 11402 is configured as multi-plate type, for example, image signals corresponding to RGB can be generated by each imaging element, and a color image can be obtained by combining the image signals. Alternatively, imaging unit 11402 may include a pair of imaging elements for acquiring image signals for the right and left eyes corresponding to three-dimensional (3D) display. With the execution of 3D display, operator 11131 can more accurately grasp the depth of living tissue at the surgical site. Note that when imaging unit 11402 is configured as multi-plate type, multiple lens units 11401 can be provided corresponding to each imaging element.
[0337] Furthermore, the imaging unit 11402 does not necessarily have to be located in the camera head 11102. For example, the imaging unit 11402 can be located directly behind the objective lens inside the lens barrel 11101.
[0338] The drive unit 11403 is implemented by an actuator and, under the control of the camera head control unit 11405, moves the zoom lens and focusing lens of the lens unit 11401 a predetermined distance along the optical axis. As a result, the magnification and focus of the image captured by the imaging unit 11402 can be appropriately adjusted.
[0339] The communication unit 11404 is implemented by a communication device for transmitting various types of information to and receiving various types of information from the CCU 11201. The communication unit 11404 transmits the image signal obtained from the imaging unit 11402 as raw data to the CCU 11201 via the transmission cable 11400.
[0340] Furthermore, the communication unit 11404 receives control signals from the CCU 11201 for controlling the drive of the camera head 11102, and supplies the control signals to the camera head control unit 11405. The control signals include, for example, information about imaging conditions, such as information for specifying the frame rate of the captured image, information for specifying the exposure value during imaging, and / or information for specifying the magnification and focus of the captured image.
[0341] Note that imaging conditions such as frame rate, exposure value, magnification, and focus can be appropriately specified by the user, or can be automatically set by the control unit 11413 of CCU 11201 based on the acquired image signal. In the latter case, endoscope 11100 has so-called automatic exposure (AE), automatic focus (AF), and automatic white balance (AWB) functions.
[0342] The camera head control unit 11405 controls the driving of the camera head 11102 based on the control signals received from the CCU 11201 via the communication unit 11404.
[0343] The communication unit 11411 is implemented by a communication device for transmitting various types of information to and receiving various types of information from the camera head 11102. The communication unit 11411 receives image signals transmitted from the camera head 11102 via a transmission cable 11400.
[0344] In addition, the communication unit 11411 sends control signals for controlling the driving of the camera head 11102 to the camera head 11102. The image signal or control signal can be transmitted via electrical communication, optical communication, etc.
[0345] The image processing unit 11412 performs various types of image processing on the image signal, which is raw data sent from the camera head 11102.
[0346] The control unit 11413 performs various types of control related to the capture of images of surgical sites, etc., performed by the endoscope 11100, and the display of the captured images obtained by capturing images of surgical sites, etc. For example, the control unit 11413 generates control signals for controlling the drive of the camera head 11102.
[0347] Furthermore, based on the image signal processed by the image processing unit 11412, the control unit 11413 causes the display device 11202 to display a captured image of the surgical site, etc. At this time, the control unit 11413 can identify various objects in the captured image using various image recognition technologies. For example, the control unit 11413 can identify surgical instruments such as forceps, specific parts of an organism, bleeding, and fog when using the energy therapy tool 11112 by detecting the edge shape, color, etc., of objects included in the captured image. When the captured image is displayed on the display device 11202, the control unit 11413 can overlay various types of surgical support information onto the image of the surgical site using the recognition results. The surgical support information is overlaid, displayed, and presented to the operator 11131, thereby reducing the burden on the operator 11131 and allowing the operator 11131 to perform surgery reliably.
[0348] The transmission cable 11400 connecting the camera head 11102 and the CCU 11201 is an electrical signal cable that supports electrical signal communication, an optical fiber that supports optical communication, or a composite cable thereof.
[0349] Here, in Figure 33In the example, wired communication is performed using transmission cable 11400, but wireless communication can also be performed between camera head 11102 and CCU 11201.
[0350] Examples of endoscopic surgical systems to which the technology according to this disclosure can be applied have been described above. The technology according to this disclosure can be applied, for example, to the CCU 11201 in the configuration described above. Specifically, the control units 220, 220B, and 220C described above can be applied to the image processing unit 11412. By applying the technology according to this disclosure to the image processing unit 11412, the shape of surfaces within the body can be measured based on in vivo images captured by the endoscope 11100.
[0351] Note that an endoscopic surgical system has been described here as an example, but the techniques based on this disclosure can be applied to, for example, microsurgical systems.
[0352] <14. Supplementary Description>
[0353] As described above, preferred embodiments of the present disclosure have been described in detail with reference to the accompanying drawings; however, the technical scope of the present disclosure is not limited to these examples. It will be apparent to those skilled in the art that various modifications or alterations can be conceived within the scope of the technical concept described in the claims, and it is naturally understood that such modifications or alterations fall within the technical scope of the present disclosure.
[0354] Furthermore, in the various processes described in the above embodiments, all or some of the processes described as automatically executed may be performed manually. Alternatively, all or some of the processes described as manually executed may be performed automatically by known methods. Moreover, unless otherwise stated, the processing procedures, specific names, and information including various data and parameters shown in the specification and drawings may be arbitrarily changed. For example, the various information shown in each drawing is not limited to the information shown.
[0355] Furthermore, each component illustrated in each device is functionally conceptual and does not necessarily have to be physically configured as shown in the figures. That is, the specific distribution / integration patterns of the individual devices are not limited to those shown in the figures. All or some of the devices below can be functionally or physically distributed / integrated in any arbitrary unit according to various loads or usage conditions.
[0356] Furthermore, the above-described embodiments and modifications can be appropriately combined as long as the processing content does not contradict each other. Additionally, while a microscope has been described as an example of an image processing device in the above embodiments, the image processing described herein can also be applied to imaging devices other than microscopes.
[0357] Furthermore, the effects described in this specification are illustrative or exemplary only, and not restrictive. That is, in addition to or instead of the effects described above, the technology according to this disclosure may exhibit other effects that will be apparent to those skilled in the art based on the description herein.
[0358] Note that the following configurations also fall within the technical scope of this disclosure. (1)
[0360] An information processing device, comprising:
[0361] The control unit acquires a captured image of the target by a sensor, extracts a flat region from the captured image based on the brightness values of the captured image, and calculates shape information about the surface shape of the target based on the flat region of the captured image and information about the sensor.
[0362] The captured image is obtained based on the reflected light from multiple light sources emitted to the target from different locations. (2)
[0364] According to the information processing apparatus of (1), the control unit acquires the captured image based on the reflected light of the light emitted simultaneously from the plurality of light sources to the target. (3)
[0366] According to the information processing apparatus of (1) or (2), the control unit sets the region in which the brightness value of the captured image is equal to or greater than a predetermined threshold as the flat region. (4)
[0368] According to the information processing apparatus of (1), the control unit acquires a captured image for each of the light sources based on the reflected light of the light emitted from one of the plurality of light sources to the target, and extracts a region where the change in brightness value among the plurality of captured images acquired for each of the light sources is less than a predetermined threshold as the flat region. (5)
[0370] The information processing apparatus according to any one of (1) to (4), wherein the control unit generates a smoothed image by performing a smoothing process on the captured image, and extracts the flat region from the captured image based on the smoothed image. (6)
[0372] The information processing apparatus according to any one of (1) to (5), wherein the control unit divides the captured image into a plurality of partitioned regions and extracts the flat region for each of the plurality of partitioned regions to calculate shape information. (7)
[0374] The information processing apparatus according to any one of (1) to (6), wherein the control unit calculates the normal information of the flat region in the captured image and inputs the normal information into a model generated by machine learning to obtain the shape information. (8)
[0376] According to the information processing apparatus described in (7), the model is generated by updating weights based on comparisons between multiple candidate images of correct answers and the output data of the model, and
[0377] The multiple candidate images for correct answers are generated by shifting the correct answer image by different numbers of pixels. (9)
[0379] According to the information processing apparatus of (8), the model calculates the least squares error between each of the plurality of candidate images for correct answers and the output data, and generates the model by updating the weights based on the minimum of the plurality of least squares errors. (10)
[0381] According to the information processing apparatus of (1), the control unit extracts the flat region based on a contrast value calculated from the brightness value of the captured image. (11)
[0383] According to the information processing apparatus of (1), the control unit divides the captured image into multiple partition regions, calculates the contrast value for each partition region in the multiple partition regions, and determines whether the partition region is flat based on the contrast value, so as to extract the flat region. (12)
[0385] According to the information processing apparatus of (10) or (11), the control unit acquires a plurality of captured images of the sensor whose depth planes are different from each other, and extracts the flat region from each of the plurality of captured images. (13)
[0387] An information processing method, comprising:
[0388] Acquire captured images of the target captured by the sensor;
[0389] Extract flat regions from the captured image based on the brightness values of the captured image; and
[0390] Shape information about the surface shape of the target is calculated based on the flat area of the captured image and information about the sensor.
[0391] The captured image is obtained based on the reflected light from multiple light sources emitted to the target from different locations.
[0392] List of reference numerals
[0393] 10-head installation unit
[0394] 100 microscopes
[0395] 150 sensors
[0396] 160 point light source
[0397] 200 Information Processing Devices
[0398] 220 Control Unit
[0399] 221 Acquisition Unit
[0400] 222 Normal Calculation Unit
[0401] 223 Depth Computing Units
[0402] 224 Display Control Unit
[0403] 225 Area Acquisition Unit
[0404] 226 Normal Frequency Separation Units
[0405] 227 deep combination units
[0406] 230 storage units
Claims
1. An information processing apparatus, comprising: The control unit acquires a captured image of the target by a sensor, extracts a flat region from the captured image based on the brightness values of the captured image, and calculates shape information about the surface shape of the target based on the flat region of the captured image and information about the sensor. The captured image is an image obtained based on reflected light emitted from multiple light sources arranged at different locations and directed to the target. The control unit sets the region where the brightness value of the captured image is equal to or greater than a predetermined threshold as the flat region.
2. The information processing apparatus according to claim 1, wherein, The control unit acquires the captured image based on the reflected light from the light emitted simultaneously from the multiple light sources to the target.
3. The information processing apparatus according to claim 1, wherein, The control unit acquires a captured image for each of the light sources based on the reflected light from the light emitted from one of the plurality of light sources to the target, and extracts the region where the change in brightness value among the plurality of captured images acquired for each of the light sources is less than a predetermined threshold as the flat region.
4. The information processing apparatus according to claim 2, wherein, The control unit generates a smoothed image by performing a smoothing process on the captured image, and extracts the flat region from the captured image based on the smoothed image.
5. The information processing apparatus according to claim 4, wherein, The control unit divides the captured image into multiple regions and extracts the flat region for each of the multiple regions to calculate shape information.
6. The information processing apparatus according to claim 5, wherein, The control unit calculates the normal information of the flat region in the captured image and inputs the normal information into a model generated by machine learning to obtain the shape information.
7. The information processing apparatus according to claim 6, wherein, The model is generated by updating weights based on comparisons between multiple candidate images of correct answers and the model's output data. The multiple candidate images for correct answers are generated by shifting the correct answer image by different numbers of pixels.
8. The information processing apparatus according to claim 7, wherein, The model calculates the least squares error between each of the plurality of candidate images for correct answers and the output data, and the model is generated by updating the weights based on the minimum of the plurality of least squares errors.
9. The information processing apparatus according to claim 1, wherein, The control unit extracts the flat area based on a contrast value calculated from the brightness value of the captured image.
10. The information processing apparatus according to claim 9, wherein, The control unit divides the captured image into multiple regions, calculates the contrast value for each of the multiple regions, and determines whether the region is flat based on the contrast value in order to extract the flat region.
11. The information processing apparatus according to claim 9, wherein, The control unit acquires multiple captured images from which the depth planes of the sensor differ from one another, and extracts the flat region from each of the multiple captured images.
12. An information processing method, comprising: Acquire captured images of the target captured by the sensor; Flat regions are extracted from the captured image based on the brightness values of the captured image; as well as Shape information about the surface shape of the target is calculated based on the flat area of the captured image and information about the sensor. The captured image is an image obtained based on reflected light emitted from multiple light sources arranged at different locations and directed to the target. Specifically, the region where the brightness value of the captured image is equal to or greater than a predetermined threshold is defined as the flat region.
13. A non-transitory computer-readable medium having a program implemented thereon, the program, when executed by a computer, causing the computer to perform an information processing method, the information processing method comprising: Acquire captured images of the target captured by the sensor; Flat regions are extracted from the captured image based on the brightness values of the captured image; as well as Shape information about the surface shape of the target is calculated based on the flat area of the captured image and information about the sensor. The captured image is an image obtained based on reflected light emitted from multiple light sources arranged at different locations and directed to the target. Specifically, the region where the brightness value of the captured image is equal to or greater than a predetermined threshold is defined as the flat region.
14. A computer program product comprising a computer program / instructions, wherein the computer program / instructions, when executed by a computer, cause the computer to perform an information processing method, the information processing method comprising: Acquire captured images of the target captured by the sensor; Flat regions are extracted from the captured image based on the brightness values of the captured image; as well as Shape information about the surface shape of the target is calculated based on the flat area of the captured image and information about the sensor. The captured image is an image obtained based on reflected light emitted from multiple light sources arranged at different locations and directed to the target. Specifically, the region where the brightness value of the captured image is equal to or greater than a predetermined threshold is defined as the flat region.
Citation Information
Patent Citations
Tip dome of microscope
JP2008253498A
Apparatus for inspecting shape of metal body, and method for inspecting shape of metal body
CN107003115A