Three-dimensional structure estimation device, three-dimensional structure estimation method, program, and recording medium
The device uses perspective images to train a learning model for accurate three-dimensional structure estimation by correcting attenuation coefficients based on signed distances, addressing the limitations of existing methods like NeRF.
Patent Information
- Application Number
- PCT/JP2025/021242
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-06-12
- Publication Date
- 2026-02-05
Smart Images

Figure JP2025021242_05022026_PF_FP_ABST
Abstract
Description
Three-dimensional structure estimation device, three-dimensional structure estimation method, program, and recording medium
[0001] The present invention relates to a technique for predicting a three-dimensional structure.
[0002] Techniques for estimating three-dimensional structures are known (Patent Document 1, Non-Patent Documents 1 to 4). For example, Patent Document 1 discloses a technique for training a neural network model used in re-rendering an image for display using NeRF (Neural Radiance Fields).
[0003] Japan Special Table No. 2024-507887
[0004] R. Zha, Y. Zhang, and H. Li, “Naf: neural attenuation fields for sparseviewcbct reconstruction,” in International Conference on Medical ImageComputing and Computer-Assisted Intervention. Springer, 2022, pp.442-452.D. R¨uckert, Y. Wang, R. Li, R. Idoughi, and W. Heidrich, “Neat: Neuraladaptive tomography,” ACM Transactions on Graphics (TOG), vol. 41,no. 4, pp. 1-13, 2022.B. Mildenhall, PP Srinivasan, M. Tancik, JT Barron, R. Ramamoorthi,and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in ECCV, 2020.P. Wang, L. Liu, Y. Liu, C. Theobalt, T. Komura, and W. Wang, “Neus:Learning neural implicit surfaces by volume rendering for multi-view reconstruction,” arXiv preprint arXiv:2106.10689, 2021
[0005] However, the method using NeRF has a problem that a large number of captured images are required to estimate the three-dimensional structure of the object with high accuracy. Moreover, NeRF assumes a visible light camera image, and when a perspective image is used, information in the depth direction of the image cannot be recognized. Furthermore, although Non-Patent Document 1 has succeeded in extending NeRF to a perspective image, it has problems of requiring a large number of captured images and poor accuracy in the boundary portion of the measurement object.
[0006] An object of one aspect of the present invention is to realize a technology for more accurately estimating the three-dimensional structure, including the internal structure, of an object without requiring a large number of captured images.
[0007] In order to solve the above problem, a three-dimensional structure estimation device according to one embodiment of the present invention comprises: an acquisition unit that acquires multiple perspective images of an object taken from different directions; a learning unit that uses the multiple perspective images to train a learning model that receives input data indicating the three-dimensional coordinates of an arbitrary point within the object and outputs first output data indicating a signed distance of the arbitrary point relative to the surface of an internal structure of the object and second output data indicating a first attenuation coefficient at the arbitrary point; a calculation unit that inputs input data indicating each point within the object to the learning model and calculates a second attenuation coefficient at each point within the object by correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data; and an estimation unit that estimates the three-dimensional structure of the object based on the second attenuation coefficient at each point within the object.
[0008] A three-dimensional structure estimation method according to one aspect of the present invention includes an acquisition step of acquiring a plurality of perspective images of an object taken from different directions; a learning step of using the plurality of perspective images to train a learning model that receives input data indicating the three-dimensional coordinates of an arbitrary point within the object and outputs first output data indicating a signed distance of the arbitrary point relative to the surface of an internal structure of the object and second output data indicating a first attenuation coefficient at the arbitrary point; a calculation step of inputting input data indicating each point within the object to the learning model and calculating a second attenuation coefficient at each point within the object by correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data; and an estimation step of estimating the three-dimensional structure of the object based on the second attenuation coefficient at each point within the object.
[0009] A program according to one aspect of the present invention is a program for causing a computer to function as a three-dimensional structure estimation device, and causes the computer to function as: an acquisition unit that acquires multiple perspective images of an object taken from different directions; a learning unit that uses the multiple perspective images to learn a learning model that receives input data indicating the three-dimensional coordinates of an arbitrary point within the object and outputs first output data indicating a signed distance of the arbitrary point relative to the surface of an internal structure of the object and second output data indicating a first attenuation coefficient at the arbitrary point; a calculation unit that inputs input data indicating each point within the object to the learning model and calculates a second attenuation coefficient at each point within the object by correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data; and an estimation unit that estimates the three-dimensional structure of the object based on the second attenuation coefficient at each point within the object.
[0010] The three-dimensional structure estimation device according to each aspect of the present invention may be realized by a computer. In this case, the control program for the three-dimensional structure estimation device, which causes the computer to operate as each part (software element) of the three-dimensional structure estimation device, thereby realizing the three-dimensional structure estimation device on a computer, and the computer-readable recording medium on which the control program is recorded, also fall within the scope of the present invention.
[0011] According to one aspect of the present invention, it is possible to more accurately estimate the three-dimensional structure of an object, including its internal structure, without requiring a large number of captured images.
[0012] FIG. 1 is a diagram showing an example of a three-dimensional structure estimated by a three-dimensional structure estimation device according to an embodiment of the present invention. FIG. 2 is a block diagram showing an example of the configuration of a three-dimensional structure estimation device according to an embodiment of the present invention. FIG. 3 is a flow diagram showing an example of the flow of a three-dimensional structure estimation method according to an embodiment of the present invention. FIG. 4 is a diagram showing a specific example of a method for capturing a fluoroscopic image according to an embodiment of the present invention. FIG. 5 is a diagram showing a specific example of a method for capturing a fluoroscopic image according to an embodiment of the present invention. FIG. 6 is a diagram showing an example of the configuration of a learning model according to an embodiment of the present invention and training processing performed by a learning unit. FIG. 7 is a diagram showing a specific example of a function according to an embodiment of the present invention. FIG. 8 is a diagram for explaining processing for estimating bone density according to an embodiment of the present invention. FIG. 9 is a diagram showing an example of the configuration of a learning
[0013] [Embodiment] (Outline of 3D Structure Estimation Device) An embodiment of the present invention will be described in detail below. The 3D construction device according to this embodiment is a device that estimates the 3D structure of an object using multiple fluoroscopic images of the object. An example of the object is, but is not limited to, the body of a person who is the subject of an examination. The object may also be, for example, a crop that is partially or completely covered with soil. A fluoroscopic image is an image obtained by viewing the object through a fluoroscopic view, and examples include two-dimensional X-ray images obtained by X-ray imaging.
[0014] 1 is a diagram showing an example of a three-dimensional structure estimated by a three-dimensional structure construction device according to this embodiment. In the example of Fig. 1, the three-dimensional structure construction device generates a three-dimensional image Img21 using multiple perspective images Img11, 12, ... obtained by capturing images of a human body from different directions.
[0015] (Configuration of 3D structure estimation device) Fig. 2 is a block diagram showing an example of the configuration of a 3D structure estimation device 10 according to this embodiment. The 3D structure estimation device 10 is an example of a computer according to the present disclosure. The 3D structure estimation device 10 includes a control unit 11, a storage unit 12, a communication unit 13, an input unit 14, and an output unit 15.
[0016] The control unit 11 executes instructions of a computer program stored in the memory unit 12. The memory unit 12 stores instructions of the computer program executed by the control unit 11. The communication unit 13 communicates with devices external to the three-dimensional structure estimation device 10 via a communication line. The specific configuration of the communication line does not limit this exemplary embodiment, but examples of the communication line include a wireless LAN (Local Area Network), a wired LAN, a WAN (Wide Area Network), a public line network, a mobile data communication network, or a combination thereof. The communication unit 13 transmits data supplied from the control unit 11 to other devices, and supplies data received from other devices to the control unit 11.
[0017] The input unit 14 is configured to receive input to the three-dimensional structure estimation device 10, and includes, for example, input devices such as a keyboard, a mouse, a touch panel, an X-ray camera (X-ray imaging device), a microphone, etc. The input unit 14 may also be configured to receive data from the input devices via an interface such as a USB (Universal Serial Bus).
[0018] The output unit 15 is a component for performing output from the three-dimensional structure estimation device 10, and includes, for example, output devices such as a display, a printer, a touch panel, a speaker, etc. The output unit 15 may also include an interface such as a USB, and may be configured to output data to the output device via the interface.
[0019] The memory unit 12 stores instructions of a computer program executed by the control unit 11. The memory unit 12 also stores various types of information referenced by the control unit 11. Examples of such information include a perspective image 121 and a learning model 122 used to estimate the three-dimensional structure of a target. Note that storing the learning model 122 in the memory unit 12 means that parameters defining the learning model 122 are stored in the memory unit 12.
[0020] The perspective images 121 are multiple perspective images of an object captured from different directions. The learning model 122 is a model generated by machine learning and used to estimate the three-dimensional structure of the object. The learning model 122 is, for example, a neural network such as an MLP (multilayer perceptron). The learning model 122 is trained to receive input data indicating the three-dimensional coordinates of an arbitrary point within the object and to output first output data indicating the signed distance of the arbitrary point relative to the surface of the internal structure of the object and second output data indicating a first attenuation coefficient at the arbitrary point. The perspective images 121 are used to train the learning model 122. Details of training the learning model 122 will be described later.
[0021] The control unit 11 includes an acquisition unit 111, a learning unit 112, a calculation unit 113, and an estimation unit 114. The acquisition unit 111, the learning unit 112, the calculation unit 113, and the estimation unit 114 are realized by the control unit 11 executing instructions of a computer program stored in the storage unit 12.
[0022] The acquisition unit 111 acquires the fluoroscopic image 121. For example, the acquisition unit 111 acquires the fluoroscopic image 121 from an X-ray imaging device connected to the input unit 14. The acquisition unit 111 may also acquire the fluoroscopic image 121 from another device via the communication unit 13. The acquisition unit 111 may also acquire the fluoroscopic image 121 by reading the fluoroscopic image 121 from a storage destination (which may be a storage device of the three-dimensional structure estimation device 10 or a storage device outside the three-dimensional structure estimation device 10) designated by the user of the three-dimensional structure estimation device 10.
[0023] The learning unit 112 learns a learning model 122 using a plurality of perspective images 121. The calculation unit 113 inputs input data indicating each point in the object to the learning model 122, and calculates a second attenuation coefficient at each point in the object by correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data.
[0024] The estimation unit 114 estimates the three-dimensional structure of the object based on the second attenuation coefficient at each point within the object. The estimation unit 114 may also estimate the properties or state of a material corresponding to each point within the object based on the second attenuation coefficient at each point within the object. Here, examples of the properties or state of a material include, but are not limited to, the proportion of muscle or fat, density, water content, and bone density when the object is a human.
[0025] 3 is a flow diagram showing an example of the flow of the three-dimensional structure estimation method. In step S11 (acquisition step), the acquisition unit 111 acquires a plurality of perspective images of the object captured from different directions. The number of perspective images acquired by the acquisition unit 111 is, for example, several tens of images, but is not limited to this.
[0026] 4 and 5 are diagrams showing specific examples of a method for capturing multiple fluoroscopic images. FIG. 4 is a diagram illustrating a case where multiple fluoroscopic images are acquired by a CT examination. In the example of FIG. 4, a patient lies on his / her back on a table 101. The table 101 then moves, causing the body part to be imaged to enter the scanner tube 102, and the subject is imaged. FIG. 5 is a diagram illustrating a case where imaging is performed by a robot equipped with an imaging device. In the example of FIG. 5, a robot 201 moves and images a subject who is stationary in a predetermined position and posture from multiple directions.
[0027] 3, the learning unit 112 uses a plurality of perspective images to learn the learning model 122. The learning process performed by the learning unit 112 will be described in detail later.
[0028] In step S13 (calculation step), the calculation unit 113 inputs input data indicating each point within the object to the learning model 122 obtained in step S12, and calculates the second attenuation coefficient at each point within the object by correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data. More specifically, as an example, the calculation unit 113 corrects the first attenuation coefficient so that the second attenuation coefficient is smaller when the signed distance corresponds to the outside of the internal structure than when the signed distance corresponds to the inside of the internal structure.
[0029] The signed distance f can be used to output a noise-free surface shape by explicitly using the boundary surface in the calculation, i.e., to determine whether the object is inside or outside. Since attenuation outside the internal structure may be noise (artifacts) that occurs when capturing a fluoroscopic image, reducing the attenuation coefficient of points outside the internal structure can reduce the noise.
[0030] For example, the calculation unit 113 calculates the second attenuation coefficient by multiplying the first attenuation coefficient by a function Ω(s, f) expressed by the following formula (1): In formula (1), the signed distance f is the output of the learning model 122, and the learnable parameter s is a parameter related to the gradient of the surface boundary.
[0031] In this case, the second damping coefficient μ calculated by the calculation unit 113 is expressed by the following formula (2): d is the first damping coefficient and μ is the second damping coefficient.
[0032] In step S14 (estimation step), the estimation unit 114 estimates the three-dimensional structure of the object based on the second attenuation coefficient μ at each point in the object. For example, the estimation unit 114 may generate a three-dimensional image showing the three-dimensional structure of the object based on the second attenuation coefficient μ at each point in the object.
[0033] (Example of the configuration of the learning model and an example of the training process) The details of the learning process of the learning model 122 will be described with reference to the drawings. FIG. 6 is a diagram showing an example of the configuration of the learning model 122 and an example of the training process performed by the learning unit 112. In the example of FIG. 6, the learning model 122 is an SDF model M sdf and ATT Model M att In the example of FIG. 6, the SDF model M sdf and ATT Model M att is an MLP (multilayer perceptron). The SDF model M sdf is a model that outputs a feature value r of each point in the object and generates a signed distance f. att is the first attenuation coefficient μ d This is a model that generates
[0034] (Input and Output of SDF Model) In the example of FIG. 6, the SDF model M sdf The input to is the position feature Γ(x) obtained by positional encoding of three-dimensional coordinates (x, y, z). Positional encoding is a technique for extracting low-frequency and high-frequency components from position information and converting them into a high-dimensional feature vector. In MLP, the high-frequency structure (more complex structure) of a target object tends to be more difficult to learn than the low-frequency structure, but performing positional encoding makes it easier to learn the complex structure.
[0035] Also, SDF Model M sdf The input to may include information indicating at least one of the imaging direction and imaging distance of the fluoroscopic image (the distance between the object and the detector in the case of X-rays).
[0036] SDF Model M sdf The output is a feature value r and a signed distance f. The feature value r is, for example, a high-dimensional (e.g., 257-dimensional) feature vector. The signed distance f is a signed value that represents the distance to the surface of a specific internal structure of the object. The signed distance f indicates a point outside the specific internal structure with a positive value, and a point inside the specific internal structure with a negative value.
[0037] (ATT model input / output) ATT model M att The input of the SDF model M sdf is the output feature value r. Also, the ATT model M att The output of the first damping coefficient μ d The first damping coefficient μ d is a value that estimates the degree of attenuation of light rays (e.g., X-rays) used to capture a fluoroscopic image at each point within the object. d can be said to correspond to an estimate of the density of the material at each point within the object.
[0038] (Overview of Learning Process) Next, the learning process by the learning unit 112 will be described. The learning unit 112 inputs, to the learning model 122 in the initial state or during learning, a position feature Γ(x) obtained by position-encoding the three-dimensional coordinates of each point of the target. The learning unit 112 also inputs the position feature Γ(x) to the SDF model M sdf The signed distance f and the first attenuation coefficient μ d The learning unit 112 then calculates a second attenuation coefficient μ for each point using the second attenuation coefficient μ for each point. Next, the learning unit 112 performs volume rendering using the second attenuation coefficient μ for each point to generate estimated images of perspective images from each direction. In volume rendering, when a virtual ray incident on an object from a certain position in a certain direction passes through the object, the degree to which the virtual ray is attenuated at each point on the object is calculated based on the second attenuation coefficient μ, thereby calculating the pixel value at each point on the perspective images from each direction. Furthermore, the learning unit 112 compares the generated perspective images with the perspective images 121 acquired by the acquisition unit 111 to calculate a loss function and adjust the parameters of the learning model 122.
[0039] At this time, the learning unit 112 renders a perspective image using both the signed distance f and the second attenuation coefficient μ. Therefore, if the learning model 122 is not trained to correctly output both the signed distance f and the second attenuation coefficient μ, an incorrect perspective image will be generated. By training the learning model 122 to minimize loss so that a correct perspective image is output, both the signed distance f and the second attenuation coefficient μ will be correctly output.
[0040] (Details of Correction of Damping Coefficient) As an example, the learning unit 112 corrects the function Ω(s, f) expressed by the above-mentioned formula (1) as the first damping coefficient μ d The second attenuation coefficient μ is calculated by multiplying μ by σ. FIG. 7 is a diagram showing a specific example of the function Ω(s, f). FIG. 7 shows the signed distance f and the function Ω(s, f) when the learnable parameter s has values of 100, 20, and 5. In FIG. 7, the horizontal axis represents position, and the vertical axis represents Ω(s, f) or the signed distance f. As shown in FIG. 7, the larger the value of the learnable parameter s, the steeper the gradient of the function Ω(s, f) near the boundary surface.
[0041] In this way, while increasing the learnable parameter s is expected to output a clearer surface shape, the gradient of the function Ω(s, f) disappears outside the vicinity of the surface, and the backpropagation calculation is not transmitted to the signed distance f, preventing progress in learning of the signed distance f. Therefore, the learning unit 112 sets a small value for the learnable parameter s in the early stages of learning so that the gradient calculation is also transmitted to the signed distance f, and gradually increases the value of the learnable parameter s as learning progresses so that a clearer surface shape is output. In this way, the signed distance f and the second attenuation coefficient μ can be learned in parallel.
[0042] In this way, by learning so that the learnable parameter s increases as the learning progresses, learning related to the output of the signed distance f progresses in the early stage of learning, and learning progresses so that the surface shape indicated by the second attenuation coefficient μ becomes clearer in the later stage of learning. Learning (updating of the model parameters) is performed repeatedly. For example, when 30 perspective images are input, this repetition involves a first learning based on the 30 perspective images, a second learning based on the same 30 perspective images, a third learning based on the same 30 perspective images, and so on. As this repetition increases, the value of the learnable parameter s increases.
[0043] (Pose refinement) Furthermore, the learning unit 112 may also correct the camera parameters of the image capture device along with other parameters during the learning (iterative calculation) process to calculate optimal camera parameters of the image capture device. This enables optimal learning even if the camera parameters of the image capture device are not accurate. This technique is also referred to as "pose refinement" in this specification.
[0044] In one embodiment, the training unit 112 performs position coding according to a coarse-to-fine approach to perform pose refinement. The coarse-to-fine approach is a technique in which training is initially performed using only the low-frequency components of the position-coded values, and then gradually training is performed using the high-frequency components as well. This allows optimal training even if the camera parameters of the image capture device are not accurate.
[0045] (Estimation of the Property or State of a Material) In the process of estimating the three-dimensional structure of the object, the estimation unit 114 may estimate the property or state of the object, such as bone density, using the second attenuation coefficient μ. att 1 is a diagram for explaining a process of estimating bone density from the output of . According to this embodiment, since the attenuation rate (density information) at each 3D data point can be acquired, the density information can also be used to measure bone density and muscle quality (muscle and fat). As an example, the estimation unit 114 may estimate the properties or state of an object from the density information using data obtained by statistics.
[0046] Furthermore, the three-dimensional structure estimation device 10 according to this embodiment can be applied to fields other than the medical field (such as human body measurements). For example, it can be applied to fields such as measuring plants. When measuring plants, as with human body measurements, a 3D image can be generated from a 2D image, and the plant surface can be accurately recognized. For example, clubroot disease, which occurs in broccoli and other plants, causes the roots to swell into nodules. Conventionally, this disease could only be detected by actually removing the target object from the soil. In contrast, this embodiment allows measurements to be performed even when the plant is planted in soil. This can reduce food waste, for example. Furthermore, this embodiment can acquire density information, allowing differences in plant conditions (such as differences in water content) to be measured.
[0047] (Configuration example and processing example of learning model when multiple internal structures are included) Fig. 9 is a diagram showing an example of the configuration of the learning model 122 and the training processing performed by the learning unit 112 when multiple internal structures of the target are considered. In the example of Fig. 9, the learning model 122 is an SDF model M sdf and multiple ATT models M att1 , M att2 and SDF Model M sdf is the signed distance f between the feature r and the surface of each of multiple internal structures (e.g., bones and muscles). 0 , f 1 ATT model M att0 is the first attenuation coefficient μ d0 is a model that generates the ATT model M att1 is the first attenuation coefficient μ d1 This is a model that generates
[0048] The SDF model M of the learning model 122 sdf From the above, multiple signed distances f 0 , f 1 In other words, when the object has multiple internal structures, the first output data output by the learning model 122 indicates the signed distance of an arbitrary point to the surface of each of the multiple internal structures. sdfFrom the above, a plurality of first damping coefficients μ d0 , μ d1 (For example, the first attenuation coefficient μ corresponding to bone d0 and the first damping coefficient μ corresponding to the muscle d1 In other words, the second output data output by the learning model 122 indicates the first attenuation coefficients corresponding to the plurality of internal structures at an arbitrary point.
[0049] First damping coefficient μ d0 , μ d1 The respective ranges of are set in advance by, for example, a user. When a plurality of signed distances are used, the first attenuation coefficients μ corresponding to the respective signed distances are set. d0 , μ d1 The range is preset based on the knowledge that bone has a lower damping coefficient than muscle, for example.
[0050] The calculation unit 113 calculates, for each point in the object, a signed distance f 0 , f 1 and determining the internal structure to which the point belongs based on the first attenuation coefficient μ d0 , μ d1 Based on this, the second damping coefficient μ 0 , μ 1 Calculate.
[0051] The calculation unit 113 further calculates the second damping coefficient μ 0 , μ 1 and the signed distance f 0 , f 1 At this time, the calculation unit 113 calculates the second attenuation coefficient μ corresponding to each point based on the above. 0 and a second damping coefficient μ corresponding to the muscle. 1 After obtaining the above, the calculation unit 113 calculates the second attenuation coefficient μ corresponding to each point. Alternatively, the calculation unit 113 may calculate the second attenuation coefficient μ after determining whether each point corresponds to a muscle or a bone.
[0052] In the example of FIG. 9, in order to separate bones and muscles, the ATT model M corresponding to the bones is used. att0 The first damping coefficient μ d0The range of values of and the ATT model M corresponding to the muscle att1 The first damping coefficient μ d1 For example, the ATT model M att0 If the activation function of the last layer of the model is ασ(x)+β (σ(x) is a sigmoid function), the range of the output value is (β, α+β). att0 , ATT Model M att1 For example, different α and β are set for β 1 = 0.1, α 1 = 3.9, β 0 = 4, α 0 = 4, then ATT model M att1 The output range of the ATT model is (0.1, 4). Most muscles have damping coefficients within this range. att1 It can be said that the reliability of the ATT model M is high. att0 The output range of is in the range of 4 to 8, which is a model for bone with a larger damping coefficient.
[0053] Effect of the embodiment As described above, according to the present embodiment, the three-dimensional structure estimation device 10 generates the learning model 122 by machine learning using a plurality of perspective images of the object captured from different directions, and calculates, for each point in the object, the signed distance f to the surface of the object's internal structure (bones, muscles, etc.) and the first attenuation coefficient μ d and the signed distance f is used to estimate the first attenuation coefficient μ d The three-dimensional structure of the target is estimated based on the correction results.
[0054] Here, the first damping coefficient μ d When the three-dimensional structure of the object is estimated using the first attenuation coefficient μ as it is without correction, the estimated boundary surface may contain many inappropriate parts. d The three-dimensional structure of the object can be estimated more accurately by using the second attenuation coefficient μ obtained by correcting the above based on the signed distance f, which is the output of the learning model 122.
[0055] Furthermore, according to the three-dimensional structure estimation device 10, highly accurate learning can be performed even if the camera parameters are not accurate, because the learning unit 112 performs pose refinement. Furthermore, according to the three-dimensional structure estimation device 10, when the imaging direction and imaging distance are known to a certain extent in advance, highly accurate learning can be performed even if the imaging direction and imaging distance are not accurate.
[0056] [Examples] Several examples of the present invention will be described below. In an experiment (Example 1) using a simulation model, 36 fluoroscopic images were generated from 3D models of the patella and skull, with the direction changed by 10 degrees. The size of each fluoroscopic image was 480 x 480. The distance from the radiation source to the subject was 40 cm. The distance from the subject to the detection panel was 10 cm.
[0057] The pixel values were normalized to the range of 0 to 1, with 1 representing background (no attenuation). In each learning session, 512 virtual rays were sampled, and 128 points were sampled for each virtual ray. The learning model 122 used had the structure shown in FIG. 6. M SDF A six-layer MLP was used as the model. The hidden size of each layer was set to 256. ATT We used an MLP with three hidden layers, each with a size of 256.
[0058] The accuracy of the obtained 3D structure in estimating the original 3D models of the patella and skull was evaluated using PSNR (Peak Signal to Noise Ratio), SSIM (Structural Similarity), LPIPS (Learned Perceptual Image Patch Similarity), and CD (Chamfer Distance). Furthermore, as comparative examples, evaluations were also performed on cases where the 3D structures were estimated using NAF (neural attenuation fields) (Comparative Example 1) described in Non-Patent Document 1, NeAT (Neural adaptive tomography) (Comparative Example 2) described in Non-Patent Document 2, NeRF (Comparative Example 3) described in Non-Patent Document 3, and NeuS (neural implicit surfaces) (Comparative Example 4) described in Non-Patent Document 4, based on the same fluoroscopic images.
[0059] In the actual experiments (Examples 2 to 4), 30 fluoroscopic images were taken of the subject's feet and head, rotating the direction by 12 degrees each time using a computer-controlled rotating panel. The distance from the radiation source to the detection panel was approximately 1 m. The distance from the radiation source to the subject was approximately 77 cm. The voltage of the X-ray source was 60 kV, and the current was 200 mA.
[0060] In Example 2, a learning model with a structure outputting one type of signed distance as shown in FIG. 6 was used, and pose refinement was not performed. In Example 3, a learning model with a structure outputting one type of signed distance as shown in FIG. 6 was used, and pose refinement was performed. In Example 4, a learning model with a structure outputting two types of signed distance as shown in FIG. 9 was used, and pose refinement was performed. Other conditions were the same as in Example 1. In pose refinement, position encoding was initially performed using only two low-frequency components, and then high-frequency components were gradually included, and position encoding was performed so that the highest frequency component was also included by the time half of the learning (iterative calculation) was performed. The initial value of the learnable parameter s was 20.
[0061] Table 1 shows the experimental results for the patella in Example 1, and Table 2 shows the experimental results for the skull. As shown in Tables 1 and 2, the estimation results obtained by the method according to this embodiment have higher estimation accuracy than the estimation results obtained by other algorithms.
[0062]
[0063]
[0064] Table 3 shows the measured values obtained from actual X-ray images in Examples 2 to 4. As shown in Table 3, estimation was possible with high accuracy from actual X-ray images, and accuracy improved by performing pose refinement, and accuracy was further improved by configuring the learning model to output multiple signed distances.
[0065] [Additional Notes] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
[0066] [Summary] A three-dimensional structure estimation device according to a first aspect of the present invention is configured to include: an acquisition unit that acquires multiple perspective images of an object taken from different directions; a learning unit that uses the multiple perspective images to train a learning model that receives input data indicating the three-dimensional coordinates of an arbitrary point within the object and outputs first output data indicating a signed distance of the arbitrary point relative to the surface of an internal structure of the object and second output data indicating a first attenuation coefficient at the arbitrary point; a calculation unit that inputs input data indicating each point within the object to the learning model and calculates a second attenuation coefficient at each point within the object by correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data; and an estimation unit that estimates the three-dimensional structure of the object based on the second attenuation coefficient at each point within the object.
[0067] According to the above configuration, the three-dimensional structure of the object can be estimated more accurately than when the first attenuation coefficient is not corrected based on the signed distance.
[0068] Furthermore, in the three-dimensional structure estimation device according to the second aspect of the present invention, in addition to the configuration of the three-dimensional structure estimation device according to the first aspect described above, the calculation unit is configured to correct the first damping coefficient so that the second damping coefficient is smaller when the signed distance corresponds to the outside of the internal structure than when the signed distance corresponds to the inside of the internal structure.
[0069] According to the above configuration, when the signed distance corresponds to the outside of the internal structure, the three-dimensional structure of the object can be estimated more accurately than when the first attenuation coefficient is not corrected so that the second attenuation coefficient is reduced.
[0070] Furthermore, in a three-dimensional structure estimation device according to a third aspect of the present invention, in addition to the configuration of the three-dimensional structure estimation device according to the first or second aspect described above, when the object has a plurality of internal structures, the first output data indicates the signed distance of the arbitrary point to the surface of each of the plurality of internal structures, the second output data indicates the first attenuation coefficient corresponding to each of the plurality of internal structures at the arbitrary point, and the calculation unit determines, for each of the points, the internal structure to which the point belongs based on the signed distance to the surface of each of the plurality of internal structures, and calculates the second attenuation coefficient based on the first attenuation coefficient corresponding to the internal structure to which the point belongs.
[0071] According to the above configuration, it is possible to more accurately estimate the three-dimensional structure of an object having multiple internal structures.
[0072] Furthermore, in a three-dimensional structure estimation device according to a fourth aspect of the present invention, in addition to the configuration of the three-dimensional structure estimation device according to any one of the first to third aspects described above, a configuration is adopted in which the estimation unit estimates the properties or state of a substance corresponding to each point within the object based on the second attenuation coefficient at each point within the object.
[0073] According to the above configuration, it is possible to estimate the properties or state of a substance within an object.
[0074] In addition, a three-dimensional structure estimation method according to a fifth aspect of the present invention includes an acquisition step of acquiring a plurality of perspective images of an object taken from different directions; a learning step of using the plurality of perspective images to train a learning model in which input data indicating the three-dimensional coordinates of an arbitrary point within the object is input, and which outputs first output data indicating a signed distance of the arbitrary point relative to the surface of an internal structure of the object and second output data indicating a first attenuation coefficient at the arbitrary point; a calculation step of inputting input data indicating each point within the object to the learning model, and correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data, thereby calculating a second attenuation coefficient at each point within the object; and an estimation step of estimating the three-dimensional structure of the object based on the second attenuation coefficient at each point within the object.
[0075] According to the above configuration, the three-dimensional structure of the object can be estimated more accurately than when the first attenuation coefficient is not corrected based on the signed distance.
[0076] In addition, a program according to a sixth aspect of the present invention is a program for causing a computer to function as a three-dimensional structure estimation device, and causes the computer to function as: an acquisition unit that acquires multiple perspective images of an object taken from different directions; a learning unit that uses the multiple perspective images to learn a learning model that receives input data indicating the three-dimensional coordinates of an arbitrary point within the object and outputs first output data indicating a signed distance of the arbitrary point relative to the surface of the internal structure of the object and second output data indicating a first attenuation coefficient at the arbitrary point; a calculation unit that inputs input data indicating each point within the object to the learning model and calculates a second attenuation coefficient at each point within the object by correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data; and an estimation unit that estimates the three-dimensional structure of the object based on the second attenuation coefficient at each point within the object.
[0077] According to the above configuration, the three-dimensional structure of the object can be estimated more accurately than when the first attenuation coefficient is not corrected based on the signed distance.
[0078] Furthermore, a recording medium according to a seventh aspect of the present invention is a program for causing a computer to function as a three-dimensional structure estimation device according to any one of the first to fourth aspects described above, and is a non-transitory computer-readable recording medium having recorded thereon a program for causing the computer to function as the acquisition unit, the learning unit, the calculation unit, and the estimation unit.
[0079] [Example of implementation by software] The functions of the three-dimensional structure estimation device 10 (hereinafter referred to as the "device") can be realized by a program that causes a computer to function as the device, and a program that causes a computer to function as each control block of the device (particularly each part included in the control unit 11).
[0080] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. The functions described in each of the above embodiments are realized by executing the program using the control device and storage device.
[0081] The program may be non-transitory and may be recorded on one or more computer-readable recording media. The recording media may or may not be included in the device. In the latter case, the program may be supplied to the device via any wired or wireless transmission medium.
[0082] Furthermore, some or all of the functions of the control blocks can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the control blocks is formed is also included in the scope of the present invention. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.
[0083] Furthermore, each process described in each of the above embodiments may be executed by AI (Artificial Intelligence). In this case, the AI may run on the control device or on another device (for example, an edge computer or a cloud server).
[0084] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
[0085] REFERENCE SIGNS LIST 10 Three-dimensional structure estimation device 11 Control unit 12 Memory unit 13 Communication unit 14 Input unit 15 Output unit 111 Acquisition unit 112 Learning unit 113 Calculation unit 114 Estimation unit 121 Fluoroscopic image 122 Learning model 201 Robot
Claims
1. A three-dimensional structure estimation device comprising: an acquisition unit that acquires multiple perspective images of an object taken from different directions; a learning unit that uses the multiple perspective images to train a learning model that receives input data indicating the three-dimensional coordinates of an arbitrary point within the object and outputs first output data indicating a signed distance of the arbitrary point relative to the surface of an internal structure of the object and second output data indicating a first attenuation coefficient at the arbitrary point; a calculation unit that inputs input data indicating each point within the object to the learning model and calculates a second attenuation coefficient at each point within the object by correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data; and an estimation unit that estimates the three-dimensional structure of the object based on the second attenuation coefficient at each point within the object.
2. The three-dimensional structure estimation device of claim 1, wherein the calculation unit corrects the first attenuation coefficient so that the second attenuation coefficient is reduced when the signed distance corresponds to the outside of the internal structure more than when the signed distance corresponds to the inside of the internal structure.
3. A three-dimensional structure estimation device as described in claim 1, wherein, when the object has multiple internal structures, the first output data indicates a signed distance of the arbitrary point to the surface of each of the multiple internal structures, the second output data indicates the first attenuation coefficient corresponding to each of the multiple internal structures at the arbitrary point, and the calculation unit determines, for each of the points, the internal structure to which the point belongs based on the signed distance to the surface of each of the multiple internal structures, and calculates the second attenuation coefficient based on the first attenuation coefficient corresponding to the internal structure to which the point belongs.
4. A three-dimensional structure estimation device according to claim 1, wherein the estimation unit estimates the properties or state of a material corresponding to each point within the object based on the second attenuation coefficient at each point within the object.
5. A three-dimensional structure estimation method comprising: an acquisition step of acquiring a plurality of perspective images of an object taken from different directions; a learning step of using the plurality of perspective images to train a learning model which receives input data indicating the three-dimensional coordinates of an arbitrary point within the object and outputs first output data indicating a signed distance of the arbitrary point relative to the surface of an internal structure of the object and second output data indicating a first attenuation coefficient at the arbitrary point; a calculation step of inputting input data indicating each point within the object to the learning model and calculating a second attenuation coefficient at each point within the object by correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data; and an estimation step of estimating the three-dimensional structure of the object based on the second attenuation coefficient at each point within the object.
6. A program for causing a computer to function as a three-dimensional structure estimation device, the program causing the computer to function as: an acquisition unit that acquires multiple perspective images of an object taken from different directions; a learning unit that uses the multiple perspective images to learn a learning model that receives input data indicating the three-dimensional coordinates of an arbitrary point within the object and outputs first output data indicating the signed distance of the arbitrary point relative to the surface of the internal structure of the object and second output data indicating a first attenuation coefficient at the arbitrary point; a calculation unit that receives input data indicating each point within the object to the learning model and calculates a second attenuation coefficient at each point within the object by correcting the first attenuation coefficient indicated by the second output data based on the signed distance indicated by the first output data; and an estimation unit that estimates the three-dimensional structure of the object based on the second attenuation coefficient at each point within the object.
7. A program for causing a computer to function as the three-dimensional structure estimation device described in claim 1, wherein a non-transitory computer-readable recording medium is recorded with a program for causing the computer to function as the acquisition unit, the learning unit, the calculation unit, and the estimation unit.
Citation Information
Patent Citations
Everyday Scene Restoration Engine
JP2019518268A
Visual odometry method, apparatus and computer-readable recording medium
JP2020013562A
Imaging apparatus, imaging method, and imaging program
JP2022024975A
Attenuation correction method for PET / MR using level set method
KR1020140139681A
Completion of Truncated Attenuation Maps Using MLAA
US20110103669A1