A three-dimensional reconstruction method of two-dimensional ultrasound images based on deep learning
By combining deep learning 3DCNN and LSTM models, the inter-frame features and spatial pose information of two-dimensional ultrasound image sequences are extracted, solving the problem that traditional two-dimensional ultrasound imaging systems have difficulty in automatically reconstructing three-dimensional images, and achieving high-accuracy three-dimensional reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING JIAOTONG UNIV
- Filing Date
- 2023-03-20
- Publication Date
- 2026-05-08
AI Technical Summary
Traditional two-dimensional ultrasound imaging systems are difficult to automate the reconstruction of three-dimensional images, requiring doctors to have extensive experience and spatial imagination. Existing three-dimensional reconstruction methods suffer from problems such as bulky mechanical devices, large cumulative errors, and high frame rate requirements.
By combining a deep learning-based 3D convolutional neural network (3DCNN) and a long short-term memory model (LSTM), 3D reconstruction is performed by extracting inter-frame features and spatial pose information from 2D ultrasound image sequences, thereby eliminating accumulated errors and improving accuracy.
It enables automated 3D ultrasound image reconstruction, improves reconstruction accuracy, reduces reliance on physician experience, and simplifies the operation process.
Smart Images

Figure CN116310032B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image processing technology, and specifically relates to a method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning. Background Technology
[0002] Ultrasound medicine has significant advantages over other technologies. Being non-invasive, harmless, painless, and intuitive, it is an indispensable imaging diagnostic method in clinical medicine and has been widely used. Traditional B-mode ultrasound imaging systems can only provide two-dimensional sequential images of the scanned object. Doctors must reconstruct these tomographic images into three-dimensional objects using their brains, requiring considerable experience and spatial imagination.
[0003] Three-dimensional ultrasound imaging has the advantages of intuitive image display, accurate measurement of medical diagnostic parameters, and wide application in medical teaching and surgical planning. Currently, there are four main methods for acquiring medical ultrasound three-dimensional images: 1) using expensive dedicated two-dimensional array ultrasound probes, but this method has a relatively narrow imaging field of view; 2) using mechanically driven one-dimensional array ultrasound probes to scan along a predetermined trajectory and acquire a fixed number of two-dimensional ultrasound images for three-dimensional reconstruction, but this method often involves bulky mechanical devices; 3) using one-dimensional array ultrasound probes to acquire two-dimensional ultrasound images, adding space and angle sensors to acquire the pose information corresponding to the two-dimensional ultrasound images, and then performing three-dimensional reconstruction, but this method is prone to cumulative errors; 4) using one-dimensional array ultrasound probes without any additional auxiliary positioning devices, estimating the relative spatial position between adjacent images based solely on the information of the ultrasound images themselves for three-dimensional reconstruction, but this method requires a high frame rate for the two-dimensional image sequence and requires a large amount of training and calibration. Summary of the Invention
[0004] To address the aforementioned problems, this invention provides a novel method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning.
[0005] The specific technical solution of this invention is as follows:
[0006] This invention provides a method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning. The three-dimensional reconstruction method includes the following steps:
[0007] S1: Extract the region of interest from the two-dimensional ultrasound image sequence;
[0008] S2: By stacking multiple consecutive frames, the region of interest is processed into three-dimensional data and used as the input to the trained three-dimensional convolutional neural network;
[0009] S3: A three-dimensional convolutional neural network extracts inter-frame features from a two-dimensional image sequence and calculates spatial pose information using different three-dimensional convolutional kernels;
[0010] S4: Extract the dependencies of spatial pose information sequences through the trained long short-term memory model, perform statistics and prediction on the output of the three-dimensional convolutional neural network, output the pose information of the two-dimensional ultrasound image sequence in space, and perform three-dimensional reconstruction to generate a three-dimensional reconstructed image.
[0011] The beneficial effects achieved by this invention are as follows:
[0012] This invention provides a novel method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning. This method uses a high-frame-rate two-dimensional ultrasound image sequence as input data for a 3DCNN-LSTM network, then uses 3DCNN to extract spatial pose information from the inter-frame features of the ultrasound images, and uses LSTM to perform statistical analysis and prediction on the spatial pose information sequence. This effectively utilizes the inter-frame features of the ultrasound images and the spatiotemporal correlation of the pose information sequence, eliminates accumulated errors, and improves the accuracy of three-dimensional reconstruction. Attached Figure Description
[0013] Figure 1 This is a flowchart of the three-dimensional reconstruction method for two-dimensional ultrasound images based on deep learning in this invention;
[0014] Figure 2 This is a flowchart of step S1 in the present invention;
[0015] Figure 3 This is a schematic diagram illustrating interpolation and fitting using the Bézier curve function within the same control window in this invention.
[0016] Figure 4 This is a schematic diagram of the 3DCNN network in this invention;
[0017] Figure 5 This is a schematic diagram of the 3DCNN-LSTM network in this invention;
[0018] Figure 6 This is a flowchart of the 3DCNN program in this invention;
[0019] Figure 7 This is a flowchart of step S3 in the present invention;
[0020] Figure 8 This is a flowchart of step S34 in the present invention;
[0021] Figure 9 This is a structural diagram of the LSTM in this invention. Detailed Implementation
[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments. The following embodiments are only used to explain the invention and are not intended to limit the scope of protection of the present invention.
[0023] This invention provides a method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning, such as... Figure 1 The three-dimensional reconstruction method includes the following steps:
[0024] S1: Extract the region of interest from the two-dimensional ultrasound image sequence;
[0025] S2: By stacking multiple consecutive frames, the region of interest is processed into three-dimensional data and used as the input to the trained three-dimensional convolutional neural network;
[0026] S3: Three-dimensional convolutional neural networks (i.e., 3DCNN) extract inter-frame features from two-dimensional image sequences and calculate spatial pose information through different three-dimensional convolutional kernels; 3D convolutional kernels can not only extract image features around the voxel to be predicted, but also extract spatial features between image frames;
[0027] S4: Extract the dependencies of spatial pose information sequences through the trained Long Short-Term Memory (LSTM) model, perform statistics and prediction on the output of the three-dimensional convolutional neural network, output the pose information of the two-dimensional ultrasound image sequence in space, and perform three-dimensional reconstruction to generate a three-dimensional reconstructed image.
[0028] This invention provides a novel method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning. This method uses a high-frame-rate two-dimensional ultrasound image sequence as input data for a 3DCNN-LSTM network, then uses 3DCNN to extract spatial pose information from the inter-frame features of the ultrasound images, and uses LSTM to perform statistical analysis and prediction on the spatial pose information sequence. This effectively utilizes the inter-frame features of the ultrasound images and the spatiotemporal correlation of the pose information sequence, eliminates accumulated errors, and improves the accuracy of three-dimensional reconstruction.
[0029] like Figure 2 As shown, step S1 in this embodiment includes the following steps:
[0030] S11: Acquire low-frame-rate two-dimensional ultrasound images obtained by a one-dimensional array ultrasound probe, and pose information corresponding to the two-dimensional ultrasound images obtained by an acousto-optic positioning system.
[0031] S12: Interpolate and fit the image sequence using the Bézier curve function, and add points to add two-dimensional ultrasound images to construct a high frame rate two-dimensional ultrasound image sequence. The expression of the Bézier curve function is as follows:
[0032]
[0033] The Bézier curve B(t) has a total of n+1 control points, namely P0, P1, ..., Pn. n t is a curve function. For example, when n = 1 and t = 0.5, the value of B(t) is exactly in the middle of the curve from P0 to P1.
[0034] S13: Extract the region of interest (ROI) for each frame in a high frame rate two-dimensional ultrasound image sequence.
[0035] In this embodiment, step S12 inserts 4 frames between every two frames of the low frame rate two-dimensional ultrasound image sequence to obtain a high frame rate two-dimensional ultrasound image sequence.
[0036] In this embodiment, a one-dimensional array ultrasound probe is used to acquire low-frame-rate two-dimensional ultrasound images. An additional acoustic-optical localization system is used to acquire the pose information corresponding to the two-dimensional ultrasound images. The image sequence is then interpolated and fitted using a Bézier curve function, and two-dimensional ultrasound images are added by interpolation points to construct a high-frame-rate two-dimensional ultrasound image sequence. The region of interest (ROI) for each frame in the image sequence is defined and extracted, and used as input data for the 3DCNN-LSTM network.
[0037] Figure 3 This is a schematic diagram illustrating interpolation and fitting using Bézier curve functions within the same control window of the present invention. In this embodiment, as shown... Figure 3 As shown, the entire 3D reconstruction process begins with the first four frames of images. Each time the control window is moved, the Bézier curve within each control window is calculated. If the last window contains fewer than four frames, a lower-order Bézier curve is used. P11, P21, P31, and M1 on the first frame represent the regions of interest (ROIs) 1, 2, and 3, respectively, and the center point of the current frame image. P11, P12, P13, and P14 represent the regions of interest ROIs 1 on four adjacent frames. The corresponding ROIs 1, ROI2, and ROI3 on four consecutive low-frame-rate 2D ultrasound images acquired using a one-dimensional array ultrasound probe are calculated, resulting in three Bézier curves B1, B2, and B3.
[0038] Between points P11 and P12 on Bézier curve B1, four insertion points 11, 12, 13, and 14 are selected along the direction of movement of the control window. Similarly, insertion points 21, 22, 23, 24 and 31, 32, 33, and 34 on the other two Bézier curves B2 and B3 can be obtained. Next, based on the principle that three points in space that are not on the same straight line can determine a plane, the normal vector expression of plane I1 and the coordinates of the midpoint ml of I1 can be obtained from 11, 21, and 31. 11, 21, and 31 are located at ROI1, ROI2, and ROI3 of I1, respectively. Similarly, the normal vector expressions of the other three planes I2, I3, and I4, and the coordinates of the midpoints m2, m3, and m4 can be obtained. Two-dimensional ultrasound images at these insertion points are acquired using a one-dimensional array ultrasound probe. Four frames are inserted between every two frames of the original low-frame-rate ultrasound image sequence to obtain a high-frame-rate two-dimensional ultrasound image sequence. Define and extract the regions of interest (ROI1, ROI2, and ROI3) for each frame in the image sequence, and use them as input data for the 3DCNN-LSTM network.
[0039] In this embodiment, in step S2, four insertion frames are inserted between every two consecutive frames. After ROI extraction processing is performed on each consecutive frame and the insertion frames, three-dimensional data is obtained. The cube composed of ROIs of multiple ultrasound images is used as the input of a three-dimensional convolutional neural network.
[0040] like Figure 7 As shown, step S3 in this embodiment includes the following steps:
[0041] S31: The three-dimensional convolutional neural network replicates the cube 6 times, with each set of 3 copies forming a feature group;
[0042] S32: Perform 3D convolution on the two-dimensional image sequence input from the two feature groups and extract deep inter-frame features;
[0043] S33: Max pooling is performed on each convolutional layer to reduce the size of the feature maps;
[0044] S34: Spatial pose information is calculated through a fully connected layer;
[0045] like Figure 8 As shown, the method for calculating spatial pose information in step S34 of this embodiment is as follows:
[0046] S341: Calculate the speckle decorrelation of three sets of regions of interest at different distances between adjacent frames, measure multiple times, and plot the distance-decorrelation calibration curve by taking the average value. The expression for calculating the speckle decorrelation is as follows:
[0047]
[0048] Where coV(X,Y) is the covariance of ROI regions X and Y corresponding to adjacent frames, and σX and σY are the standard deviations of ROI regions X and Y.
[0049] S342: The frame image moves along the X, Y, and Z directions at a fixed step size to obtain three distance-decorrelation calibration curves in the three directions;
[0050] S343: Calculate the speckle correlation of three groups of regions of interest in the current two adjacent frames;
[0051] S344: The distance between the current adjacent images is obtained from the calibration curve by looking up a table. Based on the principle of determining the unique plane in space by three non-collinear points in space, the spatial position and angle relationship of the current image relative to the previous frame image are calculated. That is, the pose information of the next frame image is calculated based on the three distances and the pose information of the previous frame image.
[0052] Figure 4 This is a schematic diagram of the 3DCNN network of the present invention. Three original consecutive frames and four inserted frames between every two consecutive frames (a total of three consecutive frames and eight inserted frames) are processed by ROI extraction to obtain three-dimensional data. The ROI size is set to 256×256. A cube with a size of 256×256×11, composed of the ROIs of 11 ultrasound images, is used as input to the 3DCNN for convolution operations.
[0053] Figure 6 This is a flowchart of the 3DCNN program of the present invention. In the 3DCNN network, a cube with a size of 256×256×11 is copied 6 times, with each set of 3 copies forming a feature group. The first feature group is responsible for extracting speckle uncorrelation features in the X, Y, and Z directions between adjacent ultrasound image frames. Since two ultrasound images are needed each time, the cube size of the first feature group is 256×256×10. The second feature group is responsible for extracting grayscale, X-direction gradient, Y-direction gradient, and other features of a single ultrasound image. Since only one image is needed each time, the cube size of the second feature group is 256×256×11. In the 3DCNN, the number of convolutional kernels in the three layers is set to 32, 64, and 128, respectively, with sizes of 9×9×3, 9×9×5, and 9×9×3. The 3D convolutional kernels can not only extract image features around the voxel to be predicted but also extract spatial features between image frames. 3D convolution is performed on the two-dimensional image sequences input to the two feature groups to extract deep inter-frame features. Each convolutional layer undergoes max pooling to reduce the feature map size to 82×82×51, 24×24×27, and 5×5×15, respectively. The spatial pose information is then calculated through a fully connected layer.
[0054] Inter-frame speckle decorrelation is achieved through the speckle noise decorrelation phenomenon and can be used to determine the spatial distance between adjacent two-dimensional ultrasound images. If two adjacent frames were acquired at the same location, their speckle patterns are identical. If the two frames undergo relative motion, their speckle decorrelation is proportional to the relative distance traveled. Speckle decorrelation involves two basic steps: calibration and distance estimation. The calibration process requires multiple measurements of adjacent frames at different distances, calculating the speckle decorrelation for three regions of interest (ROIs), namely ROI1, ROI2, and ROI3, and averaging the results to plot a "distance-decorrelation" curve.
[0055] In the process of plotting the "distance-decorrelation" calibration curve, it is necessary to move along the three directions of X, Y, and Z with a fixed step size, acquire and save two-dimensional ultrasound images, and finally obtain three "distance-decorrelation" calibration curves in the three directions.
[0056] Distance estimation involves calculating the speckle correlation of three regions of interest (ROIs) – ROI1, ROI2, and ROI3 – between two adjacent ultrasound frames, and then obtaining the distances between the current and adjacent images from the calibration curve using a lookup table. Given the distance estimates of an ultrasound image at three non-collinear positions relative to the previous frame, and based on the principle that three non-collinear points determine a unique plane in space, the spatial position and angular relationship of the current image relative to the previous frame can be calculated. In other words, based on the three distances and the pose information of the previous frame, the pose information of the next frame can be calculated.
[0057] In this embodiment, the training of the three-dimensional convolutional neural network in step S2 and the training of the long short-term memory model in step S4 include:
[0058] Initialize the learning rate, and compare the spatial pose information extracted by the 3D convolutional neural network with the pose information labels to obtain the mean squared error of the pose information loss, MSE. The expression for MSE is as follows:
[0059]
[0060] Where n is the number of samples, Ri is the true value, and Ei is the predicted value;
[0061] The mean squared error of pose information loss (MSE) and the feature sequence output by the 3D convolutional neural network are fed back into the long short-term memory model. The long short-term memory model continuously updates its parameters to reduce the MSE loss of the output pose information.
[0062] Figure 4 This is a schematic diagram of the 3DCNN network of the present invention. Figure 6This is a flowchart of a 3DCNN program. During model training and optimization, the learning rate is first initialized. The spatial pose information extracted by the 3DCNN is compared with the pose information labels to obtain the pose information loss MSE. The mean squared error (MSE) reflects the degree of difference between the predicted value and the true value. The pose information loss MSE and the feature sequence output by the 3DCNN are fed back into the LSTM. By continuously updating the parameters, the pose information loss MSE of the output is minimized.
[0063] In this embodiment, the long short-term memory model in step S4 includes a forget gate, an input gate, and an output gate. The forget gate is used to determine the discarded information in the model state at the previous moment and update the model state, using the model state at the previous moment as a parameter to update the current state.
[0064] In this embodiment, step S4 uses a multi-layer long short-term memory model combined with the features of a predetermined time to form new time series data, with each layer of LSTM having 77 neurons.
[0065] Figure 5 This is a schematic diagram of the 3DCNN-LSTM network of the present invention. Figure 9 This is a structural diagram of the LSTM of this invention. 3DCNN-LSTM comprises two independent parts: 3DCNN and LSTM. The feature sequence output of the 3DCNN is used as the input to the LSTM model. The 3DCNN is responsible for feature extraction, while the LSTM is responsible for extracting the dependencies in the spatial pose information sequence and performing statistical analysis and prediction on the 3DCNN output. An LSTM unit consists of a forget gate, an input gate, and an output gate. The forget gate is used to determine the discarded information in the unit state at the previous time step and update the unit state, using the previous time step's unit state as the parameter for updating the current state. New time series data is formed by combining multiple layers of LSTM with features from specific times. Each LSTM layer has 77 neurons, and the number of LSTM layers can be adjusted according to actual needs to improve the model's predictive ability.
[0066] In this embodiment, the three-dimensional reconstruction in step S4 is divided into pixel mapping and gap filling. During pixel mapping, each pixel of each frame of two-dimensional ultrasound image is transformed into the three-dimensional reconstruction coordinate system according to the spatial pose information of the current frame through coordinate transformation.
[0067] In this embodiment, 3D reconstruction is performed based on the spatial pose information of the output two-dimensional ultrasound image sequence. First, a 3D reconstruction coordinate system must be established; for example, the center point of the first frame of the ultrasound image is used as the origin of the 3D reconstruction coordinate system. The horizontal direction of the first frame is used as the X-axis, the vertical direction as the Y-axis, and the scanning direction as the Z-axis. 3D reconstruction mainly consists of two parts: pixel mapping and gap filling. During pixel mapping, each pixel of each frame of the two-dimensional ultrasound image is transformed into the 3D reconstruction coordinate system according to the spatial pose information of the current frame through coordinate transformation. During pixel mapping, there will still be a large number of pixels without corresponding values in the 3D reconstruction coordinate system, which need to be compensated for by gap filling. For example, the voxel values for filling can be calculated using a distance-weighted interpolation algorithm and then inserted.
[0068] Specific implementations of the subject matter have been described. Other implementations are within the scope of the following claims. For example, the activities described in the claims can be performed in a different order and still achieve the desired result. As an example, the processes described in the drawings do not necessarily require a specific order or sequence to be shown in order to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.
Claims
1. A method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning, characterized in that, The three-dimensional reconstruction method includes the following steps: S1: Extract the region of interest from the two-dimensional ultrasound image sequence; S2: By stacking multiple consecutive frames, the region of interest is processed into three-dimensional data and used as the input to the trained three-dimensional convolutional neural network; S3: A three-dimensional convolutional neural network extracts inter-frame features from a two-dimensional image sequence and calculates spatial pose information using different three-dimensional convolutional kernels; S4: Extract the dependencies of spatial pose information sequences through the trained long short-term memory model, perform statistics and prediction on the output of the three-dimensional convolutional neural network, output the pose information of the two-dimensional ultrasound image sequence in space, and perform three-dimensional reconstruction to generate a three-dimensional reconstructed image. Step S3 includes the following steps: S31: The 3D convolutional neural network replicates the cube 6 times, with each set of 3 copies forming a feature group; S32: The 3D convolution is performed on the 2D image sequence input from the two feature groups, and deep inter-frame features are extracted. S33: Max pooling is performed on each convolutional layer to reduce the size of the feature maps; S34: Spatial pose information is calculated through a fully connected layer; The method for calculating spatial pose information in step S34 is as follows: S341: Calculate the speckle decorrelation of three sets of regions of interest at different distances between adjacent frames, measure multiple times, and plot the distance-decorrelation calibration curve by taking the average value. The expression for calculating the speckle decorrelation is as follows: Where cov(X,Y) is the covariance of the ROI regions X and Y corresponding to adjacent frames, σ X and σ Y Yes, the standard deviations of the ROI regions X and Y; S342: The frame image moves along the X, Y, and Z directions with a fixed step size to obtain three distance-decorrelation calibration curves in the three directions; S343: Calculate the speckle decorrelation of three sets of regions of interest in the current two adjacent frame images; S344: The distance between the current adjacent images is obtained from the calibration curve by looking up a table. Based on the principle of determining the unique plane in space by three non-collinear points in space, the spatial position and angle relationship of the current image relative to the previous frame image are calculated. That is, the pose information of the next frame image is calculated based on the three distances and the pose information of the previous frame image.
2. The method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning as described in claim 1, characterized in that, Step S1 includes the following steps: S11: Acquire low-frame-rate two-dimensional ultrasound images obtained by a one-dimensional array ultrasound probe, and pose information corresponding to the two-dimensional ultrasound images obtained by an acousto-optic positioning system. S12: Interpolate and fit the image sequence using the Bézier curve function, and add points to add two-dimensional ultrasound images to construct a high frame rate two-dimensional ultrasound image sequence. The expression of the Bézier curve function is as follows: The Bézier curve B(t) has a total of n+1 control points, namely P0, P1, ..., P2. n t is a curvilinear function; S13: Extract the region of interest (ROI) for each frame in a high frame rate two-dimensional ultrasound image sequence.
3. The method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning as described in claim 2, characterized in that, In step S12, four frames are inserted between every two frames of the low frame rate two-dimensional ultrasound image sequence to obtain a high frame rate two-dimensional ultrasound image sequence.
4. The method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning as described in claim 1, characterized in that, In step S2, four insertion frames are inserted between every two consecutive frames. After ROI extraction processing is performed on each consecutive frame and the insertion frames, three-dimensional data is obtained. The cube composed of ROIs of multiple ultrasound images is used as the input of a three-dimensional convolutional neural network.
5. The method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning as described in claim 1, characterized in that, The training of the 3D convolutional neural network in step S2 and the training of the long short-term memory model in step S4 include: Initialize the learning rate, and compare the spatial pose information extracted by the 3D convolutional neural network with the pose information labels to obtain the mean squared error of the pose information loss, MSE. The expression for MSE is as follows: Where n is the number of samples, R0 i It is the true value, E i It is a predicted value; The mean squared error of pose information loss (MSE) and the feature sequence output by the 3D convolutional neural network are fed back into the long short-term memory model. The long short-term memory model continuously updates its parameters to reduce the MSE loss of the output pose information.
6. The method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning as described in claim 1, characterized in that, In step S4, the long short-term memory model includes a forget gate, an input gate, and an output gate. The forget gate is used to determine the discarded information in the model state at the previous time step and update the model state, using the model state at the previous time step as a parameter to update the current state.
7. The method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning as described in claim 6, characterized in that, In step S4, a new time series data is generated by combining a multi-layer long short-term memory model with the characteristics of a predetermined time period.
8. The method for three-dimensional reconstruction of two-dimensional ultrasound images based on deep learning as described in claim 7, characterized in that, The three-dimensional reconstruction in step S4 is divided into pixel mapping and gap filling. During pixel mapping, each pixel of each frame of two-dimensional ultrasound image is transformed into the three-dimensional reconstruction coordinate system according to the spatial pose information of the current frame through coordinate transformation.
Citation Information
Patent Citations
Method for positioning three-dimensional human body joints in monocular color videos
CN107392097A
Four-dimensional ultrasonic reconstruction method and system based on two-dimensional ultrasonic image
CN114581549A