Craniomaxillofacial three-dimensional reconstruction method based on free hand ultrasound
Through the craniofacial surface reconstruction method based on free hand 2D ultrasound, using binocular cameras and deep learning technology, non-invasive, real-time, high-precision craniomaxillofacial bone surface reconstruction and registration are achieved, solving the problems of invasiveness, low efficiency and insufficient accuracy in the prior art.
Patent Information
- Application Number
- CN202510766528.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing craniomaxillofacial surgical navigation registration methods have problems such as invasive operation, high production cost, low efficiency and insufficient registration accuracy, especially when soft tissue deformation is reduced during three-dimensional scanning of facial structure light.
The craniofacial surface reconstruction method based on free hand 2D ultrasound is adopted, and the position of the ultrasound probe is tracked in real time using a binocular camera, combined with deep learning segmentation and implicit neural representation network to achieve non-invasive, real-time three-dimensional reconstruction and registration.
It significantly improves the accuracy of surgical navigation registration, solves the problems of invasiveness and low efficiency, and achieves high-precision bone surface reconstruction and registration.
Smart Images

Figure CN120279197A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer simulation, and particularly to a three-dimensional craniofacial reconstruction method based on freehand ultrasound. Background Art
[0002] In modern medical practice, image-guided surgery (IGS), as a key technology, effectively assists surgeons in understanding the patient's condition by registering preoperative images such as CT or MRI with the anatomical structure of the subject during the operation and displaying virtual images on a computer screen or projecting them onto the surgical area using augmented reality technology. This greatly improves the accuracy of information. Especially in the field of craniofacial repair, IGS has broad application prospects and is particularly important for surgeries such as midfacial reconstruction after trauma. Image registration is a key step in IGS, and the accuracy of registration is crucial for the overall effect of surgical navigation. Generally, the acceptable registration error does not exceed 1 millimeter.
[0003] Existing craniofacial surgical navigation registration methods are mainly divided into two categories: feature marker-based and surface-based. Feature marker-based registration requires placing identifiable artificial markers on the subject's body, which are then associated with preoperative images. Commonly used markers include titanium screws implanted percutaneously into the bone, occlusal plates fixed to the upper and lower jaws, and metal balls pasted on the skin, etc. In contrast, surface registration uses structured light three-dimensional reconstruction to scan the exposed human surface of the surgical area, and after obtaining the three-dimensional point cloud, shape matching is performed through a surface registration algorithm. Summary of the Invention
[0004] The above methods all have certain limitations. For example, titanium screw implantation is an invasive operation that requires additional surgery to place markers before CT scanning, and the bone density of the implantation area needs to be carefully evaluated to ensure sufficient fixation strength. The occlusal plate or metal ball belongs to personalized registration markers, and their production requires additional time and cost, and the process of collecting coordinates during the operation is also very cumbersome. In addition, although facial structured light three-dimensional scanning is non-invasive, it can only scan the directly exposed human surface tissue, and the registration accuracy may be reduced due to soft tissue deformation (such as swelling after trauma).
[0005] Ultrasonic imaging has been widely used in clinical diagnosis, monitoring, and analysis due to its portability, non-invasiveness, radiation-free, low cost, and real-time capabilities. However, its potential in qualitative measurement of human structures has not been fully explored. Therefore, the present invention aims to provide a non-invasive, accurate, and real-time three-dimensional reconstruction solution for the ultrasonic bone surface for craniofacial surface reconstruction and registration, thereby improving the registration accuracy in craniofacial surgery. This method uses a stereo camera for optical tracking and obtains the pose data of a 2D ultrasonic image sequence. After deep learning segmentation, the skull structure is extracted and reconstructed into a voxel grid, and finally refined into a smooth polygon mesh through an implicit neural representation network. This method can not only significantly improve the registration accuracy of surgical navigation but also effectively solve the problems of invasiveness, low efficiency, and inaccuracy existing in the prior art.
[0006] (I) Method Overview An embodiment of the present invention provides a craniofacial surface reconstruction framework based on freehand 2D ultrasound to achieve non-invasive and high-precision registration for surgical navigation.
[0007] The method of the present invention mainly includes the following three parts: 1) Real-time volume reconstruction: Use a binocular camera to track the position of the ultrasonic probe in real-time, and at the same time, use a two-stream neural network to segment the ultrasonic image and reconstruct the voxel grid in real-time; 2) Voxel denoising: Use morphological post-processing and template matching methods to eliminate outliers, and then convert the voxel model into a point cloud; 3) Surface reconstruction: Finally, reconstruct the point cloud data into a polygon mesh through a self-supervised implicit neural representation network to obtain a smooth bone surface.
[0008] (II) Equipment and Devices The equipment and operation method used in the embodiment of the present invention are as Figure 1 shown. The subject wears a head-mounted Marker, and the binocular camera images, ultrasonic images, and three-dimensional reconstruction images are displayed in real-time on the computer software. The operator scans the subject's face with the ultrasonic probe while observing the ultrasonic image and the real-time reconstruction process on the computer screen to complete the real-time rough reconstruction. Then, denoising processing and refined surface optimization and fine reconstruction are carried out. Brief Description of the Drawings
[0009] Figure 1 It is a schematic diagram of the equipment structure used in the embodiment of the present invention, where (a) is the equipment device and (b) is the operation method.
[0010] Figure 2 It is a schematic diagram of the STDC network structure in the embodiment of the present invention, and ARM represents the Attention Refine Module.
[0011] Figure 3This is the network structure diagram of the FUNSR network in the embodiment of the present invention, where FC represents the fully connected layer.
[0012] Figure 4 This is the reconstructed result diagram in the embodiment of the present invention (the volume point cloud processing for the reconstruction results after c and later is obtained by using the b morphological post - processing denoising method): (a) Real - time volume reconstruction, (b) Morphological post - processing denoising, (c) Surface reconstruction by the marching cubes method, (d) FUNSR surface reconstruction, (e) FUNSR + smoothing loss, (f) Template matching denoising + FUNSR, (g) Template matching denoising + FUNSR + smoothing loss.
[0013] Figure 5 This is the model experimental result diagram in the embodiment of the present invention, where (a) The skull model implanted with magnetic beads and the Marker, (b) Measuring the position of the magnetic beads using a probe.
[0014] Figure 6 This is the navigation experiment reconstruction result diagram in the embodiment of the present invention, where (a) Bone surface reconstruction result, (b) CBCT skull model and magnetic beads, (c) Registration of the reconstructed surface and CT, (d) Positions of the magnetic beads (semi - transparent red) in CT and the magnetic beads (green) measured by the probe. Detailed implementation manners
[0015] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the following further details the present invention with reference to specific embodiments and the accompanying drawings. However, those skilled in the art know that the present invention is not limited to the drawings and the following embodiments.
[0016] The method of the present invention mainly includes the following three parts: (1) Real - time volume reconstruction: Use a binocular camera to track the position of the ultrasound probe in real - time, and at the same time, use a two - stream neural network to segment the ultrasound image and reconstruct the voxel grid in real - time; (2) Voxel denoising: Use morphological post - processing and template matching methods to eliminate outliers, and then convert the voxel model into a point cloud; (3) Surface reconstruction: Finally, reconstruct the point cloud data into a polygon mesh through a self - supervised implicit neural representation network to obtain a smooth bone surface. The details are as follows.
[0017] (1) Real - time volume reconstruction includes the following three steps.
[0018] 1.1. Stereo spatial positioning of the binocular camera Use a binocular camera and an X-corner marker (abbreviated as Marker) to track the spatial pose of the ultrasound probe and convert two-dimensional ultrasound images into three-dimensional volumes. Considering that the head of the subject may move during the operation, a head-mounted Marker is also designed to bind the coordinate system to the position of the subject's head to ensure precise positioning of the subject's space. By detecting the X-corners of the Marker in the images captured by the binocular camera and calculating based on the invariance of the X-corners in the y-axis direction and the parallax shift in the x-axis direction, the spatial pose of the Marker in the camera coordinate system and the coordinate transformation matrix between the camera coordinate system and the marker coordinate system are obtained.
[0019] For the process of ultrasound three-dimensional reconstruction, it is necessary to determine the position of each ultrasound image in the subject coordinate system, that is, the coordinate transformation matrix. 。The transformation matrix from the image coordinate system to the ultrasound probe marker coordinate system is obtained through calibration. 。Then, the transformation matrix from the image coordinate system to the subject coordinate system is obtained through the following coordinate transformation: 。
[0020] 1.2. Deep learning real-time segmentation To extract the three-dimensional bone surface, it is necessary to first extract the bone surface in the two-dimensional ultrasound image, so image segmentation is required. In the skull structure, areas such as the cheekbones and zygomatic arches are thinner compared to the cranial vertex (e.g., the parietal bone or frontal bone) and may not be clearly shown on a portable color Doppler ultrasound device. Traditional segmentation methods such as threshold segmentation cannot accurately separate these bones from the background. Therefore, a deep learning segmentation method is needed. The present invention selects to use an efficient two-stream neural network, the Short-Term Dense Concatenate network (STDC) to achieve real-time image segmentation.
[0021] The network structure is as Figure 2 shown, using a lightweight network structure to achieve real-time segmentation. To further improve the segmentation performance, an s-curve algorithm is applied before segmentation to enhance the contrast of the ultrasound image. First, the image is divided into small blocks, and then the contrast remapping is performed on each block separately. The pixel value r in the image block is first calculated and transformed through the following formula after normalization:
[0022] Then, it is renormalized from the range of [0, 1] to between the minimum and maximum pixel values of the image block. Among them, α = 0.9642, β = 8.594 10 -4 ,γ = 0.4962, δ = 0.07598.
[0023] During the process of acquiring ultrasound images and pose information, the operator needs to ensure that the Marker is captured and recognized by the camera, and the acquired ultrasound images are clear and continuous enough. To allow the operator to check the scan data in real time, the 3D reconstruction method is divided into two parts: online voxel reconstruction and offline surface reconstruction. The voxel reconstruction stage focuses on real-time performance and allows for the correction of deviations through repeated scans of the same area. The surface reconstruction stage focuses on generating a smooth and complete surface that is as close as possible to the actual bone surface.
[0024] 1.3 Voxel Reconstruction After obtaining the segmented 2D ultrasound images and the corresponding pose information, the 2D images are reconstructed into 3D voxels in real time. In the voxel reconstruction stage, to ensure real-time performance, the downsampled weighted pixel nearest neighbor (PNN) reconstruction algorithm is used. After obtaining the segmentation mask, the image is downsampled by a factor of 2, and the foreground part is selected as the region of interest (ROI) using a quadrilateral bounding box. Subsequently, each pixel within the ROI is transformed into coordinates and assigned to the nearest voxel grid. Since the segmentation results cannot be completely accurate during the scanning and segmentation processes, a weighted assignment algorithm is introduced to reduce the noise caused by misclassification of the segmentation algorithm. When resampling at the same spatial position, the voxel value at that position is updated using the weighted algorithm. Let dImage represent the downsampled image, represent the voxel value corresponding to the pixel (i, j), represent the weight value. The update of the voxel value is as follows:
[0025] Meanwhile, the weight is updated as follows: .
[0026] (2) Voxel Denoising After voxel reconstruction, two-step post-processing steps are used to remove the noise in the volume model. The first step is morphological post-processing. First, all voxels are traversed, and the 26 voxels adjacent to this voxel are checked. If the number of adjacent non-zero voxels is less than 6, then this voxel is judged as an outlier and eliminated. Then, connected component analysis is performed on the entire voxel model, and the connected components with a small number of voxels are eliminated, retaining the zygomatic arch and forehead regions.
[0027] The second step is template matching denoising. First, with the voxel center as the coordinate, the voxel model is converted into a point cloud. Then, the ICP algorithm is used to coarsely register this point cloud data with the CT to obtain the distance error between each point in the point cloud and its best matching point in the CT. Select 1mm as the threshold, and the points with a distance error exceeding this value are regarded as outliers and eliminated.
[0028] (3) Surface Reconstruction In the voxel reconstruction stage, even though the data has been denoised to a certain extent, due to the spatial gaps that often exist between US images during scanning, the point cloud obtained after reconstruction will have the problem of uneven density. If traditional methods such as Marching Cubes are used for surface reconstruction, the obtained surface will be uneven or have holes, which does not conform to the true shape of the bone surface. The present invention uses a self-supervised implicit neural representation (INR) network, FUNSR (Freehand 3D Ultrasound Neural Surface Reconstruction). By moving the three-dimensional query points sampled around the volumetric point cloud to iteratively learn the signed distance function (SDF), the point cloud obtained from the previous stage is converted into a smooth polygon mesh.
[0029] The structure of the network is as Figure 3 shown. For the input point cloud, a large number of query points Q are first generated within its spatial range. Using the mode of generative adversarial learning, an MLP neural network is used as the generator to learn the SDF values and gradients of the query points in Q. Then, the query points use the predicted SDF values to project along the gradient or in the direction opposite to the gradient to their nearest real points to calculate the self-supervised loss . In addition, a direction loss is added to prevent the position of the query points from oscillating on both sides of the zero level set during training. Then, an OSC-ADL module, that is, an adversarial discriminator composed of four fully connected multi-layer perceptrons and Leaky ReLU activation functions . The SDF values learned from the network are input into this module, and then the confidence level of this SDF value is generated. This enables the model to effectively reduce the points that are close to the zero level set but do not belong to the surface.
[0030] For the voxel point cloud obtained from ultrasonic scanning, there are problems such as low density, much noise, and uneven distribution. The surface obtained by directly reconstructing using the FUNSR network will be uneven due to overfitting the shape of the point cloud. Therefore, the embodiments of the present invention adjust the loss function of this network to correct this problem. A new smoothing loss is introduced. Let the magnitude of the SDF gradient of a certain point be
[0031] Take the gradient of the magnitude of this gradient, that is, the second derivative of the SDF
[0032] Let the smooth loss be
[0033] This loss function ensures the smoothness of the reconstructed surface by penalizing drastic changes in the SDF.
[0034] The loss function of the entire network, i.e., the generator loss and the discriminator loss are as follows:
[0035]
[0036] where , and are loss weights, is the confidence value predicted by the SDF learned by , and are alternately optimized during training. and represent the predicted signed distance and the value of the zero scalar field (surface constraint), respectively. The magnitude of the zero scalar field value is defined according to the batch size.
[0037] Reconstruction effect and accuracy verification Using the above method, 10 volunteers were scanned by ultrasound and the facial bone surface was reconstructed. The following Surface Registration Error (SRE) was used as a criterion to measure the matching degree between the reconstructed result and the real bone surface. The reconstructed surface and the skull CBCT model of the volunteer were respectively converted into point clouds and . This formula represents for each point in and its best matching point in The root mean square error of the distance, in millimeters. Where represents the number of points in
[0038]
[0039] Table 1 shows the SRE of 20 groups of experiments of 10 volunteers. It can be found that the error of real-time rough reconstruction has reached within 1 mm, and the average error has been successfully reduced to within 0.5 mm after fine processing. In contrast, the traditional Marching Cubes method will increase the error instead.
[0040] Figure 4The polygonal mesh of the facial bone surface reconstructed by the present invention proves that the present invention can extract the shape of the head bone surface well.
[0041] Table 1 SRE of the facial scan of the volunteer, unit: mm
[0042] The registration accuracy was verified using a model experiment. Fix the X-corner marker on the 3D printed skull model, and embed 16 magnetic beads with a diameter of 2 mm at positions such as the maxilla, zygomatic bone, zygomatic arch, supraorbital ridge, and frontal bone of the model. The shape and position of the magnetic beads are clearly visible in the CT image. Use a probe with a spherical groove with a diameter of 2 mm at the tip to touch the magnetic beads on the model, and use a binocular camera to record the positions of the magnetic beads in the subject space. Place the model in a water tank, and use an ultrasonic probe to scan and reconstruct its surface contour in real time. After surface reconstruction and registration, calculate the error between the positions of the magnetic beads recorded by the probe and the positions of the magnetic beads in the CT image to represent the alignment error between the image space and the subject space.
[0043] Use the following Target Registration Error (TRE) to quantify the alignment error between the image space and the subject space:
[0044] This formula represents each pair of corresponding points in the image space and the subject space and the distance error, represents the L2 norm, that is, the Euclidean distance. represents the number of points recorded by the probe.
[0045] Figure 6 Figure 28 shows the reconstruction results of the navigation experiment, a visual display of the surface registration results and the positions of the magnetic bead markers. The scanning range of the skull model includes the zygomatic bone, zygomatic arch, orbits, and upper jaw, basically covering the positions where the magnetic beads are embedded in the model. The positions of 15 out of 16 magnetic beads were measured using the probe in the subject space. Table 2 shows the TRE measured in the navigation experiment. The average TRE for registration using voxel direct sampling of the point cloud is 1.115 mm, while after surface reconstruction, the average TRE is reduced to less than 1 mm. This experiment once again proves that the surface reconstruction method based on deep learning can reduce the influence of noise in real-time reconstruction, reconstruct a surface closer to the real bone surface, thereby obtaining a smaller TRE and improving the registration accuracy.
[0046] Table 2 TRE of the model experiment, comparing the registration accuracy after real-time rough reconstruction and surface reconstruction, unit: mm .
Claims
1. A three-dimensional craniofacial reconstruction method based on freehand ultrasound, characterized in that, The method includes: (1) Real-time volume reconstruction: Using a binocular camera to track the position of the ultrasound probe in real time, and at the same time using a two-stream neural network to segment the ultrasound image and reconstruct the voxel grid in real time; (2) Voxel denoising: Using morphological post-processing and template matching methods to eliminate outliers, and then converting the voxel model into a point cloud; (3) Surface reconstruction: Reconstructing the point cloud data into a polygon mesh through a self-supervised implicit neural representation network to obtain a smooth bone surface.
2. The method according to claim 1, characterized in that: Step (1) Real-time volume reconstruction includes the following three steps. 1.
1. Stereo spatial positioning of the binocular camera Using a binocular camera and an X-corner marker to track the spatial pose of the ultrasound probe, and converting the two-dimensional ultrasound image into a three-dimensional volume; A head-mounted marker was designed to bind the coordinate system to the position of the subject's head; By detecting the X-corner of the marker in the images captured by the binocular camera, and based on the invariance of the X-corner in the y-axis direction and the parallax shift in the x-axis direction, the spatial pose of the marker in the camera coordinate system and the coordinate transformation matrix between the camera coordinate system and the marker coordinate system were calculated. Determine the position of each ultrasonic image in the coordinate system of the subject, that is, the coordinate transformation matrix ; Obtain the transformation matrix from the image coordinate system to the ultrasonic probe marker coordinate system through calibration , and then obtain the transformation matrix from the image coordinate system to the subject coordinate system through the following coordinate transformation: ; 1.
2. Deep learning real-time segmentation Using a two-stream neural network, namely the Short-Term Dense Concatenate network, STDC to achieve real-time image segmentation; Using a lightweight network structure to achieve real-time segmentation; Applying an s-curve algorithm to enhance the contrast of the ultrasound image before segmentation; First, the image is divided into small blocks, and then the contrast of each block is remapped separately; The pixel value r in the image block is first normalized and then transformed through the following formula: Then, it is inverse-normalized from the range of [0, 1] to between the minimum and maximum pixel values of the image patch; where, α = 0.9642, β = 8.594 10 -4 , γ = 0.4962, δ = 0.07598; During the process of obtaining the ultrasound image and pose information, ensure that the marker is captured and recognized by the camera, and the obtained ultrasound image is clear and continuous enough; The 3D reconstruction method is divided into two parts: online voxel reconstruction and offline surface reconstruction; The voxel reconstruction stage focuses on real-time performance and allows correcting deviations through repeated scanning of the same area; The surface reconstruction stage focuses on generating a smooth and complete surface as close as possible to the actual bone surface. 1.3 Voxel reconstruction After obtaining the segmented two-dimensional ultrasound image and the corresponding pose information, the two-dimensional image is reconstructed into a three-dimensional voxel in real time; the downsampled weighted pixel nearest neighbor PNN reconstruction algorithm is adopted in the voxel reconstruction stage; after obtaining the segmentation mask, the image is downsampled by a factor of 2, and the foreground part is selected using a quadrilateral bounding box as the region of interest ROI; subsequently, each pixel within the ROI is converted to coordinates and assigned to the nearest voxel grid; a weighted assignment algorithm is introduced to reduce the noise caused by misclassification of the segmentation algorithm; when resampling at the same spatial position, the voxel value at that position is updated using the weighted algorithm; let dImage denote the downsampled image, denote the voxel value corresponding to the pixel (i, j), denote the weight value, and the update of the voxel value is as follows: At the same time, the weights are updated as follows: 。 3. The method according to claim 2, wherein: (2) Voxel denoising includes: After voxel reconstruction, use two post-processing steps to remove the noise in the volume model; The first step is morphological post-processing. First, traverse all voxels and check the 26 voxels adjacent to this voxel. If the number of adjacent non-zero voxels is less than 6, then judge this voxel as an outlier and eliminate it; Then perform connected component analysis on the entire voxel model and eliminate the connected components with a small number of voxels, retaining the zygomatic arch and forehead area. The second step is template matching denoising. First, taking the voxel center as the coordinate, convert the voxel model into a point cloud, and then use the ICP algorithm to coarsely register the point cloud data with the CT to obtain the distance error between each point in the point cloud and its best matching point in the CT. Select 1mm as the threshold, and regard the points with a distance error exceeding this value as outliers and eliminate them.
4. The method according to claim 3, characterized in that: (3) Surface reconstruction includes: Using self-supervised implicit neural representation, INR network, FUNSR; iteratively learn the signed distance function SDF by moving three-dimensional query points sampled around volumetric point clouds, and convert the point cloud obtained in the previous stage into a smooth polygon mesh; For the input point cloud, a large number of query points Q are first generated within its spatial range; using the generative adversarial learning mode, an MLP neural network is used as the generator to learn the SDF values and gradients of the query points in Q; then, the query points use the predicted SDF values to project to their nearest true points along the gradient or in the direction opposite to the gradient to calculate the self-supervised loss ; in addition, a direction loss is added to prevent the positions of the query points from oscillating on both sides of the zero level set during training; then, an OSC-ADL module, that is, an adversarial discriminator composed of four fully connected multi-layer perceptrons and Leaky ReLU activation functions ; the SDF values learned from the network are input into this module, and then the confidence level of this SDF value is generated; this enables the model to effectively reduce the points that are close to the zero level set but do not belong to the surface; Introduce a new smooth loss. Let the magnitude of the SDF gradient at a certain point be Take the gradient of the magnitude of this gradient, that is, the second derivative of the SDF Let the smoothing loss be This loss function ensures the smoothness of the reconstructed surface by penalizing the drastic changes in the SDF; The loss function of the entire network, i.e., the generator loss and the discriminator loss are as follows: in, , and is the loss weight, Is The confidence value of the learned SDF prediction, and Alternate optimization during training; and They represent the predicted signed distance and the value of the zero scalar field (surface constraint), respectively; the magnitude of the zero scalar field value is defined according to the batch size.
Citation Information
Patent Citations
Projection full-convolution network three-dimensional model segmentation method based on fusion of multi-view-angle features
CN108389251A
Spinal operation auxiliary system and auxiliary equipment
CN117796905A
Circumferential synthetic aperture sonar image enhancement method based on implicit neural representation
CN118446926A
Implicit neural network magnetic resonance image reconstruction method based on sensitivity matrix constraint
CN119963678A
Registration Using Phased Array Ultrasound
US20140163377A1