Spine three-dimensional reconstruction network implementation method and system based on biplane X-ray image

By combining deep convolutional neural networks and Bézier Gaussian Triangle surface representation, the efficiency and accuracy issues of three-dimensional spinal model reconstruction in dual-plane X-ray images were solved, and a smooth, geometrically accurate three-dimensional spinal model was generated, overcoming the shortcomings of traditional methods and achieving efficient reconstruction under low-dose imaging.

CN120707775APending Publication Date: 2025-09-26BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510801543.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When reconstructing three-dimensional spinal models from dual-plane X-ray images, existing technologies lack efficiency and accuracy in fusing dual-view information to generate three-dimensional representations, making it difficult to accurately reconstruct complex boundaries. The stair-step artifacts and non-differentiable steps caused by traditional methods hinder end-to-end optimization, making it difficult to balance the accuracy of voxel predictions with the geometric smoothness of the final surface mesh.

Method used

A deep convolutional neural network is used to generate a 3D voxel probability field, which is then combined with a differentiable Bézier Gaussian Triangle surface representation for geometric optimization. The boundary prediction accuracy is improved through a spatially weighted loss function. The differentiability of the Bézier Gaussian Triangle is exploited for surface reconstruction and optimization, resulting in a smooth and structurally continuous 3D model.

Benefits of technology

It achieves efficient and accurate reconstruction of three-dimensional spinal models under low-dose imaging conditions, generating a vectorized model with smooth surface and precise geometry, overcoming the staircase artifact problem of traditional methods and improving the geometric accuracy and surface smoothness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707775A_ABST
    Figure CN120707775A_ABST
Patent Text Reader

Abstract

The invention relates to a spine three-dimensional reconstruction network implementation method and system based on a biplane X-ray image, and aims to accurately and efficiently reconstruct a three-dimensional spine model with a smooth surface from a low-dose X-ray image by combining a deep convolutional neural network and differentiable parameterized surface representation. The method comprises the following steps: firstly, preprocessing an input biplane X-ray image, and fusing the biplane X-ray image into a multi-channel three-dimensional volume representation; then, inputting the three-dimensional volume into a neural network based on a 3D U-Net architecture, and training by adopting a space weighted loss function to output a three-dimensional voxel probability field with an accurate boundary; then, a group of BG-Triple parameterized surface primitives are initialized based on the probability field; and finally, through a geometric loss function containing Chamfer distance and normal consistency, iterative optimization is directly carried out on the geometric control point of the BG-Triangle, so that a vectorized smooth three-dimensional spine model is obtained. According to the invention, accurate mapping from the voxel probability field to the smooth geometric grid is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of medical image processing, computer vision and three-dimensional reconstruction, and more specifically, to a method and system for accurately and efficiently reconstructing a three-dimensional spinal model with a smooth surface from dual-plane X-ray images by using a deep convolutional neural network to predict voxel probability fields and combining them with differentiable parameterized surface representations (such as Bézier Gaussian Triangles) for geometric optimization. Background Art

[0002] Conventional CT scans are high in radiation, while MRI scans are expensive and often performed in the non-weight-bearing position. This makes low-dose, standing-position biplane X-rays an important data source. However, recovering three-dimensional structures from two-dimensional projections presents inherent challenges, including loss of depth information and tissue overlap.

[0003] Early methods such as statistical shape models (SSMs) and template-based registration have limited robustness to complex deformities or insufficient accuracy. Deep learning methods provide end-to-end solutions. Research directions include: combining pose priors, voxel-based dense predictions, implicit neural representations, using Transformers to enhance cross-view modeling (such as Swin-X2S), generating adversarial networks (such as ReVerteR) or diffusion models to improve realism, and improving accuracy through edge-aware networks (such as EAR), attention mechanisms, synthetic data, coordinate registration, and spatial consistency constraints.

[0004] Core pain points: Existing methods still face challenges: (1) How to efficiently and accurately fuse dual-view information to generate a three-dimensional representation; (2) How to accurately reconstruct the complex boundaries of the vertebra; (3) Voxel prediction-based methods usually rely on post-processing algorithms such as MarchingCubes to generate meshes from voxel fields or point clouds, which can lead to staircase artifacts, loss of details, and are usually non-differentiable steps, hindering end-to-end optimization; (4) In a complex end-to-end framework, how to effectively balance the accuracy of voxel prediction and the geometric smoothness of the final surface mesh, posing a major technical challenge to model optimization and parameter tuning. The present invention aims to solve the above problems, in particular, to improve the surface smoothness and geometric accuracy of the reconstructed model while retaining the advantages of vectorized representation. Summary of the Invention

[0005] The present invention provides a method and system that can effectively fuse dual-plane X-ray information, accurately predict three-dimensional spinal structure (especially in the boundary area), and use advanced differentiable parameterized surface representation and geometric optimization technology to generate a vectorized three-dimensional spinal model with smooth surface, continuous structure and accurate geometry.

[0006] The overall structural diagram of the system proposed by the present invention can be referred toFigure 1 As shown, the system mainly includes the following modules:

[0007] Input preprocessing and 3D volume construction module: Input biplane X-ray image pair (anteroposterior I ap and lateral I lat The images were preprocessed, including intensity normalization (Z-score normalization) and uniform size adjustment (adjusted to 128 × 128 pixels, and I′ was obtained). ap , Il′ at ), then, as Figure 2 As shown, a fixed geometric operation (extrude each 2D view along the depth dimension D = 128) is used to form a 3D volume V ap ,V lat ):

[0008]

[0009] The two single-channel volumes are then concatenated along the channel dimension to form a two-channel three-dimensional tensor:

[0010] X 3D =concat channel (V ap ,V lat )

[0011] As the input of the subsequent network, B is the batch size, indicating the number of batches, H and W are the height and width of the image, H = W = 128.

[0012] 3D voxel probability field generation network module (based on 3D U-Net): The network structure diagram of this module can be referred to Figure 3 .

[0013] Network structure: using optimized 3D U-Net architecture

[0014] Encoder path (f enc_path ): Enter X 3D First, after the initial convolution layer f enc_init (Convolution kernel size is 3×3×3, batch normalization BN, ReLU activation, output C enc0 =8 channels) to extract shallow features:

[0015] E (0) =f enc_init (X 3D )

[0016] Then, the feature map passes through four downsampling modules f down_l(l=1,2,3,4). Each module contains two 3×3×3 convolutional layers (with BN, ReLU), and uses a stride of 2 in the first convolutional layer to achieve 2x downsampling. Multi-scale feature maps E output at each stage of the encoder path (l) are saved for skip connections:

[0017] E (l) =f down_l (E (l-1) ),l=1,2,3,4

[0018] After four downsampling stages, the deepest feature E of the encoder is obtained (4) (bottleneck), its spatial resolution is 8×8×8. The resolution change process is: 128 3 (input)→128 3 (E (0) )→64 3 (E (1) )→32 3 (E (2) )→16 3 (E (3) )→8 3 E (4) .

[0019] Decoder path (f dec_path ): Contains three upsampling modules f up_l (l=1,2,3), which aims to convert the 8 3 Features are gradually upsampled to 64 3 The target resolution of each upsampling module f up_l Receive the feature D from the previous decoding layer (l-1) (Definition D (0) =E (4) is the bottleneck feature) and fuses the skip connection feature E from the corresponding level of the encoder (4 -l) The fused features are processed by 2x upsampling and several 3×3×3 convolutional layers (including BN, ReLU) to obtain the output D (l) :

[0020] D (l) =f up_ l(Concta(D (l-1) ,E (4-l) )),l=1,2,3

[0021] (Concat represents the concatenation operation of the jump connection feature and the upsampling feature, and the fusion process can be done as follows Figure 4 The FFAG module shown in the figure). The resolution change process of the decoder is:3 (D (0) )→16 3 (D (1) )→32 3 (D (2) )→64 3 (D (3) ).

[0022] Output layer (f out ): Feature map D output by the last level of the decoder (3) (Spatial resolution 64 3 ) After the final output layer f out Processing. out Contains a 1×1×1 convolution layer (adjusting the number of channels to 1) and a Sigmoid activation function to compress the input value to between 0 and 1, representing the voxel value P at each spatial position (x, y, z) 1.0;x,y,z The probability of belonging to the spine generates the final voxel probability field P 1.0 ∈[0,1] 64×64×64 :

[0023] P 1.0 =f out (D (3) )

[0024] Training loss (supervisory P 1.0 ): Combine two loss functions for supervision:

[0025] Standard voxel loss Using binary cross entropy (BCE) loss, the predicted probability P is compared voxel by voxel. 1.0 and the true binary segmentation label Y(y x,y,z ∈{0,1}), where Y is the true label obtained by processing the original CT scan data and its corresponding 3D segmentation annotation to guide the training process, y x,y,z is the value of a voxel with coordinates (x, y, z) in Y. x,y,z =1, it means that the voxel belongs to the spine. x,y,z =0, indicating that the voxel does not belong to the spine, the total number of voxels N = 64 3 The formula is:

[0026]

[0027] Spatial weighted loss Introducing spatial weight w based on BCE x,y,z , for objects close to the real object boundary S GT The voxels with higher importance are given to improve the boundary prediction accuracy. Based on the above, higher weights w are assigned to voxels close to the boundary of real objects. x,y,z , weight w x,y,z Based on voxel (x, y, z) to real surface S GT The shortest Euclidean distance d x,y,z Calculation, when a voxel is very close to the surface (d x,y,z When it approaches 0), its weight w x,y,z will approach the maximum value; as the voxel moves away from the surface, the weight will gradually decrease, thereby achieving the purpose of giving higher importance to voxels close to the boundary. The formula is:

[0028]

[0029]

[0030] The hyperparameters α and t control the weight distribution (assuming t = 1.8, α = 1.0). Calculate d x,y,z Need to obtain S in advance GT The distance transform field. and Together, we guide the 3D U-Net decoder to generate high-quality voxel probability fields P 1.0 .

[0031] Surface reconstruction and optimization module based on Bézier Gaussian Triangle (BG-Triangle): The overall flow diagram of this module can be referred to Figure 5 This module uses BG-Triangle, an advanced differentiable parameterized surface representation technology, to generate and optimize the final mesh. The input of this stage is the voxel probability field P generated by the decoder. 1.0 ∈[0,1] 64×64×64 , the output is a vectorized 3D grid

[0032] Initial point cloud generation (S): First, we need to generate the voxel probability field P 1.0 We use the Marching Cubes algorithm to extract P 1.0 The isosurface vertices at the threshold τ = 0.5 are used as the initial point cloud S:

[0033] S = MarchingCubesVertices(P 1.0 ,τ=0.5)

[0034] It should be clear that Marching Cubes only serves as an initial reference for generating vector surface control points and does not participate in the differentiable training process of the final model.

[0035] BG-Triangle initialization: BG-Triangle is a representation based on Bézier triangle surface. A Bézier triangle surface of degree n is composed of its control points Definition (where i, j, k are non-negative integer indices associated with the control points, n is the degree of the Bézier surface i+j+k=n and i, j, k≥0). Any point S(u, v, w) on the surface can be defined by its barycentric coordinates (u, v, w) (u+v+w=1, u, v, w≥0) and the Bernstein basis function

[0036] Interpolation yields:

[0037]

[0038] During initialization, a set of BG-Triangle primitives is initialized using the generated point cloud S, primitives are assigned to points in S (or their neighborhoods), and their initial geometric control points are set. In addition to geometric control points, each BG-Triangle primitive also carries attribute information. These attributes can be obtained by using additional color control points. Perform barycentric interpolation or store in a multi-resolution two-dimensional attribute map M h and perform texture interpolation based on the barycentric coordinates (e.g. for rotating r q , zoom s q , spherical harmonic coefficient (SH)sh q Equal attributes a h (q)). The resolution of the attribute map can be adjusted according to the sensitivity of the attribute (the SH map resolution is 1x1, and the others are 3x3 triangle grids).

[0039] Optimization based on geometric loss: Unlike BG-Triangle which mainly uses photometric loss, this study uses geometric loss (see below) to directly optimize the shape of the BG-Triangle. The key lies in the geometric differentiability of the BG-Triangle representation: predicting the surface point and normals Relative to the control point {p i The gradient of} can be calculated. Therefore, we calculate (where S GT is the true model) about the control point {p i}, and use an optimizer (such as Adam) to iteratively update these control point parameters so that the prediction grid Geometrically approximate the real model S GT The loss Contains two items:

[0040] Chamfer distance Measuring two surface point sets The two-way mean squared closest point distance between :

[0041]

[0042] in, Indicates the predicted surface The set of upsampled 3D points, Represents the target real surface S GT The set of upsampled 3D points, Represents a point set The number of midpoints, Represents a point set The number of midpoints, Represents the prediction surface A point on the surface, p represents the real surface S GT A point on.

[0043] Normal consistency loss Penalize corresponding pairs The normal vector n(p) between Deviation:

[0044]

[0045] Where K is the number of corresponding point pairs and <·,·> is the cosine similarity.

[0046] Output: Finally, this stage outputs a vectorized, smooth 3D spine mesh consisting of optimized BG-Triangles The standard triangular mesh can be extracted from it by the BG-Triangle subdivision method for visualization and evaluation.

[0047] Overall training: The network is trained to optimize a composite total loss It is and (Include and ) is the weighted sum of:

[0048]

[0049] Through joint training, the 3D U-Net network parameters and the learnable parameters of the BG-Triangle module are optimized at the same time. Set λ voxel ,=1.0,λ edge =2.0,λ surface =5.0.

[0050] System composition: The present invention also provides a three-dimensional spine model reconstruction system, as described above Figure 1 As shown, the system includes:

[0051] Image receiving unit: used to receive input biplane X-ray images.

[0052] Pre-processing unit: Configure and execute the operations in step 4.1 above. For the specific process, see Figure 2 .

[0053] Voxel probability prediction unit: contains the trained 3D voxel probability field generation network (structure see Figure 3 ), configure and execute the operations in step 4.2 above.

[0054] Surface reconstruction and optimization unit: Configure and execute the above step 4.3 (for the process, see Figure 5 ), including initial surface extraction, BG-Triangle initialization and optimization based on geometric loss.

[0055] Output unit: used to output the final optimized three-dimensional spine mesh model.

[0056] Storage unit: used to store images, network models, intermediate results, and final models.

[0057] Processor: used to execute the functions of the above units.

[0058] High surface quality: By introducing BG-Triangle parameterized representation and direct optimization based on geometric loss, the stair-step effect caused by traditional methods such as Marching Cubes is overcome, generating a 3D model with smoother surface and more continuous structure.

[0059] Boundary accuracy improvement: combining spatial weighted loss function Emphasizing boundary areas during the voxel prediction stage helps improve the accuracy of the model contour.

[0060] End-to-end optimization potential: Although the optimization described so far mainly acts on the BG-Triangle parameters, its differentiability provides the possibility of achieving deeper end-to-end training in the future (such as backpropagating gradients to U-Net).

[0061] The output is a vectorized mesh model, which is more compact than voxel representation and easier to use for subsequent geometric analysis, measurement, and visualization.

[0062] The present invention provides a safe alternative that can obtain accurate three-dimensional models under low-dose imaging conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1Structural diagram of the 3D spine model reconstruction system

[0064] Figure 2 Flowchart of 3D volume construction from biplanar X-ray images

[0065] Figure 3 3D U-Net network structure diagram

[0066] Figure 4 FFAG module schematic

[0067] Figure 5 Vectorized grid generation flow chart based on BG-Triangle DETAILED DESCRIPTION

[0068] The following will refer to Figures 1 to 5 , and in combination with preferred embodiments, the specific methods for implementing the present invention are described in detail.

[0069] Network implementation details: 3D U-Net: refer to Figure 3 The detailed description is as described in Section 4, including a 4-layer encoder and a 3-layer decoder structure, convolutional layers (3×3×3 kernel, BN, ReLU), downsampling (convolution with a stride of 2), upsampling (2×2×2 transposed convolution), and skip connections. The fusion of skip connections uses the FFAG module (the principle diagram can be found in Figure 4 ).

[0070] Input preprocessing: refer to Figure 2 , the input preprocessing steps of this embodiment are as follows:

[0071] Normalization: Z-score normalization.

[0072] Resizing: Use bilinear interpolation to resize to the target resolution, such as 128×128.

[0073] 3D volume construction: The preprocessed I′ ap and I′ lat Expand along the depth axis D (D = 128) into V ap and V lat Then set V ap and V lat Concatenate along the channel dimension into a B×2×D×H×W tensor (D=H=W=128).

[0074] BG-Triangle implementation details: refer to Figure 5 The BG-Triangle implementation details of this embodiment are as follows:

[0075] Primitive definition: Based on Bézier triangle patch, composed of geometric control points {p i}Define the shape.

[0076] Attribute representation: can contain color control points Used for interpolated colors and multi-resolution attribute maps M h Used to store and interpolate other attributes (such as rotation, scaling, SH coefficients). The attribute map resolution is adjustable, for example, SH is 1x1, and other is 3x3 triangle.

[0077] Optimizer: Adam optimizer is used and the BG-Triangle learning rate is set to 0.001.

[0078] Loss function parameters:

[0079] Weight calculation requires pre-calculation of S GT The distance transform field of . The hyperparameters are set to t = 1.8, α = 1.0.

[0080] Contains Chamfer distance and normal consistency loss. Need to predict the surface and the real surface S GT Sampling point set And it is necessary to determine the corresponding point pairs for normal loss calculation.

[0081] Total loss weight (λ): λ voxel =1.0,λ edge =2.0,λ surface =5.0.

[0082] Training data generation:

[0083] Dataset: We use the public dataset VerSe 2019.

[0084] DRR simulation: Generate simulated biplanar X-ray images (e.g., 128×128 resolution) from CT data and segmentation annotations using digitally reconstructed radiographs (DRR) as network input. ap ,I lat .

[0085] Label processing: Convert 3D segmentation annotations to target voxel labels Y(64 3 resolution) and surface mesh S GT If you use Need to pre-calculate S GT The distance transform field.

[0086] Data partitioning: The training set, validation set, and test set are randomly divided at the vertebral sample level, for example, in a ratio of 8:1:1.

[0087] Training parameters:

[0088] Optimizer: Adam.

[0089] Learning rate: initial value 1×10 -3 . Decay strategy: the learning rate is multiplied by 0.1 every 50 cycles.

[0090] Batch Size: 4.

[0091] Epochs: The maximum number of epochs is 150, and training is stopped when the average Chamfer distance on the validation set is less than or equal to 2.0 mm.

[0092] Hardware / Software: The experiments were conducted on a GPU (NVIDIA RTX 3090) using the deep learning framework PyTorch. BG-Triangle operations were implemented using the public code library.

Claims

1. A three-dimensional spine model reconstruction method, characterized in that: The following steps are involved: a. Obtain a biplane X-ray image of the object to be processed; b. Preprocessing the biplane X-ray images and fusing them into a multi-channel three-dimensional volume representation through geometric expansion and stitching operations; c. Input the 3D volume representation into a convolutional neural network based on a 3D U-Net architecture, which is trained to output a 3D voxel probability field, wherein the training process uses a spatially weighted cross entropy loss. The loss function of d. Applying an isosurface extraction algorithm Marching Cubes to the three-dimensional voxel probability field to generate an initial surface point cloud; e. Based on the initial surface point cloud, initialize a set of Bézier Gaussian Triangle (BG-Triangle) parameterized surface primitives, each primitive having an optimizable geometric control point ({p i }); f. Define a Chamfer distance loss and normal consistency loss The geometric loss function Used to measure the prediction surface defined by the current BG-Triangle control points The target real surface (S GT ) between the geometric differences; g. Using the geometric differentiability of BG-Triangle, the control points ({p i }), to minimize the geometric loss function Thus, the final vectorized and smooth three-dimensional spine model is obtained.

2. The method according to claim 1, characterized in that The 3D U-Net architecture includes an encoder path (f enc_path ) for downsampling and feature extraction, and the decoder path (f dec_path ) is used for upsampling and feature recovery, and skip connections are used to fuse the features of the encoder and decoder.

3. The method according to claim 1, characterized in that The spatially weighted cross entropy loss According to the distance from the voxel to the target real surface (d x,y,z ) assigns weights (w x,y,z ), the closer the distance, the higher the weight.

4. The method according to claim 1, wherein The optimizer in step g is the Adam optimizer.

5. A three-dimensional spine model reconstruction system, characterized in that: include: A memory storing instructions and a processor configured to execute the instructions to implement the method according to any one of claims 1 to 4.

6. A computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

7. The method according to claim 1, characterized in that Input biplane X-ray image pair, anteroposterior I ap and lateral I lat ; Preprocess the image, including intensity normalization (Z-score normalization) and uniform size adjustment, adjust to 128 × 128 pixels, and obtain I′ ap , I′ a , each 2D view is expanded along the depth dimension D = 128 to form a 3D volume V through fixed geometric operations ap ,V lat : Then concat these two single-channel channel volumes along the channel dimension to form a two-channel three-dimensional tensor: X 3D =concat channel (V ap ,V lat ) Specifically, it is B×2×128×128×128, which serves as the input of the subsequent network, where B is the batch size, indicating the number of batches, H and W are the height and width of the image, and H=W=128; Adopting an optimized 3D U-Net architecture; Encoder path (f enc_path ): Enter X 3D First, after the initial convolution layer f enc_init , convolution kernel size is 3×3×3 convolution, batch normalization BN, ReLU activation, output C enc0 =8 channels, extracting shallow features: E (0) =f enc_init (X 3D ) Then, the feature map passes through four downsampling modules f down_l , l = 1, 2, 3, 4; each module contains two 3 × 3 × 3 convolutional layers, batch normalization BN, ReLU activation, and a stride of 2 is used in the first convolutional layer to achieve 2x downsampling; the multi-scale feature map E output at each stage of the encoder path (l) are saved for skip connections: AND (l) =f down_l (AND (l-1) ),l=1,2,3,4 After four downsampling stages, the deepest feature E of the encoder is obtained (4) (bottleneck), its spatial resolution is 8×8×8; the resolution change process is: 128 3 (input)→128 3 (E (0) )→64 3 (E (1) )→32 3 (E (2) )→16 3 (E (3) )→8 3 E (4) ; Decoder path (f dec_path ): Contains three upsampling modules f up_l , l=1,2,3, aiming to convert the 8 3 Features are gradually upsampled to 64 3 The target resolution of each upsampling module f up_l Receive the feature D from the previous decoding layer (l-1) (Definition D (0) =E (4) is the bottleneck feature) and fuses the skip connection feature E from the corresponding level of the encoder (4-l) The fused features are processed by 2x upsampling and several 3×3×3 convolution layers, batch normalization BN, and ReLU activation to obtain the output D (l) : D (l) =f up_l (Concat(D (l-1) ,E (4-l) )),l=1,2,3 Concat represents the concatenation operation of the jump connection feature and the upsampling feature; the resolution change process of the decoder is: 8 3 (D (0) )→16 3 (D (1) )→32 3 (D (2) )→64 3 (D (3) ); The feature map D output by the last stage of the decoder (3) After the final output layer f out Processing; out Contains a 1×1×1 convolution layer and a Sigmoid activation function to compress the input value to between 0 and 1, representing the voxel value P at each spatial position (x, y, z) 1.0;x,y,z The probability of belonging to the spine generates the final voxel probability field P 1.0 ∈[0,1] 64×64×64 : P 1.0 =f out (D (3) ) Combine two loss functions for supervision: Standard voxel loss Using binary cross entropy (BCE) loss, the predicted probability P is compared voxel by voxel. 1.0 and the true binary segmentation label Y(y x,y,z ∈{0,1}), where Y is the true label obtained by processing the original CT scan data and its corresponding 3D segmentation annotation to guide the training process, y x,y,z is the value of a voxel with coordinates (x, y, z) in Y. x,y,z =1, it means that the voxel belongs to the spine. x,y,z =0, indicating that the voxel does not belong to the spine, the total number of voxels N = 64 3 ; The formula is: Spatial weighted loss Introducing spatial weight w based on BCE x,y,z , for objects close to the real object boundary S GT The voxels with higher importance are given to improve the boundary prediction accuracy; the loss is Based on the above, higher weights w are assigned to voxels close to the boundary of real objects. x,y,z , weight w x,y,z Based on voxel (x, y, z) to real surface S GT The shortest Euclidean distance d x,y,z Calculation, when a voxel is very close to the surface (d x,y,z When it approaches 0, its weight w x,y,z will approach the maximum value; the formula is: Hyperparameters α and t control the weight distribution, t = 1.8, α = 1.0; calculate d x,y,z Need to obtain S in advance GT The distance transformation field of and Together, we guide the 3D U-Net decoder to generate the voxel probability field P 1.0 ; The final mesh is generated and optimized using BG-Triangle, an advanced differentiable parameterized surface representation technology; the input of this stage is the voxel probability field P generated by the decoder 1.0 ∈[0,1] 64×64×64 , the output is a vectorized 3D grid Initial point cloud generation (S): First, we need to generate the voxel probability field P 1.0 Extract the three-dimensional point cloud S that can represent the target surface; use the Marching Cubes algorithm to extract P 1.0 The isosurface vertices at the threshold τ = 0.5 are used as the initial point cloud S: S=MarchingCubesVertices(P 1.0 ,τ=0.5) Marching Cubes only serve as an initial reference for generating vector surface control points and do not participate in the differentiable training process of the final model; BG-Triangle initialization: BG-Triangle is a representation based on Bézier triangle surface (Bézier TriangleSurface); an n-order Bézier triangle surface is composed of its control points Definition: where i, j, k are non-negative integer indices associated with the control points, n is the degree of the Bézier surface i+j+k=n and i, j, k≥0; any point S(u, v, w) on the surface is defined by the barycentric coordinates (u, v, w) (u+v+w=1, u, v, w≥0) and the Bernstein basis function Interpolation yields: During initialization, a set of BG-Triangle primitives is initialized using the generated point cloud S, primitives are assigned to points or neighborhoods in S, and their initial geometric control points are set. In addition to the geometric control points, each BG-Triangle primitive also carries attribute information; these attributes are obtained through additional color control points. Perform barycentric interpolation or store in a multi-resolution two-dimensional attribute map M h and perform texture interpolation based on the barycentric coordinates; Using geometric loss To directly optimize the shape of BG-Triangle; the loss Contains two items: Chamfer distance ( Measuring two surface point sets The two-way mean squared closest point distance between : in, Indicates that the predicted surface The set of upsampled 3D points, Represents the target real surface S GT The set of upsampled 3D points, Represents a point set The number of midpoints, Represents a point set The number of midpoints, Represents the prediction surface A point on the surface, p represents the real surface S GT A point on Normal consistency loss Penalize corresponding pairs The normal vector n(p) between Deviation: Where K is the number of corresponding point pairs, <·,·> is the cosine similarity; Outputs a vectorized, smooth 3D spine mesh consisting of optimized BG-Triangles Optimizing a composite total loss It is and Include and The weighted sum of: Through joint training, the 3D U-Net network parameters and the learnable parameters of the BG-Triangle module are optimized at the same time; setting λ voxel ,=1.0,λ edge =2.0,λ surface =5.0.

Citation Information

Cited By

  • Three-dimensional voxel damage prediction method and device based on component consistency constraint

    CN121659396A