Spine 3D model reconstruction method based on multi-scale symmetric architecture learning

By employing a multi-scale symmetric architecture learning method, utilizing generative adversarial networks and global-local Mamba blocks, X-ray images are rotated and stitched together to form a high-channel 3D volume, solving the problem of low fidelity in reconstructing the three-dimensional structure of the spine in existing technologies and achieving higher reconstruction accuracy and detail.

CN119919579BActive Publication Date: 2025-12-30NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411977720.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-12-30
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Existing methods for reconstructing the three-dimensional structure of the spine based on X-ray images rely excessively on prior information, resulting in low fidelity of the reconstruction results. Furthermore, these methods heavily depend on human intervention, and current technologies fail to effectively address the significant information loss and weakening issues inherent in existing methods.

Method used

A 3D model reconstruction method for the spine based on multi-scale symmetric architecture learning is adopted. A high-channel 3D volume is formed by rotating and stitching X-ray images. Generative adversarial networks and global-local Mamba blocks are used for feature extraction and fusion. Combined with a U-shaped network architecture, the optimization strategy is dynamically adjusted to improve the reconstruction quality.

Benefits of technology

Without relying on prior information, it effectively preserves the input 2D X-ray image information, improves the detail and accuracy of the reconstructed 3D model of the spine, reduces the dependence on manual intervention, and enhances the fidelity of the reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919579B_ABST
    Figure CN119919579B_ABST
Patent Text Reader

Abstract

The application discloses a spine 3D model reconstruction method based on a multi-scale symmetric architecture learning, comprising the following steps: rotating a spine X-ray image to a set angle; repeating the rotated X-ray image D times and splicing to form a high-channel 3D volume; inputting the high-channel 3D volume into a generative adversarial network to obtain a 3D prediction model of the spine; calculating the standard deviation of different generated images of the 3D prediction model to evaluate the quality difference of the generated images, and selecting a global optimization or local optimization 3D prediction model of the generator based on the quality difference; and the 3D prediction model obtains a reconstruction model of the spine by minimizing the loss of the generator. The application solves the problem that the existing method for reconstructing the three-dimensional structure of the spine in the X-ray excessively depends on prior information, thereby reducing the fidelity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of biomedical image, and particularly relates to a spine 3D model reconstruction method based on a multi-scale symmetric architecture learning. BACKGROUND

[0002] Adolescent Idiopathic Scoliosis (AIS) is a common spinal disease in the adolescent population, and its etiology is usually related to bone growth imbalance or genetic factors. Since the treatment process of scoliosis is long, it is necessary to regularly evaluate the progress of the disease through imaging examination. X-Ray imaging has become the preferred method for measuring the angle of scoliosis in the clinic due to its low radiation dose, low cost, and fast imaging speed. The shape and curvature of the spine are usually evaluated by taking anteroposterior and lateral X-Ray images. However, due to the working principle of X-Ray imaging, which is to use X-rays to penetrate human tissues and project different density tissues onto a two-dimensional plane to generate images, the tissues in the image will inevitably overlap. This imaging feature makes it difficult to intuitively present the three-dimensional morphology of the anatomical structure of the spine, limiting the comprehensive analysis of complex structures.

[0003] In contrast, CT scanning can clearly show the three-dimensional structure of the spine and provide more detailed diagnostic information. However, CT scanning has a large radiation dose and high cost, which is not suitable for long-term follow-up of the adolescent population. In addition, CT imaging is usually completed in a supine position, which may affect the natural curvature of the spine and thus bias the evaluation results of scoliosis. Therefore, reconstructing the three-dimensional structure of the spine from X-Ray has become a research direction of interest in recent years.

[0004] Some early studies have proposed methods that rely on manual annotation of rib centerlines. These methods iteratively fit the spine curve through numerical analysis and interpolation techniques. However, such methods rely heavily on human intervention, resulting in inconsistent reconstruction results. With the development of deep learning, researchers have begun to use convolutional neural networks (CNN) to capture deep image features and convert two-dimensional X-Ray features into three-dimensional spine structures. Although these methods can extract semantically meaningful representations of the underlying three-dimensional scene from two-dimensional views, their asymmetric architecture may not fully capture all the information encoded in the input two-dimensional images. This loss of information can weaken the correlation between the reconstructed three-dimensional output and the original two-dimensional input. In addition, this method may lead to excessive reliance on the internal prior distribution of the model (such as Gaussian distribution), at the expense of inherent information in the two-dimensional signal, thereby reducing the fidelity of the reconstruction.

[0005] The information disclosed in this Background section is only for the purpose of increasing an understanding of the general background of the present application and should not be taken as an acknowledgement or any form of suggestion that this information forms the general knowledge of those skilled in the art before the present application. SUMMARY

[0006] To overcome the defects existing in the prior art, a spine 3D model reconstruction method based on multi-scale symmetric architecture learning is provided to solve the problem of low fidelity caused by excessive reliance on prior information in the existing method of reconstructing the three-dimensional structure of the spine in X-ray.

[0007] To achieve the above-mentioned purpose, a spine 3D model reconstruction method based on multi-scale symmetric architecture learning is provided, comprising the following steps:

[0008] Rotate multiple orthogonal X-ray images of the spine to a set angle where the shooting angles of the X-ray images are consistent;

[0009] Repeat the rotated X-ray images D times and splice the repeated X-ray images to form a high-channel 3D volume, so that the target depth of the high-channel 3D volume is adapted to the target depth of the 3D model of the spine;

[0010] Input the high-channel 3D volume into a generative adversarial network, perform preliminary feature extraction through the encoder 3D convolution of the generator of the generative adversarial network, and perform deep feature learning through the GLMamba block of the encoder to extract features and restored high-resolution features, the decoder of the generator splices the extracted features and the restored high-resolution features, and through the convolution block of the decoder, the final features are fused into continuous values to obtain the 3D prediction model of the spine;

[0011] Calculate the standard deviation of different generated images of the 3D prediction model to evaluate the quality difference of the generated images, and based on the quality difference, the generator globally optimizes or locally optimizes the 3D prediction model;

[0012] The 3D prediction model obtains the reconstructed model of the spine by minimizing the loss of the generator.

[0013] Further, the generator uses a U-shaped network as the basic architecture.

[0014] Further, the global optimization of the generator for the 3D prediction model includes using a double-regularized optimal transport learning method and a geometrically unbiased Sinkhorn divergence as a cost function.

[0015] Furthermore, the generator locally optimizes the 3D prediction model by constructing a mask matrix to mark low-quality regions and performing different levels of optimization based on the value of the local standard deviation.

[0016] The beneficial effects of this invention lie in its multi-scale symmetric architecture learning-based 3D spine model reconstruction method. By fusing a Mamba symmetric reconstruction network, it effectively preserves the information content of the input 2D x-ray without relying on other prior information. First, the 2D view is preprocessed. After repeating the x-ray to achieve the same depth as the matching 3D model, coarse registration and connection are performed to establish a 3D-to-3D symmetric architecture mapping reconstruction framework. Next, the input is fed into a designed generative adversarial network for same-dimensional translation. For the generator, a designed Global-Local Mamba (GLMamba) module effectively fuses the Visual State Space Model (VSS) and CNN, embedding them into a U-shaped structure to model the entire volume features at different scales, allowing features to better connect with context and achieve true preservation. This multi-scale symmetric architecture learning-based 3D spine model reconstruction method employs a novel dynamic discrimination strategy, simultaneously discriminating between global and local information without adding an additional decision maker, effectively increasing the detail of the generated 3D spine model. Attached Figure Description

[0017] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0018] Fig. 1 This is a schematic diagram of the framework of the generative adversarial network model according to an embodiment of the present invention.

[0019] Fig. 2 This is an architecture diagram of the Global-Local Mamba (GLMamba) module in an embodiment of the present invention.

[0020] Fig. 3 This is a diagram of the parallel Mamba-CNN hybrid (MCP) module architecture according to an embodiment of the present invention. Detailed Implementation

[0021] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0022] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0023] Reference Figs. 1 to 3 As shown, this invention provides a method for reconstructing a 3D model of the spine based on multi-scale symmetric architecture learning, comprising the following steps:

[0024] S1. Rotate multiple orthogonal X-ray images of the spine to a set angle where the X-ray images are taken from the same angle.

[0025] Two orthogonal X-ray images of the spine (typically anteroposterior and lateral views) are rotated to their respective imaging angles. Because X-ray images are acquired from different angles, rotation is necessary to align them for subsequent processing. This operation ensures that the X-ray images are consistent with their actual imaging angles to accurately reflect the actual condition of the spine.

[0026] The initial angle of the two unrotated orthogonal X-ray images is 0°. Since different X-ray images are taken from different perspectives and have their corresponding geometric and spatial physical meanings, they need to be rotated to the corresponding geometric angles before further processing.

[0027] S2. Repeat the rotated X-ray image D times and stitch the repeated X-ray images to form a high-channel 3D volume, so that the target depth of the high-channel 3D volume is adapted to the target depth of the 3D model of the spine.

[0028] Create a high-channel 3D volume that matches the target depth of the 3D model, re-establishing a symmetric architecture for the 3D model. For each 2D X-ray view, repeat D times to form a 3D data block, where D is the depth dimension of the 3D volume. Stitch these repeated views along the channel dimension to form a high-channel volume that extends into 3D space.

[0029] Specifically, the input projection for each 2D X-ray view is I. i ∈R C×H×W Where C is the number of channels, H is the height, and W is the width, each view is repeated D times to match the depth of the output 3D target.

[0030] D represents depth, which is the third dimension of a 3D volume, mathematically represented as: Get I i ∈R C ×H×W×D .

[0031] Then connect them in the channel dimension to form a higher-channel 3D volume V. cat ∈R NC×H×W×D Each 2D input view is expanded into 3D space and then input into the network to form a mapping relationship with the corresponding target 3D volume, thus forming a symmetrical architecture.

[0032] S3. Input the high-channel 3D volume into the generative adversarial network (GAN). The generator of the GAN performs preliminary feature extraction through the encoder's 3D convolution. The encoder's Global-Local Mamba (GLMamba) block performs deep feature learning to extract and recover high-resolution features. The generator's decoder concatenates the extracted and recovered high-resolution features and fuses the final features into continuous values ​​through the decoder's convolution block to obtain a 3D prediction model of the spine.

[0033] The generator uses a U-shaped network (UNet) as its basic architecture. Specifically, it employs a U-shaped network containing an initialization layer and multiple Global-Local Mamba (GLMamba) blocks. Initial feature extraction is performed through 3D convolutions, followed by deep feature learning through GLMamba blocks. At the end of each stage, downsampling layers reduce spatial resolution and increase feature dimensionality to extract higher-level feature representations. The decoding part concatenates the features extracted by the encoder with the recovered high-resolution features, integrating spatial and semantic information, and finally obtains continuous value predictions through 1×1×1 convolutional blocks.

[0034] In this embodiment, the 2D projection I is to be i The pseudo-3d volume V formed by D-fold overlap connection cat It can be directly mapped to a real 3D pattern V. However, to achieve a high-quality conversion, the key is to make full use of volume data to capture the feature information of three-dimensional space.

[0035] The generator uses unet as its basic architecture, and the encoder consists of an initialization layer and four global-local mamba (GLMamba) blocks.

[0036] Given an input image V cat ∈R NC×H×W×D V cat This is pseudo-3D data obtained during preprocessing. The initialization layer is the first step in the entire network, using 3D convolutional pairs with a kernel size of 7×7×7 and a stride of 2×2×2. cat Preliminary extraction is performed to obtain shallow feature maps.

[0037] F0 is then fed into a Global-Local Mamba (GLMamba) block for four stages of deep representation learning. Simultaneously, at the end of each stage, a downsampling layer is used to downsample the data feature dimension from dim[l] to dim[l+1], achieving a doubling of the dimension and reducing the spatial resolution to half of the original. Where l∈{1, 2, 3, 4}. This reduces the number of tokens while maintaining the feature dimensionality, thus reducing computational complexity while extracting higher-level feature representations. In this way, the encoder can learn feature representations at different scales at each stage, thereby providing richer contextual information and multi-scale perception capabilities.

[0038] For the decoding part, the features extracted by the Global-Local Mamba (GLMamba) blocks in each layer of the encoder are fed into a residual block consisting of two 3×3×3 convolutional layers. These further extracted and fused features are then concatenated along the channel dimension with the features recovered at twice the resolution in the previous stage, and then fed into the residual block together to integrate high-resolution spatial information with low-resolution semantic information, preparing for the next stage of feature recovery. Finally, a 1×1×1 convolutional block fuses the final features into continuous values ​​to obtain the final prediction result.

[0039] In this embodiment, Global-Local Mamba (GLMamba) is used to enhance the efficient fusion of convolutional blocks (CNN) and visual state space models (VSS) for the extraction of deep long-range and local semantic features of images.

[0040] Specifically, advanced residual connections are used to enhance the long-range spatial modeling capabilities of the proposed parallel Mamba-convolution (MCP) blocks. This introduces almost no new parameters or computational complexity. Next, the Global-Local Mamba (GLMamba) block uses another normalization layer (LayerNorm) to normalize. Subsequently, MLP was used to optimize the feature mapping process from input to output and integrate these features into more abstract and useful information.

[0041] For the l-th Global-Local Mamba (GLMamba) block in the above steps, the calculation process can be defined as follows: Where MCP stands for the proposed parallel Mamba-CNN hybrid block (MCP block), LN represents layer normalization, and MLP represents multilayer perceptual layer.

[0042] In this embodiment, the efficient parallel Mamba-CNN Hybrid Block (MCP block) uses the Visual State Space Model (VSS) to demonstrate Mamba's potential for long sequence modeling in image processing.

[0043] Assume the input tensor is

[0044] First, it is input into a 1×1×1 convolutional layer with C output channels for preliminary feature extraction.

[0045] Then the tensor is uniformly divided into two tensors f.CNN and f VSS Both are of the following sizes This operation has two advantages. First, halving the number of feature channels input to the CNN and Visual State Space Model (VSS) modules reduces model complexity. Second, local and non-local features can be processed independently and in parallel, which is beneficial for better feature extraction.

[0046] Then, tensor f CNN Send to the residual network to get At the same time, f VSS Send to the Visual State Space Model (VSS) module to obtain Then, connect and To restore it to a tensor of the same size as the input. And a 1×1×1 convolutional block is used to fuse local and non-local features.

[0047] Finally, a skip connection is established between f and the output to obtain the finally learned fused feature f.

[0048] The calculation process for the above steps is shown below:

[0049] f CNN ,f VSS =Split(Conv 1×1×1 (f)) (1);

[0050]

[0051] Where Split represents the splitting operation, Conv 1×1×1 This represents a 1×1×1 3D convolutional layer. Represents tensor cascade operation, Res represents residual network, and VSS represents visual state space module.

[0052] S4. Calculate the standard deviation of different generated images of the 3D prediction model to evaluate the quality difference of the generated images, and select the generator to globally optimize or locally optimize the 3D prediction model based on the quality difference.

[0053] By evaluating the quality differences in the generated images, the optimization strategy is dynamically adjusted to improve the reconstruction quality. The standard deviation between images is used to assess quality differences, and either global or local optimization is selected.

[0054] The generator-global optimized 3D prediction model includes an optimal transfer learning method using dual regularization and geometrically debiased Sinkhorn divergence as the cost function.

[0055] The generator locally optimizes the 3D prediction model by constructing a mask matrix to mark low-quality regions and performing different levels of optimization based on the value of the local standard deviation.

[0056] Specifically, since the quality differences between images can reflect the overall feature distribution differences between certain images and the real images, the standard deviation σ between different images is used to select different optimization modes. The specific calculation is as follows:

[0057]

[0058] Where, μ k The mean value among different receptive fields within an image, μ = the mean quality value of all images.

[0059] If σ is high, it indicates that there is a large range of quality differences between the images, so global optimization should be performed because images in the generative network are learned from coarse to fine. Conversely, if the generated images are relatively similar in terms of receptive field and the overall performance is relatively uniform, then local optimization should be focused on, and the feature performance of each image in local areas should be adjusted to make them more consistent.

[0060] Therefore, for global optimization, the process of directly pursuing model optimization is as follows:

[0061]

[0062] Among them, S ε It is an optimal transfer learning (OT) method that applies double regularization and uses geometrically debiased Sinkhorn divergence as the cost function to measure the difference between the generator's predicted features and the true features; G is the generator, D is the discriminator, f is the operation function on the discriminator output (using the hinge loss function), and λ is a parameter that controls the degree of regularization.

[0063] For local optimization, a mask matrix h of the same size as the discriminator output is constructed. * This is used to mark low-quality areas. Compared to the quality differences between images, the quality differences between different receptive fields within the entire image can better measure local differences. If the quality differences between different receptive fields in the same image are large, it means that some local areas in the image are generated well, while other areas may have defects. Therefore, local optimization is divided into three levels α based on the value of the local standard deviation σ. Ⅰ ,α Ⅱ ,α Ⅲ This corresponds to the region where the mask is generated. The standard deviation σ is calculated as follows:

[0064]

[0065] The higher the level (Level III), the larger the defined local optimization range (and the larger the α constant). Therefore, when the output value is lower than α, the corresponding value in the mask matrix is ​​1, and vice versa. Thus, if σ is small, the optimization process transforms into the following formula:

[0066]

[0067] in Let H represent the dot product, where H = {h1, h2, ..., h}. n ,…} is a set of mask matrices, which is equivalent to the generator finding the optimal low-quality region h during gradient descent. * , and h * It can clearly define the direction of generator optimization and attempt to adjust its parameters so that the discriminator output in these low-quality regions (after operations such as with the mask matrix) can change in a direction that is closer to the real image, thereby reducing the degree of low quality in these regions and ultimately achieving the goal of reducing loss.

[0068] The S5 3D prediction model obtains a reconstructed model of the spine by minimizing the generator's loss.

[0069] Based on feedback from the discriminator, adjust the generator's parameters. Use a loss function to measure the difference between the generator's predicted features and the true features, and optimize these parameters to reduce the discrepancy.

[0070] Specifically, the loss of the discriminator is calculated as follows:

[0071]

[0072] This invention presents a 3D spine model reconstruction method based on multi-scale symmetric architecture learning. By fusing a Mamba symmetric reconstruction network, it effectively preserves the information content of the input 2D x-ray without relying on other prior information. First, the 2D view is preprocessed by repeating the x-ray to achieve the same depth as the matched 3D model, followed by coarse registration and connection to establish a 3D-to-3D symmetric architecture mapping reconstruction framework. Next, the input is fed into a designed generative adversarial network for same-dimensional translation. For the generator, a designed Global-Local Mamba (GLMamba) block is used to effectively fuse the Visual State Space Model (VSS) and CNN, embedding them into a U-shaped structure to model the entire volume features at different scales, allowing features to better connect with context and thus be more realistically preserved. This invention's 3D spine model reconstruction method based on multi-scale symmetric architecture learning employs a novel dynamic discrimination strategy, simultaneously discriminating between global and local information without adding an additional decision maker, effectively increasing the detail of the generated 3D spine model.

[0073] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A method for spinal 3D model reconstruction based on multi-scale symmetric architecture learning, characterized in that, The method comprises the following steps: rotating a plurality of orthogonal X-ray images of the spine to a consistent set angle of the shooting angle of the X-ray images; repeating the rotated X-ray images D times and splicing the repeated X-ray images to form a high-channel 3D volume, so that the target depth of the high-channel 3D volume is adapted to the target depth of the 3D model of the spine; inputting the high-channel 3D volume into a generative adversarial network, performing preliminary feature extraction through the encoder 3D convolution of the generator of the generative adversarial network, performing deep feature learning through the GLMamba block of the encoder to extract features and restored high-resolution features, splicing the extracted features and restored high-resolution features through the decoder of the generator, and fusing the final features into continuous values through the convolution block of the decoder to obtain a 3D prediction model of the spine; calculating the standard deviation of different generated images of the 3D prediction model to evaluate the quality difference of the generated images, and selecting the 3D prediction model for global optimization or local optimization of the generator based on the quality difference; the 3D prediction model obtains a reconstructed model of the spine by minimizing the loss of the generator; using unet as the basic framework of the generator, the encoder is composed of an initialization layer and four GLMamba blocks, and the calculation process of the lth GLMamba block is defined as: where MCP denotes the proposed parallel Mamba-CNN hybrid block, LN denotes layer normalization, and MLP denotes a multi-layer perception layer; denotes the remote spatial modeling capability output of the parallel Mamba-CNN hybrid block; the parallel Mamba-CNN hybrid block step is: Assume the input tensor is , it is first input to a 1 × 1 × 1 convolutional layer with the same number of output channels C for preliminary feature extraction; then the tensor is evenly divided into two tensors f CNN and f VSS , both of which have a size of ; after that, the tensor f CNN is sent to the residual network to obtain , and f VSS is sent to the visual state space model module to obtain , then and are connected to restore the tensor to the same size as the input and pass through a 1 × 1 × 1 convolutional block to fuse local and non-local features; finally, a skip connection is established between f and the output to obtain the final learned fusion features .

2. The method of claim 1, wherein the method is based on a multi-scale symmetric architecture learning for spinal 3D model reconstruction. the global optimization of the 3D prediction model by the generator includes using the optimal transport learning method with double regularization and the geometrically unbiased Sinkhorn divergence as the cost function.

3. The method of claim 1, wherein the method is based on a multi-scale symmetric architecture learning for spinal 3D model reconstruction. the local optimization of the 3D prediction model by the generator includes constructing a mask matrix to mark low-quality areas and performing different levels of optimization according to the value of the local standard deviation.