Aorta true and false cavity segmentation and three-dimensional reconstruction method and device based on Mamba-UNet
Through the multi-scale feature extraction and fusion of the Mamba-UNet architecture, the problems of high computational complexity and low accuracy in the true and false cavity segmentation of aortic dissection are solved, and efficient and accurate segmentation and three-dimensional reconstruction are achieved to assist clinical decision-making.
Patent Information
- Application Number
- CN202510601794.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-01
AI Technical Summary
The existing medical imaging segmentation technology, especially the real and false cavity segmentation method in the aortic dissection, cannot ensure high accuracy while reducing calculation overhead, resulting in inaccurate segmentation results, affecting clinical decision-making, and unable to effectively deal with complex anatomical structures and noise interference, increasing patient risks.
Using a method based on Mamba-UNet architecture, through multi-scale feature extraction and fusion, combining the linear timing modeling of the Mamba module and the local detail capture capability of U-Net, precise segmentation of the aortic true and false cavity is carried out, and intuitive models are generated through three-dimensional reconstruction to assist clinical decision-making.
It realizes high-precision true and false aortic cavity segmentation, improves segmentation accuracy and calculation efficiency, provides reliable clinical support, and reduces the error in reverse tear risk assessment and treatment decision making.
Smart Images

Figure CN120411523A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to a method and device for aortic true and false lumen segmentation and three-dimensional reconstruction based on Mamba-UNet. Background Art
[0002] Aortic dissection is a serious cardiovascular disease. Usually, due to the tearing of the aortic intima, blood enters the space between the intima and the adventitia of the aorta, forming true and false lumens. Timely diagnosis and treatment of aortic dissection are crucial for the survival rate of patients. For the diagnosis of aortic dissection, medical imaging techniques, especially computer-aided image analysis methods, provide important support. Among these methods, the segmentation of the aortic true and false lumens is a key technical link, which helps doctors accurately identify the location of intimal tear, the boundary of the true and false lumens, and its impact on the patient's health, providing a decision-making basis for the formulation of subsequent treatment plans. However, despite the fact that existing imaging techniques can provide high-resolution three-dimensional images, the accurate segmentation of the true and false lumens remains a major challenge in current medical image analysis.
[0003] Currently, most traditional medical image segmentation methods rely on convolutional neural networks (CNNs) or the classic U-Net architecture, and these methods have achieved remarkable results in most medical image tasks. However, when applied to the field of the aorta with high anatomical variability and complex pathological structures, traditional methods have shown obvious limitations. The anatomical structure of the aorta is extremely complex. Especially in the task of true and false lumen segmentation, the morphology, size, and extension direction of intimal tears vary greatly due to individual differences among patients. Existing segmentation models often can only handle local features and lack effective global information fusion, resulting in poor segmentation accuracy and consistency. Especially at the edges of the blood vessel wall and the junction of the true and false lumens, the segmentation results are often inaccurate and cannot effectively reflect the true situation of complex lesions.
[0004] In addition, with the rapid growth of medical image data volume, especially in the application of three-dimensional medical images (such as CTA, MRA), traditional CNN- and U-Net-based models often face huge inference latency problems due to high computational complexity and large memory consumption. This problem is particularly prominent in clinical emergencies and interventional surgeries. There is an urgent need for a new technology that can significantly reduce the computational overhead and improve the inference speed while ensuring high-precision segmentation, so as to provide timely and accurate decision support for clinicians in a short time. Especially for a disease like aortic dissection that progresses rapidly and is fatal, any delay may lead to catastrophic consequences. The deficiencies of traditional methods in such high-pressure situations can no longer meet the current clinical needs.
[0005] More seriously, when faced with the complex anatomical structure of the true and false lumens of the aorta, existing technologies often fail to accurately segment the boundaries between the true and false lumens. Especially in complex vascular networks, the boundaries between the false lumen and the true lumen are often blurred. Traditional segmentation methods are easily interfered by factors such as noise and artifacts, resulting in misjudgment or even missed judgment. These inaccurate segmentation results will directly affect clinical decisions and may even lead to incorrect treatment choices, further increasing the risks for patients. In the treatment of aortic dissection, the accuracy of segmentation not only affects the identification of the lesion area but also directly impacts key treatment steps such as subsequent retrograde dissection risk assessment, tear location, anchor zone selection, etc. If precise segmentation and efficient processing cannot be achieved, the patient's life safety will be greatly threatened.
[0006] With the continuous increase in medical imaging data, traditional image segmentation methods can no longer adapt to this rapidly developing trend. For three-dimensional data, especially in situations where real-time analysis is required and support for clinical decisions is needed, the computational complexity of traditional methods often increases exponentially. This not only increases the burden on equipment but also delays the speed of clinical decision-making and reduces medical efficiency. Against this background, there is an urgent need for a new and efficient deep learning model that can reduce computational overhead and improve inference speed while ensuring segmentation accuracy, so as to address the challenges in modern medical image analysis.
[0007] Therefore, current medical image segmentation technologies, especially the segmentation of the true and false lumens in aortic dissection, can no longer meet the urgent clinical needs for real-time performance and high precision. If these key problems cannot be solved, it will greatly restrict the diagnosis and treatment efficiency of acute cardiovascular diseases such as aortic dissection and further increase the life risks of patients. In this situation, there is an urgent need for a breakthrough technology to improve segmentation accuracy and computational efficiency, provide more reliable decision support for clinical practice, and ultimately achieve more precise and timely medical intervention. Summary of the Invention
[0008] The object of the present invention is to provide a method and device for segmenting and three-dimensionally reconstructing the true and false lumens of the aorta based on Mamba-UNet, which can accurately identify the true and false lumens, generate a three-dimensional model with high precision, and assist doctors in performing retrograde dissection risk assessment, tear location, anchor zone selection, and visceral artery area identification, so as to provide reliable support for clinical decisions.
[0009] To solve the above problems, the technical solution of the present invention is as follows: A method for segmenting and three-dimensionally reconstructing the true and false lumens of the aorta based on the Mamba-UNet architecture, comprising: Obtain the three-dimensional medical image data of the patient's aorta, and preprocess the medical image data, including denoising, normalization, and data augmentation; among them, denoising uses a filtering algorithm, normalization is used to eliminate image differences generated under different scanning devices or scanning conditions, and data augmentation includes rotation, translation, and mirror processing; Input the preprocessed data into the Mamba-UNet model, and the Mamba-UNet model includes: A multi-scale feature extraction module for extracting multi-scale feature information of the aortic image; A feature fusion module for performing spatial interaction within the modality and inter-modal interaction on the extracted feature information to obtain multi-layer fusion features; A segmentation head module for segmenting the true and false cavities of the aorta according to the multi-layer fusion features; Both the multi-scale feature extraction module and the feature fusion module include a convolutional operation layer and multiple Mamba feature extraction layers, forming a specific structure for multi-scale feature extraction and fusion; Based on the segmentation result of the true and false cavities of the aorta output by the Mamba-UNet model, use a three-dimensional reconstruction method to convert the segmentation result into an intuitive three-dimensional model.
[0010] According to an embodiment of the present invention, the multi-scale feature extraction further includes: In the encoder part: Perform a convolutional operation on the input image to extract an initial feature map; Through multiple Mamba blocks, gradually extract features of different scales from the initial feature map; each Mamba block performs: Perform a normalization process on the input feature map; Map the normalized feature map into different feature representations; Perform a flattening operation on the feature map in multiple directions, including horizontal forward, horizontal backward, vertical forward, and vertical backward, to generate one-dimensional sequences, and process these sequences through a state space model, and then combine them to obtain a new feature representation; Perform a dot product operation on the processed features, and add them to the input features through a linear operation of the convolutional layer to obtain the final output features; In the decoder part: Upsample the high-level features extracted by the encoder to restore the spatial resolution of the feature map; Splice the feature maps of the corresponding levels in the encoder and the decoder through skip connections to retain local detail information; Perform a convolutional operation on the spliced feature map to further fuse multi-scale features and improve the segmentation accuracy.
[0011] According to an embodiment of the present invention, in the Mamba-UNet model, the Mamba module is embedded in the encoder part of the U-Net to enhance the ability to model long-range dependencies; The multi-scale features extracted in the encoder are transmitted to the decoder through skip connections, and the decoder fuses the multi-scale features from different levels, taking into account both the global semantic information and local detail information of the aorta, so as to facilitate accurate boundary detection and detail restoration.
[0012] According to an embodiment of the present invention, the splicing of the feature maps corresponding to the same level in the encoder and the feature maps in the decoder through skip connections further includes: Transmit the feature maps extracted from each layer of the encoder to the decoder, so that at each layer of the decoder, the feature maps corresponding to the corresponding level are obtained from the encoder; Align the feature maps of the encoder and the decoder spatially to adjust the feature maps of the decoder to the same spatial resolution as the encoder feature maps; Splice the aligned encoder feature maps and decoder feature maps in the channel dimension; Further process the spliced feature maps through a convolutional layer to fuse the feature information from the encoder and the decoder and generate a new feature representation.
[0013] According to an embodiment of the present invention, when performing feature fusion, the high-level features containing the global semantic information of the aorta extracted in the encoder are combined with the low-level features containing the edges of the vessel wall and the position of the intimal tear in the decoder, so that the model can more comprehensively understand the anatomical structure of the aorta.
[0014] According to an embodiment of the present invention, based on the fusion of multi-scale features, the boundary of the aorta is detected, and the junction of the true lumen and the false lumen is identified to improve the accuracy of true and false lumen segmentation; At the same time, based on the local detail information in the multi-scale features, the branch vessels and the position of the intimal tear in the aorta are restored.
[0015] According to an embodiment of the present invention, when processing the aortic arch, ascending aorta, descending aorta and their branches in the aorta, based on the spliced feature maps, identification and segmentation are performed to prevent omissions.
[0016] According to an embodiment of the present invention, the Dice coefficient is used to evaluate the segmentation result of the true and false lumen of the aorta.
[0017] According to an embodiment of the present invention, the risk of retrograde dissection is evaluated by calculating the contact area and relative position between the true and false lumen and the visceral artery.
[0018] An aortic true and false lumen segmentation and three-dimensional reconstruction device based on the Mamba-UNet architecture, comprising: A preprocessing module for obtaining three-dimensional medical image data of a patient's aorta and preprocessing the medical image data; A true and false lumen segmentation module for inputting the preprocessed data into a Mamba-UNet model, where the Mamba-UNet model includes: A multi-scale feature extraction unit for extracting multi-scale feature information of the aorta image; A feature fusion unit for performing spatial interaction within the modality and inter-modal interaction on the extracted feature information to obtain multi-layer fusion features; A segmentation head unit for segmenting the true and false lumens of the aorta according to the multi-layer fusion features; Both the multi-scale feature extraction unit and the feature fusion unit include a convolutional operation layer and multiple Mamba feature extraction layers, forming a specific structure for multi-scale feature extraction and fusion; A three-dimensional reconstruction module for converting the aorta true and false lumen segmentation result output by the Mamba-UNet model into an intuitive three-dimensional model by using a three-dimensional reconstruction method.
[0019] Due to the above technical solutions, the present invention has the following advantages and positive effects compared with the prior art: In an embodiment of the present invention, a method for aortic true and false lumen segmentation and three-dimensional reconstruction based on the Mamba-UNet architecture aims at the problem that the existing segmentation of the true and false lumens in aortic dissection cannot meet the urgent clinical demand for high precision. By obtaining three-dimensional medical image data of a patient's aorta and preprocessing the medical image data; inputting the preprocessed data into the Mamba-UNet model to extract multi-scale feature information of the aorta image; performing spatial interaction within the modality and inter-modal interaction on the extracted feature information to obtain multi-layer fusion features; segmenting the true and false lumens of the aorta according to the multi-layer fusion features; and based on the aorta true and false lumen segmentation result output by the Mamba-UNet model, using a three-dimensional reconstruction method to convert the segmentation result into an intuitive three-dimensional model. This method combines the linear temporal modeling ability of the Mamba model and the local detail capture advantage of UNet, and combines the prior knowledge of medical images, can accurately identify the true and false lumens, generate a high-precision three-dimensional model, assist doctors in performing retrograde dissection risk assessment, tear orifice localization, anchor zone selection and visceral artery region identification, so as to provide reliable support for clinical decision-making. Description of the Drawings
[0020] Figure 1 is a method for aortic true and false lumen segmentation and three-dimensional reconstruction based on the Mamba-UNet architecture in an embodiment of the present invention; Figure 2Schematic diagram of aortic CTA image in an embodiment of the present invention; Figure 3 High-level feature map extracted by the encoder in an embodiment of the present invention; Figure 4 Low-level feature map extracted by the decoder in an embodiment of the present invention; Figure 5 Feature map after splicing in an embodiment of the present invention; Figure 6 3D reconstruction model diagram in an embodiment of the present invention. Detailed implementation manners
[0021] The following further describes in detail a method and device for aortic true and false lumen segmentation and 3D reconstruction based on the Mamba-UNet architecture proposed by the present invention in conjunction with the accompanying drawings and specific embodiments. The advantages and features of the present invention will be clearer according to the following description and claims.
[0022] Please refer to Figure 1 , the method for aortic true and false lumen segmentation and 3D reconstruction based on the Mamba-UNet architecture provided in this embodiment includes the following steps: S1: Obtain the three-dimensional medical image data of the patient's aorta and preprocess the medical image data; S2: Input the preprocessed data into the Mamba-UNet model, and the Mamba-UNet model includes: A multi-scale feature extraction module for extracting multi-scale feature information of the aortic image; A feature fusion module for performing spatial interaction within the modality and inter-modal interaction on the extracted feature information to obtain multi-layer fusion features; A segmentation head module for segmenting the true and false lumens of the aorta according to the multi-layer fusion features; Among them, both the multi-scale feature extraction module and the feature fusion module include convolutional operation layers and multiple Mamba feature extraction layers, forming a specific structure for multi-scale feature extraction and fusion; S3: Based on the aortic true and false lumen segmentation result output by the Mamba-UNet model, use a 3D reconstruction method to convert the segmentation result into an intuitive 3D model.
[0023] Specifically, in step S1, the three-dimensional medical image data of the patient's aorta is obtained and the medical image data is preprocessed. The three-dimensional medical image data of the aorta is derived from the patient's CT angiography (CTA) or magnetic resonance angiography (MRA), and the data should have high resolution and sufficient contrast to ensure that the aortic structure is clearly visible.
[0024] Preprocess the medical image data, including denoising, normalization, and data augmentation, to improve the effect of model training. Among them, denoising uses filtering algorithms (such as median filtering, Gaussian filtering, etc.); normalization processing can eliminate image differences generated under different scanning devices or scanning conditions; data augmentation includes methods such as rotation, translation, and mirroring, which are used to expand training samples and improve the robustness of the model.
[0025] In step S2, input the preprocessed data into the Mamba-UNet model to output the segmentation results of the true and false lumens of the aorta.
[0026] Among them, the Mamba model is based on the idea of linear time series modeling and Selective State Space (SSM), and can efficiently process long sequence data. By introducing the selective state space mechanism, Mamba can dynamically and selectively propagate and forget information without increasing the computational complexity. Its computational complexity is linear (o(n)), avoiding the O(n 2 ) computational complexity in the traditional attention mechanism, and has stronger computational efficiency and scalability.
[0027] U-Net is a classic medical image segmentation network structure. Through a symmetric encoder-decoder architecture, it can efficiently extract local features and perform accurate pixel-level segmentation. Through skip connections, U-Net can combine global context information and local details to improve the segmentation accuracy.
[0028] In this embodiment, the Mamba model is integrated with the U-Net architecture, combining the advantages of the Mamba model in sequence modeling to enhance the model's ability to capture long-distance dependence information, while retaining the U-Net's ability to accurately model local details. In Mamba-UNet, the Mamba module is integrated into the encoder part of U-Net to improve the modeling ability of the complex structure of the aorta, especially the accurate recognition of the boundary between the true and false lumens.
[0029] Specifically, the Mamba-UNet model realizes multi-scale feature extraction by combining the linear time series modeling ability of the Mamba model with the local detail capture advantage of U-Net. The model mainly includes the following modules: Mamba module: responsible for capturing long-distance dependence information and enhancing the modeling ability of complex anatomical structures; U-Net module: realizes accurate segmentation of local details through the encoder-decoder structure and skip connections.
[0030] Among them, the encoder part includes the following units: Convolutional layer: performs convolutional operations on the input image to extract the initial feature map; Mamba Feature Extraction Layer: Gradually extract features of different scales through multiple Mamba blocks; each Mamba block includes: Layer Normalization Unit: Normalize the input feature map; Convolutional Layer Linear Unit: Map the normalized feature map into different feature representations; Feature Processing Unit: Flatten the feature map in multiple directions (horizontal forward, horizontal backward, vertical forward, vertical backward) to generate one-dimensional sequences, process these sequences through a state space model, and finally combine them to obtain a new feature representation; Dot Product Unit and Output Unit: Perform a dot product operation on the processed features and add them to the input features through a convolutional layer linear operation to obtain the final output features; Downsampling Layer: Perform downsampling after some Mamba blocks to reduce the spatial resolution of the feature map and further extract high-level features.
[0031] The decoder part includes the following units: Upsampling Layer: Upsample the high-level features extracted by the encoder to restore the spatial resolution of the feature map; Skip Connection: Concatenate the feature maps of the corresponding levels in the encoder with the feature maps in the decoder to retain more local detail information; Convolutional Layer: Perform a convolution operation on the concatenated feature map to further fuse multi-scale features and improve the segmentation accuracy.
[0032] When performing feature fusion and segmentation, the feature fusion module fuses the multi-scale features extracted by the encoder and the decoder, and processes the input features using specific scanning strategies (such as spatial-first scanning and modality-first scanning) to achieve more sufficient and efficient feature aggregation of the two modalities.
[0033] The segmentation head module performs the segmentation of the true and false lumens of the aorta based on the fused multi-layer features, and realizes the fusion and segmentation of features of different scales through multiple processing units (such as convolutional layers, activation layers, interpolation layers, etc.), and finally outputs the segmentation result.
[0034] By combining the Mamba model and the U-Net architecture in this way, the Mamba-UNet model can capture feature information of different scales in the aorta true and false lumen segmentation task, while taking into account global semantic modeling and local detail segmentation, so as to achieve high-precision segmentation results and provide reliable support for subsequent 3D reconstruction and clinical decision-making.
[0035] Among them, the Mamba module can efficiently process long-sequence data and capture long-range dependency information by combining the ideas of linear temporal modeling and Selective State Space (SSM). The following are the specific ways for the Mamba module to capture long-range dependency information: The Mamba module performs flattening operations on the aortic feature map in the horizontal forward, horizontal backward, vertical forward, and vertical backward directions, converting the two-dimensional feature map into four one-dimensional sequences. This multi-directional scanning strategy can capture long-range dependency relationships in different directions, ensuring a more comprehensive understanding of complex anatomical structures by the model. Each one-dimensional sequence is processed by an independent state space model, which can capture the long-range dependency relationships between elements in the sequence. In this way, the Mamba module can effectively model the relationships between pixels that are far apart in the aortic feature map.
[0036] The Mamba module also introduces a selective state space mechanism that can dynamically and selectively propagate and forget information without increasing computational complexity. This mechanism allows the model to decide which information is important and should be retained and which information can be ignored based on the currently processed features, thus capturing long-range dependency relationships more efficiently. Compared with traditional attention mechanisms, the computational complexity of the Mamba module is linear (O(n)), avoiding the quadratic complexity (O(n²)) in attention mechanisms. This makes the Mamba module more efficient in processing high-resolution medical images and able to quickly capture long-distance dependency relationships.
[0037] In the Mamba module, the four output features after being processed by the state space model are combined and dot-multiplied with the features after the SiLU operation. This feature combination and dot-multiplication operation can further enhance the model's ability to capture long-range dependency information, enabling the model to better understand the relationships between different regions in the feature map. The output of the Mamba module is added to the input features after linear operations through a convolutional layer to form a residual connection. This residual connection helps alleviate the vanishing gradient problem in deep networks and also helps the model better propagate long-range dependency information.
[0038] In the aortic true and false lumen segmentation task, these characteristics of the Mamba module can help the model more accurately identify and segment complex aortic anatomical structures. For example, the anatomical structure of the aorta has a large span from the ascending aorta to the descending aorta, with long-range structural dependency relationships. By capturing these long-range dependency relationships, the Mamba module can more accurately identify the location and extent of aortic intimal tears, thereby improving the accuracy of true and false lumen segmentation.
[0039] In summary, through technical means such as multi-directional sequence modeling, selective state space mechanism, feature fusion and enhancement, the Mamba module effectively captures long-range dependence information and provides powerful modeling capabilities for the segmentation task of complex medical images.
[0040] Furthermore, in the Mamba-UNet model, the decoder stitches together the high-level features extracted in the encoder and the low-level features in the decoder through skip connections to retain more local detailed information. The following is the specific process of how the decoder stitches feature maps through skip connections: The U-Net architecture consists of two parts: an encoder and a decoder. The encoder gradually extracts high-level features of the image through convolution and pooling operations, while the decoder gradually restores these high-level features to the resolution of the original image through upsampling and convolution operations.
[0041] Skip connections are a key feature of U-Net. They directly connect the feature map of a certain layer in the encoder to the corresponding layer's feature map in the decoder. This can transfer the local detailed information extracted in the encoder to the decoder and avoid losing these details during the downsampling process.
[0042] At each layer of the encoder, the extracted feature map is saved. At each layer of the decoder, the corresponding layer's feature map is obtained from the encoder.
[0043] Since the feature maps of the encoder and decoder may have different spatial resolutions, spatial alignment is required. Usually, the feature map of the decoder is adjusted to the same spatial resolution as the encoder's feature map through upsampling (such as bilinear interpolation or transposed convolution).
[0044] The aligned encoder and decoder feature maps are stitched together in the channel dimension. For example, if the number of channels of the encoder feature map is C and the number of channels of the decoder feature map is also C, then the number of channels of the stitched feature map will be 2C.
[0045] The stitched feature map is further processed through a convolutional layer to fuse the feature information from the encoder and decoder and generate a new feature representation.
[0046] In the Mamba-UNet model, the Mamba module is embedded in the encoder part of the U-Net to enhance the ability to model long-range dependencies. Skip connections transfer the multi-scale features extracted in the encoder to the decoder, ensuring that the decoder can make full use of these features for accurate segmentation. Through skip connections, the decoder can fuse multi-scale features from different levels, thus taking into account both global semantic information and local detail information during the segmentation process. This is particularly important in the task of aortic true and false lumen segmentation because the anatomical structure of the aorta is complex and requires precise boundary detection and detail restoration.
[0047] The following is a specific example to illustrate the effect after the concatenation of encoder and decoder feature maps, especially in the application of aortic true and false lumen segmentation.
[0048] Obtain a CT angiography (CTA) image of the aorta with a resolution of 512×512×200 and a voxel size of 0.5 mm × 0.5 mm × 1 mm. This image contains the complex anatomical structure of the aorta, including the true lumen and the false lumen.
[0049] The encoder part gradually extracts high-level features of the image through multiple layers of convolution and pooling operations. The shape of the feature map Fenc output by a certain layer of the encoder is 64×64×128 (height × width × number of channels).
[0050] The decoder part gradually restores the spatial resolution of the feature map through upsampling and convolution operations. The shape of the feature map Fdec output by a certain layer of the decoder is 32×32×128.
[0051] At this layer of the decoder, the feature map Fenc of the encoder is concatenated with the feature map Fdec of the decoder through a skip connection.
[0052] First, the feature map Fdec of the decoder is adjusted to the same spatial resolution as the encoder feature map Fenc, i.e., 64×64×128, through upsampling (such as bilinear interpolation).
[0053] The aligned encoder feature map and decoder feature map are concatenated in the channel dimension to obtain a new feature map Fconcat with a shape of 64×64×256.
[0054] The concatenated feature map Fconcat is further processed through a convolutional layer to fuse the feature information from the encoder and the decoder, generating a new feature representation Fnew.
[0055] Finally, the decoder generates a segmentation result through multiple layers of convolution and activation functions (such as Sigmoid or Softmax), and outputs the segmentation mask of the aortic true and false lumen.
[0056] Effect after splicing Through skip connections, the decoder can utilize the local detailed information extracted by the encoder, thus more accurately restoring the fine structures of the aorta in the segmentation results, such as small branch vessels and the precise location of intimal tears.
[0057] The spliced feature map integrates multi-scale information and can more accurately detect the boundary of the aorta, especially at the junction of the true lumen and the false lumen. This helps improve the accuracy of the segmentation results and avoid blurred or discontinuous boundaries.
[0058] When dealing with complex aortic anatomical structures (such as the aortic arch, ascending aorta, descending aorta and their branches), the spliced feature map can provide more comprehensive context information to help the model better understand the relationships between these structures, thus improving the accuracy of segmentation.
[0059] The accuracy of the segmentation results is evaluated by metrics such as the Dice coefficient. For example, in the aortic true and false lumen segmentation task using the Mamba-UNet model with skip connections, the Dice coefficient reached 0.92, showing excellent segmentation accuracy. Especially when dealing with complex structures such as the aortic arch and ascending aorta, the segmentation results are more accurate.
[0060] Through this method of feature map splicing, the Mamba-UNet model can provide high-precision aortic true and false lumen segmentation results for clinicians to assist in surgical planning, lesion analysis and retrograde dissection risk assessment. Doctors can use these results to more accurately judge the distribution of aortic dissection, select appropriate stent anchoring areas, evaluate the compression of visceral arteries, and thus formulate more reasonable treatment plans.
[0061] In the Mamba-UNet model, after splicing the feature maps, the model can better identify the true and false lumens of the aorta, mainly reflected in the following aspects: The spliced feature map combines the high-level features (global semantic information) extracted by the encoder with the low-level features (local detailed information) in the decoder. The fusion of this multi-scale information enables the model to utilize both global and local features simultaneously, thus more comprehensively understanding the aortic anatomical structure. Among them, the high-level feature map in the encoder provides the overall structure and semantic information of the aorta, helping the model understand the main branches and general outline of the aorta. The low-level feature map in the decoder retains more local details, such as the edges of the vessel wall and the precise location of intimal tears.
[0062] Moreover, the spliced feature map can more accurately detect the boundary of the aorta, especially at the junction of the true lumen and the false lumen. This helps to improve the accuracy of the segmentation results and avoid blurred or discontinuous boundaries. Among them, the fusion of multi-scale features enables the model to more accurately identify the location of aortic intimal tears, thereby more precisely segmenting the true lumen and the false lumen. The retention of local detailed information enables the model to recover the fine structures of the aorta, such as small branch vessels and the precise location of intimal tears.
[0063] When dealing with complex aortic anatomical structures (such as the aortic arch, ascending aorta, descending aorta and their branches), the spliced feature map can provide more comprehensive context information, helping the model to better understand the relationships between these structures, thereby improving the accuracy of segmentation. Among them, the fusion of multi-scale features enables the model to better understand the complex aortic anatomical structure and avoid missing important details during the segmentation process. The combination of global semantic information and local detailed information enables the model to more accurately identify and segment complex aortic structures.
[0064] Through more accurate true and false lumen segmentation, the model can better evaluate the risk of retrograde tear. Retrograde tear refers to the further tearing of the aortic intima during stent implantation due to the operation of guide wires, catheters or stents, and even extending proximally. Accurate true and false lumen segmentation helps to identify the location and extent of intimal tears, thereby providing an important basis for retrograde tear risk assessment.
[0065] In step S3, based on the aortic true and false lumen segmentation results output by the Mamba-UNet model, a three-dimensional reconstruction method is used to convert the segmentation results into an intuitive three-dimensional model. Three-dimensional reconstruction is the process of converting medical image data (such as CT or MRA) from two-dimensional slices into a three-dimensional model, usually using voxel data and surface reconstruction algorithms to achieve. After the aortic true and false lumen is segmented, generating a three-dimensional model helps clinicians better understand information such as the aortic anatomical structure and the expansion of the lesion area.
[0066] Voxel data: Medical images are usually obtained by three-dimensional scanners (such as CT or MRA). Each slice can be regarded as a two-dimensional image, and the entire dataset consists of multiple slices (or voxels). Each voxel represents a small unit in the image space and has a certain volume (usually a cube).
[0067] In this embodiment, the Marching Cubes algorithm is used to perform three-dimensional reconstruction on the aortic true and false lumen segmentation results. The Marching Cubes algorithm scans the voxel grid one by one, judges the surface between voxels according to the values of the voxels (such as 0 and 1 of the segmentation results), and uses interpolation techniques to insert points at the boundaries, finally forming an approximate three-dimensional surface. This algorithm is particularly suitable for three-dimensional reconstruction in medical images and can generate a three-dimensional grid with high precision.
[0068] The advantage of this algorithm is that it can efficiently generate a smooth three-dimensional surface and is applicable to the three-dimensional reconstruction of various medical images, including complex structures such as blood vessels and organs. For the true and false lumen segmentation of the aorta, the Marching Cubes algorithm can accurately convert the boundaries of the segmented inner and outer lumens into a three-dimensional model, helping doctors observe the morphology of the aorta from different angles.
[0069] The model generated through three-dimensional reconstruction can be visualized and further analyzed using specialized software (such as 3D Slicer, OsiriX, etc.). Doctors can use the three-dimensional model to accurately observe the anatomical structure of the aorta, evaluate the location and extent of lesions, and even perform virtual surgical planning and inverse tear risk assessment.
[0070] The following uses multiple embodiments to demonstrate different application scenarios of the above-mentioned method for true and false lumen segmentation and three-dimensional reconstruction of the aorta based on the Mamba-UNet architecture.
[0071] First Embodiment True and False Lumen Segmentation and Three-Dimensional Reconstruction of the Aorta Use the CT angiography (CTA) data of a certain hospital, which contains three-dimensional image data of 200 patients with aortic dissection. The resolution of each dataset is 512×512×200, and the size of each voxel is 0.5 mm × 0.5 mm × 1 mm.
[0072] Preprocess the three-dimensional image data: Denoising: Use Gaussian filtering (σ = 1.0) to remove image noise; Normalization: Normalize the image gray values to between 0 and 1 to eliminate differences under different scanning devices and scanning conditions; Data augmentation: Perform data augmentation processing on the training data, including rotation (angle ±15°), translation (±5 mm), scaling (±5%), and mirroring.
[0073] Train the Mamba-UNet model based on the preprocessed data: Loss function: Use a weighted cross-entropy loss function, and the weights are adjusted according to the distribution of different categories (true lumen and false lumen) to address the problem of data imbalance; Optimizer: Use the Adam optimizer, with a learning rate of 0.0001, a batch size of 4, and train for 50 epochs; Hardware environment: Train on an NVIDIA Tesla V100 GPU, and the training time is approximately 5 hours.
[0074] Evaluate the segmentation accuracy of the trained Mamba-UNet model: Dice coefficient: The Dice coefficient for segmenting the true and false lumens is 0.92, showing excellent segmentation accuracy, especially when dealing with complex structures such as the aortic arch and ascending aorta; Boundary detection: Through precise boundary detection, the model successfully identified the location of the aortic intimal tear and accurately evaluated the extent of the retrograde tear area.
[0075] Specifically, a CT angiography (CTA) image of the aorta is input into the Mamba-UNet model. The resolution of this image is 512×512×200, and the voxel size is 0.5 mm × 0.5 mm × 1 mm. Please refer to Figure 2 , this image contains the complex anatomical structure of the aorta, including the true lumen and false lumen. The Mamba-UNet model performs the following steps: 1. Encoder feature extraction The encoder part gradually extracts high-level features of the image through multi-layer convolution and pooling operations. Assume that the shape of the output feature map of a certain layer of the encoder is 64×64×128 (height × width × number of channels). Please refer to Figure 3 , the high-level feature map extracted by the encoder shows the main structure and some details of the aorta.
[0076] 2. Decoder feature extraction The decoder part gradually restores the spatial resolution of the feature map through upsampling and convolution operations. Assume that the shape of the output feature map of a certain layer of the decoder is 32×32×128. Please refer to Figure 4 , the low-level feature map extracted by the decoder shows the preliminary segmentation result of the aorta.
[0077] 3. Feature map concatenation At this layer of the decoder, the feature map of the encoder is concatenated with the feature map of the decoder through skip connections.
[0078] Spatial alignment: First, the feature map of the decoder is adjusted to the same spatial resolution as the encoder feature map (i.e., 64×64×128) through upsampling (such as bilinear interpolation).
[0079] Channel concatenation: The aligned encoder feature map and decoder feature map are concatenated in the channel dimension to obtain a new feature map with a shape of 64×64×256.
[0080] Please refer to Figure 5 , the concatenated feature map integrates the multi-scale information of the encoder and decoder, showing richer details and more accurate boundaries.
[0081] 4. Convolution operation The concatenated feature map is further processed through a convolutional layer to fuse the feature information from the encoder and the decoder, generating a new feature representation.
[0082] 5. Segmentation Results Finally, the decoder generates the segmentation result through multiple layers of convolution and activation functions (such as Sigmoid or Softmax), outputting the segmentation masks of the true and false lumens of the aorta. The final segmentation result shows the true and false lumens of the aorta with clear boundaries and rich details.
[0083] 3D Reconstruction: The Marching Cubes algorithm is used to perform 3D reconstruction on the segmentation result, generating a high-precision 3D model. Please refer to Figure 6 , and the volume of each model is 512×512×200. Among them, the white part is the true lumen, and the red and blue parts are the false lumens. The boundary between the true and false lumens of the aorta is clearly shown in the figure. Doctors can perform operations such as rotation and slicing through the 3D visualization interface to comprehensively evaluate the anatomical structure of the aorta.
[0084] Second Embodiment: Visceral Artery Region Identification and Retrograde Tear Risk Assessment CT angiography (CTA) data of 100 patients with aortic dissection were collected. The data resolution is 256×256×100, and the voxel size is 0.6 mm × 0.6 mm × 1 mm.
[0085] Data Preprocessing: Denoising: Median filtering (3×3) is applied for denoising; Normalization: The image is normalized so that the value range of each pixel is between 0 and 1; Data Augmentation: Affine transformation is performed on the data to increase the diversity of the dataset.
[0086] Model Architecture and Training Network Architecture: The Mamba-UNet architecture is adopted, where the Mamba module is used to extract the global features of the aorta, and U-Net is used for fine segmentation of the visceral artery region (such as the superior mesenteric artery, hepatic artery, etc.); Training Process: Loss Function: The Dice coefficient loss function and the weighted cross-entropy loss function are adopted to ensure accurate segmentation of the visceral arteries; Optimizer: The Adam optimizer is used, with a learning rate of 0.0005, a batch size of 8, and 60 epochs of training; Hardware Environment: Training is performed on an NVIDIA RTX 3090 GPU, and the training time is about 4 hours.
[0087] Segmentation Results: Dice coefficient: The Dice coefficient for visceral artery segmentation is 0.88. The model can accurately segment most visceral arteries, especially in the complex aortic arch region; Boundary detection: Successfully identified the boundary between arteries and veins, especially in the complex superior mesenteric artery region.
[0088] Inverse tear risk assessment Analysis: By combining true and false lumen segmentation with visceral artery segmentation, the model can evaluate the compression of the true and false lumens on visceral arteries and the possible inverse tear risk. By calculating the contact area and relative position between the true and false lumens and visceral arteries, the model provides an inverse tear risk score to help doctors evaluate the surgical risk.
[0089] Clinical application: This method provides a quantitative basis for inverse tear risk assessment, helping doctors make more reasonable decisions during aortic dissection surgery.
[0090] The third embodiment: Aortic arch and its branch segmentation and three-dimensional reconstruction Magnetic resonance angiography (MRA) data of 300 aortic dissection patients were used. The data resolution was 512×512×200, and the voxel size was 0.6 mm × 0.6 mm × 1 mm.
[0091] Data preprocessing: Denoising: Gaussian filtering (σ = 1.2) was used for denoising to ensure smooth images with details retained; Data augmentation: The data was augmented by random cropping, rotation, etc. to increase the generalization ability of the model.
[0092] Model architecture and training Network architecture: The Mamba-UNet model adopts a dual-channel U-Net structure to simultaneously process the aortic arch region and its main branches (such as brachiocephalic artery, left common carotid artery, etc.). The Mamba module is used to extract global information, and U-Net is used for detailed segmentation of each branch region.
[0093] Training process: Loss function: The Dice coefficient loss function was combined with Focal Loss to optimize small regions that are difficult to segment (such as branches of the brachiocephalic artery); Optimizer: The Adam optimizer was used with a learning rate of 0.0002, a batch size of 4, and trained for 70 epochs; Hardware environment: Training was performed using an NVIDIA A100 GPU, and the training time was approximately 6 hours.
[0094] Segmentation results Segmentation accuracy: Dice Coefficient: The Dice coefficient for the segmentation of the aortic arch and its main branches is 0.90, especially with high precision in the segmentation of the brachiocephalic artery and the left common carotid artery; Detail Restoration: Mamba-UNet can restore details in complex anatomical regions (such as the aortic arch and its branches), successfully segmenting tiny arterial branches.
[0095] Three-dimensional Reconstruction and Clinical Application Three-dimensional Reconstruction: Based on the segmentation results, the Marching Cubes algorithm is used to convert the segmentation results into a three-dimensional model, generating a high-quality three-dimensional model of the aortic arch and its branches.
[0096] Clinical Application: The model provides three-dimensional visualization of the aortic arch and its branches for clinicians, helping to evaluate the extent of aortic dissection, the risk of arterial rupture, and guiding the selection of stent implantation and treatment plans.
[0097] Fourth Embodiment Early Screening for Aortic Vascular Diseases CTA data of 1000 high-risk populations are used, with a data resolution of 512×512×300 and a voxel size of 0.7mm × 0.7 mm × 1 mm.
[0098] Data Preprocessing: Denoising: The non-local means denoising (NLM) algorithm is used for denoising; Normalization and Enhancement: The data is normalized, and data enhancement is performed using translation, rotation, etc. to expand the training set.
[0099] Model Architecture and Training Network Architecture: Mamba-UNet structure, where the Mamba module is used to extract long-range dependency information, and U-Net is responsible for high-precision segmentation of aortic lesion regions.
[0100] Training Process: Loss Function: The Dice coefficient loss function is combined with **Focal Loss** to handle unbalanced classes; Optimizer: The Adam optimizer is used, with a learning rate of 0.0005, a batch size of 4, and training for 50 epochs.
[0101] Segmentation Results and Clinical Application Segmentation Accuracy: The Dice coefficient is 0.85, which can effectively segment relevant lesion regions of aortic diseases (such as aortic dissection and aneurysm); Screening Effect: Through this method, the risk of aortic lesions can be identified in advance, especially in the early screening stage, providing support for the early diagnosis and treatment of aortic dissection and aortic aneurysm.
[0102] The above embodiments demonstrate various applications of the method for aortic true and false lumen segmentation and three-dimensional reconstruction based on the Mamba-UNet architecture of the present invention, including detailed information and performance in aspects such as segmentation accuracy, three-dimensional reconstruction, and risk assessment. These embodiments show that the use of this method can effectively improve the accuracy of aortic image segmentation, provide accurate decision-making support for clinical practice, and is particularly significant in aspects such as aortic disease diagnosis, surgical planning, and risk assessment.
[0103] Based on the same concept, the present invention also provides an aortic true and false lumen segmentation and three-dimensional reconstruction device based on the Mamba-UNet architecture, including: A preprocessing module for obtaining three-dimensional medical image data of the patient's aorta and preprocessing the medical image data; A true and false lumen segmentation module for inputting the preprocessed data into the Mamba-UNet model for aortic true and false lumen segmentation. The Mamba-UNet model includes: A multi-scale feature extraction unit for extracting multi-scale feature information of the aortic image; A feature fusion unit for performing spatial interaction within the modality and inter-modal interaction on the extracted feature information to obtain multi-layer fusion features; A segmentation head unit for segmenting the aortic true and false lumen based on the multi-layer fusion features; Both the multi-scale feature extraction unit and the feature fusion unit include a convolutional operation layer and multiple Mamba feature extraction layers, forming a specific structure for multi-scale feature extraction and fusion; A three-dimensional reconstruction module for converting the aortic true and false lumen segmentation result output by the Mamba-UNet model into an intuitive three-dimensional model using a three-dimensional reconstruction method.
[0104] This device is used to implement the above method for aortic true and false lumen segmentation and three-dimensional reconstruction based on the Mamba-UNet architecture, and its functions and implementation methods are similar and will not be elaborated here.
[0105] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above embodiments. Even if various changes are made to the present invention, provided that these changes fall within the scope of the claims of the present invention and their equivalent technologies, they still fall within the protection scope of the present invention.
Claims
1. A method for aortic true and false lumen segmentation and three-dimensional reconstruction based on the Mamba-UNet architecture, characterized in that, include: Acquiring three-dimensional medical imaging data of the patient's aorta and preprocessing the medical imaging data; The preprocessed data is input into the Mamba-UNet model for segmentation of the true and false lumens of the aorta. The Mamba-UNet model includes: Multi-scale feature extraction module, used to extract multi-scale feature information of aortic images; The feature fusion module is used to perform spatial interaction within the modality and interaction between modalities on the extracted feature information to obtain multi-layer fusion features; a segmentation head module, configured to segment the true and false lumens of the aorta according to the multi-layer fusion features; The multi-scale feature extraction module and the feature fusion module both include a convolution operation layer and multiple Mamba feature extraction layers, forming a specific structure for performing multi-scale feature extraction and fusion; Based on the segmentation results of the true and false lumens of the aorta output by the Mamba-UNet model, a three-dimensional reconstruction method is used to convert the segmentation results into an intuitive three-dimensional model.
2. The method for aortic true and false lumen segmentation and three-dimensional reconstruction based on Mamba-UNet according to claim 1, characterized in that The multi-scale feature extraction further comprises: In the encoder section: Perform convolution operation on the input image to extract the initial feature map; Through multiple layers of Mamba blocks, features of different scales are gradually extracted from the initial feature map; each Mamba block performs: Normalize the input feature map; Map the normalized feature map into different feature representations; Perform multi-directional flattening operations on the feature map, including horizontal forward, horizontal reverse, vertical forward, and vertical reverse, to generate one-dimensional sequences. These sequences are then processed through the state space model and combined to obtain new feature representations. Perform a dot multiplication operation on the processed features and add them to the input features through a linear operation of the convolution layer to obtain the final output features; In the decoder part: Upsample the high-level features extracted by the encoder to restore the spatial resolution of the feature map; The feature maps of the corresponding layers in the encoder are concatenated with the feature maps in the decoder through skip connections to preserve local detail information. The convolution operation is performed on the spliced feature map to further fuse multi-scale features and improve segmentation accuracy.
3. The method for segmentation and 3D reconstruction of the true and false lumens of the aorta based on Mamba-UNet according to claim 1, characterized in that: In the Mamba-UNet model, the Mamba module is embedded into the encoder part of the U-Net to enhance the modeling ability of long-distance dependencies; The multi-scale features extracted from the encoder are passed to the decoder through skip connections. The decoder then fuses multi-scale features from different levels, taking into account both the global semantic information and local detail information of the aorta, to facilitate accurate boundary detection and detail recovery.
4. The method for aortic true and false lumen segmentation and three-dimensional reconstruction based on Mamba-UNet according to claim 2, characterized in that, Splicing the feature map of the corresponding level in the encoder with the feature map in the decoder through the skip connection further includes: The feature map extracted by each layer of the encoder is passed to the decoder so that at each layer of the decoder, the feature map of the corresponding level is obtained from the encoder; Spatial alignment of the encoder and decoder feature maps to adjust the decoder feature maps to the same spatial resolution as the encoder feature maps; Concatenate the aligned encoder feature map and decoder feature map in the channel dimension; The concatenated feature map is further processed through a convolutional layer to fuse the feature information from the encoder and the decoder, generating a new feature representation.
5. The method for aortic true and false lumen segmentation and three-dimensional reconstruction based on Mamba-UNet according to claim 2, wherein During feature fusion, the high-level features containing the global semantic information of the aorta extracted from the encoder are combined with the low-level features containing the edges of the vessel wall and the location of intimal tears in the decoder, enabling the model to comprehensively understand the anatomical structure of the aorta.
6. The method for aortic true and false lumen segmentation and three-dimensional reconstruction based on Mamba-UNet according to claim 3, wherein Based on the fusion of multi-scale features, the boundary of the aorta is detected, and the junction between the true lumen and the false lumen is identified to improve the accuracy of true and false lumen segmentation. Meanwhile, based on the local detail information in the multi-scale features, the branch vessels and the location of intimal tears in the aorta are restored.
7. The method for aortic true and false lumen segmentation and three-dimensional reconstruction based on Mamba-UNet according to claim 6, characterized in that, When dealing with the aortic arch, ascending aorta, descending aorta and their branches in the aorta, identification and segmentation are performed based on the concatenated feature map to prevent omissions.
8. The method for aortic true and false lumen segmentation and three-dimensional reconstruction based on Mamba-UNet according to claim 1, characterized in that The Dice coefficient is used to evaluate the segmentation results of the true and false lumens of the aorta.
9. The method for aortic true and false lumen segmentation and three-dimensional reconstruction based on Mamba-UNet according to claim 1, wherein, By calculating the contact area and relative position between the true and false lumens and the visceral arteries, the risk of retrograde dissection is evaluated.
10. An aortic true and false lumen segmentation and three-dimensional reconstruction device based on the Mamba-UNet architecture, characterized in that, Including: A preprocessing module for obtaining the three-dimensional medical image data of the patient's aorta and preprocessing the medical image data. A true and false lumen segmentation module for inputting the preprocessed data into the Mamba-UNet model for segmenting the true and false lumens of the aorta. The Mamba-UNet model includes: A multi-scale feature extraction unit for extracting multi-scale feature information of the aortic image. A feature fusion unit for performing spatial interaction within the modality and interaction between modalities on the extracted feature information to obtain multi-layer fusion features. A segmentation head unit for segmenting the true and false lumens of the aorta according to the multi-layer fusion features. Both the multi-scale feature extraction unit and the feature fusion unit include a convolutional operation layer and multiple Mamba feature extraction layers, forming a specific structure for multi-scale feature extraction and fusion. A three-dimensional reconstruction module for converting the segmentation result into an intuitive three-dimensional model using a three-dimensional reconstruction method based on the segmentation result of the true and false lumens of the aorta output by the Mamba-UNet model.