Two-stage multi-class lung crack segmentation method based on multi-view fusion
Through a two-stage method of multi-view fusion, combined with position prior clustering and lightweight network, the robustness and accuracy problems of lung crack segmentation are solved, and efficient and convenient lung crack segmentation is achieved under limited computing resources, expanding the scope of application.
Patent Information
- Application Number
- CN202510496793.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies have poor robustness and high parameter dependence in lung fissure segmentation, making it difficult to adapt to different data sets. Traditional methods and deep learning methods have failed to effectively solve the four-category task of right lung fissure segmentation, especially the morphological complexity of the right oblique fissure, which limits the segmentation accuracy.
A two-stage multi-class lung crack segmentation method based on multi-view fusion is adopted. Combined with the prior knowledge of lung cracks, the SRPNet network is used to extract single-class lung cracks from multiple perspectives, and the lightweight MFNet network is used for fusion and classification. The position prior clustering, semantic consistency self-regularization and internal feature distillation are combined to improve the segmentation accuracy.
Under limited computing resources, more accurate lung fissure segmentation is achieved, which improves the real-time performance and accuracy of the system, facilitates integration into medical imaging analysis software, and expands the scope of application.
Smart Images

Figure CN120612334A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image intelligent computing technology, and in particular to a two-stage multi-class lung crack segmentation method based on multi-view fusion. Background Art
[0002] The lungs are important respiratory organs in the human body. Each lung is divided into several lobes by fissures. From an anatomical perspective, the right lung has one oblique fissure and one horizontal fissure, which divide the right lung into three lobes. The left lung has only one oblique fissure, which divides it into two lobes. Fissures can completely or partially separate the lobes from each other. Fissures also play a significant role in clinical diagnosis. For example, "Koster TD, Slebos D J. The fissure: interlobar collateral ventilation and implications for endoscopic therapy in emphysema [J]. International journal of chronic obstructive pulmonary disease, 2016: 765-773. doi: 10.2147 / COPD.S103807" points out that during endobronchial valve treatment and screening, incomplete lobar fissures can affect the occurrence of interlobar collateral ventilation, thereby affecting the formulation of treatment plans. In addition, "Lee S, Lee J G. The significance of pulmonary fissure completeness in video-assisted thoracoscopic surgery [J]. Journal of Thoracic Disease, 2019, 11 (Suppl 3): S420. doi: 10.21037 / jtd.2018.11.78" proposed that the integrity of the pulmonary fissure is obviously the main factor affecting the complication rate and hospitalization time after lobectomy using video-assisted thoracoscopic surgery (VATS) technology. Therefore, thoracic surgeons should carefully consider the condition of the pulmonary fissure when planning VATS lobectomy.
[0003] However, the integrity of pulmonary fissures varies greatly among individuals. "Agrawal N, Bhardwaj H, Chhabra N, et al. Pulmonary Fissures Including Accessory and Azygos Fissures and their Clinical Significance[J]. Acta Medica International, 2024, 11(1): 32-36. doi: 10.4103 / amit.amit_84_23" showed that among 25 right lungs, only 8 had complete oblique and horizontal fissures, and the other samples had varying degrees of fissure loss. In addition, "Ross JC, Nardelli P, Onieva J, et al. An open-source framework for pulmonary fissure completeness assessment [J]. Computerized Medical Imaging and Graphics, 2020, 83: 101712. doi: 10.1016 / j.compmedimag.2020.101712" points out that due to various factors such as the imperfection of computed tomography technology, the presence of image noise, uneven image intensity, and the confusion and unpredictable pathological deformation of pulmonary blood vessels in the image, pulmonary fissure segmentation is a difficult task.
[0004] Existing pulmonary fissure segmentation schemes can be roughly divided into two categories: traditional methods based on the anatomical structure information of the pulmonary fissure itself and deep learning methods based on training annotated datasets. "Xiao C, Stoel BC, Bakker ME, et al. Pulmonary fissure detection in CT images using a derivative of stick filter [J]. IEEE transactions on medical imaging, 2016, 35 (6): 1488-1500. doi: 10.1109 / TMI.2016.2517680" proposed a stick derivative (DoS) filter for crack enhancement, then introduced a branch point removal algorithm to isolate fissure patches from adhesion clutter, and then used a multi-threshold merging framework to compensate for local intensity inhomogeneity for post-processing. The Chinese patent "CN107622492B Pulmonary fissure segmentation method and system" first identifies multiple candidate pulmonary fissures, then performs region growing based on the candidate pulmonary fissures and merges the processed regions using a classification algorithm to obtain complete pulmonary fissures. "Peng Y, Xiao C. An oriented derivative of stick filter and post-processing segmentation algorithms for pulmonary fissure detection in CT images[J]. Biomedical Signal Processing and Control, 2018, 43: 278-288. doi: 10.1016 / j.bspc.2018.03.013" proposed an ODoS filter as an improvement of the existing stick derivative (DoS) filter, which enhances pulmonary fissures by incorporating stick orientation information, which will improve the filter's ability to identify clutter associated with fissure objects. "Peng Y, Zhong H, Xu Z, et al. Pulmonary lobe segmentation in CT images based on lung anatomy knowledge[J]. Mathematical Problems in Engineering, 2021, 2021(1): 5588629. doi: 10.1155 / 2021 / 5588629" introduces the alpha shape algorithm and integrates the prior knowledge of the airway and pulmonary artery to segment the pulmonary fissure. Traditional methods are sensitive to changes in data sampling and are difficult to apply to multiple datasets from different data centers.Furthermore, designing an algorithm that exploits enough global information to produce reliable results is an extremely difficult task. The reason is that due to natural anatomical variability, the location and shape of the fissures depend on global information in subtle ways.
[0005] Deep learning, on the other hand, uses multi-layer nonlinear transformations to extract high-level abstract features of data and learn the underlying distribution patterns of the data, thereby acquiring the ability to make reasonable judgments or predictions about new data. With its powerful fitting capabilities, it has emerged in various fields. Therefore, more and more people are trying to use deep learning methods to complete this task. The Chinese patent "CN114092470B A Method and Apparatus for Automatic Detection of Lung Fissures Based on Deep Learning" first uses traditional methods to obtain incomplete lung fissures, which are then used as a priori to assist in subsequent deep learning segmentation. The segmentation framework is U-Net.
[0006] "Tada DK, Teng P, McNitt-Gray M, et al. 3D patch-based CNN for fissure segmentation on CT images to quantitatively assess fissure integrity and evaluate emphysema patients for endobronchial valve treatment[C] / / MedicalImaging 2023: Computer-Aided Diagnosis. SPIE, 2023, 12465: 319-326. doi: 10.1117 / 12.2653698" A 3D U-Net was configured based on the 3D patch block to segment fissures. “Althof ZW, Gerard SE, Eskandari A, et al. Attention U-net for automated pulmonary fissure integrity analysis in lung computed tomography images[J]. Scientific reports, 2023, 13(1): 14135. doi: 10.1038 / s41598-023-41322-y” Directly segment the pulmonary fissures based on U-Net and the attention mechanism. “Fufin M, Makarov V, Alfimov VI, et al. Pulmonary Fissure Segmentation in CT Images Using Image Filtering and Machine Learning[J]. Tomography, 2024, 10(10): 1645. doi: 10.3390 / tomography10100121” A fully automatic pulmonary fissure segmentation method based on an integrated model of convolutional networks and machine learning on lung computed tomography was proposed.
[0007] “Xiao C, Stoel BC, Bakker ME, et al. Pulmonary fissure detection in CT images using a derivative of stick filter[J]. IEEE transactions on medicalimaging, 2016, 35(6): 1488-1500. doi: 10.1109 / TMI.2016.2517680” proposed a stick derivative (DoS) filter for crack enhancement, then introduced a branch point removal algorithm to isolate the crack patches from the sticky clutter, and then used a multi-threshold merging framework to compensate for local intensity inhomogeneity for post-processing. The Chinese patent “CN107622492B Pulmonary fissure segmentation method and system” first identified multiple candidate pulmonary fissures, then performed region growing based on the candidate pulmonary fissures and merged the processed regions using a classification algorithm to obtain complete pulmonary fissures. Traditional methods are sensitive to changes in data sampling and are difficult to apply to multiple data sets from different data centers. These deep learning methods mainly use general deep learning methods and do not combine prior knowledge to solve the problem of pulmonary fissure fracture. In addition, they are more likely to be classified as binary or ternary, and the curvature of the right oblique fissure in the sagittal plane varies greatly, e.g. Figure 1 As shown in the figure, the right lower lobe is morphologically in contact with the right upper lobe and right middle lobe, respectively. Because a four-class classification task can be considered in the scheme setting, the left oblique fissure, right upper oblique fissure, right horizontal fissure, and right lower oblique fissure correspond to categories 1, 2, 3, and 4, respectively.
[0008] Due to natural anatomical variability, the location and shape of the fissures depend on global information in non-obvious ways, so designing an algorithm that exploits enough global information to produce reliable results is an extremely difficult task. Summary of the Invention
[0009] To address the problems of poor robustness and high parameter dependence in traditional lung crack segmentation methods, this paper proposes a two-stage, multi-class lung crack segmentation method based on multi-perspective fusion. This method, combined with prior knowledge of lung cracks, performs clustering to effectively address the problem of lung crack fragmentation. Furthermore, we classify lung cracks into four categories (the right oblique fissure is divided into the right superior oblique fissure and the right inferior oblique fissure). The right superior oblique fissure and the right inferior oblique fissure represent the junctions of the right lower lobe, right upper lobe, and right middle lobe, respectively, further expanding its application scope.
[0010] Furthermore, the lightweight nature of this segmentation method enables more accurate segmentation results even with limited computing resources. Because the overall framework is less dependent on hardware resources, this solution can achieve good segmentation performance even on computers with limited memory and computing power. It also facilitates integration into existing medical imaging analysis software, significantly improving the system's real-time performance and accuracy, thus providing a more efficient and convenient solution for clinical applications.
[0011] The technical solution of the present invention is as follows: a two-stage multi-class lung crack segmentation method based on multi-view fusion. In the first stage, each data set is sampled according to the sagittal plane, coronal plane, and cross-section, and input into the SRPNet network for training to extract single-class lung cracks; the SRPNet network combines position prior clustering, semantic consistency self-regularization, and internal feature distillation; in the second stage, the single-class lung cracks and the original three-dimensional image are input into the lightweight MFNet network to perform lung crack fusion optimization and classification tasks.
[0012] In the first stage, a 4-fold cross-validation is performed on the data of the three orthogonal planes, namely the cross-section, coronal plane and sagittal plane. The segmentation results of each validation process are saved and used as one of the inputs of the second stage.
[0013] The first stage obtains single-class lung cracks in three orthogonal planes: coronal, sagittal, and cross-sectional planes. The SRPNet network combines position prior clustering, semantic consistency self-regularization, and internal feature distillation on the basis of the U-Net architecture. The SRPNet network includes an encoder, a decoder, and skip-layer connections of the U-Net architecture. Position prior clustering is combined with convolutional layers to form a self-attention module, which is applied to each decoder block in the decoder. The segmentation target is constrained by position prior information, and the feature maps generated by the encoder and decoder are collected. The feature maps are subjected to semantic consistency self-regularization and internal feature distillation as internal supervision.
[0014] The self-attention module obtains the contextual dependencies of the lung cracks and strengthens the relevant features by constructing position prior information and redistributing the contextual information in a clustering manner;
[0015] Perform convolution operation on the feature map to generate a single-channel image, dilate and erode the single-channel image respectively, and the eroded result is used as the foreground area A foreground , the complement of the dilated result is used as the background area A background ; Taking the foreground area and background area as the prior position, perform feature clustering to improve segmentation accuracy; Known feature map F∈R C×H×W 、A foreground ∈R 1×H×W 、A background ∈R 1×H×W, adjust the shape of F to C×HW, A foreground and A background Adjust the shape to HW×1, where C, H, and W represent the number of channels, width, and height respectively. The cluster center of each class is calculated as follows:
[0016]
[0017] The element of class class is one of the foreground area or background area, F(i)∈R C×1 is the feature vector of position i in the feature map F; reconstruct the feature map F into HW×C, fuse the cluster centers along the last dimension, and perform the feature map F and cluster center F class Perform matrix multiplication and normalize to generate affinity matrix AM:
[0018]
[0019] in, Represents the similarity between the feature vector at position j and the cluster center of the class; multiply the affinity matrix by the transposed feature vector of the cluster center to obtain the attention feature map;
[0020] The attention feature map is added to the feature map F element by element to form a new feature map The generation formula of the new feature map is as follows:
[0021]
[0022] Combine semantic consistency self-regularization and internal feature distillation to supervise the network internally, and use its cross entropy loss L ce Combined with Dice loss L dice As a joint loss function, its calculation formula is:
[0023] Loss = L dice +L ce +αL SCR +βL IFD (4)
[0024] Where α and β are L SCR With L IFD The weighting coefficient of L SCR is the loss function of semantic consistency self-regularization; L IFD is the loss function for internal feature distillation.
[0025] Semantic consistency self-regularization is to provide internal supervision to other blocks through the feature map of the last decoder block of the decoder, and establish the loss function L SCR; The other blocks include all encoder blocks and other decoder blocks except the last decoder block;
[0026]
[0027] Where M represents the total number of encoder blocks and decoder blocks, F final is the feature map of the last decoder block, are all feature maps located in the i-th layer of the m-th other block; RCS refers to the average pooling and random channel selection operations on the feature maps of other blocks; represents the number of images in the current training batch, Indicates that all images in the training batch are processed and summed.
[0028] The channels of the feature map are evenly divided into two halves, the upper half is the shallow channel features, and the lower half is the deep channel features. The internal feature distillation uses the Lp norm to distill the shallow channel features and the deep channel features, guiding the learning of useful context information to solve the feature redundancy problem. The formula is:
[0029]
[0030] Among them, F top Represents the shallow channel characteristics, F bottom Represents deep channel characteristics.
[0031] The second stage is to use the lightweight MFNet network to fuse and classify the single-class lung crack segmentation results of the three orthogonal planes in the first stage task, and finally obtain multi-class lung crack segmentation results.
[0032] The lightweight MFNet network is modified based on UNet3D, retaining the encoder structure, decoder structure and skip layer connection; the core lies in adjusting the convolution operation in the encoder and decoder of the lightweight MFNet network; adjusting the single convolution to three convolutions, and performing convolution along different axes of the three-dimensional orthogonal coordinate system respectively; for the standard 3D convolution coordinate P, the center coordinate is P i =(x i ,y i ,z i ), extract three feature maps F(P from the lth layer from the x-axis, y-axis and z-axis respectively x ), F l (P y ) and F l (P z ), expressed as:
[0033]
[0034] Among them, w(P i ) represents the position P i The weight at the lth layer convolution kernel P extracts features using the accumulation method, and then extracts n groups of features for fusion. By modifying the standard convolution kernel method, more diverse features T are obtained. l :
[0035] T l ={F l (P)1,F l (P)2,…,F l (P) n} (8).
[0036] During the training phase of the lightweight MFNet network, a random dropout strategy is introduced to prevent overfitting:
[0037]
[0038] Among them, p is the probability of random discard, d l Conforms to Bernoulli distribution; saves the best dropout strategy in the training phase and guides the model to fuse key features in the test phase; new l Perform convolution operations to fuse multiple sets of features as the overall feature map of the l-th layer lightweight MFNet network.
[0039] The present invention proposes applying a position-prior clustering self-attention module to the lung fissure segmentation task to address the problem of lung fissure fractures within a single slice. This self-attention module can be flexibly inserted into the network with minimal additional memory consumption and computational burden. Furthermore, a symmetric supervised regularization mechanism and a feature distillation module are used to improve the network, discarding irrelevant information and better preserving semantics, further enhancing model accuracy.
[0040] The overall solution of the present invention is a lightweight solution based on multi-perspective fusion. The MFNet lightweight network is proposed to integrate the multi-perspective segmentation results obtained by the SRPNet network in the first stage, so as to achieve multi-faceted attention to lung cracks, occupy low computing resources while maintaining high precision, so as to solve the problem of limited segmentation accuracy caused by the complex morphology of lung cracks under low computing resources. It is easy to integrate into existing medical image analysis software, providing a more efficient and convenient solution for clinical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 The middle blue is the horizontal fissure of the right lung, the green and yellow are the superior oblique fissure and inferior oblique fissure of the right lung respectively;
[0042] Figure 2 Flowchart of a two-stage multi-class lung crack segmentation method based on multi-view fusion;
[0043] Figure 3 This is the SRPNet network structure diagram;
[0044] Figure 4 Schematic diagram of the self-attention module based on position prior clustering;
[0045] Figure 5 This is the lightweight MFNet network structure diagram. DETAILED DESCRIPTION
[0046] This paper proposes a two-stage multi-class lung crack segmentation method based on multi-view fusion. The training process of the method is as follows: Figure 2 As shown, better segmentation results can be achieved with limited computing resources. In the first stage, each data set is sampled according to the sagittal plane, coronal plane, and cross-section, and trained using SRPNet that combines position prior clustering, semantic consistency self-regularization, and internal feature distillation to accurately extract single-class lung cracks; in the second stage, the segmentation results output by the trained network in the first stage and the original image are used as input, and a lightweight MFNet is used to perform fusion optimization and classification tasks of lung cracks. In order to avoid the situation where the segmentation results are abnormally accurate due to the reuse of the data of the first stage in the second stage, we adopt a cross-enhancement strategy, that is, the verification results of the 4-fold cross-validation process in the first stage are saved as one of the inputs of the second stage. Of course, if there is enough data, the entire data set can be divided into two, and the two stages use these two subsets respectively, which can reduce the number of models trained and the time required. The overall technical solution is as follows:
[0047] Step 1: Data preprocessing:
[0048] Data preprocessing is mainly divided into the following steps: (1) Data resampling. Because the spatial size represented by each pixel in CT or CTA images from different data sources is different, we resample them to a uniform size before training. (2) Data truncation and normalization. In order to make the network focus more on the lung fissure part, we truncate the pixel values to [-200, -1000] and then normalize them to accelerate the convergence of the model. (3) Data slicing. The three-dimensional data is sliced according to the sagittal plane, coronal plane, and cross-section, and saved as two-dimensional data for use in the first stage of training.
[0049] Step 2: First phase of training:
[0050] The training goal of the first stage is to obtain single-class lung cracks in three orthogonal planes, with the goal of fusion and classification in the second stage. Combining position prior clustering, semantic consistency self-regularization and internal feature distillation, the SRPNet network is proposed. The network structure is as follows Figure 3As shown. Using the SRPNet network for segmentation can effectively solve the problem of lung crack fractures on the two-dimensional plane and obtain good preliminary segmentation results. The core architecture of the SRPNet network is U-Net. On the basis of the U-Net architecture, it combines position prior clustering, semantic consistency self-regularization, and internal feature distillation. The SRPNet network includes the encoder, decoder and skip layer connection of the U-Net architecture. The position prior clustering is combined with the convolution layer to form a self-attention module, which is applied to each decoder block in the decoder. The segmentation target is constrained by the position prior information, and the feature maps generated by the encoder and decoder are collected. The feature maps are subjected to semantic consistency self-regularization and internal feature distillation as internal supervision.
[0051] The first is the self-attention module, which aims to obtain the contextual dependencies of lung cracks, redistribute contextual information by constructing prior positions and clustering to strengthen relevant features and ensure the continuity of segmentation results. The self-attention module module structure is as follows Figure 4 As shown in the figure. Since the lung fissure accounts for a small proportion in the CT image, the presence of several false positive points in the coarse segmentation result will cause the estimated result to deviate significantly from the true class center. We first perform a convolution operation on the feature map F to generate a single-channel image, namely the side-output result, and then dilate and erode it respectively. The eroded result is used as the foreground area A. foreground , the complement of the dilated result is used as the background area A background , and use the corrected result as the prior position to improve the accuracy of feature clustering. C×H×W 、A foreground ∈R 1×H×W 、A background ∈R 1×H×W , reshape F to C×HW and A to HW×1, where C, H, and W represent the number of channels, width, and height, respectively. The class center of each class is calculated as follows:
[0052]
[0053] The elements of the class are either foreground or background, F(i)∈R C×1 It is the feature vector of position i in the feature map F. Then the feature map F is reconstructed into HW×C, and the class center is fused along the last dimension. class Perform matrix multiplication and normalize the result to generate the affinity graph AM:
[0054]
[0055] in, Represents the similarity between the feature vector at position j and the cluster center of the class. The affinity map is then multiplied by the transposed feature vector of the class center to obtain the attention feature map, which is further added to the feature map F element by element. The generation formula of the new feature map is as follows:
[0056]
[0057] As shown in formula (2) and formula (3), two points with similar features are more likely to be assigned the same label, thus achieving continuous segmentation results of lung fissures.
[0058] The second is a joint loss based on semantic consistency regularization and internal feature distillation. It uses the feature map containing the most semantic information (i.e., the last layer of the decoder) to provide additional supervision for other blocks. By utilizing feature distillation to provide additional supervision and reduce feature redundancy, the accuracy of lung fissure segmentation is improved. Semantic consistency regularization balances the supervision between the encoder and decoder by using the feature map containing the most semantic information to provide additional supervision for the remaining blocks in the U-Net. Its formula is:
[0059]
[0060] Where M represents the total number of encoder blocks and decoder blocks, F final is the feature map of the last decoder block, are all feature maps at the i-th layer of the m-th other block. To align features in the channel and spatial dimensions, average pooling and random channel selection (RCS) are performed on the feature maps. Channel selection does not introduce additional modules, reducing computational and semantic conflicts, and the L2 norm is selected as the distance metric.
[0061] Internal feature distillation uses the Lp norm to distill information from shallow channel features to deep channel features, guiding deep features to learn useful contextual information to solve the problem of feature redundancy. Its formula is:
[0062]
[0063] in, Represents all feature maps of the i-th layer of the m-th block. In order to ensure that the number of shallow channel features and deep channel features is the same, the channel is divided into upper and lower halves, F top Represents the shallow channel characteristics, F bottom Represents deep channel characteristics.
[0064] The combination of cross entropy loss and Dice loss has been verified by many experiments. The overall joint loss function uses this standard combination plus L SCR With L IFD The weighted calculation formula is:
[0065] Loss = L dice +L ce +αL SCR +βL IFD (6)
[0066] Where α and β are L SCR With L IFD The weighting coefficient of .
[0067] Step 3: Second phase of training:
[0068] In the second phase, the lightweight MFNet network aims to fuse and classify the segmentation results from the first phase across three orthogonal planes (coronal, sagittal, and transverse). The network then uses the shallow information generated by the convolution process to recover locally erroneous results. To meet these requirements, a lightweight MFNet network was designed, enabling each neuron to fuse and repair the first-phase results using a smaller receptive field. Furthermore, the lightweight MFNet network boasts a smaller number of parameters and shorter training and testing times.
[0069] The lightweight MFNet network structure is as follows Figure 5 As shown in the figure, MFNet is adjusted based on UNet3D, retaining the basic network architecture and skip layer connection part, and then retaining two downsampling and upsampling. The input image is the segmentation result of the three perspectives in the first stage and the fusion of the original image. The number of input channels is set to 4, and the output channels are set to 8 in the first convolution. The number of channels is doubled in each subsequent downsampling process and reduced in each upsampling process. The labels of the lung fissure are set to 4 categories, so the number of output layer channels is also 4.
[0070] It is worth mentioning that the convolution blocks in the upsampling and downsampling layers of the lightweight MFNet network are different from those in the original network. Based on the linear structure of the lung cracks, we use multiple convolution kernels of different sizes to observe the structural features of the target from multiple angles and achieve feature fusion by summarizing basic standard features, thereby improving the performance of our model. Figure 5 The "Convfrom 3D" in the example represents the most basic 3×3×3 convolution kernel, and the "Conv from x" represents the n×1×1 convolution kernel. The convolution is adjusted to three times, and the convolution is performed along different axes of the three-dimensional orthogonal coordinate system. For the standard 3D convolution coordinate P, the center coordinate is P i =(x i ,y i ,z i ), extract three feature maps F(P from the lth layer from the x-axis, y-axis and z-axis respectively x ), F l (Py ) and F l (P z ).
[0071] First, set the standard 3D convolution coordinate P, the center coordinate is P i =(x i ,y i ,z i ). For each P, extract three feature maps F from the lth layer from the x-axis, y-axis, and z-axis respectively. l (P x ), F l (P y ) and F l (P z ), expressed as:
[0072]
[0073] Among them, w(P i ) represents the position P i The weights at , the features extracted by the convolution kernel P in the lth layer are calculated using the accumulation method, and then n groups of features are extracted for fusion. More diverse features are obtained by modifying the standard convolution kernel:
[0074] T l ={F l (P)1,F l (P)2,…,F l (P) n}(8)
[0075] The feature fusion of multiple templates will inevitably introduce redundant noise. Therefore, a random dropout strategy is introduced in the training phase to improve the performance of the model and prevent overfitting without adding additional computational burden, as shown in formula (9):
[0076]
[0077] Among them, p is the probability of random discard, d l The best dropout strategy is saved during the training phase and guides the model to integrate key features during the testing phase.
[0078] According to the anatomical structure of the pulmonary fissures and their relative relationship with the lung lobes, the present invention divides them into four categories (left oblique fissure, right superior oblique fissure, right inferior oblique fissure, and right horizontal fissure) based on the research of most three types of pulmonary fissures (left oblique fissure, right oblique fissure, and right horizontal fissure), which can further expand its scope of application.
[0079] Model evaluation of the present invention:
[0080] Several recent studies have focused on analyzing lung fissures in different lobes separately. However, due to class imbalance, the average value cannot accurately measure the true results, so only the weighted average of the four categories is used. To show, see formula (10). Where C i Represents the number of voxels contained in a certain class, M i is the corresponding indicator.
[0081]
[0082] The evaluation indicators adopted are as follows: precision (Pre), Dice coefficient (DSC), over-segmentation ratio (OR), under-segmentation ratio (UR), average surface distance (ASD), 95% Hausdorff Distance (HD95), surface Dice coefficient (SD). The calculation formulas of the seven indicators are shown in Equations (11)-(17).
[0083]
[0084]
[0085] Where TP represents the number of voxels correctly identified as lung fissures, TN represents the number of voxels correctly identified as background, FP represents the number of voxels incorrectly identified as lung fissures, and FN represents the number of voxels incorrectly identified as background. Pred and Gt represent the surface point sets of the segmentation result and corresponding labels, respectively. |Pred| and |Gt| represent the number of points in the set, and ‖pred-gt‖ refers to the Euclidean distance between two points.
[0086] The present invention was compared with several typical deep learning-based methods for pulmonary fissure segmentation (FissureNet, IntegrityNet) and several classic or emerging general methods (UNet3D, VNet, ResUNet3D, UNETR, TransUNet3D, UKAN3D, U-Mamba3D, and Vision-LSTM3D). Training was performed on CT and CTA datasets. Model sizes and related parameters are shown in Table 1. The present invention outperformed the other methods in terms of the Dice coefficient, undersegmentation rate, and surface Dice coefficient, reaching 81.79%, 17.56%, and 93.12%, respectively, on the CT dataset, and 81.19%, 17.09%, and 92.37%, respectively, on the CTA dataset. In the CT dataset, the present method achieved the best average surface distance of 0.24 mm. In the CTA dataset, the present method achieved the best 95% Hausdorff distance of 8.21 mm. The weighted average indicators are shown in Tables 2 and 3, and the overall indicators of each category are shown in Tables 5 and 6.
[0087] To further validate the present invention, a series of ablation studies were conducted on a CT dataset. To present the results more intuitively, the first-stage task of segmenting the entire lung fissure was changed to segmenting four types of lung fissures, consistent with the output format of the second-stage. The results presented in each stage are the highest of the three orthogonal planes, as shown in Table 4. In the table, PC represents the self-attention mechanism based on position prior clustering, SR represents semantic consistency regularization and internal distillation, and MF represents multi-view feature fusion. It can be seen that the optimal results are achieved when all frameworks are fully utilized, with an overall Dice coefficient improvement of 4.3894%. Using only a single component (PC, SR, and MF) yielded Dice coefficient improvements of 1.0071%, 0.5598%, and 2.6588%, respectively. Multi-view fusion contributed the most, demonstrating that segmentation from different viewpoints does indeed differ. These differences can be mutually leveraged during the second-stage training process, allowing for the recovery of incorrectly segmented parts to a certain extent.
[0088] Table 1 shows the different model parameters and corresponding settings
[0089] Model Flops / G Params / MB Patch_size UNet3D 2708.59 79.98 128*128*128 VNet 1146.75 237.11 128*128*128 UNETR 2238.35 557.71 128*128*128 ResUNet3D 872.64 36.25 128*128*128 IntegrityNet 1582.83 93.12 128*128*128 TransUNet3D 2016.32 500.53 128*128*128 UKAN3D 421.37 107.18 128*128*128 FissureNet 276.61 13.42 128*128*128 Vision-LSTM3D 1361.19 84.06 128*128*128 U-Mamba3D 5422.07 82.25 128*128*128 Stage-1 324.30 16.48 512*512 Stage-2 173.99 1.03 128*128*128
[0090] Table 2 shows the weighted prediction results of different models in the CT dataset
[0091]
[0092]
[0093] Table 3 Weighted prediction results of different models in the dataset
[0094]
[0095] Table 4 Ablation experiments of different components
[0096]
[0097] Table 5 Prediction results of different categories of different models in the CT dataset
[0098]
[0099]
[0100]
[0101] Table 6 Prediction results of different categories of different models in the CTA dataset
[0102]
[0103]
[0104]
Claims
1. A two-stage multi-class lung crack segmentation method based on multi-view fusion, characterized by: In the first stage, each dataset is sampled according to the sagittal, coronal, and cross-sectional planes and fed into the SRPNet network for training to extract single-class lung cracks. The SRPNet network combines position prior clustering, semantic consistency self-regularization, and internal feature distillation. In the second stage, single-class lung cracks and original three-dimensional images are input into the lightweight MFNet network to perform lung crack fusion optimization and classification tasks.
2. The two-stage multi-class lung crack segmentation method based on multi-view fusion according to claim 1 is characterized in that: In the first stage, a 4-fold cross-validation is performed on the data of the three orthogonal planes, namely the cross-section, coronal plane and sagittal plane. The segmentation results of each validation process are saved and used as one of the inputs of the second stage.
3. The two-stage multi-class lung crack segmentation method based on multi-view fusion according to claim 1 or 2, characterized in that: The first stage obtains single-class lung cracks in three orthogonal planes: coronal, sagittal, and cross-sectional planes. The SRPNet network is based on the U-Net architecture and combines position prior clustering, semantic consistency self-regularization, and internal feature distillation. The SRPNet network includes an encoder, decoder, and skip-layer connections of the U-Net architecture. It combines position prior clustering with convolutional layers to form a self-attention module, which is applied to each decoder block in the decoder. It constrains the segmentation target through position prior information, collects feature maps generated by the encoder and decoder, and performs semantic consistency self-regularization and internal feature distillation on the feature maps as internal supervision.
4. The two-stage multi-class lung crack segmentation method based on multi-view fusion according to claim 3 is characterized in that: The self-attention module obtains the contextual dependencies of the lung cracks and strengthens the relevant features by constructing position prior information and redistributing the contextual information in a clustering manner; Perform convolution operation on the feature map to generate a single-channel image, dilate and erode the single-channel image respectively, and the eroded result is used as the foreground area A foreground , the complement of the dilated result is used as the background area A background ; Taking the foreground area and background area as the prior position, perform feature clustering to improve segmentation accuracy; Known feature map F∈R C×H×W 、A foreground ∈R 1×H×W 、A background ∈R 1×H×W , adjust the shape of F to C×HW, A foreground and A background Adjust the shape to HW×1, where C, H, and W represent the number of channels, width, and height respectively. The cluster center of each class is calculated as follows: The element of class class is one of the foreground area or background area, F(i)∈R C×1 is the feature vector of position i in the feature map F; reconstruct the feature map F into HW×C, fuse the cluster centers along the last dimension, and perform the feature map F and cluster center F class Perform matrix multiplication and normalize to generate affinity matrix AM: in, Represents the similarity between the feature vector at position j and the cluster center of the class; multiply the affinity matrix by the transposed feature vector of the cluster center to obtain the attention feature map; The attention feature map is added to the feature map F element by element to form a new feature map The generation formula of the new feature map is as follows:
5. The two-stage multi-class lung crack segmentation method based on multi-view fusion according to claim 3 is characterized in that: Combine semantic consistency self-regularization and internal feature distillation to supervise the network internally, and use its cross entropy loss L ce Combined with Dice loss L dice As a joint loss function, its calculation formula is: Loss= L dice +L ce +αL SCR +βL IFD (4) Where α and β are L SCR With L IFD The weighting coefficient of L SCR is the loss function of semantic consistency self-regularization; L IFD is the loss function for internal feature distillation.
6. The two-stage multi-class lung crack segmentation method based on multi-view fusion according to claim 5 is characterized in that: Semantic consistency self-regularization is to provide internal supervision to other blocks through the feature map of the last decoder block of the decoder, and establish the loss function L SCR ; The other blocks include all encoder blocks and other decoder blocks except the last decoder block; Where M represents the total number of encoder blocks and decoder blocks, F final is the feature map of the last decoder block, are all feature maps located in the i-th layer of the m-th other block; RCS refers to the average pooling and random channel selection operations on the feature maps of other blocks; represents the number of images in the current training batch, Indicates that all images in the training batch are processed and summed.
7. The two-stage multi-class lung crack segmentation method based on multi-view fusion according to claim 5, characterized in that: The channels of the feature map are evenly divided into two halves, the upper half is the shallow channel features, and the lower half is the deep channel features. The internal feature distillation uses the Lp norm to distill the shallow channel features and the deep channel features, guiding the learning of useful context information to solve the feature redundancy problem. The formula is: Among them, F top Represents the shallow channel characteristics, F bottom Represents deep channel characteristics.
8. The two-stage multi-class lung crack segmentation method based on multi-view fusion according to claim 3 is characterized in that: The second stage is to use the lightweight MFNet network to fuse and classify the single-class lung crack segmentation results of the three orthogonal planes in the first stage task, and finally obtain multi-class lung crack segmentation results.
9. The two-stage multi-class lung crack segmentation method based on multi-view fusion according to claim 8, characterized in that: The lightweight MFNet network is modified based on UNet3D, retaining the encoder structure, decoder structure and skip layer connection; the core lies in adjusting the convolution operation in the encoder and decoder of the lightweight MFNet network; adjusting the single convolution to three convolutions, and performing convolution along different axes of the three-dimensional orthogonal coordinate system respectively; for the standard 3D convolution coordinate P, the center coordinate is P i =(x i ,y i ,z i ), extract three feature maps F(P from the lth layer from the x-axis, y-axis and z-axis respectively x ), F l (P y ) and F l (P z ), expressed as: Among them, w(P i ) represents the position P i The weight at the lth layer convolution kernel P extracts features using the accumulation method, and then extracts n groups of features for fusion. By modifying the standard convolution kernel method, more diverse features T are obtained. l : T l ={F l (P)1,F l (P)2,…,F l (P) n } (8)。 10. The two-stage multi-class lung crack segmentation method based on multi-view fusion according to claim 9, characterized in that: During the training phase of the lightweight MFNet network, a random dropout strategy is introduced to prevent overfitting: Among them, p is the probability of random discard, d l Conforms to Bernoulli distribution; saves the best dropout strategy in the training phase and guides the model to fuse key features in the test phase; new l Perform convolution operations to fuse multiple sets of features as the overall feature map of the l-th layer lightweight MFNet network.
Citation Information
Patent Citations
Methods and systems for lung fissure segmentation
CN107622492B
A Deep Learning-Based Automated Detection Method and Device for Lung Clefts
CN114092470B