Improved 3D network-based active / latent tuberculosis identification system and method

By combining the improved 3D ResUNet50 model with multi-scale attention modules and grouped convolution with channel shuffling, the accuracy and efficiency problems of CT image classification for tuberculosis in existing technologies are solved, achieving higher recognition accuracy and faster classification speed.

CN120876961BActive Publication Date: 2026-03-17GUIZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510978880.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2026-03-17
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

Existing 3D network-based active tuberculosis identification systems are insufficient in terms of accurate classification, especially in terms of the accuracy and efficiency of CT image recognition, which needs to be improved.

Method used

An improved 3D ResUNet50 classification model, combined with multi-scale attention module (MSAM) and grouped convolution with channel shuffling (GC), is used to extract features and classify CT images of pulmonary tuberculosis. Image quality is optimized through data preprocessing, and accurate segmentation is achieved using the 3D ResUNet segmentation model.

Benefits of technology

The model improved the classification accuracy and efficiency of CT images of pulmonary tuberculosis. Experimental results showed that the precision was 0.941, the recall was 0.902, the F1 score was 0.921, and the ROC was 0.970, which were 2.5, 2.0, 3.2 and 5.0 percentage points higher than the traditional model, respectively. The accuracy reached 90% in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876961B_ABST
    Figure CN120876961B_ABST
Patent Text Reader

Abstract

The application discloses an activity / inactivity tuberculosis recognition system based on an improved 3D network, which comprises an image feature acquisition module, a data preprocessing module, an image segmentation module and an image classification recognition module. The image feature acquisition module is used for acquiring lung tuberculosis CT chest image features. The data preprocessing module is used for performing data preprocessing on the acquired lung tuberculosis CT chest image features. The image segmentation module is used for segmenting the preprocessed lung tuberculosis CT chest image features to acquire lung parenchyma image features, and a segmentation model adopted is a 3D ResUNet segmentation model. The image classification recognition module is used for performing prediction classification on the lung parenchyma image features after image segmentation. An image classification model adopted is a 3D ResUNet50 classification model, and the 3D ResUNet50 classification model is combined with a multi-scale attention module to obtain image output features. Then, grouping convolution and channel shuffling are used to process the image output features, the categories of active and inactive tuberculosis are predicted, and the confidence score of the model is generated. The application can improve the CT image classification accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of tuberculosis classification and detection devices, and relates to an active / inactive tuberculosis identification system and method based on an improved 3D network. Background Technology

[0002] Pulmonary tuberculosis (PTB) is a chronic infectious disease that is primarily transmitted through the respiratory tract and is caused by Mycobacterium tuberculosis.

[0003] Traditional diagnostic methods for PTB mainly include sputum culture, sputum smear microscopy, immunological testing, and imaging examinations. Sputum culture is considered the "gold standard" for diagnosing tuberculosis and can detect drug-resistant strains; however, it is time-consuming and complex. Sputum smear microscopy, while simple, rapid, and inexpensive, has low sensitivity and is prone to missed diagnoses. Immunological testing is an important tool for diagnosing Mycobacterium tuberculosis infection; the common tuberculin skin test involves subcutaneous injection of tuberculin and observation of local skin reaction, but it cannot distinguish between active and inactive tuberculosis. In contrast, imaging examinations can directly show the spatial distribution, size, and morphological characteristics of pulmonary tuberculosis lesions, aiding in the assessment of disease severity and the formulation of treatment plans. Furthermore, approximately 30%-40% of patients with active pulmonary tuberculosis have negative sputum smear results, making diagnosis difficult using traditional microbiological testing alone. Imaging examinations can clearly show typical active lesions in these patients, effectively compensating for the limitations of microbiological testing.

[0004] Due to the rapid development of computer technology and artificial intelligence (AI), research on the application of deep learning in medical imaging has reached new heights. By combining computer vision AI with medical image diagnostic technology, organ segmentation, disease classification, and automatic lesion detection in medical images can be achieved, which helps improve diagnostic accuracy, work efficiency, and enables remote diagnosis. Some researchers have used chest X-rays to differentiate between active and inactive pulmonary tuberculosis (PTB). Compared to computed tomography (CT), chest X-rays, as a two-dimensional imaging modality, can only provide a projection of chest structures, thus limiting the visualization clarity of certain active PTB-related features (such as cavitation, tree-in-bud sign, and bronchial wall thickening). In the application of deep learning to CT datasets, scholars have proposed many good methods, but these methods all have some shortcomings. The paper "Li X, Zhou Y, Du P, et al. A deep learning system that generates quantitative CT reports for diagnosing pulmonary tuberculosis[J]. Applied Intelligence, 2021, 51:4082-4093 (A Deep Learning-Based Quantitative CT Report Generation System for Pulmonary Tuberculosis)" evaluated four fine-tuned three-dimensional convolutional neural network (CNN) models and used the best-performing model to detect and classify PTB lesion areas using CT image datasets. However, its lung parenchyma segmentation uses a fixed threshold, which has weak generalization ability and poor recognition accuracy. The paper "Duwairi R, Melhem A. A deep learning-based framework for automatic detection of drug resistance intuberculosis patients[J]. Egyptian Informatics Journal, 2023, 24(1):139-148" uses pre-trained networks such as VGG19 and ResNet for feature extraction and performs classification through cascaded convolutions and dense layers, validating the advantages of multimodal input. This method achieved a recognition accuracy of 74.13% in the multidrug resistance detection task, but only 53% in the pulmonary tuberculosis type classification task.The paper "Gao XW, James-Reynolds C, Currie E. Analysis of tuberculosis severity levels from CT pulmonary images based on enhanced residual deep learning architecture[J]" introduces an enhanced residual deep learning architecture (depth-ResNet) and proposes a new method to integrate the scores of individual blocks to predict the severity of the entire tuberculosis dataset. However, the recognition efficiency is very low, requiring 4 days for training and 2 days for testing. The paper "Nijiati M, Zhou R, Damaola M, Hu C, Li L, Qian B, et al. Deep learning based CT images automatic analysis model for active / non-active pulmonary tuberculosis differential diagnosis. Front Mol Biosci. 2022 Dec; 9:1086047. doi:10.3389 / fmolb.2022.1086047 (Application of Deep Learning-Based CT Image Automatic Analysis Model in Differential Diagnosis of Active / Non-Active Pulmonary Tuberculosis)" first applied the 3D ResNet-50 model to diagnose active pulmonary tuberculosis and achieved significant results. However, the proposed model has poor accuracy in identifying active pulmonary tuberculosis. Summary of the Invention

[0005] The technical problem to be solved by this invention is: an active / inactive pulmonary tuberculosis identification system based on an improved 3D network, which can efficiently and accurately classify CT images of active pulmonary tuberculosis.

[0006] To address this problem, the technical solution adopted by the present invention is as follows:

[0007] An active / inactive tuberculosis identification system based on an improved 3D network includes:

[0008] Image feature acquisition module, used to acquire chest CT image features of pulmonary tuberculosis;

[0009] The data preprocessing module is used to preprocess the acquired CT chest image features of pulmonary tuberculosis.

[0010] The image segmentation module is used to segment the features of the preprocessed CT chest image of pulmonary tuberculosis to obtain lung parenchymal image features. The segmentation model used by the image segmentation module is the 3D ResUNet segmentation model.

[0011] The image classification and recognition module is used to predict and classify the features of lung parenchyma images after image segmentation. The image classification and recognition module uses the 3D ResUNet50 classification model and combines the 3D ResUNet50 classification model with a Multi Scale Attention Module (MSAM) to obtain image output features. Then, group convolution and channel shuffle (GC) is used to process the image output features (group convolution and channel shuffle optimize the parameter redundancy and high computational cost caused by the 3D convolutional network and MSGC module) to obtain the classification of active and inactive pulmonary tuberculosis and generate the model's confidence score.

[0012] Furthermore, the aforementioned data preprocessing module preprocesses the original CT image, including distortion removal, binarization, resizing, and threshold adjustment. The window width and window level of the adjusted CT image are then adjusted to [1400, -500]. Next, the CT image size is adjusted to [128, 256, 256]. Then, the pixel values ​​of the image are normalized to a fixed range [0, 1]. Finally, the normalized image is enhanced by flipping, rotating, scaling, and cropping to obtain the preprocessed pulmonary tuberculosis CT chest image features.

[0013] Furthermore, the implementation method of the above-mentioned 3D ResUNet segmentation model is as follows: when using CT image features of size 128×256×256 as input, it is processed through an encoder, a bottleneck layer, and a decoder to generate segmentation probability maps with the same resolution as the input multiple times. The encoder uses residual blocks and three-dimensional max pooling (MaxPool3D) to progressively downsample the data, reducing the spatial size from 128×256×256 to 64×128×128, 32×64×64, 16×32×32, and 8×16×16; simultaneously, the number of feature channels is also... The number of convolutional dimensions increases from 1 to 32, 64, 128, 256, and 512, effectively extracting multi-scale and deep semantic features. At the bottleneck layer, the network captures global contextual information through deep dialogue. Then, the decoder uses transposed convolutions for progressive upsampling, restoring the spatial dimensions from 8×16×16 to 16×32×32, 32×64×64, 64×128×128, and 128×256×256. At each decoding stage, jump connections connect the low-level features from the corresponding encoder stage to the decoder, integrating global semantics and local details to improve segmentation accuracy. Finally, a 1×1×1 fully connected layer is used with a sigmoid activation function to map the feature map to a voxel segmentation probability map, ensuring the output dimension matches the input. After training the model using dataset 1, lung parenchyma is extracted using dataset 2, preparing for subsequent classification.

[0014] Furthermore, the conv2 implementation method of the aforementioned 3D ResUNet50 classification model is as follows: a multi-scale attention mechanism is introduced into the residual structure of 3DResNet. This mechanism first captures lesion features at multiple scales simultaneously through multi-scale convolution operations (such as 1×1×1, 3×3×3, and 5×5×5), which is particularly effective for small and unevenly distributed tuberculous lesions. Then, the output characteristics of the multi-scale convolutions are weighted to emphasize the lesion region and minimize interference from irrelevant regions. In addition, a single-channel spatial attention map is generated using 1×1×1 convolution and a sigmoid activation function to dynamically adjust the model's attention to different spatial locations in the feature map. Finally, the elements of the input feature map are multiplied by the attention map, and residual connections are used to maintain the integrity of the original features while highlighting key regional features, thus providing more discriminative feature reproduction for subsequent classification tasks.

[0015] Furthermore, the implementation method of the above grouped convolution and channel shuffling is as follows: Assume that the shape of the input feature map is C×D×H×W, where C is the number of channels, D is the depth, H is the height, W is the width, the size of the convolution kernel is K×K×K, and the number of output channels is N. In the grouped convolution, assuming that the input channels are divided into G groups, the number of channels in each group is C / G. Perform the three-dimensional convolution operation independently on each group of channels, and the computational complexity of each group is as shown in equation (2):

[0016] O2′=C / G×D×H×W×K 3 ×N / G (2)

[0017] The output feature maps of group G are concatenated by channels to obtain the final feature map, and its computational complexity is shown in equation (3):

[0018] O2 = C × D × H × W × K 3 ×N×1 / G (3)

[0019] Channel shuffling rearranges the channel order of feature maps through three steps: recombination, transposition, and flattening, enabling different groups of features to mix and interact.

[0020] A method for identifying active / inactive pulmonary tuberculosis based on an improved 3D network, the method comprising:

[0021] Obtain chest CT image features of pulmonary tuberculosis;

[0022] Data preprocessing was performed on the acquired chest CT images of pulmonary tuberculosis.

[0023] The preprocessed CT chest images of pulmonary tuberculosis are segmented to obtain lung parenchymal image features. The segmentation model used in the image segmentation module is the 3D ResUNet segmentation model.

[0024] The lung parenchyma image features after image segmentation are predicted and classified. The image classification and recognition module uses the 3D ResUNet50 classification model and combines the 3D ResUNet50 classification model with a multi-scale attention module to obtain image output features. Then, grouped convolution and channel shuffling are used to process the image output features, predict the category of active and inactive pulmonary tuberculosis, and generate the confidence score of the model.

[0025] The beneficial effects of this invention are as follows: Compared with the prior art, this invention utilizes 3D ResUNet to segment lung parenchymal CT images in the dataset, and designs a 3DMSGC-ResNet50 classification model based on the 3D ResNet50 architecture. This model introduces a multi-scale attention module (MSAM) and grouped convolution with channel shuffling (GC). This improvement enhances the network's ability to capture multi-scale lesion features and dynamically focus on key regions, thereby improving the model's classification performance. This results in higher accuracy and significantly improved classification efficiency for lung parenchymal CT images. Experimental results show that the proposed model has an accuracy of 0.941, a recall of 0.902, an F1 score of 0.921, and an ROC of 0.970, which are 2.5, 2.0, 3.2, and 5.0 percentage points higher than the ResNet50 model alone, respectively. Furthermore, the improved model achieves a prediction accuracy of 90% in 20 external test cases. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the classification model's workflow;

[0027] Figure 2 This is a dataset of lung parenchyma segmentation; in the figure, (a) coronal plane; (b) sagittal plane; (c) transverse plane;

[0028] Figure 3 Diagram of 3DResUnet structure;

[0029] Figure 4 To improve the comparison diagram of residual networks;

[0030] Figure 5 For 3D MSGC-ResNet50 networks

[0031] Figure 6 The diagram shows the segmentation of lung parenchyma in patients with active and inactive pulmonary tuberculosis. In the diagram, (a) shows the segmentation of lung parenchyma in patients with active pulmonary tuberculosis; and (b) shows the segmentation of lung parenchyma in patients with inactive pulmonary tuberculosis.

[0032] Figure 7 This is the performance metric curve of the model during training;

[0033] Figure 8 The ROC curve of the model;

[0034] Figure 9 The diagram shows the confusion matrix; in the diagram, (a) is the test set of the dataset, and (b) is the external test set. Detailed Implementation

[0035] The present invention will be further described below with reference to specific embodiments.

[0036] Example 1: As Figure 1As shown, an active / inactive pulmonary tuberculosis identification system based on an improved 3D network includes:

[0037] Image feature acquisition module, used to acquire chest CT image features of pulmonary tuberculosis;

[0038] The data preprocessing module is used to preprocess the acquired CT chest image features of pulmonary tuberculosis.

[0039] The image segmentation module is used to segment the features of the preprocessed CT chest image of pulmonary tuberculosis to obtain lung parenchymal image features. The segmentation model used by the image segmentation module is the 3D ResUNet segmentation model.

[0040] The image classification and recognition module is used to predict and classify the features of lung parenchyma images after image segmentation. The module employs a 3D ResUNet50 classification model, which is combined with a Multi Scale Attention Module (MSAM) to obtain image output features. Then, group convolution and channel shuffle (GC) are used to process these features (group convolution and channel shuffle optimize the parameter redundancy and high computational cost caused by the 3D convolutional network and MSGC module) to obtain the classification of active and inactive pulmonary tuberculosis and generate a model confidence score (the confidence score refers to a probability value predicted by the artificial intelligence model when identifying active / inactive pulmonary tuberculosis).

[0041] Furthermore, the aforementioned data preprocessing module preprocesses the original CT images, including distortion removal, binarization, resizing, and threshold adjustment.

[0042] To optimize the efficiency and stability of model training, the data preprocessing method was further optimized. First, the window width and window level of the threshold-adjusted lung parenchymal CT images were adjusted to [1400, -500] to better display the details of pulmonary tuberculosis lesions, while suppressing interference from other irrelevant tissues (such as bone and fat), thus optimizing the visualization effect of CT images and providing the model with clearer input data for lesion features. Next, the CT image size was adjusted to [128, 256, 256] to maximize image resolution within hardware resource constraints, providing richer spatial features and helping the model to better learn and distinguish target regions. Then, the pixel values ​​of the images were normalized to a fixed range [0, 1] to eliminate data scale differences, accelerate model convergence, and improve training stability. Finally, image enhancement operations such as flipping, rotating, scaling, and cropping were used to increase data diversity, improve the model's generalization ability and robustness, and avoid overfitting.

[0043] The window width and window level of the adjusted CT image were then set to [1400, -500]. Next, the CT image size was adjusted to [128, 256, 256]. Then, the pixel values ​​of the image were normalized to a fixed range of [0, 1]. Finally, the normalized image was enhanced by flipping, rotating, scaling, and cropping to obtain the preprocessed CT chest image features of pulmonary tuberculosis.

[0044] The aforementioned 3D ResUNet is a 3D model that integrates the classic U-Net and ResNet architectures. It enhances feature extraction by incorporating residual blocks and employs an encoder-decoder architecture to maintain spatial resolution while progressively restoring the input size, thus ensuring accurate segmentation. The structure of 3D ResUNet is as follows: Figure 3 As shown, the implementation method of the 3D ResUNet segmentation model is as follows: When using CT image features of size 128×256×256 as input, it is processed through an encoder, bottleneck layer, and decoder to generate segmentation probability maps with the same resolution as the input multiple times. The encoder uses residual blocks and 3D max pooling (MaxPool3D) to progressively downsample the data, reducing the spatial size from 128×256×256 to 64×128×128, 32×64×64, 16×32×32, and 8×16×16; simultaneously, the number of feature channels is also... The number of convolutional dimensions increases from 1 to 32, 64, 128, 256, and 512, effectively extracting multi-scale and deep semantic features. At the bottleneck layer, the network captures global contextual information through deep dialogue. Then, the decoder uses transposed convolutions for progressive upsampling, restoring the spatial dimensions from 8×16×16 to 16×32×32, 32×64×64, 64×128×128, and 128×256×256. At each decoding stage, jump connections connect the low-level features from the corresponding encoder stage to the decoder, integrating global semantics and local details to improve segmentation accuracy. Finally, a 1×1×1 fully connected layer is used with a sigmoid activation function to map the feature map to a voxel segmentation probability map, ensuring the output dimension matches the input. After training the model using dataset 1, lung parenchyma is extracted using dataset 2, preparing for subsequent classification.

[0045] To better extract lesion features of pulmonary tuberculosis and achieve diagnosis of active and inactive PTB, this invention proposes a method based on 3D ResNet50 (e.g., Figure 4 (a) shows the 3D MSGS-ReNet50 classification model, which combines multi-scale attention mechanisms with grouped convolution and channel shuffling. Its improved residual blocks are as follows: Figure 4As shown in (b) (taking the conv2 structure as an example, the subsequent layers are similar (i.e., the conv3-conv5 joint multi-scale attention mechanism of the 3D MSGS-ReNet50 classification model, as well as grouped convolution and channel shuffling)

[0046] like Figure 4 As shown in (a), 3D ResNet is an extension of the traditional ResNet architecture. The core concept of ResNet is residual learning, which introduces skip connections to directly add the input to the output, forming so-called "residual blocks." This structure solves the vanishing and exploding gradient problems common in deep network training, making network training more efficient and stable. Furthermore, ResNet employs batch normalization and the ReLU activation function to accelerate training and improve the model's non-linear expressive power. Table 1 summarizes the 3D ResNet18, ResNet34, ResNet50, ResNet101, and ResNet200 architectures.

[0047] Table 1. ResNet18, ResNet34, ResNet50, ResNet101, and ResNet200 architectures in 3D

[0048]

[0049]

[0050] When classifying active and inactive primary lung tumors (PTB), traditional 3D ResNet classification networks face significant challenges in feature extraction and lesion classification due to the varying sizes, complexity, and uneven distribution of PTB lesions. On one hand, PTB lesions typically exhibit features at multiple scales, and the fixed kernel size of standard 3D ResNet is insufficient to fully capture these multi-scale features. On the other hand, 3D ResNet lacks a clear attention mechanism, resulting in insufficient focus on key lesion regions and making the model susceptible to interference from other lung structures, leading to unsatisfactory classification results.

[0051] To address the aforementioned challenges, this invention introduces a multi-scale attention mechanism (MSAM) into the residual structure of 3D ResNet. The implementation methods for conv2-conv5 in the 3D ResUNet50 classification model are the same. The conv2 implementation method of the 3D ResUNet50 classification model is as follows: A multi-scale attention mechanism is introduced into the residual structure of 3D ResNet. This multi-scale attention mechanism first captures lesion features at multiple scales simultaneously through multi-scale convolution operations (such as 1×1×1, 3×3×3, and 5×5×5), which is particularly effective for small and unevenly distributed tuberculous lesions. Then, the output characteristics of the multi-scale convolution are weighted to emphasize the lesion area and minimize the interference of irrelevant areas. In addition, a single-channel spatial attention map is generated using 1×1×1 convolution and the Sigmoid activation function to dynamically adjust the model's attention to different spatial locations in the feature map. Finally, the elements of the input feature map are multiplied by the attention map, and residual connections are used to maintain the integrity of the original features while highlighting key regional features, thereby providing more discriminative feature reproduction for subsequent classification tasks.

[0052] like Figure 4 As shown in b, in order to extract important features after compression more efficiently, this invention introduces grouped convolution and channel shuffling. The introduction of grouped convolution can not only further extract features, but also has lower computational complexity compared with traditional convolution. Specifically, the implementation method of grouped convolution and channel shuffling is as follows: Assuming that the shape of the input feature map is C×D×H×W, where C is the number of channels, D is the depth, H is the height, W is the width, the size of the convolution kernel is K×K×K, and the number of output channels is N, then the computational complexity of traditional convolution is as shown in equation (1):

[0053] O1=C×D×H×W×K 3 ×N (1)

[0054] In grouped convolution, assuming the input channels are divided into G groups, the number of channels in each group is C / G. The three-dimensional convolution operation is performed independently on each group of channels, and the computational complexity of each group is as shown in equation (2).

[0055] O′2=C / G×D×H×W×K 3 ×N / G (2)

[0056] The output feature maps of group G are concatenated by channels to obtain the final feature map, and its computational complexity is shown in equation (3):

[0057] O2 = C × D × H × W × K 3 ×N×1 / G (3)

[0058] The calculation results show that the number of parameters after using grouped convolution is 1 / G of that of traditional convolution. Therefore, grouped convolution allows for the utilization of more convolution kernels with the same computational resources, increasing the model's depth and width, enhancing its expressive power, capturing richer features, and thus improving the accuracy of CT image classification. On the other hand, by independently processing different groups of input channels, it promotes feature diversity, enabling different groups of convolution kernels to learn different feature representations, enhancing the model's generalization ability, and also contributing to improved CT image classification accuracy.

[0059] Meanwhile, channel shuffling is used to solve the problem of information isolation between groups caused by grouped convolution. Channel shuffling rearranges the channel order of the feature map through three steps: recombination, transposition, and flattening, so that features from different groups can mix and interact.

[0060] Evaluation metrics for the prediction system: In the simulation experiments of this invention, "Positive" represents active pulmonary tuberculosis (positive), and "Negative" represents inactive pulmonary tuberculosis (negative). A true positive (TP) indicates a case where the true condition is positive and the prediction is also positive. A false positive (FP) indicates a case where the true condition is negative but the prediction is positive. A true negative (TN) indicates an instance where the true condition is negative and the prediction is also negative. A false negative (FN) indicates an instance where the true condition is positive but the prediction is negative.

[0061] This simulation experiment uses multiple metrics to evaluate model performance, including precision, recall, F1 score, and AUC, as shown in formulas (4)-(6):

[0062]

[0063] In the formula, AUC represents the area under the ROC curve. The AUC value is directly proportional to the model's performance. The vertical axis of the ROC curve represents the true positive rate (TPR), and the horizontal axis represents the false positive rate (FPR).

[0064]

[0065] In the formula, rank_i represents the index of the i-th sample (samples are sorted by probability score from smallest to largest); M and N represent the number of positive samples and negative samples, respectively; pc represents the index of all positive samples; M×(M+1) / 2 is a correction coefficient used to adjust the sorting of positive samples.

[0066] The classification model employs the cross-entropy loss function (CELoss) to optimize its classification performance. This loss function quantifies the difference between the model's predicted probability distribution and the true label distribution, providing a gradient direction for parameter optimization. Its mathematical expression can be represented as:

[0067]

[0068] In formula (10), N is the batch size, C represents the total number of categories, and y i,c p represents the true label value of the first sample in class c (using single-encoding, where the active PTB class is labeled 0 and other classes are labeled 1). i,c This represents the probability value that the model predicts the first sample belongs to class c.

[0069] Example 2: A method for identifying active / inactive pulmonary tuberculosis based on an improved 3D network, the method comprising:

[0070] Obtain chest CT image features of pulmonary tuberculosis;

[0071] Data preprocessing was performed on the acquired chest CT images of pulmonary tuberculosis.

[0072] The preprocessed CT chest images of pulmonary tuberculosis are segmented to obtain lung parenchymal image features. The segmentation model used in the image segmentation module is the 3D ResUNet segmentation model.

[0073] The lung parenchyma image features after image segmentation are predicted and classified. The image classification and recognition module uses the 3D ResUNet50 classification model and combines the 3D ResUNet50 classification model with a multi-scale attention module to obtain image output features. Then, grouped convolution and channel shuffling are used to process the image output features, predict the category of active and inactive pulmonary tuberculosis, and generate the confidence score of the model.

[0074] Dataset selection for training in this invention: Two datasets were used. The first is the lung CT dataset provided by the official 2020 ImageCLEF competition (hereinafter referred to as dataset 1). This dataset was used to train the lung parenchyma segmentation model. This dataset includes 217 original CT images and 217 mask images. All data are stored in NIFTI file format. Three views of a specific CT image are shown below. Figure 2 As shown in the figure. The second is the active / inactive pulmonary tuberculosis CT dataset (hereinafter referred to as dataset 2) provided by the Guiyang Public Health Treatment Center of Guizhou Province. It is a dataset of CT images collected by the Guiyang Public Health Treatment Center of Guizhou Province from January 2018 to March 2023. It contains 500 cases each of active and inactive pulmonary tuberculosis diagnosed. These cases were used as the lung parenchyma segmentation objects, and the segmentation results were used as the raw data for classification diagnosis.

[0075] To verify the effectiveness of the present invention, the following simulation was performed:

[0076] 1. Experimental Environment

[0077] The experiment was conducted in a WSL-Ubuntu 22.04 environment. The hardware configuration included one Intel(R) Core(TM) i9-12900K CPU and two NVIDIA... TM The system consisted of a GeForce RTX 4090 GPU, CUDA 11.7, 128GB of RAM, and a 64-bit operating system. Training parameters included a learning rate of 0.001, a batch size of 4, 8 worker threads, and a lung parenchymal CT scan size adjusted to 128×256×256. The classification model's dataset was divided into training and test sets in a 9:1 ratio, with the training set using K=1 K-fold cross-validation.

[0078] 2. Analysis of lung parenchyma segmentation results

[0079] Figure 6 The results of lung parenchyma segmentation from CT images of active (a) and inactive (b) pulmonary tuberculosis are presented. The segmentation results show that the algorithm exhibits high accuracy in both pathological types of lung CT. Regarding anatomical fit, the segmented lung parenchyma contours closely match the actual anatomical boundaries in the axial, coronal, and sagittal planes, without any errors such as region omissions or boundary breaks.

[0080] From a pathological perspective, the segmentation algorithm demonstrates good adaptability to different pathological features. For example, it can still preserve the complete lung parenchymal structure in the case of variable lesion areas (such as nodular shadows and cavities) in active pulmonary tuberculosis; and it does not lead to misjudgment of lung parenchymal segmentation boundaries in the case of high-density fibrosis or calcifications in inactive pulmonary tuberculosis.

[0081] In addition, from the perspective of lesion morphology, although the inflammatory infiltration area of ​​active cases presents a blurred and irregular texture, while the fibrotic area of ​​inactive cases presents a sharp and dotted morphological feature, the segmentation algorithm can accurately distinguish between lung parenchyma and diseased tissue, indicating that it has strong robustness.

[0082] The experimental results verified the generalization ability of the segmentation model for pulmonary tuberculosis at different pathological stages, providing a reliable basis for subsequent classification studies based on quantitative analysis of lung parenchyma.

[0083] 3. Classification Experiment Analysis

[0084] 3.1. Comparative Experimental Analysis of Backbone Networks

[0085] Lightweight models (ShuffleNet, ShuffleNetV2, MobileNet, MobileNetV2, and SqueezeNet) exhibit slightly lower performance metrics in 3D model classification tasks, primarily because they were designed for operation in environments with limited computational resources. Specifically, these lightweight models significantly reduce the number of parameters and computational complexity by using techniques such as depthwise separable convolutions, channel shuffling, and compression architectures. However, 3D data contains more complex spatial relationships and structural information, requiring a larger receptive field and richer feature representation capabilities, which the shallow structure and simplified feature interaction mechanisms of lightweight models struggle to meet.

[0086] Compared to lightweight models, VGG16 employs a simple linear stacking structure (3×3 convolution kernels), which is easier to understand and implement. However, this structure has significant drawbacks: on the one hand, it has a large number of parameters and high computational complexity, making it difficult to efficiently extract complex spatial relationships in 3D medical images; on the other hand, it lacks structures such as residual connections, making it prone to gradient vanishing or exploding problems, ultimately resulting in poor performance metrics.

[0087] Furthermore, ResNet models with different layers exhibited significant performance differences in classification tasks. ResNet50 and ResNet101 achieved the highest classification performance, with ResNet50 achieving Precision, Recall, F1 Score, and AUC of 0.914, 0.882, 0.898, and 0.930, respectively. ResNet101 slightly outperformed ResNet50 in Recall and F1 Score, at 0.894 and 0.903, respectively. This indicates that ResNet50 and ResNet101 have a good ability to balance precision and recall and can accurately classify most samples. In contrast, ResNet34 is slightly lower than ResNet50 and ResNet101 in Recall (0.878) and F1 Score (0.886), but its Precision reaches 0.894, indicating that it has a certain advantage in reducing the false positive rate. It is worth noting that ResNet18 and ResNet200 performed relatively poorly, with all performance metrics lower than other models. This may indicate that models that are too shallow (such as ResNet18) or too deep (such as ResNet200) may be at risk of under-capturing or overfitting features.

[0088] While ResNet101 offers a slight advantage in classification performance, its model complexity and computational cost are significantly increased, directly leading to higher memory usage and longer inference time, thus increasing costs for practical applications. Therefore, ResNet50 is a more practical choice as the backbone network for this improved network, ensuring near-optimal classification performance while offering significant advantages in computational efficiency and resource consumption. This trade-off makes ResNet50 an optimal solution that balances performance and efficiency, maximizing recognition efficiency and minimizing recognition time.

[0089] Table 2 Comparison of classification performance of different backbone networks

[0090]

[0091]

[0092] 3.2. Performance Analysis of MSGC-ResNet50

[0093] To verify the effectiveness of each improvement method and the superiority of MSGC-ResNet50, this invention uses 3DResNet50 as the baseline model M1, the model with MSAM added as M2, the model with GC added as M3, and M4 with both MSAM and GC added as M4. Ablation experiments were conducted using the same dataset. As shown in Table 3, M2 outperforms the base backbone network M1 in all performance metrics, and M3 also shows improvements over M1. M4, however, performs exceptionally well across all metrics, with an accuracy of 0.941, a recall of 0.902, an F1 score of 0.921, and a ROC of 0.970, representing improvements of 2.5, 2.0, 3.2, and 5.0 percentage points compared to ResNet50, respectively.

[0094] Table 3 Ablation Experiment

[0095] Model Precision Recall F1 Score AUC M1: 3D ResNet50 0.916 0.882 0.899 0.920 M2: 3D ResNet50+MSAM 0.937 0.898 0.917 0.970 M3: 3D ResNet50+GC 0.915 0.886 0.900 0.930 M4: 3D MSGC-ResNet50 0.941 0.902 0.921 0.970

[0096] The model's performance metrics curve during training is as follows: Figure 7 As shown, after the 80th epoch, the model's classification performance gradually stabilized, with precision remaining around 0.94, recall around 0.90, and F1-score around 0.92. This indicates that the model minimized false positives while maintaining a high true positive rate. Figure 8 The ROC curves show that the average AUC for active and inactive PTB is approximately 0.97. This indicates that the model exhibits excellent resolution at different feature thresholds, effectively distinguishing different categories, and demonstrates strong generalization ability and robustness, thus improving the accuracy of CT image recognition for active and inactive PTB.

[0097] 4. Model Validation and Analysis

[0098] In this simulation experiment, 10 cases of active pulmonary tuberculosis (PTB) and 10 cases of inactive PTB were added as reference samples for comparison with the doctors' diagnoses. The comparison results are shown in Table 4. In the 10 cases of active PTB, the ResNet50 model correctly predicted 7 cases and incorrectly predicted 3. In contrast, the MSGC-ResNet50 model correctly predicted 9 cases and incorrectly predicted 1. In the 10 cases of inactive PTB, the ResNet50 model correctly predicted 9 cases and incorrectly predicted 1. In contrast, the MSGC-ResNet50 model correctly predicted 9 cases and made 1 error. ResNet50 tends to favor the classification of inactive PTB because it lacks effective feature extraction for key lesions in active PTB. In contrast, MSGC-ResNet50 shows a more balanced performance, effectively predicting both active and inactive cases. This indicates that the addition of enhancements such as MSAM strengthens the model's ability to extract features and classify samples, enabling it to better handle the complexity of active and inactive PTB.

[0099] Overall, the AI ​​models performed well, with confidence scores mostly exceeding 0.8. In terms of overall accuracy, MSGC-ResNet50 achieved 90%, while ResNet50 reached 80%. These results demonstrate that the improved models not only perform well on the training set but also show application potential in real-world clinical settings. This further validates the model's effectiveness as an auxiliary diagnostic tool, providing reliable support for clinicians.

[0100] Table 4. Hospital Data Test Results

[0101]

[0102] In addition, to further demonstrate the model's performance in real-world diagnostic scenarios, confusion matrices were plotted for the test set of the dataset and the supplementary external test set, as shown below. Figure 9As shown in the figure, in the test set (containing 25 cases of active pulmonary tuberculosis and 25 cases of inactive pulmonary tuberculosis), the model achieved the following performance metrics: 22 true positives (TP), 23 true negatives (TN), 2 false positives (FP), and 3 false negatives (FN), resulting in a precision of 0.917 and a recall of 0.880. In the external test set (containing 10 cases of active pulmonary tuberculosis and 10 cases of inactive pulmonary tuberculosis), the model performed as follows: 9 true positives (TP), 9 true negatives (TN), 1 false positive (FP), and 1 false negative (FN), with a corresponding precision of 0.900 and a recall of 0.900. Comparative analysis shows that the external test set had a slightly lower precision than the test set. This slight difference may stem from the fact that the external test set included cases on the borderline of active tuberculosis with atypical radiographic features. This slight performance degradation is acceptable, indicating that the model has good generalization ability.

[0103] Conclusion: This invention introduces a 3D MSGC-ResNet50 model for identifying active and inactive PTB images. The workflow consists of two stages: In the first stage, CT images and corresponding mask labels are preprocessed, and then the lung parenchyma is segmented using 3D ResUNet. In the second stage, the 3D MSGC-ResNet50 model is used to extract lesion features from the lung parenchyma to facilitate the classification of active and inactive PTB.

[0104] The 3D MSGC-ResNet50 introduces MSAM, addressing the problem in standard 3D ResNet models where the fixed kernel size fails to capture multi-scale features simultaneously, leading to insufficient extraction of key features. Furthermore, without a clear attention mechanism, 3D ResNet models struggle to focus on critical regions, making them more susceptible to interference from irrelevant areas and reducing classification accuracy. MSAM combines multi-scale (e.g., 1×1×1, 3×3×3, 5×5×5) convolutions with attention mechanisms, enabling the network to more effectively capture multi-scale lesion features and dynamically focus on key regions. This, in turn, improves the model's classification performance. Additionally, to extract compressed key features more efficiently, this paper introduces grouped convolution and channel shuffling (GC). The introduction of grouped convolution not only further extracts features but also has lower computational complexity compared to traditional convolution, saving computational resources and significantly improving CT image classification efficiency.

[0105] Experimental results show that the 3D MSGC-ResNet50 model exhibits superior performance, with an ROC of 0.970, recall of 0.902, precision of 0.941, and F1 score of 0.921. These results represent improvements of 5.0, 2.0, 2.5, and 3.2 percentage points, respectively, compared to the standard 3D ResNet50 model. Furthermore, 10 cases of active and inactive PTB were used as external test sets, and the results were compared with physician diagnoses. In the external test set of 20 cases, the improved model made 2 incorrect predictions, achieving an accuracy of 90%. This demonstrates that the improved model not only performs well on the training set but also shows great potential in practical applications. This further validates the effectiveness of the model as an auxiliary diagnostic tool, providing reliable support for physicians, improving the effectiveness and accuracy of active pulmonary tuberculosis classification and identification, and enhancing the efficiency of early diagnosis of pulmonary tuberculosis.

Claims

1. Improved 3D network based activity / inactivity tuberculosis identification, characterized in that: The method comprises the following steps: An image feature acquisition module is configured to acquire lung tuberculosis CT chest image features; A data preprocessing module is configured to preprocess the acquired lung tuberculosis CT chest image features; An image segmentation module is configured to segment the preprocessed lung tuberculosis CT chest image features to acquire lung parenchyma image features, and the segmentation model used by the image segmentation module is a 3D ResUNet segmentation model; An image classification and recognition module is configured to predict and classify the lung parenchyma image features after image segmentation; wherein the image classification model used by the image classification and recognition module is a 3D ResUNet50 classification model, and the 3D ResUNet50 classification model is combined with a multi-scale attention module to obtain image output features; then, the image output features are processed by using grouped convolution and channel shuffling to predict the categories of active and non-active pulmonary tuberculosis, and generate the confidence score of the model; 2. The improved 3D network based active / inactive TB identification system as claimed in claim 1, wherein: The implementation methods of conv2-conv5 of the 3D ResUNet50 classification model are the same, and the implementation method of conv2 of the 3D ResUNet50 classification model is as follows: a multi-scale attention mechanism is introduced into the residual structure of the 3D ResNet, the multi-scale attention mechanism first captures the lesion features of multiple scales through multi-scale convolution operation, then the output characteristics of the multi-scale convolution are weighted, in addition, a single-channel spatial attention map is generated by using 1×1×1 convolution and Sigmoid activation function to dynamically adjust the attention of the model to different spatial positions in the feature map; finally, the elements of the input feature map are multiplied by the attention map, and residual connection is used to maintain the integrity of the original features while highlighting the key regional features. The data preprocessing module preprocesses the original CT image, including normalization, binarization, size adjustment and threshold adjustment, and adjusts the window width and window level of the adjusted CT image to [1400, -500], then adjusts the CT image size to [128, 256, 256], then normalizes the pixel value of the image to a fixed range [0, 1], and finally, the normalized image is enhanced by flipping, rotating, scaling and cropping to obtain the preprocessed lung tuberculosis CT chest image features.

3. The improved 3D network based active / non-active tuberculosis identification system as claimed in claim 2, wherein: The implementation method of the 3D ResUNet segmentation model is: when using a CT image feature with a size of 128*256*256 as input, processing is performed through an encoder, a bottleneck layer and a decoder, and a segmentation probability map with the same resolution as the input is generated multiple times, wherein the encoder uses residual blocks and three-dimensional maximum pooling to gradually down-sample the data, reducing the spatial size from 128*256*256 to 64*128*128, 32*64*64, 16*32*32 and 8*16*16; at the same time, the number of feature channels is also increased from 1 to 32, 64, 128, 256 and 512, in the bottleneck layer, the network captures global context information through deep dialogue, and then the decoder uses transposed convolution to gradually up-sample, restoring the spatial dimension from 8*16*16 to 16*32*32, 32*64*64, 64*128*128 and 128*256*256; in each decoding stage, the skip connection connects the low-level features of the corresponding encoder stage with the decoder, integrates global semantics and local details, and finally, a 1*1*1 convolutional full connection layer is mapped to a voxel segmentation probability map through a Sigmoid activation function, ensuring that the output dimension matches the input.

4. The improved 3D network based active / inactive TB identification system as claimed in claim 1, wherein: The implementation method of the grouped convolution and channel shuffling is as follows: assuming that the shape of an input feature map is , wherein C is the number of channels, D is the depth, H is the height, W is the width, the size of a convolution kernel is , the number of output channels is N, in the grouped convolution, assuming that the input channels are divided into G groups, the number of channels in each group is , and a three-dimensional convolution operation is independently performed on each group of channels, and the calculation complexity of each group is as shown in formula (2): (2) (2), The output feature maps of the G group are spliced in the channel to obtain the final feature map, and the computational complexity thereof is as shown in formula (3): (3), The channel shuffling rearranges the channel order of the feature map through three steps of reorganization, transposition and flattening, so that the features of different groups are mixed and interacted.

5. An improved 3D network based active / inactive tuberculosis identification method, characterized in that: The method comprises: obtaining a pulmonary tuberculosis CT chest image feature; performing data preprocessing on the obtained pulmonary tuberculosis CT chest image feature; segmenting the preprocessed pulmonary tuberculosis CT chest image feature to obtain a lung parenchyma image feature, and the segmentation model used by the image segmentation module is a 3D ResUNet segmentation model; performing prediction classification and recognition on the lung parenchyma image feature after image segmentation; wherein the image classification and recognition module uses a 3D ResNet50 classification model, and obtains image output features by combining a 3D ResUNet50 classification model with a multi-scale attention module; then, the image output features are processed by using group convolution and channel shuffling, the categories of active and non-active pulmonary tuberculosis are predicted, and the confidence score of the model is generated; The implementation method of conv2-conv5 of the 3D ResUNet50 classification model is the same, and the implementation method of conv2 of the 3D ResUNet50 classification model is as follows: a multi-scale attention mechanism is introduced into the residual structure of the 3D ResNet, in which the multi-scale attention mechanism firstly captures multiple scale lesion characteristics through multi-scale convolution operation, then the output characteristics of the multi-scale convolution are weighted, in addition, a single-channel spatial attention map is generated by using 1x1x1 convolution and Sigmoid activation function to dynamically adjust the attention of the model to different spatial positions in the feature map; finally, the elements of the input feature map are multiplied by the attention map, and the residual connection is used to maintain the integrity of the original features while highlighting the key regional features.