A fully automatic system for analyzing body tissue composition in chest CT

Through fully automatic analysis of the body tissue composition system of chest CT and using ResNet-50 and 3D U-Net+ models for slice positioning and segmentation, the problem of inability to efficiently and accurately quantify muscle and adipose tissue in chest CT images in the prior art is solved, and efficient quantitative analysis and early disease diagnosis are achieved.

CN119446431BActive Publication Date: 2025-08-15SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411428461.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-08-15
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

In the prior art, when processing large amounts of chest CT images, automation methods either focus only on improving the accuracy of chest CT image segmentation networks, or use fully automatic selection and segmentation neural networks previously developed and verified for L3 slices, and cannot efficiently and accurately quantify individual contributions to muscle and adipose tissue.

Method used

The fully automatic analysis of chest CT body tissue composition system is adopted, including positioning module, segmentation module and automatic measurement module. The ResNet-50 model is used for slice positioning, combined with the 3D U-Net+ model for segmentation of skeletal muscle and adipose tissue, and automated measurement is achieved through deep learning network.

Benefits of technology

It improves the efficiency and accuracy of quantitative analysis of muscle and adipose tissue in chest CT images, supports the early diagnosis and treatment of skeletal muscle system diseases, and improves the efficiency and accuracy of clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446431B_ABST
    Figure CN119446431B_ABST
Patent Text Reader

Abstract

This invention discloses a fully automated system for analyzing body tissue composition in chest CT scans, relating to the technical field of medical diagnostic equipment. The system automatically selects L1 slices that meet the requirements and quantitatively analyzes the muscle and fat area and density at the target level. The system specifically comprises three modules: a positioning module for locating the target L1 slice on the chest CT image; a segmentation module for segmenting the skeletal muscle and adipose tissue corresponding to the L1 axial slice on the chest CT image; and an automatic measurement module for automatically measuring body composition in clinical chest CT images. In summary, this invention not only improves the efficiency and accuracy of measuring body composition on chest CT scans, but also fills a gap in previous automated screening of single L1 slices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical diagnostic equipment, and in particular to a fully automatic system for analyzing body tissue components in chest CT. Background Art

[0002] Assessing body tissue composition throughout the body is important for a variety of clinical and research applications. It is widely accepted that body composition can independently influence health. For example, reduced skeletal muscle mass is associated with poorer quality of life, increased risk of falls and fractures, and has been considered a strong predictor of cardiovascular events and mortality. Reduced skeletal muscle mass is an independent risk factor for prolonged hospitalization in elderly patients with proximal femoral fractures and for mortality in patients receiving mechanical ventilation in the intensive care unit (ICU), colorectal cancer patients, and renal transplant recipients. Studies have demonstrated that adipose tissue mass can significantly influence the incidence of cardiovascular and metabolic diseases, various cancers, and clinical outcomes in lung transplant recipients. Furthermore, recent studies have demonstrated that the same type of adipose tissue, but distributed in different anatomical locations, such as subcutaneous adipose tissue and visceral adipose tissue, can have different effects on health and disease, necessitating the separate quantification and analysis of these body components. Therefore, an accurate, practical, and efficient method for quantifying body composition could have profound clinical benefits.

[0003] Previously, there have been many methods for quantifying body composition, such as anthropometry, including body mass index (BMI), skinfold thickness, waist circumference, etc. These are simple to operate, fast and easy to obtain, and have been widely used to assess obesity. However, these measurement methods do not provide information on the individual contribution of each tissue type to body composition, and are limited in measuring changes in human body composition. Currently, most imaging-based methods require human-computer interaction to quantify different tissues. However, when processing chest CT images of a large number of subjects, interactive methods are obviously impractical. Although related inventions have provided automated measurement software, existing automated methods either focus only on improving the accuracy of chest CT image segmentation networks, or use fully automatic selection and segmentation neural networks previously developed and verified for L3 slices. Summary of the Invention

[0004] To address the deficiencies mentioned in the above background technology, the present invention aims to provide a fully automatic system for analyzing body tissue composition in chest CT, which can automatically select the L1 section that meets the requirements and perform quantitative analysis of the muscle and fat area and density at the target level.

[0005] In a first aspect, the purpose of the present invention can be achieved by the following technical solution: a fully automatic system for analyzing body tissue composition in chest CT, comprising:

[0006] A positioning module, a segmentation module, and an automatic measurement module; the positioning module is used to acquire chest CT scan images, identify and sort the anatomical features of all vertebrae on the chest CT, obtain CT sequence images that may be predicted to be L1, find the sequence with the longest transverse process spacing in this sequence, and use it as the L1 sequence. The processed L1 sequence CT images are input into a pre-established ResNet-50 model to train an L1 positioning model, and the CT image of the optimal L1 slice is sent to the segmentation module, and the L1 positioning model is sent to the automatic measurement module.

[0007] The segmentation module is used to divide the CT image of the best L1 slice into skeletal muscle and fat tissue to obtain the divided L1 slice, input the divided L1 slice into a pre-established 3D U-Net+ model, output the optimal segmentation model, and send the optimal segmentation model to the automatic measurement module;

[0008] The automatic measurement module is used to acquire a chest CT image to be processed, input the chest CT image to be processed into the L1 positioning model to obtain a single optimal LI slice, input the single optimal LI slice into the optimal segmentation model to obtain skeletal muscle, subcutaneous fat, and visceral fat, and use the skeletal muscle, subcutaneous fat, and visceral fat to calculate the quantity and average muscle attenuation to obtain body tissue composition measurement results.

[0009] In conjunction with the first aspect, in certain implementations of the first aspect, the system further includes: a process in which the positioning module acquires a chest CT scan image:

[0010] Chest CT slices that have undergone random vertical flipping and affine transformation are input into the localization model. The model automatically identifies all anatomical features of the cervical, thoracic, and lumbar vertebrae, and selects a possible L1 sequence based on the unique anatomical features of the L1 vertebra. The CT sequence with the longest intertransverse process spacing is then selected from the L1 sequence as the optimal L1 sequence. Finally, the localization model performs further processing.

[0011] In conjunction with the first aspect, in certain implementations of the first aspect, the system further includes: a process in which the positioning module processes the CT image of the best L1 slice:

[0012] The data was processed on the Verse20 dataset. The CT image of the best L1 slice and the slices connected above and below it were used as positive samples, and the remaining slices were randomly sampled as negative samples, so that the ratio of positive samples to negative samples was 1:5.

[0013] In conjunction with the first aspect, in certain implementations of the first aspect, the system further includes: a pre-established ResNet-50 model of the positioning module is as follows:

[0014] The expression of ResNet-50 can be expressed as:

[0015] [H(x)=F(x,{W_i})+x]

[0016] in:

[0017] (H(x)) is the output of the network.

[0018] (x) is the input to the network.

[0019] (F(x,{W_i})) is the residual function, which represents the result of the input (x) after being transformed through a series of layers, and ({W_i}) represents the set of weights in the network.

[0020] (+x) is a skip connection that adds the input (x) directly to the output of the residual function (F(x,{W_i})).

[0021] In combination with the first aspect, in certain implementations of the first aspect, the system further includes: the segmentation module divides the CT image of the best L1 slice into skeletal muscle and fat tissue, including background, skeletal muscle, subcutaneous fat and visceral fat, and the division is based on the spatial distribution of skeletal muscle and fat tissue and tissue radiation attenuation characteristics.

[0022] In combination with the first aspect, in certain implementations of the first aspect, the system further includes: the pre-established 3D U-Net+ model of the segmentation module includes a 3D encoder, a 3D decoder, and an output layer.

[0023] In combination with the first aspect, in certain implementations of the first aspect, the system further includes: the 3D encoder is composed of convolution blocks, each convolution block includes two 3D convolution layers and a 3D maximum pooling layer, after the convolution layer, each convolution block is downsampled through a 2×2×2 maximum pooling layer, and the output of each convolution block is transmitted to the corresponding 3D decoder part through a skip connection;

[0024] The 3D decoder consists of convolution blocks, each of which includes a 3D upsampling layer and two 3D convolution layers. After upsampling, each convolution block uses two convolution kernels and is also equipped with BN and ReLU activation functions. The upsampled feature map is spliced with the skip connection from the corresponding convolution block of the encoder.

[0025] The output layer adopts a 3D convolutional layer to map the output of the last convolutional block of the 3D decoder to the number of target categories.

[0026] In conjunction with the first aspect, in certain implementations of the first aspect, the system further includes: a loss function constructed by the pre-established 3DU-Net+ model is as follows:

[0027] For a chest CT image I of size H×W, a multi-class cross entropy loss function is combined with a Dice loss function to segment the background, skeletal muscle, subcutaneous fat, and visceral fat, as shown in formula (1):

[0028]

[0029] Wherein, C = 4 represents background, skeletal muscle, subcutaneous fat, and visceral fat;

[0030] ω c represents the weight of the c-th type of organization;

[0031] Indicates the gold standard value of pixel i belonging to the c-th category label;

[0032] Indicates the prediction result of pixel i as the c-th class label;

[0033] H and W represent the height and width of the 2D axial image, respectively.

[0034] In combination with the first aspect, in some implementations of the first aspect, the system further includes: counting the pixel ratios of skeletal muscle, subcutaneous fat, visceral fat, and background, and then setting a small weight for skeletal muscle with a small ratio, and setting a large weight for subcutaneous and visceral fat with a large ratio, as shown in formula (2).

[0035]

[0036] Where H, W and D represent the height, width and depth of the two-dimensional image;

[0037] N c Represents the pixel count statistics of the c-th label.

[0038] In conjunction with the first aspect, in certain implementations of the first aspect, the system further includes: the process of calculating the amount and average muscle attenuation using skeletal muscle, subcutaneous fat, and visceral fat is as follows:

[0039] Calculation expression for skeletal muscle, subcutaneous fat, and visceral fat area: [A = sum_{i = 1}^{N}p_i\timesa_i]

[0040] Where (p_i) is the classification probability that pixel (i) belongs to a specific tissue, and (a_i) is the area of pixel (i).

[0041] Calculation of average muscle attenuation values for skeletal muscle, subcutaneous fat, and visceral fat:

[0042] For each pixel, if it is classified as skeletal muscle, its CT value is recorded, and the average of the CT values of all skeletal muscle pixels is calculated.

[0043] Expression: [mu = frac{sum_{i = 1}^{M}v_i}{M}]

[0044] Where (v_i) is the CT value of pixel (i) and (M) is the total number of skeletal muscle pixels.

[0045] Beneficial effects of the present invention:

[0046] The present invention fully automatically analyzes chest CT images through a deep learning network and can accurately locate the target L1 slice, which is crucial for the subsequent diagnosis of skeletal muscle system diseases. Secondly, the ResNet-50 classification model is used for slice positioning to ensure the speed and accuracy of positioning and improve the overall performance of the system. Thirdly, the 3D U-Net+ architecture is used for image segmentation, which has the advantages of a simple network structure, easy training and optimization. At the same time, the jump connection retains more image details and improves the accuracy of segmentation. These improvements are conducive to the early diagnosis and treatment of skeletal muscle system diseases (such as sarcopenia and muscle fatty degeneration), and improve the efficiency and accuracy of clinical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0048] Figure 1 This is the overall framework diagram of the system of the present invention;

[0049] Figure 2 The best first lumbar vertebra slice selected from the chest CT image of the present invention;

[0050] Figure 3 This is a schematic diagram of the positioning model training of the present invention;

[0051] Figure 4 This is a schematic diagram of the division of skeletal muscle, subcutaneous fat and visceral fat in the present invention;

[0052] Figure 5 Schematic diagram of the 3D segmentation network of the present invention;

[0053] Figure 6 The following are the muscle and fat segmentation results of 3 physical examination subjects in the present invention;

[0054] Figure 7 This is a flow chart of the application of the system of the present invention. DETAILED DESCRIPTION

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0056] Example 1:

[0057] like Figure 1 As shown, a fully automatic system for analyzing chest CT body tissue composition includes:

[0058] Positioning module, segmentation module and automatic measurement module;

[0059] The positioning module is used to collect chest CT scan images, perform anatomical feature recognition and positioning on the chest CT scan images, obtain multiple CT images predicted to be L1, process the multiple CT images predicted to be L1 slices, obtain multiple processed CT images that may be the best L1 slices, input the multiple processed CT images that may be the best L1 slices into a pre-established ResNet-50 model, train to obtain an L1 positioning model, locate and determine the multiple CT images that may be the best L1 slices based on the L1 positioning model, obtain a single CT image of the best L1 slice, send the single CT image of the best L1 slice to the segmentation module, and send the L1 positioning model to the automatic measurement module;

[0060] First, the patient underwent a chest CT scan, resulting in images of the cervical and thoracic spine, including the entire L1. The optimal L1 slice was then determined. Data was first processed on the Verse20 dataset, with the optimal L1 slice and several adjacent slices selected as positive samples. The remaining slices were randomly sampled as negative samples, maintaining a positive-to-negative sample ratio of 1:5. The resulting L1 localization model was then trained on the ResNet-50 model. The slice with the highest predicted probability from the 5-15 best slices was selected as the final L1 slice. The L1 localization model was tested on the test set, with the results shown in Table 1.

[0061] Table 1 Test results of the positioning model on 248 healthy people undergoing physical examinations

[0062]

[0063] Specifically, the specific operation process of the positioning module:

[0064] Positioning preprocessing: Data was processed on the Verse20 dataset. The best L1 slice and several slices connected above and below it were used as positive samples. The remaining slices were randomly sampled (10 slices were randomly selected) as negative samples, ensuring a positive-to-negative sample ratio of 1:5.

[0065] The localization model obtained by training the ResNet-50 model on the first lumbar vertebra slices of 200 physical examination subjects was used, and the slice with the highest prediction probability was selected from the 5-15 best slices as the final L1 slice. Figure 3 .

[0066] The pre-built ResNet-50 model of the positioning module is as follows:

[0067] The expression of ResNet-50 can be expressed as:

[0068] [H(x)=F(x,{W_i})+x]

[0069] in:

[0070] (H(x)) is the output of the network.

[0071] (x) is the input to the network.

[0072] (F(x,{W_i})) is the residual function, which represents the result of the input (x) after being transformed through a series of layers, and ({W_i}) represents the set of weights in the network.

[0073] (+x) is a skip connection that adds the input (x) directly to the output of the residual function (F(x,{W_i})).

[0074] In ResNet-50, the residual function (F(x,{W_i})) is usually composed of multiple residual blocks, each of which can be expressed as:

[0075] [F(x)=text{ReLU}(text{BN}(text{Conv}(x))+text{BN}(text{Conv}(text{ReLU}(text{BN}(text{Conv}(x))))))

[0076] in:

[0077] (text{ReLU}) is the activation function.

[0078] (text{BN}) stands for batch normalization.

[0079] (text{Conv}) is the convolution operation.

[0080] Each residual block usually contains 3 convolutional layers, where the stride of the first and third convolutional layers is 1 to maintain the size of the feature map, and the stride of the second convolutional layer is 2 for downsampling.

[0081] The structure of the ResNet-50 model is roughly as follows:

[0082] 1. A 7x7 convolutional layer with a stride of 2, followed by a max pooling layer.

[0083] 2. Three residual block sequences, each sequence contains several residual blocks.

[0084] 3. The number of residual blocks in each sequence is 3, 4, and 6 respectively.

[0085] 4. After each sequence, there is usually an average pooling layer with a stride of 2 for downsampling.

[0086] 5. The last layer is a fully connected layer for classification.

[0087] The segmentation module is used to divide the CT image of the best L1 slice into skeletal muscle and fat tissue to obtain the divided L1 slice, input the divided L1 slice into a pre-established 3D U-Net+ model, output the optimal segmentation model, and send the optimal segmentation model to the automatic measurement module;

[0088] The segmentation module extracts four labels from the automatically selected L1 single slice: background, skeletal muscle, subcutaneous fat, and visceral fat. These four categories are based on the spatial distribution of skeletal muscle and adipose tissue and their radiation attenuation characteristics. The automatically selected L1 axial slices from the chest CT image serve as input for training the segmentation network, and the output is the four-category label for that segmented region. After training, an optimal segmentation model is obtained. The segmentation network utilizes an innovative 3D U-Net+ network architecture proposed in this paper, consisting of two main components: an encoder and a decoder.

[0089] Specifically, the segmentation module

[0090] The slices belonging to L1 obtained by the positioning module are divided into skeletal muscle and fat tissue, including L1 axial skeletal muscle (see Figure 4 Label2), subcutaneous fat and visceral fat (see Figure 4 Label3 and Label4);

[0091] Segmentation preprocessing: All DICOM files were re-set to a window level range of -190 to 150 HU to ensure accurate image information within this range. Data augmentation was also performed, including random XY flips, 90-degree rotations, ±10-degree rotations around the XY axis, and random cropping.

[0092] Constructing the segmentation network architecture: The method is to use the 3D U-Net+ model (see Figure 5 ), a model based on deep learning-based 3D convolutional neural network (CNN) technology, is designed to achieve fully automated organ segmentation. The 3D U-Net+ model inherits and expands the traditional U-Net architecture for medical image segmentation to more effectively extract and integrate multi-scale image feature information. The model consists of two main components: an encoder and a decoder, which work closely together through skip connections to achieve accurate image segmentation.

[0093] in:

[0094] 3D encoder design

[0095] Convolutional block composition: The encoder consists of four carefully designed convolutional blocks, each of which contains two 3D convolutional layers and one 3D maximum pooling layer.

[0096] Max pooling layer: After the convolutional layer, each convolution block is downsampled by a 2×2×2 max pooling layer with a stride of 2 to reduce the spatial dimension of the feature map while retaining the most important feature information;

[0097] Skip connection: To preserve spatial detail information, the output of each convolutional block is passed to the corresponding decoder part through a skip connection, laying the foundation for subsequent feature fusion and detail reconstruction;

[0098] In particular, at the beginning of the second convolutional block, the present invention innovatively introduces a 4×4×4 max pooling layer. This design aims to perform deeper feature extraction on the input data. By selecting the maximum value within a local region, it captures significant features in the image and provides a more abstract and compact feature representation for subsequent processing.

[0099] 3D decoder design

[0100] Convolution block composition: The decoder also consists of four convolution blocks, each of which includes a 3D upsampling layer and two 3D convolution layers.

[0101] Upsampling layer: In the first convolutional block of the decoder, we pass a 2×2×2 upsampling layer to increase the spatial dimension of the feature map and recover the details that may be lost during the downsampling process.

[0102] Convolutional layer details: After upsampling, each convolution block uses two 3×3×3 convolution kernels and is also equipped with BN and ReLU activation functions to further refine the feature representation.

[0103] Feature fusion: By concatenating upsampled feature maps with skip connections from the corresponding convolutional blocks of the encoder, we achieve comprehensive utilization of cross-level features. This fusion strategy not only enhances the model's ability to extract multi-scale features of the input data, but also significantly improves the model's ability to represent complex structures and detailed information.

[0104] When constructing the 3D U-Net+ model, the present invention carefully designs a 4×4×4 max pooling layer at the beginning of the second convolutional block of the encoder, immediately following the first convolutional layer. The introduction of this max pooling layer aims to effectively extract deep features from the input data, capturing the most salient and important information in the image by selecting the maximum value within each local region. This operation not only reduces the dimensionality of the data but also helps the network focus on the most discriminative features, providing a more compact and abstract feature representation for subsequent network layers. Subsequently, the feature map processed by the max pooling layer from the second convolutional block is fused with the feature map from the fourth convolutional block. This feature fusion strategy combines feature information from different convolutional blocks and different layers, achieving comprehensive cross-layer feature utilization. This integration not only enhances the model's ability to capture multi-scale features of the input data but also improves its ability to represent complex structures and detailed information.

[0105] Output layer design

[0106] The output layer uses a 3D convolutional layer, which maps the output of the decoder's last convolutional block to the number of target categories. Specifically, a 1×1×1 convolution kernel is used to reduce the number of output channels to a dimension that matches the number of target categories, and a softmax activation function is used to output the probability of each voxel belonging to a different category. In this paper, the output layer can accurately distinguish between background, skeletal muscle, subcutaneous fat, and visceral fat tissue.

[0107] Constructing the loss function

[0108] For a chest CT image I of size H×W, the commonly used multi-class cross entropy loss function in medical image segmentation tasks is combined with the Dice loss function to segment the background, skeletal muscle, subcutaneous fat, and visceral fat, as shown in formula (1).

[0109]

[0110] Wherein, C = 4 represents background, skeletal muscle, subcutaneous fat, and visceral fat;

[0111] ω c represents the weight of the c-th type of organization;

[0112] Indicates the gold standard value of pixel i belonging to the c-th category label;

[0113] Indicates the prediction result of pixel i being the c-th class label.

[0114] H and W represent the height and width of the 2D axial image, respectively.

[0115] The size of skeletal muscle, subcutaneous fat, visceral fat and background is very different, which means that there is a class imbalance problem that causes the segmentation framework to be unstable. Therefore, during the training process, in the loss function, it is necessary to penalize low-confidence predictions (such as Label2) by setting weights. Specifically, first, in the training image, the pixel ratios of skeletal muscle, subcutaneous fat, visceral fat and background are counted, and then the skeletal muscle with a small ratio is set to a small weight, and the subcutaneous and visceral fat with a large ratio are set to a large weight, as shown in formula (2).

[0116]

[0117] Where H, W and D represent the height, width and depth of the two-dimensional image;

[0118] N c Represents the pixel count statistics of the c-th label;

[0119] Therefore, from ω c The prior statistics of ensure the class-balanced optimization of the loss function.

[0120] The automatic measurement module is used to acquire a chest CT image to be processed, input the chest CT image to be processed into the L1 positioning model to obtain a single optimal LI slice, input the single optimal LI slice into the optimal segmentation model to obtain skeletal muscle, subcutaneous fat, and visceral fat, and use the skeletal muscle, subcutaneous fat, and visceral fat to calculate the quantity and average muscle attenuation to obtain body tissue composition measurement results.

[0121] The automatic measurement module leverages the localization and segmentation models trained in the first two modules. It takes a new chest CT image and its corresponding L1 position as input, along with pre-set HU intervals for skeletal muscle (-29, 150), visceral fat (-150, -50), and subcutaneous fat (-190, -30). Using the deep learning model, it calculates the amount and average HU of skeletal muscle, subcutaneous fat, and visceral fat. Ultimately, this allows for automated body composition measurement.

[0122] The automatic measurement module includes:

[0123] After positioning preprocessing, the new chest CT image is input into the trained optimal positioning model to obtain a single optimal L1 slice;

[0124] The obtained single best L1 slice was pre-processed for segmentation and then input into the segmentation optimization model to obtain skeletal muscle, subcutaneous fat, and visceral fat;

[0125] At the same time, the single best L1 slice is thresholded to obtain the corresponding areas of skeletal muscle, subcutaneous fat, and visceral fat corresponding to L1;

[0126] The amount and average muscle attenuation of the obtained skeletal muscle, subcutaneous fat and visceral fat are calculated to achieve automatic measurement of body composition.

[0127] The process for calculating the amount and average muscle attenuation using skeletal muscle, subcutaneous fat, and visceral fat is as follows:

[0128] Calculation of skeletal muscle, subcutaneous fat and visceral fat area:

[0129] For each pixel, if its classification probability exceeds a certain threshold (e.g. 0.5), the pixel is considered to belong to the tissue. The total number of pixels in each category is calculated and then multiplied by the physical area of the pixel to get the number of each category. Expression: [A = sum_{i = 1}^{N}p_i\times a_i]

[0130] Where (p_i) is the classification probability that pixel (i) belongs to a specific tissue, and (a_i) is the area of pixel (i).

[0131] Calculation of average muscle attenuation values for skeletal muscle, subcutaneous fat, and visceral fat:

[0132] For each pixel, if it is classified as skeletal muscle, its CT value is recorded, and the average of the CT values of all skeletal muscle pixels is calculated.

[0133] Expression: [mu = frac{sum_{i = 1}^{M}v_i}{M}]

[0134] Where (v_i) is the CT value of pixel (i) and (M) is the total number of skeletal muscle pixels.

[0135] Specifically, the present invention will be further described below through examples:

[0136] The present invention is further described below with reference to examples and a system overall framework diagram.

[0137] Taking the clinical original images collected from the clinic as an example, the application process of the present invention for automatically measuring body components is as follows: Figure 7 shown.

[0138] Module 1 involves the selection of L1 axial slices. First, a clinical CT image set of 248 healthy subjects undergoes positioning preprocessing. This involves saving the 3D dataset as images with the extension "DICOM" and then performing a random vertical flip and affine transformation. All slices are then fed into the positioning model trained on a dataset of 200 healthy subjects. The predicted L1 slices are shown in Table 1. Accuracy represents precision, allowable error represents allowable error, and average error represents average error.

[0139] Module 2 involves the segmentation of skeletal muscle, subcutaneous fat, and visceral fat. First, the slices corresponding to L1 located in module 1 undergo segmentation preprocessing. This involves setting the window level of all DICOM files to the range of [-190HU, 150HU] and performing data augmentation operations, including random XY flips, 90-degree rotations, ±10-degree rotations around the XY axis, and random cropping. The images are then fed into a segmentation model trained on a dataset of 248 healthy subjects undergoing physical examinations to obtain segmentation results for skeletal muscle, subcutaneous fat, and visceral fat.

[0140] Table 1 and Table 2 show the test results of the positioning model and segmentation model for the data of 248 healthy subjects undergoing physical examinations. In the positioning model, the slice corresponding to L1 is used as the gold standard for training. The higher the accuracy obtained in Table 1, the better the positioning effect. In the segmentation model, skeletal muscle, subcutaneous fat and visceral fat are used as the gold standard for training. The higher the segmentation indicators Iou, Precision, Recall and Dice, the better, and the smaller the SD, the better. Table 2 calculates the segmentation indicators of skeletal muscle, subcutaneous fat and visceral fat respectively. Here, Iou is an indicator that measures the degree of overlap between the predicted results and the true labels in the image segmentation task, Precision indicates accuracy, Recall is an indicator that measures the model's ability to recognize four types of labeled samples, Dice is a commonly used indicator for evaluating image segmentation accuracy, and SD is the standard deviation.

[0141] Table 2 Test results of the segmentation model on 248 healthy people undergoing physical examinations

[0142]

[0143]

[0144] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.

[0145] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium having a computer program stored thereon, which executes the above method when executed by a processor. The storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.

[0146] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present disclosure. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0147] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present disclosure. Various changes and improvements may be made to the present disclosure without departing from the spirit and scope of the present disclosure, and such changes and improvements shall fall within the scope of the present disclosure.

Claims

1. A fully automatic system for analyzing chest CT body tissue composition, characterized by: include: Positioning module, segmentation module and automatic measurement module; The positioning module is used to collect chest CT scan images, identify and sort the anatomical features of all vertebrae in the chest CT, obtain CT sequence images predicted to be L1, obtain the sequence with the longest transverse process spacing in the sequence as the L1 sequence CT image, input the L1 sequence CT image into a pre-established ResNet-50 model, train to obtain an L1 positioning model, determine the positioning of the L1 sequence CT image based on the L1 positioning model, obtain the CT image of the best L1 slice, send the CT image of the best L1 slice to the segmentation module, and send the L1 positioning model to the automatic measurement module; The segmentation module is used to divide the CT image of the best L1 slice into skeletal muscle and fat tissue to obtain the divided L1 slice, input the divided L1 slice into a pre-established 3D U-Net+ model, output the optimal segmentation model, and send the optimal segmentation model to the automatic measurement module; The pre-built 3D U-Net+ model of the segmentation module includes a 3D encoder, a 3D decoder, and an output layer; The 3D encoder consists of convolution blocks, each of which contains two 3D convolutional layers and a 3D maximum pooling layer. After the convolution layer, each convolution block is downsampled by a 2×2×2 maximum pooling layer, and the output of each convolution block is passed to the corresponding 3D decoder part through a skip connection. The 3D decoder consists of convolution blocks, each of which includes a 3D upsampling layer and two 3D convolution layers. After upsampling, each convolution block uses two convolution kernels and is also equipped with BN and ReLU activation functions. The upsampled feature map is spliced with the skip connection from the corresponding convolution block of the encoder. The output layer uses a 3D convolutional layer to map the output of the last convolutional block of the 3D decoder to the number of target categories; The loss function of the pre-built 3D U-Net+ model is as follows: For a chest CT image I of size H×W, a multi-class cross entropy loss function is combined with a Dice loss function to segment the background, skeletal muscle, subcutaneous fat, and visceral fat, as shown in formula (1): Wherein, C = 4 represents background, skeletal muscle, subcutaneous fat, and visceral fat; ω c represents the weight of the c-th type of organization; Indicates the gold standard value of pixel i belonging to the c-th category label; Indicates the prediction result of pixel i as the c-th class label; H and W represent the height and width of the 2D axial image, respectively; The pixel ratios of skeletal muscle, subcutaneous fat, visceral fat and background are counted, and then the skeletal muscle with a small ratio is given a small weight, and the subcutaneous and visceral fat with a large ratio is given a large weight, as shown in formula (2). Where H, W and D represent the height, width and depth of the two-dimensional image; N c Represents the pixel count statistics of the c-th label; The automatic measurement module is used to acquire a chest CT image to be processed, input the chest CT image to be processed into the L1 positioning model to obtain a single optimal LI slice, input the single optimal LI slice into the optimal segmentation model to obtain skeletal muscle, subcutaneous fat and visceral fat, and calculate the amount and average muscle attenuation of skeletal muscle, subcutaneous fat and visceral fat to obtain body tissue composition measurement results; The process of calculating the amount and average muscle attenuation using skeletal muscle, subcutaneous fat and visceral fat is as follows: The calculation expressions for the areas of skeletal muscle, subcutaneous fat and visceral fat are: [A=sum_{i=1}^{N}p_i\times a_i] Where (p_i) is the classification probability that pixel (i) belongs to a specific tissue, (a_i) is the area of pixel (i); Calculation of average muscle attenuation values for skeletal muscle, subcutaneous fat, and visceral fat: For each pixel classified as skeletal muscle, record the CT value and calculate the average CT value of all skeletal muscle pixels; expression: [mu = frac{sum_{i = 1}^{M}v_i}{M}] Where (v_i) is the CT value of pixel (i) and (M) is the total number of skeletal muscle pixels.

2. The fully automatic chest CT body tissue composition analysis system according to claim 1 is characterized in that: The process of the positioning module collecting chest CT scan images: Chest CT slices that have undergone random vertical flipping and affine transformation are input into the positioning model. The model automatically identifies all anatomical features of the cervical, thoracic, and lumbar vertebrae, and selects the sequence that may be L1 based on the unique anatomical features of the L1 vertebra. Then, the CT sequence with the longest intertransverse process spacing is selected from the L1 sequence as the optimal L1 sequence, which is then further processed by the positioning model.

3. The fully automatic chest CT body tissue composition analysis system according to claim 2, characterized in that: The positioning module processes the CT image of the best L1 slice: The data was processed on the Verse20 dataset. The CT image of the best L1 slice and the slices connected above and below it were used as positive samples, and the remaining slices were randomly sampled as negative samples, so that the ratio of positive samples to negative samples was 1:

5.

4. The fully automatic chest CT body tissue composition analysis system according to claim 1, characterized in that: The pre-built ResNet-50 model of the positioning module is as follows: [H(x)=F(x,{W_i})+x] in: (H(x)) is the output of the network; (x) is the input of the network; (F(x,{W_i})) is the residual function, which represents the result of the input (x) after being transformed through a series of layers, and ({W_i}) represents the set of weights in the network; (+x) is a skip connection that adds the input (x) directly to the output of the residual function (F(x,{W_i})).

5. The fully automatic chest CT body tissue composition analysis system according to claim 1, characterized in that: The segmentation module divides the CT image of the best L1 slice into skeletal muscle and fat tissue, including background, skeletal muscle, subcutaneous fat and visceral fat, and the division is based on the spatial distribution of skeletal muscle and fat tissue and tissue radiation attenuation characteristics.

Citation Information

Patent Citations

  • Body composition automatic measurement system based on abdomen CT image and deep learning

    CN114305473A

  • Physical examination CT image data processing and analyzing system and application thereof

    CN117788435A