Abdominal muscle and fat density and area calculation method and system

The method automates the classification and segmentation of abdominal muscles and fat in CT images using deep learning models, addressing labor-intensive and error-prone manual processes to enhance precision in density and area calculations.

CN120318197APending Publication Date: 2025-07-15XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510470766.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, abdominal CT image analysis is time-consuming and labor-intensive, making it difficult to automatically screen the best third lumbar vertebra image, and the division of abdominal muscles and fat, density and area calculations are cumbersome and inaccurate.

Method used

The DenseNet121 network architecture is used for lumbar vertebra classification and vertebrae judgment, combined with the ResWAA_DNet deep network segmentation model, and connected with the cross-layer using the wide-area attention mechanism, filter the best L3 images through the optimal L3 positioning algorithm, and combine the binarized mask density algorithm and the Monte Carlo integral area algorithm for automatic segmentation and calculation.

Benefits of technology

Automatic classification of abdominal CT images and intelligent positioning of optimal L3 images are realized, which improves the accuracy of abdominal muscle and fat segmentation and the accuracy of density and area calculations, and reduces the complexity and time consumption of manual operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318197A_ABST
    Figure CN120318197A_ABST
Patent Text Reader

Abstract

The invention discloses an abdominal muscle and fat density and area calculation method and system. The method comprises the following steps: training to obtain a lumbar vertebra classification model and a vertebral process judgment model, performing lumbar vertebra classification on an abdominal region CT image by using the lumbar vertebra classification model, inputting an image of which the prediction category is L3 into the vertebral process judgment model to find an L3 image with a vertebral process, and screening an optimal L3 image; constructing a ResWAADNet model to learn mapping from an abdominal CT image to a semantic segmentation mask image thereof, emphasizing the boundary of a target area based on an edge detection operator to optimize the boundary segmentation effect of the model on the target area, and adding a wide area attention mechanism and cross-layer connection to obtain an optimal semantic segmentation result of abdominal muscle, subcutaneous fat and visceral fat of an L3 image; and obtaining the density and area of the region according to the segmentation result. According to the method, the advanced levels of automatic lumbar vertebra classification of the abdominal CT image, automatic positioning of the optimal L3 image and automatic segmentation of abdominal muscles and fat are achieved, and the accuracy of the density and area calculation result of the related region is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Generally, after obtaining the abdominal CT images collected by a medical device, the abdominal muscles and related regions on the CT images at the third lumbar vertebra level (L3) are analyzed, and the images are manually annotated to segment regions such as muscles and fat. Then, relevant medical software is used to calculate the segmented regions to obtain parameters related to the diagnosis of sarcopenia, such as the density and area of skeletal muscle and fat regions. The relevant process is very energy-consuming, and problems such as missed diagnosis and misdiagnosis may occur. Specifically:

[0003] First, CT images with medical analysis value are usually in the third lumbar vertebra region. For a single patient, there are hundreds of CT images at a time. When doctors annotate, they need to manually screen out the CT images in the third lumbar vertebra region, which is very time-consuming and laborious.

[0004] Second, in the past, the segmentation of regions related to sarcopenia was all manually performed based on rich experience. Muscles and related regions have characteristics such as complex structures and scattered positions compared to other tissues and organs. The manual annotation method is very cumbersome and requires doctors with certain experience to perform the annotation.

[0005] Third, doctors need to put the segmented images into relevant medical software to calculate key index information such as density and area. Among them, file saving and input are relatively cumbersome and require certain software operation knowledge. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a method and system for calculating the density and area of abdominal muscles and fat in view of the deficiencies in the above-mentioned prior art, so as to solve the technical problems that traditional methods cannot automatically classify abdominal CT images, are difficult to automatically screen the best third lumbar vertebra images, and cannot achieve automatic and accurate segmentation of abdominal muscles and fat and intelligent calculation of their density and area.

[0007] The present invention adopts the following technical solutions:

[0008] A method for calculating the density and area of abdominal muscles and fat includes the following steps:

[0009] Construct a deep classification model based on the DenseNet121 network architecture to classify the lumbar vertebrae and judge the presence or absence of vertebral processes for abdominal region CT images, and screen the best L3 images through the optimal L3 positioning algorithm;

[0010] Construct a ResWAA_DNet deep network segmentation model, use the CNN convolutional layer as the encoder and decoder, add a wide-area attention mechanism and cross-layer connection, and input the screened best L3 images into the ResWAA_DNet deep network segmentation model to obtain the corresponding semantic mask image;

[0011] Based on the best L3 image semantic mask map, the density and area of abdominal muscles, subcutaneous fat, and visceral fat are calculated by combining the density algorithm based on the binary mask and the area algorithm based on Monte Carlo integration.

[0012] Preferably, a deep network classification model based on the DenseNet121 network architecture is constructed to classify the lumbar vertebrae and judge the presence or absence of vertebral processes in the abdominal region CT images. The best L3 image is screened by the best L3 positioning algorithm, specifically as follows:

[0013] The CSV file with 5-class lumbar vertebra type annotation information and the classification training set are input into the DenseNet_5 model built based on the DenseNet121 deep network architecture. The cross-entropy loss function containing mitigation of class imbalance is used to guide the training of the DenseNet_5 model to obtain the lumbar vertebra classification model; the images in the classification test set are input into the lumbar vertebra classification model to obtain the lumbar vertebra classification results, the lumbar vertebra classification results are evaluated, and the confusion matrix and ROC curve of the classification results are drawn;

[0014] The images of the third lumbar vertebra L3 class in the classification dataset are windowed and center-cropped, and the processed images are divided into three categories: with vertebral processes, unclear vertebral processes, and without vertebral processes according to the presence or absence of vertebral processes. The CSV file with three-category presence or absence of vertebral process annotation information and the processed training set are input into the DenseNet_3 model built based on the DenseNet121 deep network architecture. The cross-entropy loss function containing mitigation of class imbalance is used to guide the training of the DenseNet_3 model to obtain the vertebral process judgment model; the images predicted as L3 class in the test set are input into the vertebral process judgment model to obtain the vertebral process judgment results of the L3 class images, the vertebral process judgment results are evaluated, and the confusion matrix and ROC curve of the judgment results are drawn; the best L3 image is screened by using the best L3 positioning algorithm for the L3 images predicted to have vertebral processes to obtain the best L3 image.

[0015] Preferably, an abdominal CT image dataset is constructed. According to the task, the dataset is divided into the abdominal CT image classification dataset class_data1 and the abdominal CT image segmentation dataset seg_data1. Then, through preprocessing and image augmentation operations, the augmented classification dataset Augclass_data1 and the augmented segmentation dataset Augseg_data1 are obtained. Finally, they are divided into training set, validation set, and test set.

[0016] Preferably, traverse the folders where different classes of images are located to generate the corresponding class One-Hot encoding. For the same class of images, the value 1 is assigned to the corresponding class channel, and the value 0 is assigned to the remaining channels; the image names and the corresponding class labels are saved in the CSV file;

[0017] Select the median of the number of 5 types of images from the first lumbar vertebra to the fifth lumbar vertebra as the category weight benchmark i. Divide the number of 5 types of images by i respectively to obtain the category weight coefficients, and add the obtained weight coefficients to the cross-entropy loss function to guide the model training;

[0018] Based on the DenseNet121 deep network framework, establish the lumbar vertebra classification model DenseNet_5, and obtain the predicted category label of the input image through the input layer, dense layer, transition layer and fully connected layer;

[0019] Use the Adam optimizer with weight decay to minimize the loss function to train the DenseNet_5 model and obtain the DenseNet_5 lumbar vertebra classification model;

[0020] Input the images in the test set Augclass_data1_test into the lumbar vertebra classification model DenseNet_5 to obtain the lumbar vertebra classification results of the test set, evaluate the lumbar vertebra classification results, and draw the confusion matrix and ROC curve of the classification results.

[0021] Preferably, perform windowing processing and central cropping on the L3 class images in the Augclass_data1 dataset to obtain images with a size of 224×224. Divide the processed images into three categories: with vertebral processes, unclear vertebral processes, and without vertebral processes according to the presence or absence of vertebral processes. Traverse the folders where different class images are located to generate the One-Hot encoding of the corresponding classes. Assign 1 to the corresponding category channels for the same class images, and assign 0 to the remaining channels; Save the image names and the corresponding category labels in a CSV file;

[0022] Select the median of the number of 3 types of images as the category weight benchmark i. Divide the number of 3 types of images by i respectively to obtain the category weight coefficients, and add the obtained weight coefficients to the cross-entropy loss function to guide the model training;

[0023] Based on the DenseNet121 deep network framework, establish the vertebral process judgment model DenseNet_3, and obtain the predicted category label of the input image through the input layer, dense layer, transition layer and fully connected layer;

[0024] Use the Adam optimizer with weight decay to minimize the loss function in step S302 to train the DenseNet_3 model and obtain the DenseNet_3 vertebral process judgment model;

[0025] Input the images predicted as L3 class in the test set into the DenseNet_3 vertebral process judgment model to obtain the vertebral process judgment results of the L3 class images, evaluate the vertebral process judgment results, and draw the confusion matrix and ROC curve of the classification results;

[0026] By using the fact that the best L3 image occupies more bone pixels in adjacent images and combining with the high CT value of bone pixels, 150 HU is selected as the screening threshold for bone pixels. The number of pixels with a CT value greater than the screening threshold in each L3 image with vertebral processes is counted, and the image with the largest number of pixels in the same group is determined as the best L3 image, and the image number corresponding to the best L3 image is output.

[0027] Preferably, a ResWAA_DNet deep network segmentation model is constructed as follows:

[0028] The annotated mask image in the segmentation training set Augseg_data1_train is mapped from a single-channel grayscale image to a four-channel semantic mask image containing the background class through One-Hot encoding. The semantic mask image and the training image are input into the ResWAA_DNet deep U-shaped network model built based on the CNN architecture and the attention mechanism. The ResWAA_DNet deep U-shaped network model is trained using the hard Dice coefficient loss function with edge enhancement, and the ResWAA_DNet deep network segmentation model is obtained. The best L3 image obtained by screening is input into the ResWAA_DNet deep network segmentation model to obtain its semantic mask image.

[0029] Preferably, the Laplacian of Gaussian operator is used to perform edge detection on the real muscle channel semantic mask image, and the Gaussian standard deviation σ of the LoG convolution kernel is set to 5;

[0030] The pixel positions with pixel intensity greater than 0.001 in the image obtained by edge detection are regarded as outer edges, and the pixel positions with pixel intensity less than -0.001 are regarded as inner edges;

[0031] The real and predicted semantic mask images are flattened into vectors, and then the pixels at the corresponding positions of the inner and outer edges in the two images are flattened into vectors again. The three vectors are concatenated in the order of mask image, inner edge, and outer edge, and the Dice loss Loss between the vectors obtained after concatenation is calculated as follows:

[0032] Loss = EE_DiceLoss(σ(Y), G)

[0033] where Y is the predicted mask image output by the model, G is the real semantic mask image, EE_DiceLoss is the cross-entropy loss based on edge enhancement, and σ(·) is the Sigmoid activation function.

[0034] Preferably, the ResWAA_DNet deep U-shaped network model is specifically as follows:

[0035] The input image resolution is 1×512×512. The ResWAA_DNet deep U-shaped network model includes an encoder, a decoder, and skip connections, and a residual structure is introduced in the encoding and decoding stages;

[0036] In the input stage of each skip connection, a wide-area attention mechanism WAA and a channel attention mechanism SE are introduced. The WAA module performs strip pooling on the input image in the height and width directions to obtain two sets of matrices, performs information fusion in the channel direction and then restores the original channel size, multiplies the two sets of matrices to obtain a weight matrix with the same size as the input image, and multiplies the weight matrix with the original image to obtain the final output result; The image processed by the WAA module is input into the SE module, and then corresponding weight extraction is performed on the channel direction through the squeeze-and-excitation mechanism and weighted to the original image. A DSC module is added between the architecture layers for cross-layer connection to make more full use of image features;

[0037] The encoding and decoding infrastructure of the ResWAA_DNet deep U-shaped network model is processed by two 3×3 convolutions combined with batch normalization and ReLU activation functions. In the encoding stage, the change of the image channels is 1-32-64-128-256-512, and in the decoding stage, the change of the image channels is 512-256-128-64-32-4. The final output result is used to calculate the loss with the four-channel semantic mask image.

[0038] Preferably, based on the best L3 image semantic mask image, the density and area of abdominal muscles, subcutaneous fat, and visceral fat are calculated by combining the density algorithm based on the binary mask and the area algorithm based on Monte Carlo integration. Specifically:

[0039] According to the predicted semantic mask image corresponding to the best L3 image, the mask images corresponding to muscles, subcutaneous fat, and visceral fat are extracted from the predicted segmentation grayscale image according to the grayscale values corresponding to different regions;

[0040] The mask images of muscles, subcutaneous fat, and visceral fat are respectively binarized to obtain binarized label images. The binarized label images are dot-multiplied with the corresponding matrices of the original CT data to obtain density masks. The pixels in the masks are summed and then averaged to obtain the density of the corresponding regions;

[0041] Adopting the Monte Carlo integration idea, the ratio of the number of pixel points contained in muscles, subcutaneous fat, and visceral fat to the total number of pixel points is regarded as the ratio of the areas of the three to the total area. Then, the actual pixel point spacing is extracted from the original CT data, the actual total area of the image is calculated, and the areas of the relevant regions are obtained according to the ratio.

[0042] In a second aspect, an embodiment of the present invention provides a system for calculating the density and area of abdominal muscles and fat, including:

[0043] A data module constructs a deep network classification model based on the DenseNet121 network architecture to classify lumbar vertebrae and determine the presence or absence of vertebral processes in abdominal region CT images, and filters the best L3 image through the optimal L3 positioning algorithm;

[0044] A network module constructs a ResWAA_DNet deep network segmentation model, uses the CNN convolutional layer as the encoder and decoder, adds a wide-area attention mechanism and cross-layer connection, and inputs the filtered best L3 image into the ResWAA_DNet deep network segmentation model to obtain the corresponding semantic mask image;

[0045] An output module calculates the density and area of abdominal muscles, subcutaneous fat, and visceral fat based on the semantic mask image of the best L3 image, in combination with the density algorithm based on a binary mask and the area algorithm based on Monte Carlo integration.

[0046] In a third aspect, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method for calculating the density and area of abdominal muscles and fat are implemented.

[0047] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium including a computer program. When the computer program is executed by a processor, the steps of the above method for calculating the density and area of abdominal muscles and fat are implemented.

[0048] In a fifth aspect, a chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method for calculating the density and area of abdominal muscles and fat are implemented.

[0049] In a sixth aspect, an embodiment of the present invention provides an electronic device including a computer program. When the computer program is executed by the electronic device, the steps of the above method for calculating the density and area of abdominal muscles and fat are implemented.

[0050] Compared with the prior art, the present invention has at least the following beneficial effects:

[0051] A method for calculating the density and area of abdominal muscles and fat, which uses a deep network classification model based on the DenseNet121 network architecture and a deep network segmentation model ResWAA_DNet. First, a deep network classification model based on the DenseNet121 network architecture is constructed to classify the lumbar vertebra to which the abdominal region CT image belongs and judge whether there is a vertebral process, and the best L3 image is selected through the optimal L3 positioning algorithm. Secondly, an abdominal muscle and fat segmentation model ResWAA_DNet is constructed. Using the CNN convolutional layer as the encoder, considering the relatively complex shape characteristics of the segmentation region, a wide-area attention mechanism and cross-layer connection are added, which can better extract the target features of the image and enrich the information exchange between features, further improving the segmentation effect of the target region and reducing the relative error between the predicted density and area and the true value later. Based on the edge detection operator, the boundary of the target region is emphasized, which can further optimize the boundary segmentation effect of the model on the target region and is conducive to the density and area prediction later being closer to the true value.

[0052] Furthermore, the contrast between the background and foreground of abdominal CT images is usually low, resulting in the target boundary and the image background being difficult to distinguish. Considering that the original images are all grayscale images, windowing and histogram equalization operations are performed on the pixel values of the images to enhance the contrast between the target region and the background of the image, making the target region more prominent and easy to distinguish.

[0053] Furthermore, the deep network model DenseNet121 is used to classify the lumbar vertebra region and judge whether there is a vertebral process. Combining the idea of transfer learning, the encoder weight parameters are initialized with the pre-trained Densenet weight file, which alleviates the "data starvation" problem. The dense connection technology is applied, that is, each layer will pass the obtained feature map to all subsequent layers through channel splicing, not just to the next layer. Through this feature transfer method, the subsequent layers can better reuse the features and achieve better classification results with fewer parameters. On the other hand, due to the application of dense connections, each layer of the model can directly obtain gradient information during backpropagation, greatly alleviating the problem of gradient disappearance, so that deeper network layers can be applied for learning, thereby improving the accuracy of lumbar vertebra region classification and the judgment of whether there is a vertebral process.

[0054] Furthermore, through the optimal L3 positioning algorithm, the model can automatically and intelligently find the number of the best third lumbar vertebra L3 image, saving the time for professional doctors to manually find the best L3 image and improving the work efficiency of doctors.

[0055] Furthermore, a CNN convolutional layer is used as the encoder and decoder of the deep network, and the Wide Area Attention mechanism (WAA) is adopted to promote the model to better capture long-range dependencies. The Squeeze-and-Excitation (SE) mechanism is used to accurately obtain the weight coefficients of different channels for the target, so as to screen out valuable information from a large amount of data information, thereby improving the segmentation performance of the model. In addition, the DSC module is used for cross-layer connection to achieve multi-feature fusion of the model and make better use of the feature information extracted by the model.

[0056] Furthermore, for the semantic mask map predicted by the ResWAA_DNet deep network, the Dice coefficient that emphasizes the mask edge is used to measure the loss, guiding the model to better retain the boundary characteristics of the target area. Based on the edge detection operator, a larger weight is assigned to the target boundary in the loss calculation, enabling the network to focus on the segmentation effect of the target boundary.

[0057] Furthermore, Adam with weight decay is used as the optimizer, which alleviates the model's convergence to local minima to a certain extent and improves the convergence speed. To address model overfitting, weight decay is also introduced into the optimizer to prevent the network weight parameters from being too large and reduce the complexity of the model.

[0058] It can be understood that the beneficial effects of the second to sixth aspects above can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.

[0059] In summary, the present invention uses DenseNet121 as the lumbar classification and vertebral protrusion judgment model, and utilizes its advantage of feature reuse to obtain better classification results under the premise of less data volume. The optimal L3 positioning algorithm further improves the prediction accuracy of the target image; ResWAA_DNet is used as the segmentation model, and by combining the Wide Area Attention mechanism (WAA) and the Squeeze-and-Excitation (SE) mechanism with the CNN-based encoder and decoder, the prediction accuracy of the semantic mask of the target area is improved; the loss function with edge enhancement is used to guide model training, making the boundary segmentation result of the model for the target area more accurate; the density and area intelligent calculation algorithm can accurately calculate the medical indicators of the abdominal muscle and fat areas.

[0060] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings

[0061] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced below. Obviously, the following described drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0062] Figure 1 This is the overall flowchart of the present invention;

[0063] Figure 2 This is the original CT image and its labeled mask image;

[0064] Figure 3 This is the original CT image and its contrast-enhanced image;

[0065] Figure 4 This is the network structure diagram of DenseNet121;

[0066] Figure 5 This is the confusion matrix and ROC curve of the lumbar spine classification result;

[0067] Figure 6 This is the original CT image and its preprocessed vertebral process judgment image;

[0068] Figure 7 This is the confusion matrix and ROC curve of the vertebral process judgment result;

[0069] Figure 8 This is an overview of the WAA module;

[0070] Figure 9 This is the network structure diagram of the ResWAA_DNet model;

[0071] Figure 10 These are the semantic mask maps predicted by five segmentation models;

[0072] Figure 11 This is a comparison chart of the prediction results of five segmentation models;

[0073] Figure 12 This is the Bland-Altman analysis chart of the density and area calculation results of muscle, subcutaneous fat, and visceral fat;

[0074] Figure 13 This is the linear regression analysis chart of the density and area calculation results of muscle, subcutaneous fat, and visceral fat;

[0075] Figure 14 This is a schematic diagram of the computer device provided by an embodiment of the present invention;

[0076] Figure 15 This is a block diagram of a chip provided by an embodiment of the present invention.

[0077] Among them, 60. Computer device; 61. Processor; 62. Memory; 63. Computer program; 600. Electronic device; 610. Processing unit; 620. Storage unit; 6201. Random access storage unit; 6202. Cache storage unit; 6203. Read-only storage unit; 6204. Program / utilities; 6205. Program module; 630. Bus; 640. Display unit; 650. Input / output interface; 660. Network adapter; 700. External device. Detailed implementation manner

[0078] The present invention provides a method for calculating the density and area of abdominal muscles and fat. A dataset is constructed based on the abdominal CT images of liver cirrhosis patients in the First Affiliated Hospital of Xi'an Jiaotong University and image augmentation is performed; based on the DenseNet121 model, a lumbar classification model DenseNet_5 and a vertebral process judgment model DenseNet_3 are trained. The DenseNet_5 is used to classify the lumbar vertebra to which the abdominal region CT image belongs, and the image with the predicted class of L3 is input into the DenseNet_3 model to find the L3 image with vertebral processes. The best L3 image is screened through the optimal L3 positioning algorithm; a ResWAA_DNet model is constructed to learn the mapping from the abdominal CT image to its semantic segmentation mask image. The CNN convolutional layer is used as the encoder and decoder, and the edge detection operator is used to emphasize the boundary of the target region to optimize the boundary segmentation effect of the model on the target region. In view of the complex shape characteristics of the segmentation region, a wide-area attention mechanism and cross-layer connection are added to further improve the segmentation effect of the target region. Finally, the semantic segmentation results of abdominal muscles, subcutaneous fat, and visceral fat of the best L3 image are obtained; according to the segmentation results, the density and area of the above regions are obtained by using the density algorithm based on the binary mask and the area algorithm based on Monte Carlo integration. The present invention reaches an advanced level in the automatic classification of lumbar vertebrae in abdominal CT images, the automatic positioning of the best L3 image, and the automatic segmentation of abdominal muscles and fat, and improves the accuracy of the calculation results of the density and area of related regions.

[0079] Embodiment 1

[0080] Please refer to Figure 1 , a method for calculating the density and area of abdominal muscles and fat according to the present invention, includes the following steps:

[0081] S1. Construct an abdominal CT image dataset based on the CT images of 97 liver cirrhosis patients at different times from the First Affiliated Hospital of Xi'an Jiaotong University. According to the tasks, divide the dataset into an abdominal CT image classification dataset class_data1 and an abdominal CT image segmentation dataset seg_data1. Perform preprocessing and image augmentation operations on the two types of datasets to obtain an augmented classification dataset Augclass_data1 and an augmented segmentation dataset Augseg_data1. Divide the two augmented datasets into a training set (_train), a validation set (_val), and a test set (_test);

[0082] S101. Dataset acquisition

[0083] The dataset of the present invention is derived from the abdominal CT images collected from 97 liver cirrhosis patients at different times from the First Affiliated Hospital of Xi'an Jiaotong University. This dataset is divided into an abdominal CT image classification dataset class_data1 and an abdominal CT image segmentation dataset seg_data1 according to two different tasks of automatic classification and automatic segmentation. The specific data composition is shown in Table 1:

[0084] Table 1 Composition of the original abdominal CT image dataset

[0085]

[0086] S102. Standard preprocessing

[0087] For the abdominal CT image classification dataset, open the dcm files of each patient at different times through the 3D_slicer medical software. According to the shape characteristics of the lumbar vertebrae, divide the CT images into images of the first lumbar vertebra to the fifth lumbar vertebra L1-L5, as shown in Table 2. Divide the dcm files into 5 categories from L1 to L5 and save them in the corresponding folders. Generate a CSV file containing 5-category annotation information through programming software, and then further convert the dcm files into corresponding png format images and save them;

[0088] Table 2 Shape characteristics of the first lumbar vertebra to the fifth lumbar vertebra

[0089]

[0090]

[0091] For the abdominal CT image segmentation dataset, two professional doctors from the First Affiliated Hospital of Xi'an Jiaotong University used the 3D_slicer medical software to label the CT images of the L3 and L4 regions with corresponding medical analysis value for the abdominal muscles, subcutaneous fat, and visceral fat regions. Three types of regions, namely muscles (CT value range: -29 HU - 150 HU), subcutaneous fat (CT value range: -190 HU - -30 HU), and visceral fat (CT value range: -150 HU - -50 HU), were segmented from the CT images, and the segmented results were saved as medical files in nii format for storage. Subsequently, the nii format medical files were further converted into corresponding png format images and saved; Figure 2 The original CT image and its labeled mask image are shown.

[0092] S103. Medical CT Image Preprocessing

[0093] In the abdominal CT image dataset, the contrast between the image background and the foreground is low, making it difficult to distinguish the target boundary from the image background. Considering that the original images are all grayscale images, windowing operations are performed on the pixel values of the images to enhance the contrast between the target region and the background of the images. Windowing technology is a class of methods for adjusting the gray-level components of CT images, which includes two important parameters: window width and window center. The window width refers to the CT value range of the target image, which affects the contrast of the image. A larger window width will make the transition between the bright and dark organ structures in the image less obvious, making the different tissues blurred. The window center represents the midpoint of the CT value of the image, which affects the brightness of the image. A smaller window center will result in a brighter image. The calculation process is as follows:

[0094]

[0095] Among them, WW is the window width and WC is the window center; for classification task images, (window width, window center) is set to (250, 45), and for segmentation task images, (window width, window center) is set to (400, 0).

[0096] After the windowing operation, histogram equalization is used to process the images. Histogram equalization can transform the histogram of the image into an approximate distribution through the cumulative distribution function to enhance the contrast of the image and improve the convergence speed of network training. The cumulative distribution function is as follows:

[0097]

[0098] Among them, k is a certain gray level of the image, n is the total number of pixels in the image, and n j is the number of pixels with gray level j.

[0099] Figure 3Shows the original image (left column) and its contrast-enhanced image (right column). It can be seen that the target area of the enhanced image is more prominent and easier to distinguish.

[0100] S104. Image augmentation and dataset division

[0101] To increase the amount of training data, alleviate the "data hunger" phenomenon, and enhance the generalization performance of the model, image augmentation operations are simultaneously performed on the classification and segmentation training images and the corresponding labeled mask images. Then, the augmented classification images are scaled to images of size 224×224, and the segmentation images and labeled mask images are kept at size 512×512 to adapt to the model input. The image augmentation operations include random rotation, adding noise, and horizontal flipping. These operations are simultaneously performed on the training images and their corresponding labeled mask images to obtain the augmented classification dataset Augclass_data1 and the augmented segmentation dataset Augseg_data1. The two augmented datasets are divided into a training set (_train), a validation set (_val), and a test set (_test); the specific data composition is shown in Table 3. The augmented training set becomes three times the size of the original training set data volume, and the number of images from L1 to L5 also becomes three times the original accordingly.

[0102] Table 3 Composition of the augmented abdominal CT image dataset

[0103]

[0104] S2. Input the CSV file with the annotation information of 5 lumbar spine types and the classification training set Augclass_data1_train into the DenseNet_5 deep network classification model built based on the DenseNet121 deep network architecture. Use the cross-entropy loss function containing measures to alleviate class imbalance to guide the model training, and perform validation on the classification validation set Augclass_data1_val to obtain the lumbar spine classification model DenseNet_5; input the images in the classification test set Augclass_data1_test into the DenseNet_5 model to obtain the lumbar spine classification results. Use evaluation metrics such as Accuracy, Precision, Recall, and F1 value to evaluate the lumbar spine classification results, and draw the confusion matrix and ROC curve of the classification results;

[0105] The specific steps are as follows:

[0106] S201. Traverse the folders where different classes of images are located to generate the corresponding class One-Hot encodings. Specifically, for images of the same class, assign 1 to the corresponding class channel and 0 to the remaining channels; save the image names and their corresponding class labels in a CSV file;

[0107] S202. Select the median of the number of images in 5 categories as the category weight benchmark \(i\), divide the number of images in the 5 categories by \(i\) respectively to obtain the category weight coefficients, so that the images with a larger number correspond to smaller coefficients, and the images with a smaller number have higher weight coefficients to alleviate the category imbalance phenomenon. Add the obtained weight coefficients to the cross-entropy loss function to guide the model training;

[0108] S203. Establish a lumbar spine classification model DenseNet_5 based on the DenseNet121 deep network framework, and obtain the predicted category label of the input image through the input layer, dense layer, transition layer and fully connected layer;

[0109] Please refer to Figure 4 , and the specific process of constructing the DenseNet121 deep network is as follows:

[0110] S2031. The resolution of the input image is 224×224×3. The initial input layer adjusts the image to a size of 56×56×64 through 64 7×7 convolutional kernels and 3×3 max-pooling kernels;

[0111] S2032. The DenseNet121 core architecture composed of four dense layers and three transition layers is responsible for processing the results output by the initial input layer. Among them, the dense layer is composed of multiple groups of 1×1 and 3×3 convolutional kernels connected in series, and the number of groups is [6, 12, 24, 16] respectively. The 1×1 convolutional kernel is responsible for fusing channel information, and the 3×3 convolutional kernel is used for feature extraction; the transition layer is composed of multiple groups of 1×1 convolutional kernels and 2×2 average pooling kernels, and the number of groups is [128, 256, 512] respectively, which is mainly used to control the output channel number and size of the feature map;

[0112] S2033. After the image is processed by the dense layer and the transition layer, its size is adjusted to 7×7×1024. The image is sent into the classification layer composed of 1024 7×7 average pooling kernels and a fully connected layer, and the final prediction probability of each category is obtained in combination with the Softmax function;

[0113] S204. Use the Adam optimizer with weight decay to minimize the loss function in step S202 to train the DenseNet_5 model. Set the betas of the Adam function in the torch toolkit to (0.9, 0.999), the weight decay to 0.0005, the number of training rounds to 200, introduce an early stopping mechanism, set the acceptable threshold to 0.001, the acceptable maximum number of rounds to 25, the BatchSize to 32, and the initial value of the learning rate to 0.0001. Train to obtain the lumbar spine classification model DenseNet_5;

[0114] S205. Input the images in the classification test set Augclass_data1_test into the DenseNet_5 model to obtain the lumbar classification results of the test set. Use evaluation metrics such as Accuracy, Precision, Recall, and F1 value to evaluate the lumbar classification results, and draw the confusion matrix and ROC curve of the classification results;

[0115] The formula for calculating the accuracy Accuracy is as follows:

[0116]

[0117] Among them, TP represents the number of samples of true positives, FN represents the number of samples of false negatives, FP represents the number of samples of false positives, and TN represents the number of samples of true negatives;

[0118] The formula for calculating the precision Precision is as follows:

[0119]

[0120] The formula for calculating the recall Recall is as follows:

[0121]

[0122] The formula for calculating the F1 value is as follows:

[0123]

[0124] The results of the evaluation metrics of the lumbar classification model on the test set are shown in Table 4:

[0125] Table 4 Results of the evaluation metrics of the lumbar classification model on the test set

[0126]

[0127] The confusion matrix and ROC curve of the lumbar classification results are shown in Figure 5 .

[0128] S3. Window the images of the third lumbar vertebra L3 class in the classification dataset Augclass_data1 and perform central cropping. Divide the processed images into three categories: with processus vertebral, indistinct processus vertebral, and without processus vertebral according to the presence or absence of processus vertebral. Input the CSV file with the annotation information of the three categories of the presence or absence of processus vertebral and the processed training set Augclass_data1_train into the DenseNet_3 deep network classification model built based on the DenseNet121 deep network architecture. Use the cross-entropy loss function containing mitigation of class imbalance to guide the model training, and perform validation on the processed validation set Augclass_data1_val to obtain the processus vertebral judgment model DenseNet_3. Input the images predicted as L3 class in the test set of step S2 into the DenseNet_3 model to obtain the processus vertebral judgment results of the L3 class images. Use evaluation indicators such as Accuracy, Precision, Recall, and F1 value to evaluate the processus vertebral judgment results, and draw the confusion matrix and ROC curve of the judgment results. Screen the best L3 images from the L3 images predicted as having processus vertebral using the best L3 localization algorithm;

[0129] The specific steps are as follows:

[0130] S301. Window the images of the third lumbar vertebra L3 class in the classification dataset Augclass_data1 and perform central cropping to make the shape of the main lumbar region more prominent, obtaining images of size 224×224. Divide the processed images into three categories: with processus vertebral, indistinct processus vertebral, and without processus vertebral according to the presence or absence of processus vertebral. Traverse the folders where different categories of images are located to generate the One-Hot encoding of the corresponding categories. Specifically, assign 1 to the corresponding category channel for images of the same category and 0 to the remaining channels. Save the image names and their corresponding category labels in a CSV file;

[0131] S302. Select the median of the numbers of the three categories of images as the category weight benchmark i. Divide the numbers of the three categories of images by i respectively to obtain the category weight coefficients, so that images with a larger number correspond to smaller coefficients and images with a smaller number have higher weight coefficients to mitigate the class imbalance phenomenon. Add the obtained weight coefficients to the cross-entropy loss function to guide the model training;

[0132] S303. Establish the processus vertebral judgment model DenseNet_3 based on the DenseNet121 deep network framework, and obtain the predicted category labels of the input images through the input layer, dense layer, transition layer, and fully connected layer;

[0133] S304. Use the Adam optimizer with weight decay to minimize the loss function in step S302 to train the DenseNet_3 model. Set the betas of the Adam function in the torch toolkit to (0.9, 0.999), the weight decay to 0.0005, the number of training epochs to 200, introduce an early stopping mechanism, set the acceptable threshold to 0.001, the maximum acceptable number of epochs to 25, the BatchSize to 32, and the initial learning rate to 0.0001. Train to obtain the vertebral protrusion judgment model DenseNet_3;

[0134] S305. Input the images predicted as the L3 class in the test set of step S2 into the DenseNet_3 model to obtain the vertebral protrusion judgment results of the L3 class images. Use evaluation metrics such as Accuracy, Precision, Recall, and F1 value to evaluate the vertebral protrusion judgment results, and draw the confusion matrix and ROC curve of the classification results;

[0135] The evaluation metric results of the vertebral protrusion judgment model DenseNet_3 on the test set are shown in Table 5:

[0136] Table 5 Evaluation Metric Results of the Vertebral Protrusion Classification Model on the Test Set

[0137]

[0138] The confusion matrix and ROC curve of the vertebral protrusion judgment results are shown in Figure 7 ;

[0139] S306. Use the optimal L3 positioning algorithm to screen the optimal L3 images for the L3 images predicted to have vertebral protrusions; the specific operations are as follows. The third lumbar vertebra and the vertebral protrusion being obvious are the gold standards for determining the optimal L3 image in medicine. Based on this standard, it can be considered that the optimal L3 image has the characteristic of occupying more bone pixels in adjacent images. Using this characteristic and the feature that bone pixels have a relatively high CT value (unit: Hounsfield Unit, abbreviated as HU), select 150HU as the screening threshold for bone pixels (the CT values of the remaining tissue components in the abdomen are usually lower than 150HU). Count the number of pixels with a CT value greater than the screening threshold in each L3 class image with vertebral protrusions, and determine the image with the largest number of pixels in the same group as the optimal L3 image, and output the image number corresponding to the optimal L3 image; use this algorithm to predict the optimal L3 images of the patient sequences in the test set Augclass_data1_test.

[0140] Please refer to Table 6 (the numbers without units in the table represent the numbers of the images in the CT sequence), and the specific positioning process and results are as follows:

[0141] S3061. Since the image spacing in the CT sequence is in millimeters and adjacent images have very similar medical features, there may be multiple L3 images with medical analysis value in the CT sequence collected from a patient at one time, because the error between the medical indicators calculated from these images can be ignored. Usually, a 5-mm distance above and below the best L3 image selected by the doctor is used as the acceptable range, and any image within the acceptable range can be regarded as the best L3 image for medical analysis. The specific definition of the acceptable range is as follows: in the CT sequence with a 10-mm spacing, the acceptable range only contains the best L3 image marked by the doctor; in the CT sequence with a 5-mm spacing, the acceptable range is the best L3 image marked by the doctor and the two adjacent images above and below it; in the CT sequence with a 1-mm spacing, the acceptable range is the best L3 image marked by the doctor and the five adjacent images above and below it.

[0142] S3062. Obtain the spacing of the patient's CT sequence and count the number of pixels with a CT value greater than 150 HU in the L3 images predicted to have vertebral protrusions in this CT sequence. Obtain the image number with the largest number of pixels, and then determine whether this number falls within the acceptable range defined in S3061. If it falls within the range, it means the prediction is correct; otherwise, the prediction is incorrect. Count the prediction results of the best L3 images in the CT sequences of each patient in the test set Augclass_data1_test. There are a total of 21 patients' CT sequence data in the test set, and the final result is that the predicted best L3 image numbers of all 21 patients fall within the acceptable range.

[0143] Table 6 Prediction Results of the Best L3 Image Numbers (First 10 Test Patients)

[0144]

[0145] S4. Map the annotated mask image in the segmented training set Augseg_data1_train from a single-channel grayscale image to a four-channel semantic mask image including the background class through One-Hot encoding. Input the semantic mask image and the training image into the ResWAA_DNet deep U-shaped network model built based on the CNN architecture and the attention mechanism. Use the hard Dice coefficient loss function with edge enhancement to guide the training of the ResWAA_DNet model. Validate it on the segmentation validation set Augseg_data1_val to obtain the abdominal muscle and fat segmentation model ResWAA_DNet. Input the segmented test set Augseg_data1_test into the ResWAA_DNet model to obtain the predicted semantic mask image, and obtain the grayscale images of abdominal muscles, subcutaneous fat, and visceral fat through inverse One-Hot encoding. Use evaluation metrics such as Dice coefficient, Jaccard coefficient, Precision, and Recall to measure the model segmentation effect. Input the best L3 image of each patient selected in step S3 into the ResWAA_DNet model to obtain its semantic mask image and the grayscale images of abdominal muscles, subcutaneous fat, and visceral fat.

[0146] Please refer to Figures 8 to 9 , the specific steps are as follows:

[0147] S401. Map the grayscale label from [H, W] to [H, W, K]. Each category corresponds to a specific channel. There are four categories in multi-object segmentation including the background (the grayscale value corresponding to the background is 0, muscle is 85, subcutaneous fat is 170, and visceral fat is 255). For each category, use the One-Hot encoding method. Each category is mapped to 1 on the corresponding channel, and the rest of the channels are mapped to 0 to obtain a four-channel semantic mask image;

[0148] S402. Use edge detection to obtain the inner and outer edges of the region boundary, and combine the inner and outer edge pixels with the semantic mask image to construct a hard Dice loss function with edge enhancement;

[0149] The specific process is as follows:

[0150] S4021. Use the Laplacian of Gaussian (LoG) to perform edge detection on the real muscle channel semantic mask image, which is implemented using the gaussian_laplace function in the Scipy toolkit. Set the Gaussian standard deviation σ of the LoG convolution kernel to 5;

[0151] S4022. Consider the pixel positions with pixel intensity greater than 0.001 in the image obtained by edge detection in step S4021 as the outer edge, and the pixel positions with pixel intensity less than -0.001 as the inner edge;

[0152] S4023. Flatten the real and predicted semantic mask images into vectors, then flatten the pixels at the corresponding positions of the inner and outer edges in the two images into vectors again. Concatenate the three vectors in the order of the mask image, inner edge, and outer edge, and calculate the Dice loss between the vectors obtained after concatenation.

[0153] S4024. The final loss function is as follows:

[0154] Loss = EE_DiceLoss(σ(Y), G)

[0155] where Y is the predicted mask image output by the model, G is the real semantic mask image, EE_DiceLoss is the cross-entropy loss based on edge enhancement, and σ(·) is the Sigmoid activation function.

[0156] S403. Introduce the residual structure into the encoding and decoding modules of U-Net, construct the wide-area attention WAA module, introduce the WAA module and the channel attention module SE into the skip connection mechanism of U-Net, and introduce the depthwise separable convolution (DSC) to establish cross-layer connections, thereby constructing the ResWAA_DNet deep network segmentation model, and input the CT image to predict the corresponding semantic mask image.

[0157] Please refer to Figures 8 to 9 , the construction process of the wide-area attention mechanism WAA and the ResWAA_DNet model is as follows:

[0158] S4031. The input image resolution is 1×512×512. The overall model architecture is improved based on U-Net, mainly including an encoder, a decoder, and skip connections, and the residual structure is introduced in the encoding and decoding stages.

[0159] S4032. The overall model includes four skip connections. The wide-area attention mechanism WAA and the channel attention mechanism SE are introduced in the input stage of each skip connection. Among them, WAA performs strip pooling on the input image in the height and width directions to obtain two groups of matrices, performs information fusion in the channel direction and then restores the original channel size, multiplies the two groups of matrices to obtain a weight matrix with the same size as the input image, and multiplies the weight matrix with the original image to obtain the final output result; the image processed by the WAA module is input into the SE module, and then the squeeze-and-excitation mechanism is used to extract the corresponding weights in the channel direction and weight them to the original image, and the DSC module is added between the architecture layers for cross-layer connection to make more full use of the image features.

[0160] S4033. The encoding and decoding infrastructure of the model is processed by two 3×3 convolutions combined with batch normalization and ReLU activation functions. During the encoding stage, the change in the number of image channels is 1-32-64-128-256-512, and during the decoding stage, the change in the number of image channels is 512-256-128-64-32-4. The final output result is used to calculate the loss with the four-channel semantic mask image of S401;

[0161] S404. The Adam optimizer with weight decay is used to minimize the loss function in step S402 to train the ResWAA_DNet model. The weight decay of the Adam function in the torch toolkit is set to 0.00001, the number of training epochs is set to 200, an early stopping mechanism is introduced, the acceptable threshold is set to 0.001, the acceptable maximum number of epochs is set to 25, the BatchSize is set to 4, and the initial learning rate is set to 0.0001; The abdominal muscle and fat segmentation model ResWAA_DNet is trained;

[0162] S405. The test set Augseg_data1_test is input into the ResWAA_DNet segmentation model to obtain the predicted semantic mask image, and the predicted segmentation grayscale image is obtained through inverse One-Hot encoding; Evaluation metrics such as Dice coefficient, Jaccard coefficient, Precision, and Recall are used to measure the model segmentation effect;

[0163] The calculation formula of the Dice coefficient is as follows:

[0164]

[0165] The calculation formula of the Jaccard coefficient is as follows:

[0166]

[0167] The evaluation metric results of the abdominal muscle and fat segmentation model on the segmentation test set are shown in Table 7:

[0168] Table 7 Performance comparison of five models on the test set

[0169]

[0170] The ResWAA-DNet segmentation model provided by the present invention is superior to the other four models in three types of regions and four indicators, except for the Recall indicator on subcutaneous fat. Compared with the advanced model ResU-Net, the ResWAA-DNet model has increased by 0.5% in the average Dice coefficient, 0.92% in the average Jaccard coefficient, 0.74% in the average Precision, and 0.23% in the average Recall. The present invention has reached the advanced level of automatic segmentation of muscle and fat in abdominal CT images.

[0171] To visually compare the segmentation effects of the five models on abdominal CT images, four CT images were selected in the test set, and the corresponding annotation masks and the prediction masks of the five models were visualized, as well as the regions where the prediction masks are different from the annotation masks, specifically as Figure 10 、 Figure 11 shown. It can be seen that the segmentation fineness of the automatic segmentation method for abdominal CT images provided by the present invention is better than that of other models, further verifying the quantitative results shown in Table 7;

[0172] S406. Input the best L3 image of each patient screened in step S3 into the ResWAA_DNet model to obtain its predicted semantic mask image and the gray-scale images of abdominal muscle, subcutaneous fat, and visceral fat.

[0173] S5. Design a density algorithm based on a binary mask and an area algorithm based on Monte Carlo integration, and calculate the density and area of abdominal muscle, subcutaneous fat, and visceral fat by combining with the predicted semantic mask image corresponding to the best L3 image obtained in step S4; use statistical methods such as relative error, Bland-Altman evaluation method, and linear regression analysis to analyze the prediction effects of the density and area of the abdominal muscle and fat regions.

[0174] S501. According to the predicted semantic mask image corresponding to the best L3 image obtained in step S4, extract the mask images corresponding to muscle, subcutaneous fat, and visceral fat from the predicted segmentation gray-scale image according to the gray-scale values corresponding to different regions;

[0175] S502. Design a density algorithm based on a binary mask, specifically as follows: perform binary processing on the mask images of muscle, subcutaneous fat, and visceral fat (the target region is mapped to 1, and the background region is mapped to 0) to obtain binary label images, perform a dot product operation on the binary label images and the corresponding matrix of the original CT data to obtain a density mask, sum the pixels in the mask, and then take the average value to obtain the density of the corresponding region;

[0176] S503. Design an area algorithm based on Monte Carlo integration. Since muscle, subcutaneous fat, and visceral fat are all irregular regions, and these regions are all distributed in an image with a regular resolution of 512×512, using the Monte Carlo integration idea, the ratio of the number of pixel points contained in muscle, subcutaneous fat, and visceral fat (the number of 1s in the binary label in S502) to the total number of pixel points (512×512) is regarded as the ratio of the areas of the three to the total area. Then, extract the actual pixel point spacing from the original CT image, thereby calculating the actual total area of the image, and obtaining the areas of the three relevant regions according to the ratio;

[0177] S504. Use statistical methods such as relative error, Bland - Altman evaluation method, and linear regression analysis to analyze the density and area prediction effects of abdominal muscle and fat regions;

[0178] Among them, the relative error calculation formula is as follows:

[0179]

[0180] The true, predicted results, and relative errors of the density and area of the three regions on the best L3 image of the test set patients are shown in Tables 8 and 9:

[0181] Table 8 Calculation results of the density of muscle, subcutaneous fat, and visceral fat on the best L3 image (the first 10 test patients)

[0182]

[0183] Table 9 Calculation results of the area of muscle, subcutaneous fat, and visceral fat on the best L3 image (the first 10 test patients)

[0184]

[0185] It can be seen from the two result tables that the density algorithm based on binary mask and the area algorithm based on Monte Carlo integration are relatively accurate in calculating the density and area of the best L3 image. Except that the relative error of the visceral fat area in sequence 1 exceeds 5%, the relative errors of all other results are within 5%, and most of the relative errors do not exceed 1%, reflecting good calculation effects;

[0186] The Bland - Altman evaluation method was proposed by Bland and Altman in 1986 and is used to evaluate and analyze two different measurement methods to calculate their consistency or difference when measuring the same thing. The calculation results will be shown through a Bland - Altman plot. The horizontal axis is the average value calculated by the two measurement methods for the sample, and the vertical axis is the difference between the two measurement methods. The limits of agreement are drawn based on 95% of the data distribution, representing the acceptable range of measurement differences. The relevant calculation formula is as follows:

[0187]

[0188] Lower Limit of Agreement = Mean Difference - 1.96 × SD of Differences

[0189] Upper Limit of Agreement = Mean Difference + 1.96 × SD of Differences

[0190] where D i is the difference between the two measurement methods on the i-th sample. If the average value of the differences between the two measurement methods is near 0 and most of the sample points in the figure fall within the limits of agreement, then it can be considered that the two measurement methods have good consistency.

[0191] Linear regression analysis is a statistical method used to study the relationship between two or more variables. It is usually used to evaluate the linear relationship between two sets of data and is often applied to the verification of measurement methods in scientific research. The fitting validity and statistical significance are evaluated through R 2 and P values, where the calculation formula of R 2 is as follows:

[0192]

[0193] If R 2 is closer to 1, it indicates that the fitting of the two methods is better. When the P value is less than 0.05, it shows that there is a significant relationship between the two methods.

[0194] To better demonstrate the good generalization performance of the model in calculating density and area, not only the density and area of the best L3 image in the test set are calculated, but also the relevant extractions of density and area for the remaining images in the test set are performed. Then, relevant statistical analyses are carried out on the calculation results. The Bland - Altman analysis and linear regression analysis diagrams of the density and area of abdominal muscles, subcutaneous fat, and visceral fat in the test set are shown in Figure 12 , Figure 13 , from which it can be seen that the predicted results of the density and area of this model are very close to the actual results, having high practical value.

[0195] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "platform" here.

[0196] Embodiment 2

[0197] The present invention provides a system for calculating the density and area of abdominal muscles and fat, which can be used to implement the above method for calculating the density and area of abdominal muscles and fat. Specifically, the system for calculating the density and area of abdominal muscles and fat includes a data module, a network module, and an output module.

[0198] Among them, the data module constructs a deep network classification model based on the DenseNet121 network architecture to classify the lumbar vertebrae and judge the presence or absence of vertebral processes in the abdominal region CT image, and screens the best L3 image through the optimal L3 positioning algorithm;

[0199] The network module constructs a ResWAA_DNet deep network segmentation model, uses the CNN convolutional layer as the encoder and decoder, adds a wide-area attention mechanism and cross-layer connection, and inputs the screened best L3 image into the ResWAA_DNet deep network segmentation model to obtain the corresponding semantic mask image;

[0200] The output module calculates the density and area of abdominal muscles, subcutaneous fat, and visceral fat based on the semantic mask image of the best L3 image, in combination with the density algorithm based on the binary mask and the area algorithm based on the Monte Carlo integral.

[0201] Embodiment 3

[0202] The present invention provides a terminal device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Graphics Processing Units (GPU), Tensor Processing Units (TPU), Digital Signal Processors (DSP), Application Specific Integrated Circuits (ASIC), Field-Programmable Gate Arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiments of the present invention can be used for the operation of the method for calculating the density and area of abdominal muscles and fat, including:

[0203] Construct a deep network classification model based on the DenseNet121 network architecture to classify the lumbar vertebrae and judge the presence or absence of vertebral processes in the abdominal region CT image, and screen the best L3 image through the optimal L3 positioning algorithm; construct a ResWAA_DNet deep network segmentation model, use the CNN convolutional layer as the encoder and decoder, add a wide-area attention mechanism and cross-layer connection, and input the screened best L3 image into the ResWAA_DNet deep network segmentation model to obtain the corresponding semantic mask image; based on the semantic mask image of the best L3 image, combine the density algorithm based on the binary mask and the area algorithm based on Monte Carlo integration to calculate the density and area of abdominal muscles, subcutaneous fat, and visceral fat.

[0204] Please refer to Figure 14 , the terminal device is a computer device. The computer device 60 in this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the computer program 63 is executed by the processor 61, it implements the method for calculating the density and area of abdominal muscles and fat in the embodiment. To avoid repetition, it will not be elaborated here one by one. Alternatively, when the computer program 63 is executed by the processor 61, it implements the functions of each model / unit in the system for calculating the density and area of abdominal muscles and fat in the embodiment. To avoid repetition, it will not be elaborated here one by one.

[0205] The computer device 60 can be a computing device such as a desktop computer, a notebook, a handheld computer, and a cloud server. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art can understand that Figure 14 merely examples of the computer device 60, which do not constitute a limitation on the computer device 60, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0206] The so-called processor 61 may be a central processing unit (CPU), or may also be other general-purpose processors, a graphics processing unit (GPU), a tensor processing unit (TPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0207] The memory 62 may be an internal storage unit of the computer device 60, such as the hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk equipped on the computer device 60, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0208] Furthermore, the memory 62 may also include both an internal storage unit of the computer device 60 and an external storage device. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 may also be used to temporarily store data that has been output or will be output.

[0209] Please refer to Figure 15 , the terminal device is an electronic device 600, and the electronic device 600 is presented in the form of a general computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.

[0210] Among them, the storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the method section of this specification. For example, the processing unit 610 can execute the steps as shown in Figure 1 .

[0211] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.

[0212] The storage unit 620 may further include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0213] The bus 630 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.

[0214] The electronic device 600 can also communicate with one or more external devices 700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device (such as a router, a modem) that enables the electronic device 600 to communicate with one or more other computing devices. Such communication can be carried out through the input / output interface 650. Moreover, the electronic device 600 can also communicate with one or more networks (such as a local area network, a wide area network, and / or a public network, such as the Internet) through the network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.

[0215] Example 4

[0216] The present invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. It can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The computer-readable storage medium provides a storage space that stores the operating system of the terminal. And, in this storage space, one or more instructions suitable for being loaded and executed by a processor are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that more specific examples of the computer-readable storage medium here include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0217] The computer-readable storage medium also includes data signals propagated in a baseband or as part of a carrier wave, which carry readable program codes. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or component. The program codes contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, radio frequency, etc., or any suitable combination of the above.

[0218] The program codes for performing the operations of the present invention can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages - such as Java, C++, etc., and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program codes can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network or a wide area network, or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0219] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the method for calculating the density and area of abdominal muscles and fat in the above embodiments; the one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps:

[0220] Construct a deep network classification model based on the DenseNet121 network architecture to classify the lumbar vertebrae and determine the presence or absence of vertebral processes in the abdominal region CT image, and screen the best L3 image through the optimal L3 positioning algorithm; construct a ResWAA_DNet deep network segmentation model, use the CNN convolutional layer as the encoder and decoder, add a wide-area attention mechanism and cross-layer connection, and input the selected best L3 image into the ResWAA_DNet deep network segmentation model to obtain the corresponding semantic mask image; based on the semantic mask image of the best L3 image, calculate the density and area of abdominal muscles, subcutaneous fat, and visceral fat by combining the density algorithm based on the binary mask and the area algorithm based on Monte Carlo integration.

[0221] The databases involved in the embodiments provided in the present application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., and is not limited thereto. The processors involved in the embodiments provided in the present application may be a general-purpose processor, a central processor, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., and is not limited thereto.

[0222] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0223] Compared with the advanced model ResU-Net, the ResWAA-DNet model has improved by 0.5% in the average Dice coefficient, 0.92% in the average Jaccard coefficient, 0.74% in the average Precision, and 0.23% in the average Recall. The present invention has reached an advanced level in the automatic segmentation of abdominal CT images.

[0224] In summary, for the method and system for calculating the density and area of abdominal muscles and fat of the present invention, the abdominal CT images are preprocessed to enhance the image contrast; based on the DenseNet121 model, the lumbar vertebra classification model DenseNet_5 and the vertebral process judgment model DenseNet_3 are trained. The DenseNet_5 is used to classify the lumbar vertebra to which the abdominal region CT image belongs, and the image with the predicted class of L3 is input into the DenseNet_3 model to find the L3 image with vertebral processes, and the best L3 image is screened through the optimal L3 positioning algorithm; secondly, the ResWAA_DNet model learns the mapping from the abdominal CT image to its semantic segmentation mask image, and obtains the semantic segmentation results of abdominal muscles, subcutaneous fat, and visceral fat of the best L3 image. According to the segmentation results, the density and area of the above regions are obtained by using the density algorithm based on the binary mask and the area algorithm based on the Monte Carlo integral; among them, the ResWAA_DNet segmentation model uses the CNN convolutional layer as the encoder and decoder, and emphasizes the boundary of the target region based on the edge detection operator to optimize the boundary segmentation effect of the model on the target region; in view of the complex shape characteristics of the segmentation region, a wide-area attention mechanism and cross-layer connection are added to further improve the segmentation effect of the target region. The present invention reaches an advanced level in the automatic classification of lumbar vertebrae in abdominal CT images, the automatic positioning of the best L3 image, and the automatic segmentation of abdominal muscles and fat, and improves the accuracy of the density and area prediction results of related regions.

[0225] The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.

Claims

1. A method for calculating the density and area of abdominal muscles and fat, characterized in that, It includes the following steps: Construct a deep network classification model based on the DenseNet121 network architecture to classify the lumbar vertebrae and judge the presence or absence of vertebral processes in abdominal region CT images, and screen the best L3 images through the optimal L3 positioning algorithm; Construct a ResWAA_DNet deep network segmentation model, use the CNN convolutional layer as the encoder and decoder, add a wide-area attention mechanism and cross-layer connection, and input the screened best L3 images into the ResWAA_DNet deep network segmentation model to obtain the corresponding semantic mask image; Based on the semantic mask image of the best L3 image, calculate the density and area of abdominal muscles, subcutaneous fat, and visceral fat by combining the density algorithm based on the binary mask and the area algorithm based on Monte Carlo integration.

2. The method for calculating the density and area of abdominal muscles and fat according to claim 1, wherein Construct a deep network classification model based on the DenseNet121 network architecture to classify the lumbar vertebrae and judge the presence or absence of vertebral processes in abdominal region CT images, and screen the best L3 images through the optimal L3 positioning algorithm. Specifically: Input the CSV file with 5-class lumbar vertebra type annotation information and the classification training set into the DenseNet_5 model built based on the DenseNet121 deep network architecture, use the cross-entropy loss function containing mitigation of class imbalance to guide the training of the DenseNet_5 model to obtain the lumbar vertebra classification model; input the images in the classification test set into the lumbar vertebra classification model to obtain the lumbar vertebra classification results, evaluate the lumbar vertebra classification results, and draw the confusion matrix and ROC curve of the classification results; Perform windowing processing and central cropping on the L3-class images in the classification dataset, divide the processed images into three categories: with vertebral processes, unclear vertebral processes, and without vertebral processes according to the presence or absence of vertebral processes, input the CSV file with three-category presence or absence of vertebral process annotation information and the processed training set into the DenseNet_3 model built based on the DenseNet121 deep network architecture, use the cross-entropy loss function containing mitigation of class imbalance to guide the training of the DenseNet_3 model to obtain the vertebral process judgment model; input the images predicted as L3-class in the test set into the vertebral process judgment model to obtain the vertebral process judgment results of the L3-class images, evaluate the vertebral process judgment results, and draw the confusion matrix and ROC curve of the judgment results; use the optimal L3 positioning algorithm to screen the L3 images predicted as having vertebral processes to obtain the best L3 images.

3. The method for calculating the density and area of abdominal muscles and fat according to claim 2, characterized in that, Construct an abdominal CT image dataset, divide the dataset into an abdominal CT image classification dataset class_data1 and an abdominal CT image segmentation dataset seg_data1 according to the task, then perform preprocessing and image augmentation operations to obtain an augmented classification dataset Augclass_data1 and an augmented segmentation dataset Augseg_data1, and finally divide them into a training set, a validation set, and a test set.

4. The method for calculating the density and area of abdominal muscles and fat according to claim 2, wherein, Traverse the folders where different classes of images are located to generate the corresponding class One-Hot encoding, assign 1 to the corresponding class channel for images of the same class, and assign 0 to the remaining channels; save the image names and the corresponding class labels in a CSV file; Select the median of the number of 5 types of images from the first lumbar vertebra to the fifth lumbar vertebra as the category weight benchmark i. Divide the number of the 5 types of images by i respectively to obtain the category weight coefficients, and add the obtained weight coefficients to the cross-entropy loss function to guide the model training; Based on the DenseNet121 deep network framework, establish the lumbar vertebra classification model DenseNet_5, and obtain the predicted class label of the input image through the input layer, dense layer, transition layer and fully connected layer; Use the Adam optimizer with weight decay to minimize the loss function to train the DenseNet_5 model and obtain the DenseNet_5 lumbar vertebra classification model; Input the images in the test set Augclass_data1_test into the lumbar vertebra classification model DenseNet_5 to obtain the lumbar vertebra classification results of the test set, evaluate the lumbar vertebra classification results, and draw the confusion matrix and ROC curve of the classification results.

5. The method for calculating the density and area of abdominal muscles and fat according to claim 2, characterized in that, Perform windowing processing and central cropping on the L3-class images in the Augclass_data1 dataset to obtain images with a size of 224×224. Divide the processed images into three categories: with vertebral processes, inconspicuous vertebral processes, and without vertebral processes according to the presence or absence of vertebral processes. Traverse the folders where the images of different categories are located to generate the corresponding One-Hot encoding. Assign 1 to the corresponding category channel for the images of the same category, and assign 0 to the remaining channels; Save the image names and the corresponding category labels in a CSV file; Select the median of the number of 3 types of images as the category weight benchmark i. Divide the number of the 3 types of images by i respectively to obtain the category weight coefficients, and add the obtained weight coefficients to the cross-entropy loss function to guide the model training; Based on the DenseNet121 deep network framework, establish the vertebral process judgment model DenseNet_3, and obtain the predicted class label of the input image through the input layer, dense layer, transition layer and fully connected layer; Use the Adam optimizer with weight decay to minimize the loss function to train the DenseNet_3 model and obtain the DenseNet_3 vertebral process judgment model; Input the images predicted as L3-class images in the test set into the DenseNet_3 vertebral process judgment model to obtain the vertebral process judgment results of the L3-class images, evaluate the vertebral process judgment results, and draw the confusion matrix and ROC curve of the classification results; Utilize that the best L3 image occupies more bone pixels in adjacent images. Combining that bone pixels have a high CT value, select 150HU as the screening threshold for bone pixels. Count the number of pixels with a CT value greater than the screening threshold in each L3-class image with vertebral processes. Determine the image with the largest number of pixels in the same group as the best L3 image, and output the image number corresponding to the best L3 image.

6. The method for calculating the density and area of abdominal muscles and fat according to claim 1, characterized in that Construct the ResWAA_DNet deep network segmentation model, specifically as follows: The annotated mask images in the split training set Augseg_data1_train are mapped from single-channel grayscale images to four-channel semantic mask images containing the background class through One-Hot encoding. The semantic mask images and the training images are input into the ResWAA_DNet deep U-shaped network model built based on the CNN architecture and the attention mechanism. The hard Dice coefficient loss function with edge enhancement is used to guide the training of the ResWAA_DNet deep network model to obtain the ResWAA_DNet deep network segmentation model; the best L3 image selected is input into the ResWAA_DNet deep network segmentation model to obtain its semantic mask image.

7. The method for calculating the density and area of abdominal muscles and fat according to claim 6, characterized in that The Laplacian of Gaussian operator is used to detect the edges of the real muscle channel semantic mask image, and the Gaussian standard deviation σ of the LoG convolution kernel is set to 5. The pixel positions with pixel intensity greater than 0.001 in the image obtained by edge detection are regarded as outer edges, and the pixel positions with pixel intensity less than -0.001 are regarded as inner edges. The real and predicted semantic mask images are flattened into vectors. Then, the pixels at the corresponding positions of the inner and outer edges in the two images are respectively flattened into vectors again. The three vectors are concatenated in the order of mask image, inner edge, and outer edge. The Dice loss Loss between the vectors obtained after concatenation is calculated as follows: Loss = EE_DiceLoss(σ(Y), G) where Y is the predicted mask image output by the model, G is the real semantic mask image, EE_DiceLoss is the cross-entropy loss based on edge enhancement, and σ(·) is the Sigmoid activation function.

8. The method for calculating the density and area of abdominal muscles and fat according to claim 6, characterized in that The ResWAA_DNet deep U-shaped network model is specifically as follows: The input image resolution is 1×512×512. The ResWAA_DNet deep U-shaped network model includes an encoder, a decoder, and skip connections, and a residual structure is introduced in the encoding and decoding stages. In the input stage of each skip connection, a wide-area attention mechanism WAA and a channel attention mechanism SE are introduced. The WAA module performs strip pooling on the input image in the height and width directions to obtain two groups of matrices, performs information fusion in the channel direction and then restores the original channel size, multiplies the two groups of matrices to obtain a weight matrix with the same size as the input image, and multiplies the weight matrix with the original image to obtain the final output result; the image processed by the WAA module is input into the SE module, and then the squeeze-and-excitation mechanism is used to extract the corresponding weights in the channel direction and weight them to the original image. A DSC module is added between the architecture layers for cross-layer connection to make more full use of the image features. The encoding and decoding infrastructure of the ResWAA_DNet deep U-shaped network model is processed by two 3×3 convolutions combined with batch normalization and ReLU activation functions. In the encoding stage, the change of the image channels is 1-32-64-128-256-512, and in the decoding stage, the change of the image channels is 512-256-128-64-32-4. The final output result is used to calculate the loss with the four-channel semantic mask image.

9. The method for calculating the density and area of abdominal muscles and fat according to claim 1, characterized in that, Based on the best L3 image semantic mask map, the density and area of abdominal muscles, subcutaneous fat, and visceral fat are calculated by combining the density algorithm based on the binary mask and the area algorithm based on Monte Carlo integration. Specifically: According to the predicted semantic mask map corresponding to the best L3 image, the mask maps corresponding to muscles, subcutaneous fat, and visceral fat are respectively extracted from the predicted segmentation grayscale image according to the grayscale values corresponding to different regions; The mask maps of muscles, subcutaneous fat, and visceral fat are respectively binarized to obtain binary label images. After performing a dot product operation on the binary label images and the corresponding matrices of the original CT data, a density mask is obtained. The pixels within the mask are summed and then averaged to obtain the density of the corresponding region; Adopting the idea of Monte Carlo integration, the ratio of the number of pixels contained in muscles, subcutaneous fat, and visceral fat to the total number of pixels is regarded as the ratio of the areas of the three to the total area. Then, the actual pixel point spacing is extracted from the original CT data, the actual total area of the image is calculated, and the area of the relevant region is obtained according to the ratio.

10. A density and area calculation system for abdominal muscles and fat, characterized in that, Including: A data module that constructs a deep network classification model based on the DenseNet121 network architecture, classifies the lumbar vertebrae and determines the presence or absence of vertebral processes in abdominal region CT images, and screens the best L3 image through the best L3 positioning algorithm; A network module that constructs a ResWAA_DNet deep network segmentation model, uses the CNN convolutional layer as the encoder and decoder, adds a wide-area attention mechanism and cross-layer connection, and inputs the screened best L3 image into the ResWAA_DNet deep network segmentation model to obtain the corresponding semantic mask map; An output module that calculates the density and area of abdominal muscles, subcutaneous fat, and visceral fat based on the best L3 image semantic mask map by combining the density algorithm based on the binary mask and the area algorithm based on Monte Carlo integration.