Skeletal muscle radiomics feature extraction method and system for malnutrition assessment
Patent Information
- Application Number
- CN202610856516.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-11
AI Technical Summary
[0005]然而,现有利用胸部CT在T12层面评估营养不良的研究,多局限于骨骼肌指数(SMI)、骨骼肌放射密度(SMD)等宏观形态学参数
1.营养不良引起的骨骼肌变化属于微弱信号,若使用常规的特征筛选参数,极易被噪声干扰或直接遗漏;针对这一问题,本发明摒弃了传统的肿瘤组学流程, 针对性地采用“方差阈值-互信息-L1正则化”这一特定筛选流程,并通过严格的参数适配,从看似均一的T12水平骨骼肌影像中锁定24个核心特征;相较于单一筛选方法,该流程能够从海量数据中更精准地剔除噪声特征并保留与营养状况强相关的非线性特征,确保了输出特征集的高信噪比;经验证,这些特征(如wavelet-HLL_firstorder_Median等)直接指向肌肉微观结构的改变(如肌间脂肪浸润、肌纤维排列紊乱),实现了对肉眼无法分辨病变的高精度量化,打破了现有技术对于骨骼肌影像无深层信息的临床偏见;
Smart Images

Figure CN122737643A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical image processing and artificial intelligence technology, specifically to a method and system for extracting skeletal muscle radiomics features for malnutrition assessment. Background Technology
[0002] Malnutrition is a common clinical problem among hospitalized patients, characterized by insufficient intake or impaired utilization of energy, protein, and other nutrients, leading to altered body composition, decreased physiological function, and adverse clinical outcomes. Studies show that the incidence of malnutrition among hospitalized patients is as high as 20%-50%, and it is closely related to prolonged hospital stays, increased complications, higher mortality rates, and increased medical costs. Therefore, early and accurate identification and timely intervention of malnourished patients are crucial for improving prognosis and enhancing the quality of medical care.
[0003] However, the application of traditional nutritional assessment tools in hospitalized patients has many limitations. Body Mass Index (BMI) cannot distinguish between fat and muscle content in body composition and is easily affected by edema and dehydration. Serological indicators such as albumin (ALB) and prealbumin (PA) are suppressed by inflammatory factors during acute inflammation, resulting in decreased specificity and making it difficult to accurately reflect nutritional status. Composite nutritional assessment scales, such as Patient-Generated Subjective Global Assessment (PG-SGA), Nutritional Risk Screening 2002 (NRS 2002), and MiniNutritional Assessment-Short Form (MNA-SF), are more comprehensive, but they rely to some extent on the patient's subjective recollection and the assessor's experience, are time-consuming, and are subject to subjective bias; their standardization needs improvement. Therefore, there is an urgent clinical need for a new, objective, quantitative, and reproducible assessment method to compensate for the shortcomings of existing tools and achieve accurate and efficient prediction of malnutrition risk.
[0004] Skeletal muscle, as the largest protein storage depot in the human body, is a core pathophysiological feature of malnutrition, and changes in its quality are a key characteristic. Computed tomography (CT) scans can accurately assess the area and density of skeletal muscle and are considered the gold standard for assessing sarcopenia. Previous studies have largely focused on measuring the cross-sectional area of skeletal muscle at the third lumbar vertebra (L3) level and calculating the skeletal muscle index (L3-SMI) to assess overall muscle mass. However, chest CT scans are more commonly used in clinical practice, especially for respiratory and cardiovascular diseases and routine physical examinations. Multiple studies have confirmed that skeletal muscle area and density measured by chest CT at the twelfth thoracic vertebra (T12) level can also effectively reflect the overall muscle status and are significantly correlated with the prognosis of various diseases. This opportunistic screening method requires no additional scanning; chest CT alone can provide valuable information for nutritional assessment.
[0005] However, existing studies using chest CT at the T12 level to assess malnutrition are mostly limited to macroscopic morphological parameters such as skeletal muscle index (SMI) and skeletal muscle radiometric density (SMD). Since the microstructural changes in skeletal muscle caused by malnutrition (such as intramuscular fat infiltration and disordered muscle fiber arrangement) are weak signals, easily masked by noise, conventional feature extraction and screening methods struggle to capture effective deep radiomic features. Therefore, accurately extracting radiomic features reflecting microscopic pathological changes from noisy T12 images remains a significant technological gap. Summary of the Invention
[0006] To address the problems existing in the prior art, this invention proposes a method and system for extracting skeletal muscle radiomics features for malnutrition assessment. This method deeply mines the multi-level features of texture, morphology, and wavelet transform within skeletal muscle at the T12 level, and combines them with a specific parameter screening process to accurately capture and quantify changes in muscle microstructure that are indistinguishable to the naked eye, providing valuable decision-making basis for subsequent assessment and diagnosis of malnutrition.
[0007] In a first aspect, a method for extracting skeletal muscle radiomics features for malnutrition assessment is characterized by comprising: Obtain chest CT imaging data of the subject to be evaluated; Global preprocessing is performed on the chest CT image data. Based on the globally preprocessed chest CT image data, the target skeletal muscle region is located at the level of the 12th thoracic vertebral body. The target skeletal muscle region is extracted at the level of the 12th thoracic vertebral body, and a region of interest mask is generated. The image data within the region of interest mask is discretized in grayscale to construct its discretized voxel space. Using globally preprocessed chest CT image data as input images, skeletal muscle radiomics features are extracted in batches within the discretized voxel space of the region of interest mask. Based on the nutritional status assessment results of the subjects to be assessed, nutritional status classification labels corresponding to the skeletal muscle radiomics features are assigned; and an original radiomics feature set for all subjects to be assessed is constructed based on the skeletal muscle radiomics features and their nutritional status classification labels. Multi-level dimensionality reduction and screening were performed on the original radiomics feature set to obtain the target skeletal muscle radiomics features.
[0008] In one possible implementation of the first aspect, the global preprocessing of the chest CT image data, locating the target skeletal muscle region at the level of the twelfth thoracic vertebra based on the globally preprocessed chest CT image data, and generating a region of interest mask at the level of the twelfth thoracic vertebra, includes: Spatial resampling is performed on the chest CT image data, and a linear interpolation algorithm is used to make the chest CT image data reach a preset isotropic voxel spacing. Based on the resampled chest CT image data, the target skeletal muscle region was located at the level of the 12th thoracic vertebra and extracted at this level. A binary label matrix was generated to identify the spatial location of the target skeletal muscle region as an initial region of interest mask. The initial region of interest mask is spatially resampled, and the nearest neighbor interpolation algorithm is used to obtain the region of interest mask. The region of interest mask is spatially aligned with the spatially resampled chest CT image data and maintains binary attributes.
[0009] In one possible implementation of the first aspect, the extraction of the target skeletal muscle region at the level of the twelfth thoracic vertebral body includes: The target skeletal muscle region is automatically generated using image processing algorithms and / or manually delineated using human-computer interaction methods. In one possible implementation of the first aspect, the global preprocessing of the chest CT image data further includes: The spatially resampled chest CT image data is filtered and transformed to generate derived images containing different spatial frequency components and / or grayscale gradients.
[0010] In one possible implementation of the first aspect, the grayscale discretization processing of the image data within the region of interest mask to construct its discretized voxel space includes: For the image voxel data within the region of interest mask, grayscale discretization is performed according to a preset fixed bin width, and continuous CT values are discretized into a finite number of gray levels to construct the discretized voxel space of the region of interest mask.
[0011] In one possible implementation of the first aspect, the batch-extracted skeletal muscle radiomics features include morphological features, first-order statistical features, and higher-order texture features.
[0012] In one possible implementation of the first aspect, the multi-level dimensionality reduction screening of the original radiomics feature set to obtain target skeletal muscle radiomics features includes: Calculate the variance of each skeletal muscle radiomics feature in the original radiomics feature set, and remove invalid features with variances of zero or less than a preset threshold to obtain the first candidate feature set; Calculate the mutual information value between each skeletal muscle radiomics feature and the nutritional status classification label in the first candidate feature set, sort the skeletal muscle radiomics features in descending order according to the mutual information value, and retain the skeletal muscle radiomics features that rank first by a preset number or a preset proportion as the second candidate feature set. A logistic regression model with L1 regularization is constructed. Through grid search and 5-fold cross-validation, each candidate value in the preset regularization penalty coefficient candidate set is substituted into the logistic regression model for iterative evaluation. With the goal of optimizing the model classification performance, the optimal regularization penalty coefficient is determined from the regularization penalty coefficient candidate set. The optimal regularization penalty coefficient is used to fit the second candidate feature set, and skeletal muscle radiomics features with non-zero regression coefficients are extracted as target skeletal muscle radiomics features.
[0013] In one possible implementation of the first aspect, after obtaining the target skeletal muscle radiomics features through screening, the method further includes: Construct a feature evaluation model and train the feature evaluation model based on the training set; The importance / contribution ranking of the target skeletal muscle radiomics features output by the trained feature evaluation model is obtained. Based on importance / contribution, the top-ranked skeletal muscle radiomics features are retained.
[0014] In one possible implementation of the first aspect, the feature evaluation model is constructed based on a machine learning algorithm selected from random forest, extreme gradient boosting, lightweight gradient boosting machine, and neural network.
[0015] Secondly, this application provides a system for performing the skeletal muscle radiomics feature extraction method for malnutrition assessment, comprising: The image acquisition module is used to acquire chest CT image data of the subject to be evaluated. The image preprocessing and ROI localization module performs global preprocessing on the chest CT image data, locates the target skeletal muscle region at the level of the twelfth thoracic vertebra based on the globally preprocessed chest CT image data, extracts the target skeletal muscle region at the level of the twelfth thoracic vertebra, and generates a region of interest mask. The discretized voxel space construction module is used to perform grayscale discretization processing on the image data within the region of interest mask to construct its discretized voxel space; The radiomics feature extraction module uses globally preprocessed chest CT image data as input images to extract skeletal muscle radiomics features in batches within the discretized voxel space of the region of interest mask. The labeling and feature set construction module assigns nutritional status classification labels corresponding to the skeletal muscle radiomics features based on the nutritional status assessment results of the subjects to be evaluated; and constructs the original radiomics feature set for all subjects to be evaluated based on the skeletal muscle radiomics features and their nutritional status classification labels. The target feature dimensionality reduction and screening module is used to perform multi-level dimensionality reduction and screening on the original radiomics feature set to screen out the target skeletal muscle radiomics features.
[0016] Compared with existing methods, the present invention has the following advantages: 1. Malnutrition-induced changes in skeletal muscle are weak signals, easily susceptible to noise interference or direct omission when using conventional feature screening parameters. To address this issue, this invention abandons the traditional tumoromics process and specifically employs a screening process of "variance thresholding-mutual information-L1 regularization." Through rigorous parameter adaptation, it identifies 24 core features from seemingly homogeneous T12-level skeletal muscle images. Compared to single screening methods, this process can more accurately remove noisy features from massive datasets while retaining nonlinear features strongly correlated with nutritional status, ensuring a high signal-to-noise ratio in the output feature set. Verification has shown that these features (such as wavelet-HLL_firstorder_Median, etc.) directly point to changes in muscle microstructure (such as intramuscular fat infiltration and disordered muscle fiber arrangement), achieving high-precision quantification of lesions that are indistinguishable to the naked eye, breaking the clinical bias of existing technologies regarding the lack of deep information in skeletal muscle images. 2. This invention analyzes existing chest CT imaging data from clinical diagnosis and treatment. Without increasing additional examination costs and radiation dose, it can provide a reliable technical basis for large-scale, non-invasive, and objective assessment of malnutrition by utilizing the radiomics characteristics of skeletal muscle at the T12 level. This invention has significant clinical application value. Attached Figure Description
[0017] Figure 1A flowchart of a skeletal muscle radiomics feature extraction method for malnutrition assessment provided for an embodiment of the present invention; Figure 2 An architecture diagram of a skeletal muscle radiomics feature extraction system for malnutrition assessment provided for an embodiment of the present invention; Figure 3(a) is a horizontal plane schematic diagram of a T12 level skeletal muscle provided by an embodiment of the present invention; Figure 3(b) is a coronal view of a T12 level skeletal muscle provided by an embodiment of the present invention; Figure 4 A heatmap of correlation of target image omics features provided for embodiments of the present invention; Figure 5(a) shows the confusion matrix of a pure radiomics model based on the construction provided by an embodiment of the present invention on the test set; Figure 5(b) shows the confusion matrix of a pure radiomics model based on the construction provided by an embodiment of the present invention on the test set; Figure 5(c) shows the confusion matrix of a pure radiomics model based on the construction provided by an embodiment of the present invention on the test set; Figure 5(d) shows the confusion matrix of a pure radiomics model based on the construction provided by an embodiment of the present invention on the test set; Figure 5(e) shows the confusion matrix of a pure radiomics model based on the construction provided by an embodiment of the present invention on the test set; Figure 5(f) shows the confusion matrix of a pure radiomics model based on the construction provided by an embodiment of the present invention on the test set; Figure 6 ROC curves of a pure radiomics model constructed based on six different types of machine learning algorithms provided for embodiments of the present invention on a test set; Figure 7 The PR curve of a pure radiomics model built based on six different types of machine learning algorithms provided for embodiments of the present invention on a test set. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will now be described in further detail with reference to the accompanying drawings.
[0019] It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0020] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0021] It should be understood that the term "and / or" used in this article is merely a description of the same field for related objects, indicating that three relationships can exist.
[0022] like Figure 1 As shown, this application provides a method for extracting skeletal muscle radiomics features for malnutrition assessment, including: S100: Acquire chest CT image data of the subject to be evaluated; S200: Perform global preprocessing on chest CT image data, locate the level of the 12th thoracic vertebra (T12) based on the globally preprocessed chest CT image data, extract the target skeletal muscle region at the level of the T12 vertebral body, that is, the three-dimensional region of interest (ROI) of skeletal muscle, and generate the ROI mask. S300: Perform grayscale discretization on the image data within the region of interest mask to construct its discretized voxel space; S400: Using a radiomics feature extraction tool, with globally preprocessed chest CT image data as input images, skeletal muscle radiomics features limited to the region of interest mask are extracted in batches within the discretized voxel space of the region of interest mask; the skeletal muscle radiomics features include morphological features, first-order statistical features, and higher-order texture features. S500: Based on the nutritional status assessment results of the subjects to be assessed, assign nutritional status classification labels corresponding to the skeletal muscle radiomics features; construct the original radiomics feature set of all subjects to be assessed based on the skeletal muscle radiomics features and their nutritional status classification labels. S600: Performs multi-level dimensionality reduction screening on the original radiomics feature set, removes redundant features, and selects target skeletal muscle radiomics features that are significantly correlated with nutritional status assessment results.
[0023] In one possible implementation, step S200 includes the following: S210: Spatial resampling is performed on the chest CT image data, and a linear interpolation algorithm is used to make the chest CT image data reach a preset isotropic voxel spacing; for example, the linear interpolation algorithm can be trilinear interpolation or B-spline interpolation; S220: Based on the resampled chest CT image data, locate the target skeletal muscle region at the T12 vertebral level, extract the target skeletal muscle region at the T12 vertebral level, and generate a binary label matrix to identify the spatial location of the target skeletal muscle region as an initial region of interest mask. S230: Spatially resample the initial region of interest (ROI) mask and use a nearest neighbor interpolation algorithm to obtain a ROI mask that is spatially strictly aligned with the resampled chest CT image data and maintains binary attributes. This mask is used to subsequently limit the feature extraction range. In some possible implementations, in step S220, the target skeletal muscle region can be automatically generated by an image processing algorithm, obtained through a manual delineation method based on human-computer interaction, or obtained through a semi-automatic method combining an image processing algorithm and a manual delineation method. For example, the image processing algorithm can be a traditional image processing algorithm based on CT value thresholding, or a modern semantic segmentation algorithm based on a deep learning model (such as U-Net, nnU-Net, etc. network architectures). The semi-automatic method can be: first, use a traditional image processing algorithm based on CT value thresholding for preliminary segmentation, and then manually confirm, correct, or finely delineate the segmentation results based on prior anatomical knowledge to obtain the target skeletal muscle region.
[0024] Furthermore, step S200 also includes filtering and transformation processing, specifically: A filtering transformation is performed on spatially resampled chest CT image data to highlight the microscopic texture features in the chest CT image data. The filtering transformation generates derived images containing different spatial frequency components and / or gray-level gradients. The derived images are used to enhance the features of the chest CT image data from dimensions such as multi-scale spatial frequency, gray-level gradient, and local texture patterns, so as to comprehensively capture the heterogeneity and implicit information of the internal microstructure of skeletal muscle, thereby improving the representation ability of subsequent radiomics feature sets and the robustness of prediction models. For example, the filtering transformation includes at least one of wavelet transform, Laplace transform of Gaussian, square transform, square root transform, logarithmic transform, exponential transform, and gradient transform.
[0025] In one possible implementation, step S300 includes the following: For the image voxel data within the region of interest mask after spatial resampling, grayscale discretization is performed according to a preset fixed bin width, and continuous CT values are discretized into a finite number of gray levels to construct the discretized voxel space of the region of interest mask. It should be noted that the image data within the region of interest mask consists of multiple image voxel data arranged in three-dimensional space. Each image voxel data contains gray value information reflecting the physical properties of the tissue. The skeletal muscle radiomics features are calculated based on the spatial distribution and gray value of the image voxel data.
[0026] In one possible implementation, step S600 sequentially employs variance thresholding, mutual information, and L1 regularization to perform multi-level dimensionality reduction and screening on the original image omics feature set, specifically including the following: S610: Stability screening based on variance threshold method; For each skeletal muscle radiomics feature in the original radiomics feature set, its variance in all samples is calculated. After removing invalid features with a variance of zero or less than a preset threshold, the first candidate feature set is obtained. Specifically, the expression for the variance is: ; in, Represents the original image omics feature set The Middle Variance of individual skeletal muscle radiomic features; This represents the total number of samples, i.e., the total number of objects to be evaluated in the training set, where i represents the index number of the sample, i=1,2,…,N; Indicates the first The first sample The actual feature value of each skeletal muscle radiomics feature; Represents the original image omics feature set The Middle Mean value of each skeletal muscle radiomics feature. ; The filtering rule is: if (or less than the preset minimum threshold) This indicates that the value of this feature is almost the same across all samples, and it does not contribute to the classification task. Therefore, this feature should be removed from the original image omics feature set. Eliminate features from the remaining features to form the first candidate feature set. .
[0027] S620: Relationship pre-screening based on mutual information method; For the first candidate feature set after variance filtering, the mutual information (MI) value between each skeletal muscle radiomics feature and the nutritional status classification label is calculated. The skeletal muscle radiomics features are sorted in descending order according to the MI value, and the skeletal muscle radiomics features that rank first by a preset number or a preset proportion are retained as the second candidate feature set. Specifically, the expression for the mutual information value is: ; in, Represents the first candidate feature set The Middle skeletal muscle radiomics features (random variables) Nutritional status classification labels Mutual information values between them; Represents the first candidate feature set The Middle The set of all possible values for a skeletal muscle radiomics feature. Indicates the first The specific values of each skeletal muscle radiomics feature; This represents the set of all possible values for the nutritional status classification label. Indicates the specific values that the nutrition status classification label can take; Indicates the first The skeletal muscle radiomics feature is set to a value of The marginal probability distribution; The marginal probability distribution representing the nutritional status classification label y; Indicates the first The skeletal muscle radiomics feature is set to a value of And the joint probability distribution of the nutritional status classification label y; Represents the first candidate feature set Index number of radiomic features of skeletal muscle; Selection rule: For the first candidate feature set All skeletal muscle radiomics features are categorized by The values are sorted in descending order, with a preset ratio of [value]. (Preferably 30%), retain the values before sorting. The features constitute the second candidate feature set. .
[0028] S630: Sparsity filtering based on L1 regularization (LASSO): The second candidate feature set As input, construct a logistic regression model that includes an L1 regularization term; The objective function of the logistic regression model is: ;
[0029] Among them, the prediction probability of logistic regression With linear combination The relationship is: ; in, Indicates the total number of samples; Represents the second candidate feature set The total number of radiomic features of skeletal muscle. Represents the second candidate feature set Index number of radiomic features of skeletal muscle; Indicates the first The nutritional status classification label for each sample is either 0 (no malnutrition) or 1 (malnutrition). The model predicts the first... The probability that a sample belongs to the positive class (i.e., malnutrition); Indicates the first The sample at the th Actual observations on skeletal muscle radiomics features; This represents the intercept term (the constant term regression coefficient of the model); Represents the partial regression coefficient, i.e., the first... The weighting coefficients corresponding to each skeletal muscle radiomics feature, if If so, the feature is removed; This represents the regularization penalty coefficient; By using grid search and 5-fold cross-validation, each candidate value in the preset regularization penalty coefficient candidate set is substituted into the logistic regression model for iterative evaluation. With the goal of achieving the best classification performance of the model on the validation set, the optimal regularization penalty coefficient is determined from the candidate set. The optimal regularization penalty coefficient is then used to fit the second candidate feature set, and skeletal muscle radiomics features with non-zero regression coefficients are extracted as target skeletal muscle radiomics features.
[0030] Furthermore, after obtaining the target skeletal muscle radiomics features through step S600, the following content is also included: Construct a feature evaluation model and train the feature evaluation model based on the training set; Obtain the importance / contribution ranking or weight coefficient ranking of the target skeletal muscle radiomics features output by the trained feature evaluation model; Based on the sorting, the skeletal muscle radiomics features ranked first by a predetermined number are retained.
[0031] For example, the feature evaluation model can be constructed using machine learning algorithms such as Random Forest (RF), eXtreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), and neural networks.
[0032] like Figure 2As shown, based on the same inventive concept, the present invention also provides a system for performing the skeletal muscle radiomics feature extraction method for malnutrition assessment, comprising the following processing modules: The image acquisition module is used to acquire chest CT image data of the subject to be evaluated. The image preprocessing and ROI localization module performs global preprocessing on the chest CT image data, locates the target skeletal muscle region at the level of the twelfth thoracic vertebra based on the globally preprocessed chest CT image data, extracts the target skeletal muscle region at the level of the twelfth thoracic vertebra, and generates a region of interest mask. The discretized voxel space construction module is used to perform grayscale discretization processing on the image data within the region of interest mask to construct its discretized voxel space; The radiomics feature extraction module uses globally preprocessed chest CT image data as input images to extract skeletal muscle radiomics features in batches within the discretized voxel space of the region of interest mask. The labeling and feature set construction module assigns nutritional status classification labels corresponding to the skeletal muscle radiomics features based on the nutritional status assessment results of the subjects to be evaluated; and constructs the original radiomics feature set for all subjects to be evaluated based on the skeletal muscle radiomics features and their nutritional status classification labels. The target feature dimensionality reduction and screening module is used to perform multi-level dimensionality reduction and screening on the original radiomics feature set to screen out the target skeletal muscle radiomics features.
[0033] It should be understood that the above division of the processing modules in the system of this application is based on their logical functions. In practical applications, these processing modules can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the processing modules in the system can be implemented by a processor calling software; for example, the system includes a processor connected to memory, the memory storing instructions, and the processor calling the instructions stored in memory to implement any of the above methods or to implement the functions of each processing module in the system. The processor is a general-purpose processor, such as a central processing unit or a microprocessor, and the memory is either internal or external to the system.
[0034] The following detailed explanation of the implementation process of the present invention is provided in conjunction with an application example, and the feasibility and effectiveness of the present invention are illustrated.
[0035] I. Scanning Equipment and Methods Chest scans were performed using multi-slice spiral CT scanners (including 16-slice, 64-slice, and 128-slice models). Scanning parameters: tube voltage 120kV, tube current using automatic milliampere modulation technology, image reconstruction matrix 512×512, reconstruction slice thickness 5mm, and slice interval 5mm. All images were exported and stored in DICOM format. All subjects were in a supine position with arms raised, and the scan was performed while holding their breath after a deep inspiration. The scan range extended from the apex of the lung to the posterior costophrenic angle. Images were exported and stored in DICOM format.
[0036] II. Data Collection and Preprocessing A retrospective collection of chest CT images of inpatients at a hospital from 2024 to 2026 was conducted. Spatial resampling of chest CT images to 1×1×1mm 3 Isotropic voxels (using B-spline interpolation); based on this, multiple filter transformations are applied to the image to comprehensively capture the implicit information under different spatial scales and grayscale mapping modes; Based on whether there is a "malnutrition" assessment record in the medical record, patients are divided into a malnutrition group (label=1) and a non-malnutrition group (label=0); the assessment of malnutrition is determined by a clinical nutritionist based on the GLIM standard or NRS 2002 score ≥3 points and combined with clinical manifestations. A total of 257 patients were ultimately included, including 130 malnourished patients and 127 non-malnourished patients. The patient data were anonymized by removing or replacing information that could directly or indirectly identify individuals.
[0037] III. Skeletal Muscle ROI Delineation and Preprocessing Semi-automatic segmentation using thresholding was performed using the Segmentation Editor module of 3D Slicer software (version 5.8.1); 3D delineation of skeletal muscle ROIs was performed on axial CT images at the T12 vertebral level; the specific steps are as follows: (1) Locate the level of the T12 vertebral body in the chest CT image; (2) Set an appropriate CT value threshold range (usually -29 to 150 HU) to include muscle tissue and exclude bone and fat tissue; (3) Delineate the ROI of skeletal muscles, covering the skeletal muscle groups visible at the level of the T12 vertebra, including the erector spinae, latissimus dorsi, and abdominal muscles. (4) Check and correct the boundaries of skeletal muscle ROIs layer by layer to ensure that the muscle contour is complete and accurate; (5) Set the fixed box width to 25Hu and perform grayscale discretization on the skeletal muscle ROI; (6) Save the outlined and preprocessed skeletal muscle ROI as an NRRD format file, and store them in different folders according to the malnourished group (label 1) and the non-malnourished group (label 0); Figures 3(a) and 3(b) show two different orientations of the ROI of skeletal muscle at the T12 level in a patient.
[0038] IV. Feature Extraction and Screening In the Anaconda Jupyter Notebook environment, skeletal muscle radiomics features were extracted in batches from each ROI using the Python language and the PyRadiomics toolkit (version 3.0.1). Feature extraction parameters were set according to PyRadiomics' official recommendations and combined with the characteristics of skeletal muscle imaging, and image normalization and 3D feature extraction modes were enabled. The extraction process is as follows: (1) First, traverse the folders corresponding to each group, filter out all original image files, and automatically match the corresponding ROI mask file according to the file name; (2) For each pair of original images and ROI masks, the PyRadiomics feature extractor with pre-configured parameters is called to calculate the skeletal muscle radiomics features one by one. The final feature table contains 1634 columns including image number and label, excluding 37 columns of diagnostic information, and the number of effective features is 1595. The extracted skeletal muscle radiomics features can be divided into three categories according to their attributes: morphological features, first-order statistical features, and second-order and higher-order texture features. (3) Based on the different image source folders, the program automatically assigns corresponding diagnostic labels to each patient (0 for the non-malnourished group and 1 for the malnourished group). (4) Finally, the patient’s image number, classification label and all extracted features are merged into one row of data, and each case is added to the total data table and exported as a CSV file for subsequent feature screening and modeling.
[0039] The following multi-step progressive feature selection strategy is adopted: after feature selection is completed on the training set, the rules are applied to the test set: (1) First, calculate the variance of each feature in the training set and remove constant features with zero variance. Such features have a constant value in all samples and have no discriminative power; after this step, the number of features is reduced from 1595 to 1593, and 2 invalid features are eliminated; (2) Calculate the mutual information values between the remaining 1593 features and the malnutrition label, sort them in descending order and retain the features of the first preset proportion to narrow the search range for subsequent fine screening; In this embodiment, the performance of the feature subsets selected by subsequent L1 regularization in the classification task was systematically compared on the training set when the first 20%, 30%, and 40% of the features were retained; Experimental results show that when the first 30% are retained, the final model has the best evaluation performance and is the most stable; Therefore, retaining the first 30% of the features compresses the feature space to 477; (3) By applying a penalty term for the absolute value of the regression coefficients in the logistic regression loss function, some unimportant feature coefficients are compressed to zero, thereby achieving feature selection; In this embodiment, a logistic regression model with L1 regularization term is used, and the optimal regularization strength parameter C is determined by 5-fold cross-validation combined with grid search; After cross-validation, the optimal C value is 0.2, and under this parameter, the model retains 24 features with non-zero coefficients; The final 24 selected skeletal muscle radiomics features include 1 shape feature, 5 first-order statistical features, and 18 texture features. All features underwent filtering processing such as wavelet transform, Laplace transform of Gaussian, or exponential transform. The feature list is summarized by category as shown in Table 1. Table 1. Category distribution of radiomic features of target skeletal muscle
[0040] In Table 1, GLCM represents the Gray Level Co-occurrence Matrix, GLRLM is the Gray Level Run Length Matrix, GLSZM represents the Gray Level Size Zone Matrix, GLDM represents the Gray Level Dependence Matrix, and NGTDM represents the Neighboring Gray Tone Difference Matrix. The absolute values of the correlation coefficients among the 24 selected target skeletal muscle radiomics features were mostly below 0.7, with no obvious collinearity redundancy. The feature combinations were both independent and representative, providing concise and effective input variables for subsequent model construction.
[0041] Furthermore, the radiomics features of the 24 target skeletal muscle were sorted according to the absolute value of feature importance, as detailed in Table 2; the sorting in Table 2 is based on the absolute value of the L1 logistic regression coefficient, reflecting the feature importance under the linear model; Table 2. Radiomic features of 24 target skeletal muscles (sorted by absolute value of feature importance)
[0042] In Table 2, `wavelet-HLL_firstorder_Median` represents the first-order median of the high-low-low wavelet subband, `wavelet-LLL_glszm_LowGrayLevelZoneEmphasis` represents the gray-level region size matrix of the low-low-low wavelet subband, emphasizing the low gray-level region, `log-sigma-5-0-mm-3D_glcm_Imc2` represents the correlation coefficient 2 of the information measure of the gray-level co-occurrence matrix at a scale of 5 mm, `log-sigma-5-0-mm-3D_gldm_DependenceVariance` represents the dependency variance of the gray-level dependency matrix at a scale of 5 mm, `log-sigma-5-0-mm-3D_glcm_Correlation` represents the correlation of the gray-level co-occurrence matrix at a scale of 5 mm, and `wavelet-HLH_firstorder_Mean`... `wavelet-LHL_glcm_Idmn` represents the first-order mean of the high-low-high wavelet subband; `wavelet-LHL_glcm_Idmn` represents the normalized inverse difference moment of the gray-level co-occurrence matrix of the low-high-low wavelet subband; `exponential_gldm_DependenceVariance` represents the dependency variance of the gray-level dependency matrix after exponential transformation; `wavelet-LLL_glcm_MCC` represents the maximum correlation coefficient of the gray-level co-occurrence matrix of the low-low-low wavelet subband; `log-sigma-5-0-mm-3D_glcm_MaximumProbability` represents the maximum probability of the gray-level co-occurrence matrix of the Gaussian Laplace filter at a scale of 5 mm; `wavelet-LLL_glszm_LargeAreaEmphasis` represents the large area emphasis of the gray-level region size matrix of the low-low-low wavelet subband; `wavelet-LLL_glszm_ZoneVariance`... `log-sigma-3-0-mm-3D_firstorder_Kurtosis` represents the region variance of the gray-level region size matrix in the low-low-low wavelet subband; `log-sigma-3-0-mm-3D_firstorder_Kurtosis` represents the first-order kurtosis of the Gaussian Laplacian filter at a scale of 3 mm; `log-arithm_ngtdm_Busyness` represents the busyness of the neighborhood gray-level difference matrix after logarithmic transformation; `original_glcm_Imc2` represents the correlation coefficient 2 of the gray-level co-occurrence matrix information measure of the original image; `log-sigma-1-0-mm-3D_firstorder_Median` represents the first-order median of the Gaussian Laplacian filter at a scale of 1 mm; and `original_shape_MinorAxisLength` represents the length of the secondary axis of the three-dimensional shape of the original image.`log-sigma-3-0-mm-3D_glcm_SumAverage` represents the sum of the gray-level co-occurrence matrix and the average value of the Gaussian Laplacian filter at a scale of 3 mm. `exponential_firstorder_RootMeanSquared` represents the first-order root mean square after the exponential transformation. `squareroot_glrlm_GrayLevelNonUniformity` represents the gray-level run-length matrix gray-level non-uniformity after the square root transformation. `log-sigma-3-0-mm-3D_glszm_LargeAreaHighGrayLevelEmphasis` represents the large area high gray-level emphasis of the Gaussian Laplacian filter gray-level region size matrix at a scale of 3 mm. `log-sigma-3-0-mm-3D_glrlm_GrayLevelNonUniformity` represents the gray-level run-length matrix gray-level non-uniformity at a scale of 3 mm. `log-sigma-3-0-mm-3D_ngtdm_Coarseness` represents the gray-level non-uniformity of the Gaussian Laplacian filter gray-level run-length matrix at a scale of 3 mm. The roughness of the neighborhood gray-level difference matrix of the Gaussian Laplace filter at a scale of mm, denoted by log-sigma-3-0-mm-3D_glcm_Idn, represents the normalized inverse difference of the gray-level co-occurrence matrix of the Gaussian Laplace filter at a scale of 3 mm. The aforementioned features encompass both original and derived images, enabling a comprehensive and quantitative characterization of the heterogeneity and microstructural changes in the target skeletal muscle region across different spatial scales and frequency domains. For example... Figure 4 As shown, a correlation heatmap was plotted for the radiomics features of the 24 selected target skeletal muscle features; in the figure, the darker the color, the stronger the correlation between features; the lighter the color, the weaker the correlation.
[0043] VI. Feature Validity Verification To further verify the effectiveness of the 24 extracted target skeletal muscle radiomics features, this embodiment uses these 24 radiomics features and employs machine learning algorithms such as Logistic Regression (LR), Support Vector Machine (SVM), Random Forest (RF), NaiveBayes (NB), eXtreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM) to construct six pure radiomics models. Performance validation was performed on the training set (with 5-fold cross-validation) and the test set. The results show that each model achieved excellent performance.
[0044] Specifically, Figures 5(a) to 5(f) The confusion matrices of these six pure radiomics models on the test set are presented respectively, which intuitively reflects the misclassification distribution of each model; Figure 6 The ROC curves of each model on the test set are presented, along with the area under the ROC curve (AUC). This metric directly reflects the model's overall discriminative ability and generalization performance in distinguishing between positive and negative samples. Data shows that on the training set, the SVM model performed best (accuracy 0.922, AUC 0.972), followed by the LR model (accuracy 0.899, AUC 0.973). On the test set, the LightGBM model exhibited the best generalization ability, achieving an accuracy of 0.936 and an AUC as high as 0.976.
[0045] Furthermore, considering the prevalent problem of significant disparities in the number of positive and negative samples (i.e., class imbalance) in real clinical data, and the susceptibility of traditional assessment metrics to interference from the majority class, this embodiment further introduces the PR curve and its mean precision (AP) metric for in-depth evaluation, such as... Figure 7 As shown in the figure. Simulation results show that all models achieved extremely high AP values (0.956~0.981). This excellent AP value fully demonstrates the following two points: First, these 24 radiomics features can effectively overcome the model bias caused by class imbalance when identifying minority class samples. The model performance did not degrade due to the large number of negative samples, demonstrating extremely high robustness. Second, the model maintains extremely high precision while maintaining high recall, that is, it can effectively reduce misdiagnosis and avoid unnecessary medical intervention for healthy individuals without missing patients who truly need intervention.
[0046] In summary, this embodiment demonstrates that the 24 feature combinations extracted by this invention have strong clinical application value and can provide reliable and robust data support for subsequent precision medicine assessment.
[0047] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for extracting skeletal muscle radiomics features for malnutrition assessment, characterized in that, include: Obtain chest CT imaging data of the subject to be evaluated; Global preprocessing is performed on the chest CT image data. Based on the globally preprocessed chest CT image data, the target skeletal muscle region is located at the level of the 12th thoracic vertebral body. The target skeletal muscle region is extracted at the level of the 12th thoracic vertebral body, and a region of interest mask is generated. The image data within the region of interest mask is discretized in grayscale to construct its discretized voxel space. Using globally preprocessed chest CT image data as input images, skeletal muscle radiomics features are extracted in batches within the discretized voxel space of the region of interest mask. Based on the nutritional status assessment results of the subjects to be assessed, nutritional status classification labels corresponding to the skeletal muscle radiomics features are assigned; and an original radiomics feature set for all subjects to be assessed is constructed based on the skeletal muscle radiomics features and their nutritional status classification labels. Multi-level dimensionality reduction and screening were performed on the original radiomics feature set to obtain the target skeletal muscle radiomics features.
2. The method for extracting skeletal muscle radiomics features for malnutrition assessment according to claim 1, characterized in that, The process involves global preprocessing of chest CT image data, locating the target skeletal muscle region at the level of the twelfth thoracic vertebra based on the preprocessed chest CT image data, and generating a region of interest mask at the level of the twelfth thoracic vertebra. This includes: Spatial resampling is performed on the chest CT image data, and a linear interpolation algorithm is used to make the chest CT image data reach a preset isotropic voxel spacing. Based on the resampled chest CT image data, the target skeletal muscle region was located at the level of the 12th thoracic vertebra and extracted at this level. A binary label matrix was generated to identify the spatial location of the target skeletal muscle region as an initial region of interest mask. The initial region of interest mask is spatially resampled, and the nearest neighbor interpolation algorithm is used to obtain the region of interest mask. The region of interest mask is spatially aligned with the spatially resampled chest CT image data and maintains binary attributes.
3. The method for extracting skeletal muscle radiomics features for malnutrition assessment according to claim 2, characterized in that, The extraction of the target skeletal muscle region at the level of the twelfth thoracic vertebra includes: The target skeletal muscle region is automatically generated using image processing algorithms and / or manually delineated using human-computer interaction methods.
4. The method for extracting skeletal muscle radiomics features for malnutrition assessment according to claim 2, characterized in that, The global preprocessing of chest CT image data also includes: The spatially resampled chest CT image data is filtered and transformed to generate derived images containing different spatial frequency components and / or grayscale gradients.
5. The method for extracting skeletal muscle radiomics features for malnutrition assessment according to claim 1, characterized in that, The step of performing grayscale discretization processing on the image data within the region of interest mask to construct its discretized voxel space includes: For the image voxel data within the region of interest mask, grayscale discretization is performed according to a preset fixed bin width, and continuous CT values are discretized into a finite number of gray levels to construct the discretized voxel space of the region of interest mask.
6. The method for extracting skeletal muscle radiomics features for malnutrition assessment according to claim 1, characterized in that, The batch-extracted skeletal muscle image omics features include morphological features, first-order statistical features, and higher-order texture features.
7. The method for extracting skeletal muscle radiomics features for malnutrition assessment according to claim 1, characterized in that, The process of performing multi-level dimensionality reduction and screening on the original radiomics feature set to obtain target skeletal muscle radiomics features includes: Calculate the variance of each skeletal muscle radiomics feature in the original radiomics feature set, and remove invalid features with variances of zero or less than a preset threshold to obtain the first candidate feature set; Calculate the mutual information value between each skeletal muscle radiomics feature and the nutritional status classification label in the first candidate feature set, sort the skeletal muscle radiomics features in descending order according to the mutual information value, and retain the skeletal muscle radiomics features that rank first by a preset number or a preset proportion as the second candidate feature set. A logistic regression model with L1 regularization is constructed. Through grid search and 5-fold cross-validation, each candidate value in the preset regularization penalty coefficient candidate set is substituted into the logistic regression model for iterative evaluation. With the goal of optimizing the model classification performance, the optimal regularization penalty coefficient is determined from the regularization penalty coefficient candidate set. The optimal regularization penalty coefficient is used to fit the second candidate feature set, and skeletal muscle radiomics features with non-zero regression coefficients are extracted as target skeletal muscle radiomics features.
8. The method for extracting skeletal muscle radiomics features for malnutrition assessment according to claim 1, characterized in that, After obtaining the target skeletal muscle radiomics features through screening, the process also includes: Construct a feature evaluation model and train the feature evaluation model based on the training set; The importance / contribution ranking of the target skeletal muscle radiomics features output by the trained feature evaluation model is obtained. Based on importance / contribution, the top-ranked skeletal muscle radiomics features are retained.
9. The method for extracting skeletal muscle radiomics features for malnutrition assessment according to claim 8, characterized in that, The feature evaluation model is constructed based on one of the following machine learning algorithms: random forest, extreme gradient boosting, lightweight gradient boosting machine, and neural network.
10. A system for performing the skeletal muscle radiomics feature extraction method for malnutrition assessment as described in any one of claims 1-9, characterized in that, include: The image acquisition module is used to acquire chest CT image data of the subject to be evaluated. The image preprocessing and ROI localization module performs global preprocessing on the chest CT image data, locates the target skeletal muscle region at the level of the twelfth thoracic vertebra based on the globally preprocessed chest CT image data, extracts the target skeletal muscle region at the level of the twelfth thoracic vertebra, and generates a region of interest mask. The discretized voxel space construction module is used to perform grayscale discretization processing on the image data within the region of interest mask to construct its discretized voxel space; The radiomics feature extraction module uses globally preprocessed chest CT image data as input images to extract skeletal muscle radiomics features in batches within the discretized voxel space of the region of interest mask. The label assignment and feature set construction module assigns nutritional status classification labels corresponding to the skeletal muscle radiomics features based on the nutritional status assessment results of the object to be evaluated. Based on the skeletal muscle radiomics features and their nutritional status classification labels, an original radiomics feature set for all subjects to be evaluated was constructed. The target feature dimensionality reduction and screening module is used to perform multi-level dimensionality reduction and screening on the original radiomics feature set to screen out the target skeletal muscle radiomics features.