Method and system for acquiring head and neck squamous cell carcinoma immune predicted value based on machine learning

Through a machine learning-based method, combining clinical data and magnetic resonance image data, imagingomics and deep learning features are extracted, and the problem of inaccurate evaluation of immune responses in head and neck squamous cell carcinoma in the prior art is solved, more accurate immune prediction values ​​are achieved, and individual differences are improved.

CN120221107APending Publication Date: 2025-06-27SOUTHERN MEDICAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510221892.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to accurately evaluate the immune response of squamous cell carcinoma in the head and neck, and traditional imagingomics methods ignore individual differences, resulting in poor prediction results.

Method used

Using a machine learning-based method, we obtain clinical data and magnetic resonance image data of head and neck squamous cell carcinoma objects, perform preprocessing and feature extraction, combine imagingomics and deep learning features, and use risk prediction models to predict to obtain more accurate immune prediction values.

Benefits of technology

It improves the accuracy of immune prediction of squamous cell carcinoma in the head and neck, and can more effectively screen out patients who may benefit from immunotherapy, avoiding misjudgment caused by tumor heterogeneity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120221107A_ABST
    Figure CN120221107A_ABST
Patent Text Reader

Abstract

The invention relates to a machine learning-based head and neck squamous cell carcinoma immune prediction value acquisition method and system, and the machine learning-based head and neck squamous cell carcinoma immune prediction value acquisition method comprises five steps. Compared with the prior art, the method has the beneficial effects that 1, the problem that a traditional clinical medical method only depends on low-dimensional characteristics and is difficult to comprehensively evaluate tumor immune response is solved, and misjudgment caused by tumor heterogeneity is avoided; and 2, a radiomics feature extraction method is generally relatively fixed and ignores individual differences, but the method of the invention strives to explore more comprehensive features by combining a deep learning method, thereby improving the prediction accuracy. And 3, the most relevant diagnostic features are reserved through statistical feature screening and machine learning feature selection, and the weight value of the most relevant diagnostic features is calculated, so that the problem of insufficient interpretability of a deep learning method in medicine is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and particularly relates to a method and system for obtaining an immune prediction value of head and neck squamous cell carcinoma based on machine learning. Background Art

[0002] Head and neck tumors are the eighth most common cancers globally, mainly occurring in the oral cavity, pharynx, larynx, salivary glands, nasal cavity, and paranasal sinuses. Among them, head and neck squamous cell carcinoma (hereinafter referred to as head and neck squamous cancer) is the most common pathological type. This type of cancer is also relatively common globally, especially with a higher incidence rate among populations with common high-risk factors such as smoking and alcohol consumption. Most patients are diagnosed at an advanced stage, and local recurrence and metastasis are the main reasons for treatment failure. Due to the complex anatomical and physiological structure of the head and neck, head and neck squamous cell carcinoma (hereinafter referred to as HNSCC) has become a group of highly heterogeneous malignant tumors. Therefore, exploring effective treatment options is crucial for improving the prognosis of HNSCC patients.

[0003] Tumor immunotherapy is a treatment method that activates or enhances the body's own immune system to identify and attack tumor cells. Immune checkpoint inhibitors, such as cytotoxic T lymphocyte antigen-4 (CTLA-4), programmed death receptor 1 (PD-1), and their ligand (PD-L1) monoclonal antibodies, have significantly improved the clinical treatment effects of multiple solid tumors. Immunotherapy represented by anti-PD-1 monoclonal antibodies has shown excellent clinical efficacy in the first-line and second-line treatments of recurrent or metastatic head and neck squamous cancer. For head and neck squamous cell carcinoma, early research results have proven that neoadjuvant immunotherapy has a good pathological remission rate and significant tumor downstaging potential. However, head and neck squamous cancer is a highly heterogeneous tumor. For example, there are obvious individual differences in the response of patients to immunotherapy, the occurrence of drug resistance, treatment-related adverse reactions, etc. These characteristics lead to a generally low overall effective rate of immunotherapy. A considerable number of patients cannot benefit from immunotherapy, and immunotherapy may cause serious adverse reactions. Therefore, it is crucial to screen out the most likely beneficiaries from head and neck squamous cancer patients receiving immunotherapy.

[0004] Magnetic resonance imaging is the preferred imaging examination method for evaluating head and neck soft tissue lesions. It can clearly show the extent of lesion expansion and internal components, and has the advantages of high-resolution soft tissue imaging, multi-directional sectional imaging, no bone artifacts, and no ionizing radiation. Traditional imaging evaluation methods usually rely on the visual judgment of operators. This method is easily affected by the subjective opinions of operators, qualitative or semi-quantitative analysis, and the evaluation results may be restricted by the examination equipment, scanning sequences, and operator experience.

[0005] In recent years, the combination of medical image-assisted diagnosis and big data technology has given rise to a new imaging analysis method called "radiomics". This method efficiently extracts a large number of imaging features from tomographic scan images of head and neck squamous cell carcinoma, converting visual images into quantitative radiomic features, providing a new approach for the objective quantitative evaluation of tumors. However, the feature extraction process of traditional radiomics is relatively fixed, often ignoring the individual differences of patients and making it difficult to comprehensively reflect the heterogeneity of head and neck squamous cell carcinoma. Deep learning has made remarkable progress in medical image analysis, especially in natural image classification tasks, and can extract richer high-level semantic information than radiomics methods. However, due to the significant differences between magnetic resonance images of head and neck squamous cell carcinoma and traditional natural images, and the relatively small sample size, the performance of deep learning models still faces challenges and it is difficult to achieve ideal results. Due to the heterogeneity of tumors, a single feature extraction method often cannot comprehensively and accurately evaluate the immune response of tumors.

[0006] Therefore, in view of the deficiencies of the existing technology, it is very necessary to provide a method and system for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning to solve the deficiencies of the existing technology. Summary of the Invention

[0007] The first object of the present invention is to avoid the deficiencies of the existing technology and provide a method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning. The prediction value obtained by this method for obtaining the immune prediction value of head and neck squamous cell carcinoma is more accurate.

[0008] The above object of the present invention is achieved by the following technical measures:

[0009] Provide a method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning, including the following steps:

[0010] S1. Obtain the composition of data pairs of head and neck squamous cell carcinoma subjects, where the data pairs are composed of clinical data and corresponding magnetic resonance image data, and the magnetic resonance image data is a magnetic resonance image and a corresponding mask image, and the mask image is a labeled region of interest ROI;

[0011] S2. Preprocess the clinical data and magnetic resonance images obtained in S1 respectively to obtain preprocessed clinical data and preprocessed magnetic resonance images;

[0012] S3. Extract radiomic features and deep learning features from the preprocessed magnetic resonance images obtained in S2;

[0013] S4. Remove redundant and irrelevant features from the radiomic features and deep learning features obtained in S3 respectively to obtain screened radiomic features and screened deep learning features; and remove redundant and irrelevant features from the preprocessed clinical data obtained in S2 to obtain screened clinical data;

[0014] S5. Input the screened radiomics features, screened deep learning features, and screened clinical data obtained in S4 into the trained risk prediction model for prediction to obtain the trained risk prediction model. The risk prediction model is trained using a training set, and the training set consists of data pairs of different head and neck squamous cell carcinoma subjects. Each data pair in the training set consists of clinical data and corresponding magnetic resonance image data.

[0015] In S2, the preprocessing of the clinical data is a cleaning operation and a standardization operation;

[0016] The cleaning operation of the clinical data is carried out according to the following steps:

[0017] A1. Check all fields in the clinical data and delete the records with missing numerical values;

[0018] A2. Check the numerical range of the clinical data. When the numerical value exceeds the reasonable range, it is defined as an outlier, and the outlier is deleted, replaced, or marked;

[0019] A3. Check the duplication of the clinical data. When duplicate record data entries appear, retain the optimal data entry and delete the other data entries.

[0020] Preferably, the standardization operation of the above clinical data is Min - Max standardization operation or Z - Score standardization operation.

[0021] Preferably, the above Min - Max standardization operation is to scale the clinical data to a fixed interval according to formula (1);

[0022]

[0023] In formula (1), x is the original clinical data, x ′ is the scaled clinical data, x max and x min are the maximum and minimum values of this variable respectively.

[0024] Preferably, the above Z - Score standardization operation is to convert the clinical data into the standard normal distribution form after formula (2);

[0025]

[0026] In formula (2), x is the original clinical data, x ′ is the scaled clinical data, μ is the mean of the variable, and σ is the standard deviation of the variable.

[0027] In step S2, the preprocessing of the magnetic resonance images involves performing cleaning operations, resampling operations, and normalization operations.

[0028] Preferably, the cleaning operation of the above-mentioned magnetic resonance images is to delete the magnetic resonance images with low image quality or with mismatched dimensions between the magnetic resonance images and the mask images.

[0029] Preferably, the cleaning operation of the above-mentioned magnetic resonance images is carried out according to the following steps:

[0030] C1. Check the quality of all magnetic resonance images and delete the non-compliant magnetic resonance images;

[0031] C2. Check whether the dimensions of all magnetic resonance images match the corresponding mask images, and delete the group of magnetic resonance image data when the dimensions are inconsistent.

[0032] Preferably, the resampling operation of the above-mentioned magnetic resonance images is to resample the magnetic resonance images and the corresponding mask images to a fixed voxel size.

[0033] Preferably, the normalization operation of the above-mentioned magnetic resonance images is to map the pixel intensity values of the magnetic resonance images to a specific interval; the method of the normalization operation is Min-Max normalization or Z-Score normalization.

[0034] Preferably, the above-mentioned Min-Max normalization linearly transforms the pixel intensity values of the magnetic resonance images to a fixed range through Equation (3);

[0035]

[0036] In Equation (3), I is the original pixel intensity value of the magnetic resonance image, I ′ is the normalized pixel intensity value, I max and I min are the maximum and minimum pixel values of the image respectively.

[0037] Preferably, the above-mentioned Z-Score normalization subtracts the mean of the pixel intensity and divides by the standard deviation of the pixel intensity of the magnetic resonance image through Equation (4) to convert the data into the following form;

[0038]

[0039] In Equation (4), I is the original pixel intensity value of the magnetic resonance image, I ′ is the normalized pixel intensity value, μ is the mean of the pixel intensity, and σ is the standard deviation of the pixel intensity.

[0040] In step S3, the radiomics features and deep learning features in the preprocessed magnetic resonance images are extracted through the feature extraction module.

[0041] Preferably, the above-mentioned feature extraction module includes a radiomics feature extraction sub-module and a deep learning feature extraction sub-module.

[0042] Preferably, the above-mentioned radiomics feature extraction sub-module is the Software Environment for Radiomics Analysis (SERA) standardized based on the Image Biomarker Standardization Initiative (IBSI).

[0043] Preferably, the above-mentioned deep learning feature extraction sub-module uses the preprocessed magnetic resonance images as a fine-tuning training set, and then updates the weights of the pre-trained 3D-ResNet18 model in MedicalNet through the fine-tuning training set to obtain the deep learning feature extraction sub-module.

[0044] In the step S4, the method for deleting redundant and irrelevant features in the preprocessed clinical data obtained in S2 is as follows:

[0045] E1. Conduct univariate analysis, that is, conduct univariate statistical tests or univariate logistic regression analysis on each clinical variable respectively to evaluate its correlation with the efficacy of immunotherapy, and then proceed to E2;

[0046] E2. Conduct multivariate logistic regression analysis, and use stepwise regression to construct the final model to identify variables with independent predictive value;

[0047] E3. Calculate the p-value and the odds ratio (OR) value of the 95% confidence interval for each variable in E1 and E2 to screen out the variables in the clinical data, and these clinical variables are the screened clinical data.

[0048] In the step S4, the method for deleting redundant and irrelevant features in the radiomics features and deep learning features obtained in S2 and S3 is as follows:

[0049] F1. Combine the radiomics features and deep learning features into a feature matrix X, and \(X\in R^{n\times p}\), where n is the number of samples and p is the number of features; in the feature matrix, each radiomics feature and each deep learning feature have been standardized so that the mean is 0 and the standard deviation is 1; n×p where \(y_{i}\) is the label value of the \(i\)-th feature, \(X_{ij}\) is the value of the \(j\)-th feature of the \(i\)-th sample, \(\beta_{j}\) is the regression coefficient of the \(j\)-th feature, \(\lambda\) is the regularization parameter and \(\lambda\geq0\), where \(\lambda\) is determined by cross-validation.

[0050] F2. Construct a LASSO regression model for screening, where the objective function of the LASSO regression model is represented by;

[0051]

[0052] where, \(y\) o is the label value of the \(i\)-th feature, \(X\) ij is the value of the \(j\)-th feature of the \(i\)-th sample, \(\beta\) j is the regression coefficient of the \(j\)-th feature, \(\lambda\) is the regularization parameter and \(\lambda\geq0\), where \(\lambda\) is determined by cross-validation, is the regression coefficient vector estimated in the LASSO regression model; β is the regression coefficient; n is the number of samples; p is the number of features;

[0053] F3. In the LASSO regression model, select the non - zero regression coefficients β j corresponding features;

[0054] F4. After normalizing the absolute values of the non - zero coefficients of the features selected by Equation (6), the normalized absolute values are used as the weight values of each feature, and these selected features are used as the screened radiomics features or the screened deep - learning features;

[0055]

[0056] where S is the set of feature indices corresponding to all non - zero coefficients, ω j is the weight value of feature j; Σ k∈S |β k | is the sum of the absolute values of the regression coefficients of all selected key features.

[0057] Preferably, the above - mentioned risk prediction model includes a feature fusion sub - module, a data balancing sub - module, and a classification modeling sub - module.

[0058] Preferably, the training of the above - mentioned risk prediction model is carried out by the following steps:

[0059] G.1. Input the data in the training set into the feature fusion sub - module for splicing to form a complete comprehensive feature vector; the feature fusion sub - module uses horizontal splicing for splicing; the training set includes data of different category objects, and each category of object data includes screened radiomics features, screened deep - learning features, and screened clinical data;

[0060] G.2. The data balancing sub - module uses the SMOTE method to expand the comprehensive feature vectors corresponding to the minority classes to obtain synthetic comprehensive feature vectors;

[0061] G.3. Input the comprehensive feature vector obtained in G.1 and the synthetic comprehensive feature vector obtained in G.2 into the classification modeling sub - module for training to obtain a trained classification modeling sub - module; the classification modeling sub - module is a linear classification model.

[0062] Preferably, the training of the above - mentioned classification modeling sub - module uses the log - likelihood function as the loss function to evaluate the difference between the prediction and the true label, maximizes the log - likelihood, and then updates the parameters through the gradient descent update formula until the training ends.

[0063] Preferably, the above - mentioned log - likelihood function is represented by Equation (7);

[0064]

[0066] Among them, w is the weight vector of the classification modeling sub-module, and b is the bias.

[0067] Preferably, the above gradient descent update formula is represented by Equation (8);

[0068]

[0069] Among them, η is the learning rate, is the gradient of w and b.

[0070] Preferably, the above S5 is specifically carried out through the following steps:

[0071] H.1. Concatenate the screened radiomics features, screened deep learning features, and screened clinical data input features obtained in S4 into the feature fusion sub-module to form a complete comprehensive feature vector;

[0072] H.2. Input the comprehensive feature vector obtained in H.1 into the trained classification modeling sub-module for prediction to obtain the immunological prediction value of head and neck squamous cell carcinoma.

[0073] Preferably, the above immunological prediction value of head and neck squamous cell carcinoma is represented by Equation (9);

[0074]

[0075] The second object of the present invention is to provide an immunological prediction value prediction system for head and neck squamous cell carcinoma to avoid the deficiencies of the prior art. The immunological prediction value prediction system for head and neck squamous cell carcinoma can obtain a prediction value, and this prediction value is more accurate than the prior art.

[0076] The above object of the present invention is achieved by the following technical measures:

[0077] Provide an immunological prediction value prediction system for head and neck squamous cell carcinoma, and adopt the above method for obtaining the immunological prediction value of head and neck squamous cell carcinoma based on machine learning.

[0078] The immunological prediction value prediction system for head and neck squamous cell carcinoma is provided with:

[0079] Input unit - used to collect the clinical data and magnetic resonance image data of head and neck squamous cell carcinoma patients;

[0080] Preprocessing module - preprocess the clinical data and magnetic resonance images to obtain preprocessed clinical data and preprocessed magnetic resonance images;

[0081] Feature extraction module - extract radiomics features and deep learning features from the preprocessed magnetic resonance images;

[0082] Feature screening module - deleting redundant and irrelevant features in radiomics features and deep learning features, and deleting redundant and irrelevant features in preprocessed clinical data;

[0083] Multimodal prediction unit - combining the screened radiomics features, the screened deep learning features and the screened clinical data, then predicting to obtain the immune prediction value of head and neck squamous cell carcinoma, and finally evaluating the immune prediction value of head and neck squamous cell carcinoma.

[0084] A method and system for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning. The method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning includes the following steps: S1. Obtain the composition of data pairs of head and neck squamous cell carcinoma subjects. The data pairs are composed of clinical data and corresponding magnetic resonance image data. The magnetic resonance image data is a magnetic resonance image and a corresponding mask image, and the mask image is a labeled region of interest ROI; S2. Preprocess the clinical data and magnetic resonance images obtained in S1 respectively to obtain preprocessed clinical data and preprocessed magnetic resonance images; S3. Extract radiomics features and deep learning features from the preprocessed magnetic resonance images obtained in S2; S4. Remove redundant and irrelevant features from the radiomics features and deep learning features obtained in S3 respectively to obtain the screened radiomics features and the screened deep learning features correspondingly; and remove redundant and irrelevant features from the preprocessed clinical data obtained in S2 to obtain the screened clinical data; S5. Input the screened radiomics features, the screened deep learning features and the screened clinical data obtained in S4 into a trained risk prediction model for prediction to obtain the trained risk prediction model. The risk prediction model is trained by a training set, and the training set is composed of data pairs of different head and neck squamous cell carcinoma subjects. Each pair of data pairs in the training set is composed of clinical data and corresponding magnetic resonance image data. Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Solve the problem that traditional clinical medicine methods only rely on low-dimensional features and it is difficult to comprehensively evaluate the tumor immune response, and avoid misjudgment caused by tumor heterogeneity. 2. The radiomics feature extraction method is usually relatively fixed and ignores individual differences. The present invention attempts to discover more comprehensive features by combining deep learning methods, thereby improving the accuracy of prediction. 3. Retain the most relevant diagnostic features through statistical feature screening and machine learning feature selection, and calculate their weight values, effectively solving the problem of insufficient interpretability of deep learning methods in medicine. Description of the Drawings

[0085] The present invention is further illustrated by the accompanying drawings, but the content in the drawings does not constitute any limitation to the present invention.

[0086] Figure 1It is a flowchart of a method for obtaining an immune prediction value of head and neck squamous cell carcinoma based on machine learning.

[0087] Figure 2 It is a structural diagram of the 3D-ResNet18 model.

[0088] Figure 3 It is a process diagram of the radiomics features and deep learning features in the selection of LASSO model parameters in Example 2.

[0089] Figure 4 It is a feature diagram of the radiomics features and deep learning features that affect the immune efficacy response and are retained after screening by the LASSO model in Example 2.

[0090] Figure 5 It is a comparison diagram of the ROC curves of the four models in the training set in Example 2.

[0091] Figure 6 It is a comparison diagram of the ROC curves of the four models in the internal test set in Example 2.

[0092] Figure 7 It is a comparison diagram of the ROC curves of the external validation sets of the four models in Example 2. Detailed implementation manners

[0093] The technical solution of the present invention will be further described in conjunction with the following embodiments.

[0094] Example 1

[0095] A method for obtaining an immune prediction value of head and neck squamous cell carcinoma based on machine learning, as Figure 1 shown, includes the following steps:

[0096] S1. Obtain the composition of data pairs of head and neck squamous cell carcinoma subjects. The data pairs are composed of clinical data and corresponding magnetic resonance image data, where the magnetic resonance image data is a magnetic resonance image and a corresponding mask image, and the mask image is a labeled region of interest ROI; among them, the clinical data can be at least one of the basic information of head and neck squamous cell carcinoma subjects, disease stage, or blood indexes.

[0097] S2. Preprocess the clinical data and magnetic resonance images obtained in S1 respectively to obtain preprocessed clinical data and preprocessed magnetic resonance images;

[0098] S3. Extract radiomics features and deep learning features from the preprocessed magnetic resonance images obtained in S2;

[0099] S4. Remove redundant and irrelevant features from the radiomics features and deep learning features obtained in S3, respectively, to obtain the screened radiomics features and the screened deep learning features; and remove redundant and irrelevant features from the preprocessed clinical data obtained in S2 to obtain the screened clinical data.

[0100] S5. Input the screened radiomics features, the screened deep learning features, and the screened clinical data obtained in S4 into the trained risk prediction model for prediction to obtain the trained risk prediction model. The risk prediction model is trained through a training set, and the training set is composed of data pairs of different head and neck squamous cell carcinoma subjects. Each data pair in the training set consists of clinical data and the corresponding magnetic resonance image data.

[0101] In S2, the preprocessing of clinical data is cleaning operation and standardization operation.

[0102] The cleaning operation of clinical data is carried out according to the following steps:

[0103] A1. Check all fields in the clinical data and delete the records with missing numerical values.

[0104] A2. Check the numerical range of the clinical data. When the numerical value exceeds the reasonable range, it is defined as an outlier, and the outlier is deleted, replaced, or marked.

[0105] A3. Check the duplication situation of the clinical data. When there are duplicate record data entries, retain the optimal data entry and delete the other data entries. The optimal data entry can be the most complete and accurate data entry.

[0106] It should be noted that the so-called missing numerical values in the present invention usually manifest as that some indicators of some objects are not recorded. For example, there are numerical vacancies in some key indicators (such as tumor stage). For the objects with missing key indicators, their data records will be directly deleted. Secondly, for the problem of numerical anomalies, a rationality analysis is carried out on all numerical data. Combining the actual meaning and the normal reference range of each indicator, screen whether there are outliers exceeding the reasonable range in the data. For the detected outliers, processing methods such as deletion, replacement, or marking can be adopted according to the specific situation. For obvious logical errors or outliers beyond physical possibility (such as negative height or extremely deviated data), they are directly deleted or corrected to ensure the accuracy of the data. Finally, detect and process the possible duplication situation in the data. Check the data records through the unique identifier of the object (such as ID or hospital number) to check whether there are duplicate entries. For the data redundancy caused by multiple records, give priority to retaining the complete and accurate data entries and deleting the duplicate or unnecessary records.

[0107] The standardization operation of clinical data is Min - Max standardization operation or Z - Score standardization operation.

[0108] It should be noted that the process of standardizing clinical data includes the following: for continuous variables, appropriate standardization methods are selected for processing in combination with their characteristics and analysis requirements to eliminate the influence of different variable dimensions and value ranges, and to ensure the stability and accuracy of model training and analysis results.

[0109] Among them, the Min-Max standardization operation scales the clinical data to a fixed interval according to Equation (1);

[0110]

[0111] In Equation (1), x is the original clinical data, x ′ is the scaled clinical data, x max and x min are the maximum and minimum values of this variable respectively. The fixed interval is usually [0,1]. The Min-Max standardization operation is applicable to variables with a fixed range, can retain the relative magnitude of the data and make it more suitable for use in range-sensitive models.

[0112] The Z-Score standardization operation converts the clinical data into the form of a standard normal distribution through Equation (2);

[0113]

[0114] In Equation (2), x is the original clinical data, x ′ is the scaled clinical data, μ is the mean of the variable, and σ is the standard deviation of the variable. The form of the standard normal distribution is specifically a mean of 0 and a standard deviation of 1. The Z-Score standardization operation is applicable to variables without a fixed value range or with a relatively wide distribution, and can more effectively eliminate the dimensional difference and make the data comparable between different variables.

[0115] In S2, the preprocessing of magnetic resonance images is to perform cleaning operations, resampling operations, and normalization operations;

[0116] The cleaning operation of magnetic resonance images is to delete the magnetic resonance images that do not meet the specifications or the dimensions of the magnetic resonance images do not match the mask images;

[0117] The cleaning operation of magnetic resonance images is carried out according to the following steps:

[0118] C1. Check the quality of all magnetic resonance images and delete the magnetic resonance images that do not meet the specifications; among them, the images that do not meet the specifications are manifested as problems such as blurred images, excessive noise, and serious artifacts.

[0119] C2. Check whether the dimensions of all magnetic resonance images match the corresponding mask images. When the dimensions are inconsistent, delete the group of magnetic resonance image data.

[0120] Substandard magnetic resonance images refer to images that are blurred, have excessive noise, or severe artifacts. For blurred images, the Laplace transform can be used to calculate the image sharpness. If the variance is lower than the set threshold (e.g., <100), it is determined as a blurred image. For excessive noise, the signal-to-noise ratio (SNR) or the peak signal-to-noise ratio (PSNR) can be calculated. If the SNR is lower than 10 dB or the PSNR is lower than 30 dB, it is considered that there is excessive noise. For severe artifact cases, the structural similarity index (SSIM) can be used to evaluate the image quality. If the SSIM is lower than 0.8, there may be severe artifacts. At the same time, the Fourier transform can be used to analyze the distribution of high-frequency noises such as circular artifacts and stripe artifacts. If the abnormal high-frequency components exceed the set threshold (e.g., more than 5%), it is determined as severe artifacts.

[0121] It should be noted that substandard images show problems such as blurred images, excessive noise, and severe artifacts. Substandard images may be caused by imaging device failures, patient movement, or improper imaging parameter settings. Such substandard images will affect subsequent feature extraction and model analysis. Therefore, it is necessary to screen and delete the images with the above problems to ensure the quality and consistency of the dataset. In the present invention, checking whether the dimensions of the magnetic resonance images match the corresponding mask images. Since the mask image is usually used to label the region of interest (ROI), its dimensions should be exactly the same as those of the original magnetic resonance image. If the dimensions of the two do not match, the feature extraction work cannot be carried out normally. Therefore, for the case of dimension mismatch, it is necessary to carefully check the data source and processing flow to confirm the root cause of the problem. For images with dimension mismatch that cannot be solved by correction, they should be removed from the dataset. Through the cleaning operation of magnetic resonance images, substandard images can be removed, which can effectively improve the reliability of the data and the accuracy of the analysis results, providing a high-quality imaging data basis for subsequent research and modeling.

[0122] The resampling operation of magnetic resonance images is to resample the magnetic resonance images and the corresponding mask images to a fixed voxel size. The purpose of the resampling operation of magnetic resonance images is to adjust the spatial resolution of the images, make the pixel sizes of different images consistent, so as to facilitate subsequent image registration and analysis.

[0123] The normalization operation of magnetic resonance images is to map the pixel intensity values of magnetic resonance images to a specific interval; the methods of the normalization operation are Min-Max normalization or Z-Score normalization. The purpose of the normalization operation of magnetic resonance images is to reduce or eliminate the inconsistency in the gray-scale information of the same tissue caused by factors such as acquisition and imaging of different devices. Mapping the pixel intensity values of magnetic resonance images to a specific interval aims to reduce or eliminate the problem of inconsistent gray-scale information caused by differences in different imaging devices, scanning parameters, and other imaging conditions. The main purpose of this normalization process is to eliminate the bias introduced by device or environmental differences in data preprocessing, so as to ensure higher consistency in the gray-scale values of the same tissue in the image.

[0124] Among them, Min-Max normalization linearly transforms the pixel intensity values of magnetic resonance images to a fixed range through Equation (3);

[0125]

[0126] In Equation (3), I is the original pixel intensity value of the magnetic resonance image, and I ′ is the normalized pixel intensity value, I max and I min are the maximum and minimum pixel values of the image respectively. The fixed range can be [0,1]. Min-Max normalization is applicable to pixel values with a fixed range and can retain the relative gray-scale distribution information of the image.

[0127] Among them, Z-Score normalization subtracts the mean of the pixel intensity and divides by the standard deviation of the pixel intensity values of magnetic resonance images through Equation (4) to convert the data into the following form;

[0128]

[0129] In Equation (4), I is the original pixel intensity value of the magnetic resonance image, and I ′ is the normalized pixel intensity value, μ is the mean of the pixel intensity, and σ is the standard deviation of the pixel intensity. The mean in the form of the present invention is 0 and the standard deviation is 1. Z-Score normalization is applicable to data with a wide distribution range and requires eliminating the influence of dimensions, and can improve the comparability between different data sets.

[0130] In S3, radiomics features and deep learning features are extracted from the preprocessed magnetic resonance images through the feature extraction module. The feature extraction module includes a radiomics feature extraction sub-module and a deep learning feature extraction sub-module.

[0131] The radiomics feature extraction sub-module is the SERA software package for radiomics analysis standardization environment based on the Image Biomarker Standardization Initiative IBSI.

[0132] It should be noted that the standardized environment software package SERA for radiomics analysis based on the Image Biomarker Standardization Initiative (IBSI) calculates radiomics features according to the guidelines of the Image Biomarker Standardization Initiative (IBSI). The SERA package can calculate 487 IBSI-standardized features, including: 79 first-order features (morphological features, statistical features, histogram features, and intensity histogram features), 272 higher-order two-dimensional features, and 136 three-dimensional features. In addition, it also calculates 10 moment invariant features not included in the IBSI standard. Therefore, different radiomics features can be manually selected for calculation, and these features constitute different feature subsets, that is, different radiomics feature vectors. When using the standardized environment software package SERA for radiomics analysis based on the Image Biomarker Standardization Initiative (IBSI) in the present invention, the input path is adjusted to the location of the preprocessed magnetic resonance image on the local machine, and the number of features is adjusted to 94. The present invention extracts radiomics features from the preprocessed magnetic resonance image, including 8 first-order features and 86 higher-order texture features, and stores these features to form a radiomics feature vector.

[0133] The deep learning feature extraction sub-module uses the preprocessed magnetic resonance image as a fine-tuning training set, and then updates the weights of the pre-trained 3D-ResNet18 model in MedicalNet through the fine-tuning training set to obtain the deep learning feature extraction sub-module. The deep learning feature extraction sub-module uses a convolutional neural network to perform multi-layer feature extraction on the input image, and finally obtains 512 deep learning features through the global average pooling layer, and stores these features to form a deep learning feature vector.

[0134] It should be noted that the deep learning feature extraction sub-module is integrated by different convolutional neural networks for classification tasks. The process from the input image to the output deep learning feature vector can be simply summarized as follows: First, the processed magnetic resonance image is input into the network, and the size is adjusted to fit the network structure to ensure the consistency of the input data. In the convolutional layer, the convolutional kernel is used to gradually scan and extract features from the local area of the image. The low-level convolutional layer extracts basic edge and texture information, the middle-level convolutional layer identifies more complex shapes and patterns, and the high-level convolutional layer extracts global semantic features. After each convolutional operation, a non-linear transformation is introduced through an activation function such as ReLU to enhance the network's ability to express complex features. Subsequently, the pooling layer reduces the dimensionality of the feature map, retains the main information, reduces the computational burden, and enhances the translational invariance of the features. After multiple layers of convolution and pooling are stacked, the network gradually learns multi-scale and multi-level high-order non-linear features, and finally converts the extracted feature map into a one-dimensional feature vector through the fully connected layer. These features can comprehensively characterize the spatial structure, texture pattern, and semantic information of the image.

[0135] It should also be noted that the pre-trained 3D-ResNet18 model of the present invention is the 3D-ResNet18 model that comes with MedicalNet, such as Figure 2 , the model has been pre-trained. The present invention directly uses the pre-trained 3D-ResNet18 model, and then retrains it through the fine-tuning training set of the present invention. The obtained deep learning feature extraction submodule extracts high-dimensional feature maps from the last convolutional layer. These feature maps are converted into feature vectors through global average pooling, and these vectors are stored to form deep learning feature vectors. The fine-tuning training set of the present invention is also a magnetic resonance image for training processed by S1 and S2.

[0136] In S4, the method for removing redundant and irrelevant features from the preprocessed clinical data obtained in S2 is as follows:

[0137] E1. Perform univariate analysis, that is, perform univariate statistical test or univariate logistic regression analysis on each clinical variable to evaluate its correlation with the efficacy of immunotherapy, and proceed to E2;

[0138] E2. Based on the variables screened by univariate analysis, multivariate logistic regression analysis was performed and the final model was constructed using stepwise regression to identify variables with independent predictive value;

[0139] E3. Calculate the p value of each variable in E1 and E2 and the OR value with 95% confidence interval to screen out the variables of clinical data. This part of clinical variables is the clinical data after screening. Generally, the significance level (such as p<0.05) is set to screen out the variables that may be related. This part of clinical variables is the clinical data after screening.

[0140] In S4, the method for removing redundant and irrelevant features from the radiomics features and deep learning features obtained in S3 is as follows:

[0141] F1. Combine radiomics features and deep learning features into a feature matrix X, where X∈R n×p , where n is the number of samples and p is the number of features; in the feature matrix, each radiomics feature and each deep learning feature are standardized so that the mean is 0 and the standard deviation is 1;

[0142] F2, construct LASSO regression model screening, where the objective function of the LASSO regression model is represented by;

[0143]

[0144] Among them, y i is the label value of the i-th sample, X ij is the value of the jth feature of the i-th sample, β jis the regression coefficient for the j-th feature, λ is the regularization parameter and λ ≥ 0, where λ is determined by cross-validation, is the vector of regression coefficients estimated in the LASSO regression model, β is the regression coefficient, n is the number of samples, and p is the number of features. The regression coefficient β represents the magnitude of the influence of the feature variable on the predicted variable. The number of samples n is the number of observed samples in the dataset, and the number of features p is the number of features used for modeling in the dataset. And a sample refers to a head and neck squamous cell carcinoma subject.

[0145] It should be noted that LASSO is a linear regression method with L1 regularization. By imposing a penalty term on the regression coefficients, some coefficients are forced to shrink to zero, thus achieving feature selection; and the determination of λ is to select the optimal λ through cross-validation (usually k-fold cross-validation) to balance the fitting ability of the model and the sparsity of the features; in cross-validation, the performance of the model is evaluated for each λ value, and the λ that minimizes the prediction error (such as the mean squared error MSE or the classification error rate) is selected;

[0146] F3. In the LASSO regression model, select the non-zero regression coefficients β j corresponding features;

[0147] F4. After normalizing the absolute values of the non-zero coefficients of the features selected by Equation (6), the normalized absolute values are used as the weight values for each feature, and these selected features are used as the screened radiomics features or the screened deep learning features;

[0148]

[0149] where S is the set of feature indices corresponding to all non-zero coefficients, ω j is the weight value of feature j, Σ k∈S |β k | is the sum of the absolute values of the regression coefficients of all selected key features. Here, the key features refer to the features corresponding to non-zero regression coefficients, and Σ k∈S |β k | is used for normalization so that the sum of the weight values ω j of all features is 1, thus enabling the comparison of the importance of different features.

[0150] It should be noted that the feature screening of the present invention includes screening out features with high representativeness and stability through statistical feature screening and machine learning feature selection.

[0151] The risk prediction model includes a feature fusion sub-module, a data balancing sub-module, and a classification modeling sub-module;

[0152] The training of the risk prediction model is carried out by the following steps:

[0153] G.1. Input the data in the training set into the feature fusion sub-module for splicing to form a complete comprehensive feature vector. The feature fusion sub-module performs splicing in a horizontal splicing manner. The training set includes data of different category objects, and each category object data includes screened imaging omics features, screened deep learning features, and screened clinical data.

[0154] G.2. The data balancing sub-module uses the SMOTE method to augment the comprehensive feature vectors corresponding to the minority categories to obtain synthetic comprehensive feature vectors.

[0155] G.3. Input the comprehensive feature vector obtained in G.1 and the synthetic comprehensive feature vector obtained in G.2 into the classification modeling sub-module for training to obtain the trained classification modeling sub-module. The classification modeling sub-module is a linear classification model.

[0156] It should be noted that the main idea behind the SMOTE method is to bridge the gap between the minority samples and the majority samples by generating synthetic samples. The following is a step-by-step description of the working principle of SMOTE:

[0157] 1. Identify minority samples: The first step involves identifying the samples in the dataset that belong to the minority class.

[0158] 2. Identify K-nearest neighbors: For each minority sample, SMOTE identifies its K-nearest neighbors in the feature space. Usually, the Euclidean distance metric is used to measure the similarity between data points.

[0159] 3. Synthetic sample generation: Once the neighbors are identified, SMOTE selects a random neighbor and calculates the difference between the feature vector of the minority sample and its selected neighbor. Then, this difference is multiplied by a random number between 0 and 1 and added to the feature vector of the minority sample. This process creates new synthetic samples that lie on the line segment between the minority sample and its selected neighbor. Repeat the process of generating synthetic samples until the desired class balance level is reached.

[0160] The classification modeling sub-module uses the log-likelihood function as the loss function to evaluate the difference between the prediction and the true label, maximizes the log-likelihood, then updates the parameters through the gradient descent update formula, and finally trains until the training ends. The classification modeling sub-module of the present invention uses a logistic regression classifier and adopts five-fold cross-validation to solve the overfitting problem. The logistic regression classifier is a commonly used linear classification model, and based on the input feature vector x and the corresponding classification label y, it trains the model parameters by maximizing the maximum likelihood estimation (MLE). The specific training end condition of the present invention is that the AUC of the validation set reaches 0.7.

[0161] The log-likelihood function is represented by Equation (7);

[0162]

[0163] Among them, w is the weight vector of the classification modeling sub-module, and b is the bias term.

[0164] The gradient descent update formula is represented by Equation (8);

[0165]

[0166] Among them, η is the learning rate, is the gradient of w and b.

[0167] S5 is specifically carried out through the following steps:

[0168] H.1. Concatenate the screened radiomics features, screened deep learning features, and screened clinical data input features obtained in S4 into the feature fusion sub-module to form a complete comprehensive feature vector;

[0169] H.2. Input the comprehensive feature vector obtained in H.1 into the trained classification modeling sub-module for prediction to obtain the immune prediction value of head and neck squamous cell carcinoma;

[0170] Immune prediction value of head and neck squamous cell carcinoma is represented by Equation (9);

[0171]

[0172] It should be noted that the definition of the classification modeling sub-module of the present invention is as follows: Assume that the comprehensive feature vector is x ∈ R p , and the output of the classification modeling sub-module is the probability of predicting to belong to the positive class (such as effective immunotherapy) where the probability can also be represented by the following formula:

[0173]

[0174] Among them, σ(z) is the sigmoid function and is the weight vector of the model, and b is the bias.

[0175] Compared with the prior art, the beneficial effects of the method for obtaining the immunological prediction value of head and neck squamous cell carcinoma based on machine learning are as follows: 1. It solves the problem that traditional clinical medicine methods rely only on low-dimensional features and are difficult to comprehensively evaluate the tumor immune response, and avoids misjudgment caused by tumor heterogeneity. 2. The method for extracting radiomics features is usually relatively fixed and ignores individual differences. However, the present invention combines deep learning methods to strive to discover more comprehensive features, thereby improving the accuracy of prediction. 3. By performing statistical feature screening and machine learning feature selection to retain the most relevant diagnostic features and calculating their weight values, it effectively solves the problem of insufficient interpretability of deep learning methods in medicine.

[0176] Example 2

[0177] An application of the method for obtaining the immunological prediction value of head and neck squamous cell carcinoma based on machine learning as in Example 1. In this example, data of 186 head and neck squamous cell carcinoma subjects from the Sun Yat-sen Memorial Hospital of Sun Yat-sen University were collected as the training set and the internal test set, and the division ratio of the training set and the internal test set was 7:3. Data of 96 head and neck squamous cell carcinoma subjects from the Sun Yat-sen University Cancer Center were also collected as the external validation set.

[0178] Among them, the voxel resampling for the resampling of the magnetic resonance image and the mask image in this example is all 1×1×5 mm. And the voxel intensity of the magnetic resonance image is normalized to the range of 0-255. Moreover, the mask image is a Mask image outlined by the operator to determine the tumor area in the magnetic resonance image, and this area is the ROI area. Then, the pixel values of the ROI area are kept unchanged, and the pixel values of other irrelevant areas are set to 0, so as to obtain a magnetic resonance image conducive to the extraction of radiomics features. To facilitate the extraction of deep learning features, it is also necessary to select the layer with the largest ROI cross-sectional area in the magnetic resonance labeled image, and combine the upper and lower layers of this layer as the three-channel input image, and use the minimum bounding rectangle to crop the magnetic resonance image, and adjust the size of the cropped image to 224×224, so as to obtain a magnetic resonance image conducive to the extraction of deep learning features.

[0179] In this example, based on the standardized environment package SERA for radiomics analysis of the Image Biomarker Standardization Initiative (IBSI), a total of 94 features were extracted, namely: 8 first-order features and 86 high-order texture features, and these features together form a radiomics feature vector.

[0180] For the screening of clinical data in this embodiment, in univariate analysis, appropriate statistical methods are selected according to the type of variables. For example, for categorical variables (such as gender, stage, etc.), chi-square test or Fisher's exact test is usually used to evaluate the association between variables and treatment outcomes; for continuous variables (such as age, BMI, etc.), t-test, Mann-Whitney U test or one-way analysis of variance (ANOVA) can be used to compare the differences between different treatment outcome groups. Through analysis, the p-value of each variable can be obtained to determine whether it is statistically significant (usually with p < 0.05 as the threshold). Based on univariate analysis, all variables with statistical significance or clinical significance are included in the multivariate analysis model. Multivariate analysis mainly uses a logistic regression model. Through stepwise regression, the independent effect of each variable after adjusting for other factors is evaluated, and at the same time, the OR value, 95% confidence interval and p-value of each variable are calculated. Finally, independent clinical factors significantly associated with treatment outcomes are screened out from the multivariate analysis. These factors are considered to have a significant impact on treatment outcomes alone after adjusting for other possible confounding variables.

[0181] In this embodiment, for radiomics features and deep learning features, the LASSO method is used for statistical feature screening, and the weight value of each feature is calculated to retain the features that are most critical for predicting the efficacy of immunotherapy. Among them, Figure 3 is the process diagram of parameter selection of radiomics features and deep learning features in the LASSO model in this embodiment, showing the change of the number of features with the increase of the regularization parameter λ. Figure 4 are the features that affect the immune efficacy response retained after screening by the LASSO model for radiomics features and deep learning features in this embodiment.

[0182] In this embodiment, the validation set is used to calculate the sensitivity and specificity corresponding to each possible threshold, and the threshold with the maximum Youden index is determined as the optimal classification threshold according to these results. Then, the optimal threshold is applied to the predicted probability of the test sample to judge the category. When the predicted probability is higher than or equal to the optimal threshold, it is judged as the positive class (effective immunotherapy); otherwise, it is judged as the negative class (ineffective immunotherapy).

[0183] Moreover, the risk prediction model of this embodiment may also include a performance evaluation sub-module. The role of the performance evaluation sub-module is to analyze and evaluate the prediction performance and clinical applicability of the model, including using prediction performance indicators to evaluate the prediction performance of the model and evaluating the clinical applicability of the model from aspects such as discrimination, calibration and clinical benefit.

[0184] Among them, the prediction performance of the model is evaluated on the test data, using indicators such as accuracy, AUC (area under the ROC curve), sensitivity, specificity, etc.

[0185] Finally, the performance evaluation sub-module evaluates the prediction results, using the area under the curve (AUC), decision curve analysis (DCA), nomogram, etc. to evaluate the effectiveness of the model and its clinical applicability.

[0186] To compare the impact of the fusion of multi-modal features on the performance of the prediction model, we compared four different-input logistic regression models, namely: the logistic regression model using only clinical information as input (C); the radiomics model using only radiomics features as input (R); the logistic regression model using clinical information + radiomics features as input (C+R); the logistic regression model using clinical information + radiomics features + deep learning features as input (C+R+ResNet18).

[0187] This example uses Figures 5 - 7 For the comparison of the ROC curves of these four models in the training set, internal test set, and external validation set. Among them, Figure 5 represents the comparison of the ROC curves of these four models in the training set; Figure 6 is the comparison of the ROC curves of these four models in the internal test set; Figure 7 is the comparison of the ROC curves of these four models in the external validation set.

[0188] Through Figures 5 to 7 It can be seen that in the training set, the AUC values of the C model, R model, C+R model, and C+R+ResNet18 model are: 0.712, 0.702, 0.717, and 0.779 respectively; in the internal test set, the AUC values of the C model, R model, C+R model, and C+R+ResNet18 model are: 0.703, 0.705, 0.715, and 0.749 respectively; in the external validation set, the AUC values of the C model, R model, C+R model, and C+R+ResNet18 model are: 0.662, 0.644, 0.663, and 0.729 respectively.

[0189] Thus, it can be seen that the integration of clinical information, radiomics features, and deep learning features in this embodiment effectively improves the overall performance of the prediction model and more accurately distinguishes the individual objects that are effective for immune prediction.

[0190] Example 3

[0191] A head and neck squamous cell carcinoma immune prediction value prediction system is carried out by using the method for obtaining the head and neck squamous cell carcinoma immune prediction value based on machine learning in Example 1 or 2.

[0192] The head and neck squamous cell carcinoma immune prediction value prediction system is provided with:

[0193] Input unit - used to collect clinical data and magnetic resonance image data of subjects with head and neck squamous cell carcinoma;

[0194] Preprocessing module - preprocesses the clinical data and magnetic resonance images to obtain preprocessed clinical data and preprocessed magnetic resonance images;

[0195] Feature extraction module - extracts radiomics features and deep learning features from the preprocessed magnetic resonance images;

[0196] Feature screening module - deletes redundant and irrelevant features in the radiomics features and deep learning features, and deletes redundant and irrelevant features in the preprocessed clinical data;

[0197] Multimodal prediction unit - combines the screened radiomics features, screened deep learning features, and screened clinical data, then predicts to obtain the immune prediction value of head and neck squamous cell carcinoma, and finally evaluates the immune prediction value of head and neck squamous cell carcinoma.

[0198] Compared with the prior art, the beneficial effects of the immune prediction value prediction system for head and neck squamous cell carcinoma of the present invention are as follows: 1. It solves the problem that traditional clinical medicine methods rely only on low-dimensional features and are difficult to comprehensively evaluate the tumor immune response, and avoids misjudgment caused by tumor heterogeneity. 2. The radiomics feature extraction method is usually relatively fixed and ignores individual differences, while the present invention attempts to discover more comprehensive features by combining deep learning methods, thereby improving the accuracy of prediction. 3. By statistically screening features and machine learning feature selection, the most relevant diagnostic features are retained and their weight values are calculated, effectively solving the problem of insufficient interpretability of deep learning methods in medicine.

[0199] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for obtaining immune prediction value of head and neck squamous cell carcinoma based on machine learning, characterized in that: The steps include: S1. Acquire a data pair of a head and neck squamous cell carcinoma object, wherein the data pair is composed of clinical data and corresponding magnetic resonance image data, wherein the magnetic resonance image data is a magnetic resonance image and a corresponding mask image, and the mask image is a marked region of interest ROI; S2, preprocessing the clinical data and magnetic resonance image obtained in S1 respectively to obtain preprocessed clinical data and preprocessed magnetic resonance image; S3, extracting radiomics features and deep learning features from the preprocessed magnetic resonance images obtained in S2; S4, respectively removing redundant and irrelevant features from the imaging omics features and deep learning features obtained in S3, and obtaining the screened imaging omics features and the screened deep learning features accordingly; and removing redundant and irrelevant features in the pre-processed clinical data obtained by S2 to obtain the screened clinical data; S5. Input the post-screening imaging features, post-screening deep learning features and post-screening clinical data obtained in S4 into the trained risk prediction model to obtain the immune prediction value of head and neck squamous cell carcinoma. The risk prediction model is trained by a training set, and the training set is composed of data pairs of different head and neck squamous cell carcinoma subjects. Each pair of data in the training set is composed of clinical data and corresponding magnetic resonance image data.

2. The method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning according to claim 1, characterized in that: The preprocessing of clinical data in S2 includes cleaning operation and standardization operation; The clinical data cleaning operation is performed according to the following steps: A1. Check all fields in the clinical data and delete records with missing values; A2. Check the numerical range of the clinical data. When the numerical value exceeds the reasonable range, define it as an abnormal value, and delete, replace or mark the abnormal value; A3. Check the duplication of the clinical data. When duplicate data entries are found, keep the best data entry and delete the other data entries. The standardization operation of the clinical data is a Min-Max standardization operation or a Z-Score standardization operation; The Min-Max normalization operation is to scale the clinical data to a fixed interval by formula (1); In formula (1), x is the original clinical data, x′ is the scaled clinical data, and x max and x min are the maximum and minimum values ​​of the variable respectively; The Z-Score standardization operation is to convert the clinical data into a standard normal distribution form after passing the formula (2); In formula (2), x is the original clinical data, x′ is the scaled clinical data, μ is the mean of the variable, and σ is the standard deviation of the variable.

3. The method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning according to claim 1, characterized in that: In said S2, the preprocessing of the magnetic resonance image includes performing a cleaning operation, a resampling operation and a normalization operation; The cleaning operation of the magnetic resonance image is performed according to the following steps: C1. Check the quality of all MRI images and delete those that do not meet the standards; C2, checking whether the dimensions of all magnetic resonance images and corresponding mask images match, and deleting the group of magnetic resonance image data when the dimensions are inconsistent; The resampling operation of the magnetic resonance image is to resample the magnetic resonance image and the corresponding mask image to a fixed voxel size; The normalization operation of the magnetic resonance image is to map the pixel intensity value of the magnetic resonance image to a specific interval; the method of the normalization operation is Min-Max normalization or Z-Score normalization; The Min-Max normalization linearly transforms the pixel intensity value of the magnetic resonance image to a fixed range through formula (3); In formula (3), I is the original pixel intensity value of the magnetic resonance image, I′ is the normalized pixel intensity value, and I max and I min are the maximum and minimum pixel values ​​of the image respectively; The Z-Score normalization subtracts the mean value of the pixel intensity from the magnetic resonance image through formula (4) and divides it by the standard deviation to convert the data into the following form: In formula (4), I is the original pixel intensity value of the magnetic resonance image, I′ is the normalized pixel intensity value, μ is the mean of the pixel intensity, and σ is the standard deviation of the pixel intensity.

4. The method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning according to claim 3, characterized in that: In S3, the imaging omics features and deep learning features in the preprocessed magnetic resonance image are extracted by a feature extraction module; The feature extraction module includes a radiomics feature extraction submodule and a deep learning feature extraction submodule; The radiomics feature extraction submodule is a radiomics analysis standardization environment software package SERA based on the image biomarker standardization program IBSI; The deep learning feature extraction submodule uses the preprocessed magnetic resonance image as a fine-tuning training set, and then updates the weights of the pre-trained 3D-ResNet18 model in MedicalNet through the fine-tuning training set to obtain the deep learning feature extraction submodule.

5. The method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning according to claim 4, characterized in that: In S4, the method for removing redundant and irrelevant features in the pre-processed clinical data obtained in S2 is as follows: E1. Perform univariate analysis, that is, perform univariate statistical test or univariate logistic regression analysis on each clinical variable to evaluate its correlation with the efficacy of immunotherapy, and proceed to E2; E2. Multivariate logistic regression analysis was performed and the final model was constructed using stepwise regression to identify variables with independent predictive value; E3. Calculate the p value and 95% confidence interval of each variable of E1 and E2 to screen out the variables of clinical data. This part of clinical variables is the clinical data after screening.

6. The method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning according to claim 5, characterized in that In S4, the method for removing redundant and irrelevant features in the radiomics features and deep learning features obtained in S3 is as follows: F1. Combine radiomics features and deep learning features into a feature matrix X, where X∈R n×p , where n is the number of samples and p is the number of features; In the feature matrix, each radiomics feature and each deep learning feature were standardized so that the mean was 0 and the standard deviation was 1; F2, construct LASSO regression model screening, where the objective function of the LASSO regression model is represented by; Among them, y i is the label value of the i-th sample, X ij is the value of the jth feature of the i-th sample, β j is the regression coefficient of the jth feature, λ is the regularization parameter and λ≥0, where λ is determined by cross-validation, is the regression coefficient vector estimated in the LASSO regression model, β is the regression coefficient, n is the number of samples, and p is the number of features; F3. In the LASSO regression model, select a non-zero regression coefficient β j Corresponding features; F4, after the absolute values ​​of the non-zero coefficients in the features selected by formula (6) are normalized, the normalized absolute values ​​are used as the weight values ​​of each feature, and these selected features are used as the post-screening imaging omics features or post-screening deep learning features; Among them, S is the feature index set corresponding to all non-zero coefficients, ω j is the weight value of feature j, ∑ k∈S |β k | is the sum of the absolute values ​​of the regression coefficients of all selected key features.

7. The method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning according to claim 6, characterized in that: The risk prediction model includes a feature fusion submodule, a data balancing submodule and a classification modeling submodule; The training of the risk prediction model is carried out by the following steps: G.

1. Input the data in the training set into the feature fusion submodule for splicing to form a complete comprehensive feature vector; the feature fusion submodule adopts a horizontal splicing method for splicing; the training set includes object data of different categories, and each category of object data includes post-screening imaging omics features, post-screening deep learning features and post-screening clinical data; G.2, the data balancing submodule uses the SMOTE method to expand the comprehensive feature vector corresponding to the minority category to obtain a synthetic comprehensive feature vector; G.

3. Input the comprehensive feature vector obtained in G.1 and the synthetic comprehensive feature vector obtained in G.2 into the classification modeling submodule for training to obtain a trained classification modeling submodule; the classification modeling submodule is a linear classification model.

8. The method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning according to claim 7, characterized in that: The classification modeling submodule training uses the log-likelihood function as a loss function to evaluate the difference between the prediction and the true label, maximizes the log-likelihood, and then updates the parameters through the gradient descent update formula, and finally trains until the training is completed; The log-likelihood function is expressed by formula (7); Among them, w is the weight vector of the classification modeling submodule, and b is the bias term; The gradient descent update formula is expressed by formula (8); Where η is the learning rate, is the gradient of w and b.

9. The method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning according to claim 8, characterized in that: The S5 is specifically performed through the following steps: H.

1. Splice the post-screening radiomics features, post-screening deep learning features, and post-screening clinical data input feature fusion submodules obtained in S4 into a complete comprehensive feature vector; H.

2. Input the comprehensive feature vector obtained in H.1 into the post-training classification modeling submodule for prediction to obtain the immune prediction value of head and neck squamous cell carcinoma; The immunopredictive value of head and neck squamous cell carcinoma It is expressed by formula (9); 10. A prediction system for immune prediction value of head and neck squamous cell carcinoma, characterized in that: The method for obtaining the immune prediction value of head and neck squamous cell carcinoma based on machine learning as described in any one of claims 1 to 9 is used; The settings are: Input unit - used to collect clinical data and magnetic resonance image data of head and neck squamous cell carcinoma subjects; Preprocessing module: preprocessing clinical data and magnetic resonance images to obtain preprocessed clinical data and preprocessed magnetic resonance images; Feature extraction module - extracts radiomics features and deep learning features from preprocessed MRI images; Feature screening module - removes redundant and irrelevant features in radiomics features and deep learning features, as well as redundant and irrelevant features in preprocessed clinical data; Multimodal prediction unit - merges the post-screening radiomics features, post-screening deep learning features and post-screening clinical data, and then predicts the immune prediction value of head and neck squamous cell carcinoma, and finally evaluates the immune prediction value of head and neck squamous cell carcinoma.