Breast cancer radiodermatitis prediction system construction method and prediction system

By constructing a prediction system for radiodermatitis of breast cancer, combining target area omists, nearby histoothiologies and clinical characteristics, a mixed prediction model is established, which solves the problems of poor consistency and individualized treatment of radiodermatitis prediction in the prior art, and achieves the formulation of individualized treatment plans for breast cancer patients and improves radiotherapy tolerance.

CN120356644APending Publication Date: 2025-07-22HANGZHOU CANCER HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311676770.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art lacks reliable and rapid methods to evaluate the risk of radiodermatitis in breast cancer patients, making it difficult to implement individualized treatment plans, and the predictive model of radiodermatitis is poorly consistent in different research centers and lacks practical operability.

Method used

A prediction system for radiodermatitis of breast cancer is constructed, and by establishing the original sample set, deleting invalid features, constructing a subset of feature of target area, adjacent histomics, and clinical features, and fusing the final sample set, a hybrid prediction model is constructed based on artificial neuron networks to realize quantitative analysis of the patient's surface skin and skin-related tissues.

Benefits of technology

The individualized prediction of radiodermatitis in breast cancer patients is achieved, and the decision-making guidance of individualized radiotherapy plans is provided, which improves the patient's radiotherapy tolerance and treatment effect, and reduces the probability of radiodermatitis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356644A_ABST
    Figure CN120356644A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of prediction system construction, in particular to a breast cancer radiodermatitis prediction system construction method and a prediction system. The method comprises the following steps: S1, establishing an original sample set; s2, deleting invalid features; s3, obtaining an optimal target area omics feature subset; s4, obtaining an optimal target region adjacent histomics feature subset; s5, obtaining an optimal clinical feature subset; s6, fusing the optimal target region omics feature subset, the optimal target region adjacent tissue omics feature subset and the optimal clinical feature subset to obtain a final sample set, and constructing a hybrid prediction model based on the final sample set; the constructed hybrid prediction model is the breast cancer radiodermatitis prediction system. The prediction system is constructed based on the method. According to the method, mining of related features can be well realized, and a prediction system can be well constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of prediction system construction, and specifically, to a method for constructing a breast cancer radiation dermatitis prediction system and a prediction system. Background Art

[0002] For breast cancer patients with stage II or above, postoperative adjuvant radiotherapy has become the standard treatment method. The target areas of breast cancer radiotherapy mainly include the breast, chest wall, axilla, supraclavicular and internal mammary lymph nodes, etc. According to statistics, about 23% of patients develop irreversible radiation dermatitis (RD) of varying degrees during or after radiotherapy, which has become the main toxic and side effect seriously affecting the treatment tolerance and quality of life of patients. How to predict and intervene early in patients who may develop radiation dermatitis during breast cancer radiotherapy, reduce the incidence of radiation dermatitis, and improve the radiotherapy tolerance of patients has become a major problem faced by breast cancer radiotherapy.

[0003] Regarding the research and prediction of factors related to radiation dermatitis caused by breast cancer radiotherapy, at home and abroad, it is mainly achieved by monitoring the temperature changes on the surface of the breast during radiotherapy. The Radiation Oncology Center of the University of the French West Indies in Martinique established the correlation between the probability of radiation dermatitis and temperature changes, and found that compared with patients without or with mild dermatitis, the differences in the maximum, minimum, and average temperatures between the two breasts of patients with grade 2 or higher radiation dermatitis are higher. In addition to temperature-related factors, the surgical method before radiotherapy for breast cancer radiotherapy is also related to radiation dermatitis. Foreign studies have shown that compared with patients who have undergone total mastectomy, the incidence of moderate and severe skin dermatitis in patients who receive radiotherapy after breast-conserving surgery is higher. However, the research results of Zhang Qian in China are just the opposite. In the observation of 80 breast cancer patients after radiotherapy, none of the patients after breast-conserving surgery developed grade 2 or higher acute radiation dermatitis, and all cases occurred in patients who received modified radical mastectomy. In addition, some studies have shown that the incidence of radiation dermatitis is highly correlated with the radiation dose of the breast. Acute radiation dermatitis reactions such as epithelial exfoliation and ulcers can occur when the skin is irradiated with 20-40 Gy. Although most acute radiation dermatitis will resolve in 2-3 weeks, late toxicity usually persists and has a continuous negative impact on the quality of life. However, the consistency of these research results in different research centers is poor, and at the same time, there is a lack of practical operability for individualized quantitative screening of high-risk patients. Therefore, if the risk of radiation dermatitis in patients can be reliably and quickly evaluated, more targeted individualized radiotherapy can be carried out; secondly, if early prediction of radiation dermatitis can be achieved, targeted intervention can be carried out before the start of radiotherapy, improving the radiotherapy tolerance of patients and ensuring the curative effect.

[0004] In the past, tumors have always been the research objects. By taking advantage of the significant differences in treatment responses among different patients caused by the heterogeneity existing between different tumors and within the same tumor in terms of time and space, the exploration of the internal causes of treatment heterogeneity and solutions for cancer individuals has become a research hotspot in imaging radiomics. Screening out potential patients who can benefit from treatment through imaging radiomics technology and avoiding ineffective treatment is also an urgent need for the development of precision medicine. In clinical practice, we have found that for patients with similar ages, highly similar radiotherapy dose distributions, and the same treatment sites, the probability and degree of radiation dermatitis vary completely. This suggests that in addition to being related to the above external factors, radiation dermatitis is also related to certain internal characteristics of the skin and surrounding tissues at the radiotherapy site of the patient. There is also a certain degree of heterogeneity in normal tissues such as the skin among patients. If the related technologies of imaging radiomics are used to quantitatively analyze the individual skin tissues of breast cancer patients and a quantitative, stable, and reliable prediction model is established by combining external radiotherapy-related factors, it may be possible to judge the probability and severity of radiation dermatitis in a patient before radiotherapy begins, providing a scientific basis for individualized precision treatment.

[0005] Based on this, the present invention provides a method for constructing a prediction system for breast cancer radiation dermatitis and a prediction system. Summary of the Invention

[0006] The present invention provides a method for constructing a prediction system for breast cancer radiation dermatitis, which can overcome certain defects of the prior art.

[0007] According to a method for constructing a prediction system for breast cancer radiation dermatitis of the present invention, it has the following steps:

[0008] S1. Establish an original sample set. The original sample set has multiple samples, and each sample has a target region omics feature set, a target region adjacent tissue omics feature set, a clinical feature set, and a label. The target region omics feature set has all the omics features of multiple target regions, and the target region adjacent tissue omics feature set has all the omics features of multiple target region adjacent regions;

[0009] S2. Process the target region omics feature set and the target region adjacent tissue omics feature set to delete invalid features, and eliminate irrelevant features in the target region omics feature set and the target region adjacent tissue omics feature set based on feature engineering; process the clinical feature set to delete invalid features;

[0010] S3. Construct a feature subset of the target region omics feature set, establish a first prediction model based on the target region omics features, and obtain the optimal target region omics feature subset through the first prediction model;

[0011] S4. Construct a feature subset of the omics features of the tissues adjacent to the target area, establish a second prediction model based on the omics features of the tissues adjacent to the target area, and obtain the optimal omics feature subset of the tissues adjacent to the target area through the second prediction model;

[0012] S5. Construct a feature subset of the clinical features, establish a third prediction model based on the clinical features, and obtain the optimal clinical feature subset through the third prediction model;

[0013] S6. Fuse the optimal target area omics feature subset, the optimal omics feature subset of the tissues adjacent to the target area, and the optimal clinical feature subset to obtain the final sample set, and construct a hybrid prediction model based on the final sample set; the constructed hybrid prediction model is the breast cancer radiation dermatitis prediction system.

[0014] Through the above method, it is possible to achieve non-invasive and reproducible quantitative analysis of all tissues of the body surface skin and the planning target area highly related to the skin of breast cancer patients, and through combining the clinical data information of patients, it is possible to better achieve in-depth exploration of relevant features. By establishing the optimal prediction parameters or the optimal combination of prediction parameters, it is possible to establish a prediction system. Thus, it can provide decision-making guidance for the formulation of individualized radiotherapy plans and treatment screening stratification plans for breast cancer patients; at the same time, it will also provide new concepts and research ideas for other tumors dominated by radiotherapy and chemotherapy.

[0015] Preferably, in step S2, the processing of the target area omics feature set and the omics feature set of the tissues adjacent to the target area to delete invalid features includes deleting null value features and unbalanced features in the target area omics feature set and the omics feature set of the tissues adjacent to the target area. Therefore, it is possible to better achieve feature screening.

[0016] Preferably, in step S2, the elimination of irrelevant features in the target area omics feature set and the omics feature set of the tissues adjacent to the target area based on feature engineering includes the Mann-Whitney U (MWU) test step, the correlation detection step, the LASSO dimensionality reduction processing step, the logistic regression analysis step, and the multiple linearity test step of the variance inflation factor VIF. Based on this, it is possible to better achieve the deletion of omics features irrelevant to the prediction results.

[0017] Preferably, in the MWU test step, extract the values of each omics feature in all samples with radiation dermatitis to construct the first omics feature sequence corresponding to the omics feature, and at the same time extract the values of each omics feature in all samples without radiation dermatitis to construct the second omics feature sequence corresponding to the omics feature, and eliminate the omics features corresponding to the first omics feature sequence and the second omics feature sequence with similar data distribution characteristics. Thus, it is possible to better achieve the elimination of useless omics features.

[0018] Preferably, in the correlation detection step, the values of each omics feature in all samples are extracted to construct a third omics feature sequence corresponding to the omics feature, and the omics feature corresponding to the variance of the third omics feature sequence approaching 0 is removed; and if the Pearson correlation coefficient of the third omics feature sequences of any two corresponding omics features is not less than 0.9, the omics feature corresponding to the smaller variance of the third omics feature sequence among the any two corresponding omics features is removed. The correlation detection mainly detects the variance of a single omics feature and the Pearson correlation coefficient of any two omics features, so as to better remove the omics features that contribute less to the prediction result and the redundant omics features.

[0019] Preferably, in the logistic regression analysis step, it is implemented by a binary logistic regression method.

[0020] Preferably, the clinical feature set is processed to delete invalid features, the t-test and MUW test are performed on the clinical feature sequence, and the clinical features with P-values not exceeding 0.5 in both tests are retained. Thus, the screening of useful variables in the clinical feature set can be better achieved.

[0021] Preferably, in step S3, step S4 and step S5, the construction of the feature subsets of the target region omics feature set, the target region adjacent tissue omics feature set and the clinical feature set is completed based on the genetic algorithm, and the optimization of each feature subset is realized based on the wrapper algorithm.

[0022] Preferably, in step S3, step S4 and step S5, the first prediction model, the second prediction model and the third prediction model are constructed based on the artificial neural network. Therefore, it has better universality.

[0023] In addition, the present invention also provides a breast cancer radiation dermatitis prediction system, which is constructed based on any one of the above breast cancer radiation dermatitis prediction system construction methods. Description of the Drawings

[0024] Figure 1 It is a schematic diagram of a method for constructing a breast cancer radiation dermatitis prediction system in Example 1.

[0025] Figure 2 It is a schematic diagram of the correlation between the index "PTV100PD.F6.IH_GaussFit1GaussMean" in Example 1 and the probability of occurrence of grade II or above radiation dermatitis.

[0026] Figure 3Schematic diagram of the correlation of the index "PTV105PD.F4.ID_LocalEntropyMax" in Example 1 with the probability of occurrence of radioactive dermatitis above grade II.

[0027] Figure 4 Schematic diagram of the correlation of the index "PTV108PD.F1.GOH0.975Quantile" in Example 1 with the probability of occurrence of radioactive dermatitis above grade II.

[0028] Figure 5 Schematic diagram of the correlation of the index "PTV108PD.F2._GLCM2590.7_IV" in Example 1 with the probability of occurrence of radioactive dermatitis above grade II.

[0029] Figure 6 Schematic diagram of the correlation of the index "PTV108PD.F8.ShapeNumberOfObjects" in Example 1 with the probability of occurrence of radioactive dermatitis above grade II.

[0030] Figure 7 Schematic diagram of the correlation of the index "SKIN30Gy.F1.GOH_MAD" in Example 1 with the probability of occurrence of radioactive dermatitis above grade II.

[0031] Figure 8 Schematic diagram of the correlation of the index "SKIN30Gy.F2._GLCM25225.4Contrast" in Example 1 with the probability of occurrence of radioactive dermatitis above grade II.

[0032] Figure 9 Schematic diagram of the correlation of the index "SKIN30Gy.F4.ID_LocalRangeMax" in Example 1 with the probability of occurrence of radioactive dermatitis above grade II.

[0033] Figure 10 Schematic diagram of the correlation of the index "SKIN30Gy.F6.IHGaussFit1GaussStd" in Example 1 with the probability of occurrence of radioactive dermatitis above grade II.

[0034] Figure 11 Schematic diagram of the correlation of the index "Quadrant.positions." in Example 1 with the probability of occurrence of radioactive dermatitis above grade II.

[0035] Figure 12Schematic diagram of the correlation between the indicator "T.Stage" in Example 1 and the probability of occurrence of radioactive dermatitis above grade II.

[0036] Figure 13 Schematic diagram of the correlation between the indicator "Hormone.therapy.yes.no." in Example 1 and the probability of occurrence of radioactive dermatitis above grade II. Detailed implementation manners

[0037] To further understand the content of the present invention, the present invention will be described in detail in combination with embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.

[0038] Example 1

[0039] As seen in Figure 1 , this embodiment provides a method for constructing a breast cancer radioactive dermatitis prediction system, which has the following steps:

[0040] S1. Establish an original sample set. The original sample set has multiple samples, and each sample has a target region omics feature set, a target region adjacent tissue omics feature set, a clinical feature set, and a label; the target region omics feature set has all the omics features of multiple target regions, and the target region adjacent tissue omics feature set has all the omics features of multiple target region adjacent regions;

[0041] S2. Process the target region omics feature set and the target region adjacent tissue omics feature set to delete invalid features, and eliminate irrelevant features in the target region omics feature set and the target region adjacent tissue omics feature set based on feature engineering; process the clinical feature set to delete invalid features;

[0042] S3. Construct a feature subset of the target region omics feature set, establish a first prediction model based on the target region omics features, and obtain the best target region omics feature subset through the first prediction model;

[0043] S4. Construct a feature subset of the target region adjacent tissue omics feature set, establish a second prediction model based on the target region adjacent tissue omics features, and obtain the best target region adjacent tissue omics feature subset through the second prediction model;

[0044] S5. Construct a feature subset of the clinical feature set, establish a third prediction model based on the clinical features, and obtain the best clinical feature subset through the third prediction model;

[0045] S6. Integrate the best target region omics feature subset, the best target region adjacent tissue omics feature subset, and the best clinical feature subset to obtain a final sample set, and construct a hybrid prediction model based on the final sample set; the constructed hybrid prediction model is the breast cancer radioactive dermatitis prediction system.

[0046] Through the above method, non-invasive and reproducible quantitative analysis can be achieved for all tissues of the body surface skin of breast cancer patients and the planning target volume highly related to the skin. Moreover, by combining the clinical data information of the patients, in-depth exploration of relevant features can be better achieved. By establishing the best prediction parameters or the best combination of prediction parameters, a prediction system can be established. Thereby, it can provide decision-making guidance for formulating individualized radiotherapy plans and treatment screening and stratification plans for breast cancer patients; at the same time, it will also provide new concepts and research ideas for other tumors mainly treated with radiotherapy and chemotherapy.

[0047] In step S1 of this embodiment, 90 breast cancer patients who received radiotherapy at the patentee from April 2019 to April 2020 were selected. Subsequently, the original 4DCT image data and clinical information of each patient were collected. For the original 4DCT image data, specific regions of interest were delineated, and then quantitative analysis and extraction of radiomics features were performed on the delineated region images, so as to obtain the target region radiomics feature set and the radiomics feature set of tissues adjacent to the target region for each patient.

[0048] It can be immediately seen that the target region radiomics feature set, the radiomics feature set of tissues adjacent to the target region, the clinical feature set, and the label of each patient constitute a sample. Among them, if a patient has radiation dermatitis, the corresponding label is "yes" and can be identified by, for example, "1", and if a patient does not have radiation dermatitis, the corresponding label is "no" and can be identified by, for example, "0".

[0049] In this embodiment, for the target region radiomics feature set, the radiomics features in 3 target regions of "R-PTV_100Pd", "R-PTV_105Pd", and "R-PTV_108Pd" are selected. For the radiomics feature set of tissues adjacent to the target region, the radiomics features in 3 target adjacent regions of "R-SKIN_20Gy", "R-SKIN_30Gy", and "R-SKIN_40Gy" are selected; for the clinical feature set, the physical characteristics of the patients (such as age, BMI, etc.) and clinical characteristics such as drug dosage are selected. In this embodiment, 38 variables are selected to construct the clinical feature set.

[0050] As follows, the originally obtained target region radiomics feature set and the radiomics feature set of tissues adjacent to the target region are respectively:

[0051] data.radiomics.patient.(1)R-PTV_100Pd;

[0052] data.radiomics.patient.(2)R-PTV_105Pd;

[0053] data.radiomics.patient.(3)R-PTV_108Pd;

[0054] data.radiomics.patient.(4)R-SKIN_20Gy;

[0055] data.radiomics.patient.(5)R-SKIN_30Gy;

[0056] data.radiomics.patient.(6)R-SKIN_40Gy.

[0057] In step S2 of this embodiment, the processing of the target region omics feature set and the target region adjacent tissue omics feature set to delete invalid features includes deleting null value features and unbalanced features in the target region omics feature set and the target region adjacent tissue omics feature set.

[0058] As follows, the target region omics feature set and the target region adjacent tissue omics feature set after the above processing are respectively:

[0059] PTV_100Pd - 812;

[0060] PTV_105Pd - 789;

[0061] PTV_108Pd - 674;

[0062] SKIN_20Gy - 684;

[0063] SKIN_30Gy - 657;

[0064] SKIN_40Gy - 664.

[0065] The above indicates that the "R-PTV_100Pd", "R-PTV_105Pd", "R-PTV_108Pd", "R-SKIN_20Gy", "R-SKIN_30Gy" and "R-SKIN_40Gy" regions respectively retain 812, 789, 674, 684, 657 and 664 omics features.

[0066] In step S2 of this embodiment, the elimination of irrelevant features in the target region omics feature set and the target region adjacent tissue omics feature set based on feature engineering includes the MWU test step, the correlation detection step, the LASSO dimensionality reduction processing step, the logistic regression analysis step, and the multiple linearity test step of the variance inflation factor VIF. Based on this, the deletion of omics features irrelevant to the prediction result can be better achieved.

[0067] Among them, in the MWU test step, the values of each omics feature in all samples with radiation dermatitis are extracted to construct the first omics feature sequence of the corresponding omics feature. At the same time, the values of each omics feature in all samples without radiation dermatitis are extracted to construct the second omics feature sequence of the corresponding omics feature. The corresponding omics features with similar data distribution characteristics in the first omics feature sequence and the second omics feature sequence are removed. Thus, the removal of useless omics features can be better achieved.

[0068] Among them, the MWU test is the Mann-Whitney U test, which is a mature existing test method and will not be elaborated in this embodiment. The omics features described in this embodiment include the omics features of each target area and each area adjacent to the target area in each target area omics feature set and each target area adjacent tissue omics feature set.

[0069] As follows, after the removal in the MWU test step, the omics features of the "R-PTV_100Pd", "R-PTV_105Pd", "R-PTV_108Pd", "R-SKIN_20Gy", "R-SKIN_30Gy", and "R-SKIN_40Gy" regions are respectively left with "95", "159", "66", "329", "247", and "153".

[0070] PTV_100Pd - 812 → 95;

[0071] PTV_105Pd - 789 → 159;

[0072] PTV_108Pd - 674 → 66;

[0073] SKIN_20Gy - 684 → 329;

[0074] SKIN_30Gy - 657 → 247;

[0075] SKIN_40Gy - 664 → 153.

[0076] Among them, in the correlation detection step, the values of each omics feature in all samples are extracted to construct the third omics feature sequence of the corresponding omics feature. The corresponding omics features with variances of the third omics feature sequence close to 0 (such as ≤0.05) are removed; and if the Pearson correlation coefficients of the third omics feature sequences of any two corresponding omics features are not less than 0.9, the corresponding omics feature with a smaller variance in the third omics feature sequence of the any two corresponding omics features is removed. The correlation detection mainly detects the variance of a single omics feature and the Pearson correlation coefficients of any two omics features, so that the removal of omics features with low contribution to the prediction result and redundant omics features can be better achieved.

[0077] As shown below, after variance-based elimination, the omics features in the regions of "R-PTV_100Pd", "R-PTV_105Pd", "R-PTV_108Pd", "R-SKIN_20Gy", "R-SKIN_30Gy", and "R-SKIN_40Gy" remain "90", "153", "51", "322", "241", and "151" respectively; after Pearson correlation coefficient-based elimination, the omics features in the corresponding regions remain "41", "28", "32", "124", "37", and "45" respectively.

[0078] PTV_100Pd - 812 → 95 → 90 → 41;

[0079] PTV_105Pd - 789 → 159 → 153 → 28;

[0080] PTV_108Pd - 674 → 66 → 51 → 32;

[0081] SKIN_20Gy - 684 → 329 → 322 → 124;

[0082] SKIN_30Gy - 657 → 247 → 241 → 37;

[0083] SKIN_40Gy - 664 → 153 → 151 → 45.

[0084] As shown below, after the above-mentioned LASSO dimensionality reduction processing step, the omics features in the corresponding regions remain "23", "12", "8", "43", "14", and "29" respectively.

[0085] PTV_100Pd - 812 → 95 → 90 → 41 → 23;

[0086] PTV_105Pd - 789 → 159 → 153 → 28 → 12;

[0087] PTV_108Pd - 674 → 66 → 51 → 32 → 8;

[0088] SKIN_20Gy - 684 → 329 → 322 → 124 → 43;

[0089] SKIN_30Gy - 657 → 247 → 241 → 37 → 14;

[0090] SKIN_40Gy - 664 → 153 → 151 → 45 → 29.

[0091] Among them, in the above-mentioned logistic regression analysis step, it is implemented by the binary logistic regression method.

[0092] As follows, after logistic regression analysis, the omics features in the corresponding regions are respectively left with "9", "5", "6", "17", "9" and "5".

[0093] Among them, in the multiple linearity test step of the variance inflation factor VIF, a multiple linearity test of the variance inflation factor VIF is performed on the feature sequences of the omics features obtained through the logistic regression analysis step, and the corresponding omics features with VIF > 10 are removed.

[0094] As follows, for the target region "PTV_100Pd", the omics feature "F2._GLCM25270.4_Corr" is removed.

[0095] Omics feature VIF Remark F2._GLCM25225.1_Corr 8.175 Retain F2._GLCM25270.4_Corr 20.364 Exclude F2._GLCM25270.7_Corr 8.608 Retain F4.ID_LocalStdMedian 8.945 Retain F4.ID_10Percentile 5.046 Retain F4.ID_0.025Quantile 5.471 Retain F4.ID_Range 1.407 Retain F6.IHGaussFit2GaussMean 1.311 Retain F8.ShapeMax3DDiameter 1.551 Retain

[0096] As follows, for the target region "PTV_105Pd", all omics features are retained.

[0097] Omics feature VIF Remark F2._GLCM25.333.7_Corr 1.208 Retain F4.ID_LocalEntropyMax 1.080 Retain F4.ID_LocalEntropyMin 1.195 Retain F7.NID25Coarseness 1.145 Retain F8.ShapeMeanBreadth 1.123 Retain

[0098] As follows, for the target region "PTV_108Pd", all omics features are retained.

[0099] Omics feature VIF Remark F1.GOH0.975Quantile 1.277 Retain F2._GLCM25180.1Dissimilarity 1.831 Retain F2._GLCM2590.7_IV 1.301 Retain F2._GLCM25225.1_MP 1.703 Retain F8.ShapeNumberOfObjects 1.292 Retain F8.ShapeSphericalDisproportion 1.513 Retain

[0100] As follows, for the region adjacent to the target region "SKIN_20Gy", the omics features F1.GOH..0.5Quantile, F1.GOHSkewness, F5.IH0.975Quantile and F8.ShapeConvex are removed.

[0101] Omics feature VIF Remark F1.GOH..0.5Quantile 11.650 Exclude F1.GOHSkewness 10.290 Exclude F2._GLCM25270.7_AC 5.360 Retain F2._GLCM25180.1Contrast 6.014 Retain F2._GLCM25225.4Contrast 3.749 Retain F2._GLCM25180.4Energy 4.063 Retain F2._GLCM25270.1IMC1 7.703 Retain F2._GLCM250.1IMC2 7.683 Retain F2._GLCM25180.4SumVariance 4.358 Retain F2._GLCM25180.7Variance 8.875 Retain F2._GLCM25270.7Variance 6.202 Retain F4.ID_GlobalMin 1.102 Retain F4.ID_LocalEntropyStd 9.470 Retain F5.IH0.975Quantile 10.854 Exclude F8.ShapeConvex 11.225 Exclude F8.ShapeConvexHullVolume3D 1.827 Retain F8.ShapeMeanBreadth 7.401 Retain

[0102] As follows, for the region adjacent to the target region "SKIN_30Gy", all omics features are retained.

[0103] Omics feature VIF Remark F1.GOH_MAD 1.351 Retain F2._GLCM25225.4Contrast 1.411 Retain F2._GLCM25315.1Contrast 2.793 Retain F2._GLCM25.333.1_IV 1.437 Retain F4.ID_LocalRangeMax 1.286 Reserved F6.IHGaussFit1GaussStd 1.255 Reserved F8.ShapeMax3DDiameter 1.075 Reserved F8.ShapeOrientation 2.273 Reserved

[0104] As follows, for the region adjacent to the target region "SKIN_40Gy", all omics features are retained.

[0105] Omics feature VIF Remarks F2._GLCM25270.7Dissimilarity 1.359 Reserved F2._GLCM25225.7_IV 1.047 Reserved F3.GLRLM_25...0_HGLRE 9.805 Reserved F5.IH.5Percentile 1.834 Reserved F5.IH90Percentile 8.489 Reserved

[0106] Among them, the clinical feature set is processed to delete invalid features, a t-test and a MUW test are performed on the clinical feature sequences, and the clinical features with P values not exceeding 0.5 in both tests are retained. Thus, the screening of useful variables in the clinical feature set can be better achieved.

[0107] As follows are the results of t-test and MUW test for each clinical feature in the clinical feature set.

[0108]

[0109] After elimination, the finally retained clinical features are as follows.

[0110]

[0111]

[0112] Among them, in steps S3, S4 and S5, the construction of feature subsets of the target region omics feature set, the target region adjacent tissue omics feature set and the clinical feature set is completed based on the genetic algorithm, and the optimization of each feature subset is realized based on the wrapper algorithm. Thus, the construction of feature subsets can be better achieved.

[0113] Among them, in steps S3, S4 and S5, the first prediction model, the second prediction model and the third prediction model are constructed based on the artificial neural network. Therefore, it can have better universality.

[0114] In S6, the decision tree algorithm, the random forest algorithm and the support vector machine algorithm are respectively used to construct the hybrid prediction model.

[0115] In S6, the feature subset adopted by the decision tree algorithm includes the following features:

[0116] “SKIN30.F6.IHGaussFit1GaussStd”;

[0117] “SKIN30.F8.ShapeMax3DDiameter”;

[0118] “PTV108.F1.GOH0.975Quantile”;

[0119] “PTV105.F4.ID_LocalEntropyMax”;

[0120] “PTV108.F2._GLCM25180.1Dissimilarity”;

[0121] “SKIN20.F2._GLCM25225.4Contrast”;

[0122] “SKIN30.F2._GLCM25225.4Contrast”;

[0123] “SKIN20.F8.ShapeMeanBreadth”;

[0124] “PTV105.F8.ShapeMeanBreadth”;

[0125] “PTV100.F2._GLCM25270.7_Corr”;

[0126] “PTV108.F8.ShapeNumberOfObjects”;

[0127] “PTV100.F4.ID_LocalStdMedian”;

[0128] “PTV100.F8.ShapeMax3DDiameter”;

[0129] “PTV100.F6.IHGaussFit1GaussMean”;

[0130] “PTV105.F2._GLCM25.333.7_Corr”;

[0131] “SKIN30.F1.GOH_MAD”;

[0132] “PTV100.F4.ID_Range”;

[0133] “SKIN30.F4.ID_LocalRangeMax”;

[0134] “PTV108.F2._GLCM2590.7_IV”;

[0135] “SKIN20.F8.ShapeConvexHullVolume3D”;

[0136] “Laterality”;

[0137] “Quadrant.positions.”;

[0138] “Histologic.type”;

[0139] “Overall..Stage”;

[0140] “T.Stage”;

[0141] “PR”;

[0142] “Hormone.therapy.yes.no.”;

[0143] "RT.method";

[0144] "Fractionation.regimen.Gy.fx.";

[0145] "EQD2_all";

[0146] "lotion.application.yes.no.";

[0147] "Skin.condiction.before.RT";

[0148] "SKIN_V20";

[0149] "SKIN_V30".

[0150] In S6, the feature subsets adopted by the random forest algorithm include the following features:

[0151] "SKIN20.F4.ID_GlobalMin";

[0152] "SKIN30.F6.IHGaussFit1GaussStd";

[0153] "PTV105.F4.ID_LocalEntropyMax";

[0154] "SKIN20.F2._GLCM25225.4Contrast";

[0155] "SKIN30.F8.ShapeMax3DDiameter";

[0156] "SKIN30.F2._GLCM25225.4Contrast";

[0157] "PTV108.F1.GOH0.975Quantile";

[0158] "PTV108.F2._GLCM25180.1Dissimilarity";

[0159] "PTV100.F4.ID_LocalStdMedian";

[0160] "PTV105.F8.ShapeMeanBreadth";

[0161] "Laterality";

[0162] "Quadrant.positions.";

[0163] "Histologic.type";

[0164] "Overall..Stage";

[0165] "T.Stage";

[0166] "PR";

[0167] "Hormone.therapy.yes.no.";

[0168] "RT.method";

[0169] "Fractionation.regimen.Gy.fx.";

[0170] "EQD2_all";

[0171] "lotion.application.yes.no.";

[0172] "Skin.condiction.before.RT";

[0173] "SKIN_V20";

[0174] "SKIN_V30".

[0175] In S6, the feature subset adopted by the support vector machine algorithm includes the following features:

[0176] "PTV100.F2._GLCM25225.1_Corr";

[0177] "PTV100.F2._GLCM25270.7_Corr";

[0178] "PTV100.F4.ID_0.025Quantile";

[0179] "PTV100.F8.ShapeMax3DDiameter";

[0180] "PTV105.F2._GLCM25.333.7_Corr";

[0181] "PTV105.F4.ID_LocalEntropyMax";

[0182] "PTV105.F4.ID_LocalEntropyMin";

[0183] “PTV105.F7.NID25Coarseness”;

[0184] “PTV108.F1.GOH0.975Quantile”;

[0185] “PTV108.F2._GLCM25180.1Dissimilarity”;

[0186] “PTV108.F2._GLCM2590.7_IV”;

[0187] “PTV108.F2._GLCM25225.1_MP”;

[0188] “PTV108.F8.ShapeSphericalDisproportion”;

[0189] “SKIN20.F2._GLCM25180.4Energy”;

[0190] “SKIN20.F2._GLCM25270.1IMC1”;

[0191] “SKIN20.F2._GLCM250.1IMC2”;

[0192] “SKIN20.F4.ID_GlobalMin”;

[0193] “SKIN20.F4.ID_LocalEntropyStd”;

[0194] “SKIN30.F2._GLCM25315.1Contrast”;

[0195] “SKIN30.F2._GLCM25.333.1_IV”;

[0196] “SKIN30.F8.ShapeMax3DDiameter”;

[0197] “SKIN40.F2._GLCM25270.7Dissimilarity”;

[0198] “SKIN40.F2._GLCM25225.7_IV”;

[0199] “Laterality”;

[0200] “Quadrant.positions.”;

[0201] “Histologic.type”;

[0202] “Overall..Stage”;

[0203] “T.Stage”;

[0204] “PR”;

[0205] “Hormone.therapy.yes.no.”;

[0206] “RT.method”;

[0207] “Fractionation.regimen.Gy.fx.”;

[0208] “EQD2_all”;

[0209] “lotion.application.yes.no.”;

[0210] “Skin.condiction.before.RT”;

[0211] “SKIN_V20”;

[0212] “SKIN_V30”.

[0213] Subsequently, prospective studies were conducted on the decision tree algorithm, random forest algorithm, and support vector machine algorithm, respectively, to determine the feature sets adopted by the final sample set, which are as follows:

[0214] “PTV100PD.F6.IH_GaussFit1GaussMean”

[0215] “PTV105PD.F4.ID_LocalEntropyMax”

[0216] “PTV108PD.F1.GOH0.975Quantile”;

[0217] “PTV108PD.F2._GLCM2590.7_IV”

[0218] “PTV108PD.F8.ShapeNumberOfObjects”;

[0219] “SKIN30Gy.F1.GOH_MAD”;

[0220] “SKIN30Gy.F2._GLCM25225.4Contrast”;

[0221] “SKIN30Gy.F4.ID_LocalRangeMax”;

[0222] “SKIN30Gy.F6.IHGaussFit1GaussStd”;

[0223] “Quadrant.positions.”

[0224] “T.Stage”;

[0225] “Hormone.therapy.yes.no.”

[0226] seen in Figure 2-13 are the correlations of the above 12 indicators with the probability of occurrence of radioactive dermatitis above grade II, respectively.

[0227] In addition, this embodiment also provides a radioactive dermatitis prediction system for breast cancer, which is constructed based on the method of this embodiment.

[0228] It is easy to understand that those skilled in the art can combine, split, and reorganize the embodiments of this application based on one or several embodiments provided by this application to obtain other embodiments, and these embodiments do not exceed the protection scope of this application.

[0229] The above schematically describes the present invention and its implementation manners. This description is not restrictive, and what is shown in the embodiments is only part of the implementation manners of the present invention. The actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. A method for constructing a prediction system for breast cancer radiation dermatitis, comprising the following steps: S1. Establish an original sample set, which has multiple samples, and each sample has a target area omics feature set, a target area adjacent tissue omics feature set, a clinical feature set, and a label; the target area omics feature set has all the omics features of multiple target areas, and the target area adjacent tissue omics feature set has all the omics features of multiple target area adjacent areas; S2. Process the target area omics feature set and the target area adjacent tissue omics feature set to delete invalid features, and eliminate irrelevant features in the target area omics feature set and the target area adjacent tissue omics feature set based on feature engineering; Process the clinical feature set to delete invalid features; S3. Construct a feature subset of the target area omics feature set, establish a first prediction model based on the target area omics features, and obtain the best target area omics feature subset through the first prediction model; S4. Construct a feature subset of the target area adjacent tissue omics feature set, establish a second prediction model based on the target area adjacent tissue omics features, and obtain the best target area adjacent tissue omics feature subset through the second prediction model; S5. Construct a feature subset of the clinical feature set, establish a third prediction model based on the clinical features, and obtain the best clinical feature subset through the third prediction model; S6. Fuse the best target area omics feature subset, the best target area adjacent tissue omics feature subset, and the best clinical feature subset to obtain a final sample set, and construct a hybrid prediction model based on the final sample set; the constructed hybrid prediction model is the prediction system for breast cancer radiation dermatitis.

2. The method for constructing a breast cancer radiation dermatitis prediction system according to claim 1, wherein: In step S2, the process of processing the target area omics feature set and the target area adjacent tissue omics feature set to delete invalid features includes deleting null value features and unbalanced features in the target area omics feature set and the target area adjacent tissue omics feature set.

3. A method for constructing a breast cancer radiation dermatitis prediction system according to claim 2, characterized in that: In step S2, the elimination of irrelevant features in the target area omics feature set and the target area adjacent tissue omics feature set based on feature engineering includes the MWU test step, the correlation detection step, the LASSO dimensionality reduction processing step, the logistic regression analysis step, and the multiple linearity test step of the variance inflation factor VIF in sequence.

4. A method for constructing a breast cancer radiation dermatitis prediction system according to claim 3, characterized in that: In the MWU test step, extract the values of each omics feature in all samples with radiation dermatitis to construct a first omics feature sequence corresponding to the omics feature, and at the same time extract the values of each omics feature in all samples without radiation dermatitis to construct a second omics feature sequence corresponding to the omics feature, and eliminate the omics features corresponding to the first omics feature sequence and the second omics feature sequence with similar data distribution characteristics.

5. A method for constructing a breast cancer radiation dermatitis prediction system according to claim 4, characterized in that: In the correlation detection step, extract the values of each omics feature in all samples to construct a third omics feature sequence corresponding to the omics feature, and eliminate the omics features corresponding to the third omics feature sequence with a variance close to 0 (such as ≤0.05); and if the Pearson correlation coefficient of the third omics feature sequences of any two corresponding omics features is not less than 0.9, then eliminate the omics feature corresponding to the smaller variance of the third omics feature sequences of the two corresponding omics features.

6. A method for constructing a breast cancer radiation dermatitis prediction system according to claim 5, characterized in that: In the above-mentioned logistic regression analysis step, it is implemented by the binary logistic regression method.

7. A method for constructing a breast cancer radiation dermatitis prediction system according to claim 6, characterized in that: The above-mentioned processing of the clinical feature set to delete invalid features, performs t-test and MUW test on the clinical feature sequence, and retains the clinical features with P-values not exceeding 0.5 in both tests.

8. A method for constructing a breast cancer radiation dermatitis prediction system according to claim 7, characterized in that: In step S3, step S4 and step S5, the construction of the feature subsets of the target region omics feature set, the target region adjacent tissue omics feature set and the clinical feature set is completed based on the genetic algorithm, and the optimization of each feature subset is realized based on the wrapper algorithm.

9. A method for constructing a breast cancer radiation dermatitis prediction system according to claim 8, characterized in that: In step S3, step S4 and step S5, the first prediction model, the second prediction model and the third prediction model are constructed based on the artificial neural network.

10. A breast cancer radiation dermatitis prediction system, which is constructed based on the construction method of a breast cancer radiation dermatitis prediction system described in any one of claims 1-9.