Method for predicting risk of chronic actinic dermatitis based on risk prediction model

By integrating multi-dimensional indicators and a dynamic risk-calibrated risk prediction model, the problems of difficult early identification, low prediction accuracy, and insufficient intervention in chronic actinic dermatitis have been solved, achieving highly accurate and personalized risk prediction and intervention.

CN121075657BActive Publication Date: 2026-02-06川北医学院附属医院 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511604789.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-02-06
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

In existing technologies, risk prediction methods for chronic actinic dermatitis lack systematicity and accuracy, fail to effectively integrate multi-dimensional indicators, resulting in difficulties in early identification, low prediction accuracy, insufficient personalized intervention, and failure to consider individual differences and skin type differences, leading to a high misjudgment rate.

Method used

A risk prediction model-based approach was adopted, which integrates photobiology, immune function, molecular markers and lifestyle indicators through hierarchical preprocessing, multimodal fusion model and dynamic risk calibration. Combined with Fitzpatrick skin type and photosensitivity history, the incidence probability of the prediction model was dynamically adjusted to output personalized risk level and core risk factor ranking.

Benefits of technology

It improves the accuracy and personalization of risk prediction for chronic actinic dermatitis, reduces the misjudgment rate, enables early risk identification and personalized intervention, and reduces the waste of medical resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075657B_ABST
    Figure CN121075657B_ABST
Patent Text Reader

Abstract

The present application relates to a chronic actinic dermatitis risk prediction method based on a risk prediction model, which first acquires the multi-dimensional indexes of photobiology, immune function, molecular markers, clinical basis and life behavior of the object to be predicted and establishes a data subset; through stratified preprocessing, correction and credibility scoring, a standard data set is obtained; through variance analysis and recursive feature elimination, risk indexes are extracted, the comprehensive correlation degree of the indexes is calculated by combining the Pearson correlation coefficient and the mutual information value, the high correlation degree indexes are screened to construct a risk feature vector; the vector is input into a multi-modal fusion risk prediction model constructed by a feature attention mechanism and an improved gradient boosting tree algorithm, and the incidence probability, risk level and core risk factor ranking are output; the risk level is also calibrated in combination with skin typing and photosensitive contact history, and a preset prevention and control measure table is called to perform personalized prevention and control. The present application realizes accurate prediction and personalized prevention and control guidance of CAD incidence risk, and improves the early identification rate and intervention accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data processing, and particularly relates to a chronic actinic dermatitis pathogenic risk prediction method based on a risk prediction model. BACKGROUND

[0002] Chronic actinic dermatitis (CAD) is a type of idiopathic photosensitive skin disease characterized by chronic inflammation at light-exposed sites. Its pathogenesis is complex and its clinical diagnosis and treatment are difficult, and it has become an important threat to the skin health of middle-aged and elderly male groups. According to clinical statistics, the global incidence of CAD is increasing year by year, and it is more common in tropical and subtropical regions or summer. Some severe cases can progress to cutaneous T-cell lymphoma, significantly reducing the quality of life of patients and increasing the medical burden.

[0003] The pathogenesis of CAD is related to the synergistic action of multiple factors. Currently known ultraviolet (especially UVA, UVB) irradiation is the core inducer, but its pathogenic process also involves multiple mechanisms such as immune dysfunction, skin barrier damage, and abnormal molecular regulation. Clinical studies have shown that CAD patients have peripheral blood T cell subpopulation imbalance (CD3 + CD4 + ratio decreases, CD3 + CD8 + ratio increases), Th1 cell overactivation (IFN-γ / IL-4 ratio increases), and abnormal expression of skin tissue-specific lncRNA (such as RP11-356I2.4). The discreteness and correlation of these indicators increase the difficulty of early diagnosis. In addition, there are individual differences in the sensitivity of CAD patients to ultraviolet light. Even in an indoor environment, 50% UVA that can penetrate ordinary glass may still induce the disease, further increasing the complexity of risk prediction.

[0004] Current clinical assessment of CAD risk mainly relies on single indicator detection and physician experience, and lacks systematicness and precision. For example, the minimal erythema dose (MED) determined by a solar simulator can reflect skin photosensitivity, but it does not combine key dimensions such as immune function and molecular markers, and is prone to miss potential patients with normal photosensitivity but immune disorder; traditional photopatch tests can only identify known photosensitizing substance contact history and cannot quantify the impact of environmental photosensitive load (such as residential ultraviolet radiation intensity and sunscreen adherence) on the onset. In addition, existing methods do not consider the differentiated risk of skin types (Fitzpatrick classification), and the mean UVA-MED of type III skin patients is significantly lower than that of type IV skin, but a unified risk determination standard is used, resulting in a high misdiagnosis rate.

[0005] The incidence of CAD has the characteristics of chronicity and recurrence. The existing technology focuses on single diagnosis, and does not construct a dynamic risk model based on long-term index monitoring. For example, the CD3 + CD4 + / CD3 + CD8 + When the ratio returns to 1.2 or more, the probability of disease remission, but the traditional method does not include such dynamic indicators in the risk update mechanism.

[0006] In the existing CAD related research, the feature screening depends on single variance analysis (ANOVA) or single variable correlation analysis. For example, there is a segmented correlation between the environmental light sensitivity load index and the incidence of risk, and the Pearson correlation coefficient cannot capture such non-linear relationship, resulting in that the contribution of the index in the traditional model is underestimated. At the same time, without removing low-contribution indicators, there is redundancy in the model input, such as height, weight and other indicators that have no significant correlation with CAD, which reduces the prediction accuracy.

[0007] The existing CAD risk prediction model mostly uses a single algorithm, without considering the modal difference of multi-dimensional indicators. The weight contribution of photobiological indicators and molecular markers is different, with the former being about 0.35 and the latter being about 0.3. A single model cannot dynamically allocate dimension weight.

[0008] The incidence probability output by the existing model is not adjusted twice in combination with individual skin type and light sensitivity contact history, resulting in that the risk of patients with type III skin + light sensitivity contact history is underestimated, and the risk of patients with type IV skin + no light sensitivity contact history is overestimated. This lack of calibration directly affects the accuracy of risk level determination, and further leads to mismatched intervention measures.

[0009] With the development of precision medicine, the demand for CAD risk prediction in clinical practice has shifted from qualitative diagnosis to quantitative stratification + dynamic intervention. The present application aims to provide a prediction method that integrates multi-dimensional indicators, optimizes feature screening, and has risk calibration and intervention guidance capability, in order to solve the problems of early identification difficulty, low prediction accuracy and insufficient personalized intervention in the prior art. Through the innovative design of hierarchical preprocessing, multi-modal fusion model and dynamic risk calibration, the accuracy and clinical adaptability of CAD incidence risk prediction can be effectively improved. SUMMARY

[0010] The purpose of the present application is to provide a risk prediction model based method for predicting the incidence of chronic photodermatitis, in order to solve the above technical problems in the prior art.

[0011] To solve the above technical problems, the technical solution adopted by the present application is as follows:

[0012] The risk prediction model based method for predicting the incidence of chronic photodermatitis comprises the following steps:

[0013] S1: Obtain the photobiological indicators, immune function dynamic indicators, molecular marker combination indicators, clinical basic indicators, and life behavior risk indicators of the object to be predicted within a specified period, and establish corresponding data subsets respectively;

[0014] S2: Perform targeted preprocessing on the data subsets through hierarchical preprocessing, and then perform correction and credibility scoring respectively. Retain data with a score greater than or equal to a specified threshold, and merge them into a data set;

[0015] S3: Extract risk indicators from the data set: first filter out low-variation irrelevant indicators, then gradually eliminate indicators with low prediction contribution from the remaining indicators, and finally retain a specified number of candidate indicators;

[0016] S4: Calculate the correlation of each candidate indicator with CAD incidence risk: Pearson correlation quantifies linear correlation, mutual information value quantifies nonlinear correlation, and the comprehensive correlation of each indicator is calculated through a weighted method;

[0017] S5: Sort the candidate indicators according to the comprehensive correlation from high to low, select indicators with a comprehensive correlation greater than or equal to a threshold to construct a risk feature vector, and assign weights to each dimension according to the comprehensive correlation;

[0018] S6: Input the risk feature vector into the pre-constructed multi-modal fusion risk prediction model, and the multi-modal fusion risk prediction model outputs the incidence probability of the object to be predicted. According to the incidence probability, the incidence risk level and the core risk factor ranking are obtained.

[0019] Preferably, it also includes a risk level calibration process, which is as follows:

[0020] S7: According to the Fitzpatrick skin type standard, divide it into type III and type IV, and determine it through the history of contact with photosensitizing substances or no history of contact with photosensitizing substances. Based on skin type and photosensitizer contact history, set groups and adjust the probability adjustment coefficient respectively:

[0021] S8: Determine the grouping of the object to be predicted, obtain the corresponding adjustment coefficient, calibrate the incidence probability predicted by the model based on the adjustment coefficient, and set the incidence probability boundary constraint. According to the adjusted probability value, combined with the risk level division standard, the final incidence risk level is determined;

[0022] S9: Based on the final incidence risk level and the core risk factor ranking, call the preset risk prevention and control measure table to perform corresponding risk prevention and control measures on the object to be predicted.

[0023] Preferably, the photobiological indicators in step S1 are the minimal erythema dose of long-wave ultraviolet light UVA-MED, the minimal erythema dose of medium-wave ultraviolet light UVB-MED and the erythema regression time ERT measured by a sunlight simulator, the dynamic indicators of immune function are the proportion of peripheral blood T cell subsets detected by flow cytometry, the CD3 + CD4 + / CD3 + CD8 + ratio and the Th1 / Th2 cytokine balance state, the combination of molecular markers includes the serum IFN-γ, IL-12, sFas / sFasL, skin tissue-specific lncRNA and TNFAIP3 gene promoter region methylation rate, the clinical basic indicators include age, gender, Fitzpatrick skin type and history of previous photosensitive diseases, and the life behavior risk indicators include daily sun exposure time, sunscreen measure compliance, photosensitizer intake frequency and environmental photosensitive load index.

[0024] Among the indicators of CD3 + CD4 + / CD3 + CD8 + ratio, CD3 + CD4 + , that is, the CD3 + CD4 + proportion, refers to the percentage of Th cells that simultaneously express CD4 molecules among the total T lymphocytes expressing CD3 molecules in the peripheral blood; CD3 + CD8 + , that is, the CD3 + CD8 + proportion, refers to the percentage of Tc cells that simultaneously express CD8 molecules among the total T lymphocytes expressing CD3 molecules in the peripheral blood; CD3 + CD4 + / CD3 + CD8 + ratio refers to the division of CD3 + CD4 + and CD3 + CD8 + , to obtain a numerical indicator reflecting the balance state of T cell subsets.

[0025] Preferably, the 5 data subsets in step S2 are preprocessed in a hierarchical preprocessing manner, and the specific process of correction and credibility scoring is as follows:

[0026] S21: Box-Cox transformation is performed on the photobiological indicators to eliminate the dimension;

[0027] The dynamic index of immune function is removed from abnormal values by Z-score method;

[0028] The combination index of molecular markers is filled with missing values by multiple interpolation method;

[0029] The clinical basic index and the life behavior risk index are converted into classification variables by one-hot encoding;

[0030] Each index is standardized;

[0031] S22: Verify the data distribution of each data subset by 3σ criterion, and correct the abnormal subset by local weighted regression;

[0032] S23: Generate 0-10 score data reliability score based on the accuracy of the detection instrument and the stability of the sample collection, and retain the data with score ≥8, and merge to obtain the standardized data set.

[0033] Preferably, the specific process of step S3 is as follows:

[0034] S31: Obtain the index data of the healthy subjects with the same clinical basic index as the to-be-predicted object as the healthy control group, and take the data set in step S2 as the CAD case group;

[0035] S32: Calculate the overall variance, within-group variance, within-group variance, and between-group variance of each index, calculate the F value by ANOVA formula, F value = between-group variance / in-group variance;

[0036] The following screening criteria are set for index screening:

[0037] The overall variance of the index is greater than 0.1;

[0038] F value > 3.85;

[0039] The index meets the above two points at the same time, otherwise it is removed;

[0040] S33: Further index screening is performed by incremental index elimination method RFE.

[0041] Preferably, the specific process of calculating the correlation degree of each candidate index and the risk of CAD in step S4 is as follows:

[0042] S41: Quantify the linear correlation between each candidate index and the risk of CAD by calculating Pearson correlation coefficient;

[0043] S42: Quantify the nonlinear correlation between each candidate index and the risk of CAD by calculating mutual information value;

[0044] S43: Calculate the comprehensive correlation degree based on the Pearson correlation coefficient and the mutual information value: adopt the weighted formula of linear correlation weight 0.6 + nonlinear correlation weight 0.4.

[0045] S44: For all candidate indicators screened by ANOVA, repeat steps S41-S43 to calculate the R 综合 .

[0046] Preferably, the specific process of step S5 is as follows:

[0047] S51: Calculate the R 综合 According to the order from high to low, filter the indicators with R 综合 ≥0.6 to construct the risk feature;

[0048] S52: Construct the risk feature vector based on the screened 8-12 indicators, and assign the basic weight to each dimension according to the comprehensive correlation degree.

[0049] Preferably, the multi-modal fusion risk prediction model of step S6 is constructed by a feature attention mechanism + an improved gradient boosting tree algorithm, the risk feature vector of the historical case is used as the training sample, the feature attention mechanism dynamically allocates the basic weight of each dimension in the risk feature vector, and the improved gradient boosting tree algorithm optimizes the model training efficiency by introducing an adaptive learning rate;

[0050] After the risk feature vector of the to-be-predicted object is input into the multi-modal fusion risk prediction model, the multi-modal fusion risk prediction model outputs the incidence probability of the to-be-predicted object, and the incidence risk level and the core risk factor ranking are obtained according to the incidence probability, and the risk level includes low risk, medium risk and high risk.

[0051] The beneficial effects of the present application include:

[0052] The chronic actinic dermatitis incidence risk prediction method based on the risk prediction model provided by the present application solves the core pain points of single index, insufficient precision and blind intervention in the existing CAD risk prediction through multi-dimensional index integration, refined feature screening, multi-modal model construction and personalized risk calibration, and realizes significant breakthroughs in clinical practicability, prediction accuracy and personalized adaptability.

[0053] 1. The photobiological index, immune function index, molecular marker, clinical basic index and life behavior index are integrated into a unified system, and the cause-inducing-immunity-molecule-behavior pathogenic chain of CAD is comprehensively covered. For example, through the synergistic detection of UVA-MED and TNFAIP3 methylation rate, the recognition rate of early asymptomatic patients is improved, and the traditional single index is avoided.

[0054] 2. Two-step strategy of filtering low-variation indicators by ANOVA + removing low-contribution indicators by RFE, combined with comprehensive correlation calculation of Pearson correlation coefficient of linear correlation + mutual information value of nonlinear correlation, accurately retaining 8-12 core features, improving the effective correlation rate of model input features, and at the same time controlling data quality through 3σ criterion and reliability score to ensure the reliability of model input.

[0055] 3. The model is constructed by using feature attention mechanism + improved gradient boosting tree, the feature attention mechanism can dynamically allocate the weight of each dimension, and the improved gradient boosting tree optimizes the training efficiency through adaptive learning rate, which can not only capture the linear correlation between indicators, but also identify the nonlinear correlation, so as to improve the prediction accuracy of the model.

[0056] 4. According to the risk difference of Fitzpatrick classification, the grouping adjustment coefficient of skin type-light sensitive contact history is set, and the incidence probability is optimized through a quadratic calibration formula. This calibration mechanism reduces the misdiagnosis rate of type III skin patients and reduces the misdiagnosis rate of type IV skin patients, solving the adaptability problem of traditional unified standard.

[0057] 5. The model output not only includes the risk level, but also analyzes and generates the core risk factor ranking, which clearly identifies the key driving factors of patient onset, facilitates targeted intervention based on the core risk factor ranking, avoids the one-size-fits-all approach of clinical intervention, and reduces the waste of medical resources, such as unnecessary genetic testing for medium-risk patients.

[0058] 6. Promote the transformation of CAD prevention and control from passive diagnosis to active early warning: early risk identification, delay disease progression, and early warning of potential patients with normal photosensitivity but immune disorders or abnormal molecular markers can identify risks 6-12 months before clinical lesions appear. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 The flowchart of the chronic photodermatitis onset risk prediction method based on the risk prediction model of the present application.

[0060] Figure 2 The logic principle diagram of the chronic photodermatitis onset risk prediction based on the risk prediction model of the present application.

[0061] Figure 3 The architecture diagram of the multi-modal fusion risk prediction model of the present application. DETAILED DESCRIPTION

[0062] The following will be combined with the drawings Figures 1-3 Further detailed description of the present application:

[0063] Example 1:

[0064] Reference is made to the drawingsFigure 1 and Figure 2 As shown in the figure, the chronic actinic dermatitis incidence risk prediction method based on the risk prediction model comprises the following steps:

[0065] S1: Obtain the photobiological index, immune function dynamic index, molecular marker combination index, clinical basic index and life behavior risk index of the to-be-predicted object within a specified time period, and establish corresponding five data subsets respectively.

[0066] S2: The five data subsets are preprocessed in a hierarchical preprocessing manner, and then corrected and scored for credibility, and data equal to or greater than a specified scoring threshold is retained to form a data set.

[0067] S3: Risk index extraction is performed on the data set: low-variation irrelevant indexes are filtered out, and indexes with low prediction contribution are gradually removed from the remaining indexes, and finally 15-20 candidate indexes are retained.

[0068] S4: Calculate the correlation degree of each candidate index and the incidence risk of CAD: Pearson correlation quantifies linear correlation, with a value range of [-1, 1], and the closer the absolute value is to 1, the higher the correlation degree is, mutual information value quantifies nonlinear correlation, with a value range of [0, log2(n)], n is the sample size, and the larger the value is, the higher the correlation degree is, and the comprehensive correlation degree of each index is calculated by a weighted formula, with a value range of [0, 1].

[0069] S5: Sort the candidate indexes according to the comprehensive correlation degree from high to low, select the indexes with a comprehensive correlation degree greater than or equal to 0.6 to construct a risk feature vector, the vector dimension is 8-12, each dimension is assigned a weight according to the correlation degree, and the weight of the index with the highest correlation degree is set to 0.2, and the weights of the remaining indexes decrease in proportion to the correlation degree.

[0070] S6: Input the risk feature vector into the pre-constructed multi-modal fusion risk prediction model, and the multi-modal fusion risk prediction model outputs the incidence probability of the to-be-predicted object, and the incidence risk level (low risk, medium risk, high risk) and core risk factor ranking are obtained according to the incidence probability.

[0071] In this embodiment, the photobiological index in step S1 is the minimum erythema dose of long-wave ultraviolet light UVA-MED, the minimum erythema dose of medium-wave ultraviolet light UVB-MED and the erythema regression time ERT measured by a sunlight simulator, the immune function dynamic index is the peripheral blood T cell subsets (CD3 + CD4 + proportion, CD3 + CD8 + proportion) detected by flow cytometry, CD3 + CD4 + / CD3 + CD8+ The ratio and Th1 / Th2 cytokine balance state (IFN-γ / IL-4 ratio) refers to the value obtained by dividing the concentration of IFN-γ (interferon gamma) in serum by the concentration of IL-4 (interleukin-4) in serum; the combination of molecular markers includes serum IFN-γ, IL-12, sFas / sFasL (ABC-ELISA method), skin tissue-specific lncRNA (RP11-356I2.4, LNC_000310, qRT-PCR method), and TNFAIP3 gene promoter region methylation rate (bisulfite sequencing method, detection region chr6: 1378200-1378500); the clinical basic indicators include age, gender, Fitzpatrick skin type, and history of previous photosensitive diseases, and the life behavior risk indicators include daily average sun exposure time, sunscreen measure compliance, photosensitizer intake frequency, and environmental photosensitive load index (calculation formula: index = UV radiation intensity of residence × sun exposure time × indoor glass UVA penetration rate).

[0072] In another embodiment of the present embodiment, a risk level calibration process is also included, specifically as follows:

[0073] S7: According to the Fitzpatrick skin type standard, it is divided into type III and type IV, type III is light brown skin, easy to sunburn and moderate sunburn after sun exposure, and type IV is brown skin, not easy to sunburn and easy to sunburn after sun exposure. The ultraviolet tolerance threshold of the two types of skin is significantly different, and the UVA-MED of type skin is 20%-30% higher than that of type III.

[0074] Through the past medical history, medication history (such as taking quinolone antibiotics and tetracycline drugs), occupational exposure history (such as long-term contact with asphalt and coal tar), and photopatch test results, it is determined to have a clear exposure history (exposure history ≥ 1 time and photopatch test positive) or no photosensitive exposure history (no exposure history or photopatch test negative).

[0075] Based on skin type and photosensitizer exposure history, grouping is set, and a probability adjustment coefficient is set respectively K :

[0076] Type III has a clear exposure history, and the adjustment coefficient K is 1.10-1.15;

[0077] Type III has no clear exposure history, and the adjustment coefficient K is 1.00-1.08;

[0078] Type IV has a clear exposure history, and the adjustment coefficient K is 1.05-1.08;

[0079] Type IV has no clear exposure history, and the adjustment coefficientK 0.92-0.95.

[0080] S8: Determine the group of the object to be predicted, obtain the corresponding adjustment coefficient, and based on the adjustment coefficient, adjust the incidence probability predicted by the model P init Calibration is performed, and the calibration formula is as follows:

[0081] P adj P init ×K ;

[0082] wherein, P adj is the calibrated incidence probability;

[0083] and set the boundary constraint of the incidence probability:

[0084] If P adj > 1.0, then P adj = 1.0 (judged as extremely high risk);

[0085] If P adj < 0.05, then P adj = 0.05 (judged as extremely low risk), to ensure the rationality of the subsequent risk level determination;

[0086] According to the adjusted probability value P adj , combined with the risk level division standard, the final incidence risk level is determined: low risk: P adj < 0.25;

[0087] Medium risk: 0.25 ≤ P adj ≤ 0.75;

[0088] High risk: P adj > 0.75.

[0089] S9: Based on the final incidence risk level and the ranking of the core risk factors, the preset risk prevention and control measure table is called to execute the corresponding risk prevention and control measures on the object to be predicted. The risk prevention and control measure table is shown in the following Table 1.

[0090] Table 1 Risk prevention and control measure table

[0091] ​obtaining the final risk level of the to-be-predicted object from the multi-modal fusion model output result, and screening the comprehensive correlation degree R from the core risk factor ranking output by the model 综合 The top 3 factors, and the abnormal state of each factor, such as TNFAIP3 methylation rate < 20%, RP11-356I2.4 expression lower than 50% of the healthy mean value.

[0092] For example: the core factors of a high-risk object are TOP3:

[0093] 1. TNFAIP3 methylation rate < 20%, R 综合 = 0.75;

[0094] 2. RP11-356I2.4 expression is low, R 综合 = 0.72;

[0095] 3. (IFN-γ) / (IL-4) > 2.0, R 综合 = 0.68;

[0096] Form the "high risk-TNFAIP3+RP11-356I2.4+IFN-γ / IL-4" combination label of the object.

[0097] In the risk prevention and control measure table, first locate the vertical column according to the final risk level, such as the high risk column, and then locate the horizontal row according to the TOP3 core risk factors, such as the "TNFAIP3 methylation rate+RP11-356I2.4+IFN-γ / IL-4" row, and the intersection cell is the exclusive prevention and control scheme of the object.

[0098] Example 2

[0099] On the basis of example 1, the specific process of correcting and credibility scoring of the 5 data subsets through hierarchical preprocessing in step S2 is as follows:

[0100] S21: Box-Cox transformation is performed on the photobiological index to eliminate the dimension;

[0101] The abnormal values of the immune function dynamic index are removed by Z-score method (Z-score absolute value > 3);

[0102] The missing values of the molecular marker combination index are filled by multiple imputation method (based on the same disease stage and age group distribution);

[0103] The clinical basic index and life behavior risk index are converted into classification variables by one-hot encoding;

[0104] The indexes are standardized.

[0105] S22: Verify the data distribution of each data subset by the 3σ criterion, 99.7% of the data falls within the mean ± 3 times the standard deviation, and the abnormal subset is corrected by local weighted regression LOESS.

[0106] S23: Generate a 0-10 data reliability score based on the detection instrument accuracy and sample collection stability, retain data with a score ≥8, and merge to obtain a standardized data set.

[0107] The specific process of step S3 is as follows:

[0108] S31: Obtain the data of each index of healthy subjects with the same clinical basic index as the object to be predicted as the healthy control group, and the data set in step S2 as the CAD case group.

[0109] S32: Calculate the overall variance, within-group variance, within-group variance, and between-group variance of each index, and calculate the F value by the ANOVA formula:

[0110] F value = between-group variance / within-group variance, between-group variance reflects the difference between the two groups, and within-group variance reflects the data fluctuation within the group.

[0111] The following screening criteria are set for index screening:

[0112] The overall variance of the index is greater than 0.1, ensuring that the index has sufficient variation;

[0113] F value > 3.85, ensuring that the difference between the two groups is statistically significant;

[0114] The index meets the above two points at the same time, otherwise it is rejected, and the number of retained indexes is M;

[0115] Take CD3 + CD4 + ratio index as an example, assuming that the sample size of the case group n1 = 200, the sample size of the control group n2 = 200, and the basic parameters are as follows:

[0116] n1: sample size of the case group; n2: sample size of the control group; total sample size N = n1 + n2 = 400;

[0117] : mean of the case group index, such as mean of the case group CD3 + CD4 + ratio 30.2%;

[0118] : mean of the control group index, such as mean of the healthy group CD3 + CD4 + ratio 42.5%;

[0119] : total sample CD3+ CD4 + The average percentage of the proportional indicator is 36.35%: (200 × 30.2% + 200 × 42.5%) / 400 = 36.35;

[0120] Case group number i CD3 of each sample + CD4 + Proportional index values, such as CD3 in case 1 of the case group + CD4 + =28.5%;

[0121] : Control group j CD3 of each sample + CD4 + Proportional index values, such as CD3 in case 1 of the control group. + CD4 + =41.8%;

[0122] k: number of groups, k=2.

[0123] The process of calculating the variance between groups is as follows:

[0124] First calculate the sum of squared deviations between groups from the mean. SSB Between-group degrees of freedom dfB Then, the variance between groups can be obtained by comparing the two values:

[0125] Between-group sum of squared differences SSB This measures the sum of squared deviations of each group's mean from the overall mean, using sample size weighting to ensure fairness. The specific formula is as follows:

[0126] ;

[0127] Substitute the above data:

[0128] SSB = 200 × (30.2 - 36.35) 2 +200×(42.5-36.35) 2

[0129] =200×(-6.15) 2 +200×(6.15) 2

[0130] =200×37.82+200×37.82

[0131] =7564+7564=15128;

[0132] Between-group degrees of freedom (dfB): the amount of information from independent groups, expressed by the formula: dfB =k-1=1;

[0133] MSB= 15128 / 1 = 15128. SSB / dfB

[0134] The within-group variance reflects the degree of fluctuation of the indicators within the same group (such as within the case group), the smaller the fluctuation, the stronger the stability of the indicators, and the higher the reliability.

[0135] The calculation process of the within-group variance MSW is as follows:

[0136] First, calculate the sum of squares of within-group deviations SSW and the within-group degrees of freedom dfW , and then get the within-group variance through the ratio of the two:

[0137] Sum of squares of within-group deviations SSW : measures the sum of squares of deviations of each sample within the group from the mean of the group, excluding the interference of the group difference.

[0138] ;

[0139] Substitute the above data:

[0140] SSW = 850 + 920 = 1770;

[0141] Within-group degrees of freedom dfW is the number of independent samples within the group, and the formula is: dfW=N-k = 400 - 2 = 398;

[0142] Within-group variance MSW = SSW / dfW= 1770 / 398 ≈ 4.45.

[0143] Total variance: reflects the overall variation of all samples (case group + control group), which is a comprehensive reflection of the between-group variance and the within-group variance. First, calculate the total sum of squares of deviations SST and the total degrees of freedom dfT , and then get the total variance through the ratio of the two.

[0144] The total sum of squares of deviations SST measures the sum of squares of deviations of all samples from the total mean, which is equal to the sum of the between-group sum of squares and the within-group sum of squares:

[0145] That is, SST = SSB + SSW = 15128 + 1770 = 16898;

[0146] Total degrees of freedom dfT is the number of total independent samples, dfT = N - 1 = 400 - 1 = 399;

[0147] Total variance σ 2 = 16898 / 399 = 42.35;​

[0148] The following screening criteria are set for index screening:

[0149] The total variance of the index is greater than 0.1 (to ensure that the index has sufficient variation);

[0150] F value = 15128 / 4.45 = 3400 > 3.85, indicating that the difference between groups is significant, and the index can be retained;

[0151] S33: Further index screening is performed by incremental index elimination method RFE:

[0152] In each iteration process of the logistic regression model, the 5% of the M indexes with the lowest importance score are removed, and the specific process is as follows:

[0153] First iteration:

[0154] The logistic regression model is trained with the M indexes, and the importance score of each index is calculated by the absolute value of the model coefficient. The higher the score, the greater the contribution to the prediction. Assuming N = 22;

[0155] Remove the 1 index with the lowest score (5% x 22 = 1), leaving 21 indexes;

[0156] The logistic regression model is retrained with the 21 indexes, and the 5-fold cross-validation AUC value is calculated. Assuming it increases from 0.82 to 0.83.

[0157] 2-5 iterations:

[0158] Repeat the process of training the logistic regression model, calculating the importance, removing the lowest score index, and verifying the AUC. Remove 1 low contribution index each time, and the number of indexes decreases from 21 to 18.

[0159] If the AUC value does not improve after an iteration, such as 0.85, then continue to remove; if the AUC value does not improve for 3 consecutive times, such as 0.86 for the 3rd-5th time, then stop iteration;

[0160] Assuming that after 5 iterations, the number of indexes decreases from 22 to 18, and the AUC value stabilizes at 0.86 (an increase of 0.04 from the initial 22 indexes), then the 18 indexes are the candidate indexes after RFE screening.

[0161] Example 3:

[0162] On the basis of example 1 or example 2, the specific process of calculating the correlation degree of each candidate index with the risk of CAD in step S4 is as follows:

[0163] Data preparation: 500 cases of CAD were diagnosed, which met the criteria of disease course > 1 year, rash distribution in light exposed parts, and clear photosensitivity symptoms; 500 healthy control samples had no history of photosensitivity diseases and had not taken photosensitive drugs recently, with a total sample size of N≥1000. The samples were pretreated. The "candidate index-disease state" data set was constructed, in which the disease state was a binary variable, the case group was marked as 1, and the control group was marked as 0.

[0164] Definition of core parameters:

[0165] X: Numerical vector of candidate indicators, such as UVA-MED values of 500 cases + 500 controls, a total of 1000 data points;

[0166] Y: Disease state, 1000 data points, 1 = CAD case, 0 = healthy control;

[0167] r: Pearson correlation coefficient, value [-1, 1], the larger the absolute value |r|, the stronger the linear correlation;

[0168] MI(X, Y): Mutual information value of indicator X and disease state Y, value [0, log2n], when n=1000, log21000≈9.97, the larger the value, the stronger the nonlinear correlation;

[0169] maxMI: The maximum value of all candidate feature mutual information values, used for normalization, so that the mutual information value is mapped to the range [0, 1];

[0170] R 综合 : Comprehensive correlation value [0, 1], the closer the value to 1, the stronger the correlation between the feature and the risk of CAD.

[0171] S41: Quantify the linear correlation between each candidate indicator and the risk of CAD by calculating the Pearson correlation coefficient:

[0172] ;

[0173] Where n is the total sample size, is the product sum of indicator X and disease state Y, ∑X is the total sum of indicator X, ∑Y is the total sum of disease state, is the square sum of indicator X, ∑Y 2 is the square sum of disease state.

[0174] Take the UVA-MED feature as an example: Assuming that the mean UVA-MED of 500 cases (Y=1) is 8.2 J / cm², the mean UVA-MED of 500 controls (Y=0) is 18.5 J / cm², and the total sample size n=1000, the calculation is:

[0175] ∑X=500×8.2+500×18.5=13350;

[0176] ∑Y=500×1+500×0=500;

[0177] ∑XY=500×8.2×1+500×18.5×0=4100;

[0178] ∑X 2 =500×(8.2 2 +variance)+500×(18.5) 2 (+variance) = 218925; Variance = 4.96;

[0179] ∑Y 2 =500×12+500×02=500.

[0180] Substituting into the formula, we get:

[0181] r=(1000×4100-13350×500) / ([1000×218925-13350 2 [1000×500-5002]) 1 / 2 ≈-0.78;

[0182] Interpretation of results: r≈-0.78, the negative sign indicates that UVA-MED is negatively correlated with the risk of CAD, that is, the lower the UVA-MED, the higher the risk of CAD. |r|=0.78, the linear association is relatively strong.

[0183] S42: Quantifying the nonlinear association between each candidate indicator and the risk of CAD incidence by calculating mutual information values:

[0184] Discretization process: First, the continuous feature X is discretized into K intervals. For example, UVA-MED is divided into 5 intervals with an interval of 5J / cm²: [0-5), [5-10), [10-15), [15-20), and ≥20J / cm². Ensure that each interval contains enough samples, with ≥50 samples.

[0185] Calculate the marginal probabilities and joint probabilities:

[0186] Marginal probability P( X=x i Feature X takes the value of x i The sample proportion is given by the formula: P( X=x i The characteristic X takes the value of ) = x i Number of samples / n;

[0187] Marginal probability P( Y=yj ) : the sample proportion of the feature X taking value y j

[0188] P( X=y j ) : the sample number of the feature X taking value

[0189] P( X=x i , P( Y=y j )) : the sample proportion of the feature X taking value x i and the disease state Y taking value y j , P( X =x i , P( Y=y j )) : the sample number of the feature X taking value x i and the disease state Y taking value y j / n.

[0190] The mutual information value is calculated as follows:

[0191]

[0192] If X and Y are independent, P( X=x i , P( Y=y j )) = P( X=x i ) · P( Y=y j ), in which case log2(1) = 0, MI = 0.

[0193] The stronger the correlation between X and Y, the greater the ratio and the higher the MI value.

[0194] Take the UVA-MED index as an example:

[0195] Suppose the UVA-MED is discretized into 5 intervals, and the joint probability and marginal probability of each interval are calculated, as shown in Table 2 below:

[0196] Table 2 Joint probability and marginal probability of each interval

[0197]

[0198] Substitute into the formula:

[0199] ​​MI(X, Y) = 0.1425 x log2(0.1425 / (0.15 x 0.5)) + 0.0075 x log2(0.0075 / (0.15 x 0.5)) + 0.0950 x log2(0.0950 / (0.10 x 0.5))

[0200] + 0.0950 x log2(0.0950 / (0.10 x 0.5))

[0201] ≈ 4.2;

[0202] MI ≈ 4.2, maxMI ≈ 9.97 when n = 1000, indicating that UVA-MED has a strong non-linear correlation with the risk of CAD.

[0203] S43: Calculate the comprehensive correlation degree based on the Pearson correlation coefficient and the mutual information value:

[0204] Since the MI values of different features are in different ranges, depending on the feature discretization interval and sample distribution, they need to be normalized to the range [0, 1] first. The formula is: MI 归一化 = MI(X, Y) / maxMI;

[0205] If the maxMI of all candidate features is approximately 9.97, then the MI 归一化 of UVA-MED is MI(X, Y) / maxMI = 4.2 / 9.97 ≈ 0.42.

[0206] Using the weighted formula of linear correlation weight 0.6 + non-linear correlation weight 0.4, based on clinical data verification, the linear correlation contributes more to the prediction of CAD risk, therefore, the formula is that the linear correlation weight is larger:

[0207] R 综合 = 0.6 x |r| + 0.4 x MI 归一化 ;

[0208] Taking the data of UVA-MED instance: R 综合 = 0.6 x 0.78 + 0.4 x 0.42 = 0.636;

[0209] Then the comprehensive correlation degree of UVA-MED is approximately 0.64 (≥ 0.6), which meets the high correlation degree feature standard and can be included in the risk feature vector.

[0210] S44: For all candidate indicators after ANOVA screening (such as UVA-MED, CD3 + CD4 + ratio, TNFAIP3 methylation rate, etc.), repeat steps S41-S43 to calculate their R 综合 .

[0211] The specific process of step S5 is as follows:

[0212] S51: Calculate R for all candidate indicators calculated in step S4 综合 Filter the indicators R in descending order 综合 Indicators with R ≥ 0.6 are used to build the risk feature, usually 8-12 indicators are retained;

[0213] S52: Build a risk feature vector based on the filtered 8-12 indicators, i.e. the vector dimension is 8-12, and assign a basic weight to each dimension according to the comprehensive correlation degree, and the weight of the indicator with the highest correlation degree is set to 0.2, and the weights of the remaining indicators decrease in proportion to the correlation degree.

[0214] Referring to Figure 3 , the multi-modal fusion risk prediction model is built based on feature attention mechanism + improved gradient boosting tree algorithm, the risk feature vector of historical cases is used as training sample, the feature attention mechanism dynamically allocates the basic weight of each dimension in the risk feature vector, the improved gradient boosting tree algorithm optimizes the model training efficiency by introducing adaptive learning rate, the initial learning rate is 0.05, and the learning rate is dynamically decayed according to the formula learning rate = 0.05 × [1-iteration number / 1000 × 0.6] with the number of iterations). After inputting the risk feature vector of the object to be predicted into the multi-modal fusion risk prediction model, the multi-modal fusion risk prediction model outputs the incidence probability of the object to be predicted, and the incidence risk level (low risk, medium risk, high risk) and core risk factor ranking are obtained according to the incidence probability.

[0215] The specific process of the trained multi-modal fusion risk prediction model outputting the incidence probability of the object to be predicted and the core risk factor ranking is as follows:

[0216] Input layer:

[0217] The risk feature vector of the object to be predicted, i.e. the filtered 8-12 high correlation indicators such as UVA-MED, TNFAIP3 methylation rate, etc., each dimension with a basic weight, is input into the input layer of the model, completing the preliminary reception and format adaptation of the data, ensuring that the subsequent modules can directly call the indicator data in the vector.

[0218] Feature attention mechanism dynamically adjusts feature weight:

[0219] Based on the mean square error loss function of the model, the gradient of each risk feature dimension is calculated, and the weight of each feature is dynamically adjusted according to the gradient size:

[0220] If a certain indicator (such as TNFAIP3 methylation rate) has a large impact on the loss function (large absolute value of gradient), it means that this feature is more critical to the prediction result, and its weight will be increased;

[0221] If a certain feature (such as some clinical basic indicators) has a small impact on the loss function (small absolute value of gradient), its weight will be reduced.

[0222] Every 50 model iterations, the weights of each feature are recalculated and updated to better match the prediction needs of the current data.

[0223] Improved gradient boosting tree calculation of incidence probability:

[0224] First, hyperparameter optimization, the model pre-completion of hyperparameter optimization configuration:

[0225] The number of trees is set to 1000 to ensure the model has enough learning ability;

[0226] The maximum depth of each tree is set to 8 layers to avoid overfitting;

[0227] The minimum number of split samples is set to 20 to ensure that the node splitting has enough data support.

[0228] Then, adaptive learning rate update:

[0229] The initial learning rate is set to 0.05, and then it is dynamically adjusted according to the formula learning rate iteration times. As the number of iterations increases, the learning rate gradually decreases, allowing the model to learn quickly in the early stages of training and fine-tune in the later stages, improving training efficiency and stability.

[0230] L1 regularization constraint: Through the L1 regularization unit (penalty coefficient set to 0.01), the model parameters are constrained. This can effectively suppress the interference of abnormal values on the model, making the model pay more attention to the relationship between general and stable risk features and incidence probability, and enhancing the model's generalization ability.

[0231] Incidence probability calculation: Based on the adjusted feature weights, the gradient boosting tree is improved to calculate the node splitting of each tree based on the squared error loss function. By combining the prediction results of all trees, the preliminary incidence probability of the object to be predicted is obtained.

[0232] Core risk factor ranking generation: Using the SHAP (SHapley Additive exPlanations) value analysis method, the contribution of each risk feature to the prediction result of the incidence probability is calculated. The higher the contribution, the more critical the feature in the prediction. According to the contribution from high to low, the risk features are sorted to generate the core risk factor ranking, to clarify which indicators (such as TNFAIP3 methylation rate, UVA-MED, etc.) are the main factors leading to high incidence risk of the object to be predicted.

Claims

1. A method for predicting the risk of onset of chronic actinic dermatitis based on a risk prediction model, characterized by, Comprise the following steps: S1: Obtain the photobiological indicators, immune function dynamic indicators, molecular marker combination indicators, clinical basic indicators and life behavior risk indicators of the object to be predicted within a specified period of time, and establish corresponding data subsets respectively; S2: The data subsets are preprocessed by hierarchical preprocessing, then corrected and scored for reliability, and data with a score greater than or equal to a specified threshold are retained to form a data set; S3: Risk indicator extraction is performed on the data set: first, filter out low-variation irrelevant indicators, then gradually eliminate indicators with low prediction contribution from the remaining indicators, and finally retain a specified number of candidate indicators; S4: Calculate the correlation of each candidate indicator with CAD incidence risk: Pearson correlation quantifies linear correlation, mutual information value quantifies nonlinear correlation, and the comprehensive correlation of each indicator is calculated by weighting; S5: Sort the candidate indicators according to the comprehensive correlation from high to low, select indicators with a comprehensive correlation greater than or equal to a threshold to construct a risk feature vector, and assign weights to each dimension according to the comprehensive correlation; S6: Input the risk feature vector into the pre-constructed multi-modal fusion risk prediction model, and the multi-modal fusion risk prediction model outputs the incidence probability of the object to be predicted, and the incidence risk level and core risk factor ranking are obtained according to the incidence probability; The specific process of calculating the correlation of each candidate indicator with CAD incidence risk in step S4 is as follows: S41: Calculate the linear correlation between each candidate indicator and CAD incidence risk by Pearson correlation; S42: Quantify the nonlinear correlation between each candidate indicator and CAD incidence risk by calculating the mutual information value; S43: Calculate the comprehensive correlation based on the Pearson correlation coefficient and the mutual information value: use a weighting formula with a linear correlation weight of 0.6 and a nonlinear correlation weight of 0.4; S44: For all the candidate indicators after ANOVA screening, repeat steps S41-S43 to calculate the R2of each 综合 ; The specific process of step S5 is as follows: S51: Calculate R for all candidate indicators calculated in step S4 综合 Select indicators with R ranked from high to low 综合 Indicators with R ≥ 0.6 are used to construct the risk feature S52: Based on the selected 8-12 indicators, construct a risk feature vector and assign a basic weight to each dimension according to the comprehensive correlation; It also includes a risk level calibration process, which is as follows: S7: According to the Fitzpatrick skin type standard, divide it into type III and type IV, determine it by the light-sensitive contact history or no light-sensitive contact history of the object to be predicted, set the grouping based on the skin type and light-sensitive substance contact history, and set the probability adjustment coefficient respectively: S8: Determine the grouping of the object to be predicted, obtain the corresponding adjustment coefficient, calibrate the incidence probability predicted by the model based on the adjustment coefficient, set the incidence probability boundary constraint, and determine the final incidence risk level according to the adjusted probability value and the risk level division standard; S9: Based on the final incidence risk level and the core risk factor ranking, call the pre-set risk prevention and control measures table to perform corresponding risk prevention and control measures on the object to be predicted.

2. The method of claim 1, wherein the method is a method of predicting the risk of developing chronic actinic dermatosis based on a risk prediction model, characterized by, The photobiological indicators in step S1 are the minimum erythema dose of long-wave ultraviolet light UVA-MED, the minimum erythema dose of medium-wave ultraviolet light UVB-MED and the erythema regression time ERT measured by a sunlight simulator, the dynamic indicators of immune function are the proportion of peripheral blood T cell subsets detected by flow cytometry, the proportion of CD3 + CD4 + / CD3 + CD8 + , the Th1 / Th2 cytokine balance state, the combination of molecular markers includes the serum IFN-γ, IL-12, sFas / sFasL, skin tissue-specific lncRNA and TNFAIP3 gene promoter region methylation rate, the clinical basic indicators include age, gender, Fitzpatrick skin type and history of previous photosensitivity diseases, the life behavior risk indicators include daily average sun exposure time, sunscreen measure compliance, photosensitizer intake frequency and environmental photosensitivity load index.

3. The method of claim 1, wherein the method is a method of predicting the risk of developing chronic actinic dermatosis based on a risk prediction model, characterized by, The specific process of step S2 is as follows: S21: Box-Cox transformation is performed on the photobiological indicators to eliminate dimensions; Z-score method is used to remove outliers from immune function dynamic indicators; The molecular marker combination index is filled with missing values by multiple imputation; The clinical basic index and the life behavior risk index are converted into classification variables by one-hot encoding; Each index is standardized; S22: Verify the data distribution of each data subset by 3σ criterion, and correct the abnormal subset by local weighted regression; S23: Generate a 0-10 score data reliability score based on the detection instrument accuracy and sample collection stability, retain the data with a score greater than or equal to 8, and merge to obtain a standardized data set.

4. The method of claim 3, wherein the method is a method of predicting the risk of developing chronic actinic dermatosis based on a risk prediction model, characterized by, The specific process of step S3 is as follows: S31: Obtain the index data of the healthy subjects with the same clinical basic index as the to-be-predicted object as the healthy control group, and take the data set in step S2 as the CAD case group; S32: Calculate the overall variance, within-group variance, and between-group variance of each index, and calculate the F value by the ANOVA formula, F value = between-group variance / within-group variance; The following screening criteria are set for index screening: The overall variance of the index is greater than 0.1; The F value is greater than 3.85; Both of the above two points are met, otherwise, they are rejected; S33: Further index screening is performed by incremental index elimination method RFE.

5. The method of claim 1, wherein the method is a method of predicting the risk of developing chronic actinic dermatosis based on a risk prediction model, characterized by, The multi-modal fusion risk prediction model of step S6 is constructed by feature attention mechanism + improved gradient boosting tree algorithm, the risk feature vector of the historical cases is taken as the training sample, the feature attention mechanism dynamically allocates the basic weight of each dimension in the risk feature vector, and the improved gradient boosting tree algorithm optimizes the model training efficiency by introducing an adaptive learning rate; After the risk feature vector of the to-be-predicted object is input into the multi-modal fusion risk prediction model, the multi-modal fusion risk prediction model outputs the incidence probability of the to-be-predicted object, and the incidence risk level and core risk factor ranking are obtained according to the incidence probability, the risk level includes low risk, medium risk, and high risk.

Citation Information

Patent Citations

  • A predictive model of the risk of depression after a stroke based on magnetic resonance spectroscopy imaging

    BE1029568A1

  • Application of protein marker in preparation of product for predicting future coronary heart disease onset risk of subject

    CN119959557A