Immunotherapy prognosis evaluation system based on combined penalty likelihood modeling

By adopting a combined punishment likelihood modeling strategy in tumor immunotherapy, the association between multi-endpoint data and clonal mutation characteristics is solved based on a generalized linear mixed model, and the limitations of the TMB evaluation method and the complexity of the data analysis of the immunotherapy cohort are achieved, achieving more accurate prognostic evaluation and personalized treatment strategies.

CN120108718AActive Publication Date: 2025-06-06NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202510135091.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-02
Filing Date
2025-02-06
Publication Date
2025-06-06
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

Existing tumor mutation burden (TMB) assessment methods cannot fully capture the intrinsic complex heterogeneity of the tumor, resulting in insufficient perfection in predicting the effectiveness of immunotherapy. Meanwhile, the integration of multi-endpoint data and the high degree of collinearity of clonal genomic data in the immunotherapy cohort also challenge the effectiveness of traditional data analysis methods.

Method used

A joint punishment likelihood modeling strategy was adopted, and random effects were introduced in multi-endpoint data based on generalized linear mixed model (GLMM), and the potential associations between clonal mutation characteristics were processed through the penalty term, multi-endpoint data were integrated and model construction under limited sample size conditions were optimized.

Benefits of technology

It improves the statistical effectiveness and interpretability of the model, enhances the efficiency of statistical inference and the scalability of the model, and provides more comprehensive and accurate information support for clinical decision-making of tumor immunotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108718A_ABST
    Figure CN120108718A_ABST
Patent Text Reader

Abstract

The invention relates to an immunotherapy prognosis evaluation system based on combined penalty likelihood modeling, and belongs to the technical field of precision medicine. Based on a generalized linear hybrid model, various types of clinical endpoints, such as objective remission rate and disease-progression-free lifetime, are effectively integrated. By introducing the random effect, the model can capture potential correlation between end points, and the statistical effectiveness of analysis is enhanced. For the colinearity between clonality mutation features, key variables are screened through a penalty likelihood method, the influence of multiple colinearity is reduced, and the explanatory property and prediction accuracy of the model are improved. In order to overcome the challenge of limited sample size, cross-subgroup data fusion likelihood modeling is introduced, the strategy not only retains the model specificity of each subgroup, but also improves the adaptability of the model to small sample data and the prediction accuracy through information sharing among the subgroups.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an immunotherapy prognosis evaluation system for joint penalty likelihood modeling, and belongs to the technical field of precision medicine. Background Art

[0002] Immune checkpoint inhibitor (ICI) therapy has fundamentally changed the treatment strategy for patients with advanced cancer by activating the patient's own T cells to attack tumor cells. However, the proportion of patients who benefit from ICI treatment is relatively low, prompting the scientific community to conduct in-depth research on biomarkers that affect efficacy. Tumor mutation burden (TMB) is defined as the number of non-synonymous mutations in the coding region of the tumor cell genome. Its correlation with the efficacy of immunotherapy has been widely studied and recognized. Studies have revealed that due to the new antigens generated by tumor mutations, they can be presented by MHC proteins on the surface of cancer cells, allowing T cells to recognize these new antigens and eliminate target cells in a targeted manner. Therefore, a higher TMB means the generation of more new antigens, thereby increasing the probability of T cell recognition of targets and improving the efficacy of immunotherapy. Retrospective evidence shows that TMB can promote effective anti-tumor immune responses and ultimately enable immunotherapy to produce sustained clinical efficacy. Despite this, TMB has obvious limitations as a predictor of prognosis in immunotherapy. The currently commonly adopted TMB calculation method treats all types of mutations as equivalent, ignoring the differences in the effects of different mutations on immunotherapy. This single-value-based assessment fails to fully capture the complex heterogeneity of tumors, which plays a key role in driving effective anti-tumor immune responses. Therefore, although TMB provides a certain degree of predictive value, as a biomarker, its perfection in predicting the effect of immunotherapy needs to be improved. The current TMB assessment method based on the number of overall cell mutations cannot fully explain the expression of new antigens and their complex mechanism of action in immunotherapy response, making it difficult to provide comprehensive and precise treatment guidance for all ICI indications.

[0003] There are a series of complex computational problems in the analysis of genomic clonal data of immunotherapy cohort patients, which not only challenge traditional data analysis methods, but are also crucial to improving the efficacy of immunotherapy and formulating personalized treatment strategies. The integration of multiple end-point data in immunotherapy clinical practice and the high collinearity of clonal genomic data together constitute the main computational difficulties of this study. The integration of multiple end-point data in immunotherapy clinical practice, the high collinearity of clonal genomic data, and the limited cohort sample size together constitute the main computational difficulties of this study.

[0004] First, ideal clinical trials usually focus on a single efficacy endpoint in order to draw clear and unambiguous statistical conclusions. However, given the diversity of evaluation needs, multiple endpoints are often involved in tumor immunotherapy practice. Traditionally, cancer treatment is mainly based on the anti-tumor cytotoxic effect of the drug, with tumor shrinkage as the main criterion, making the objective response rate (ORR) a key efficacy indicator. However, due to the unique safety and efficacy characteristics of immunotherapy, the focus of clinical evaluation has shifted from ORR to overall survival (OS) or progression-free survival (PFS). Immunotherapy efficacy endpoints are divided into two categories: discrete ORR and continuous time-to-event (TTE, including OS, PFS, etc.). Multi-endpoint data in immunotherapy studies can provide a multi-dimensional perspective for evaluating treatment effects. However, there are often inherent correlations between these endpoints, and ignoring these correlations may lead to misjudgment of treatment effects. For example, some immunotherapies may not significantly improve ORR in the early stage, but can significantly prolong the patient's OS. Relying solely on a single endpoint may not fully reflect the true effect of the treatment. This multidimensional data structure requires us to develop statistical models that can handle multiple related endpoints simultaneously and overcome the limitations of traditional single endpoint analysis.

[0005] Secondly, the analysis of clonal characteristic data of tumors faces the problem of high collinearity. The clonal characteristics of tumors, including the existence of different clonal populations and their differences in gene expression and mutation characteristics, provide key information for understanding tumor development and treatment response. However, there are often strong correlations between these characteristics, which leads to multicollinearity problems when using traditional statistical models, which not only increases the difficulty of model selection and parameter estimation, but also may affect the accuracy and stability of model predictions. Therefore, the analysis of clonal genomic data needs to consider the complex dependencies between these variables and seek advanced statistical methods that can overcome the problem of multicollinearity.

[0006] Finally, the most common major challenge faced by tumor immunotherapy research is the limited sample size. Due to the cost and potential side effects of immunotherapy itself, the number of patients included in each clinical trial or research cohort is usually small. This problem is particularly significant when considering different molecular subtypes or the heterogeneity of treatment response. This limitation will increase the risk of overfitting of the model, which may weaken the model's ability to generalize to new data. At the same time, the limited sample size also restricts the application of complex models, because such models rely on sufficient data to ensure the accuracy of parameter estimation. The observational data covers multiple subgroups, and there may be biological differences between subgroups, which makes the data and their association patterns between different research groups complex and varied. Applying a single regression model to all data may lead to improper model specification and even serious errors. Building a model for each subgroup independently faces the problem that the sample size of a single subgroup is much smaller than the overall sample size, especially when more characteristic variables are involved, which further increases the challenge of parameter estimation. Therefore, it is urgent to develop more flexible, statistically efficient and scalable models to meet the actual needs of immunotherapy genomic mutation clonality research. Summary of the invention

[0007] When analyzing the genomic clonal mutation data of patients in the immunotherapy cohort, the integration of clinical multi-endpoint data is of great significance for enhancing the prognostic prediction effect of immunotherapy. However, the high collinearity between clonal mutation features and the limitation of cohort sample size may challenge the stability and accuracy of the immunotherapy prognostic prediction model.

[0008] Based on the clonal TMB index, the present invention applies random effects in a generalized linear mixed model (GLMM) to jointly model multi-endpoint data, and handles the potential association between clonal mutation features through a penalized likelihood strategy. The introduction of the penalty operator causes the likelihood function to present a non-convex feature, which constitutes a major computational challenge in the optimization process. The joint penalized likelihood modeling strategy proposed in the present invention aims to efficiently integrate clinical multi-endpoint data to accurately predict the prognosis of patients in immunotherapy cohorts based on clonal mutation features. Based on the GLMM framework, the present invention introduces random effects to reveal the potential connection between different clinical endpoints; by embedding penalty terms in the joint likelihood, the collinearity problem between tumor clonal features is solved, and key variables that have a significant impact on treatment response are screened out; by controlling the coefficient differences between specific subgroups, the observed data of different immunotherapy cohorts are effectively integrated to ensure the model specificity of each subgroup and data sharing between subgroups. The present invention enhances the statistical power and interpretability of the model by integrating multiple end point data, dealing with collinearity issues and optimizing model construction under limited sample size conditions, thereby enhancing the efficiency of statistical inference and the scalability of the model, providing more comprehensive and accurate information support for clinical decision-making in tumor immunotherapy.

[0009] In the first technical solution of the present invention, a system for evaluating the prognosis of immunotherapy using joint penalized likelihood modeling comprises:

[0010] A data grouping module is used to divide the samples into K different subgroups according to the source of the patient data;

[0011] A clinical data acquisition module is used to obtain the tumor remission status and survival data information of the sample;

[0012] The sequencing module is used to obtain tumor sequencing information of the sample and perform clonal analysis to obtain the mutation characteristics of each subclone population;

[0013] An evaluation and analysis module is used to solve the joint likelihood function with the penalty term introduced to obtain the maximum likelihood estimation parameter value;

[0014] The evaluation and analysis module uses the status of tumor remission, survival data information and mutation characteristics as input parameters;

[0015] The joint likelihood function with the penalty term includes the joint likelihood function and the penalty term for the fixed effect parameters and / or the penalty term for the cross-subgroup data; the joint likelihood function, the penalty term for the fixed effect parameters and the penalty term for the cross-subgroup data are linearly connected.

[0016] The joint likelihood function is derived from the model representing the objective response rate endpoint and the model representing the time-to-event endpoint through the shared joint random effect ω k,i connect.

[0017] The joint likelihood function is:

[0018]

[0019] Where: f(R k,i ∣ω k,i ; Θ) represents the total likelihood function based on the objective response rate endpoint; f(T k,i ,δ k,i ∣ω k,i ; Θ) represents the total likelihood function based on the time to the end of the event; f(ω k ; Θ) is the density function of the normal distribution; Θ represents the family of parameters that need to be estimated.

[0020] The total likelihood function for the objective response rate endpoint is:

[0021]

[0022] Where: α k,m is a fixed effect coefficient vector of size p×1 for each subclone m, α k =[α k,1 ,...,αk,M ];ω k,i is the random effect corresponding to patient i in subgroup k.

[0023] The total likelihood function based on the time to the event endpoint is:

[0024]

[0025] Where: h(t) is the instantaneous risk of an event occurring within the time interval [t, t+dt) when the survival time reaches time point t; h 0 (t) is the baseline risk; β k,m is the fixed effect coefficient matrix of size p × 1 for each subclone m, β k =[β k,1 ,...,β k,M ].

[0026] X k,i,m is the p-dimensional mutation signature of the mth subclone, and the mutation signature includes tumor mutation burden, average variant allele frequency, and cancer cell fraction.

[0027] The penalty term for the fixed effect parameter is a Lasso or Ridge or Elastic Net penalty term.

[0028] The penalty term for cross-subgroup data is used to control the differences between subgroup-specific coefficient vectors.

[0029] The expression of the penalty term for cross-subgroup data is:

[0030]

[0031] τ k,k′ =1-d(k,k′) / d max

[0032] γ is a hyperparameter used to control the amount of penalty shrinkage; τ k,k′ is a parameter used to adjust the strength of fusion between different subgroups; α k′,m and β k′,m is the fixed effect coefficient corresponding to the objective response rate endpoint and the time-to-event endpoint in group k′; α k,m and β k,m is the fixed effects coefficient corresponding to the objective response rate endpoint based on time to event in group k.

[0033] d(k,k′) is calculated in one of the following ways:

[0034] The first one:

[0035]

[0036] Where: ——The mean of the p-dimensional mutation signature from the m-th subclone of the k-th subgroup; refers to the corresponding feature mean in another subgroup k′;

[0037] Second type:

[0038]

[0039] —The covariance matrix between the mutational signatures of the mth subclone in subgroup k and subgroup k′ can be calculated using the observed data of the cohort patients;

[0040] The third type:

[0041]

[0042] Where: and are the estimated distributions of the mutational signatures of the mth subclone in subgroup k and subgroup k′, respectively; d max is the maximum distance between groups.

[0043] The objective response rate endpoint was expressed as tumor response status, and the time-to-event endpoint was expressed as survival data information.

[0044] The sources of patient data described refer to different clinical trial cohorts.

[0045] The state of tumor remission refers to complete remission, partial remission, disease stabilization and disease progression; the survival data information includes survival time and event occurrence indicators.

[0046] In the above evaluation system, if there are penalty terms for fixed effect parameters and penalty terms for cross-subgroup data at the same time, the joint likelihood function, the penalty terms for fixed effect parameters, and the penalty terms for cross-subgroup data are linearly connected; when there is only a penalty term for fixed effect parameters, it is also linearly connected with the joint likelihood function; when there is only a penalty term for cross-subgroup data, it is also linearly connected with the joint likelihood function.

[0047] In the second technical solution of the present invention, the following is an immunotherapy prognosis evaluation system modeled by a joint likelihood function and a joint penalized likelihood containing only penalty terms constructed for fixed effect parameters:

[0048] A data grouping module is used to divide the samples into K different subgroups according to the source of the patient data;

[0049] A clinical data acquisition module is used to obtain the tumor remission status and survival data information of the sample;

[0050] The sequencing module is used to obtain tumor sequencing information of the sample and perform clonal analysis to obtain the mutation characteristics of each subclone population;

[0051] An evaluation and analysis module is used to solve the joint likelihood function with the penalty term introduced to obtain the maximum likelihood estimation parameter value;

[0052] The evaluation and analysis module uses the status of tumor remission, survival data information and mutation characteristics as input parameters;

[0053] The joint likelihood function with the penalty term includes the joint likelihood function and the penalty term for the fixed effect parameters; the joint likelihood function and the penalty term for the fixed effect parameters are linearly connected.

[0054] In this technical solution, the joint likelihood function used in the evaluation and analysis module and the relevant definitions of the penalty terms for the fixed effect parameters are the same as those in the first technical solution.

[0055] The beneficial effect of the present invention is that a joint penalized likelihood modeling strategy is proposed based on the clonal mutation data of patients in the immunotherapy cohort, which aims to solve the problems caused by multi-endpoint data integration, high collinearity and limited sample size through a unified framework. The strategy is based on a generalized linear mixed model (GLMM), which can not only handle multiple types of clinical endpoints, but also capture the potential associations between different endpoints by introducing random effects. In addition, by introducing a penalty term in the model, the collinearity problem of clonal data can be effectively handled, and key variables that have a significant impact on treatment response can be screened out, thereby improving the interpretability and prediction accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1It is a graph of the model performance analysis results; among which, (A) ROC curve comparison of HAPenalize, Pooled Analysis and Separate Analysis for ORR prediction in the experimental cohort; (B) ROC curve comparison of HAPFusion, Pooled Analysis and Separate Analysis for ORR prediction in the experimental cohort; (C) ROC curve comparison of HAPFusion and Separate Analysis for ORR prediction in each subgroup of the experimental cohort; (D) Kaplan-Meier survival curve comparison based on HAPenalize risk grouping; (E) Kaplan-Meier survival curve comparison based on Pooled Analysis risk grouping.

[0057] Figure 2 It is a system module diagram. DETAILED DESCRIPTION

[0058] The following examples are used to illustrate the regression modeling framework of high-dimensional clonal structural features and clinical multi-endpoints in the patient cohort grouping environment. The statistical framework is used to merge retrospective data of the same cancer type, similar subtype, and different cohorts to promote information sharing between cohorts.

[0059] Specifically, consider K different subgroups (i.e., K patient cohorts), each of which has a specific sample size n k The total sample size is Assume that the patient's genomic mutation information comes from M tumor subclones, each of which corresponds to p features or predictors. The specific symbol definitions are shown below.

[0060] For patient i (i=1,…,n k ), set R k,i As the response variable, it represents the patient's tumor response status. According to the recognized solid tumor response evaluation criteria (RECIST) 1.1, the status of tumor response is usually classified into complete response (CR), partial response (PR), stable disease (SD) and progressive disease (PD). In the model setting, let R k,i =1 represents CR and PR, let R k,i=0 represents SD and PD.

[0061] Survival analysis usually involves the follow-up time before the target event occurs, which is also called survival data. The notable feature of survival data is that it is often affected by censoring. Censoring means that some individuals in the study did not experience an event (such as death, disease recurrence, etc.) during the observation period, but due to loss of follow-up or early termination of the study, the relevant events of all patients cannot be observed. Usually, the setting As the actual survival time of patient i, C k,i is the censoring time, T k,i represents the observed time of an event (such as tumor recurrence, progression, death, etc.), which is defined as If the ith patient is not censored during follow-up, then his or her complete event time can be observed. On the contrary, if the patient is lost during the follow-up period, possibly due to loss of follow-up or other reasons, the At this time, the event occurrence indicator is defined as Where I(·) is the indicative function. Therefore, the observed data of the survival outcome is finally given by (T k,i ,δ k,i )composition.

[0062] Assume X k,i represents an observed covariate matrix of size M×p, where the mth row vector is the p-dimensional mutation signature X from the mth subclone k,i,m . Mutation signatures usually include clonal TMB (i.e., TMB belonging to the clone), the average allele frequency of all mutations in the clone, and the cancer cell fraction (CCF) of the clone. Clonal mutation signatures are calculated using Pyclone software. Generalized linear mixed effects models usually include fixed effects and random effects, and are often used to analyze clinical trial data involving repeated measurements or clustered data. Random effects take into account the different unobservable characteristics of each individual, characterizing individual specificity while depicting the potential association between different endpoints.

[0063] The binary response variable, objective response rate (ORR), was modeled using a generalized linear mixed model (GLMM) with a logit link function:

[0064]

[0065] Where: α k,m ——a fixed effect coefficient vector of size p×1 for each subclone m, α k =[α k,1 ,...,α k,M ];ω k,i——corresponds to the random effect of patient i in subgroup h.

[0066] The Cox proportional hazards model was used to process the survival data, and a generalized linear mixed model (GLMM) with a log link function was used to model the hazard function for patient i at time t:

[0067]

[0068] Where: k(t) is the instantaneous risk of an event occurring within the time interval [t, t+dt) when the survival time reaches time point t; h is 0 (t)——baseline hazard; β k,m ——A fixed effect coefficient matrix of size p×1 for each subclone m, β k =[β k,1 ,...,β k,M ].

[0069] Model (1) representing the objective response rate (ORR) endpoint and model (2) representing the time-to-event (TTE) endpoint were combined by sharing the joint random effect ω k,i Connection. Random effect ω of cohort k k Follows a normal distribution with zero mean ω k,i It represents both the variation in the individual level within the cohort and the inherent correlation between different endpoints of the same individual. k,i , the two sub-models are conditionally independent. Therefore, the joint likelihood function based on the observed data can be written as:

[0070]

[0071] Where: f(R k,i ∣ω k,i ; Θ) represents the total likelihood function based on the ORR endpoint; f(T k,i ,δ k,i ∣ω k,i ; Θ) represents the total likelihood function based on the TTE endpoint; f(ω k ; Θ) is the density function of the normal distribution; Θ represents the family of parameters that need to be estimated. Thus, the joint likelihood function can be rewritten as:

[0072]

[0073] Where: ——Cumulative baseline hazard function.

[0074] The maximum likelihood estimate of the parameter Θ can be obtained by maximizing the joint likelihood function (4). However, there is no analytical solution for the integral of the joint likelihood function (4) with respect to the random effects, so a numerical integration method is needed to obtain an approximate solution, such as the Gauss-Hessian quadrature formula. After obtaining the approximate numerical integral with respect to the random effect ω, the Newton-Raphson algorithm can be used to estimate the parameter value under the maximum likelihood. The algorithm updates the parameter estimate through the first-order derivative (Score) and the second-order derivative (Hessian matrix) of the likelihood function, thereby achieving optimization.

[0075] The biological complexity of tumors leads to cell division and evolution to form multiple subclones with genetic heterogeneity. The mutations accumulated by these subclones at different locations in the genome affect their survival and proliferation, making some subclones dominant in the tumor microenvironment. During the clonal evolution process, the gene mutations of different subclones show correlation, resulting in collinearity problems in tumor clonality data, reflecting the complex genetic relationship and dynamic changes between subclones.

[0076] The present invention adopts a statistical method of introducing penalty terms to deal with complexity. By applying regularization techniques, such as Least Absolute Shrinkage and Selection Operator (LASSO), Ridge Regression, or Elastic Net, constraints are imposed on the input features, which can significantly reduce the impact of multicollinearity. These methods introduce additional penalty terms (e.g., l 1 or 2 norm), limiting the parameter size, thereby reducing estimation uncertainty and improving the generalization ability of the model.

[0077] Taking the joint model (4) as an example, the penalty term is introduced with K = 1 for simplification. The joint modeling of ORR and TTE is as follows:

[0078]

[0079] Lasso adds l to the loss function 1 The norm penalty term (i.e., the sum of the absolute values ​​of the parameters) is used for variable selection and model regularization, so that some fixed effect characteristic coefficients that are not predictive of clinical outcomes can be accurately shrunk to zero. Introducing Lasso penalty for fixed effects in the joint likelihood function, such a penalized likelihood can be expressed as:

[0080]

[0081] Where: ——The logarithmic joint likelihood of the penalty term; λ——A hyperparameter used to control the amount of penalty shrinkage. In specific implementation, by adjusting the constraint strength of the penalty term, a suitable balance can be found between model complexity and goodness of fit.

[0082] In addition, based on l 2 Ridge regression with the norm is used to avoid modeling scenarios where Lasso is not applicable. Ridge regression effectively reduces the high variance problem caused by collinearity by imposing a penalty on the sum of squares of coefficients, thereby avoiding overfitting and improving the stability of model estimation. The Ridge penalty is introduced for the fixed effects in the joint likelihood, and the penalized likelihood can be written as:

[0083]

[0084] In addition, Elastic Net is a l that integrates Lasso 1 Norm and Ridge's l 2 The norm penalty regularization method introduces these two penalty terms at the same time to make up for their respective limitations and combine their advantages. Elastic Net uses the variable selection ability of Lasso and the parameter stability of Ridge to effectively handle the high collinearity problem of data while maintaining the sparsity of the model. The Elastic Net penalty is introduced for the fixed effect in the joint likelihood, and the penalized likelihood can be written as:

[0085]

[0086] Where: r∈[0,1]——l 1 and l 2 The relative weight of the penalty. The penalty parameter λ is related to r (i.e. l 1 and l 2 The selection of the proportion of the maximum likelihood estimate (the ratio of the maximum likelihood estimate) needs to be carefully tuned through methods such as cross-validation. After the penalty term is introduced in the joint likelihood, the derivation of the parameters of the maximum likelihood estimate MLE requires an iterative algorithm, such as coordinate descent or gradient-based optimization methods. When using the Ridge penalty, the gradient descent method or other more efficient numerical optimization methods such as the conjugate gradient method are usually used to iteratively estimate the parameters. The fixed effect coefficients α and β are subject to Ridge penalty. After each iteration, check whether the parameter update amount or the change of the objective function is lower than the predetermined threshold. If the convergence condition is met (for example, the parameter change is less than a small positive number ∈ or the maximum number of iterations is reached), the iteration is stopped.

[0087] The above is the expression under the condition of K=1. Under other queue conditions such as K=2, 3, etc., the corresponding fixed effect parameters can also be penalized by L1, L2 or Elastic Net.

[0088] When analyzing clinical retrospective data on tumor immunotherapy, it should be ideally assumed that all observed samples are consistent and follow the same statistical model. Reliable inference results can be provided through sufficient data volume. However, due to the high cost and potential side effects of immunotherapy, researchers are often limited by problems such as heterogeneity of cohort data and small sample size. The samples in these studies often involve multiple subgroups, each of which may be based on different regression models. In biomedical problems, sample sets representing disease subtypes, etc. may differ in basic biology, and the relationship between observed features and responses of interest is also different. Therefore, the data structure and correlation patterns between different study subgroups may show high complexity and variability. Applying a single statistical model to the overall data may lead to model inadaptability, resulting in significant estimation errors. A fusion modeling strategy for cross-subgroup data is further proposed, which controls the coefficient differences between subgroups in the likelihood function through a penalty operator to promote information sharing between subgroups.

[0089] Taking the joint likelihood (7) with Ridge penalty as an example, the fusion likelihood model of cross-subgroup data is proposed as follows:

[0090]

[0091] Where: λ, γ, τ are hyperparameters used to control the amount of penalty shrinkage; the last term represents the fusion penalty between subgroups, which takes into account the differences between subgroup-specific coefficient vectors; the control parameter τ k,k′ Used to adjust the strength of fusion between different subgroups. k′,m and β k′,m is the fixed effect coefficient corresponding to ORR and TTE endpoint in group k′; α k,m and β k,m is the fixed effect coefficient corresponding to ORR and TTE endpoint in group k.

[0092] By default, all fusion parameters τ can be set to 1 to achieve non-weighted fusion. Depending on the differences between subgroups, the fusion parameter τ can be set to different values ​​to achieve weighted fusion and introduce differences. In weighted fusion, the parameter τ can be set by cross-validation on observed data in a similar way to λ and γ, but this may be cumbersome in practice. As an alternative, the present invention proposes a more interpretable method to set the fusion parameter τ by a feature-based distance metric d(k, k′) k,k′ , thus allowing for tighter integration between subgroups with similar characteristics. Different distance formulas are given below.

[0093] In order to measure the differences in clonal mutation characteristics between different subgroups, the Euclidean distance can be used to measure the differences between the mean characteristics of samples in each subgroup:

[0094]

[0095] Where: ——The mean of the p-dimensional mutation signature from the m-th subclone in the k-th subgroup (data have been normalized). refers to the corresponding feature mean in another subgroup k′.

[0096] Furthermore, considering the potential correlation between clonal features, the Mahalanobis distance considering the covariance matrix V is used for calculation:

[0097]

[0098] Where: ——Subgroup k and subgroup k ′ The covariance matrix between the mth subclonal mutation signatures in can be calculated using the observed data of the cohort patients.

[0099] The symmetric KL divergence (Kullback-Leibler divergence) is used to measure the distance between the feature differences between different subgroups. KL divergence, also known as relative entropy, is a method to measure the difference between two probability distributions. KL divergence is asymmetric and is used to measure the cost of simulating p from q given two probability distributions p and q.

[0100] The solution process of the fused maximum likelihood inference of cross-subgroup data is as follows:

[0101]

[0102]

[0103] The complete solution steps for collinearity penalized likelihood modeling are as follows:

[0104]

[0105]

[0106] Clinical cohort patient data analysis

[0107] Individualized strategies for tumor treatment are gradually gaining widespread attention, among which the study of clonal heterogeneity is of key significance for revealing the biological behavior and therapeutic response of tumors. This clonal heterogeneity not only records the development process of tumors, but also may affect the patient's response to immunotherapy.

[0108] Therefore, the present invention collects and analyzes patient data from multiple immunotherapy clinical cohorts, covering multiple cancer types. The present invention combines multiple clinical endpoints, such as objective response rate (ORR) and progression-free survival (PFS), to evaluate the impact of clonal mutation characteristics of the patient's genome on the prognosis of its treatment. Table 1 lists the information summary of the clinical cohort patients used in the present invention.

[0109] The present invention retrospectively analyzed samples from 238 patients who received ICI monotherapy. Participants included 64 patients with recurrent / metastatic nasopharyngeal carcinoma (NPC) and 73 patients with non-small cell lung cancer (NSCLC) from Sun Yat-sen University Cancer Center (SYUCC), and 75 patients with non-small cell lung cancer and 26 patients with melanoma from the Second Affiliated Hospital of Xi'an Jiaotong University who were sequenced at Geneplus Beijing Institute.

[0110] To further validate the clinical application of the model, we collected a total of 830 patient samples from publicly published literature, including their genomic mutation profiles and clinical pathological data, thereby expanding the dataset. According to the research source and disease type, the validation group included 467 patients with non-small cell lung cancer divided into 4 subgroups and 363 patients with melanoma divided into 4 subgroups.

[0111] Table 1 Summary of clinical cohort patient information

[0112]

[0113] For high-throughput genome sequencing data, the present invention uses the PyClone algorithm to accurately identify and quantify the clonal mutation characteristics in each patient's tumor, aiming to reveal the potential association between these clonal markers and the effect of immunotherapy. The specific bioinformatics data analysis process is as follows:

[0114] 1. Data preprocessing

[0115] Raw sequencing data (FASTQ format) of each patient sample is collected, and each dataset undergoes rigorous integrity and potential corruption checks to ensure data integrity.

[0116] Quality control (QC) was performed using the FastQC tool to assess the sequencing quality of each sample. Based on the FastQC report, possible problems in the sequencing data were checked, such as too low sequencing quality, too much sequencing contamination, or too many adapter sequences. Data cleaning included quality filtering and adapter sequence removal using Cutadapt. Low-quality reads, short reads, and duplicate sequences were removed to ensure the accuracy of subsequent analysis.

[0117] 2. Data Analysis

[0118] Alignment

[0119] Use the BWA tool to align the cleaned reads with the reference genome to ensure that the alignment results are accurate. Generate a BAM format alignment file for subsequent analysis.

[0120] Variant Calling

[0121] Mutect2 was used to detect single nucleotide variants (SNVs) and small insertions / deletions (Indels) in tumor samples. CNVkit was used for copy number variation (CNV) analysis. After detection, potential false positives were filtered out to retain high-quality variant data.

[0122] Clonal Analysis

[0123] Prepare necessary input files for PyClone software, including variant data, variant allele frequencies, and sample depth information. Run PyClone analysis to infer clonal populations and their relative proportions in tumor samples.

[0124] The clonal TMB, average variant allele frequency, and cancer cell fraction (CCF) characteristics of each subclone obtained above were input into the model as input variables, and the model was solved and parameters were optimized.

[0125] This workflow integrates advanced bioinformatics tools and methods to enable robust analysis of clonal mutations in clinical cohorts, ensuring high fidelity in characterizing tumor heterogeneity.

[0126] Based on the clinical endpoint information and genomic clonal mutation characteristics of the collected patients, the present invention compares the proposed collinearity penalty likelihood model (HAPenalize) and the fusion likelihood model (HAPFusion) with the traditional methods, pooled analysis and separate analysis to evaluate their superiority in clinical prognosis prediction. In the collinearity penalty likelihood model (HAPenalize), the Ridge regularization strategy is used in the penalty term, that is, the modeling method of joint likelihood (7) is adopted, and the difference between different subgroups is measured by KL divergence, that is, d(k,k′) is calculated using formula (11). Pooled analysis (Pooled Analysis) and separate analysis (Separate Analysis) are used as comparison baselines, and traditional Logistic regression is used to predict the clinical ORR endpoint, and traditional Cox proportional hazard regression is used to group the clinical TTE endpoint by risk, and no penalty term is included in the regression model. The set analysis method ignores the differences between subgroups of patient data during the regression modeling process, merges all subgroup data into a single data set, and only sets up a common regression model without incorporating the two penalty terms; the individual analysis method treats each subgroup as an independent data set during the regression modeling process, sets up a separate regression model for each subgroup of patients, and does not consider data sharing between subgroups.

[0127] Specifically, Figure 1 The advantages of the present invention are intuitively demonstrated by depicting the receiver operating characteristic curves (ROC) of the three methods of collinearity penalty likelihood model (HAPenalize), fusion likelihood model (HAPFusion), pooled analysis and separate analysis in predicting the ORR endpoint. The AUC value of HAPenalize for ORR prediction is as high as 0.67, and the AUC value of HAPFusion for ORR prediction is as high as 0.69 ( Figure 1 B), surpassing the 0.58 of the combined analysis method and the 0.65 of the individual analysis method. This data not only proves the high accuracy and reliability of the present invention in predicting the prognosis of tumor remission in patients, but also reveals that its ROC curve is closer to the ideal state - that is, the perfect combination of low false positive rate and high true positive rate. This advantage shows that the present invention can enhance the statistical power of the model and achieve more accurate predictions by integrating multiple clinical endpoints and implementing penalty constraints on collinear genomic features. Figure 1The C of the method proposed in the present invention confirms the robustness of the method proposed in the present invention. In the subgroup data composed of different genes and clinical environments, the predictive ability of the HAPFusion proposed in the present invention exceeds the separate analysis method (Separate Analysis). The specific data are as follows: in the Geneplus_Lung subgroup, the AUC of HAPFusion reached 0.62, which is higher than the 0.50 of the separate analysis method; in the Geneplus_Mel subgroup, the AUC of HAPFusion was 0.76, compared with 0.61 for the separate analysis method; in the SYUCC_Lung subgroup, the AUC of HAPFusion was 0.67, while the separate analysis method was 0.64; finally, in the SYUCC_NPC subgroup, the AUC of HAPFusion reached 0.70, far exceeding the 0.55 of the separate analysis method. Secondly, in the separate analysis method, due to insufficient sample size in each subgroup, the problem of model convergence occurred when the Cox proportional hazard regression was used to independently analyze each subgroup. Therefore, in the prognostic analysis of TTE endpoints, the HAPenalize method was compared with the set analysis method. Figure 1 D shows the Kaplan-Meier survival curve of progression-free survival (PFS) after calculating the patient risk based on HAPenalize and dividing the patients into high-risk group and low-risk group according to the median risk. Figure 1 E shows the PFS survival curve after the patients were grouped based on the Cox model in the set analysis method. The comparison of the survival curves of the two methods clearly shows that the present invention has significant advantages in distinguishing the progression-free survival (PFS) of high-risk and low-risk patients. Specifically, the Log-rank test p value of the HAPenalize method and the HAPFusion method is less than 0.0001, which is lower than the traditional set analysis method (p=0.0001) that does not include a penalty term, which fully demonstrates the excellent performance of the present invention in patient survival prognosis stratification.

[0128] Table 2 lists in detail the predictive performance indicators for the ORR endpoint, further verifying the high predictive accuracy of the present invention. In the SYUCC and Geneplus experimental cohorts, the logarithmic loss of HAPenalize for predicting the ORR endpoint was only 0.5262, and the accuracy was as high as 76.48%, which was significantly better than the two traditional methods. The latter had logarithmic losses of 0.8349 and 0.7543 when predicting the ORR endpoint, and the accuracy was only 52.86% and 53.16%, respectively, showing higher losses and poorer data fit. Table 2 Comparison of predictive performance evaluation of HAPenalize, HAPFusion, Pooled Analysis, and Separate Analysis

[0129]

[0130] In the non-small cell lung cancer validation cohort, HAPenalize had an AUC value of 0.69 in ORR prediction, while HAPFusion had an AUC value of 0.71 in ORR, which was significantly better than the 0.66 of the pooled analysis method without penalty and the 0.65 of the single analysis method. This result not only shows that HAPenalize has a high accuracy in predicting tumor remission prognosis, but also further consolidates its application potential in the field of non-small cell lung cancer.

[0131] There are significant differences in progression-free survival (PFS) between high-risk and low-risk groups calculated by HAPenalize and HAPFusion. Under the HAPenalize method, the progression-free survival of patients in the high-risk group was significantly lower than that in the low-risk group. The p-value of the log-rank test was extremely low (p=0.0009), while that of HAPFusion was p=0.0003. This result fully supports the clinical value of HAPenalize and HAPFusion in risk stratification. In contrast, although the Kaplan-Meier survival curve obtained by using the Cox model to calculate the risk grouping in the set analysis method also reflects the survival disadvantage of patients in the high-risk group, the p-value of the log-rank test has increased, which further highlights the advantages of HAPenalize and HAPFusion in predicting patient survival prognosis.

[0132] Finally, the performance comparison data in Table 2 further validates the prediction performance of HAPenalize in non-small cell lung cancer. The logarithmic loss of HAPenalize is only 0.5961, and the accuracy is 70.19%, while the logarithmic loss of HAPFusion is only 0.5741, and the accuracy is 71.80%, which is significantly better than the traditional method. These results show that HAPenalize has made significant breakthroughs in improving prediction accuracy and clinical applicability, providing important support for the precision treatment of patients with non-small cell lung cancer.

[0133] In the analysis of the melanoma validation cohort, the AUC value of HAPenalize was as high as 0.75 and the AUC value of HAPFusion was as high as 0.77, which were significantly better than the 0.65 of the combined analysis method and 0.58 of the individual analysis method, highlighting the high accuracy and reliability of HAPenalize in predicting melanoma treatment response.

[0134] In addition, in the analysis of the survival difference between the high-risk group and the low-risk group under the HAPenalize and HAPFusion methods, the progression-free survival of the high-risk group was significantly worse than that of the low-risk group, and the p-value of the log-rank test was extremely low, which further proved the clinical value of HAPenalize in risk stratification. In contrast, the Kaplan-Meier survival curve generated by the set analysis method showed a relatively weak difference between the groups, further highlighting the advantages of HAPenalize. The robustness and wide applicability of HAPFusion were further verified by comparing the ROC curves of HAPFusion with those of separate analysis methods in four melanoma subgroups (i.e., Liu et al. Nat Medicine, Riaz et al. Cell, Roh et al. Sci Transl Med, and Van Allen et al. Science). In these subgroups, HAPFusion performed very well, with AUC values ​​of 0.65 vs. 0.63, 0.66 vs. 0.56, 0.98 vs. 0.93, and 0.72 vs. 0.69, respectively. These data strongly demonstrate the superior performance of HAPFusion in different clinical contexts.

[0135] The performance comparison data in Table 2 also further verifies the excellent prediction performance of HAPenalize in melanoma. Its lower logarithmic loss (0.5898) and higher accuracy (66.39%), HAPFusion logarithmic loss (0.5458) and accuracy (69.69%) all show that HAPenalize and HAPFusion have achieved good results in improving prediction accuracy.

Claims

1. An immunotherapy prognosis evaluation system based on joint penalized likelihood modeling, characterized in that: include: A data grouping module is used to divide the samples into K different subgroups according to the source of the patient data; A clinical data acquisition module is used to obtain the tumor remission status and survival data information of the sample; The sequencing module is used to obtain tumor sequencing information of the sample and perform clonal analysis to obtain the mutation characteristics of each subclone population; An evaluation and analysis module is used to solve the joint likelihood function with the penalty term introduced, obtain the maximum likelihood estimation parameter value, and evaluate the prognosis of immunotherapy; The evaluation and analysis module uses the status of tumor remission, survival data information and mutation characteristics as input parameters; The joint likelihood function with the penalty term includes the joint likelihood function, the penalty term for the fixed effect parameters, and the penalty term for the cross-subgroup data; the joint likelihood function, the penalty term for the fixed effect parameters, and the penalty term for the cross-subgroup data are linearly connected.

2. The immunotherapy prognosis evaluation system of joint penalized likelihood modeling according to claim 1, characterized in that: The joint likelihood function is derived from the model representing the objective response rate endpoint and the model representing the time-to-event endpoint through the shared joint random effect ω k,i The joint likelihood function is: Where: f(R k,i |ω k,i ; Θ) represents the total likelihood function based on the objective response rate endpoint; f(T k,i , δ k,i |ω k,i ; Θ) represents the total likelihood function based on the time to the end of the event; f(ω k ; Θ) is the density function of the normal distribution; Θ represents the family of parameters that need to be estimated.

3. The immunotherapy prognosis evaluation system of joint penalized likelihood modeling according to claim 2, characterized in that: The total likelihood function for the objective response rate endpoint is: Where: α k,m is a fixed effect coefficient vector of size p×1 for each subclone m, α k =[α k,1 , …, α k,M ];ω k,i is the random effect corresponding to patient i in subgroup k.

4. The immunotherapy prognosis evaluation system of joint penalized likelihood modeling according to claim 2, characterized in that: The total likelihood function based on the time to the event endpoint is: Where: h(t) is the instantaneous risk of an event occurring in the time interval [t, t+dt) when the survival time reaches time point t; h0(t) is the baseline risk; βk,m is the fixed effect coefficient matrix of size p×1 for each subclone m, βk=[β k,1 , ..., β k,M ]; X k,i,m is the p-dimensional mutation signature of the mth subclone, and the mutation signature includes tumor mutation burden, average variant allele frequency, and cancer cell fraction.

5. The immunotherapy prognosis evaluation system of joint penalized likelihood modeling according to claim 1, characterized in that: The penalty term for the fixed effect parameter is a Lasso or Ridge or ElasticNet penalty term.

6. The immunotherapy prognosis evaluation system of joint penalized likelihood modeling according to claim 1, characterized in that: The penalty term for cross-subgroup data is used to control the difference between subgroup-specific coefficient vectors; the expression of the penalty term for cross-subgroup data is: t k,k′ =1-d(k,k′) / d max θ is a hyperparameter used to control the amount of penalty shrinkage; τ k,k′ is a parameter used to adjust the strength of fusion between different subgroups; α k′,m and β k′,m is the fixed effect coefficient corresponding to the objective response rate endpoint and the time-to-event endpoint in group k′; α k,m and β k,m is the fixed effects coefficient corresponding to the objective response rate endpoint based on time to event in group k.

7. The immunotherapy prognosis evaluation system of joint penalized likelihood modeling according to claim 6, characterized in that: d(k, k′) is calculated in one of the following ways: The first one: Where: ——The mean of the p-dimensional mutation signature from the m-th subclone of the k-th subgroup; refers to the corresponding feature mean in another subgroup k′; Second type: —The covariance matrix between the mutational signatures of the mth subclone in subgroup k and subgroup k′ can be calculated using the observed data of the cohort patients; The third type: Where: and are the estimated distributions of the mutational signatures of the mth subclone in subgroup k and subgroup k′, respectively; d max is the maximum distance between groups.

8. The immunotherapy prognosis evaluation system of joint penalized likelihood modeling according to claim 1, characterized in that: The objective response rate endpoint was expressed as tumor response status, and the time-to-event endpoint was expressed as survival data information.

9. The immunotherapy prognosis evaluation system of joint penalized likelihood modeling according to claim 1, characterized in that: The sources of patient data described refer to different clinical trial cohorts.

10. The immunotherapy prognosis evaluation system of joint penalized likelihood modeling according to claim 1, characterized in that: The state of tumor remission refers to complete remission, partial remission, disease stabilization and disease progression; the survival data information includes survival time and event occurrence indicators.

Citation Information

Patent Citations

  • Spectrometer based upon tunable on-chip laser and spectrum measurement method

    KR1020210049660A

  • Diagnostic markers predictive of outcomes in colorectal cancer treatment and progression and methods of use thereof

    US20090275057A1

  • Methods of detecting somatic and germline variants in impure tumors

    US20190362808A1

  • Machine-learning with respect to multi-state model of an illness

    US20220293270A1

  • Disease prognosis prediction system based on deep semi-supervised multi-task learning survival analysis

    WO2021203796A1

Cited By

  • Hypocephalus correlation analysis method and system based on multi-dimensional data

    CN121393933A