Non-small cell lung cancer immunotherapy risk assessment method and system and storage medium
By constructing an optimal strategy tree model and a dynamic monitoring mechanism, the problem of insufficient individualized strategies in the immunotherapy of non-small cell lung cancer was solved, which improved the precision and safety of personalized treatment, discovered potential biomarkers, and provided data-driven hypotheses for clinical research.
Patent Information
- Application Number
- CN202511399938.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-01-09
AI Technical Summary
Current immunotherapy regimens for non-small cell lung cancer lack individualized considerations, making it difficult to accurately match treatment strategies and adjust them in a timely manner. This may lead to undertreatment or overtreatment, and there is a lack of effective dynamic monitoring methods.
By constructing an optimal strategy tree model, subgrouping is performed based on patients' baseline data and dynamic trajectory data to generate personalized treatment strategies. Deviation is monitored in real time during treatment and dynamically adjusted. Treatment plans are optimized by combining counterfactual prediction models and multi-objective clinical utility functions.
It enables personalized treatment strategy recommendations, improves the accuracy and safety of the treatment process, reduces the risk of treatment delays and overtreatment, enhances resource utilization efficiency, and discovers potential biomarkers, providing data-driven hypotheses for clinical research.
Smart Images

Figure CN121306531A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical artificial intelligence technology, specifically to a method, system, and storage medium for risk assessment of immunotherapy for non-small cell lung cancer. Background Technology
[0002] Non-small cell lung cancer (NSCLC) is one of the most common and deadliest malignant tumors worldwide. In recent years, the application of immune checkpoint inhibitors (ICIs), represented by inhibitors of programmed death receptor-1 (PD-1) and its ligand (PD-L1), has significantly improved the survival prognosis of some patients with advanced NSCLC and has become one of the standard treatment methods in this field.
[0003] However, clinical practice shows that there is significant individual heterogeneity in patient response to immune checkpoint inhibitor therapy. Some patients experience long-term and significant survival benefits, while others respond poorly or even experience hyperprogression. Furthermore, the occurrence of immune-related adverse events (irAEs) introduces uncertainty and risk into treatment. Currently, clinical practice relies primarily on a few biomarkers, such as PD-L1 expression levels, to guide treatment decisions, but their predictive accuracy is limited, leading to the continued use of standardized, fixed-cycle treatment strategies.
[0004] This standardized treatment approach struggles to precisely match each patient's complex tumor microenvironment and immune status, potentially leading to undertreatment or overtreatment. More critically, there is a lack of an effective, prospective assessment system to dynamically monitor individualized patient responses after treatment begins. Current efficacy assessments rely primarily on regular imaging examinations, which are often lagging and fail to identify patients whose response trajectories have deviated from expectations early in treatment. Therefore, by the time treatment is deemed ineffective or severe toxicity occurs, patients may have already incurred unnecessary treatment burdens and risks.
[0005] Therefore, this invention proposes a method, system, and storage medium for risk assessment of immunotherapy for non-small cell lung cancer to address the shortcomings of existing technologies. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a method, system, and storage medium for risk assessment of immunotherapy for non-small cell lung cancer. This solves the problem that immunotherapy regimens for non-small cell lung cancer often adopt a standardized model with fixed cycles, which lacks sufficient consideration of individual patient heterogeneity and cannot make timely and precise strategy adjustments based on the patient's actual dynamic response during treatment. This may lead to some patients receiving insufficient or excessive treatment, affecting the final clinical benefit.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for risk assessment of immunotherapy for non-small cell lung cancer, comprising the following steps:
[0008] Obtain a historical patient dataset containing baseline data and dynamic trajectory data of multiple historical patients;
[0009] An optimal strategy tree model is constructed based on historical patient datasets. The optimal strategy tree model divides patients into multiple subgroups and determines an initial optimal immunotherapy course recommendation strategy for each subgroup.
[0010] Baseline data of new patients are obtained, and the subgroup to which the new patient belongs and the recommended initial optimal immunotherapy course strategy corresponding to the subgroup are determined according to the optimal strategy tree model.
[0011] During the process of treating new patients according to the recommended strategy of the initial optimal immunotherapy course, dynamic trajectory data of new patients are continuously acquired to form individual dynamic feature vectors.
[0012] Calculate the deviation between the individual dynamic feature vector and the subgroup average dynamic response trajectory vector of the subgroup to which the new patient belongs;
[0013] When the deviation exceeds the subgroup's preset deviation threshold, the updated treatment adjustment suggestions are output by combining the new patient's baseline data and the new patient's dynamic trajectory data.
[0014] Preferably, a counterfactual prediction model is trained based on the historical patient dataset to generate a prediction utility matrix for each historical patient under different preset treatment strategies;
[0015] Using the predicted utility matrix as input, the optimal policy tree model is constructed through a recursive segmentation process aimed at maximizing the total expected utility of the group.
[0016] Preferably, the counterfactual prediction model is trained based on a multi-objective clinical utility function, which comprehensively quantifies the clinical value of different treatment strategies by weighting and combining expected survival benefits and adverse event risks.
[0017] Preferably, the deviation is calculated by measuring the weighted Euclidean distance between the individual dynamic feature vector and the subgroup average dynamic response trajectory vector.
[0018] Preferably, the subgroup average dynamic response trajectory vector is predetermined during the model training phase by calculating the average of the dynamic characteristics of all patients belonging to the same subgroup in the historical patient dataset at each assessment time point.
[0019] Preferably, the baseline data includes the patient's age, the longest diameter of the tumor, and the ratio of derived neutrophils to lymphocytes; the dynamic trajectory data includes the depth of tumor shrinkage and records of immune-related adverse events.
[0020] Preferably, the method further includes:
[0021] After the optimal strategy tree model completes the subgrouping of patients, a correlation analysis is performed within each subgroup on predefined secondary biomarkers and the actual clinical outcomes of patients in order to backward mine secondary biomarkers related to the efficacy or toxicity of the subgroup.
[0022] Preferably, the method further includes:
[0023] Generate and output a comprehensive evaluation report, which includes a recommended strategy for the initial optimal immunotherapy course and a risk-benefit balance analysis diagram comparing the expected survival benefits and adverse event risks under different treatment strategies.
[0024] This invention also provides a risk assessment system for immunotherapy of non-small cell lung cancer, the system comprising:
[0025] The data acquisition and integration module is used to acquire historical patient datasets containing baseline data and dynamic trajectory data of multiple historical patients, as well as to acquire baseline data and dynamic trajectory data of new patients.
[0026] The model training and optimization module is used to construct an optimal strategy tree model based on the historical patient dataset. The optimal strategy tree model divides patients into multiple subgroups and determines an initial optimal immunotherapy course recommendation strategy for each subgroup.
[0027] The subgrouping and strategy generation module is used to determine the subgroup to which a new patient belongs and the recommended initial optimal immunotherapy course strategy corresponding to the subgroup based on the optimal strategy tree model.
[0028] The dynamic monitoring and adjustment module is used to continuously acquire the dynamic trajectory data of new patients to form individual dynamic feature vectors during the treatment of new patients according to the initial optimal immunotherapy course recommendation strategy, calculate the deviation between the individual dynamic feature vector and the subgroup average dynamic response trajectory vector of the subgroup to which the new patient belongs, and output updated treatment adjustment suggestions when the deviation exceeds the preset deviation threshold for the subgroup, combining the baseline data of the new patient and the dynamic trajectory data of the new patient.
[0029] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a method for risk assessment of immunotherapy for non-small cell lung cancer.
[0030] This invention provides a method, system, and storage medium for risk assessment of immunotherapy for non-small cell lung cancer.
[0031] It has the following beneficial effects:
[0032] 1. This invention uses an optimal strategy tree model to subgroup patients and generate individualized initial treatment recommendations. Then, during the treatment process, the actual response trajectory of the patient is quantitatively compared with the average response trajectory of the specific subgroup to which the patient belongs. This setting makes the benchmark for dynamic monitoring no longer a general group of data, but a group of patients with similar baseline characteristics, thereby improving the accuracy of risk assessment and the targeting of subsequent treatment strategy adjustments.
[0033] 2. This invention establishes a closed-loop adaptive adjustment mechanism by calculating the deviation of individual responses from subgroup benchmarks and setting trigger thresholds. This mechanism can proactively identify individual patients whose treatment responses deviate from expectations, thereby enabling timely intervention and dynamic optimization of treatment strategies. It avoids treatment delays or overtreatment that may occur due to waiting for clear clinical endpoint events, effectively improving the safety and resource utilization efficiency of the entire treatment process.
[0034] 3. This invention not only serves the clinical decision-making of individual patients, but also provides a framework for knowledge discovery. By conducting independent biomarker association analysis on patient subgroups with clinical homogeneity divided by the model, it is possible to reverse-engineer potential efficacy or toxicity predictive biomarkers for specific populations from real-world data. This provides data-driven and valuable research hypotheses for subsequent clinical research and promotes the long-term development of the field of precision immunotherapy. Attached Figure Description
[0035] Figure 1 This is a structural block diagram of the non-small cell lung cancer immunotherapy risk assessment system of the present invention;
[0036] Figure 2 This is a flowchart of the risk assessment method for non-small cell lung cancer immunotherapy according to the present invention;
[0037] Figure 3 This is a schematic diagram of the optimal strategy tree structure of the present invention;
[0038] Figure 4 This is an example diagram of a visualization interface for the risks and benefits of this invention.
[0039] The modules include: 110. Data acquisition and integration module; 120. Model training and optimization module; 130. Subgroup division and strategy generation module; 140. Risk assessment and visualization module; 150. Dynamic monitoring and adjustment module; and 160. Knowledge discovery module. Detailed Implementation
[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0041] See attached document Figure 1 , Figure 1 This is a structural block diagram of a risk assessment system for non-small cell lung cancer immunotherapy according to an embodiment of the present invention. The present invention provides a method and system for risk assessment of non-small cell lung cancer immunotherapy. This system is used to generate personalized immunotherapy cycle recommendations based on the patient's multidimensional data, and to dynamically assess and adjust the risks during the treatment process.
[0042] The execution environment of the methods and systems of this invention may include, but is not limited to, servers, personal computers, workstations, or virtual computing resources based on cloud computing platforms. These computing devices typically include at least one processor, memory, communication interface, and input / output devices for human-computer interaction. The memory stores computer program instructions, and when the processor executes these instructions, it implements the risk assessment method of this invention. The functions of each step of the method may be implemented by one or more software modules running on a general operating system (e.g., Linux, Windows), and the collected and generated patient data may be structured and stored and managed using a database system (e.g., a relational database or a non-relational database).
[0043] See attached document Figure 1 The non-small cell lung cancer immunotherapy risk assessment system provided in this embodiment of the invention may include: a data acquisition and integration module 110, a model training and optimization module 120, a subgroup division and strategy generation module 130, a risk assessment and visualization module 140, a dynamic monitoring and adjustment module 150, and a knowledge discovery module 160.
[0044] The data acquisition and integration module 110 is used to acquire baseline data and dynamic trajectory data of patients during treatment from medical information systems, laboratory testing systems, or manual input interfaces. The baseline data includes pre-treatment indicators such as patient age, longest tumor diameter (LD), and derived neutrophil-to-lymphocyte ratio (dNLR). The dynamic trajectory data includes tumor shrinkage depth during treatment, records of immune-related adverse events, and dynamic changes in key blood indicators. This module is also responsible for cleaning, standardizing, and time-aligning the acquired raw data to generate a structured dataset that can be used by subsequent models.
[0045] The model training and optimization module 120 is responsible for building and training the core artificial intelligence model required by the present invention. This module receives the historical patient dataset processed by the data acquisition and integration module 110, builds and trains the counterfactual prediction model and the optimal strategy tree model. One specific function of this module is to train the counterfactual prediction model based on a predefined multi-objective clinical utility function to generate a prediction utility matrix for each patient under different preset treatment strategies.
[0046] The subgrouping and strategy generation module 130 uses the optimal strategy tree model trained by the model training and optimization module 120 to accurately subgroup newly input patients. When the baseline data of a new patient is received, the module determines the characteristic subgroup to which the patient belongs through model calculation, and generates an initial optimal immunotherapy course recommendation strategy based on this classification.
[0047] The risk assessment and visualization module 140 performs risk-benefit analysis and visualization of individualized treatment strategies. This module receives the recommended strategies output by the subgroup division and strategy generation module 130, and further calculates the expected survival benefits and potential adverse event risks when selecting this strategy and other alternative strategies. Finally, it presents the complex assessment results to clinical users in an intuitive graphical interface such as decision tree path diagram and risk-benefit balance diagram.
[0048] The dynamic monitoring and adjustment module 150 is used to adaptively adjust the treatment strategy during the actual treatment of the patient. This module continuously receives the patient's dynamic trajectory data, calculates the deviation between the individual response trajectory and the average response trajectory of the subgroup. When the deviation exceeds a preset threshold, the module is activated, calls the dynamic adjustment model, and outputs updated treatment adjustment suggestions based on the patient's latest data.
[0049] The knowledge discovery module 160 is used to perform exploratory data mining within the divided patient subgroups to discover new biological associations. This module is activated after the subgroup division is completed and performs independent association analysis on the patient data within each subgroup. It aims to reverse-engineer secondary biomarkers that are significantly related to the efficacy or toxicity of the subgroup, in addition to the core division features, and output the discovered strong associations as research hypotheses.
[0050] See attached document Figure 2 , Figure 2 This is a flowchart of a method for risk assessment of immunotherapy for non-small cell lung cancer according to an embodiment of the present invention. The method may specifically include the following steps:
[0051] S201: Perform multidimensional data acquisition and preprocessing. This step acquires the patient data required for model building and individualized assessment. The acquired dataset consists of two main parts: baseline data before the patient begins immunotherapy, and dynamic trajectory data collected at different time points after the start of treatment. The baseline data specifically includes the patient's age, the longest diameter of the tumor (LD) measured by imaging methods such as computed tomography (CT), and a blood indicator reflecting the body's immune inflammatory status, namely the derived neutrophil to lymphocyte ratio (dNLR). The dynamic trajectory data specifically includes the tumor shrinkage depth (DoR) assessed at the end of each treatment cycle, the occurrence and severity of immune-related adverse events (irAEs), and the dynamic changes in the aforementioned blood indicators. The acquired raw data is cleaned, missing value processing is performed, and data standardization or normalization operations are conducted to form a structured dataset for use in subsequent steps.
[0052] S202: Offline construction and training of the core AI model. This step, based on a training dataset containing a large amount of complete historical patient data, constructs and optimizes the core model for risk assessment. First, a multi-objective clinical utility function U is defined to comprehensively quantify the expected clinical value of a specific treatment strategy d for a specific patient. Then, based on this utility function, a counterfactual prediction model, such as a counterfactual random forest, is trained to generate a predicted utility matrix for each patient under all preset treatment strategies (e.g., short, medium, and long courses). Finally, using this utility matrix as input, an Optimal Policy Tree (OPST) model is trained. This model recursively partitions the baseline feature space to form a series of decision rules, ultimately dividing the patient population into multiple subgroups with different clinical characteristics and determining an optimal treatment strategy for each subgroup that maximizes the average expected utility of its members.
[0053] S203: Perform individualized initial assessment for new patients upon receiving baseline data X from a new patient. i Then, the data is input into the optimal policy tree model that has been trained in step S202; the model, based on its internal decision rules, processes the patient's feature vector X. i The decision tree is propagated along its path until a leaf node is reached, at which point the subgroup S to which the patient belongs is determined. k Meanwhile, the optimal treatment strategy associated with this subgroup This was determined as the initial recommended treatment for the patient;
[0054] S204: Triggering of dynamic monitoring and reassessment during treatment. After the patient begins treatment according to the initial recommended plan generated in step S203, the system enters the dynamic monitoring phase; during this phase, the system continuously acquires the patient's dynamic trajectory data Z during the treatment process.i (t); The system compares the patient's actual response trajectory with the subgroup S to which the patient belongs. k Average response trajectory The two are compared and the degree of deviation between them is quantified; when the deviation exceeds a specific threshold θ set for that subgroup... k When this happens, the reassessment mechanism is triggered;
[0055] S205: Adaptive adjustment of the treatment strategy is executed. After the reassessment mechanism in step S204 is triggered, the system calls a pre-trained dynamic adjustment model; this model receives all available information about the patient, namely their initial baseline data X. i and all dynamic trajectory data Z up to the current time point t i (t) is taken as input; based on this updated information, the model reassesses the risk and benefit and outputs an adjusted treatment recommendation, such as suggesting early termination of the current treatment or extending the treatment period;
[0056] S206: Perform reverse mining of subgroup-specific biomarkers. This step aims to discover new knowledge from the already segmented subgroups; after the optimal strategy tree completes the subgrouping of the patient population, the system targets each subgroup S... k Patient data within the study were used to independently perform association analyses; these analyses aimed to explore secondary biomarkers M in addition to the core features used for model construction (age, LD, dNLR). sec (e.g., PD-L1 expression levels, specific gene mutation status) and this subgroup of patients receiving their recommended treatment course There is a potential strong association between the actual clinical outcomes and the results.
[0057] S207: Generate and output a comprehensive assessment report. This step integrates the analysis results of all the preceding steps to generate a visualized comprehensive decision support report for clinical users. The report includes: the patient's subgroup affiliation and its characteristics; the initial optimal treatment course recommendation; the risk-benefit balance analysis diagram for different treatment courses; any dynamic adjustment suggestions generated in step S205 during treatment; and specific biomarker information related to the patient's subgroup discovered in step S206.
[0058] See attached document Figure 2 The risk assessment method of the present invention begins with step S201, which is to perform multidimensional data acquisition and preprocessing. The purpose of this step is to acquire and organize a high-quality, structured dataset necessary for subsequent model training and individualized assessment. The multidimensional data explicitly includes two aspects: static baseline data of the patient before treatment and dynamic trajectory data during the treatment process.
[0059] One specific implementation of step S201 may include the following sub-steps:
[0060] S2011: Perform baseline data collection and standardization. This step focuses on obtaining raw data from the patient prior to initial immune checkpoint inhibitor therapy. Specific data collected include: patient demographics, such as chronological age; imaging features reflecting tumor burden, such as the longest diameter (LD) of the target lesion measured by computed tomography (CT); and hematological parameters reflecting the body's immune inflammatory status, such as the derived neutrophil-to-lymphocyte ratio (dNLR). The specific calculation method for dNLR is as follows:
[0061]
[0062] Where, N abs WBC represents the absolute neutrophil count, while WBC represents the total white blood cell count.
[0063] After obtaining the raw baseline data, data preprocessing is necessary to ensure data quality and improve subsequent model performance. Preprocessing may include handling missing values; for example, missing data items can be filled using the mean, median, or model-based imputation methods. Further preprocessing includes data standardization to eliminate the adverse effects of differences in units and numerical ranges between different features on model training. A specific standardization method is Z-score standardization, which converts the value of each feature into a distribution with a mean of 0 and a standard deviation of 1. Its calculation formula is: Where, x ij μ is the raw value of the j-th baseline feature for the i-th patient in the dataset. j σ is the mean of the j-th feature of all patients in the training dataset; j x is the standard deviation of this feature; ij ′ represents the standardized feature value; after this step, a standardized baseline feature vector X is generated.
[0064] S2012: Perform dynamic trajectory data acquisition and time-series alignment processing. This step is used to capture information on the evolution of the patient's physical condition and treatment response over time during immunotherapy. The dynamic data items acquired specifically include: tumor shrinkage depth (DoR) for evaluating efficacy, defined as the maximum percentage reduction in tumor diameter compared to baseline; immune-related adverse events (irAEs) for evaluating safety, whose occurrence and severity are recorded and coded according to clinically accepted standards (e.g., Common Adverse Event Evaluation Criteria for Clinical Assessment); and dynamic changes in key hematological parameters (e.g., dNLR) after each treatment cycle.
[0065] Since the assessment time points for different patients are not entirely consistent on the calendar, time alignment processing needs to be performed on the collected dynamic trajectory data in order to construct a unified time series model. One specific alignment implementation method is to define a series of discrete logical time nodes {t1, t2, ..., t...}. k ,...}, where each node t k For a specific evaluation event, such as the evaluation after the end of the kth treatment cycle, the data collected for each patient on a specific calendar date is then mapped to the nearest logical time node. For cases where data is missing at certain logical time nodes, specific time-series data interpolation techniques can be used, such as the Last-Observation-Carried-Forward (LOCF) method or linear interpolation, to generate a complete and aligned dynamic trajectory data sequence Z(t). This step ensures that the dynamic data from different patients are comparable in the time dimension.
[0066] See attached document Figure 2 After performing step S201, the method of the present invention further includes step S202, namely, performing offline construction and training of the core artificial intelligence model; the purpose of this step is to generate a decision model capable of individualized risk assessment and treatment strategy recommendation based on a training dataset containing a large amount of complete historical patient data; the specific implementation of this step may include:
[0067] S2021: Define and implement a multi-objective clinical utility function. To enable the model's decision-making to comprehensively balance the efficacy benefits and safety risks in clinical treatment, this embodiment introduces a multi-objective clinical utility function U. The goal of this function is to assign a quantified comprehensive value score to the potential outcomes of each patient under different treatment strategies. This function U is related to the prediction of overall survival (OS). pred And predicting the risk of serious immune-related adverse events R irAE The function is formally expressed as:
[0068] U(d|X)=f(OS pred (d|X),R irAE (d|X));
[0069] Where X is the patient's baseline feature vector, f is a pre-defined utility mapping function, and d represents a specific treatment strategy (e.g., short, medium, or long treatment courses), selected from a pre-defined set of treatment strategies.
[0070] In one specific embodiment, the utility mapping function f can be implemented as a weighted linear combination:
[0071] f = w os OS pred -w irae ·R irAE ;
[0072] Wherein, the weighting coefficient w os and w irae These are non-negative weighting coefficients assigned to overall survival and adverse event risk, respectively. The values of these weighting coefficients can be pre-set based on clinical expert consensus, or they can be obtained through data-driven methods, such as finding the optimal weight combination that maximizes the value of the final strategy during model training using optimization algorithms such as grid search. In other embodiments, the utility mapping function f is not limited to a linear combination, but can also take the form of a non-linear function, such as a sigmoid function, to more precisely simulate the non-linear trade-off between risk and benefit in clinical decision-making.
[0073] S2022: Construct and train a counterfactual prediction model. In real-world clinical data, for any given patient, only the outcome under the actual treatment strategy received (factual outcome) can be observed, while the outcome if other treatment strategies were received (counterfactual outcome) remains unknown. To address this issue, this embodiment constructs and trains a counterfactual prediction model. The underlying implementation of this model can be a counterfactual random forest model. This model modifies the traditional random forest to estimate the potential outcomes under different interventions. In other embodiments, the counterfactual prediction model can also be other machine learning models designed for causal inference, such as causal Bayesian networks, Gaussian process models for causal inference, or deep learning models with specific network structures, such as Treatment Agnostic Representation Network (TARNet).
[0074] The training process of the model is as follows: The historical patient dataset processed in step S201 is used as input, which contains the baseline feature vector X of each patient. i The actual treatment strategies received i and its actual clinical outcomes (e.g., actual overall survival and whether serious irAEs occur); the model was trained to apply each possible treatment strategy. Predict their corresponding OS respectively pred (d|X i ) and R irAE (d|X i After training, for any patient i in the dataset, the model can output a predicted utility vector. This vector contains the predicted utility values for the patient under all preset treatment strategies:
[0075]
[0076] in, It is a predicted utility estimate calculated based on the utility function U defined in S2021 under strategy d, and its value range is the preset set of treatment strategies D = {D short D medium D long};D short D medium D long These represent pre-defined, discrete treatment strategies, referring to short, medium, and long treatment courses, respectively.
[0077] S2023: Construct the Optimal Strategy Tree (OPST) and complete the patient subgroup division. In order to transform the complex utility information output by the counterfactual prediction model into intuitive decision rules that clinicians can directly understand and apply, this embodiment further constructs an optimal strategy tree model. The goal of this model is to learn an optimal treatment strategy π, which can map any patient's baseline feature X to a treatment strategy d that can bring the maximum expected utility to it.
[0078] The model is constructed as a recursive segmentation process aimed at maximizing the total expected utility of the population; the predicted utility matrix is generated from the baseline feature vectors X and S2022 of all training patients. As input, the OPST algorithm, at each node, searches for a splitting rule (e.g., dNLR < 2.1) that maximizes the expected utility gain of the child nodes after splitting by traversing all baseline features and all possible split points. This process continues until a preset stopping condition is met (e.g., the number of samples within a node is too small or the utility gain is no longer significant). This process aims to find an optimal policy π. * The strategy maximizes the strategy value function V(π):
[0079]
[0080] in, This represents the mathematical expectation over the overall distribution of patient characteristics. This represents the expected utility that a patient with characteristic X can obtain when following strategy π;
[0081] The constructed optimal strategy tree is a binary tree structure, where each leaf node represents a patient subgroup S with similar clinical characteristics. k Furthermore, each leaf node is uniquely assigned an optimal treatment strategy. This strategy was determined during model training to be effective for the subgroup S. k Treatment strategies that deliver the highest average predictive utility to patients within the body.
[0082] See attached document Figure 2 After completing the construction and training of the core artificial intelligence model in step S202, the method of the present invention further includes step S203, which is to perform individualized treatment strategy generation and risk assessment based on the trained model. This step aims to apply the offline trained model to new individual patients to generate specific and actionable initial treatment recommendations and present relevant risk and benefit information in an intuitive way.
[0083] One specific implementation of step S203 may include the following sub-steps:
[0084] S2031: Receive and process baseline data from new patients. When it is necessary to evaluate a new non-small cell lung cancer patient, first, according to the method in step S2011, collect the patient's baseline data and perform preprocessing and standardization operations completely consistent with those used during model training to generate the standardized baseline feature vector X for the new patient. new ;
[0085] S2032: Perform subgroup assignment determination for the new patient, and assign the baseline feature vector X of the new patient. new As input, it is fed into the root node of the Optimal Policy Tree (OPST) model that has been constructed in step S2023; starting from the root node, the model, according to the decision rules set at that node, applies the following to X. new The system uses the corresponding feature values in the data to make a judgment; for example, if the rule for the root node is "dNLR<2.1", then the system will extract X. new Find the corresponding dNLR value and determine whether it is less than 2.1;
[0086] S2033: Recursively propagate along the decision tree path, see appendix. Figure 3 Based on the veracity of the judgment result in S2032, the system will determine the patient's feature vector X. new The decision is propagated to the corresponding child node (e.g., if the decision is true, it propagates to the left child node; if it is false, it propagates to the right child node). This process is repeated at each subsequent non-leaf node, that is, the decision rule is applied at each node and the next path is selected based on the decision result, until a leaf node is finally reached.
[0087] S2034: Generates an initial optimal treatment strategy. The final reached leaf node uniquely identifies the patient subgroup S to which the new patient belongs. k Because during the model building phase, each leaf node (i.e., each subgroup) is associated with an optimal treatment strategy that maximizes the expected utility of that subgroup. Therefore, the system will use this strategy This serves as the initial individualized treatment recommendation for the new patient;
[0088] S2035: Quantification and presentation of risk-benefit information. To assist clinical decision-making, the system not only outputs a single optimal strategy recommendation, but also performs a comprehensive quantitative analysis of the potential outcomes under different treatment options and presents them visually. The system uses the counterfactual prediction model trained in step S2022 to calculate the new patient's outcomes under all preset treatment strategies. The expected survival benefit (e.g., predicted median overall survival) and adverse event risk (e.g., predicted probability of serious irAEs) are considered.
[0089] S2036: Generate visual output, presenting the quantified risk-benefit information from S2035 to clinical users in one or more intuitive formats through a graphical user interface. (Refer to Appendix) Figure 4 The specific subordinate features implemented by this functional generalization may include:
[0090] One approach is to generate a risk-benefit comparison chart, such as a bar chart where each group of bars corresponds to a treatment strategy (short, medium, or long treatment duration). Each group contains two bars of different colors, representing the predicted survival benefit and the predicted toxicity risk under that strategy, allowing users to intuitively compare the trade-offs between different strategies.
[0091] Another approach is to generate a decision path diagram that visually represents the complete transmission path of the patient data in the optimal policy tree in S2033, highlighting each node it passes through, each decision rule, and the final subgroup it is assigned to, in order to enhance the transparency and interpretability of the model's decision-making process.
[0092] Another approach is to generate a comprehensive information table that lists all alternative treatment strategies in a structured manner, clearly labeling each strategy with its corresponding predicted survival outcome, predicted adverse event risk, and comprehensive utility score. The optimal strategy recommended by the system is also specially marked. Through at least one of the above visualization methods, the complex model output can be quickly understood and effectively utilized by clinicians, thereby providing strong decision support for their final treatment plan.
[0093] See attached document Figure 2 After completing the initial assessment of the new patient in step S203, the method of the present invention may further include steps S204 and S205, namely, performing dynamic monitoring and adaptive adjustment of the treatment process; the function of this step is to make the treatment strategy not static, but able to provide dynamic and intelligent feedback and optimization based on the patient's actual individualized response during the treatment process.
[0094] One specific implementation of this process may include the following sub-steps:
[0095] S2041: Calculate the deviation of the dynamic response trajectory. After the patient begins treatment according to the initial recommended plan generated in step S203, the method of this invention enters the dynamic monitoring stage. In this stage, the system continuously and periodically receives and records the patient's dynamic trajectory data at each assessment time point t, forming their individual dynamic feature vector Z. i (t); To quantify whether the patient's treatment response is as expected, the system needs to calculate the individual response trajectory and the subgroup S to which the patient belongs. k Deviation P between the expected response trajectory i (t).
[0096] The expected response trajectory of a subgroup is specifically represented as a subgroup average dynamic response trajectory vector. This vector is generated during the model training phase by calculating all data belonging to subgroup S in historical data. k The average dynamic characteristics of the patient at time point t are predetermined;
[0097] Deviation P i The calculation of (t) is performed using the metric vector Z. i (t) and vector This is achieved through the distance or difference between them; in one specific embodiment, this deviation can be obtained by calculating a weighted Euclidean distance:
[0098]
[0099] Where: m is the total number of features in the dynamic feature vector; z ij (t) is the value of the j-th dynamic feature of patient i at time point t (e.g., tumor shrinkage depth DoR or dynamic dNLR value); It is subgroup S k The average value of the j-th dynamic feature at time point t; w j is the weighting coefficient assigned to the j-th dynamic feature. This weight is used to adjust the importance of different features when calculating the overall deviation. For example, the DoR, which reflects direct changes in the tumor, can be given a higher weight than blood indicators. In another embodiment, all weights w j All values can be set to 1, in which case the formula degenerates into the standard Euclidean distance; in other embodiments, other distance metrics such as Mahalanobis distance can also be used to consider the correlation between different dynamic features.
[0100] S2051: Trigger determination for implementing the adaptive adjustment mechanism; the system calculates the deviation P at each evaluation time point t. i (t) will then be combined with a subgroup S. k Preset deviation threshold θ k Compare; the threshold θk It is a pre-defined subgroup S k The independently set value can be determined based on statistical principles. For example, it can be set to the 90th or 95th percentile of the deviation distribution observed in the historical training data for this subgroup, in order to identify individuals who statistically significantly deviate from the group's average response. The adaptive adjustment mechanism is triggered if and only if the calculated deviation meets the following condition:
[0101] P i (t)>θ k ;
[0102] If the deviation does not exceed the threshold, the system determines that the patient's response is within the expected range, continues to implement the original treatment plan, and is continuously monitored.
[0103] S2052: Adaptive adjustment of the treatment strategy. After the adaptive adjustment mechanism in S2051 is triggered, the system calls a pre-trained dynamic adjustment model dedicated to strategy adjustment. The specific implementation of the dynamic adjustment model can be an independent classification or regression model, such as a logistic regression model, support vector machine, gradient boosting decision tree, or a small neural network.
[0104] The model receives all available information about the patient, namely their initial baseline feature vector X. i Compared with all dynamic trajectory data Z up to the current time point t i (t) serves as its combined input feature; based on this comprehensive and updated information, the model reassesses the risk and benefit, and outputs an updated and specific treatment adjustment recommendation; this recommendation may include, but is not limited to: maintaining the current treatment regimen and increasing monitoring frequency, recommending immediate termination of treatment to avoid potential serious toxicity or ineffective treatment, or recommending a comprehensive assessment after extending the treatment period by a specific number of cycles. This adjustment recommendation is then integrated into the comprehensive assessment report of step S207 for clinical users' decision-making reference.
[0105] See attached document Figure 2 After completing the core model construction and subgroup division, the method of the present invention may further include step S206, namely, performing reverse mining of subgroup-specific biomarkers. This step is a knowledge discovery process, the purpose of which is to use the established, clinically homogeneous patient subgroups to explore in depth the potential association between other biological factors and treatment outcomes, in addition to the core features used to divide the subgroups, thereby providing data-driven hypotheses for clinical research.
[0106] One specific implementation of step S206 may include the following sub-steps:
[0107] S2061: Perform association analysis within subgroups. This step is performed for each patient subgroup S defined by the optimal strategy tree constructed in step S2023. k Data analysis is conducted independently; the datasets analyzed are limited to those belonging to that specific subgroup S. k The patient sample; the core of the analysis is to evaluate a predefined set of secondary biomarkers M sec The statistical strength of the association between the patient's actual clinical outcome Y and the patient's actual clinical outcome Y;
[0108] Secondary biomarker M sec This may include, but is not limited to: tumor proportion score (TPS) of programmed death-ligand-1 (PD-L1), tumor mutation burden (TMB), specific gene mutation status (e.g., KRAS or EGFR mutation status, represented as a binary variable), or other hematological indicators; the actual clinical outcome Y may include the patient's actual overall survival (OS), progression-free survival (PFS), or a binary record of whether serious (≥ grade 3) immune-related adverse events (irAEs) occurred;
[0109] The specific sub-technical implementation of association analysis varies depending on the type of outcome variable: for survival-related outcomes (such as OS or PFS), a Cox proportional hazards regression model can be used for analysis; this model is used to assess the impact of one or more minor biomarkers on patient survival risk; its mathematical expression is:
[0110] h(t|M sec )=h0(t)exp(β T M sec );
[0111] Among them, h(t|M sec ) is given a secondary biomarker vector M sec In the case of , h0(t) is the baseline hazard rate at time point t; β is the vector of regression coefficients to be estimated, and each component β j This reflects the magnitude of the logarithmic hazard rate of the j-th minor biomarker; by performing a statistical significance test on β (e.g., Wald test), it can be determined whether a specific biomarker is significantly associated with survival outcomes.
[0112] For binary outcomes (e.g., whether severe irAEs occur), a logistic regression model can be used for analysis; this model is used to assess the impact of secondary biomarkers on the probability of event occurrence; its mathematical expression is:
[0113]
[0114] Where p is the probability of the event occurring. Similarly, by performing a significance test on the regression coefficient vector β, we can determine which markers are significant predictors of the event's occurrence.
[0115] S2062: Perform knowledge discovery and hypothesis generation. This step, based on the correlation analysis results of S2061, automatically generates structured knowledge and research hypotheses. When the statistical model test results in S2061 show that in a certain subgroup S... k In one instance, biomarker M was needed. sec,j With a certain clinical outcome Y l When the correlation between the two reaches a preset statistical significance level (e.g., p-value less than 0.05), the system determines that a strong correlation has been found;
[0116] The system then solidified this finding into a structured knowledge entry; this entry explicitly records the contextual information of the association, including: the patient subgroup S to which it belongs. k Definition characteristics, secondary biomarkers involved M sec,j Related clinical outcomes Y l The direction and strength of the association (e.g., quantified by the hazard ratio or odds ratio), and the corresponding statistical confidence level (e.g., p-value and confidence interval);
[0117] These structured knowledge entries constitute the research hypotheses output by this invention. For example, the system might output a hypothesis: "For subgroup 7 characterized by high dNLR and long LD, the presence of KRAS gene mutations is significantly associated with prolonged overall survival after receiving long-term immunotherapy (HR = 0.6, p = 0.03)." Such hypotheses are not directly used to guide current patient treatment decisions, but rather serve as a high-level analytical result, pointing clinical researchers to valuable research directions. For example, they can be used to design prospective clinical trials targeting specific subgroups to verify the predictive value of this biomarker.
[0118] See attached document Figure 1 The present invention also provides an embodiment of a risk assessment system for immunotherapy of non-small cell lung cancer, which is configured to perform the aforementioned risk assessment method. Each functional module of the system is a physical or logical entity that implements each step of the method.
[0119] The data acquisition and integration module 110 performs the aforementioned method step S201. This module includes a data interface for communicating with external data sources. The specific implementation of the data interface may include: a database connector for connecting to a hospital information system (HIS) or electronic medical record (EMR) system, an application programming interface (API) for calling data from a laboratory information system (LIS), or a file import component for parsing and reading structured data files (e.g., CSV or XML format). This module further includes a data preprocessing unit, which has embedded algorithmic logic for performing data cleaning, missing value imputation, and data standardization based on the Z-score standardization formula. This module also includes a time series alignment unit, which performs regularization and alignment of dynamic trajectory data collected on different calendar dates according to logical time nodes such as treatment cycles, and finally generates standardized baseline feature vectors and dynamic trajectory data sequences that can be used by other modules of the system. The output of this module is transmitted to the model training and optimization module 120 and the subgroup division and strategy generation module 130.
[0120] Model training and optimization module 120, whose function is to execute the aforementioned method step S202, typically runs in an offline environment. The internal structure of this module may include: a utility function configuration unit, allowing system administrators or researchers to set or adjust the weight coefficients w in a multi-objective clinical utility function. os and w irae The module includes: a counterfactual model training unit with an embedded machine learning algorithm library for training counterfactual random forests or other causal inference models, which receives historical datasets and generates a prediction utility matrix; and an optimal policy tree construction unit that implements a recursive segmentation algorithm aimed at maximizing the total expected utility of the population. The final output of this module is not data for a single patient, but a trained, fixed, and directly callable core artificial intelligence model file, such as a serialized optimal policy tree model, which is stored and used by the subgrouping and policy generation module 130 during online evaluation.
[0121] The subgrouping and strategy generation module 130 performs the individualized assessment in step S203 of the aforementioned method. This module includes a model loading unit for reading and loading the trained model generated by the model training and optimization module 120 from the storage medium. At its core is an inference engine. When a standardized baseline feature vector of a new patient is received, the inference engine executes the decision rules fixed in the optimal strategy tree model. Through continuous judgment of the patient's feature values, it completes the path traversal of the patient in the decision tree, thereby determining the final subgroup to which the patient belongs and extracting the optimal treatment strategy associated with that subgroup as the initial recommendation. This module outputs the generated subgroup attribution information and initial strategy suggestions to the risk assessment and visualization module 140.
[0122] The risk assessment and visualization module 140 performs risk quantification and presentation in step S203 and report generation in step S207. This module includes a risk calculation unit that, upon receiving new patient data, invokes a trained counterfactual prediction model to calculate the patient's expected survival benefit and adverse event risk under all alternative treatment strategies. The module further includes a graphics rendering engine that converts numerical information and decision path information output by the subgrouping and strategy generation module 130 into graphical interface elements. This engine generates risk-benefit comparison bar charts, decision path diagrams, and comprehensive information data tables, integrating these visualization results into a comprehensive assessment report displayed on the user terminal's display device.
[0123] The dynamic monitoring and adjustment module 150 is designed to execute the aforementioned method steps S204 and S205, forming a closed-loop feedback mechanism for the system. This module includes a data stream receiving unit for continuously receiving and updating the patient's dynamic trajectory data during treatment. Internally, the module includes a deviation calculation unit that calculates the deviation between the individual response trajectory and the subgroup average trajectory in real time based on a preset distance metric formula (e.g., weighted Euclidean distance). The module also includes a trigger unit that compares the calculated deviation with a preset threshold. When the trigger condition is met, a dynamic adjustment model execution unit is activated. This execution unit calls an independent, pre-trained dynamic adjustment model, takes all the patient's latest data as input, generates updated treatment adjustment suggestions, and transmits these suggestions to the risk assessment and visualization module 140 for display.
[0124] The knowledge discovery module 160, which executes the aforementioned method step S206, is an offline analysis module for exploratory data analysis. This module includes a data filtering unit that extracts patient subsets belonging to specific subgroups from the complete dataset based on the partitioning results of the optimal strategy tree. At its core is a statistical analysis engine that incorporates various statistical modeling methods, such as Cox proportional hazards regression and logistic regression. This engine can automatically and in batches test the association between predefined secondary biomarkers and clinical outcomes on each subgroup data subset. The module also includes a hypothesis generation unit. When the statistical analysis engine finds a strong association reaching a significant level, this unit is responsible for formatting the finding into a structured text format containing information such as subgroup characteristics, biomarkers, outcomes, association strength, and confidence levels, forming a data-driven research hypothesis report.
[0125] The present invention also provides a computer-readable storage medium having a computer program stored thereon; when executed by a processor, the computer program is configured to implement the non-small cell lung cancer immunotherapy risk assessment method of any of the foregoing embodiments;
[0126] Specifically, when the instructions contained in the computer program are executed by the processor, they can implement all or part of the steps of the method, including: performing multidimensional data acquisition and preprocessing; performing offline construction and training of the core artificial intelligence model; performing individualized initial assessments on new patients to generate treatment strategies; performing dynamic monitoring of the treatment process and triggering adaptive adjustments of the treatment strategy when preset conditions are met; and performing reverse mining of subgroup-specific biomarkers to generate research hypotheses.
[0127] A computer-readable storage medium can be a non-transitory storage medium; a computer-readable storage medium can be any entity or device capable of storing a computer program, for example, a storage medium can include, but is not limited to: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0128] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for risk assessment of immunotherapy for non-small cell lung cancer, characterized in that, Includes the following steps: Obtain a historical patient dataset containing baseline data and dynamic trajectory data of multiple historical patients; An optimal strategy tree model is constructed based on historical patient datasets. The optimal strategy tree model divides patients into multiple subgroups and determines an initial optimal immunotherapy course recommendation strategy for each subgroup. Baseline data of new patients are obtained, and the subgroup to which the new patient belongs and the recommended initial optimal immunotherapy course strategy corresponding to the subgroup are determined according to the optimal strategy tree model. During the process of treating new patients according to the recommended strategy of the initial optimal immunotherapy course, dynamic trajectory data of new patients are continuously acquired to form individual dynamic feature vectors. Calculate the deviation between the individual dynamic feature vector and the subgroup average dynamic response trajectory vector of the subgroup to which the new patient belongs; When the deviation exceeds the subgroup's preset deviation threshold, the updated treatment adjustment suggestions are output by combining the new patient's baseline data and the new patient's dynamic trajectory data.
2. The method for risk assessment of non-small cell lung cancer immunotherapy according to claim 1, characterized in that, A counterfactual prediction model is trained based on the historical patient dataset to generate a prediction utility matrix for each historical patient under different preset treatment strategies. Using the predicted utility matrix as input, the optimal policy tree model is constructed through a recursive segmentation process aimed at maximizing the total expected utility of the group.
3. The method for risk assessment of non-small cell lung cancer immunotherapy according to claim 2, characterized in that, The counterfactual prediction model is trained based on a multi-objective clinical utility function, which comprehensively quantifies the clinical value of different treatment strategies by weighting the expected survival benefit and the risk of adverse events.
4. The method for risk assessment of immunotherapy for non-small cell lung cancer according to claim 1, characterized in that, The deviation is calculated by measuring the weighted Euclidean distance between the individual dynamic feature vector and the subgroup average dynamic response trajectory vector.
5. The method for risk assessment of immunotherapy for non-small cell lung cancer according to claim 1, characterized in that, The subgroup average dynamic response trajectory vector is predetermined during the model training phase by calculating the average dynamic characteristics of all patients belonging to the same subgroup in the historical patient dataset at each assessment time point.
6. The method for risk assessment of immunotherapy for non-small cell lung cancer according to claim 1, characterized in that, The baseline data includes the patient's age, the longest diameter of the tumor, and the ratio of derived neutrophils to lymphocytes; the dynamic trajectory data includes the depth of tumor shrinkage and records of immune-related adverse events.
7. The method for risk assessment of immunotherapy for non-small cell lung cancer according to claim 1, characterized in that, The method further includes: After the optimal strategy tree model completes the subgrouping of patients, a correlation analysis is performed within each subgroup on predefined secondary biomarkers and the actual clinical outcomes of patients in order to backward mine secondary biomarkers related to the efficacy or toxicity of the subgroup.
8. The method for risk assessment of immunotherapy for non-small cell lung cancer according to claim 1, characterized in that, The method further includes: Generate and output a comprehensive evaluation report, which includes a recommended strategy for the initial optimal immunotherapy course and a risk-benefit balance analysis diagram comparing the expected survival benefits and adverse event risks under different treatment strategies.
9. A risk assessment system for immunotherapy of non-small cell lung cancer, applied to the method described in any one of claims 1-8, characterized in that, The system includes: The data acquisition and integration module is used to acquire historical patient datasets containing baseline data and dynamic trajectory data of multiple historical patients, as well as to acquire baseline data and dynamic trajectory data of new patients. The model training and optimization module is used to construct an optimal strategy tree model based on the historical patient dataset. The optimal strategy tree model divides patients into multiple subgroups and determines an initial optimal immunotherapy course recommendation strategy for each subgroup. The subgrouping and strategy generation module is used to determine the subgroup to which a new patient belongs and the recommended initial optimal immunotherapy course strategy corresponding to the subgroup based on the optimal strategy tree model. The dynamic monitoring and adjustment module is used to continuously acquire the dynamic trajectory data of new patients to form individual dynamic feature vectors during the treatment of new patients according to the initial optimal immunotherapy course recommendation strategy, calculate the deviation between the individual dynamic feature vector and the subgroup average dynamic response trajectory vector of the subgroup to which the new patient belongs, and output updated treatment adjustment suggestions when the deviation exceeds the preset deviation threshold for the subgroup, combining the baseline data of the new patient and the dynamic trajectory data of the new patient.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.
Citation Information
Cited By
Therapeutic effect evaluation and remote follow-up visit platform for ultrasonic interventional treatment of mastitis
CN121565361A
Efficacy evaluation and remote follow-up platform for ultrasound intervention treatment of mastitis
CN121565361B