Non-small cell lung cancer anti-pd-1 individualized treatment decision support system and method thereof
By constructing a dynamic knowledge graph of multi-source heterogeneous data and an individualized treatment decision support system based on domestic large language models, the problems of low response rate, difficult side effect management, and heavy economic burden of NSCLC immunotherapy have been solved. Precision treatment and resource optimization have been achieved, and patient safety and the standardization of primary healthcare have been improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU PROVINCE HOSPITAL (THE FIRST AFFILIATED HOSPITAL OF NANJING MEDICAL UNIVERSITY)
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-10
AI Technical Summary
Current technologies for immunotherapy of non-small cell lung cancer (NSCLC) have low response rates, difficulty in predicting efficacy, complex treatment pathways, and difficulties in managing side effects, resulting in a heavy economic burden. Furthermore, traditional CDSS lacks Chinese language support and multimodal fusion, leading to resource waste and increased pressure on medical insurance.
We will construct a decision support system for individualized anti-PD-1 treatment of non-small cell lung cancer based on multi-source heterogeneous clinical data. By integrating dynamic knowledge graphs and domestic large language models, we can achieve intelligent screening of indications for immunotherapy, efficacy prediction, side effect warning and optimal treatment recommendation. Combined with individualized treatment pathways and cost-effectiveness analysis, we will develop an intelligent interactive system to optimize decision-making through incremental learning and multimodal fusion.
To improve the precision of treatment, reduce the rate of ineffective treatment, reduce the economic burden on patients, enhance toxicity warning, improve patient safety and the efficiency of doctor-patient collaboration, promote the standardization of immunotherapy in primary hospitals, and narrow the urban-rural medical gap.
Smart Images

Figure CN122369781A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical artificial intelligence, clinical decision support systems (CDSS) and precision medicine, and in particular to a decision support system and method for individualized anti-PD-1 treatment of non-small cell lung cancer. Background Technology
[0002] Currently, non-small cell lung cancer (NSCLC), as one of the malignant tumors with the highest incidence and mortality rates worldwide, has entered the era of immunotherapy, especially PD-1 / PD-L1 immunotherapy, which has greatly expanded the survival horizons of patients with advanced lung cancer. However, existing technologies still have the following prominent problems:
[0003] Low response rate and difficulty in predicting efficacy of immunotherapy: Currently, the efficacy rate of PD-1 immunotherapy in the NSCLC population is only 20%-30%, and a large proportion of patients still do not see significant benefits after treatment, or even experience accelerated progression. How to accurately screen potential beneficiaries before treatment using biomarkers (such as PD-L1 expression, TMB, MSI, etc.) is a pressing clinical challenge.
[0004] Treatment pathways are complex and side effect management is difficult: Immune-related adverse events (irAEs) caused by immunotherapy are characterized by their insidious nature, wide range of organs affected, and rapid progression. Currently, their diagnosis still mainly relies on patient self-reporting and human experience, lacking intelligent analysis and dynamic monitoring mechanisms.
[0005] The heavy economic burden and pressure on the medical insurance fund are significant challenges: the annual treatment cost for a single NSCLC patient receiving PD-1 immunotherapy is generally between 150,000 and 250,000 yuan, and ineffective treatments lead to serious waste of resources. At the same time, the pressure on the sustainable operation of the medical insurance fund is becoming increasingly prominent.
[0006] Existing decision support systems (CDSS) have significant limitations: traditional rule-driven CDSS suffers from static, rigid, and poor scenario adaptability; international solutions are mostly trained on English data, lacking adaptation to Chinese medical language habits, medical insurance policy constraints, and local treatment standards. A truly end-to-end personalized treatment platform that focuses on immunotherapy scenarios, is based on domestically developed large-scale models for local deployment, and integrates multiple modalities is still lacking. Summary of the Invention
[0007] To address the problems existing in the prior art, this invention provides a decision support system and method for personalized anti-PD-1 treatment of non-small cell lung cancer. It integrates multi-source heterogeneous clinical data (electronic medical records, imaging, genomics, pathology, medical insurance, etc.) to construct a standardized and scalable tumor knowledge base and dynamic knowledge graph. Based on a domestically developed large language model, it integrates multimodal medical information to achieve intelligent screening of immunotherapy indications, efficacy prediction, side effect warning, and optimal treatment recommendation. It constructs an intelligent matching engine for personalized treatment pathways and clinical trials, taking into account both efficacy and cost-effectiveness analysis. It develops an intelligent interaction and toxicity management system for doctors and patients to improve patient self-management and doctor-patient collaboration efficiency. Through clinical trials and multi-center promotion, it verifies and optimizes system performance, promoting the downward flow of high-quality medical resources.
[0008] The objective of this invention is achieved through the following technical solutions.
[0009] A decision support system for individualized anti-PD-1 treatment of non-small cell lung cancer, based on multi-source clinical data integration and artificial intelligence, includes:
[0010] The data acquisition and standardization module is used to collect multi-source data from patients receiving anti-PD-1 therapy at our hospital and partner hospitals within a specified time period. This multi-source data includes electronic health records, imaging data, gene sequencing data, tumor microenvironment characteristics, blood biomarkers, medical insurance policy documents, and clinical guidelines. It is used for:
[0011] Using the HL7 FHIR standard, structured data from clinical texts are extracted using natural language processing tools, mapping rules between texts and FHIR resources are established, and parameterized tables or feature scales are formed.
[0012] Image data is parsed based on the DICOM protocol, and metadata and pixel data are parsed in the order of data elements. The data is then stored in FHIR and PACS in a mixed manner, with both structured data and large pixel data.
[0013] Using the SNOMED CT terminology system, and with the help of natural language processing technology, deep learning models and medical ontology, the target, staining results, expression sites and positive ratio information in the immunohistochemistry report are extracted in a structured manner and mapped into standard SNOMED CT codes.
[0014] The dynamic knowledge graph construction module is used to integrate standardized data into the Neo4j graph database to construct a dynamic knowledge graph. The dynamic knowledge graph stores structured data, including clinical guidelines, drug indications, real-world data, and medical insurance rules, in terms of nodes and relationships.
[0015] The incremental learning module is used to capture and integrate new knowledge in real time through active learning technology, including the latest clinical trials and drug resistance mechanism research. Through change detection and capture, incremental knowledge extraction, and conflict detection and resolution, the system can continuously update the knowledge base content without retraining the entire model.
[0016] An intelligent decision support module based on a large language model is used to provide decision support for immunotherapy, including:
[0017] A localized large language model unit is used to deploy and run a domestically developed large language model fine-tuned through transfer learning, wherein the fine-tuning injects domain knowledge related to anti-PD-1 treatment;
[0018] A multimodal input processing unit is used to fuse image features, genomic data, clinical text, and economic parameters using a Transformer-based cross-modal encoder, and to perform multimodal fusion through a multi-head self-attention mechanism;
[0019] The immunotherapy decision optimization unit is used to perform indication screening, efficacy prediction, side effect prediction, conflict detection, and decision optimization.
[0020] The personalized treatment plan and clinical trial matching unit is used to perform treatment plan recommendations and cost optimization, as well as to match patients with clinical trials through semantic matching algorithms;
[0021] The patient education and interaction module is used to generate personalized educational content, provide intelligent question-and-answer services, and conduct intelligent toxicity management.
[0022] The clinical validation and optimization module is used to clinically validate the system through randomized controlled trials and supports lightweight adaptation to primary healthcare institutions.
[0023] In the data acquisition and standardization module:
[0024] When parsing image data based on the DICOM protocol, to accurately identify blurred tumor boundaries, a boundary loss function based on an integral architecture is used during training to minimize the distance between the predicted boundary and the labeled boundary. This boundary loss function L... BD As shown below: ,in Boundary loss function, Ω: Image domain Spatial location The probability value output by the model. : Distance transformation function for the true labeled boundary.
[0025] The immunotherapy decision optimization unit includes:
[0026] The indication screening subunit is used to execute a stratified decision-making model based on PD-L1, TMB, and MSI, and to make stratified decisions based on the predicted response to anti-PD-1 treatment.
[0027] The efficacy prediction subunit is used to dynamically predict the response rate using a time-series Transformer model. Its model input is a time-series feature sequence. The output is the future state. The study also incorporates positional encoding and uses an XGBoost model to analyze drug resistance risk. The objective function of the XGBoost model is shown below:
[0028] , ;in, The loss function represents the difference between predicted and actual values. For the first The complexity penalty term for trees, where, The number of leaf nodes, and These represent structural complexity and weight regularization coefficients, respectively. The weights are those of the leaf nodes;
[0029] The side effect prediction subunit is used to model the risk of immune-related adverse reactions using a graph neural network. It constructs a patient-drug-symptom triple graph G=(V,E) and learns node representations based on the update rules of a simplified graph convolutional network, as shown below:
[0030] ;in, Represents a node In the Layer feature representation, For its set of neighboring nodes, For learnable weight matrix, The normalization coefficient is... The activation function is nonlinear; this method captures the complex relationships between patients, drugs, and adverse reactions through multi-layer information propagation, thereby achieving accurate modeling of the risk of immune-related adverse reactions.
[0031] The conflict detection and decision optimization subunit is used to jointly verify logical contradictions through a rule engine and a large language model, and to penalize contradictory errors. Its verification process is formalized as follows:
[0032] ;
[0033] in: Indicates the first Rule 1 and These represent the premise and conclusion of the rule, respectively. This indicates a logical implication relationship; Represents a set of patient contextual information. This indicates that the current patient status meets the prerequisites of the rule; This indicates the conclusions predicted or generated by the model. This represents a semantic similarity function used to measure the consistency between two conclusions. This is the threshold for similarity determination.
[0034] The individualized treatment plan and clinical trial matching unit includes:
[0035] The treatment plan recommendation and cost optimization subunit is used to make tumor treatment plan decisions through a hierarchical decision model, combined with TNM staging, driver gene status, and PD-L1 expression; and to perform cost-benefit analysis, budget impact model calculation, and dynamic adjustment of out-of-pocket ratio through an economic analysis engine to achieve cost optimization.
[0036] The dynamic prognostic prediction subunit is used to combine the COX proportional hazards model and the XGBoost model for dynamic prognostic prediction.
[0037] The clinical trial matching subunit is used to extract standard features of clinical trials using the BERT model and perform semantic matching based on the cosine similarity between the patient feature embedding vector and the trial-incorporated standard embedding vector. The similarity calculation is as follows:
[0038] ;in, An embedding vector representing patient features. The embedding vector represents the criteria for inclusion in clinical trials; Representing vectors and The inner product; and Representing vectors respectively and The norm; The cosine similarity between the two is used to measure the degree of similarity between patient characteristics and trial inclusion criteria in the semantic space.
[0039] It also uses simulated blockchain technology to achieve cross-institutional data privacy protection for decentralized recruitment.
[0040] The patient education and interaction module includes:
[0041] The personalized educational content generation unit is used to generate differentiated popular science content using a multilingual large language model and supports low-resource languages.
[0042] The intelligent question-answering unit is used for intent recognition and slot filling using a bidirectional long short-term memory network-conditional random field model, and for sentiment analysis using a RoBERTa model. The classification prediction of the RoBERTa model is shown in the formula. As shown in the formula, the classification loss is... As shown; where, This represents the class probability distribution vector predicted by the model. This represents the weight matrix of the classification layer. This represents the feature representation vector input to the classification layer. Indicates the bias term. This represents the normalization exponential function, used to map the output to a probability distribution; Represents the classification loss function. The first one representing the true label Class value, Represents the first element in the prediction probability vector. The probability of a class;
[0043] The intelligent toxicity management unit is used to analyze patient complaints through natural language processing and combine them with laboratory indicators to achieve real-time monitoring and early warning; and to provide a doctor-patient collaboration platform including a doctor-side decision dashboard and a patient-side APP.
[0044] In the multimodal input processing unit, the core formula of the multi-head self-attention mechanism is as follows:
[0045] ;
[0046] ;
[0047] ,in, , , These represent the query, key, and value matrices, respectively. This represents the product of the query matrix and the transpose of the key matrix; This represents the dimension of the key vector, used to scale the inner product result; This represents the normalized exponential function, used to map attention weights to a probability distribution; This represents the result of attention calculation; Indicates the first The output of each attention head; , , They represent the first The linear transformation weight matrix used to generate queries, keys, and values in each attention head; This indicates a concatenation operation of multiple attention head outputs; This represents the linear transformation weight matrix after multi-head output; This indicates the number of attention heads.
[0048] In the efficacy prediction subunit, the position encoding of the time series Transformer model is as follows:
[0049] , ;in, and Representing positions respectively First With the The encoded value of the dimension; Indicates the position index in the sequence; Indicates the feature dimension index; This represents the total dimension of the positional encoding vector.
[0050] The clinical validation and optimization module also includes a primary healthcare adaptation unit, used for:
[0051] Lightweight models are developed through model distillation and deployed on edge computing devices, allowing samples with confidence levels below a threshold to seek help from the cloud.
[0052] Provide a remote MDT collaboration platform to deploy hybrid inference pipelines in primary healthcare information systems and synchronize multimodal data.
[0053] An immunotherapy decision support method based on the above system includes the following steps:
[0054] Data acquisition and standardization steps: Collect multi-source clinical data and standardize them using the HL7 FHIR standard, DICOM protocol and SNOMED CT terminology system;
[0055] Knowledge base construction and update steps: Construct a dynamic knowledge graph and update the knowledge base content in real time through an incremental learning mechanism;
[0056] Intelligent decision support steps: Based on a locally deployed large language model, multimodal input data is fused and processed, and indication screening, efficacy prediction, side effect prediction, conflict detection and decision optimization are performed to generate individualized treatment plans, match clinical trials, and optimize costs.
[0057] Patient interaction and management steps: Generate personalized educational content, provide intelligent Q&A services, and conduct intelligent toxicity management;
[0058] Clinical validation and optimization steps: Conduct clinical validation through randomized controlled trials, and optimize the system based on the validation results.
[0059] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method described above.
[0060] Compared with existing technologies, the advantages of this invention are: improved treatment precision and standardization: through large language models and knowledge graphs, it accurately identifies the population that benefits from immunotherapy, avoids ineffective treatment, and promotes the transformation of NSCLC treatment from a "uniform path" to an "individualized" model.
[0061] Reduce patients' economic burden: It is estimated that the incidence of ineffective treatment can be reduced by 30%. Based on an average annual treatment cost of 200,000 yuan per patient, each patient can save about 50,000 yuan in direct costs. After its promotion, it can save patients hundreds of millions of yuan in total.
[0062] Improving the efficiency of medical insurance payments: The system integrates medical insurance policy parameters and cost-benefit analysis models (ICER, QALY), which is expected to save hundreds of millions of yuan in medical insurance funds.
[0063] Enhance toxicity early warning and patient safety: Achieve early detection and intervention of irAEs, reduce the incidence of severe complications, and improve patients' quality of life.
[0064] Promoting the downward flow of high-quality medical resources: Lightweight models and remote MDT collaboration can increase the standardization rate of immunotherapy in primary hospitals by more than 30%, narrowing the urban-rural medical gap.
[0065] Enhancing patient compliance and engagement: The intelligent interactive system provides multilingual health education, toxicity self-assessment, and personalized Q&A, significantly improving patients' understanding, trust, and participation in the treatment process. Attached Figure Description
[0066] Figure 1 Screenshot of the immunohistochemical results of LLM analysis of breast cancer pathology images from our hospital.
[0067] Figure 2 Screenshot of the results of LLM analysis of high-risk factors after early NSCLC surgery in our hospital. Detailed Implementation
[0068] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0069] Multi-source clinical data integration and knowledge base construction:
[0070] 1. Data Collection: Collect data on patients who received anti-PD-1 therapy from our hospital and partner hospitals between January 1, 2015 and January 1, 2025. Integrate electronic health records (EHR), imaging data (CT / MRI), gene sequencing data (PD-L1, TMB, MSI), tumor microenvironment characteristics, blood biomarkers and medical insurance policies, NCCN / CSCO documents to construct a basic dataset and knowledge base.
[0071] 2. Standardization process:
[0072] 1) Adopt the HL7 FHIR standard to unify clinical text and structured data: Use NLP tools to extract structured data, design mapping rules between text and FHIR resources, and establish parameterized tables or feature scales in sequence.
[0073] 2) Parsing image data based on the DICOM protocol: Metadata and pixel data are parsed in the order of data elements and stored in FHIR and PACS in a mixed mode of structured data and large pixel data. In order to accurately identify the blurred tumor boundary, the distance between the predicted boundary and the labeled boundary is minimized during the training process by using a boundary loss based on an integral architecture. The boundary loss function L_BD is shown in Equation (1).
[0074] Formula (1); where Ω: Boundary Loss function; Ω: Image domain. Spatial location (pixel / voxel index). : The probability value output by the model (softmax / sigmoid output). : Distance transformation function for true labeled boundaries (Signed Distance Map, SDM).
[0075] 3) Annotating Immunohistochemistry Reports Using the SNOMED CT Terminology System: Utilizing Natural Language Processing (NLP) technology, key information from pathology reports, such as targets (e.g., PD-L1, Ki-67), staining results (positive / negative), expression sites, and positive rates, is extracted in a structured manner and mapped to standard SNOMED CT codes. This process is then collaboratively completed using deep learning models (e.g., BERT, MedLex) and medical ontology (e.g., SNOMED CT browser, national terminology extension packages). The result is standardized, interoperable structured data, providing precise semantic support for tumor knowledge graph construction, clinical decision support, and large-scale model training.
[0076] 4) Knowledge Base Architecture Design: Construct a dynamic knowledge graph (Neo4j graph database) that integrates NCCN / CSCO guidelines, FDA / EMA drug indications, real-world data (such as the Flatiron Health database), and healthcare regulations. The dynamic knowledge graph stores structured data in the graph, with data stored as nodes and relationships.
[0077] 5) Incremental learning mechanism: The knowledge base content (such as the latest clinical trials and drug resistance mechanism research) is updated in real time through active learning technology. The key lies in change detection and capture, incremental knowledge extraction, and conflict detection and resolution, enabling the system to continuously integrate new knowledge without retraining the entire model.
[0078] Development of an LLM-based Intelligent Decision Support System
[0079] 1. Localized LLM training and domain adaptation: Based on the domestic LLM, DeepSeekR1(32B) is used for transfer learning, injecting knowledge in the field of immunotherapy, including anti-PD-1 therapy-related guidelines and expert consensus, management of immunotherapy-related adverse reactions, and the latest clinical trial conclusions for self-learning.
[0080] 2. Multimodal Input Processing: A Transformer-based cross-modal encoder was used to fuse image features (extracted by ResNet-50), genomic data, clinical text, and cost-effectiveness parameters. Multi-head self-attention was employed for multimodal fusion.
[0081] The core formula generates a query (Q), key (K), and value (V) vector for each input vector through a linear transformation:
[0082] , formula 2
[0083] , formula 3
[0084] , formula 4;
[0085] in, , , These represent the query, key, and value matrices, respectively. This represents the product of the query matrix and the transpose of the key matrix; This represents the dimension of the key vector, used to scale the inner product result; This represents the normalized exponential function, used to map attention weights to a probability distribution; This represents the result of attention calculation; Indicates the first The output of each attention head; , , They represent the first The linear transformation weight matrix used to generate queries, keys, and values in each attention head; This indicates a concatenation operation of multiple attention head outputs; This represents the linear transformation weight matrix after multi-head output; This indicates the number of attention heads.
[0086] 3. Localized Deployment: A domestically developed LLM model was selected, and Ubuntu 22.04 LTS + NVIDIA CUDA 12.1 was installed on an intranet server. The DeepSeek-R1 image file was imported via a secure USB drive and loaded into a private Docker repository. After further development, data integration with HIS / PACS will be attempted to verify decision response time.
[0087] 4. Immunotherapy Decision Optimization Module:
[0088] 1) Indication screening: Stratified decision-making model (based on PD-L1, TMB, MSI, etc.) to make stratified decisions based on predicted anti-PD-1 treatment responsiveness.
[0089] 2) Treatment efficacy prediction: The response rate is dynamically predicted using a time series Transformer model. The core formula is:
[0090] Used for predicting the efficacy of immunotherapy, the model input is a time-series feature sequence. The output is the future state. Introducing positional encoding (PE): , Formulas 5 and 6; where, and Representing positions respectively First With the The encoded value of the dimension; Indicates the position index in the sequence; Indicates the feature dimension index; This represents the total dimension of the positional encoding vector.
[0091] The system is then trained using the Transformer architecture described above to output the probability of the patient's immune response in the next cycle. XGBoost was used to analyze drug resistance risk. The key formulas for building a drug resistance model using XGBoost are as follows:
[0092] , Formulas 7 and 8; where, The loss function represents the difference between predicted and actual values. For the first The complexity penalty term for trees, where, The number of leaf nodes, and These represent structural complexity and weight regularization coefficients, respectively. The weights are those of the leaf nodes.
[0093] 3) Side Effect Prediction: Graph Neural Networks (GNNs) are used to model the risk of irreversible adverse events (irAEs). A high risk level can prompt physicians to intervene early. The core formula is:
[0094] Constructing a patient-drug-symptom ternary group diagram The node is represented as
[0095] Update rules (GCN simplified version):
[0096] Formula 9; where, Represents a node In the Layer feature representation, For its set of neighboring nodes, For learnable weight matrix, The normalization coefficient is... This is a non-linear activation function. This method captures the complex relationships between patients, drugs, and adverse reactions through multi-layered information propagation, thereby achieving accurate modeling of the risk of immune-related adverse reactions.
[0097] 5. Conflict Detection and Decision Optimization: The rule engine (Drools) and LLM jointly verify logical contradictions. Inconsistent errors are penalized to reduce subsequent recommendations. The entire verification process of the joint logic check can be formalized as follows:
[0098]
[0099] Indicates the first Rule 1 and These represent the premise and conclusion of the rule, respectively. This indicates a logical implication relationship; As a premise (e.g., "PD-L1 expression < 1%)", For conclusions (e.g., "The use of PD-1 inhibitor monotherapy is not recommended"), Represents a set of patient contextual information. This indicates that the current patient status meets the prerequisites of the rule; This indicates the conclusions predicted or generated by the model. This represents a semantic similarity function used to measure the consistency between two conclusions. This is the threshold for similarity determination.
[0100] like Figure 1 LLM analysis of the immunohistochemical results of breast cancer pathology images from our hospital and Figure 2 The LLM analysis revealed the results of high-risk factors following early NSCLC surgery at our hospital.
[0101] Personalized treatment plans matched with clinical trials
[0102] 3. Treatment plan recommendations and cost optimization:
[0103] 1) Hierarchical decision-making model (TNM staging, driver gene status, PD-L1 expression). The problem of tumor immunotherapy is decomposed into multiple levels and sub-problems. Through layer-by-layer analysis and solution, a decision on the tumor treatment plan is finally made.
[0104] 2) Economic Analysis Engine (Medical Insurance Reimbursement Ratio, Out-of-Position Calculation). An economic analysis engine is introduced to help clinicians, medical insurance payers, and patients find the optimal decision-making path between efficacy and cost. This includes three key technologies: cost-benefit analysis, budget impact model, and dynamic adjustment of out-of-pocket payment ratio. This upgrades immunotherapy decision-making from a "simply biomarker-driven" approach to a "value-based care" model, balancing efficacy, cost, and social equity.
[0105] 3) Dynamic Prognostic Prediction: Combining the Cox proportional hazards model with XGBoost. The Cox proportional hazards model includes the following key assumptions: Proportional hazards assumption: The impact of covariates on risk is proportional (verifiable through the Schoenfeld residual test). Linear additivity: Covariates are linearly related to log hazard. It plays a core role in quantifying the risk-reward ratio in this system. It serves as evidence supporting recommendations, linking the Cox hazard ratio to guidelines. For the XGBoost model: An ensemble learning algorithm based on gradient boosting decision trees (GBDT) iteratively trains weak classifiers (decision trees) and combines them with weights to ultimately form a strong predictive model. Its core formula is:
[0106] , formula 10
[0107] 4. Clinical trial matching:
[0108] 1) Semantic Matching Algorithm: BERT model extracts standard features from clinical trials. A standard clinical trial terminology database (such as MedDRA, WHO Drug) is collected. Features are extracted through the process of "BERT Tokenizer → BERT model → [CLS]-tagged hidden states → feature vectors". An attention mechanism is used to enhance the weight of key features. Based on the cosine similarity between the text feature vector and the standard terminology feature vector, the best semantic association is selected. The core formula is: Let the patient feature embedding vector be u, and the trial inclusion standard embedding vector be v:
[0109] Formula 11; where, An embedding vector representing patient features. The embedding vector represents the criteria for inclusion in clinical trials; Representing vectors and The inner product; and Representing vectors respectively and The norm (usually L2 norm); The cosine similarity between the two is used to measure the similarity between patient characteristics and trial inclusion criteria in the semantic space.
[0110] 2) Decentralized Recruitment: Simulating Blockchain Technology for Cross-Institutional Data Privacy Protection. A decentralized clinical trial participant recruitment platform is built using simulated blockchain technology, enabling cross-institutional data sharing while ensuring patient privacy and security, thus solving problems such as inefficiency and data silos in traditional recruitment models.
[0111] Patient education, intelligent question-and-answer and toxicity management interactive system
[0112] 1. Personalized Educational Content Generation: LLM generates differentiated science popularization content, supporting multiple languages (such as Tibetan and Uyghur). It adopts a multilingual LLM (such as mT5, BLOOM, or NLLB) as its infrastructure. For low-resource languages like Tibetan and Uyghur, it combines Hybrid Expert (MoE) technology to achieve efficient parameter utilization and uses an adapter for lightweight fine-tuning optimization.
[0113] 2. Intelligent Question Answering System:
[0114] 1) Intent Recognition and Slot Filling: BiLSTM-CRF Model. Bidirectional Long Short-Term Memory Network-Conditional Random Field is a classic sequence labeling model, particularly suitable for intent recognition and slot filling tasks in natural language processing. This model can simultaneously capture contextual features and dependencies between labels.
[0115] 2) Sentiment Analysis: RoBERTa Model. As an improved version of BERT, it performs excellently in sentiment analysis tasks through more refined training strategies. This solution uses RoBERTa as its basic architecture to perform sentiment polarity analysis for different application scenarios. A dynamic masking mechanism significantly improves the model's robustness, supporting large-scale training and processing of longer sequences. Its core formula is:
[0116] Attention mechanism: ,Formula 2
[0117] Sentence representation extraction: , formula 12
[0118] Classification prediction: , Formula 13
[0119] Classification loss: , Formula 14
[0120] in, This represents the class probability distribution vector predicted by the model. This represents the weight matrix of the classification layer. This represents the feature representation vector input to the classification layer. Indicates the bias term. This represents the normalization exponential function, used to map the output to a probability distribution; Represents the classification loss function. The first one representing the true label Class-based value (usually one-hot encoded). Represents the first element in the prediction probability vector. The probability of a class.
[0121] 3. Intelligent toxicity management:
[0122] 1) Real-time monitoring and early warning: NLP analysis of patient complaints and laboratory indicators. Integrating natural language processing (patient complaint text) with structured data...
[0123] 2) Based on (laboratory indicators), construct a real-time early warning system for disease deterioration to achieve early identification and intervention of clinical risks.
[0124] 3) Doctor-Patient Collaboration Platform: Doctor-side decision dashboard and patient-side app. A multi-platform architecture is developed, with cross-platform development supported by Flutter for the mobile app. A dedicated medical UI component library is developed to ensure the accuracy of medical information transmission and expression. The server-side uses Spring Security for medical data filtering to ensure the security of medical information transmission. At the data layer, medical data is stored digitally, and real-time data is synchronized with Firebase to ensure the timeliness and security of medical information.
[0125] Clinical validation and optimization of decision support systems
[0126] 1. Clinical Pilot: Conduct a randomized controlled trial (RCT) with our hospital. Evaluation indicators include indication screening accuracy, progression-free survival (PFS), out-of-pocket patient costs, and irreversible adverse events (irAEs) identification rate. The experimental and control groups will undergo interventions according to the following workflows: "Patient enrollment → automated CDSS screening → multimodal risk prediction → treatment suggestion generation → physician decision-making → intelligent follow-up," and "routine treatment → physician experience judgment → standard protocol selection → manual follow-up," respectively. The accuracy of the central blinded assessment decision support system will be evaluated.
[0127] 2. Adaptation to primary healthcare:
[0128] 1) Lightweight model development (parameter compression): For primary healthcare institutions that do not have large-scale integrated systems, compressed models that have undergone model distillation are specifically equipped to reduce the number of parameters while preserving the model's inference speed and accuracy as much as possible. At the same time, edge computing is used for deployment, and samples with confidence levels below the threshold can be retrieved from the cloud.
[0129] The remote MDT collaboration platform deploys a hybrid inference pipeline in primary healthcare information systems to synchronize multimodal data while optimizing performance metrics.
Claims
1. A decision support system for individualized anti-PD-1 treatment of non-small cell lung cancer, characterized in that, Based on the integration of multi-source clinical data and artificial intelligence, including: The data acquisition and standardization module is used to collect multi-source data from patients receiving anti-PD-1 therapy at our hospital and partner hospitals within a specified time period. This multi-source data includes electronic health records, imaging data, gene sequencing data, tumor microenvironment characteristics, blood biomarkers, medical insurance policy documents, and clinical guidelines. It is used for: Using the HL7 FHIR standard, structured data from clinical texts are extracted using natural language processing tools, mapping rules between texts and FHIR resources are established, and parameterized tables or feature scales are formed. Image data is parsed based on the DICOM protocol, and metadata and pixel data are parsed in the order of data elements. The data is then stored in FHIR and PACS in a mixed manner, with both structured data and large pixel data. Using the SNOMED CT terminology system, and with the help of natural language processing technology, deep learning models and medical ontology, the target, staining results, expression sites and positive ratio information in the immunohistochemistry report are extracted in a structured manner and mapped into standard SNOMEDCT codes. The dynamic knowledge graph construction module is used to integrate standardized data into the Neo4j graph database to construct a dynamic knowledge graph. The dynamic knowledge graph stores structured data, including clinical guidelines, drug indications, real-world data, and medical insurance rules, in terms of nodes and relationships. The incremental learning module is used to capture and integrate new knowledge in real time through active learning technology, including the latest clinical trials and drug resistance mechanism research. Through change detection and capture, incremental knowledge extraction, and conflict detection and resolution, the system can continuously update the knowledge base content without retraining the entire model. An intelligent decision support module based on a large language model is used to provide decision support for immunotherapy, including: A localized large language model unit is used to deploy and run a domestically developed large language model fine-tuned through transfer learning, wherein the fine-tuning injects domain knowledge related to anti-PD-1 treatment; A multimodal input processing unit is used to fuse image features, genomic data, clinical text, and economic parameters using a Transformer-based cross-modal encoder, and to perform multimodal fusion through a multi-head self-attention mechanism; The immunotherapy decision optimization unit is used to perform indication screening, efficacy prediction, side effect prediction, conflict detection, and decision optimization. The personalized treatment plan and clinical trial matching unit is used to perform treatment plan recommendations and cost optimization, as well as to match patients with clinical trials through semantic matching algorithms; The patient education and interaction module is used to generate personalized educational content, provide intelligent question-and-answer services, and conduct intelligent toxicity management. The clinical validation and optimization module is used to clinically validate the system through randomized controlled trials and supports lightweight adaptation to primary healthcare institutions.
2. The decision support system for individualized anti-PD-1 treatment of non-small cell lung cancer according to claim 1, characterized in that, In the data acquisition and standardization module: When parsing image data based on the DICOM protocol, to accurately identify blurred tumor boundaries, a boundary loss function based on an integral architecture is used during training to minimize the distance between the predicted boundary and the labeled boundary. This boundary loss function L... BD As shown below: ,in Boundary loss function, Ω: Image domain Spatial location The probability value output by the model. : Distance transformation function for the true labeled boundary.
3. The decision support system for individualized anti-PD-1 treatment of non-small cell lung cancer according to claim 1, characterized in that, The immunotherapy decision optimization unit includes: The indication screening subunit is used to execute a stratified decision-making model based on PD-L1, TMB, and MSI, and to make stratified decisions based on the predicted response to anti-PD-1 treatment. The efficacy prediction subunit is used to dynamically predict the response rate using a time-series Transformer model. Its model input is a time-series feature sequence. The output is the future state. The study also incorporates positional encoding and uses an XGBoost model to analyze drug resistance risk. The objective function of the XGBoost model is shown below: , ;in, The loss function represents the difference between predicted and actual values. For the first The complexity penalty term for trees, where, The number of leaf nodes, and These represent structural complexity and weight regularization coefficients, respectively. The weights are those of the leaf nodes; The side effect prediction subunit is used to model the risk of immune-related adverse reactions using a graph neural network. It constructs a patient-drug-symptom triple graph G=(V,E) and learns node representations based on the update rules of a simplified graph convolutional network, as shown below: ;in, Represents a node In the Layer feature representation, For its set of neighboring nodes, For learnable weight matrix, The normalization coefficient is... The activation function is nonlinear; this method captures the complex relationships between patients, drugs, and adverse reactions through multi-layer information propagation, thereby achieving accurate modeling of the risk of immune-related adverse reactions. The conflict detection and decision optimization subunit is used to jointly verify logical contradictions through a rule engine and a large language model, and to penalize contradictory errors. Its verification process is formalized as follows: ; in: Indicates the first Rule 1 and These represent the premise and conclusion of the rule, respectively. This indicates a logical implication relationship; Represents a set of patient contextual information. This indicates that the current patient status meets the prerequisites of the rule; This indicates the conclusions predicted or generated by the model. This represents a semantic similarity function used to measure the consistency between two conclusions. This is the threshold for similarity determination.
4. The decision support system for individualized anti-PD-1 treatment of non-small cell lung cancer according to claim 1, characterized in that, The individualized treatment plan and clinical trial matching unit includes: The treatment plan recommendation and cost optimization subunit is used to make tumor treatment plan decisions through a hierarchical decision model, combined with TNM staging, driver gene status, and PD-L1 expression; and to perform cost-benefit analysis, budget impact model calculation, and dynamic adjustment of out-of-pocket ratio through an economic analysis engine to achieve cost optimization. The dynamic prognostic prediction subunit is used to combine the COX proportional hazards model and the XGBoost model for dynamic prognostic prediction. The clinical trial matching subunit is used to extract standard features of clinical trials using the BERT model and perform semantic matching based on the cosine similarity between the patient feature embedding vector and the trial-incorporated standard embedding vector. The similarity calculation is as follows: ;in, An embedding vector representing patient features. The embedding vector represents the criteria for inclusion in clinical trials; Representing vectors and The inner product; and Representing vectors respectively and The norm; The cosine similarity between the two is used to measure the degree of similarity between patient characteristics and trial inclusion criteria in the semantic space. It also uses simulated blockchain technology to achieve cross-institutional data privacy protection for decentralized recruitment.
5. The decision support system for individualized anti-PD-1 treatment of non-small cell lung cancer according to claim 1, characterized in that, The patient education and interaction module includes: The personalized educational content generation unit is used to generate differentiated popular science content using a multilingual large language model and supports low-resource languages. The intelligent question-answering unit is used for intent recognition and slot filling using a bidirectional long short-term memory network-conditional random field model, and for sentiment analysis using a RoBERTa model. The classification prediction of the RoBERTa model is shown in the formula. As shown in the formula, the classification loss is... As shown; where, This represents the probability distribution vector of the classes predicted by the model. This represents the weight matrix of the classification layer. This represents the feature representation vector input to the classification layer. Indicates the bias term. This represents the normalization exponential function, used to map the output to a probability distribution; Represents the classification loss function. The first one representing the true label Class value, Represents the first element in the prediction probability vector. The probability of a class; The intelligent toxicity management unit is used to analyze patient complaints through natural language processing and combine them with laboratory indicators to achieve real-time monitoring and early warning; and to provide a doctor-patient collaboration platform including a doctor-side decision dashboard and a patient-side APP.
6. The decision support system for individualized anti-PD-1 treatment of non-small cell lung cancer according to claim 1, characterized in that, In the multimodal input processing unit, the core formula of the multi-head self-attention mechanism is as follows: ; ; ,in, , , These represent the query, key, and value matrices, respectively. This represents the product of the query matrix and the transpose of the key matrix; This represents the dimension of the key vector, used to scale the inner product result; This represents the normalized exponential function, used to map attention weights to a probability distribution; This represents the result of attention calculation; Indicates the first The output of each attention head; , , They represent the first The linear transformation weight matrix used to generate queries, keys, and values in each attention head; This indicates a concatenation operation of multiple attention head outputs; This represents the linear transformation weight matrix after multi-head output; This indicates the number of attention heads.
7. The decision support system for individualized anti-PD-1 treatment of non-small cell lung cancer according to claim 1, characterized in that, In the efficacy prediction subunit, the position encoding of the time series Transformer model is as follows: , ;in, and Representing positions respectively First With the The encoded value of the dimension; Indicates the position index in the sequence; Indicates the feature dimension index; This represents the total dimension of the positional encoding vector.
8. The decision support system for individualized anti-PD-1 treatment of non-small cell lung cancer according to claim 1, characterized in that, The clinical validation and optimization module also includes a primary healthcare adaptation unit, used for: Lightweight models are developed through model distillation and deployed on edge computing devices, allowing samples with confidence levels below a threshold to seek help from the cloud. Provide a remote MDT collaboration platform to deploy hybrid inference pipelines in primary healthcare information systems and synchronize multimodal data.
9. An immunotherapy decision support method based on the system according to any one of claims 1 to 8, characterized in that, Includes the following steps: Data acquisition and standardization steps: Collect multi-source clinical data and standardize them using the HL7 FHIR standard, DICOM protocol and SNOMED CT terminology system; Knowledge base construction and update steps: Construct a dynamic knowledge graph and update the knowledge base content in real time through an incremental learning mechanism; Intelligent decision support steps: Based on a locally deployed large language model, multimodal input data is fused and processed, and indication screening, efficacy prediction, side effect prediction, conflict detection and decision optimization are performed to generate individualized treatment plans, match clinical trials, and optimize costs. Patient interaction and management steps: Generate personalized educational content, provide intelligent Q&A services, and conduct intelligent toxicity management; Clinical validation and optimization steps: Conduct clinical validation through randomized controlled trials, and optimize the system based on the validation results.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in claim 9.