Method and system for managing side effects of medication for cancer patients based on big data analysis

By constructing a time series analysis sub-model, a gene-drug interaction analysis sub-model and a risk score generation sub-model, the adaptability and accuracy of side effects prediction during drug use in tumor patients is solved, and accurate prediction and personalized management of side effects risks in tumor patients are achieved.

CN119446400BActive Publication Date: 2025-05-13XIAMEN COBBLESTONE NETWORK TECH CO LTD

Patent Information

Application Number
CN202510035788.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-13
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The prior art has problems such as adaptability and low prediction accuracy in predicting side effects caused by tumor patients during medication use.

Method used

Using the drug side effect management method for tumor patients based on big data analysis, we use the collection and preprocessing of patient data and drug data, and construct a time series analysis sub-model, gene-drug interaction analysis sub-model and risk score generation sub-model to predict the risk scores for patients with different types of side effects, and provide personalized adjustment suggestions based on the risk assessment results.

Benefits of technology

Effectively predict and manage the risk of side effects that may occur in the medication process by tumor patients, improve the accuracy of side effects prediction, adapt to the needs of side effects management of medication during the treatment stage, reduce the risk of treatment interruptions and complications caused by side effects, and optimize the patient's medication experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446400B_ABST
    Figure CN119446400B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for managing the side effects of medication for tumor patients based on big data analysis, the method comprising the following steps: collecting and preprocessing patient data and drug data, extracting patient characteristics and drug characteristics, wherein the patient data includes genomic data provided by the patient; constructing a prediction model, including a time series analysis sub-model, a gene-drug interaction analysis sub-model and a risk score generation sub-model, for using drug characteristics and patient characteristics as inputs of the prediction model to predict the risk score of different types of side effects in the patient; inputting the basic information, medication information and genomic data of the current patient into the trained prediction model to obtain a risk assessment of various side effects in the current patient; according to the risk assessment results, selecting a personalized adjustment suggestion suitable for the patient from a preset intervention program library and providing it to the current patient. The present invention can effectively improve the scientificity and accuracy of the management of medication side effects during the treatment stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of drug side effects, and in particular relates to a method and system for managing drug side effects in tumor patients based on big data analysis. Background Art

[0002] Regarding the study of drug side effects, a Chinese invention patent application with publication number CN112216396A discloses a method for predicting drug-side effect relationships based on graph neural networks. The method includes the following steps: collecting drug data for preprocessing, establishing relationships between drugs and drugs, and drug side effects and side effects; constructing a drug-side effect heterogeneous network model; using graph neural networks to vectorize nodes and multi-relationship edges in the network; combining point vectors and edge vectors to represent the structural characteristics of a node in the network; dividing the data set into a training set and a test set; using the obtained node vectors as input for machine learning training to train the parameters of the model; and using the node vectors represented by the graph neural network method on the test set to predict the drug-side effect relationship.

[0003] The above-mentioned scheme focuses on the side effects of drugs during the development of new drugs in the biomedical field. There is a lack of targeted research on the side effects of cancer patients during the treatment of marketed drugs. If it is used to predict side effects during the treatment process, there are problems such as low adaptability and low prediction accuracy. Summary of the invention

[0004] The present invention provides a method and system for managing side effects of medication for tumor patients based on big data analysis, aiming to solve the problems of low adaptability and low prediction accuracy in the prediction of side effects generated during the treatment process in the prior art.

[0005] In order to solve the above technical problems, the medication side effect management method proposed by the present invention comprises the following steps:

[0006] Collect and pre-process patient data and drug data, and extract patient characteristics and drug characteristics, wherein the patient data includes genomic data provided by the patient;

[0007] Constructing a prediction model, the prediction model including a time series analysis sub-model, a gene-drug interaction analysis sub-model and a risk score generation sub-model, for taking drug characteristics and patient characteristics as inputs of the prediction model to predict the risk scores of patients for different types of side effects;

[0008] Input the current patient's basic information, medication information, and genomic data into the trained prediction model to obtain the risk assessment of various side effects for the current patient;

[0009] Based on the risk assessment results, personalized adjustment suggestions suitable for the patient are selected from the preset intervention plan library and provided to the current patient.

[0010] Preferably, the time series analysis sub-model is an autoregressive integrated moving average model, and the construction method is:

[0011] Collect patients' medication data and physiological index data, and format them into time series;

[0012] Calculate drug metabolic kinetics curves and physiological index change curves based on time series;

[0013] The autoregressive integrated moving average model is used to model the drug metabolic kinetic curve and the physiological index change curve;

[0014] Based on the trained autoregressive integrated moving average model, the risk of side effects in future time periods is predicted, and the changes in side effects in the short term are predicted.

[0015] Preferably, the pharmacokinetic curve includes a pharmacokinetic curve and a pharmacodynamic curve, and the pharmacokinetic curve adopts a single compartment model and is expressed as follows:

[0016]

[0017] In the formula, It's time The blood concentration of For dosage, is the distribution volume, To obtain the elimination rate constant, the drug concentration data were fitted using the nonlinear least squares method to obtain the elimination rate constant; the maximum blood drug concentration, peak time, half-life, and overall drug exposure were calculated based on the pharmacokinetic curve model;

[0018] The pharmacodynamic curve adopts the Emax model, which is expressed as follows:

[0019]

[0020] In the formula, It's the efficacy of the medicine. is the maximum effect, is the drug concentration that produces a half-maximal effect, is the drug concentration, fitted by experimental data and .

[0021] Preferably, the gene-drug interaction analysis sub-model is a classification model based on extreme gradient boosting, and the input data of the classification model are gene characteristics, drug characteristics and patient characteristics; the gene characteristics are extracted from genome data, including whether there are mutations in the genome and the impact of gene mutations; the drug characteristics are extracted from drug molecule descriptors in drug data and the binding affinity between drugs and target genes;

[0022] The classification model includes an input layer, a boosted tree layer and a classification target output layer. The input feature dimension of the input layer is the sum of the dimensions of gene features, drug features and patient features; the boosted tree layer includes several decision trees, nodes are split according to the feature importance of gene-drug interaction, residuals are calculated in each split iteration, and leaf node prediction values ​​are updated; the classification target output layer outputs the classification probability of gene-drug interaction, and generates a risk score based on the classification probability and a preset weight coefficient.

[0023] Preferably, the processing flow of the risk score generation sub-model is:

[0024] The output features of the time series analysis sub-model were standardized, and the output features of the gene-drug interaction analysis sub-model were mapped;

[0025] Assign weights to the output features of the time series analysis submodel and the gene-drug interaction analysis submodel;

[0026] The comprehensive risk score is calculated based on the assigned weights as the final risk score output.

[0027] Preferably, the output features of the time series analysis sub-model are standardized by:

[0028]

[0029] In the formula, For the corresponding The standardized value, is the predicted value of short-term side effect changes, is the mean of the output results of the time series analysis sub-model, The standard deviation of the output results of the time series analysis sub-model;

[0030] The output features of the gene-drug interaction analysis sub-model are mapped by:

[0031]

[0032] In the formula, Score the risk level after mapping, The gene-drug interaction risk score output by the gene-drug interaction analysis sub-model, is the minimum score output by the gene-drug interaction analysis sub-model, The maximum score output by the gene-drug interaction analysis submodel.

[0033] Preferably, the comprehensive risk score is calculated as follows:

[0034]

[0035] In the formula, For the comprehensive risk score, is the weight of the time series analysis sub-model, is a nonlinear exponential parameter used to control the sensitivity of different features to the final score. Assigned manually or automatically based on preset rules.

[0036] Preferably, the drug data is collected in a drug database, including the known side effects of the drug, the probability of occurrence, the severity and the corresponding management plan; the patient data is collected in a patient communication platform, including patient information, drug name, dosage, medication time, complications, physiological indicators, patient feedback, and also includes content actively posted by patients on the platform, as well as self-related content mentioned by patients in platform chats.

[0037] Preferably, the genomic data includes genes related to drug metabolism, drug resistance genes, and gene variations that may be associated with drug side effects.

[0038] Another aspect of the present invention is to provide a management system for the side effects of medication for tumor patients based on big data analysis. The system is used to implement the above-mentioned method for managing the side effects of medication, and comprises:

[0039] Data collection and preprocessing module, used to collect patient data and drug data, and clean, format and extract features from the data;

[0040] The time series analysis sub-model is used to calculate the drug metabolic kinetics curve and the physiological index change curve, and use the autoregressive integrated moving average model to predict the change trend of side effects in the short term;

[0041] The gene-drug interaction analysis sub-model, based on the extreme gradient boosting algorithm, analyzes the interaction between gene features, drug features, and patient features, and outputs the gene-drug interaction risk level;

[0042] The risk score generation sub-model is used to integrate the output features of the time series analysis sub-model and the output features of the gene-drug interaction analysis sub-model, and generate the final comprehensive risk score after feature standardization, mapping and weighted calculation;

[0043] The intervention plan recommendation module is used to retrieve personalized adjustment suggestions suitable for patients from the preset intervention plan library based on the comprehensive risk score and feed back the results to the patients;

[0044] The user interaction module provides an interface for users to input basic patient information, medication information, and genomic data, and displays the risk score and personalized adjustment suggestions generated by the system;

[0045] Model training and updating module, which is used to dynamically update the time series analysis sub-model, gene-drug interaction analysis sub-model and risk score generation sub-model according to the newly collected data to optimize the model prediction performance;

[0046] The system interface module provides interfaces with external drug databases, genome analysis tools, and medical information systems for data exchange and system integration.

[0047] Compared with the prior art, the present invention has the following technical effects:

[0048] 1. The medication side effect management method proposed in the present invention can effectively predict and manage the side effect risks that may occur in cancer patients during medication by combining time series analysis, gene-drug interaction analysis and risk score generation, accurately predict short-term side effect changes, improve the accuracy of side effect prediction, and has stronger adaptability to medication side effect management during the treatment stage.

[0049] 2. The medication side effect management method proposed in this invention comprehensively integrates patient genomic data, drug properties and patient characteristics, and accurately identifies the impact of gene variation on drug metabolism and side effects. By integrating the results of time series analysis and gene-drug interaction analysis, the scientificity and accuracy of risk assessment in individualized medication side effect management are significantly improved.

[0050] 3. The medication side effect management method proposed in the present invention can dynamically adjust the side effect risk assessment and intervention strategy by real-time monitoring of the patient's physiological indicators and medication feedback, combined with the prediction results of the time series analysis model and the gene-drug interaction model, to adapt to the needs of complex clinical scenarios and improve the flexibility and applicability of medication management.

[0051] 4. The medication side effect management method proposed in the present invention can effectively reduce the risks of treatment interruption and serious complications caused by side effects, and optimize the patient's medication experience. While ensuring the safety of patients, it provides scientific decision-making support for medical staff, helps the smooth implementation of tumor treatment, and ultimately significantly improves the overall quality of life of patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic diagram of the process of the side effect management method of the present invention. DETAILED DESCRIPTION

[0053] In order to make the objectives, technical solutions and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in combination with specific embodiments of the present application and with reference to the accompanying drawings.

[0054] Embodiment 1

[0055] This embodiment is a method for managing the side effects of medication for tumor patients based on big data analysis. Figure 1 As shown, the following steps are included:

[0056] Collect and pre-process patient data and drug data, and extract patient characteristics and drug characteristics, wherein the patient data includes genomic data provided by the patient; wherein the genomic data includes genes related to drug metabolism (such as CYP450 family genes), drug resistance genes, and gene variations that may be related to drug side effects.

[0057] The drug data is collected in a drug database, including but not limited to known side effects of the drug, probability of occurrence, severity and corresponding management plans, as well as relevant drug-gene interaction information.

[0058] The patient data is collected on the patient communication platform, including but not limited to patient information, drug name, dosage, medication time, complications, physiological indicators, patient feedback, as well as content actively posted by patients on the above-mentioned patient communication platform, and self-related content mentioned by patients in chats on the above-mentioned patient communication platform. The text information provided by patients when actively posting on the communication platform includes: medication experience, describing the side effects of drug use; physical condition, mentioning recent health changes or symptoms; help information, asking other patients or professionals for advice on drug side effects or relief methods; timeline information, the correspondence between posting time and medication time. The patient's chat content is the patient's own situation mentioned in conversations with other users or platform robots, including: symptoms and medication mentioned when chatting with others; expressing concerns about certain symptoms; describing intervention factors in daily life; mentioned psychological feelings, etc.

[0059] Specifically, the sources of drug data may include, but are not limited to: public databases (such as FDA Adverse Event Reporting System FAERS), drug instructions, medical literature, clinical trial data, hospital electronic medical record systems (authorization required), etc. The following information is obtained, including but not limited to: drug name, chemical structure, mechanism of action, indications, contraindications, adverse reactions (including frequency, severity, associated symptoms, etc.), interactions, pharmacokinetic data, clinical trial data, relevant domestic and foreign guidelines and literature information, etc.

[0060] It can be understood by those skilled in the art that the preprocessing includes data cleaning, data standardization, and data fusion. Data cleaning is used to remove redundant information and only retain gene locus data related to drug metabolism, drug resistance, and side effects. Data standardization performs structured processing on basic patient information and medication records to ensure consistency with the format of the database and analysis module. Data fusion integrates patient information, genomic data, and drug-gene interaction data in the drug database to form a multi-dimensional analysis input.

[0061] A prediction model is constructed, which includes a time series analysis sub-model, a gene-drug interaction analysis sub-model and a risk score generation sub-model, which is used to use drug characteristics and patient characteristics as inputs of the prediction model to predict the risk scores of patients for different types of side effects.

[0062] Among them, the time series analysis sub-model combines the patient's medication time point and dosage to construct a dynamic medication pattern and predict the time-dependent side effects that may be caused by the drug. In this embodiment, the time series analysis sub-model is an AutoRegressive Integrated Moving Average (ARIMA) model, and the construction and use method is:

[0063] Collect patients' medication data and physiological index data, format and construct them into time series, align them, fill in missing values ​​and eliminate noise;

[0064] Calculate the drug metabolism kinetics curve (PK / PD data) and physiological index change curve based on the treated time series;

[0065] The autoregressive integrated moving average model is used to model the drug metabolic kinetic curve and the physiological index change curve;

[0066] Based on the trained autoregressive integrated moving average model, the risk of side effects in future time periods is predicted, and the changes in side effects in the short term are predicted.

[0067] Specifically, the drug metabolism kinetic curve includes a pharmacokinetic curve (PK) and a pharmacodynamic curve (PD). PK / PD data is used to quantitatively describe the change of drug concentration over time (PK) and the relationship between drug concentration and effect or side effect (PD). It is used as a feature in the model to predict the risk of side effects. The pharmacokinetic curve adopts a single-compartment model and is expressed as follows:

[0068]

[0069] In the formula, It's time The blood concentration of For dosage, is the distribution volume, To obtain the elimination rate constant, the drug concentration data were fitted using the nonlinear least squares method to obtain the elimination rate constant; the maximum blood drug concentration, peak time, half-life, and overall drug exposure were calculated based on the pharmacokinetic curve model;

[0070] The pharmacodynamic curve adopts the Emax model, which is expressed as follows:

[0071]

[0072] In the formula, It's the efficacy of the medicine. is the maximum effect, is the drug concentration that produces a half-maximal effect, is the drug concentration, fitted by experimental data and .

[0073] In some other embodiments of the present invention, the time series analysis sub-model adopts a Long Short-Term Memory (LSTM) network.

[0074] In this embodiment, the gene-drug interaction analysis sub-model is a multi-classification model based on Extreme Gradient Boosting (XGBoost) to predict the impact of gene variation on drug metabolism and side effect risk. XGBoost is an efficient gradient boosting tree algorithm that can handle complex nonlinear relationships and is suitable for gene-drug interaction scenarios. The input data of the classification model are gene characteristics, drug characteristics, and patient characteristics; the gene characteristics are extracted from genomic data, including whether there are mutations in the genome and the impact of gene mutations; the drug characteristics are extracted from the drug molecule descriptors in the drug data and the binding affinity between the drug and the target gene;

[0075] The classification model includes an input layer, a boosting tree layer and a classification target output layer. The input features of the input layer are patient genome data (such as variation information of CYP450, UGT, and SLCO family genes), drug gene interaction databases such as drug metabolism genes, target genes, and association relationships and drug structure characteristics such as SMILES expressions, molecular feature vectors, etc.; the input feature dimension is the sum of the dimensions of gene features, drug features, and patient features; the boosting tree layer includes several decision trees, which split nodes according to the feature importance of gene-drug interaction, calculate residuals in each split iteration, and update leaf node prediction values; the classification target output layer outputs the classification probability of gene-drug interaction, and generates a risk score based on the classification probability and a preset weight coefficient.

[0076] Specifically, the objective function of XGBoost is set to multi:softprob, which is used to output the probability that each sample belongs to each category. The output of each sample is a probability vector, and the sum of the probabilities of all categories is 1. At the same time, a weight coefficient is assigned to each category. In this embodiment, the probability of four categories is output, and the assigned weight coefficients are 1, 4, 7, and 10, respectively, representing the risk-free category, low-risk category, medium-risk category, and high-risk category. Finally, according to the probability of the XGBoost classification output and the assigned weight coefficient, the risk score of the final output of the classification target output layer is weighted. If the probability of the XGBoost model output is [0.25, 0.55, 0.15, 0.05], combined with the assigned weight coefficient, the final output risk score is 0.25*1+0.55*4+0.15*7+0.05*10=4.

[0077] The construction of the XGBoost model includes steps such as data collection, data preprocessing, data segmentation, and model training.

[0078] The data collected includes genomic data, drug data and patient data.

[0079] In the data preprocessing step, for gene features, it is necessary to check whether the gene variant exists, using a Boolean value to represent it, and to check the impact of the gene variant, using the functional score of non-synonymous mutations. For drug features, extract drug molecular descriptors (such as molecular weight, LogP value, number of hydrogen bond donors / acceptors) and the binding affinity of the drug to the target gene. Then construct gene-drug interaction features, such as the metabolic impact weight of the gene variant and the interaction score of the drug metabolizing enzyme.

[0080] The constructed feature data were randomly divided into a training set (70%), a validation set (15%), and a test set (15%) to ensure balanced distribution of samples from different patients.

[0081] During the model training process, this embodiment selects logarithmic loss as the loss function, uses F1 score as the evaluation indicator of the model, and adjusts parameters such as the tree depth, learning rate, number of weak learners, and feature sampling rate in the XGBoost model.

[0082] The classification model outputs the risk category of gene-drug interaction:

[0083] Low risk, normal metabolism, no significant risk of side effects;

[0084] Medium risk, mild side effects may occur and need to be monitored;

[0085] High risk, significant risk of side effects, recommendation to adjust medication or dosage.

[0086] In some other embodiments of the present invention, the gene-drug interaction analysis sub-model adopts a random forest model or a graph neural network (GNN).

[0087] When constructing the gene-drug interaction analysis sub-model in this embodiment, data from the DGIdb database (Drug-Gene Interaction database) or other related databases can be used, wherein DGIdb is a drug-gene interaction database that provides information on the association between genes and their known or potential drugs. The genes in it are mainly oncogenes, but there are also genes related to other diseases (such as Alzheimer's disease, heart disease, diabetes, etc.). DGIdb has a total of more than 14,000 drug-gene interactions, involving 2,600 genes and 6,300 drugs targeting these genes, as well as 6,700 other genes.

[0088] In this embodiment, the processing flow of the risk score generation sub-model is as follows:

[0089] The output features of the time series analysis sub-model were standardized, and the output features of the gene-drug interaction analysis sub-model were mapped;

[0090] Assign weights to the output features of the time series analysis submodel and the gene-drug interaction analysis submodel;

[0091] The comprehensive risk score is calculated based on the assigned weights as the final risk score output.

[0092] The method for standardizing the output features of the time series analysis sub-model is as follows:

[0093]

[0094] In the formula, For the corresponding The standardized value, is the predicted value of short-term side effect changes, is the mean of the output results of the time series analysis sub-model, The standard deviation of the output results of the time series analysis sub-model;

[0095] The output features of the gene-drug interaction analysis sub-model are mapped by:

[0096]

[0097] In the formula, Score the risk level after mapping, The gene-drug interaction risk score output by the gene-drug interaction analysis sub-model, is the minimum score output by the gene-drug interaction analysis sub-model, The maximum score output by the gene-drug interaction analysis submodel.

[0098] The calculation method of the comprehensive risk score is:

[0099]

[0100] In the formula, For the comprehensive risk score, is the weight of the time series analysis sub-model, is a nonlinear exponential parameter used to control the sensitivity of different features to the final score. Assigned manually or automatically based on preset rules.

[0101] In some other embodiments of the present invention, the calculation method of the comprehensive risk score is a weighted average method, which assigns different weights to the time series analysis sub-model and the gene-drug interaction analysis sub-model, and the sum of their weights is 1. The weighted values ​​of the two are calculated as the final output risk score.

[0102] The current patient's basic information, medication information and genomic data are input into the trained prediction model to obtain a risk assessment of various side effects for the current patient.

[0103] If the comprehensive risk score is high (greater than 0.6) during the real-time evaluation, it means that the risk of side effects in the short term increases, and the gene interaction shows a high risk. It is recommended to adjust the drug or dosage immediately, that is, to enter the personalized adjustment suggestion generation stage, generate personalized suggestions for the patient, and communicate and handle the matter between the patient and the doctor.

[0104] If the comprehensive risk score is a medium score (0.3–0.6), it means that the side effects do not change significantly, but there is a certain risk of genetic interaction, and it is recommended to closely monitor the patient's response.

[0105] If the comprehensive risk score is low (less than 0.3), it means that both side effect changes and gene interactions show low risk, and it is recommended to maintain the current medication regimen.

[0106] Among them, the dividing points for high scores, medium scores and low scores are 0.6 and 0.3. Different dividing points can be selected according to different situations.

[0107] Based on the risk assessment results, such as the high scores mentioned above, personalized adjustment suggestions suitable for the patient are selected from the preset intervention plan library and provided to the current patient.

[0108] Specifically, according to the side effect risk score, the corresponding management plan in the drug database is queried. At the same time, combined with the patient's genetic variation, health status and side effect risk, targeted medication adjustment suggestions are generated, such as dose adjustment, combination therapy or alternative drug regimens, and other non-drug intervention suggestions.

[0109] Personalized adjustment suggestions include drug intervention suggestions, as well as personalized non-drug intervention suggestions in combination with health management guidelines. It should be noted that drug intervention suggestions need to be implemented based on the doctor's advice after comprehensive evaluation by the patient and the doctor.

[0110] Non-drug interventions include dietary management, which recommends a low-fat, high-protein diet for side effects caused by specific drugs (such as liver damage); physical activity, which recommends moderate exercise therapy based on the patient's physiological data to relieve symptoms such as fatigue or muscle atrophy; psychological counseling, which recommends psychological counseling or relaxation therapy for patients who suffer from anxiety or depression due to long-term side effects. Patients are also advised to communicate and evaluate with their doctors.

[0111] Ultimately, personalized intervention suggestions are pushed to patients through the patient-side application of the communication platform, including text descriptions and visual side effect risk prompts. The patient synchronizes the personalized medication suggestions to the doctor, who confirms and implements them. In this embodiment, feedback data from patients after medication, including actual side effects and efficacy, can also be collected to optimize the above prediction model.

[0112] Embodiment 2

[0113] This embodiment is a management system for the side effects of medication for tumor patients based on big data analysis. The system is used to implement the method for managing the side effects of medication as described in the first embodiment, including:

[0114] The data collection and preprocessing module is used to collect patient data and drug data, and clean, format and extract features from the data.

[0115] The time series analysis sub-model is used to calculate the drug metabolic kinetics curve and the physiological index change curve, and the autoregressive integrated moving average model is used to predict the changing trend of side effects in the short term.

[0116] The gene-drug interaction analysis sub-model, based on the extreme gradient boosting algorithm, analyzes the interaction between gene characteristics, drug characteristics and patient characteristics, and outputs the gene-drug interaction risk level.

[0117] The risk score generation sub-model is used to integrate the output features of the time series analysis sub-model and the output features of the gene-drug interaction analysis sub-model. After feature standardization, mapping and weighted calculation, the final comprehensive risk score is generated.

[0118] The intervention plan recommendation module is used to retrieve personalized adjustment suggestions suitable for patients from the preset intervention plan library based on the comprehensive risk score, and feed back the results to the patients.

[0119] The user interaction module provides users with an interface for inputting basic patient information, medication information, and genomic data, and displays the risk scores and personalized adjustment suggestions generated by the system.

[0120] The model training and updating module is used to dynamically update the time series analysis sub-model, gene-drug interaction analysis sub-model and risk score generation sub-model according to the newly collected data to optimize the model prediction performance.

[0121] The system interface module provides interfaces with external drug databases, genome analysis tools, and medical information systems for data exchange and system integration.

[0122] The above is only a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several modifications and improvements without departing from the inventive concept of the present invention, which all belong to the protection scope of the present invention.

Claims

1. A method for managing the side effects of medication for cancer patients based on big data analysis, characterized in that: The following steps are involved: Collect and pre-process patient data and drug data, and extract patient characteristics and drug characteristics, wherein the patient data includes genomic data provided by the patient; Constructing a prediction model, the prediction model includes a time series analysis sub-model, a gene-drug interaction analysis sub-model and a risk score generation sub-model; wherein the time series analysis sub-model is an autoregressive integrated moving average model, and the construction method is: collecting the patient's medication data and physiological index data, formatting and constructing them into a time series, calculating the drug metabolic kinetic curve and the physiological index change curve according to the time series, using the autoregressive integrated moving average model to model the drug metabolic kinetic curve and the physiological index change curve, and based on the trained autoregressive integrated moving average model, predicting the risk of side effects in the future time period, and predicting the changes in side effects in the short term; The gene-drug interaction analysis sub-model is a classification model based on extreme gradient boosting, and the input data of the classification model are gene features, drug features and patient features; the gene features are extracted from genomic data; the drug features are extracted from drug molecule descriptors in drug data and the binding affinity between drugs and target genes; The classification model includes an input layer, a boosted tree layer and a classification target output layer. The input feature dimension of the input layer is the sum of the dimensions of gene features, drug features and patient features. The boosted tree layer includes a number of decision trees. The nodes are split according to the feature importance of gene-drug interaction. The residual is calculated in each split iteration and the leaf node prediction value is updated. The classification target output layer outputs the classification probability of gene-drug interaction and generates a risk score according to the classification probability and a preset weight coefficient. The processing flow of the risk score generation sub-model is as follows: standardizing the output features of the time series analysis sub-model, mapping the output features of the gene-drug interaction analysis sub-model, assigning weights to the output features of the time series analysis sub-model and the gene-drug interaction analysis sub-model, and calculating the comprehensive risk score according to the assigned weights as the final side effect risk score output; Input the current patient's basic information, medication information, and genomic data into the trained prediction model to obtain a comprehensive risk score; Based on the comprehensive risk score, personalized adjustment suggestions suitable for the patient are selected from the preset intervention plan library and provided to the current patient.

2. The method for managing side effects of medication for tumor patients based on big data analysis according to claim 1, characterized in that: The drug metabolism kinetics curve includes a pharmacokinetic curve and a pharmacodynamic curve. The pharmacokinetic curve adopts a single-compartment model and is expressed as follows: In the formula, It's time The blood concentration of For dosage, is the distribution volume, To obtain the elimination rate constant, the drug concentration data were fitted using the nonlinear least squares method to obtain the elimination rate constant; the maximum blood drug concentration, peak time, half-life, and overall drug exposure were calculated based on the pharmacokinetic curve model; The pharmacodynamic curve adopts the Emax model, which is expressed as follows: In the formula, It's the efficacy of the medicine. is the maximum effect, is the drug concentration that produces a half-maximal effect, is the drug concentration, fitted by experimental data and .

3. The method for managing side effects of medication for tumor patients based on big data analysis according to claim 1, characterized in that: The method for standardizing the output features of the time series analysis sub-model is as follows: In the formula, For the corresponding The standardized value, is the predicted value of short-term side effect changes, is the mean of the output results of the time series analysis sub-model, The standard deviation of the output results of the time series analysis sub-model; The output features of the gene-drug interaction analysis sub-model are mapped by: In the formula, Score the risk level after mapping, The gene-drug interaction risk score output by the gene-drug interaction analysis sub-model, is the minimum score output by the gene-drug interaction analysis sub-model, The maximum score output by the gene-drug interaction analysis submodel.

4. The method for managing side effects of medication for tumor patients based on big data analysis according to claim 3, characterized in that: The calculation method of the comprehensive risk score is: In the formula, For the comprehensive risk score, is the weight of the time series analysis sub-model, is a nonlinear exponential parameter used to control the sensitivity of different features to the final score. Assigned manually or automatically based on preset rules.

5. A drug side effect management system for cancer patients based on big data analysis, characterized in that: The system is used to implement the method according to any one of claims 1 to 4, comprising: Data collection and preprocessing module, used to collect patient data and drug data, and clean, format and extract features from the data; The time series analysis sub-model is used to calculate the drug metabolic kinetics curve and the physiological index change curve, and use the autoregressive integrated moving average model to predict the change trend of side effects in the short term; The gene-drug interaction analysis sub-model, based on the extreme gradient boosting algorithm, analyzes the interaction between gene features, drug features, and patient features, and outputs a risk score; The risk score generation sub-model is used to integrate the output features of the time series analysis sub-model and the output features of the gene-drug interaction analysis sub-model, and generate the final comprehensive risk score after feature standardization, mapping and weighted calculation; The intervention plan recommendation module is used to retrieve personalized adjustment suggestions suitable for patients from the preset intervention plan library based on the comprehensive risk score and feed back the results to the patients; The user interaction module provides an interface for users to input basic patient information, medication information, and genomic data, and displays the risk score and personalized adjustment suggestions generated by the system; Model training and updating module, which is used to dynamically update the time series analysis sub-model, gene-drug interaction analysis sub-model and risk score generation sub-model according to the newly collected data to optimize the model prediction performance; The system interface module provides interfaces with external drug databases, genome analysis tools, and medical information systems for data exchange and system integration.

Citation Information

Patent Citations

  • Method for predicting drug-side effect relationship based on graph neural network

    CN112216396A

  • Untoward drug reaction monitoring and early warning method

    CN118280603A

  • Adverse drug reaction reporting method

    CN118315082A

  • Multi-index comprehensive scoring patient risk assessment method and system

    CN119049718A

Cited By

  • Intelligent administration instruction execution management system and method driven by controlled cardiovascular parameters

    CN122474258A