Medical medicine curative effect evaluation method based on big data analysis of electronic health record

By integrating multi-source heterogeneous data and dynamic modeling, a subgroup-specific efficacy benchmark matrix is ​​constructed. Individualized assessment reports are generated using deep learning and reinforcement learning, which solves the problems of ignoring individual differences and poor adaptability in traditional assessment methods, and realizes accurate individualized assessment and dynamic adjustment of the efficacy of internal medicine drugs.

CN121075701APending Publication Date: 2025-12-05THE 13TH PEOPLES HOSPITAL OF CHONGQING (CHONGQING GERIATRIC HOSPITAL)

Patent Information

Application Number
CN202511174764.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Traditional methods for evaluating the efficacy of internal medicine drugs cannot effectively integrate multi-source heterogeneous data, cannot achieve individualized evaluation, are difficult to reflect the real-world effects of medication, and have poor universality of evaluation results, making it impossible to track fluctuations in physiological indicators in real time and dynamically adjust medication regimens.

Method used

By extracting data such as gene polymorphism, epigenetic markers, and gut microbiota profiles from electronic health records, and combining them with drug metabolism enzyme activity data for spatiotemporal alignment, hierarchical clustering to form subgroups, using transfer learning to construct a efficacy benchmark matrix, and generating real-time evaluation vectors through deep learning, combined with reinforcement learning for parameter optimization, a personalized evaluation report is generated.

Benefits of technology

It enables precise and individualized assessment of the efficacy of internal medicine drugs, and can track short-term physiological indicators, mid-term symptom improvement and long-term prognostic risks in real time, improving the accuracy and adaptability of the assessment and providing individualized medication recommendations and risk warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075701A_ABST
    Figure CN121075701A_ABST
Patent Text Reader

Abstract

The invention discloses an internal medicine drug curative effect evaluation method based on big data analysis of an electronic health record, and the method comprises the steps: extracting basic health data, diagnosis and treatment time sequence data and drug intervention data from the electronic health record, and carrying out the time-space alignment to generate a dynamic feature set; subgroups are obtained based on disease typing standard hierarchical clustering, and historical data and real world data are fused through transfer learning to construct a subgroup curative effect reference matrix; collecting data after medication in real time, and generating an evaluation vector containing short-term physiological response, middle-term symptom improvement and long-term prognosis risk through deep learning; dynamically matching the evaluation vector with the reference matrix, and introducing an individual weight coefficient to correct deviation; taking the deviation correction value as input, constructing a self-adaptive evaluation model through reinforcement learning, and performing iterative optimization; and generating an individualized report containing the curative effect level, the medication suggestion and the risk early warning, and quantifying the curative effect level through a fuzzy comprehensive evaluation method. According to the method, individual differences are accurately captured, full-cycle dynamic evaluation is realized, and the curative effect evaluation accuracy and the clinical decision-making efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of internal medicine drug efficacy evaluation, and particularly relates to an internal medicine drug efficacy evaluation method based on electronic health record big data analysis. BACKGROUND

[0002] Under the background of population aging and increasing incidence of chronic diseases, the evaluation of drug treatment effect for internal diseases (such as hypertension, diabetes, coronary heart disease, etc.) is facing severe challenges. In clinical practice, drug efficacy is not only affected by basic characteristics such as patient age and gender, but also closely related to individual biological markers such as genetic polymorphism, epigenetic characteristics, and intestinal flora. Meanwhile, dynamic factors such as combined drug regimen, medication compliance, and changes in comorbidities can also significantly change the efficacy. The traditional efficacy evaluation mode based on small sample clinical trials is difficult to reflect the drug effect of complex patient population in the real world due to strict control of enrollment conditions, leading to problems such as poor efficacy, frequent adverse reactions, and even misjudgment of treatment regimen for some patients. In addition, the popularity of electronic health records (EHR) has accumulated a large amount of patient data, but existing technologies lack the ability to deeply integrate and dynamically analyze multi-source heterogeneous data (such as genetic data, time-series physiological indicators, and drug metabolism data), and cannot achieve individualized efficacy evaluation. Therefore, there is an urgent need for a big data analysis method based on electronic health records, which can accurately evaluate the efficacy of internal medicine drugs in the real world through multi-dimensional data fusion and intelligent algorithms, provide scientific basis for clinical drug decision-making, and ultimately improve treatment effect and reduce medical risks.

[0003] Traditional internal medicine drug efficacy evaluation mainly relies on two types of technical solutions: one is randomized controlled trial (RCT), which compares the efficacy difference between the test group and the control group by strictly screening homogeneous patients, and its advantages are rigorous test design and high data reliability, which is the "gold standard" for drug approval. However, its disadvantages are significant, such as excluding complex conditions such as comorbidities and special age groups from the enrolled patients, making it difficult to extrapolate the results to real clinical scenarios, and the test cycle is long and costly, which cannot timely reflect the long-term drug effect or rare adverse reactions. The second is single-center small sample retrospective analysis, which collects medical record data from a single medical institution and uses statistical methods (such as t-test and chi-square test) to evaluate efficacy. Its advantages are simple operation and relatively low cost. However, it has problems such as limited sample size, single data dimension (mostly relying on diagnosis and medication records, lacking molecular biology data), and strong regional limitations, making it difficult to capture the influence of individual differences on efficacy, and the universality of the evaluation results is poor. In addition, traditional solutions are mostly static evaluations, which cannot track dynamic data such as physiological indicator fluctuations and symptom changes after medication, making it difficult to achieve dynamic monitoring and regimen adjustment of efficacy.

[0004] With the rise of big data technology, the efficacy evaluation technology based on real world data (RWD) has gradually developed. The existing technologies are mostly focused on the structured analysis of electronic medical record (EMR) data, and the association between drugs and efficacy is mined through statistical models or simple machine learning algorithms (such as logistic regression, cluster analysis). For example, some technologies extract the data of patients' diagnoses, medications, laboratory tests, etc., to construct an efficacy prediction model to assist doctors in judging the drug applicability; another technology focuses on drug adverse reaction monitoring, and identifies the potential relationship between drugs and adverse events through association rule mining. These technologies have broken through the limitations of traditional RCT to some extent, and can reflect the drug effect in real clinical scenarios; but there are still obvious deficiencies: first, the data dimension is limited, and it mostly depends on conventional clinical data, without integrating molecular level data such as genetic polymorphism and epigenetic markers, making it difficult to explain the individual difference mechanism of efficacy; second, the analysis method is static, lacking dynamic modeling of time series data (such as continuous blood pressure and blood glucose monitoring values), and unable to capture the trend of efficacy over time; third, the model adaptability is poor, without considering dynamic factors such as patient medication compliance and changes in comorbidities, and the evaluation results are easily interfered by confounding variables; fourth, there is a lack of subgroup-specific analysis, and the differences in efficacy of different disease subgroups are not identified, making it difficult to achieve individualized evaluation.

[0005] Chinese patent application (application number 202410962036.1) discloses a preparation evaluation system based on electronic medical record data, which collects patient diseases, drug prescriptions, medication history and laboratory test results through an electronic medical record data acquisition module, calculates drug interaction probability, adverse event incidence, efficacy ratio and other indicators through a combination drug evaluation unit, a safety evaluation unit and an effectiveness evaluation unit respectively, and finally outputs the evaluation indicators through a comprehensive evaluation unit. Although this technology realizes preparation evaluation based on electronic medical record data, it has serious shortcomings: first, the data dimension is limited to conventional clinical data (such as prescriptions and test results), without including molecular biology data such as genetic polymorphism and drug metabolism enzyme activity, which cannot reveal the nature of individual differences in efficacy; second, the evaluation method is static statistical analysis (such as calculating incidence and ratio), lacking dynamic modeling of time series data, and unable to generate short-term, medium-term and long-term efficacy trend evaluation; third, it does not perform disease subgroup stratification analysis, and uses a unified standard to evaluate all patients, ignoring the differences in efficacy of different subgroups; fourth, the model has no self-adaptation ability, and the parameters are not iteratively optimized through reinforcement learning algorithm, making it difficult to adapt to changes brought by new data. SUMMARY

[0006] Based on the above technical problems, the present application discloses a medical drug efficacy evaluation method based on electronic health records big data analysis, specifically including:

[0007] S1, extracting the basic health data, diagnosis and treatment time series data and drug intervention data of the patient from the electronic health record, generating a dynamic feature set after spatiotemporal alignment of the data;

[0008] S2, based on the preset internal medicine disease typing standard, stratified clustering the patient data in the dynamic feature set to obtain several subgroups, for each subgroup, fusing the historical clinical research efficacy data and the real world data of the subgroup through transfer learning to construct the corresponding efficacy benchmark matrix of each subgroup;

[0009] S3, real-time collection of the electronic health record data of the patient after drug use, feature mapping of the collected data through a deep learning model to generate a real-time evaluation vector containing short-term physiological indicator responsiveness, medium-term symptom improvement trend and long-term prognosis risk;

[0010] S4, dynamically matching the real-time evaluation vector with the corresponding subgroup efficacy benchmark matrix to calculate the matching deviation value, which is corrected by introducing a patient individual dynamic weight coefficient;

[0011] S5, using a reinforcement learning algorithm, taking the deviation correction value after dynamic matching as input and taking the multidimensional evaluation index of drug efficacy as output to construct an adaptive efficacy evaluation model, which is iteratively optimized by continuously receiving new electronic health record data;

[0012] S6, based on the output result of the efficacy evaluation model, generating an individualized evaluation report containing drug efficacy grade, optimal drug adjustment suggestion and potential risk warning, the efficacy grade in the report being converted into a quantifiable score value by a fuzzy comprehensive evaluation method.

[0013] Preferably, the basic health data in S1 includes genetic polymorphism data, epigenetic markers and intestinal flora spectrum; the diagnosis and treatment time series data includes continuously recorded physiological indicator fluctuation values and symptom semantic vectors according to time stamps; the drug intervention data includes drug metabolism enzyme activity data and blood drug concentration monitoring values.

[0014] Preferably, in S1, the data is spatiotemporally aligned to generate a dynamic feature set, specifically: taking the first drug use time of the patient as the reference time stamp t0, mapping the genetic polymorphism data G, epigenetic markers E and intestinal flora spectrum M in the basic health data into static feature vectors S0 = [G, E, M] at the reference time point; for the physiological indicator fluctuation values P(t i ) and symptom semantic vectors Q(t i ) recorded in the diagnosis and treatment time series data according to time stamps t i , through a time interpolation function This is converted to a normalized time coordinate relative to a baseline time, where Δt is a preset time interval; for drug metabolism enzyme activity data C(t) in drug intervention data... j ) and blood drug concentration monitoring value D(t) j ), through the spatial alignment function S(t) j )=ω1×C(t j )+ω2×D(t j Feature fusion is performed, where ω1 and ω2 are the corresponding weight coefficients; the dynamic feature integration formula is used. Generate a dynamic feature set F(t) covering the baseline time and subsequent monitoring time, where α, β, and γ are the corresponding fusion coefficients, n is the number of diagnosis and treatment time series data, and m is the number of drug intervention data.

[0015] Preferably, in step S2, for each subgroup, the efficacy data from historical clinical studies are fused with the real-world data of that subgroup through transfer learning to construct the efficacy benchmark matrix corresponding to each subgroup. Specifically, this involves extracting the efficacy data H = {h1, h2, ..., h...} corresponding to the disease subgroup in historical clinical studies. p} and subgroup real-world data R={r1,r2,…,r q}, feature mapping function through transfer learning Feature alignment is performed on historical data and real-world data, and the matrix element values ​​corresponding to different drugs within the fused subgroup are calculated using the following formula: λ is the fusion weight coefficient, M ij Using the element in the i-th row and j-th column of the efficacy benchmark matrix as an example, after traversing all drugs and indicators, an efficacy benchmark matrix is ​​constructed that includes the theoretical onset threshold and adverse reaction warning threshold for different drugs within subgroups. n r k represents the number of drugs within the subgroup. r This refers to the number of indicators.

[0016] Preferably, in step S3, the electronic health record data of the patient after medication is collected in real time, and a real-time evaluation vector is generated by feature mapping of the collected data through a deep learning model. Specifically, the collected time series data after medication X = {x1, x2, ..., x...} is used to generate a real-time evaluation vector. t The input consists of a deep learning model with convolutional layers and a long short-term memory network. The convolutional layers are filtered by filter K. s Extracting local variation features of short-term physiological indicators F s The formula is: T s For a short time window, the short-term physiological indicator response R is obtained after processing with an activation function. s During the forward propagation of the Long Short-Term Memory (LSTM) network, the intermediate time window T is calculated through a gating mechanism.m feature accumulation within h t-1 The hidden state from the previous moment is used to obtain the mid-term symptom improvement trend R after time-series smoothing. m The output layer of a deep learning model uses a fully connected layer to process a long-term time window T. l Weighted aggregation of features within w t Using time decay weights, and combining them with the prognostic risk prediction function, the long-term prognostic risk R is generated. l By concatenating vectors V = [R s ,R m ,R l Generate a real-time evaluation vector V.

[0017] Preferably, in step S4, the matching deviation value is calculated by dynamically matching the real-time assessment vector with the corresponding subgroup efficacy benchmark matrix. Specifically, this involves determining the relationship between each element in the real-time assessment vector V and the efficacy benchmark matrix. The corresponding drug row vector M i =[m i1 ,m i2 ,…,m ik The dimensional correspondence is determined by calculating the vector space matching degree using cosine similarity. Weights are assigned based on the clinical importance of each dimension of the indicators. j Through the deviation calculation formula The initial matching deviation value B is obtained. i , where v j -m ij To evaluate the absolute deviation of the vector from the benchmark matrix in the j-th dimension in real time, (1-S i The spatial matching degree correction factor is obtained by traversing all drug row vectors to obtain the matching deviation value set B = {B1, B2, ..., B} for each drug. n}

[0018] Preferably, the matching deviation value in S4 is corrected by introducing a patient-specific dynamic weighting coefficient. Specifically, the patient-specific dynamic weighting coefficient is generated based on real-time patient medication adherence data and comorbidity change rate data. This is achieved by acquiring real-time patient medication adherence data A and comorbidity change rate data C, normalizing them to obtain A′ and C′, and generating the patient-specific dynamic weighting coefficient W using the weighting coefficient fusion formula W=θ·A′+(1-θ)·(1-C′), where θ is the weight allocation factor. The deviation correction formula B′ is then used. i =B i ·(1-W) for matching deviation value B i The correction is performed to obtain the corrected deviation value B′. i .

[0019] Preferably, in step S5, constructing an adaptive efficacy evaluation model specifically involves: setting the dynamically matched set of deviation correction values ​​B′={B′1,B′2,…,B′}. n As an environmental state input, the multidimensional evaluation index of drug efficacy Y={y1,y2,…,y} is used. k As output, a reward function is defined. Evaluate the output effect, where α i As the indicator weight, θ represents the expected efficacy index, β is the error penalty coefficient, and the model parameters θ are updated using the policy gradient algorithm. * The formula is: Where γ s As a discount factor, The strategy function continuously receives new electronic health record data for iterative optimization, enabling the efficacy evaluation model to stably output multi-dimensional evaluation indicators consistent with the actual efficacy under the input of deviation correction values, thus constructing an adaptive efficacy evaluation model.

[0020] Preferably, in step S5, generating an individualized assessment report based on the output of the efficacy assessment model specifically involves: extracting the multidimensional assessment indicators of drug efficacy output by the efficacy assessment model, comparing them with preset efficacy level classification standards, and then using a formula... The drug efficacy grade L is determined, where δ(p,l) is the characteristic function of index p conforming to grade l. Combining patient genetic polymorphism data, drug-metabolizing enzyme activity data, and short-term response and mid-term improvement trends in real-time assessment vectors, the medication adjustment plan with the highest matching degree is selected. This is based on the adverse reaction warning threshold T and the long-term prognostic risk value R. l Using the risk probability formula P = λ p ·I(R l >T)+(1-λ p )·exp(-γ p ·(TR l Calculate the probability P of the potential risk occurring, where I is the indicator function and λ is the value of the indicator function. p γ p To adjust parameters, an individualized assessment report is generated that includes drug efficacy level, optimal medication adjustment recommendations, and potential risk warnings.

[0021] Preferably, in step S6, the efficacy level reported is transformed from multidimensional indicators into quantifiable scores using a fuzzy comprehensive evaluation method. Specifically, this involves determining the set of multidimensional indicators for drug efficacy evaluation and the efficacy level evaluation set, and then using a membership function μ. ij =exp(-k μ ·|x i -c j |) Construct a fuzzy relation matrix R = [μij ] m×n wherein x i is the actual value of the ith index, c j is the central value of the jth grade, k μ is the membership adjustment coefficient, the weight vector is determined according to the clinical importance of each index, the comprehensive evaluation vector S = {s1, s2, …, s n is obtained by taking the grade score corresponding to the maximum membership degree or by weighted calculation to obtain a quantifiable efficacy grade score value, wherein t j is the quantifiable score value corresponding to the efficacy grade evaluation set.

[0022] Compared with the prior art, the technical scheme of the application has the following technical effects:

[0023] The application integrates basic health data such as gene polymorphism, epigenetic markers and intestinal flora spectrum, combines intervention data such as drug metabolism enzyme activity and blood drug concentration, generates a dynamic feature set through space-time alignment, accurately captures individual biological differences of patients, divides disease subgroups through hierarchical clustering, and uses transfer learning to construct a subgroup-specific efficacy benchmark matrix, thereby solving the problem of individual difference neglect caused by “one-size-fits-all” in traditional evaluation, making the efficacy evaluation more in line with the actual situation of patients, and providing accurate basis for individualized medication.

[0024] The application collects and maps data after medication in real time through a deep learning model, generates an evaluation vector containing short-term physiological index response, medium-term symptom improvement trend and long-term prognosis risk, and realizes whole-cycle efficacy tracking from the moment to the long term through dynamic matching and deviation correction, thereby solving the problem that static evaluation cannot capture changes in efficacy over time, and enabling timely discovery of efficacy fluctuations to provide dynamic data support for adjusting medication regimens.

[0025] The application uses a reinforcement learning algorithm to construct a model with deviation correction values as input and multi-dimensional efficacy indicators as output, iteratively optimizes parameters by continuously receiving new data, solves the problem of poor adaptability of traditional models and difficulty in coping with dynamic changes in clinical data, enables the evaluation model to evolve with clinical practice, and improves the adaptability and evaluation accuracy of the model to complex clinical scenarios.

[0026] The application converts multi-dimensional indicators into quantifiable efficacy grades through fuzzy comprehensive evaluation, generates a report combining optimal medication adjustment suggestions and risk warnings, solves the problem of fragmented evaluation results and high clinical application threshold, provides clear efficacy judgment, medication guidance and risk prompts for doctors, helps to quickly develop accurate treatment strategies, and improves clinical decision-making efficiency.

[0027] The above description is only a summary of the technical solutions of the present application. In order to make the technical means of the present application more clearly understood, and to enable the above and other purposes, characteristics and advantages of the present application to be more apparent and easy to understand, the following will be described in detail with reference to the preferred embodiments of the present application and in conjunction with the accompanying drawings.

[0028] The above and other purposes, advantages and characteristics of the present application will be more apparent to those skilled in the art from the following detailed description of specific embodiments of the present application in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without any creative effort. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual proportions.

[0030] According to the description of the drawings in the document and the corresponding technical content, the titles of the drawings are as follows:

[0031] Figure 1 Flow chart of internal medicine drug efficacy evaluation method based on electronic health record big data analysis;

[0032] Figure 2 Deep neural network model diagram based on electronic health record big data analysis;

[0033] Figure 3 Dynamic feature set change diagram of statin therapy for patients with coronary heart disease;

[0034] Figure 4 Statin efficacy benchmark matrix diagram for subgroup 1 (high baseline LDL-C + combined diabetes);

[0035] Figure 5 Three-dimensional scatter plot for evaluating statin efficacy for target patients;

[0036] Figure 6 Reinforcement learning model parameter iterative optimization curve diagram;

[0037] Figure 7 Dynamic feature set change diagram of safety for rheumatoid arthritis patients using immunosuppressants;

[0038] Figure 8 Immunosuppressant safety benchmark matrix diagram for subgroup A (liver function borderline abnormal + HLA-B*5801 positive);

[0039] Figure 9 Three-dimensional scatter plot of immunosuppressant safety assessment vector for Subgroup A patients;

[0040] Figure 10 Matched results plot for drug safety for Subgroup A (borderline abnormal liver function + HLA-B*5801 positive) patients. DETAILED DESCRIPTION

[0041] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the accompanying drawings for the embodiments of the present application to make a clear and complete description of the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. In the following description, specific details such as specific configurations and components are provided only to help a comprehensive understanding of the embodiments of the present application. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. In addition, in order to be clear and concise, the description of known functions and structures is omitted in the embodiments.

[0042] It should be understood that the "one embodiment" or "the embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "one embodiment" or "the embodiment" appearing throughout the specification does not necessarily mean the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.

[0043] In addition, reference numerals and / or letters can be repeated in different examples in the present application. Such repetition is for the purpose of simplification and clarity, and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0044] The term "and / or" herein is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, B exists alone, and A and B exist simultaneously. The term "and" herein is a description of another association relationship of the associated objects, which means that there can be two relationships, for example, A and B can mean that A exists alone and A and B exist simultaneously. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it.

[0045] The term "at least one" herein is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, at least one of A and B can mean that A exists alone, A and B exist simultaneously, and B exists alone.

[0046] It is also need to point out that, in this article, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the term "includes", "contains" or any other variant thereof is intended to cover non-exclusive inclusion.

[0047] Embodiment 1

[0048] This embodiment mainly describes a medical drug efficacy evaluation method based on electronic health records big data analysis, such as Figure 1 as shown, including:

[0049] S1, extracting the basic health data, diagnosis and treatment time series data and drug intervention data of the patient from the electronic health records, and generating a dynamic feature set after time and space alignment of the data;

[0050] S2, based on the preset internal medicine disease typing standard, the patient data in the dynamic feature set is stratified clustered to obtain several subgroups, and for each subgroup, the efficacy data of the historical clinical research is fused with the real world data of the subgroup through transfer learning, and the efficacy benchmark matrix corresponding to each subgroup is constructed;

[0051] S3, real-time collection of electronic health record data after the patient takes medicine, feature mapping of the collected data through a deep learning model, and generation of a real-time evaluation vector containing short-term physiological indicator responsiveness, medium-term symptom improvement trend and long-term prognosis risk;

[0052] S4, dynamically matching the real-time evaluation vector with the corresponding subgroup efficacy benchmark matrix, calculating the matching deviation value, and the matching deviation value is corrected by introducing a patient individual dynamic weight coefficient;

[0053] S5, using a reinforcement learning algorithm, taking the deviation correction value after dynamic matching as input, and taking the multi-dimensional evaluation index of drug efficacy as output, constructing an adaptive efficacy evaluation model, and iteratively optimizing the parameters by continuously receiving new electronic health record data;

[0054] S6, based on the output result of the efficacy evaluation model, generating an individualized evaluation report containing drug efficacy level, optimal drug adjustment suggestion and potential risk warning, and the efficacy level in the report is converted into a quantifiable score value by a fuzzy comprehensive evaluation method.

[0055] Further, the basic health data in S1 includes genetic polymorphism data, epigenetic markers and intestinal flora spectrum; the diagnosis and treatment time series data includes the physiological indicator fluctuation value and symptom semantic vector recorded continuously according to the time stamp; the drug intervention data includes drug metabolism enzyme activity data and blood drug concentration monitoring value.

[0056] Furthermore, in S1, the data is spatiotemporally aligned to generate a dynamic feature set. Specifically, using the patient's first medication time as the baseline timestamp t0, the gene polymorphism data G, epigenetic marker E, and gut microbiota profile M in the basic health data are mapped to static feature vectors S0 = [G, E, M] at the baseline time point, respectively; for the diagnosis and treatment time series data, according to the timestamp t... i Recorded physiological index fluctuation value P(t) i ) and symptom semantic vector Q(t) i ), through time interpolation function This is converted to a normalized time coordinate relative to a baseline time, where Δt is a preset time interval; for drug metabolism enzyme activity data C(t) in drug intervention data... j ) and blood drug concentration monitoring value D(t) j ), through the spatial alignment function S(t) j )=ω1×C(t j )+ω2×D(t j Feature fusion is performed, where ω1 and ω2 are the corresponding weight coefficients; the dynamic feature integration formula is used. Generate a dynamic feature set F(t) covering the baseline time and subsequent monitoring time, where α, β, and γ are the corresponding fusion coefficients, n is the number of diagnosis and treatment time series data, and m is the number of drug intervention data.

[0057] Furthermore, in S2, for each subgroup, transfer learning is used to fuse the efficacy data from historical clinical studies with the real-world data of that subgroup to construct the efficacy benchmark matrix corresponding to each subgroup. Specifically, the efficacy data H = {h1, h2, ..., h...} corresponding to the disease subgroup in historical clinical studies is extracted. p} and subgroup real-world data R={r1,r2,…,r q}, feature mapping function through transfer learning Feature alignment is performed on historical data and real-world data, and the matrix element values ​​corresponding to different drugs within the fused subgroup are calculated using the following formula: λ is the fusion weight coefficient, M ij Using the element in the i-th row and j-th column of the efficacy benchmark matrix as an example, after traversing all drugs and indicators, an efficacy benchmark matrix is ​​constructed that includes the theoretical onset threshold and adverse reaction warning threshold for different drugs within subgroups. n r k represents the number of drugs within the subgroup. r This refers to the number of indicators.

[0058] Furthermore, such as Figure 2As shown, the electronic health record data of the patient after medication in S3 is collected in real time, and the collected data is mapped to features by a deep learning model to generate a real-time evaluation vector, specifically: the collected time series data X = {x1, x2, …, xT} after medication is input into a deep learning model composed of a convolutional layer and a long short-term memory network, the convolutional layer filters the input data through a filter K t} to extract short-term physiological indicators, and the long short-term memory network is used to extract the feature accumulation amount in the medium-term time window T s} through a gating mechanism in the forward propagation process. s The formula is: T s The short-term time window is processed by an activation function to obtain the short-term physiological indicator response R s ; in the forward propagation process of the long short-term memory network, the feature accumulation amount in the medium-term time window T m is calculated through a gating mechanism. h t-1 The previous hidden state is processed by a time smoothing process to obtain the medium-term symptom improvement trend R m ; the output layer of the deep learning model weights and aggregates the features in the long-term time window T l through a fully connected layer. w t The time decay weight is combined with the prognosis risk prediction function to generate the long-term prognosis risk R l , and the real-time evaluation vector V is generated by vector splicing V = [R s , R m , R l ].

[0059] Further, in S4, the real-time evaluation vector is dynamically matched with the corresponding subgroup efficacy benchmark matrix to calculate the matching deviation value, specifically: the dimension correspondence between each element in the real-time evaluation vector V and the corresponding drug row vector M i = [m i1 , m i2 , …, m ik ] in the efficacy benchmark matrix is determined, and the vector space matching degree is calculated by cosine similarity. Based on the clinical importance of each dimension indicator, a weight w j is assigned, and the preliminary matching deviation value B i is obtained by the deviation calculation formula , where v j -m ij is the absolute deviation of the real-time evaluation vector and the benchmark matrix in the jth dimension, (1-S i ) is the space matching degree correction factor, and after traversing all drug row vectors, the matching deviation value set B = {B1, B2, …, B n} for each drug is obtained.

[0060] Furthermore, in S4, the matching deviation value is corrected by introducing a patient-specific dynamic weighting coefficient. Specifically, the individual dynamic weighting coefficient is generated based on real-time patient medication adherence data and comorbidity change rate data. This is achieved by acquiring real-time patient medication adherence data A and comorbidity change rate data C, normalizing them to obtain A′ and C′, and generating the individual dynamic weighting coefficient W using the weighting coefficient fusion formula W=θ·A′+(1-θ)·(1-C′), where θ is the weight allocation factor. The deviation correction formula B′ is then used. i =B i ·(1-W) for matching deviation value B i The correction is performed to obtain the corrected deviation value B′. i .

[0061] Furthermore, in S5, an adaptive efficacy assessment model is constructed, specifically by: dynamically matching the set of deviation correction values ​​B′={B′1,B′2,…,B′ n As an environmental state input, the multidimensional evaluation index of drug efficacy Y={y1,y2,…,y} is used. k As output, a reward function is defined. Evaluate the output effect, where α i As the indicator weight, θ represents the expected efficacy index, β is the error penalty coefficient, and the model parameters θ are updated using the policy gradient algorithm. * The formula is: Where γ s As a discount factor, The strategy function continuously receives new electronic health record data for iterative optimization, enabling the efficacy evaluation model to stably output multi-dimensional evaluation indicators consistent with the actual efficacy under the input of deviation correction values, thus constructing an adaptive efficacy evaluation model.

[0062] Furthermore, in S5, a personalized assessment report is generated based on the output of the efficacy assessment model. Specifically, this involves extracting the multidimensional assessment indicators of drug efficacy output by the efficacy assessment model, comparing them with the preset efficacy level classification standards, and then using a formula... The drug efficacy grade L is determined, where δ(p,l) is the characteristic function of index p conforming to grade l. Combining patient genetic polymorphism data, drug-metabolizing enzyme activity data, and short-term response and mid-term improvement trends in real-time assessment vectors, the medication adjustment plan with the highest matching degree is selected. This is based on the adverse reaction warning threshold T and the long-term prognostic risk value R. l Using the risk probability formula P = λ p ·I(R l >T)+(1-λ p )·exp(-γ p ·(TRl Calculate the probability P of the potential risk occurring, where I is the indicator function and λ is the value of the indicator function. p γ p To adjust parameters, an individualized assessment report is generated that includes drug efficacy level, optimal medication adjustment recommendations, and potential risk warnings.

[0063] Furthermore, the efficacy levels in the S6 report are transformed from multidimensional indicators into quantifiable scores using a fuzzy comprehensive evaluation method. Specifically, this involves determining the set of multidimensional indicators for drug efficacy assessment and the efficacy level evaluation set, and then using the membership function μ... ij =exp(-k μ ·|x i -c j |) Construct a fuzzy relation matrix R = [μ ij ] m×n , where x i c is the actual value of the i-th indicator. j Let k be the center value of the j-th level. μ The membership adjustment coefficient is used to determine the weight vector based on the clinical importance of each indicator. A comprehensive evaluation vector S = {s1, s2, ..., s...} is obtained through fuzzy comprehensive evaluation. n}, take the grade score corresponding to the highest membership degree or calculate it through weighted average. A quantifiable efficacy rating score was obtained, where t j This represents the quantitative score corresponding to the efficacy rating set.

[0064] This detailed description of implementation integrates multi-dimensional health data to generate a dynamic feature set, combines hierarchical clustering and transfer learning to construct a subgroup-specific benchmark matrix, utilizes deep learning to generate real-time assessment vectors, and optimizes them through a reinforcement learning model to achieve accurate matching and bias correction, generating a quantitative, individualized assessment report. This accurately captures individual differences, achieves dynamic assessment throughout the entire lifecycle, improves the accuracy of efficacy assessment and the efficiency of clinical decision-making, and effectively solves the problems of traditional assessments ignoring individual differences, being static, and having poor model adaptability, providing scientific support for the efficacy and safety assessment of internal medicine drugs.

[0065] Based on Example 1, this example describes in detail the assessment of accurate prediction of drug efficacy based on big data analysis of electronic health records, taking the accurate prediction assessment of statin efficacy in patients with coronary heart disease as an example, specifically:

[0066] Five hundred coronary artery disease patients who took statins from the cardiology department of a tertiary hospital between January and December 2023 were selected. All patients had complete electronic health records. The data collection scope is as follows:

[0067] Baseline health data: Extract patient genetic polymorphism data (e.g., SLC01B1 genetic polymorphism, related to statin metabolism), epigenetic markers (e.g., vascular endothelial cell miRNA-126 expression levels), intestinal flora spectrum (e.g., relative abundance of Akkermansia genus);

[0068] Diagnosis and treatment time series data: Physiological indicator fluctuation values (e.g., low-density lipoprotein cholesterol (LDL-C), total cholesterol (TC), triglycerides (TG)) recorded continuously by timestamp (8:00, 16:00 daily), and patient self-reported "chest tightness" and "chest pain" symptom semantic vectors (symptom descriptions converted to 0-10 severity scores through natural language processing);

[0069] Drug intervention data: Statin metabolizing enzyme (e.g., CYP3A4) activity detection values (unit: nmol / min / mg protein), blood drug concentration monitoring values (e.g., atorvastatin blood trough concentration, unit: ng / mL).

[0070] Taking the time when the patient first takes statins as the baseline timestamp t0 (e.g., March 15, 2023, 8:00), the above data is spatio-temporally aligned: the baseline health data is mapped to the static feature vector at t0 (e.g., SLC01B1*1b genotype is recorded as 1, *5 genotype is recorded as 0; Akkermansia relative abundance ≥5% is recorded as 1, <5% is recorded as 0); LDL-C values and symptom scores in the diagnosis and treatment time series data are converted to normalized time coordinates relative to t0 through time interpolation (e.g., t0+7 days is recorded as 1, t0+14 days is recorded as 2); drug metabolizing enzyme activity and blood drug concentration are fused into drug intervention feature values with weights (ω1=0.6, ω2=0.4) to generate a dynamic feature set covering t0 to t0+90 days, as shown in the following table: Figure 3 The blue curve in the figure shows the trend of LDL-C gradually decreasing from 4.8 mmol / L to 1.9 mmol / L, directly reflecting the lipid-lowering effect of the drug; the red curve in the figure shows the symptom score decreasing from 8 to 1, reflecting the continuous improvement of the patient's discomfort symptoms; the green curve in the figure shows the change of the drug intervention feature value first increasing and then stabilizing, reflecting the action process of the drug in the body.

[0071] Based on the LDL-C baseline value in the dynamic feature set (≥4.14 mmol / L for high baseline, <4.14 mmol / L for low baseline), whether or not combined with diabetes (yes / no), the 500 patients are divided into 4 subgroups using hierarchical clustering algorithm:

[0072] Subgroup 1: High baseline LDL-C + combined with diabetes (120 cases);

[0073] Subgroup 2: high baseline LDL-C + without diabetes (130 cases);

[0074] Subgroup 3: low baseline LDL-C + with diabetes (110 cases);

[0075] Subgroup 4: low baseline LDL-C + without diabetes (140 cases).

[0076] For subgroup 1 (high baseline LDL-C + with diabetes), two types of data are fused by transfer learning:

[0077] Historical clinical research data (such as statin efficacy data of the same type of patients in the ODYSSEY OUTCOMES trial, including the proportion of LDL-C reduction ≥ 50%, the incidence of major adverse cardiovascular events (MACE));

[0078] Real-world data of 120 patients in this subgroup (such as actual LDL-C reduction after medication, MACE occurrence during 3-month follow-up).

[0079] The efficacy benchmark matrix constructed after fusion contains the indicators of four drugs: atorvastatin (20 mg, 40 mg), rosuvastatin (10 mg, 20 mg): theoretical onset threshold (LDL-C reduction ≥ 40%), efficacy maintenance threshold (continuous 4-week LDL-C < 1.8 mmol / L), long-term prognosis threshold (MACE incidence < 5% within 1 year), as shown in the table below. Figure 4 The matrix row represents four statins (atorvastatin 20 mg, atorvastatin 40 mg, rosuvastatin 10 mg, rosuvastatin 20 mg), and the column represents three key evaluation indicators (theoretical onset threshold: LDL-C reduction ≥ 40%; efficacy maintenance threshold: continuous 4-week LDL-C < 1.8 mmol / L; long-term prognosis threshold: MACE incidence < 5% within 1 year). The cell color gradually changes from red (< 30%) to green (≥ 70%), reflecting the compliance probability of each drug on the corresponding indicators. Among them, rosuvastatin 20 mg shows strong lipid-lowering effect on the theoretical onset threshold (78%) and efficacy maintenance threshold (72%) in this subgroup of patients, showing deep green. Atorvastatin 40 mg has a compliance probability of 65% (light green) in the long-term prognosis threshold, which is slightly better than rosuvastatin 10 mg (58%). Atorvastatin 20 mg is red to yellow (30%-50%) in all three indicators, indicating that its overall efficacy on the current subgroup of patients is weak, allowing clinicians to quickly identify the efficacy differences of different statins in this subgroup.

[0080] Real-time data collection was performed on one 65-year-old male patient in subgroup 1 (SLCO1B1*1b genotype, Akkermansia relative abundance 6.2%, taking rosuvastatin 20 mg) after medication, and electronic health record data at t0+7 days, t0+14 days, t0+30 days, and t0+60 days were obtained:

[0081] t0+7 days: LDL-C decreased from 4.8 mmol / L to 3.2 mmol / L (decrease of 33.3%), and symptom score decreased from 8 to 5;

[0082] t0+14 days: LDL-C decreased to 2.6 mmol / L (decrease of 45.8%), and symptom score decreased to 3;

[0083] t0+30 days: LDL-C decreased to 2.1 mmol / L (decrease of 56.2%), symptom score remained at 3, and ultrasound showed that coronary plaque volume decreased by 5%;

[0084] t0+60 days: LDL-C stabilized at 2.0 mmol / L, symptom score was 2, and plaque volume decreased by 8%.

[0085] The above data was input into a deep learning model composed of a convolutional layer and a long short-term memory network (LSTM): the convolutional layer extracted the LDL-C decrease feature from t0+7 days to t0+14 days, obtaining the short-term physiological indicator responsiveness (0.82, with a full score of 1.0); the LSTM calculated the symptom improvement and plaque change trend from t0+14 days to t0+30 days through a gating mechanism, obtaining the medium-term symptom improvement trend (0.76); the fully connected layer weighted and aggregated the data from t0+30 days to t0+60 days (time decay weight w t increasing over time), combined with a coronary heart disease risk prediction model to obtain the long-term prognosis risk (0.12, i.e., the probability of MACE occurrence within 1 year was 12%). The generated real-time evaluation vector was [0.82, 0.76, 0.12], as Figure 5 shown, which intuitively presented the efficacy evaluation results of a certain coronary heart disease patient taking rosuvastatin 20 mg, with the X-axis representing the short-term physiological indicator responsiveness (value range 0-1, corresponding to the LDL-C decrease feature from t0+7 to t0+14 days), which was 0.82 for this patient; the Y-axis represented the medium-term symptom improvement trend (value range 0-1, corresponding to the symptom score and plaque change from t0+14 to t0+30 days), which was 0.76 for this patient; and the Z-axis represented the long-term prognosis risk (value range 0-1, corresponding to the prediction of MACE occurrence probability from t0+30 to t0+60 days), which was 0.12 for this patient; Figure 5The end point of the vector [0.82, 0.76, 0.12] is marked with a black dot, and the corresponding values of each coordinate axis are connected by a dashed line to clearly show that it is located in the area of high short-term responsiveness (right side of the X-axis), medium-term improvement trend (right side of the middle of the Y-axis), and low long-term risk (bottom of the Z-axis). The coordinate system background is divided into areas by light grid lines, and each axis is labeled with the index name and value range. The overall reflects the patient's efficacy performance in different time dimensions after taking the drug.

[0086] The real-time evaluation vector of the current patient is dynamically matched with the efficacy benchmark matrix of subgroup 1: Figure 4 The cosine similarity of short-term responsiveness (0.82) and theoretical onset threshold of rosuvastatin 20 mg (0.80) is 0.97, the similarity of medium-term trend (0.76) and efficacy maintenance threshold (0.70) is 0.92, and the similarity of long-term risk (0.12) and prognosis threshold (0.10) is 0.89, resulting in a preliminary matching deviation value (0.05). Introduce individual dynamic weight coefficient (W = 0.12) to correct the deviation: the patient's medication adherence data (A = 0.95, A' = 0.90 after normalization), the rate of change of comorbidities (fasting blood glucose fluctuation from ±2.0 mmol / L to ±1.2 mmol / L, C' = 0.30 after normalization), through the formula W = 0.7 × 0.90 + 0.3 × (1-0.30) = 0.78, the corrected deviation value is 0.05 × (1-0.78) = 0.011.

[0087] The corrected deviation value is input into the reinforcement learning model, with "LDL-C target rate" "plaque reversal rate" "MACE incidence" as output indicators, and the reward function is defined (reward value +0.8 when deviation value <0.02). The model iteratively optimizes parameters through the policy gradient algorithm, and after receiving the new data of the patient at t0+90 days (LDL-C = 1.9 mmol / L, symptom score 1), the parameter update adjusts the long-term prognosis risk prediction value from 0.12 to 0.09, as shown in Figure 6 The prediction error of the reinforcement learning model gradually decreases from 0.06 to 0.019 through 50 iterations, and further decreases after the introduction of new data at t0+90 days (40th iteration), as shown by the key node error values marked by red markers and the new data introduction points marked by green dashed lines. The optimization process achieved by the model through continuous learning is reflected, and the error is finally stabilized within 0.02, below the threshold line of 0.03, verifying the adaptive ability of the model.

[0088] Based on the multi-dimensional indicators of model output (LDL-C target rate 92%, plaque regression rate 8%, MACE incidence 9%), the efficacy grade was converted by fuzzy comprehensive evaluation method: set "excellent (80-100 points), good (60-79 points), medium (40-59 points), poor (<40 points)" four level standards, calculate the membership degree (LDL-C target rate membership "excellent" degree 0.9, plaque regression rate membership "good" degree 0.7, MACE incidence membership "excellent" degree 0.8), after weighting, the comprehensive score is 86 points, corresponding to "excellent" level.

[0089] Therefore, the report recommends maintaining the rosuvastatin 20mg dose, monitoring LDL-C and liver enzymes every 3 months, and suggests considering reducing to 10mg if LDL-C <1.4mmol / L, with the potential risk of myalgia (incidence 3.2%).

[0090] This embodiment describes in detail the precise prediction and evaluation of the efficacy of statins in patients with coronary heart disease, the generation of dynamic feature sets by integrating genetic polymorphisms, intestinal flora spectrum, and other basic health data with dynamic diagnosis and treatment, and drug intervention data, the construction of subgroup-specific efficacy benchmark matrix by hierarchical clustering and transfer learning, which precisely captures the efficacy differences of different subgroups of patients to statins, and solves the problem of ignoring individual biological characteristics and subgroup specificity in traditional evaluation; using deep learning models to generate real-time evaluation vectors containing short-term physiological indicator response, medium-term symptom improvement trend, and long-term prognosis risk, after dynamic matching and reinforcement learning model optimization, the accuracy of efficacy prediction is significantly improved, with the prediction error of LDL-C target rate, plaque regression rate, etc. reduced to less than 0.02, and the "excellent" level of efficacy evaluation and individualized medication recommendations generated by fuzzy comprehensive evaluation method provide clear and operable decision-making basis for clinicians, effectively supporting the precise adjustment of statin treatment plans for patients with coronary heart disease, and the long-term prognosis risk prediction (such as MACE incidence) helps to intervene potential risks in advance, improving treatment safety and effectiveness.

[0091] Based on Example 1, this embodiment describes the evaluation of drug safety monitoring based on electronic health record big data analysis, taking the safety monitoring evaluation of immunosuppressive agents for rheumatoid arthritis patients as an example, specifically:

[0092] 400 rheumatoid arthritis patients using immunosuppressive agents in a rheumatology department from January 2023 to December 2023 were selected, and electronic health record data collection focused on safety-related dimensions:

[0093] Baseline health data: extract patient genetic polymorphism data (e.g., HLA-B*5801 gene associated with azathioprine allergy; TPMT gene associated with azathioprine myelosuppression risk), epigenetic markers (e.g., TNF-α promoter methylation level in peripheral blood mononuclear cells), baseline liver and kidney function values (e.g., initial ALT value 25 U / L, initial creatinine value 70 μmol / L);

[0094] Diagnosis and treatment timing data: safety indicators recorded by timestamp (every Monday, Thursday) (e.g., white blood cell count (WBC), platelet count (PLT), ALT, AST), and patient self-reported adverse reaction symptom semantic vectors (0-10 score) such as "rash" and "oral ulcer";

[0095] Drug intervention data: immune suppressant enzyme (e.g., TPMT) activity detection value (unit: nmol / h / mL), blood drug concentration monitoring value (e.g., methotrexate blood drug trough concentration, unit: μmol / L).

[0096] Take the time when the patient first uses the immune suppressant as the baseline timestamp t0 (e.g., May 20, 2023), and align the data in space and time: map the HLA-B*5801 genotype (positive as 1, negative as 0), TPMT genotype (wild type as 1, mutant type as 0) and other basic data to the static feature vector at time t0; convert the WBC, ALT and other indicators every week into normalized time coordinates (t0+7 days as 1, t0+14 days as 2, and so on) through time interpolation; fuse the metabolic enzyme activity and blood drug concentration into a drug intervention safety feature value with weights (ω1=0.5, ω2=0.5) to generate a dynamic feature set as shown in Figure 7 Figure 7 The blue curve in the middle represents the fluctuation of white blood cell count (WBC), with an initial value of 4.5×10 9 / L, which shows a slow downward trend, and at t0+120 days, it drops to 3.2×10 9 / L, which is always maintained at the edge of the normal range (3.5-9.5×10 9 / L); the orange curve reflects the change of alanine aminotransferase (ALT), which gradually rises from the baseline of 32 U / L to 58 U / L, showing a continuous rising trend of liver function indicators; the purple curve is the drug intervention safety feature value (fusion of metabolic enzyme activity and blood drug concentration), which shows a rising and then stable trend, with the peak value appearing at t0+42 days, highlighting the development trajectory of the potential liver damage risk of ALT rising.

[0097] ​Based on baseline liver and kidney function (ALT ≥ 30 U / L is considered borderline abnormal for liver function, < 30 U / L is considered normal; creatinine ≥ 80 μmol / L is considered borderline abnormal for kidney function, < 80 μmol / L is considered normal) and HLA-B*5801 genotype (positive / negative) in the dynamic feature set, stratified clustering was used to divide 400 patients into 4 safety risk subgroups:

[0098] Subgroup A: Borderline abnormal liver function + HLA-B*5801 positive (90 cases);

[0099] Subgroup B: Borderline abnormal liver function + HLA-B*5801 negative (110 cases);

[0100] Subgroup C: Borderline renal function abnormalities + HLA-B*5801 positivity (80 cases);

[0101] Subgroup D: borderline renal function abnormalities + HLA-B*5801 negative (120 cases).

[0102] For high-risk subgroup A (borderline abnormal liver function + HLA-B*5801 positive), a safety benchmark matrix was constructed by integrating historical clinical research data (such as the incidence of rash and myelosuppression of azathioprine in similar patients) with real-world data from 90 patients in this subgroup (the degree of WBC decrease and ALT increase after medication) through transfer learning. The matrix includes safety thresholds for four drugs: methotrexate (10 mg / week, 15 mg / week), azathioprine (50 mg / day, 100 mg / day), and a white blood cell count warning threshold (WBC < 3.0 × 10⁻⁶). 9 / L), liver damage warning threshold (ALT increase >2 times from baseline), allergic reaction warning threshold (rash score ≥5 points), such as Figure 8 As shown, the safety risk differences of the four drugs are presented through a color gradient. The rows of the matrix represent the drugs and dosages (methotrexate 10 mg / week, methotrexate 15 mg / week, azathioprine 50 mg / day, azathioprine 100 mg / day), and the columns correspond to the three key safety indicators (white blood cell count warning threshold: WBC < 3.0 × 10⁻⁶). 9 / L; liver injury warning threshold: ALT increased by >2-fold from baseline; allergic reaction warning threshold: rash score >5 points). The cell color gradually changes from green (low risk, <30%), yellow (medium risk, 30%-50%) to red (high risk, >50%), and the risk probability of each drug is clearly marked: azathioprine 100 mg / day is red (62%) in the allergic reaction warning threshold, indicating a high risk of allergy; the liver injury warning threshold of methotrexate 15 mg / week is yellow (45%), with a medium risk; methotrexate 10 mg / week is green (white blood cells 28%, liver injury 25%, and allergy 8%) in all three indicators, with the best overall safety, and the color scale bar on the right side of the matrix marks the risk probability interval, which facilitates quick identification of high-risk drugs and provides a safety reference for drug selection in the high-risk subgroup.

[0103] A 52-year-old female patient (HLA-B*5801 positive, TPMT mutant, initial ALT 32 U / L, using methotrexate 10 mg / week) in subgroup A was monitored in real time after medication, and safety data from t0+7 days to t0+84 days were collected:

[0104] t0+7 days: WBC 4.2 x 10 9 / L, ALT 35 U / L, no rash (score 0);

[0105] t0+21 days: WBC 3.8 x 10 9 / L, ALT 45 U / L (40.6% higher than baseline), rash score 1;

[0106] t0+42 days: WBC 3.5 x 10 9 / L, ALT 52 U / L (62.5% increase), rash score 2;

[0107] t0+84 days: WBC 3.2 x 10 9 / L, ALT 58 U / L (81.2% increase), rash score 3.

[0108] Data is input into the deep learning model: the convolutional layer extracts the short-term change characteristics of WBC and ALT from t0+7 to t0+21 days, obtaining the short-term safety risk response (0.35, reflecting the initial indicator fluctuation risk); the LSTM calculates the rash score and enzyme change trend from t0+21 to t0+42 days, obtaining the medium-term safety risk trend (0.52); the fully connected layer aggregates data from t0+42 to t0+84 days, combined with the immunosuppressant risk prediction model to obtain the long-term safety risk (0.68, indicating an increased risk of liver injury). The generated real-time safety evaluation vector is [0.35, 0.52, 0.68], as shown in Figure 9As shown, the real-time safety assessment vector of the patient in subgroup A after using methotrexate 10 mg / week is a three-dimensional scatter plot, the X-axis represents the short-term safety risk response (0-1, corresponding to the fluctuation of white blood cells and ALT from t0+7 to t0+21 days), the index of this patient is 0.35, which is at a low level; the Y-axis represents the medium-term safety risk trend (0-1, corresponding to the rash score and enzymatic change from t0+21 to t0+42 days), the index is 0.52, indicating a moderate risk; the Z-axis represents the long-term safety risk (0-1, corresponding to the cumulative risk of liver damage and allergy from t0+42 to t0+84 days), the index is 0.68, indicating a higher risk, Figure 9 The vector endpoint [0.35, 0.52, 0.68] is marked with a black dot, and is connected to the three coordinate axes by dashed lines, clearly showing its position in the three-dimensional space. The short-term risk is low, the medium-term risk is moderate, and the long-term risk is high, which is consistent with the trend of persistent increase in ALT (from 32 U / L to 58 U / L).

[0109] Match the real-time safety assessment vector of the current patient with the safety benchmark matrix of subgroup A ( Figure 8 ): Calculate the cosine similarity of the short-term risk response (0.35) and the white blood cell warning threshold of methotrexate 10 mg / week (0.40) (0.93), the similarity of the medium-term risk trend (0.52) and the liver damage warning threshold (0.50) (0.96), the similarity of the long-term risk (0.68) and the allergy warning threshold (0.70) (0.91), and obtain the preliminary matching deviation value (0.07). Introduce the individual dynamic weight coefficient (W = 0.22) to correct the deviation: the patient's medication compliance (A = 0.98, normalized A' = 0.95), the change rate of comorbidities (the increase of ALT from 0.4 U / L / day to 0.7 U / L / day, normalized C' = 0.65), through the formula W = 0.6 × 0.95 + 0.4 × (1-0.65) = 0.73, the corrected deviation value is 0.07 × (1-0.73) = 0.019.

[0110] Input the corrected deviation value into the reinforcement learning model, take the "white blood cell decrease incidence", "liver damage incidence", "allergic reaction incidence" as the output indicators, and define the reward function (reward value +0.9 when the deviation value <0.02). After receiving the new data at t0+105 days (ALT 65 U / L, rash score 4), the model iteratively optimizes, and the long-term safety risk prediction value is adjusted from 0.68 to 0.75, as shown in Figure 10 As shown, the four immunosuppressive regimens of methotrexate 10 mg / week, methotrexate 15 mg / week, azathioprine 50 mg / day, and azathioprine 100 mg / day are presented in matrix form, and the white blood cell decrease warning threshold (WBC <3.0 × 10 9The matching of three safety warning thresholds: liver injury warning threshold (ALT increased by >2 times compared with baseline), allergic reaction warning threshold (skin rash score ≥5 points); the matrix cells are distinguished by color gradient, green represents high matching (low risk), yellow represents medium matching (moderate risk), and red represents low matching (high risk). Among them, methotrexate 10 mg / week is green (high matching) in all three warning thresholds, showing the best safety; methotrexate 15 mg / week is green in white blood cell and allergic threshold, and yellow in liver injury threshold; azathioprine 50 mg / day is mostly yellow in the three thresholds; azathioprine 100 mg / day is mainly red, indicating high risk. The right side of the matrix indicates the correspondence between color and matching degree (risk level), clearly showing the safety adaptation differences of different drug regimens in subgroup A patients (liver function critical abnormality + HLA-B*5801 positive).

[0111] Based on the safety indicators (liver injury incidence 75%, white blood cell decrease incidence 12%, allergic reaction incidence 8%) output by the model, the safety level is converted to "low risk (<30 points), moderate risk (30-60 points), high risk (>60 points)" by fuzzy comprehensive evaluation method, and the score is 68 points, corresponding to "high risk".

[0112] The report suggests that the dose of methotrexate be reduced to 7.5 mg / week, and ALT and blood routine should be monitored every 2 weeks to prompt the potential risk of liver injury (incidence 75%), and a backup drug (leflunomide) is recommended.

[0113] This embodiment describes in detail the safety monitoring and evaluation of immunosuppressive agents for rheumatoid arthritis patients. By integrating basic health data such as HLA-B*5801 gene polymorphism and TPMT genotype, combined with time series safety indicators such as white blood cell count and ALT, and drug metabolism data, a dynamic feature set is generated, which accurately captures the individual safety risk differences of high-risk subgroups (such as liver function critical abnormality + HLA-B*5801 positive), solving the problem of ignoring the influence of genes and baseline health status on safety in traditional evaluation. At the same time, by using hierarchical clustering and transfer learning to construct subgroup-specific safety benchmark matrix, the warning thresholds of different immunosuppressive agents are determined (such as azathioprine 100 mg / day, with an allergic risk of 62% in subgroup A), and by combining deep learning to generate evaluation vectors containing short-term, medium-term and long-term risks, the prediction error of liver injury and other risks is reduced to 0.015 after optimization by reinforcement learning model, significantly improving the risk warning accuracy. The generated high-risk evaluation report and dose adjustment suggestion (such as methotrexate reduced to 7.5 mg / week) provide clear safety guidance for clinicians, reducing the probability of adverse reactions caused by immunosuppressive agents, and improving the safety and rationality of treatment.

[0114] The above merely describes the preferred embodiments of the present application, and is not intended to limit the protection scope of the present application. The present application can have various changes and modifications for those skilled in the art; any change, modification, replacement, integration and parameter change of the embodiments within the spirit and principle of the present application, by conventional substitution or capable of realizing the same function without departing from the principle and spirit of the present application, all fall within the protection scope of the present application.

Claims

1. A method for evaluating the efficacy of internal medicine drugs based on big data analysis of electronic health records, characterized by, The method comprises the following steps: S1, extracting the basic health data, diagnosis and treatment time series data and drug intervention data of the patient from the electronic health record, and generating a dynamic feature set after spatiotemporal alignment of the data; S2, based on the preset internal medicine disease typing standard, the patient data in the dynamic feature set is stratified clustered to obtain several subgroups, and for each subgroup, the efficacy data of the historical clinical research is fused with the real world data of the subgroup through transfer learning to construct the corresponding efficacy benchmark matrix of each subgroup; S3, real-time collection of the electronic health record data of the patient after drug use, feature mapping of the collected data through a deep learning model to generate a real-time evaluation vector containing short-term physiological indicator response, medium-term symptom improvement trend and long-term prognosis risk; S4, dynamically matching the real-time evaluation vector with the corresponding subgroup efficacy benchmark matrix to calculate the matching deviation value, and the matching deviation value is corrected by introducing the patient individual dynamic weight coefficient; S5, using reinforcement learning algorithm, taking the deviation correction value after dynamic matching as input and taking the multidimensional evaluation index of drug efficacy as output, constructing an adaptive efficacy evaluation model, and continuously receiving new electronic health record data for parameter iterative optimization; S6, based on the output result of the efficacy evaluation model, generating an individualized evaluation report containing drug efficacy grade, optimal drug adjustment suggestion and potential risk warning, and the efficacy grade in the report is converted into a quantifiable score value by fuzzy comprehensive evaluation method. 2.The medical drug efficacy evaluation method based on electronic health record and big data analysis according to claim 1, wherein, The basic health data in S1 includes genetic polymorphism data, epigenetic markers and intestinal flora spectrum; the diagnosis and treatment time series data includes physiological indicator fluctuation value and symptom semantic vector recorded continuously according to time stamp; the drug intervention data includes drug metabolism enzyme activity data and blood drug concentration monitoring value. 3.The medical drug efficacy evaluation method based on electronic health record and big data analysis according to claim 1 or 2, characterized in that, In S1, the data is spatiotemporally aligned to generate a dynamic feature set. Specifically, using the patient's first medication time as the baseline timestamp t0, the gene polymorphism data G, epigenetic marker E, and gut microbiota profile M in the basic health data are mapped to static feature vectors S0 = [G, E, M] at the baseline timestamp; for the diagnosis and treatment time series data, according to the timestamp t0... i Recorded physiological index fluctuation value P(t) i ) and symptom semantic vector Q(t) i ), through time interpolation function This is converted to a normalized time coordinate relative to a baseline time, where Δt is a preset time interval; for drug metabolism enzyme activity data C(t) in drug intervention data... j ) and blood drug concentration monitoring value D(t) j ), through the spatial alignment function S(t) j )=ω1×C(t j )+ω2×D(t j Feature fusion is performed, where ω1 and ω2 are the corresponding weight coefficients; the dynamic feature integration formula is used. Generate a dynamic feature set F(t) covering the baseline time and subsequent monitoring time, where α, β, and γ are the corresponding fusion coefficients, n is the number of diagnosis and treatment time series data, and m is the number of drug intervention data. 4.The medical drug efficacy evaluation method based on electronic health record and big data analysis according to claim 1, wherein, In S2, for each subgroup, transfer learning is used to fuse the efficacy data from historical clinical studies with the real-world data of that subgroup to construct the efficacy benchmark matrix corresponding to each subgroup. Specifically, the efficacy data H = {h1, h2, ..., h...} corresponding to the disease subgroup in historical clinical studies are extracted. p } and subgroup real-world data R={r1,r2,…,r q }, feature mapping function through transfer learning Feature alignment is performed on historical data and real-world data, and the matrix element values ​​corresponding to different drugs within the fused subgroup are calculated using the following formula: λ is the fusion weight coefficient, M ij Using the element in the i-th row and j-th column of the efficacy benchmark matrix as an example, after traversing all drugs and indicators, an efficacy benchmark matrix is ​​constructed that includes the theoretical onset threshold and adverse reaction warning threshold for different drugs within subgroups. n r k represents the number of drugs within the subgroup. r This refers to the number of indicators. 5.The medical drug efficacy evaluation method based on electronic health record and big data analysis according to claim 1, wherein, The electronic health record data of the patient after taking medicine in S3 is collected in real time, and the collected data is mapped to features by a deep learning model to generate a real-time evaluation vector, specifically: the collected time series data X={x1, x2, …, x t} is input into a deep learning model composed of a convolutional layer and a long short-term memory network, the convolutional layer calculates the local change feature F s of the short-term physiological index through the filter K s , and the formula is: T s is the short-term time window, and the short-term physiological index response R s is obtained after being processed by an activation function; in the forward propagation process of the long short-term memory network, the feature accumulation amount in the medium-term time window T m is calculated through a gating mechanism h t-1 is the hidden state at the previous moment, and the medium-term symptom improvement trend R m is obtained after time series smoothing processing; the output layer of the deep learning model weights and aggregates the features in the long-term time window T l through a fully connected layer w t is the time decay weight, and the long-term prognosis risk R l is generated by combining the prognosis risk prediction function; and the real-time evaluation vector V is generated by vector splicing V=[R s , R m , R l ]. 6.The medical drug efficacy evaluation method based on electronic health record and big data analysis according to claim 1, wherein, The real-time evaluation vector is dynamically matched with the corresponding sub-group efficacy benchmark matrix in S4 to calculate a matching deviation value, specifically: determining the dimension correspondence of each element in the real-time evaluation vector V and the drug row vector M in the efficacy benchmark matrix i = [m i1 ,m i2 ,…,m ik ]The dimension correspondence is calculated by cosine similarity Based on the clinical importance of each dimension index, a weight w j is assigned, and the preliminary matching deviation value B is obtained by the deviation calculation formula i , where v j -m ij is the absolute deviation of the real-time evaluation vector and the benchmark matrix in the jth dimension, (1-S i ) is a space matching degree correction factor, and after traversing all drug row vectors, a set of matching deviation values for each drug B = {B1, B2, …, B n} is obtained. 7.The medical drug efficacy evaluation method based on electronic health record big data analysis according to claim 1 or 5, characterized in that, The matching deviation value in the S4 is corrected by introducing a patient individual dynamic weight coefficient, specifically: the individual dynamic weight coefficient is generated based on patient medication compliance real-time data and comorbidity change rate data, by obtaining patient medication compliance real-time data A and comorbidity change rate data C, normalized processing to obtain A' and C', through the weight coefficient fusion formula W = θ·A' + (1-θ)·(1-C') to generate the individual dynamic weight coefficient W, wherein θ is the weight distribution factor, and the deviation correction formula B' i i ·(1-W) is used to correct the matching deviation value B i , to obtain the corrected deviation value B' i .​ 8.The medical drug efficacy evaluation method based on electronic health record and big data analysis according to claim 1, wherein, The adaptive efficacy assessment model constructed in S5 specifically involves: setting the dynamically matched deviation correction value set B′={B′1,B′2,…,B′…} n As an input to the environmental state, the multidimensional evaluation index of drug efficacy Y={y1,y2,…,y k As output, a reward function is defined. Evaluate the output effect, where α i As the indicator weight, θ represents the expected efficacy index, β is the error penalty coefficient, and the model parameters θ are updated using the policy gradient algorithm. * The formula is: Where γ s π is the discount factor. θ* (a|s) is the strategy function. By continuously receiving new electronic health record data for iterative optimization, the efficacy evaluation model can stably output multi-dimensional evaluation indicators consistent with the actual efficacy under the input of deviation correction value, thus constructing an adaptive efficacy evaluation model. 9.The medical drug efficacy evaluation method based on electronic health record and big data analysis according to claim 1, wherein, The output result of the efficacy evaluation model in S5 is used to generate an individualized evaluation report, specifically: extracting the multi-dimensional evaluation index of drug efficacy output by the efficacy evaluation model, comparing it with the preset efficacy level classification standard, and determining the drug efficacy level L through the formula δ(p, l) is the characteristic function of index p meeting level l, combining the patient's genetic polymorphism data, drug metabolism enzyme activity data, and short-term response and medium-term improvement trend in the real-time evaluation vector, screening out the highest matching drug adjustment scheme, and determining the adverse reaction warning threshold T and long-term prognosis risk value R l , the risk probability formula P = λ p ·I(R l >T)+(1-λ p )·exp(-γ p ·(T-R l )) is used to calculate the potential risk probability P, where I is the indicator function, λ p and γ p are adjustment parameters, and an individualized evaluation report containing drug efficacy level, optimal drug adjustment suggestion, and potential risk warning is formed. 10.The medical drug efficacy evaluation method based on electronic health record and big data analysis according to claim 1 or 9, characterized in that, The efficacy rating in the report in S6 is transformed into quantifiable scores using a fuzzy comprehensive evaluation method, which involves: determining the set of multidimensional indicators for drug efficacy assessment and the efficacy rating set, and then using the membership function μ. ij =exp(-k μ ·|x i -c j |) Construct a fuzzy relation matrix R = [μ ij ] m×n , where x i c is the actual value of the i-th indicator. j Let k be the center value of the j-th level. μ The membership adjustment coefficient is used to determine the weight vector based on the clinical importance of each indicator. A comprehensive evaluation vector S = {s1, s2, ..., s...} is obtained through fuzzy comprehensive evaluation. n }, take the grade score corresponding to the highest membership degree or calculate it through weighted average. A quantifiable efficacy rating score was obtained, where t j This represents the quantitative score corresponding to the efficacy rating set.

Citation Information

Patent Citations

  • Preparation evaluation system based on electronic medical record data

    CN118919097A

Cited By

  • Medication method and system based on old people physical rehabilitation tracking data feedback

    CN121416132A

  • Neurosurgery patient health data intelligent management method

    CN121687357A

  • Coronary heart disease risk assessment system based on gene polymorphism of helicobacter pylori infected patient

    CN121885224A