Large model-based drug data mining method, device, and medium

By using a large-model-based drug data mining method, baseline characteristics of patient cohorts are obtained and analyzed, a comparable cohort is constructed, and drug data mining is performed using a large language model. This solves the problems of confounding bias and single analytical dimension in clinical drug practice, and realizes fully automated and efficient drug value assessment.

CN122117476APending Publication Date: 2026-05-29BEIJING HUIJI ZHIYI TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING HUIJI ZHIYI TECH CO LTD
Filing Date
2026-02-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies suffer from confounding biases in clinical drug practice, resulting in low reliability of conclusions and limited analytical dimensions, making it difficult to adapt to diverse clinical expressions and achieve full-process automation.

Method used

By using a large-model-based drug data mining approach, clinical data of sample patients are obtained, propensity scores are predicted based on baseline features, a comparable patient cohort is constructed, and drug data mining is performed using a large language model, including analysis of medication behavior, efficacy, safety, and cost-effectiveness.

Benefits of technology

It enables fully automated analysis of drugs in the real world, controls confounding biases, improves the reliability and efficiency of drug data mining, and provides comprehensive and comparable drug value assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117476A_ABST
    Figure CN122117476A_ABST
Patent Text Reader

Abstract

The application provides a drug data mining method, device and medium based on a large model, which comprises the following steps: obtaining clinical data of sample patients, wherein the sample patients comprise first patients receiving a target drug produced by a first manufacturer and second patients receiving a target drug produced by a second manufacturer; predicting a propensity score of the sample patients based on baseline characteristics of the sample patients, generating a first patient queue of the first manufacturer and a second patient queue of the second manufacturer based on the propensity score; performing drug data mining on the clinical data of the first patient queue and the second patient queue respectively to obtain drug data mining results of the first patient queue and the second patient queue; and generating a control mining result of the target drug produced by the first manufacturer based on the drug data mining results of the first patient queue and the second patient queue. The method, device and medium provided by the application guarantee the reliability and scientificity of the control analysis and improve the drug data mining efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, device and medium for drug data mining based on a large model. Background Technology

[0002] Currently, the monitoring of the long-term efficacy and safety of drugs in real clinical practice is usually achieved through rule-based systems, manual statistical analysis methods, or traditional machine learning methods using Natural Language Processing (NLP) tools.

[0003] However, in observational studies, the above methods simply compare indicators between groups while ignoring patient differences, leading to confounding bias in the comparison process and directly affecting the reliability of the conclusions. Summary of the Invention

[0004] This invention provides a drug data mining method, device, and medium based on a large model to address the low reliability of drug data mining and analysis in related technologies.

[0005] This invention provides a drug data mining method based on a large model, comprising: Acquire clinical data from sample patients, including a first patient receiving treatment with the target drug manufactured by a first manufacturer and a second patient receiving treatment with the target drug manufactured by a second manufacturer; Based on the baseline characteristics of the sample patients, a propensity score is predicted for the sample patients. Based on the propensity score, a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer are generated. The propensity score is a predicted probability that the sample patients are inclined to receive treatment with the target drug produced by the first manufacturer. Drug data mining was performed on the clinical data of the first patient cohort and the second patient cohort respectively to obtain the drug data mining results of the first patient cohort and the second patient cohort. Based on the drug data mining results of the first patient cohort and the second patient cohort, control mining results of the target drug produced by the first manufacturer are generated.

[0006] According to a drug data mining method based on a large model provided by the present invention, the step of generating a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer based on the propensity score includes: For each first patient, a second patient whose propensity score is close to that of the first patient is selected from all second patients and used as a control patient for the first patient; The first patient queue is generated based on the first patient, and the second patient queue is generated based on the control patients of the first patient.

[0007] According to a drug data mining method based on a large model provided by the present invention, the step of generating a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer based on the propensity score includes: The treatment weight for the first patient is determined based on the propensity score of the first patient, and the control weight for the second patient is determined based on the propensity score of the second patient. The lower the propensity score, the greater the treatment weight, and the higher the propensity score, the greater the control weight. The first patient queue is generated based on the first patient and the processing weight of the first patient, and the second patient queue is generated based on the second patient and the control weight of the second patient.

[0008] According to a drug data mining method based on a large model provided by the present invention, the step of performing drug data mining on clinical data of the first patient cohort and the second patient cohort respectively includes: Perform at least two of the following analyses on the clinical data of the first patient cohort and the second patient cohort: medication behavior analysis, efficacy analysis, safety analysis, and cost-effectiveness analysis.

[0009] According to a drug data mining method based on a large model provided by the present invention, medication behavior analysis is performed on clinical data of a first patient cohort and a second patient cohort, respectively, including: Perform at least one of the following on the clinical data of the first patient cohort and the second patient cohort: adherence analysis, medication switching analysis, concomitant medication analysis, and medication change treatment analysis. The compliance analysis includes statistics on the percentage of time spent using medication; the medication switching analysis includes statistics on the percentage of patients switching and / or the frequency of medication switching; the concomitant medication analysis includes statistics on the incidence of concomitant medication and / or the statistics on medication combinations; and the medication change treatment analysis includes statistics on the medication change rate and / or the identification and ranking of medication change pathways.

[0010] According to a drug data mining method based on a large model provided by the present invention, effectiveness analysis is performed on clinical data from the first patient cohort and the second patient cohort, respectively, including: Based on a large language model, the efficacy indicators of the target drug and the changing trends of the efficacy indicators are mined from medical knowledge information; Based on the large language model, the values ​​of the efficacy indicators are extracted from the clinical data of the first patient cohort and the second patient cohort, respectively, and the efficacy rates of the target drugs produced by the first manufacturer and the second manufacturer are determined by the changing trends of the values ​​and the changing trends of the efficacy indicators.

[0011] According to a drug data mining method based on a large model provided by the present invention, safety analysis is performed on clinical data of the first patient cohort and the second patient cohort, respectively, including: Based on a large language model, safety indicators of the target drug and the changing trends of these safety indicators when adverse reactions occur are mined from medical knowledge information. Based on the large language model, the values ​​of the safety indicators are extracted from the clinical data of the first patient cohort and the second patient cohort, respectively. The adverse reaction rates of the target drugs produced by the first manufacturer and the second manufacturer are determined by the changing trends of the values ​​and the changing trends when adverse reactions occur.

[0012] According to a drug data mining method based on a large model provided by the present invention, economic analysis is performed on the clinical data of the first patient cohort and the second patient cohort, respectively, including: Statistical amounts of economic indicators were extracted from the clinical data of the first patient cohort and the second patient cohort, respectively. The economic indicators include at least one of the following: target drug cost, out-of-pocket cost of target drug, total medical expenses, examination fees, testing fees, and treatment fees.

[0013] The present invention also provides a drug data mining device based on a large model, comprising: The data acquisition unit is used to acquire clinical data of sample patients, including a first patient receiving treatment with a target drug manufactured by a first manufacturer and a second patient receiving treatment with a target drug manufactured by a second manufacturer. A sample balancing unit is used to predict the propensity score of the sample patients based on the baseline characteristics of the sample patients, and to generate a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer based on the propensity score, wherein the propensity score is a predicted probability that the sample patients are inclined to receive treatment with the target drug produced by the first manufacturer. The data mining unit is used to perform drug data mining on the clinical data of the first patient cohort and the second patient cohort respectively, and obtain the drug data mining results of the first patient cohort and the second patient cohort. The data comparison unit is used to generate comparison mining results of the target drug produced by the first manufacturer based on the drug data mining results of the first patient cohort and the second patient cohort.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the drug data mining method based on a large model as described above.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the drug data mining method based on a large model as described above.

[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the drug data mining method based on a large model as described above.

[0017] This invention provides a large-scale model-based drug data mining method, device, and medium. It extracts baseline features from sample patients and assigns propensity score to these features. The resulting propensity score is then applied to construct patient cohorts for comparison of target drugs from different manufacturers, resulting in balanced and comparable patient cohorts. Drug data mining is performed on each patient cohort, and the mining results are compared. This approach controls confounding bias at its source, ensuring the reliability and scientific rigor of the comparative analysis. Furthermore, it automates the entire process from data acquisition, cohort construction, sample balancing, data mining, to result generation, effectively improving the efficiency of drug data mining and significantly enhancing the timeliness of real-world research on target drugs. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is one of the flowcharts of the drug data mining method based on a large model provided by the present invention.

[0020] Figure 2 This is a flowchart illustrating the effectiveness analysis method provided by the present invention.

[0021] Figure 3 This is a flowchart illustrating the security analysis method provided by the present invention.

[0022] Figure 4 This is the second flowchart of the drug data mining method based on a large model provided by the present invention.

[0023] Figure 5 This is a schematic diagram of the structure of the drug data mining device based on a large model provided by the present invention.

[0024] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0026] All actions involving the acquisition of signal information or data in this invention are carried out in compliance with the relevant data protection laws and policies of the country where the device is located, and with the authorization granted by the owner of the device.

[0027] Traditional clinical trials for new drugs (such as randomized controlled trials, RCTs) are conducted in strictly controlled environments, and their conclusions have limitations when generalized to complex and heterogeneous patient populations in the real world. This is especially true for drugs that have passed the "generic drug consistency evaluation," where assessments of their long-term efficacy, safety, and continued equivalence to the original drug in real-world clinical practice are typically achieved through rule-based systems, manual statistical analysis methods, or traditional machine learning methods using natural language processing tools.

[0028] Rule-based systems, in particular, incorporate core indicators and judgment logic predefined through expert knowledge. For example, they can extract specific test values ​​from medical records and set thresholds to determine improvement in treatment efficacy. However, these systems require manual data annotation and rule writing, and the rules are often rigid and difficult to adapt to diverse clinical expressions.

[0029] Manual statistical analysis involves medical experts manually screening patient clinical records, extracting key indicators (such as laboratory results), and performing statistical analysis. This method relies heavily on human intervention, is time-consuming, highly subjective, and cannot be applied on a large scale.

[0030] Traditional machine learning methods extract structured information using natural language processing tools, but this requires a large amount of labeled data to train the model beforehand, and the generalization ability is limited, making it difficult to cover the complex context of real-world data.

[0031] The above-mentioned methods generally suffer from the following four shortcomings: First, it suffers from poor flexibility and adaptability. Predefined rule bases or dedicated natural language processing models often struggle to cover the complex context of real-world data. When dealing with new drugs, new efficacy indicators, or different data sources, the system requires experts to redefine rules or retrain the model, resulting in high maintenance costs and poor scalability.

[0032] Secondly, the above methods usually focus only on a single aspect such as effectiveness, with a single analytical dimension and a lack of systematicity.

[0033] Furthermore, it is difficult to effectively control for confounding bias. In observational studies, ignoring differences in patient baseline characteristics can lead to unreliable conclusions.

[0034] Finally, existing processes are mostly fragmented and not end-to-end. Modules such as data extraction, indicator calculation, and statistical analysis are usually separate, requiring a lot of manual data conversion and connection in between. This not only increases the workload but also makes it very easy to introduce human error into the process.

[0035] To address the aforementioned problems, embodiments of the present invention provide a drug data mining method based on a large model. Figure 1 This is one of the flowcharts illustrating the drug data mining method provided by the present invention, such as... Figure 1 As shown, the method includes: Step 110: Obtain clinical data of sample patients, including a first patient receiving treatment with the target drug manufactured by a first manufacturer and a second patient receiving treatment with the target drug manufactured by a second manufacturer.

[0036] Here, the target drug refers to the drug that requires data mining, or it can be understood as the drug that requires value assessment. In observational studies, relevant data on the same target drug generated by different manufacturers can be compared. In this embodiment of the invention, different manufacturers are divided into a first manufacturer and a second manufacturer, where each manufacturer can include one or more manufacturers. The first manufacturer can be considered as the manufacturer whose target drug needs to be value assessed, and the second manufacturer can be considered as the manufacturer whose target drug needs to be compared with the target drug produced by the first manufacturer. For example, each generic drug manufacturer can be designated as the first manufacturer, and the original drug manufacturer as the second manufacturer for comparison; or, for example, any one of all generic drug manufacturers and original drug manufacturers can be designated as the first manufacturer, and all other manufacturers as the second manufacturers for comparison. This embodiment of the invention does not specifically limit this approach.

[0037] For the target drug, patients receiving the target drug treatment can be recorded as sample patients, and those receiving the target drug treatment from the first manufacturer can be recorded as first patients, and those receiving the target drug treatment from the second manufacturer can be recorded as second patients.

[0038] Therefore, clinical data from sample patients can be collected. This clinical data may include at least one of the following: basic patient information, medical records, clinical observation data, and economic data. Basic patient information may include age, gender, etc.; medical records may include medical orders, diagnoses, surgical procedures, etc.; clinical observation data may include unstructured or semi-structured text such as laboratory reports and examination reports; and economic data may include expense details, settlement lists, and medical insurance payment information. Similarly, the clinical data collected from the sample patients may include clinical data from the first patient and clinical data from the second patient.

[0039] Step 120: Based on the baseline characteristics of the sample patients, predict the propensity score of the sample patients; based on the propensity score, generate a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer, wherein the propensity score is a predicted probability that the sample patients are inclined to receive treatment with the target drug produced by the first manufacturer.

[0040] Specifically, in traditional observational studies, drug data mining can be performed on the clinical data of the first patient and the clinical data of the second patient separately, and the drug data mining results of the first patient and the second patient can be directly compared. However, this approach obviously ignores the differences between the first patient and the second patient. The group consisting of the first patient and the group consisting of the second patient may not be comparable. This approach introduces confounding factors during the comparison process, making the comparison conclusions unreliable.

[0041] To address this issue, in this embodiment of the invention, baseline features are extracted from the sample patients, and propensity score is calculated based on the baseline features. The resulting propensity score is then applied to construct a patient cohort for comparison with target drugs from different manufacturers, thereby obtaining a balanced and comparable patient cohort.

[0042] For any sample patient, the baseline characteristics refer to the inherent attributes of the patient that will not change in the short term during treatment. For example, baseline characteristics may include at least one of demographic characteristics, clinical characteristics, and medical resource characteristics. Demographic characteristics may include age, gender, etc.; clinical characteristics may include daily medication dosage, comorbidities, etc.; and medical resource characteristics may include hospital level, insurance type, etc.

[0043] To compare the differences in drug data mining results from different manufacturers and achieve a fair comparison of samples, after obtaining the baseline characteristics of the sample patients, the predicted probability that a sample patient is inclined to receive the target drug treatment produced by the first manufacturer can be estimated based on these baseline characteristics. This predicted probability is then recorded as the propensity score for that sample patient. For example, a patient feature matrix can be constructed using the baseline characteristics of all sample patients, and this matrix can be input into a classification model such as logistic regression to obtain the propensity score for each sample patient output by the model. The propensity score here ranges from 0 to 1, representing the probability that, given the patient's baseline characteristics, the patient will receive the target drug treatment produced by the first manufacturer, i.e., the probability that the patient will be assigned to the first patient cohort of the first manufacturer.

[0044] After obtaining the propensity score for each sample patient, a first patient cohort from the first manufacturer and a second patient cohort from the second manufacturer can be formed based on these scores. For example, a first patient cohort can be constructed based on the first patients in the sample patients, and a second patient cohort can be constructed by finding second patients from the sample patients whose propensity scores are close to those of the first patients in the first patient cohort. This makes the distribution of propensity scores in the first and second patient cohorts similar, meaning that the baseline characteristic distributions of the patients in the first and second patient cohorts are nearly identical. Alternatively, the weights of the sample patients can be determined based on their propensity scores, and then the first and second patient cohorts can be constructed based on these weighted sample patients. The existence of these weights makes the overall baseline characteristic distributions of the first and second patient cohorts consistent, thus making the first and second patient cohorts comparable.

[0045] Step 130: Based on a large language model, perform drug data mining on the clinical data of the first patient cohort and the second patient cohort respectively to obtain the drug data mining results of the first patient cohort and the second patient cohort.

[0046] Specifically, after obtaining the first patient cohort for the first manufacturer and the second patient cohort for the second manufacturer, drug data mining can be performed on the clinical data of the first patient cohort and the clinical data of the second patient cohort, respectively.

[0047] The drug data mining described here can be achieved through a large language model, or through a large language model and a statistical analysis engine. In this embodiment of the invention, the large language model (LLM) is also referred to as a large model. For example, by leveraging the powerful natural language understanding capabilities of the large language model and the powerful data analysis capabilities of the statistical analysis engine, information related to evaluation indicators for the target drug can be comprehensively mined from clinical data. For instance, the statistical analysis engine can be used to mine drug use behavior and / or safety data from clinical data, while the large language model can be used to mine efficacy and / or safety data from clinical data.

[0048] Thus, we can obtain the drug data mining results for the first patient cohort and the drug data mining results for the second patient cohort.

[0049] Step 140: Based on the drug data mining results of the first patient cohort and the second patient cohort, generate the control mining results of the target drug produced by the first manufacturer.

[0050] Specifically, the drug data mining results of the first patient cohort and the second patient cohort can be compared to obtain the control mining results of the target drug generated by the first manufacturer, with the target drug produced by the second manufacturer as the control.

[0051] In the method provided in this embodiment of the invention, baseline features are extracted from sample patients, and propensity score is calculated based on these baseline features. The resulting propensity score is then applied to construct patient cohorts for comparison with target drugs from different manufacturers, thereby obtaining balanced and comparable patient cohorts. Drug data mining is performed on each patient cohort, and the mining results are compared. This approach controls confounding bias at its source, ensuring the reliability and scientific rigor of the comparative analysis. Furthermore, the entire process from data acquisition, cohort construction, sample balancing, data mining to result generation is automated, effectively improving the efficiency of drug data mining and significantly enhancing the timeliness of real-world research on target drugs.

[0052] Based on the above embodiments, step 120, which involves generating a first patient cohort for the first manufacturer and a second patient cohort for the second manufacturer based on the propensity score, includes: For each first patient, a second patient whose propensity score is close to that of the first patient is selected from all second patients and used as a control patient for the first patient; The first patient queue is generated based on the first patient, and the second patient queue is generated based on the control patients of the first patient.

[0053] Specifically, for the case where each generic drug manufacturer is designated as the first manufacturer and the original drug manufacturer is designated as the second manufacturer for control, in order to screen and balance two groups of samples for each first manufacturer and the second manufacturer, and to construct comparable treatment and control groups to eliminate selection bias in observational studies, control patients can be matched for each first patient of the first manufacturer.

[0054] After obtaining the propensity score for each sample patient, for each first patient, a second patient with a propensity score close to that of the first patient can be selected from all the second patients and used as the first patient's control patient. It can be understood that a second patient with a propensity score close to that of the first patient, i.e., a second patient with similar baseline characteristics to the first patient, can also be understood as a second patient comparable to the first patient, and here referred to as the first patient's control patient.

[0055] For example, you can find second patients from all the second patients whose propensity score difference from that of the first patient is less than a difference threshold.

[0056] For example, propensity scores can be generated for all sample patients. Perform logit transformation, i.e. Therefore, the logit transformed value of the propensity score of the first patient can be used to perform K-nearest neighbor matching on the logit transformed values ​​of each second patient. This allows for finding the control patient with the closest propensity score for each first patient. Furthermore, a maximum allowed difference in logit transformed values ​​can be set for K-nearest neighbor matching, for example, 0.25, allowing only control patients with logit transformed value differences within 0.25 to be matched with the first patient. The logit transformation of the propensity score aims to convert the propensity score between 0 and 1 to the entire real number axis (-∞ to +∞), making the distribution closer to a normal distribution and easier for the matching algorithm to process.

[0057] After obtaining the control patients for each first patient, each first patient can be placed into the first patient cohort, and the control patients matched with the first patients can be placed into the second patient cohort. The resulting second patient cohort can be regarded as the control group of the first patient cohort.

[0058] Therefore, when comparing the drug data mining results of the first patient cohort and the second patient cohort, the drug data mining results of each first patient in the first patient cohort can be averaged on each indicator, and the drug data mining results of each second patient in the second patient cohort can be averaged on each indicator. Thus, the mean values ​​of each indicator in the first patient cohort and the mean values ​​of each indicator in the second patient cohort can be compared at the indicator level.

[0059] Based on the above embodiments, step 120, which involves generating a first patient cohort for the first manufacturer and a second patient cohort for the second manufacturer based on the propensity score, includes: The treatment weight for the first patient is determined based on the propensity score of the first patient, and the control weight for the second patient is determined based on the propensity score of the second patient. The lower the propensity score, the greater the treatment weight, and the higher the propensity score, the greater the control weight. The first patient queue is generated based on the first patient and the processing weight of the first patient, and the second patient queue is generated based on the second patient and the control weight of the second patient.

[0060] Specifically, for the case where any one of the generic drug manufacturers and the original drug manufacturers is taken as the first manufacturer and all other manufacturers are taken as the second manufacturer as the control, the baseline characteristic differences between patient cohorts of different manufacturers can be eliminated by sample weighting, thereby obtaining comparable control mining results between manufacturers.

[0061] For the first patient, the weight (treatment weight) of the first patient in the first patient cohort can be determined based on the patient's propensity score. Here, the propensity score and treatment weight are inversely proportional; a higher propensity score results in a lower treatment weight, and vice versa. For example, suppose the propensity score is... Processing weights Therefore, by weighting the first patient, the first patient with a lower propensity score, that is, the first patient whose baseline characteristics are closer to the second patient who serves as the control, can be given a higher weight in the first patient cohort, making the first patient cohort as a whole closer to the second patient cohort that serves as the control cohort.

[0062] For the second patient, the weight of the second patient in the second patient cohort, i.e., the control weight, can be determined based on the second patient's propensity score. Here, the propensity score is directly proportional to the control weight; a higher propensity score results in a larger control weight, and a lower propensity score results in a smaller control weight. For example, suppose the propensity score is... weight comparison Therefore, by weighting the second patient, the second patient with a higher propensity score, that is, the second patient whose baseline characteristics are closer to those of the first patient, can be given a higher weight in the second patient cohort, making the second patient cohort as a whole closer to the first patient cohort.

[0063] Based on this, a first patient cohort can be constructed by combining the first patient and the treatment weight of the first patient, and a second patient cohort can be constructed by combining the second patient and the control weight of the second patient. The first patient cohort and the second patient cohort constructed in this way tend to be consistent in the baseline feature distribution.

[0064] Therefore, the treatment weights of each first patient in the first patient cohort and the control weights of each second patient in the second patient cohort can provide a researchable patient screening method for subsequent indicator comparisons between manufacturers. For example, when comparing the drug data mining results of the first patient cohort and the second patient cohort, the drug data mining results of each first patient in the first patient cohort can be weighted and summed on each indicator according to the treatment weight of each first patient, and the drug data mining results of each second patient in the second patient cohort can be weighted and summed on each indicator according to the control weight of each second patient. Thus, the weighted sums corresponding to each indicator in the first patient cohort and the weighted sums corresponding to each indicator in the second patient cohort can be compared at the indicator level.

[0065] Based on any of the above embodiments, step 130, which involves performing drug data mining on the clinical data of the first patient cohort and the second patient cohort respectively, includes: Perform at least two of the following analyses on the clinical data of the first patient cohort and the second patient cohort: medication behavior analysis, efficacy analysis, safety analysis, and cost-effectiveness analysis.

[0066] Specifically, for any patient cohort in the first or second patient cohort, drug data mining can be performed on the clinical data of that patient cohort. This can be manifested as at least two of the following: medication behavior analysis, efficacy analysis, safety analysis, and cost-effectiveness analysis of the clinical data of that patient cohort, thereby achieving a comprehensive and automated evaluation of the patient cohort.

[0067] Medication behavior analysis refers to the automated calculation of indicators such as adherence, medication switching, concomitant medication, and medication change treatment based on structured clinical data. Efficacy analysis extracts numerical values ​​of indicators reflecting the efficacy of the target drug from clinical data, and then judges whether the values ​​improve after medication, i.e., whether the medication is effective, by analyzing changes in these values, and thus statistically analyzing the efficacy rates of target drugs from different manufacturers. Safety analysis analyzes whether the use of the target drug causes adverse reactions based on clinical data, and then judges the safety of target drugs from different manufacturers by statistically analyzing the adverse reaction rates. Cost-effectiveness analysis can extract various costs from clinical data to compare the cost differences when receiving target drug treatment from different manufacturers.

[0068] Therefore, the drug data mining results generated for any patient cohort may include at least two of the following: medication behavior analysis results, efficacy analysis results, safety analysis results, and economic analysis results. Furthermore, the aforementioned medication behavior analysis results, efficacy analysis results, safety analysis results, and economic analysis results may be at the patient level or at the patient cohort level, and the embodiments of the present invention do not specifically limit this.

[0069] The method provided in this invention combines medication behavior analysis, efficacy analysis, safety analysis, and economic analysis to achieve a complete and comprehensive value assessment of the target drug. This overcomes the limitations of related technologies that only focus on a single dimension and can provide three-dimensional and comprehensive real-world evidence for drug value assessment.

[0070] Furthermore, in the aforementioned drug use behavior analysis, efficacy analysis, safety analysis, and economic analysis, large-scale language models can be invoked. As a general and adaptive information extraction engine, large-scale language models replace traditional rule-based or pre-trained dedicated models, enabling flexible analysis of multiple drugs with a single model and breaking through the bottlenecks of rigidity and difficulty in expansion of traditional methods.

[0071] Based on any of the above embodiments, in step 130, medication behavior analysis is performed on the clinical data of the first patient cohort and the second patient cohort, including: Perform at least one of the following on the clinical data of the first patient cohort and the second patient cohort: adherence analysis, medication switching analysis, concomitant medication analysis, and medication change treatment analysis. The compliance analysis includes statistics on the percentage of time spent using medication; the medication switching analysis includes statistics on the percentage of patients switching and / or the frequency of medication switching; the concomitant medication analysis includes statistics on the incidence of concomitant medication and / or the statistics on medication combinations; and the medication change treatment analysis includes statistics on the medication change rate and / or the identification and ranking of medication change pathways.

[0072] Specifically, for any one of the first and second patient cohorts, the clinical data of that patient cohort are analyzed for medication behavior, which may include at least one of the following: adherence analysis, medication switching analysis, concomitant medication analysis, and medication change treatment analysis.

[0073] The adherence analysis aims to quantify the continuity of patients' adherence to prescribed medications during hospitalization. Adherence analysis assesses the consistency of target drug treatment by calculating the ratio of actual usage time to expected usage time. Considering that non-adherence in a hospital setting is usually not caused by the patient but stems from clinical decisions (such as discontinuation due to side effects), surgical fasting, or administrative reasons (such as drug shortages), the percentage of target drug usage time is used to characterize the adherence analysis results.

[0074] Here, the formula for calculating the percentage of time spent using medication can be: Switching therapy analysis focuses on switching the same target drug between different manufacturers. The purpose of switching therapy analysis is to assess the stability of the hospital's drug supply chain and the impact of switching on treatment consistency and potential clinical risks. Potential clinical risks referred to here may include bioequivalence concerns.

[0075] Medication switching analysis can be achieved by statistically analyzing the patient-level switching rate and / or the drug-level switching frequency. The patient-level switching rate can be calculated using the following formula: Drug-level switching frequency can be obtained by counting the number of times the target drug has switched manufacturers in all medication events.

[0076] Concomitant medication refers to the use of other drugs belonging to the same pharmacological class or treating the same indication during the treatment with the target drug. The purpose of concomitant medication analysis is to identify potential unnecessary duplicate medications or non-standard combination therapy regimens, thereby avoiding risks that may lead to drug overdose, increased adverse reaction risks, and additional medical costs.

[0077] Concomitant medication analysis can be achieved by statistically analyzing the incidence of concomitant medication and / or drug combinations. The incidence of concomitant medication can be calculated using the following formula: In addition, statistics on drug combinations can specifically involve discovering, statistically analyzing, and ranking the most common combinations of target drugs and concomitant drugs in clinical data.

[0078] For example, a frequency statistics algorithm can be used to automatically calculate the frequency and percentage of all occurrences of the "target drug-companion drug" combination, and then automatically sort them in descending order of frequency, outputting the top N most common drug combinations. Here, the value of N can be any positive integer, such as N=5, N=10, etc.

[0079] For example, for the identified high-frequency drug combinations, namely the top N most common drug combinations in the output above, a large language model can be invoked based on medical knowledge information such as clinical guidelines and the instructions for use of the target drugs to analyze whether the drug combination is a reasonable combination of drugs or an unreasonable duplicate drug use. For unreasonable duplicate drug use, the large language model can automatically generate a potential risk description for such drug combinations, thereby providing doctors with more in-depth clues for review.

[0080] Switching therapy refers to a situation where a patient initially uses a target drug, but the subsequent treatment regimen is changed to another drug of the same class. The purpose of switching therapy analysis is to track treatment failure, intolerance, or ineffectiveness of the initial treatment, and it provides important real-world evidence for assessing drug efficacy and safety.

[0081] Dressing change treatment analysis can be achieved by statistically analyzing the dressing change rate and / or identifying and ranking dressing change paths. The dressing change rate can be calculated using the following formula: In addition, the identification and sorting of dressing change routes can specifically involve discovering and sorting the most common dressing change sequences in clinical data.

[0082] For example, a sequence pattern mining algorithm can be used to automatically analyze the medication sequences of all patients, statistically analyze the frequency and proportion of medication change paths such as "drug A → drug B" and "drug A → drug C", and automatically sort them in descending order of frequency, outputting the top N most common medication change paths. Here, N can be any positive integer, such as N=5, N=10, etc.

[0083] For example, for the identified high-frequency medication change paths—the top N most common paths output above—a large language model can be invoked to infer and generate possible reasons for the medication change based on medical knowledge such as clinical guidelines, the pharmacological properties of the target drug and the replacement drug, and common adverse reactions. This provides crucial insights for clinical review. For instance, "drug A → drug B" might be related to the adverse reaction of drug A causing cough, aiming to find alternative treatment options.

[0084] Therefore, by analyzing the medication behavior of each manufacturer's patient cohorts, the medication behavior analysis results of each manufacturer's target drugs can be obtained.

[0085] In the process of generating control mining results based on drug data mining results from the first and second patient cohorts, it is possible to compare whether there are significant differences between the drug use behavior analysis results of the first and second patient cohorts, thereby providing a comprehensive and quantitative basis for evaluating the clinical use patterns and stability of target drugs produced by various manufacturers.

[0086] Based on any of the above embodiments, in step 130, an effectiveness analysis is performed on the clinical data of the first patient cohort and the second patient cohort, including: Based on a large language model, the efficacy indicators of the target drug and the changing trends of the efficacy indicators are mined from medical knowledge information; Based on the large language model, the values ​​of the efficacy indicators are extracted from the clinical data of the first patient cohort and the second patient cohort, respectively, and the efficacy rates of the target drugs produced by the first manufacturer and the second manufacturer are determined by the changing trends of the values ​​and the changing trends of the efficacy indicators.

[0087] Specifically, for any patient cohort in the first or second patient cohort, the effectiveness analysis of the clinical data of that patient cohort can be performed based on a large language model.

[0088] First, for a target drug, a large-scale language model can be invoked to learn from medical knowledge information such as medical guidelines and literature and generate efficacy indicators for the target drug. These efficacy indicators reflect the effectiveness of the target drug; for example, if the target drug is an antihypertensive drug, the large-scale language model could output efficacy indicators such as blood pressure and heart rate. Furthermore, the large-scale language model can also generate trends in the changes of efficacy indicators that reflect the effectiveness of the target drug. For example, if the target drug is an antihypertensive drug, the large-scale language model could output trends such as a decrease in blood pressure.

[0089] Here, to analyze the efficacy indicators and trends of different drugs, prompt engineering can be used to dynamically adapt large language models to different drugs without the need for predefined rules.

[0090] After obtaining the efficacy indicators and their trends for the target drug, a large-scale language model can be used to extract specific numerical values ​​of the efficacy indicators from clinical data. Furthermore, by determining whether the trends in these efficacy indicator values ​​before and after drug use are consistent with the trends in efficacy indicators reflecting the drug's effectiveness, the effectiveness of the target drug for the patient can be judged. For example, a large-scale language model can extract blood pressure measurements from clinical examination reports. If the blood pressure decreases from 140 / 90 mmHg to 120 / 80 mmHg before and after medication, it can be determined that this trend aligns with the effectiveness of the antihypertensive drug, thus indicating that the antihypertensive drug is effective for the patient corresponding to this clinical data.

[0091] In the aforementioned process, when using a large language model to extract specific values ​​of efficacy indicators from clinical data, accuracy can be improved through multiple rounds of validation. These multiple rounds of validation can be achieved through methods such as repeated extraction and consistency checks. Furthermore, when using a large language model to determine the changing trends of efficacy indicators in clinical data, multiple inference validations based on clinical logic can also be performed to reduce misjudgments.

[0092] Based on this, the efficacy rate of the target drug produced by each manufacturer can be calculated. The formula for calculating the efficacy rate is as follows: Therefore, the efficacy rate of the target drug produced by each manufacturer can be used as the efficacy analysis result of each manufacturer's patient cohort.

[0093] In the process of generating control mining results based on drug data mining results from the first and second patient cohorts, it is possible to compare whether there are significant differences between the efficacy analysis results of the first and second patient cohorts. For example, a statistical t-test can be used to compare the efficacy differences of target drugs produced by different manufacturers, thereby automatically identifying a list of drugs with significant differences. Furthermore, the resulting drug list and associated clinical data clues can be further reviewed and confirmed by physicians, ultimately generating efficacy evaluation reports for target drugs produced by various manufacturers, such as efficacy evaluation reports for original drugs and generic drugs.

[0094] Based on any of the above embodiments Figure 2 This is a flowchart illustrating the effectiveness analysis method provided by the present invention, as shown below. Figure 2 As shown, the efficacy indicators of the target drug can first be generated based on a large model. Both clinical data and efficacy indicators are then input into the large model. The model extracts the treatment duration of the target drug and the values ​​of the efficacy indicators from the clinical data. Subsequently, the effectiveness of the target drug can be determined based on the efficacy indicator values ​​before and after the treatment duration, as determined by the large model. These conclusions are then output after multiple validations for use in the statistical analysis of efficacy indicators.

[0095] Based on any of the above embodiments, in step 130, a safety analysis is performed on the clinical data of the first patient cohort and the second patient cohort, including: Based on a large language model, safety indicators of the target drug and the changing trends of these safety indicators when adverse reactions occur are mined from medical knowledge information. Based on the large language model, the values ​​of the safety indicators are extracted from the clinical data of the first patient cohort and the second patient cohort, respectively. The adverse reaction rates of the target drugs produced by the first manufacturer and the second manufacturer are determined by the changing trends of the values ​​and the changing trends when adverse reactions occur.

[0096] Specifically, for any patient cohort in the first or second patient cohort, a safety analysis of the clinical data of that patient cohort can be performed based on a large language model.

[0097] First, for a target drug, a large-scale language model can be invoked to learn from medical knowledge information such as medical guidelines and literature and generate safety indicators for the target drug. These safety indicators reflect the safety of the target drug. For example, if the target drug is a lipid-lowering drug, the safety indicators output by the large-scale language model could include ALT (Alanine Aminotransferase) and AST (Aspartate Aminotransferase). Furthermore, the large-scale language model can also generate trends in safety indicators that reflect adverse reactions. For example, if the target drug is a lipid-lowering drug, the trend output by the large-scale language model could include an ALT / AST increase greater than 3 times the normal value.

[0098] Here, to analyze the safety indicators achieved by different drugs and the changing trends of these indicators when adverse reactions occur, prompting engineering can be used to dynamically adapt large language models to different drugs without the need for predefined rules.

[0099] After obtaining the safety indicators of the target drug and their trends during adverse reactions, a large-scale language model can be used to extract specific values ​​of the safety indicators from clinical data. Furthermore, by comparing the trends of these safety indicator values ​​before and after drug use with those during adverse reactions, it can be determined whether the patient experienced an adverse reaction. For example, a large-scale language model can extract ALT / AST levels from clinical test results and determine if an ALT / AST level exceeding three times the normal value indicates an adverse reaction, thus assessing the safety of the lipid-lowering drug for the patient corresponding to that clinical data.

[0100] In the aforementioned process, when using a large language model to extract specific values ​​of safety indicators from clinical data, accuracy can be improved through multiple rounds of verification. These multiple rounds of verification can be implemented through methods such as repeated extraction and consistency checks. Furthermore, when using a large language model to determine the changing trends of safety indicators in clinical data, multiple inference verifications based on clinical logic can also be performed to reduce misjudgments.

[0101] Based on this, the adverse reaction rate of the target drug produced by each manufacturer can be calculated. The formula for calculating the adverse reaction rate is as follows: Therefore, the adverse reaction rate of the target drug produced by each manufacturer can be used as the result of the safety analysis of each manufacturer's patient cohort.

[0102] In the process of generating control mining results based on drug data mining results from the first and second patient cohorts, it is possible to compare whether there are significant differences between the safety analysis results of the first and second patient cohorts. For example, a statistical t-test can be used to compare the differences in adverse reaction rates of target drugs produced by different manufacturers, thereby automatically identifying a list of drugs with significant differences. Furthermore, the resulting drug list and associated clinical data clues can be further reviewed and confirmed by physicians, ultimately generating safety assessment reports for target drugs produced by various manufacturers, such as safety assessment reports for original drugs and generic drugs.

[0103] Based on any of the above embodiments Figure 3 This is a flowchart illustrating the security analysis method provided by the present invention, as shown below. Figure 3 As shown, safety indicators for the target drug can first be generated based on a large model. Both clinical data and safety indicators are then input into the large model. The model extracts the duration of treatment with the target drug and the values ​​of the safety indicators from the clinical data. Subsequently, the large model can be used to determine whether the patient experiences adverse reactions based on the safety indicator values ​​before and after the treatment. These conclusions are then output after multiple validations for use in the statistical analysis of safety indicators.

[0104] Based on any of the above embodiments, in step 130, an economic analysis is performed on the clinical data of the first patient cohort and the second patient cohort, including: Statistical amounts of economic indicators were extracted from the clinical data of the first patient cohort and the second patient cohort, respectively. The economic indicators include at least one of the following: target drug cost, out-of-pocket cost of target drug, total medical expenses, examination fees, testing fees, and treatment fees.

[0105] Specifically, for any patient cohort in the first or second patient cohort, an economic analysis is performed on the clinical data of that patient cohort. This can involve extracting statistical amounts of various economic indicators from the clinical data, thereby achieving a systematic quantification of the direct medical costs related to the target drug treatment. This allows for a comprehensive assessment of the economic impact of the target drug treatment from multiple perspectives, including patients, hospitals, and health insurance payers.

[0106] Economic indicators include at least one of the following: target drug cost, out-of-pocket cost of target drug, total medical expenses, examination fees, laboratory fees, and treatment costs.

[0107] The cost of the target drug refers to the cost of the target drug used by the patient during hospitalization, which can be calculated using the following formula: The out-of-pocket cost of a target drug is the portion of the cost of that drug that the patient is responsible for according to medical insurance reimbursement policies. The out-of-pocket cost of a target drug can be calculated using the following formula: Total medical expenses are the total medical costs incurred by a patient during their hospitalization, and are the most comprehensive indicator for measuring the economic burden of illness. Total medical expenses can be obtained by summarizing all medical expenses on the patient's bill, such as the sum of medication costs, examination fees, laboratory test fees, treatment fees, bed fees, nursing fees, and surgical fees.

[0108] Examination fees refer to the costs incurred for diagnostic procedures using imaging techniques, such as CT (Computed Tomography), MRI (Magnetic Resonance Imaging), and ultrasound. Examination fees can be obtained by summarizing the fees for examination items in the billing system.

[0109] Laboratory testing fees refer to the costs incurred through laboratory testing methods, such as those for complete blood counts, comprehensive biochemical tests, and genetic testing. Laboratory testing fees can be obtained by summarizing the costs of each testing item in the settlement statement.

[0110] Treatment costs refer to the expenses for treatment procedures other than medications, examinations, and tests, such as those incurred for surgery and interventional treatments. Treatment costs can be obtained from the treatment procedure items in the summary settlement list.

[0111] In the process of conducting economic analysis on clinical data, all cost items and amounts can be automatically extracted directly from the cost details and settlement list of clinical data. The extracted cost items and amounts are then mapped and categorized based on the hospital's billing item catalog and internal mapping rules. Specifically, the extracted cost items and amounts are categorized into the aforementioned economic indicators, thereby enabling the statistical calculation of economic indicators for patients.

[0112] Therefore, the statistical amounts of each economic indicator corresponding to the patients in each manufacturer's patient cohort can be included in the economic analysis results of each manufacturer.

[0113] In the process of generating control mining results based on drug data mining results from the first and second patient cohorts, it is possible to compare whether there are significant differences between the economic analysis results of the first and second patient cohorts. For example, in the case where each generic drug manufacturer is designated as the first manufacturer and the original drug manufacturer as the second manufacturer (control), the average values ​​of various economic indicators for all patients in the corresponding patient cohort can be calculated for comparison for each manufacturer.

[0114] For example, in the scenario where any one of the generic drug manufacturers and the original drug manufacturer is designated as the first manufacturer, and all other manufacturers are designated as the second manufacturer (control), for each manufacturer, the costs of various economic indicators for each patient in the corresponding patient cohort can be weighted and summed to obtain a weighted average for comparison. Here, the formula for calculating the weighted average can be: By calculating a weighted average, we can obtain more comparable estimates of economic indicators that balance out confounding factors for comparison.

[0115] When comparing the economic analysis results of the first and second patient cohorts to determine if there are significant differences, for continuous variables such as total cost and drug cost, a weighted t-test can be used to compare the statistical significance of differences between groups. The resulting comparison mining results can include the difference values, confidence intervals, and p-values ​​obtained from comparing economic indicators. Furthermore, economic indicators with significant differences can be highlighted. For example, "The total medical cost of the generic drug group A is significantly lower than that of the original drug group, with a difference of -¥1500, P<0.05." In addition, a list of drug manufacturers with significant economic differences and detailed cost breakdowns can be presented to researchers; for example, the generic drug group saves ¥800 in drug costs but has no significant difference in examination costs. Researchers can then combine this with clinical knowledge to determine the rationality and clinical significance of these economic differences, ultimately forming a comprehensive pharmacoeconomic evaluation conclusion.

[0116] Based on any of the above embodiments Figure 4 This is the second flowchart of the drug data mining method based on a large model provided by the present invention, as shown below. Figure 4 As shown, the drug data mining method based on large models can be applied to the value assessment of target drugs. This method aims to build an end-to-end automated process, using a large language model as the intelligent engine and advanced statistical methods as the correction means, to achieve closed-loop processing from multi-source heterogeneous data input to multi-dimensional decision support report output. The overall process can be divided into the following four main stages: The first stage involves data preparation and standardization.

[0117] The goal of this phase is to transform raw, messy real-world data into a clean, organized, and analytically usable structured data pool. This is mainly reflected in the multi-source data acquisition and integration, which involves the automated collection of clinical data from multiple patient samples from hospital information systems.

[0118] The second phase involves the automated construction and bias control of patient cohorts.

[0119] The goal of this phase is to accurately and fairly select comparable patient cohorts from all sample patients based on analytical needs.

[0120] During this process, clinical data from sample patients can be initially screened using data inclusion criteria. These criteria may include: 1. Patients using the target drug; 2. Excluding patients taking both generic and original drugs concurrently during the treatment cycle; 3. Excluding patients exceeding the daily dosage limit. Sample patients meeting these inclusion criteria can be automatically selected using SQL (Structured Query Language) or a rule engine.

[0121] Based on this, baseline characteristics of each sample patient can be extracted, and propensity score of each sample patient can be calculated based on the baseline characteristics of each sample patient, thereby generating patient cohorts for each manufacturer to achieve sample balance.

[0122] Phase 3: Perform multi-dimensional parallel analysis.

[0123] The goal of this phase is to utilize large language models and statistical analysis engines to conduct a comprehensive and automated evaluation of the patient cohort generated in the previous phase. Specifically, this will enable parallel implementation of medication behavior analysis, efficacy analysis, safety analysis, and cost-effectiveness analysis.

[0124] Phase 4: Perform AI-based result synthesis.

[0125] The goal of this phase is to transform the analysis results into insights and reports that can directly support decision-making.

[0126] This process allows for the analysis of differences between various manufacturers, enabling the creation of a drug value assessment report. This report can be a structured summary of results automatically generated using AI technology, highlighting key findings (e.g., "Manufacturer A's efficacy is no different from the original drug, but its daily cost is 15% lower"). Alternatively, the report can be a draft of a multi-dimensional assessment report generated using AI technology, incorporating charts, statistical results, and data sources.

[0127] Researchers can review these reports, paying particular attention to highlighted discrepancies, and combine this with their clinical expertise to make a final judgment and confirmation. This process frees human experts from the tedious task of data analysis, allowing them to focus on the most valuable decision-making processes.

[0128] The resulting report can provide efficient real-world evidence for hospital drug selection, medical insurance catalog adjustments, and clinical drug use guideline development, thus completing a full closed loop from data to evidence to decision-making.

[0129] In the method provided in this embodiment of the invention, the reliability of the analytical basis is ensured through the first stage; the scientific nature of the research results is guaranteed from the source by embedding advanced statistical methods through the second stage; a comprehensive, efficient, and adaptive evaluation of medication behavior, efficacy, safety, and economy is achieved through the third stage; and finally, the data evidence is transformed into decision insights that can be directly used by doctors through the fourth stage.

[0130] The large-model-based drug data mining apparatus provided by the present invention will be described below. The large-model-based drug data mining apparatus described below can be referred to in correspondence with the large-model-based drug data mining method described above.

[0131] Figure 5 This is a schematic diagram of the structure of the drug data mining device based on a large model provided by the present invention, as shown below. Figure 5 As shown, the device includes: The data acquisition unit 510 is used to acquire clinical data of sample patients, including a first patient receiving treatment with a target drug manufactured by a first manufacturer and a second patient receiving treatment with a target drug manufactured by a second manufacturer. The sample balancing unit 520 is used to predict the propensity score of the sample patients based on the baseline characteristics of the sample patients, and generate a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer based on the propensity score, wherein the propensity score is a predicted probability that the sample patients are inclined to receive treatment with the target drug produced by the first manufacturer. The data mining unit 530 is used to perform drug data mining on the clinical data of the first patient cohort and the second patient cohort based on a large language model, and obtain the drug data mining results of the first patient cohort and the second patient cohort. The data comparison unit 540 is used to generate comparison mining results of the target drug produced by the first manufacturer based on the drug data mining results of the first patient cohort and the second patient cohort.

[0132] In the apparatus provided in this invention, baseline features are extracted from sample patients, and propensity score is assigned to them based on these features. The resulting propensity score is then applied to construct patient cohorts for comparison with target drugs from different manufacturers, resulting in balanced and comparable patient cohorts. Drug data mining is performed on each patient cohort, and the mining results are compared. This approach controls confounding bias at its source, ensuring the reliability and scientific rigor of the comparative analysis. Furthermore, the entire process from data acquisition, cohort construction, sample balancing, data mining to result generation is automated, effectively improving the efficiency of drug data mining and significantly enhancing the timeliness of real-world research on target drugs.

[0133] Based on any of the above embodiments, the sample equalization unit is specifically used for: For each first patient, a second patient whose propensity score is close to that of the first patient is selected from all second patients and used as a control patient for the first patient; The first patient queue is generated based on the first patient, and the second patient queue is generated based on the control patients of the first patient.

[0134] Based on any of the above embodiments, the sample equalization unit is specifically used for: The treatment weight for the first patient is determined based on the propensity score of the first patient, and the control weight for the second patient is determined based on the propensity score of the second patient. The lower the propensity score, the greater the treatment weight, and the higher the propensity score, the greater the control weight. The first patient queue is generated based on the first patient and the processing weight of the first patient, and the second patient queue is generated based on the second patient and the control weight of the second patient.

[0135] Based on any of the above embodiments, the data mining unit is specifically used for: Perform at least two of the following analyses on the clinical data of the first patient cohort and the second patient cohort: medication behavior analysis, efficacy analysis, safety analysis, and cost-effectiveness analysis.

[0136] Based on any of the above embodiments, the data mining unit is specifically used for: Perform at least one of the following on the clinical data of the first patient cohort and the second patient cohort: adherence analysis, medication switching analysis, concomitant medication analysis, and medication change treatment analysis. The compliance analysis includes statistics on the percentage of time spent using medication; the medication switching analysis includes statistics on the percentage of patients switching and / or the frequency of medication switching; the concomitant medication analysis includes statistics on the incidence of concomitant medication and / or the statistics on medication combinations; and the medication change treatment analysis includes statistics on the medication change rate and / or the identification and ranking of medication change pathways.

[0137] Based on any of the above embodiments, the data mining unit is specifically used for: Based on a large language model, the efficacy indicators of the target drug and the changing trends of the efficacy indicators are mined from medical knowledge information; Based on the large language model, the values ​​of the efficacy indicators are extracted from the clinical data of the first patient cohort and the second patient cohort, respectively, and the efficacy rates of the target drugs produced by the first manufacturer and the second manufacturer are determined by the changing trends of the values ​​and the changing trends of the efficacy indicators.

[0138] Based on any of the above embodiments, the data mining unit is specifically used for: Based on a large language model, safety indicators of the target drug and the changing trends of these safety indicators when adverse reactions occur are mined from medical knowledge information. Based on the large language model, the values ​​of the safety indicators are extracted from the clinical data of the first patient cohort and the second patient cohort, respectively. The adverse reaction rates of the target drugs produced by the first manufacturer and the second manufacturer are determined by the changing trends of the values ​​and the changing trends when adverse reactions occur.

[0139] Based on any of the above embodiments, the data mining unit is specifically used for: Statistical amounts of economic indicators were extracted from the clinical data of the first patient cohort and the second patient cohort, respectively. The economic indicators include at least one of the following: target drug cost, out-of-pocket cost of target drug, total medical expenses, examination fees, testing fees, and treatment fees.

[0140] Figure 6An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a drug data mining method based on a large model, the method including: Acquire clinical data from sample patients, including a first patient receiving treatment with the target drug manufactured by a first manufacturer and a second patient receiving treatment with the target drug manufactured by a second manufacturer; Based on the baseline characteristics of the sample patients, a propensity score is predicted for the sample patients. Based on the propensity score, a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer are generated. The propensity score is a predicted probability that the sample patients are inclined to receive treatment with the target drug produced by the first manufacturer. Based on a large language model, drug data mining was performed on the clinical data of the first patient cohort and the second patient cohort respectively to obtain the drug data mining results of the first patient cohort and the second patient cohort. Based on the drug data mining results of the first patient cohort and the second patient cohort, control mining results of the target drug produced by the first manufacturer are generated.

[0141] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0142] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the drug data mining method based on a large model provided by the above methods, the method comprising: Acquire clinical data from sample patients, including a first patient receiving treatment with the target drug manufactured by a first manufacturer and a second patient receiving treatment with the target drug manufactured by a second manufacturer; Based on the baseline characteristics of the sample patients, a propensity score is predicted for the sample patients. Based on the propensity score, a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer are generated. The propensity score is a predicted probability that the sample patients are inclined to receive treatment with the target drug produced by the first manufacturer. Based on a large language model, drug data mining was performed on the clinical data of the first patient cohort and the second patient cohort respectively to obtain the drug data mining results of the first patient cohort and the second patient cohort. Based on the drug data mining results of the first patient cohort and the second patient cohort, control mining results of the target drug produced by the first manufacturer are generated.

[0143] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the drug data mining method based on a large model provided by the methods described above, the method comprising: Acquire clinical data from sample patients, including a first patient receiving treatment with the target drug manufactured by a first manufacturer and a second patient receiving treatment with the target drug manufactured by a second manufacturer; Based on the baseline characteristics of the sample patients, a propensity score is predicted for the sample patients. Based on the propensity score, a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer are generated. The propensity score is a predicted probability that the sample patients are inclined to receive treatment with the target drug produced by the first manufacturer. Based on a large language model, drug data mining was performed on the clinical data of the first patient cohort and the second patient cohort respectively to obtain the drug data mining results of the first patient cohort and the second patient cohort. Based on the drug data mining results of the first patient cohort and the second patient cohort, control mining results of the target drug produced by the first manufacturer are generated.

[0144] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0145] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A drug data mining method based on a large model, characterized in that, include: Acquire clinical data from sample patients, including a first patient receiving treatment with the target drug manufactured by a first manufacturer and a second patient receiving treatment with the target drug manufactured by a second manufacturer; Based on the baseline characteristics of the sample patients, a propensity score is predicted for the sample patients. Based on the propensity score, a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer are generated. The propensity score is a predicted probability that the sample patients are inclined to receive treatment with the target drug produced by the first manufacturer. Based on a large language model, drug data mining was performed on the clinical data of the first patient cohort and the second patient cohort respectively to obtain the drug data mining results of the first patient cohort and the second patient cohort. Based on the drug data mining results of the first patient cohort and the second patient cohort, control mining results of the target drug produced by the first manufacturer are generated.

2. The drug data mining method based on a large model according to claim 1, characterized in that, The step of generating a first patient cohort for the first manufacturer and a second patient cohort for the second manufacturer based on the propensity score includes: For each first patient, a second patient whose propensity score is close to that of the first patient is selected from all second patients and used as a control patient for the first patient; The first patient queue is generated based on the first patient, and the second patient queue is generated based on the control patients of the first patient.

3. The drug data mining method based on a large model according to claim 1, characterized in that, The step of generating a first patient cohort for the first manufacturer and a second patient cohort for the second manufacturer based on the propensity score includes: The treatment weight for the first patient is determined based on the propensity score of the first patient, and the control weight for the second patient is determined based on the propensity score of the second patient. The lower the propensity score, the greater the treatment weight, and the higher the propensity score, the greater the control weight. The first patient queue is generated based on the first patient and the processing weight of the first patient, and the second patient queue is generated based on the second patient and the control weight of the second patient.

4. The drug data mining method based on a large model according to any one of claims 1 to 3, characterized in that, The step of performing drug data mining on the clinical data of the first patient cohort and the second patient cohort respectively includes: Perform at least two of the following analyses on the clinical data of the first patient cohort and the second patient cohort: medication behavior analysis, efficacy analysis, safety analysis, and cost-effectiveness analysis.

5. The drug data mining method based on a large model according to claim 4, characterized in that, Medication behavior analysis was performed on the clinical data of the first patient cohort and the second patient cohort, including: Perform at least one of the following on the clinical data of the first patient cohort and the second patient cohort: adherence analysis, medication switching analysis, concomitant medication analysis, and medication change treatment analysis. The compliance analysis includes statistics on the percentage of time spent using medication; the medication switching analysis includes statistics on the percentage of patients switching and / or the frequency of medication switching; the concomitant medication analysis includes statistics on the incidence of concomitant medication and / or the statistics on medication combinations; and the medication change treatment analysis includes statistics on the medication change rate and / or the identification and ranking of medication change pathways.

6. The drug data mining method based on a large model according to claim 4, characterized in that, Validity analysis was performed on the clinical data of the first patient cohort and the second patient cohort, including: Based on a large language model, the efficacy indicators of the target drug and the changing trends of the efficacy indicators are mined from medical knowledge information; Based on the large language model, the values ​​of the efficacy indicators are extracted from the clinical data of the first patient cohort and the second patient cohort, respectively, and the efficacy rates of the target drugs produced by the first manufacturer and the second manufacturer are determined by the changing trends of the values ​​and the changing trends of the efficacy indicators.

7. The drug data mining method based on a large model according to claim 4, characterized in that, Safety analysis was performed on the clinical data of the first patient cohort and the second patient cohort, including: Based on a large language model, safety indicators of the target drug and the changing trends of these safety indicators when adverse reactions occur are mined from medical knowledge information. Based on the large language model, the values ​​of the safety indicators are extracted from the clinical data of the first patient cohort and the second patient cohort, respectively. The adverse reaction rates of the target drugs produced by the first manufacturer and the second manufacturer are determined by the changing trends of the values ​​and the changing trends when adverse reactions occur.

8. The drug data mining method based on a large model according to claim 4, characterized in that, Economic analysis was performed on the clinical data of the first patient cohort and the second patient cohort, including: Statistical amounts of economic indicators were extracted from the clinical data of the first patient cohort and the second patient cohort, respectively. The economic indicators include at least one of the following: target drug cost, out-of-pocket cost of target drug, total medical expenses, examination fees, testing fees, and treatment fees.

9. A drug data mining device based on a large model, characterized in that, include: The data acquisition unit is used to acquire clinical data of sample patients, including a first patient receiving treatment with a target drug manufactured by a first manufacturer and a second patient receiving treatment with a target drug manufactured by a second manufacturer. A sample balancing unit is used to predict the propensity score of the sample patients based on the baseline characteristics of the sample patients, and to generate a first patient cohort of the first manufacturer and a second patient cohort of the second manufacturer based on the propensity score, wherein the propensity score is a predicted probability that the sample patients are inclined to receive treatment with the target drug produced by the first manufacturer. The data mining unit is used to perform drug data mining on the clinical data of the first patient cohort and the second patient cohort based on a large language model, and to obtain the drug data mining results of the first patient cohort and the second patient cohort. The data comparison unit is used to generate comparison mining results of the target drug produced by the first manufacturer based on the drug data mining results of the first patient cohort and the second patient cohort.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the drug data mining method based on a large model as described in any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the drug data mining method based on a large model as described in any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the drug data mining method based on a large model as described in any one of claims 1 to 8.