An intelligent drug clinical trial scheme generation method and system
By standardizing multi-source data and grading evidence quality in the drug clinical trial protocol generation tool, quantifying molecular prediction bias by combining measured data, and using multi-agent collaborative reasoning to generate Phase I clinical trial protocols, the problems of opaque evidence quality and insufficient credibility in existing technologies are solved, and the transparency and reliability of the protocols are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
- Filing Date
- 2026-04-23
- Publication Date
- 2026-07-24
AI Technical Summary
Existing drug clinical trial protocol generation tools lack credible quantitative assessment and literature evidence quality assessment, resulting in opaque decision-making basis, especially in high-risk, high-uncertainty first-in-human clinical trial design scenarios, where decision reliability is insufficient.
After receiving multi-source data of candidate molecules and standardizing them, the quality of literature evidence is graded by combining a research type classification strategy and a time decay weighting mechanism. Local bias of molecular prediction data is quantified by using measured data of reference drugs with the same target. A phase I clinical trial protocol is generated by a multi-agent collaborative reasoning mechanism, and a multi-dimensional weighted scoring mechanism is introduced for comprehensive feasibility scoring.
It achieves transparency and traceability of clinical trial protocols, improves the reliability and verifiability of decisions, ensures that the quality of evidence and the credibility of predictive data are explicitly conveyed throughout the process, and the generated protocols carry clear confidence level descriptions.
Smart Images

Figure CN122091265B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of healthcare informatics technology, and in particular to an intelligent method and system for generating drug clinical trial protocols. Background Technology
[0002] The design of clinical trial protocols is a crucial step in new drug development, requiring comprehensive consideration of multi-dimensional expertise, including target biology, molecular safety assessment, clinical data from similar drugs, and regulatory guidelines. Currently, the main tools in the field of clinical decision support for drug development exhibit significant functional limitations. The first category is molecular property prediction tools. While these tools can output predicted values for the physicochemical properties and safety of molecules, their function is limited to numerical output and cannot further translate the predictions into specific clinical guidance. Researchers still need to rely on their own experience to interpret the values and then manually formulate protocols by combining literature, clinical trial databases, and regulatory guidelines. This process is time-consuming, highly dependent on individual knowledge, and inherently subjective. The second category is literature retrieval and knowledge management tools. Although they provide a vast amount of scientific literature and clinical data, they lack the ability to integrate with molecular prediction data and, more importantly, the ability to automatically translate retrieved knowledge into specific clinical protocol designs.
[0003] Existing research has significant gaps in key technologies, hindering the output quality of clinical development decision support systems. Firstly, there is a lack of quantitative assessment mechanisms for the quality of literature evidence. Current protocols, when analyzing target drugability, typically input all literature retrieved from databases into a large language model for aggregation. A multicenter phase III clinical trial and a single-cell line in vitro experiment are treated identically within the system. The fundamental principles of evidence-based medicine indicate that there are fundamental differences in the reliability of clinical translation between different types of research evidence. Aggregating mixed-quality literature indiscriminately leads to insufficient traceability and reliability of drugability assessment conclusions. Secondly, there is a systematic neglect of the reliability of molecular prediction values. Current protocols directly input the output values of third-party prediction tools as definitive facts into the clinical interpretation stage. However, numerous published studies have shown that the prediction errors of commonly used predictive indicators are highly uneven and non-negligible. When the system provides a definitive clinical interpretation of the predicted values, it actually ignores the potentially wide range of possible true values, corresponding to entirely different dosing strategies.
[0004] Chinese patent application CN113096795A discloses a multi-source data-assisted clinical decision support system and method. This system compares standard clinical pathway data clusters with real-world clinical pathway data clusters by setting up medical / pharmaceutical data storage units, clinical pathway / evidence storage units, information input units, comparison processor units, and treatment result output units. This provides medical professionals with treatment plan suggestions and corresponding clinical and real-world research evidence support. However, this system primarily focuses on matching and displaying existing clinical pathways and real-world data clusters. The "clinical research data" or "real-world research data" upon which its evidence support module relies is treated as a holistic, ungraded evidence package. This means that when providing evidence, the system does not quantitatively assess or grade the research type (e.g., randomized controlled trials, in vitro studies, or computational predictions), publication date, or direct relevance to the current decision problem of the underlying literature or data supporting the evidence. Therefore, the quality of evidence on which the diagnostic and treatment recommendations are based is opaque, making it difficult for doctors to accurately weigh the reliability of the recommendations based on their own experience. This lack of information on the quality of evidence may lead to insufficient reliability of the decision-making basis, especially in high-risk, high-uncertainty first-in-human (Phase I) clinical trial design scenarios. Summary of the Invention
[0005] In view of this, the present invention provides an intelligent method for generating drug clinical trial protocols, in order to solve the technical problems in the prior art such as the lack of credible quantitative assessment of molecular prediction data, insufficient systematic assessment of the quality of literature evidence, and insufficient automation in generating clinical trial protocols.
[0006] The technical solution of this invention is implemented as follows: On the one hand, the present invention provides an intelligent method for generating drug clinical trial protocols, comprising the following steps: S1. Receive multi-source data of candidate molecules and perform standardization processing to obtain a standardized dataset. The multi-source data includes structural data, target information, indication information, and molecular prediction data generated based on the structural data of the candidate molecules. S2. Based on the target and indication information in the standardized dataset, retrieve relevant literature, use the research type classification strategy to classify the quality of the retrieved literature evidence, and combine the time decay weighting mechanism to calculate the comprehensive evidence quality score of the target-indication pair, and obtain the target evidence quality report. S3. Based on the target information in the standardized dataset, retrieve the measured data of the reference drug for the same target, use the measured data as an error proxy, quantify the local deviation of each molecular prediction data, attach a confidence label to each prediction index, and obtain the composite prediction data structure. S4. Using the credibility labels in the target evidence quality report and composite prediction data structure as explicit constraints, the clinical trials of candidate molecules are parameterized and generated through a multi-agent collaborative reasoning mechanism to obtain a preliminary Phase I clinical trial protocol. S5. A multi-dimensional weighted scoring mechanism is used to conduct a comprehensive feasibility score on the Phase I clinical trial protocol. Based on the comprehensive feasibility score, a graded decision-making suggestion is generated, and the final intelligent drug clinical trial protocol is output.
[0007] Based on the above technical solutions, the preferred method for constructing the standardized dataset in step S1 includes: The format of the SMILES structure string of the candidate molecule is validated and converted, and the published experimental data of the reference drug with the same target are retrieved from public databases based on the target information. The retrieval results are cached and stored. The molecular prediction data output by the third-party molecular property prediction tool submitted by the user is formatted and converted. The molecular prediction data includes solubility, permeability, plasma protein binding rate, liver microsomal clearance rate, cardiotoxicity risk, CYP inhibitory activity and hepatotoxicity risk rating. The cached measured data of reference drugs targeting the same target are merged with the predicted data of candidate molecules to form a standardized dataset.
[0008] Based on the above technical solutions, preferably, step S2 includes the following steps: S21. Retrieve relevant literature based on target and indication information in a standardized dataset to form a literature collection; assign level labels to each article in the collection according to a research type hierarchical strategy. L1 corresponds to human clinical data, L2 corresponds to in vivo animal data, L3 corresponds to in vitro mechanism data, and L4 corresponds to computational prediction data. S22. Statistically count the number of successful cases of drugs progressing from the preclinical to the clinical stage reported in literature at all levels, and calculate the success rate. S23. Calculate the weighted contribution score for each document in the document collection, and sum the weighted contribution scores of all documents to obtain the overall evidence quality score; at the same time, calculate the proportion of the sum of the weighted contribution scores of each level of documents to the overall evidence quality score, and construct the evidence structure distribution vector. S24. The credibility marker of target evidence is determined by the combined value of the weighted contribution score of L1-level literature to the total score of comprehensive evidence quality and the clinical advancement rate. S25. Integrate the overall evidence quality score, evidence structure distribution vector, clinical progress rate, target evidence credibility markers, and L1 literature list to form a target evidence quality report.
[0009] Based on the above technical solutions, preferably, each document mentioned in step S23... Weighted contribution score The calculation formula is: ; in, Document level The corresponding basic weights are assigned according to the evidence pyramid principle of evidence-based medicine; This represents the number of years since the publication of the document. This is the time decay coefficient; For the first The semantic relevance score of each document to the current target-indication pair is calculated as follows: a target keyword set and an indication keyword set are constructed separately, the frequency of each keyword set in the document title and body text is counted, the results are weighted and merged according to the target dimension weight and the indication dimension weight, and the merged relevance score of all documents is normalized to obtain the semantic relevance score of the document.
[0010] Based on the above technical solutions, preferably, step S3 specifically includes: S31. For candidate molecules and all reference drugs with the same target cached in the standardized dataset, calculate Morgan molecular fingerprints based on SMILES structure, and calculate the structural similarity between candidate molecules and each reference drug according to Tanimoto coefficient, forming a local chemical neighborhood set of candidate molecules. S32. For each molecule prediction index, for reference drugs with known measured values in the local chemical neighborhood set, calculate the prediction deviation multiple of each reference drug on the index. The prediction deviation multiple is the ratio of the predicted value to the known measured value. Using the structural similarity between the candidate molecule and each reference drug as the weight, calculate the weighted deviation estimate and weighted deviation standard deviation of the index in the local neighborhood of the candidate molecule. S33. Label each indicator with a confidence level based on the weighted bias estimate and the weighted bias standard deviation; S34. For each indicator, perform data accessibility classification. Based on the classification results, perform corresponding processing on the credibility label. Then, attach the credibility label, weighted deviation estimate, weighted deviation standard deviation, data accessibility level, and predicted value estimation interval to the original predicted value one by one to form a composite predicted data structure.
[0011] Based on the above technical solutions, preferably, the data accessibility classification in step S34 is performed according to the following rules: Grade A: For a certain prediction indicator, the number of valid reference drugs with published measured values in the local chemical neighborhood set is no less than 3, and the structural similarity of at least 1 reference drug is not lower than the structural similarity threshold. The confidence label is forcibly assigned to Grade A, indicating that the data is sufficient. Grade B: The number of effective reference drugs is 1 to 2, or the number is no less than 3 but the highest structural similarity in the local neighborhood is lower than the sufficient structural similarity threshold. The confidence label is forcibly assigned to Grade B, indicating that the data is limited. Grade C: The number of valid reference drugs is 0, or there are no reference drugs in the local neighborhood that meet the minimum threshold of structural similarity. In this case, bias estimation is not performed, and the confidence label is forcibly assigned to Grade C, indicating insufficient data.
[0012] Based on the above technical solutions, preferably, step S4 specifically includes: S41. The administration route of the candidate molecule is determined based on the combination of permeability prediction value, confidence label and liver extraction rate in the composite prediction data structure. The liver extraction rate is calculated from liver microsomal clearance rate based on the Well-Stirred model. S42. A dual-pathway approach is used to calculate the starting dose of the candidate molecule: Pathway A calculates the system clearance rate based on pharmacokinetic parameters using the Well-Stirred model, and then estimates the starting dose of Pathway A at the target steady-state plasma concentration; Pathway B extracts the approved clinical starting dose of the reference drug from the target evidence quality report, and calculates the starting dose of Pathway B by combining the activity differences and structural similarities between the candidate molecule and the reference drug; the dose calculation path is determined according to the target evidence credibility marker and credibility label to obtain the final recommended starting dose; S43. Based on the route of administration and the initial dose as basic parameters, determine the increment fold level and generate a dose escalation plan based on the risk level and confidence label of the safety indicators in the composite prediction data structure. S44. Generate corresponding security monitoring schemes based on the risk level and credibility labels of each security and PK index in the composite prediction data structure. S45. Validate the generated route of administration, final recommended starting dose, dose escalation scheme, and safety monitoring scheme against clinical guidelines, and output the validated parameter set as the preliminary Phase I clinical trial protocol.
[0013] Based on the above technical solutions, the preferred calculation formula for the dual-path initial dose in step S42 is as follows: Pathway A, based on the Well-Stirred model, calculates the hepatic system clearance rate using predicted values of hepatic microsomal clearance and plasma protein binding, and then extrapolates the starting dose at the target steady-state plasma concentration. ; in, To recommend a predicted steady-state peak concentration not exceeding a certain threshold, CL was calculated from the predicted liver microsomal clearance value using the Well-Stirred model. The dosing interval is denoted by F, the predicted oral bioavailability is denoted by SF, and the safety factor is SF. Pathway B extracts the approved clinical starting dose of the reference drug from the target evidence quality report and calculates it using the following formula: ; in, The reference drug's initial dose during Phase I; For reference drug in vitro activity values; The in vitro activity value of the candidate molecule; The Tanimoto structural similarity coefficient between the candidate molecule and the reference drug, with a value range of... ; The activity correction index; The similarity index is conservatively reduced.
[0014] Based on the above technical solutions, preferably, step S5 specifically includes: S51. A multi-dimensional weighted scoring mechanism was used to comprehensively score the feasibility of the Phase I clinical trial protocol. The scoring dimensions included evidence of target druggability, molecular safety, pharmacokinetic characteristics and differential potential. The credibility label in the target evidence quality report and composite prediction data structure was introduced as a moderating factor for each dimension score. S52. Based on the comprehensive feasibility score and credibility label, generate hierarchical decision-making suggestions according to preset thresholds; S53. Integrate the comprehensive feasibility score, the sub-scores of each dimension, and the hierarchical decision-making suggestions to output the final intelligent drug clinical trial plan.
[0015] The present invention also provides an intelligent drug clinical trial protocol generation system to implement the above-described method, comprising: The data acquisition module is used to acquire structural data, target information, indication information, and molecular prediction data of candidate molecules. The data processing module is used to perform format conversion, field mapping, unit unification and standardization on the acquired data to obtain a standardized dataset; The knowledge retrieval module is used to retrieve related literature based on target information and indication information in a standardized dataset, and to retrieve measured data of reference drugs with the same target based on the target information. The data evaluation module is used to grade the quality of the retrieved related literature and calculate the comprehensive evidence quality score of the target-indication pair to generate a target evidence quality report; and to perform a deviation quantification evaluation of the molecular prediction data based on the measured data of the reference drug, and to attach a confidence label to each prediction indicator to generate a composite prediction data structure. The protocol generation module is used to generate Phase I clinical trial parameters for candidate molecules based on the target evidence quality report and composite prediction data structure, and obtain a preliminary Phase I clinical trial protocol. The decision evaluation module is used to comprehensively assess the feasibility of preliminary Phase I clinical trial protocols and generate tiered decision recommendations. The data storage module is used to store the acquired data, standardized datasets, search results, evaluation results, generated clinical trial protocols, and hierarchical decision-making recommendations. The report output module is used to integrate the preliminary Phase I clinical trial protocol and tiered decision-making recommendations to output an intelligent drug clinical trial protocol.
[0016] The present invention has the following advantages over the prior art: (1) This invention constructs an explicit information quality transmission mechanism that runs through the entire process. Under the framework of multi-agent collaborative reasoning, the results of literature evidence quality assessment and the quantification results of molecular prediction data credibility are taken as unavoidable constraints and transmitted node by node to the determination of the route of administration, the calculation of the initial dose, the dose escalation strategy and the comprehensive scoring system. Compared with the existing clinical decision support system that directly inputs the original literature and predicted values into the large language model for free interpretation, this invention makes each clinical parameter output have traceable information quality basis. The generated Phase I clinical trial protocol carries a clear confidence level description, which helps to improve the transparency and verifiability of clinical development decisions.
[0017] (2) This invention addresses the problem of mixed-quality literature being treated with equal weights in the quality assessment of druggability evidence. By combining a research type-level labeling system (levels L1 to L4) with semantic relevance scoring and time decay weighting, the retrieved literature set is quantified into an evidence distribution vector with a clear quality structure and a target evidence credibility marker (TER). The differences in evidence quality structure are structurally transmitted through the mandatory constraint of the downstream safety factor (SF) on the TER marker, so that the quality differences of literature evidence generate quantifiable conservative gaps in the calculation of clinical parameters, thus solving the technical problem of implicit loss of literature evidence quality information during the process of transmission.
[0018] (3) This invention addresses the problem of the inability to systematically assess the reliability of black-box third-party molecular property prediction tools. It utilizes publicly available measured data of reference drugs targeting the same target in drug development scenarios and employs an error proxy strategy weighted by local chemical neighborhood similarity. Without accessing the tool's training set, it estimates the deviation multiple and standard deviation of each prediction indicator, generating a quantitatively based three-level credibility label. Simultaneously, it introduces a three-level classification mechanism of data accessibility (A / B / C). When measured reference data within the chemical neighborhood is insufficient, explicit downgrading labels replace silent failure handling, ensuring the integrity of the quality transfer chain in the case of data sparsity and filling the technical gap in the quantification of the credibility of black-box prediction tool outputs. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of the intelligent drug clinical trial protocol generation method of the present invention; Figure 2 This is a flowchart of the document evidence quality assessment process for this invention; Figure 3 This is a flowchart illustrating the generation process of the composite predictive data structure of the present invention. Figure 4 A logic diagram for generating the dual-path initial dose calculation and dose increment scheme of the present invention is provided. Figure 5 The following is a framework diagram of the intelligent drug clinical trial protocol generation system of the present invention. Detailed Implementation
[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0022] like Figure 1 As shown, the present invention provides an intelligent method for generating drug clinical trial protocols, comprising the following steps: S1. Receive multi-source data of candidate molecules and perform standardization processing to obtain a standardized dataset. The multi-source data includes structural data, target information, indication information, and molecular prediction data generated based on the structural data of the candidate molecules. S2. Based on the target and indication information in the standardized dataset, retrieve relevant literature, use the research type classification strategy to classify the quality of the retrieved literature evidence, and combine the time decay weighting mechanism to calculate the comprehensive evidence quality score of the target-indication pair, and obtain the target evidence quality report. S3. Based on the target information in the standardized dataset, retrieve the measured data of the reference drug for the same target, use the measured data as an error proxy, quantify the local deviation of each molecular prediction data, attach a confidence label to each prediction index, and obtain the composite prediction data structure. S4. Using the credibility labels in the target evidence quality report and composite prediction data structure as explicit constraints, the clinical trials of candidate molecules are parameterized and generated through a multi-agent collaborative reasoning mechanism to obtain a preliminary Phase I clinical trial protocol. S5. A multi-dimensional weighted scoring mechanism is used to conduct a comprehensive feasibility score on the Phase I clinical trial protocol. Based on the comprehensive feasibility score, a graded decision-making suggestion is generated, and the final intelligent drug clinical trial protocol is output.
[0023] In one embodiment of the present invention, step S1 includes: performing format verification and conversion on the SMILES structure string of the candidate molecule, and retrieving published experimental data of the reference drug with the same target from a public database based on the target information, and caching and storing the search results; performing format verification and conversion on the molecular prediction data output by the third-party molecular property prediction tool submitted by the user, the molecular prediction data including solubility, permeability, plasma protein binding rate, liver microsomal clearance rate, cardiotoxicity risk, CYP inhibitory activity and hepatotoxicity risk rating; merging the cached experimental data of the reference drug with the same target and the candidate molecule prediction data to form a standardized dataset.
[0024] Specifically, users submit analysis requests via a web interface or API, inputting the SMILES structural string of candidate molecules, target identifier (UniProt ID or target name), indication text, and development stage markers. The system validates the SMILES string for validity, including structural integrity checks and standardized representation conversion. The target identifier, after parsing, is used to retrieve reference drugs targeting the same target from public molecular databases. The search scope is limited to approved drugs or candidate drugs in clinical trials targeting that target. The search results include in vitro activity data (e.g., IC50) and measured in vitro molecular properties of the reference drugs. These results are cached and stored using the target identifier as the key for subsequent steps. Data output from third-party molecular property prediction tools submitted by users is merged with the cached data after format validation, forming a complete standardized dataset containing candidate molecules and reference drugs targeting the same target. This standardized dataset serves as the unified data source for all subsequent analysis steps, ensuring that all agents in steps two through five perform inference based on consistent input.
[0025] In one embodiment of the present invention, such as Figure 2 As shown, step S2 includes the following steps: S21. Retrieve relevant literature based on target and indication information in a standardized dataset to form a literature collection; assign level labels to each article in the collection according to a research type hierarchical strategy. ,in Corresponding human clinical data, Corresponding data from animals, Corresponding in vitro mechanism data, Corresponding calculation and prediction data; Specifically, the literature collection was constructed using target information and indication text as search terms, covering publicly available academic databases and clinical trial registration platforms. The literature collection... Each document in The research type is labeled one by one by a deterministic classifier based on keyword rules. The classifier matches according to the following priority order. Once a match is found, the corresponding level is assigned and no further matching is performed. The primary criteria for Level 1 (human clinical trials) are meeting any of the following: A registration number in the format NCT\d{8} is matched using regular expressions in the title or abstract; "phase I", "phase II", "phase III", "phase 1", "phase 2", or "phase 3" is matched (case-insensitive, must be a complete phrase, variants such as "phase I-like" are not matched); "randomized controlled trial" or "randomised controlled trial" is matched (complete phrase match); or "first-in-human" or "first in human" is matched. If none of the above primary criteria are met, but all three of the following auxiliary criteria are also met, the trial is also classified as Level 1. The condition is that the word "patients" or "subjects" is matched, and the word is not filtered by the excluded phrase list (excluded phrases include background descriptive phrases such as "relevant to patients", "for patients with", and "benefit for patients", and the system performs prefix matching on excluded phrases), and at least one of the following is matched: "dose escalation", "cohort", "adverse event", "AE", "PK study", "pharmacokinetics in humans", "human volunteer", and "healthy subject". Level (In vivo animal studies): Fields such as "in vivo", "mouse model", "rat", "xenograft", "animal study", and "pharmacokinetics in rats" were hit. Level (In vitro cell research): Matches fields such as "in vitro", "cellline", "IC50", "inhibition assay", and "HEK293". Level (Computational Prediction and Review): Matches fields such as "computational," "in silico," "review," "meta-analysis," and "docking," or fails to match any of these levels. When a single record triggers keywords at multiple levels, the highest level is used. This classification process is entirely based on rule matching and does not rely on large language models. Each document receives a unique level label. The results are reproducible.
[0026] S22. For the set of clinical development records of drugs targeting the same target output in step S1, extract the highest clinical development stage field for each record, count the number of compounds that have completed Phase I or higher clinical trials and the total number of compounds in all records targeting the same target, and calculate the clinical progress ratio. ; Specifically, the clinical advancement rate is defined as: ; in, A collection of clinical development records for drugs targeting the same target The number of compounds that have completed Phase I or higher clinical trials; For set The total number of compounds recorded for the same target site; As an independent validation signal for the overall clinical feasibility of the target, it is used in conjunction with the literature scoring system for the final credibility rating.
[0027] S23. Calculate the weighted contribution score for each document in the document collection. The total score for the overall quality of evidence is obtained by summing the weighted contribution scores of all documents. Simultaneously, the proportion of the sum of weighted contribution scores for each level of literature to the total score of overall evidence quality is calculated, and an evidence structure distribution vector is constructed. ; Specifically, for each document in the document collection First, calculate its semantic relevance score with the current target-indication pair. Specifically, this includes: first constructing a set of target keywords. With indication keyword set , Includes the official gene name of the target, UniProt ID, target alias, and protein family name. This includes standard medical terminology for the indications, hyponyms and hypernyms from the MeSH thesaurus, and commonly used clinical synonyms, both derived from the corresponding annotation fields in the standardized dataset of step S1. Then, respectively... and Calculate the weighted keyword frequency: ; ; in, As an indicator function, the keyword frequency weight is multiplied by 3 when it appears in the title. Keywords The word frequency count in the main text; the frequencies from the two dimensions are combined by weight to form a comprehensive raw relevance score: ; in The target dimension weights reflect the stronger discriminative power of target name hits on druggability.
[0028] Finally, for all The original scores of the document were normalized using min-max normalization: ; when When all documents have the same original score (which usually occurs when there are very few search results), let all To avoid the denominator being zero.
[0029] Based on this, for each document Calculate its weighted contribution score : ; in, For the first Weighted contribution score of each document; The weights for each document are assigned according to the evidence pyramid principle of evidence-based medicine. (Human clinical data) (Data from animals) (In vitro cell data) (Computational Prediction and Overview); The normalized semantic relevance score above has a range of values. ; Let be the time decay constant, where The knowledge half-life (in years) is adaptively determined based on the keyword classification results of the target indication text. The system's preset empirical initial value is: for oncology targets... Nervous system target selection Metabolic disease target selection Users can override the above default values in the system configuration interface. When the user does not provide an override value, the system will use the preset value and indicate it in the footer of the output report. For the first The number of years from the publication date of the document to the current retrieval date, taken as a non-negative real number. The overall score for the quality of evidence is: ; in This represents the total number of articles in the collection. Percentage of hierarchical evidence structures for: ; satisfy Thus, the evidence structure distribution vector is constructed. It reflects the quality structure composition of evidence for drugability.
[0030] S24. Based on the combined value of the weighted contribution score of L1-level literature to the total evidence quality score and the clinical advancement ratio, and combined with the target development level (TDL) for cross-validation, output a target evidence quality report and evidence credibility indicators. .
[0031] Specifically, the system introduces target development level (TDL) to cross-validate the evidence structure vector. TDL is divided into four levels: Tclin (approved drug), Tchem (highly active chemical probe), Tbio (biological activity data), and Tdark (virtually no research). The prior expectations for each level are as follows: Tclin corresponds to... Tchem corresponds to Tbio corresponds Tdark corresponds to If the evidence structure vector deviates significantly from the prior expectation of TDL, an "evidence anomaly" label will be triggered in the target evidence quality report and written into the output field as an uncertainty signal.
[0032] The final target evidence quality report includes a total evidence quality score. Evidence structure distribution vector Clinical progress rate TDL verification results and key points Literature list (hit) level and The top three document identifiers). Based on this, according to and Jointly determine the indicators of the credibility of evidence (Target EvidenceReliability): and The time stamp is marked as "highly reliable"; or It is marked as "medium confidence"; and It is marked as "low trust".
[0033] In one embodiment of the present invention, such as Figure 3 As shown, step S3 specifically includes: S31. For candidate molecules and all reference drugs with the same target cached in the standardized dataset, calculate Morgan molecular fingerprints based on the SMILES structure, and calculate the structural similarity between candidate molecules and each reference drug according to the Tanimoto coefficient, forming a local chemical neighborhood set of candidate molecules.
[0034] Specifically, for candidate molecules and all reference drugs targeting the same target cached in the standardized dataset, a 2048-bit Morgan molecular fingerprint (radius 2, equivalent to ECFP4 representation) was calculated based on the SMILES structure. The Morgan fingerprint was calculated using a deterministic algorithm, with each molecule corresponding to a unique binary vector. For candidate molecules With the Reference drug Structural similarity is calculated using the Tanimoto coefficient: ; in Morgan fingerprint vector for candidate molecules. For reference drugs Morgan's fingerprint vector, A higher similarity value indicates a greater similarity between the two molecular structures. The values are ranked by similarity. (default The reference drug constitutes the local chemical neighborhood set of the candidate molecule. and filter Reference drugs.
[0035] S32. For each molecule prediction index, for reference drugs with known measured values in the local chemical neighborhood set, calculate the prediction deviation multiple of each reference drug on the index; using the structural similarity between the candidate molecule and each reference drug as the weight, calculate the weighted deviation estimate and weighted deviation standard deviation of the index in the local neighborhood of the candidate molecule.
[0036] Specifically, for each molecular prediction index For local chemical neighborhood sets For reference drugs with known measured values, their SMILES values are input into the same third-party prediction tool to obtain predicted values. And extract the reference drug from the standardized dataset in terms of indicators Known measured values on (If a certain reference drug corresponds to the index) If no published measured values are available, exclude from the valid neighborhood. Calculate the prediction bias factor for each reference drug on this indicator: ; in For reference drugs In terms of indicators The ratio of predicted values to measured values. This indicates that the prediction and the actual measurement are completely consistent. As weights, calculate the candidate molecules based on the index Weighted bias estimate and weighted deviation standard deviation : ; ; in For candidate molecules in indicators The local weighted bias estimate reflects the direction and magnitude of the systematic bias of the prediction tool in the chemical neighborhood of candidate molecules; The weighted standard deviation reflects the dispersion of the bias of each reference drug within the neighborhood. The smaller the value, the more stable and reliable the bias estimate. According to... The actual estimated interval of the candidate molecule's predicted value can be deduced from the deviation estimation interval: the upper bound of the deviation corresponds to the original predicted value divided by... The lower bound of the deviation corresponds to the original predicted value divided by The two endpoints constitute the actual estimated range of the predicted value of this indicator.
[0037] S33, based on and A three-level confidence label is generated for each prediction indicator. High confidence: Exists within the neighborhood. Reference drugs, and ,and Low reliability: No satisfaction found within the neighborhood. Effective reference drugs, or ,or Medium confidence: An intermediate case that does not meet the above two conditions.
[0038] In the above weighted bias estimate With weighted deviation standard deviation Based on this, the present invention further introduces a neighborhood deviation consistency coefficient. This is used to quantify the structural fit between the deviation factors of various reference drugs within a local chemical neighborhood, as a basis for... An independent supplementary verification signal.
[0039] For local chemical neighborhood set The set of valid reference drugs with known measured values (let the number of elements be ). ), each reference drug Structural similarity Its deviation multiple Treating them as a pair of paired observations, calculate their Spearman rank correlation coefficient. : ; in, For reference drugs of The descending ranking of the value within the neighborhood and The difference in ascending order of values (i.e., the absolute magnitude of the deviation multiple from 1); For effective reference drug quantity; ,in This indicates that the reference drug with higher structural similarity has a bias factor closer to 1 (i.e., more accurate prediction), showing positive consistency. This indicates the presence of a reverse anomaly, meaning that the reference drug with a more similar structure has a larger bias, suggesting the existence of a predictive systematic failure region within this chemical subspace. This indicates that there is no linear relationship between similarity and deviation, and the prior assumptions of the current weighted estimation are invalid.
[0040] based on Define neighborhood deviation consistency coefficient : ; in, The consistency benchmark threshold indicates that the weighted deviation estimation can only obtain high reliability when there is a certain strong positive correlation between similarity and deviation. Steepness parameter, controlled Follow The rate of change; , hour hour , hour .Will It is written into the composite prediction data structure as a consistency check signal.
[0041] In the preferred embodiment described above, the three-level credibility labeling rules in step S33 are adjusted to be executed sequentially according to the following priority order: (1) Forced low-trust degradation (highest priority): If Effective (i.e.) )and (correspond That is, the deviation and similarity within the neighborhood are inversely distributed, regardless of and Regardless of the numerical value, the credibility label of this indicator is forcibly assigned a low credibility value, and the reason field is written as "the deviation multiple within the chemical neighborhood is inversely distributed with the structural similarity, indicating that there is a risk of predictive systematic failure in the current substructure space"; after executing this clause, subsequent judgment rules are skipped.
[0042] (2) Low reliability: Satisfies any of the following conditions: No such condition exists in the neighborhood. Valid reference drugs; or ; .
[0043] (3) High reliability: Simultaneously satisfying all of the following conditions: There exists in the neighborhood Reference drugs; and ;and ;and .like because If the value is null, the condition is considered not met and no high-confidence label is assigned.
[0044] (4) Moderately credible: The intermediate case that does not meet any of the above conditions.
[0045] when hour, Recorded as null, rule (1) does not participate in the judgment, and rule (3) involves The condition item becomes invalid synchronously, and the process reverts to... and The dual-indicator annotation logic is consistent with the basic scheme, and the corresponding field in the composite prediction data structure is noted as "the consistency coefficient was not calculated due to insufficient number of valid reference drugs." This degradation logic is consistent with the data accessibility level B in step S34. ) and Class C ( The processing rules are fully compatible and do not introduce additional risks of erroneous downgrades: In the B-level scenario The result will inevitably be null, and the standard deviation adjustment rule for Grade B will still prevail; in Grade C, a low confidence level will be directly enforced. No calculation required. , The order column used in the calculation is written into the corresponding field of the composite forecast data structure and output in a JSON attachment in the report appendix for independent review by users or regulatory agencies.
[0046] S34. For each indicator, data accessibility is graded according to the following rules, and the confidence labels are processed accordingly based on the grading results. The confidence labels are divided into three levels: A, B, and C. Level A: For a given prediction indicator, the number of valid reference drugs with published measured values in the local chemical neighborhood set is no less than 3, and the structural similarity (in terms of the Tanimoto coefficient) of at least one reference drug is no less than the sufficient structural similarity threshold. The confidence label is forcibly assigned to Level A, indicating sufficient data. Level B: The number of valid reference drugs is 1-2, or no less than 3, but the highest structural similarity in the local neighborhood is lower than the sufficient structural similarity threshold. The confidence label is forcibly assigned to Level B, indicating limited data. Level C: The number of valid reference drugs is 0, or there are no reference drugs in the local neighborhood that meet the minimum structural similarity threshold. In this case, bias estimation is not performed, and the confidence label is forcibly assigned to Level C, indicating insufficient data.
[0047] Specifically, Grade A (sufficient data), number of effective neighboring reference drugs. And at least one reference drug meets the requirements. The original process was followed to output a three-level confidence rating; Level B (limited data), number of effective neighboring reference drugs. ,or However, the highest similarity to Tanimoto in the neighborhood. In this case, the bias estimation is feasible but lacks stability. The confidence level should be reassessed after adjusting the weighted standard deviation using the small sample variance correction formula. ; in The number of effective neighboring reference drugs, in order to To provide a sufficient sample benchmark, The time standard deviation is moderately amplified; in the B-level case, the upper limit of the confidence level labeling results is moderately confident, even if and If the high-confidence numerical condition is met, a high-confidence label is not output, and the data is written to the composite data structure field. and Grade C (Insufficient data), valid neighborhood is empty ( (or no satisfaction is found after filtering) For the reference drug, deviation estimation is not performed in this case. and The record is empty, and the confidence label is forcibly assigned to low confidence, with the reason stated as "There is a lack of published measured data for this indicator in the chemical field, making it impossible to perform local error proxy estimation".
[0048] credibility labels and weighted bias estimates Weighted deviation standard deviation Data accessibility level (A / B / C), number of effective neighboring reference drugs Highest Tanimoto similarity The predicted value and the actual estimated interval are added to the original predicted value one by one to form a composite predicted data structure.
[0049] In black-box scenarios where access to the training set or internal parameters of third-party prediction tools is unavailable, this invention utilizes known measured data from reference drugs targeting the same target as an error reference for the local chemical neighborhood, generating quantitatively based bias estimates and confidence labels for each prediction indicator. A three-tiered (A / B / C) data accessibility mechanism ensures that when measured data in the public database is insufficient, the system handles the situation with explicit downgrade labeling rather than silent failure, maintaining the integrity and traceability of the entire quality transfer chain.
[0050] In one embodiment of the present invention, such as Figure 4 As shown, step S4 specifically includes: S41. Based on the Caco-2 permeability prediction value (characterizing oral absorption capacity), Caco-2 permeability confidence label, and liver extraction rate in the composite prediction data structure. The combination of conditions (characterizing the intensity of the first-pass effect) determines the route of administration for candidate molecules, wherein the liver extraction rate Based on the Well-Stirred model, liver microsomal clearance rate The calculation shows that: ; in Hepatic microsomal clearance rate (unit: mL / min / kg); This represents the fraction of free drug (dimensionless, calculated from plasma protein binding rate). (PPB is the percentage of plasma protein binding). For hepatic blood flow, an empirical value of 80 L / h is used (converted to a uniform unit and then entered). The rule for determining the route of administration is as follows: if the predicted Caco-2 osmotic pressure is higher than... cm / s and the credibility label of this indicator is high or medium. If oral administration is recommended, then if Caco-2 permeability is lower than 100%, oral administration is recommended. cm / s or If the confidence level of the above core indicators is low, the administration route will be indicated in the protocol pending confirmation by actual test data, and both oral and injection alternatives will be provided.
[0051] S42. A dual-pathway approach is used to calculate the starting dose of the candidate molecule: Pathway A calculates the system clearance rate based on pharmacokinetic parameters using the Well-Stirred model, and then estimates the starting dose of Pathway A at the target steady-state plasma concentration; Pathway B extracts the approved clinical starting dose of the reference drug from the target evidence quality report, and calculates the starting dose of Pathway B by combining the activity differences and structural similarities between the candidate molecule and the reference drug; the dose calculation path is determined according to the target evidence credibility marker and credibility label to obtain the final recommended starting dose.
[0052] Specifically, based on the credibility indicators of target evidence ( )and The confidence level label determines the dosage calculation path. Path A, based on pharmacokinetic parameters, calculates and predicts the systemic clearance rate in the human body using the Well-Stirred model. : ; in This represents hepatic blood flow (approximately 80 L / h). Based on this, the dosing interval is determined according to the predicted half-life. Predicting half-life Depend on Calculation, where To predict the distribution volume (unit L, from (and plasma protein binding rate estimation). h time h (QD scheme) Between 6 and 12 hours h (BID scheme) h time (TID protocol). The formula for calculating the starting dose for pathway A is: ; in To recommend a predicted steady-state peak concentration (in ng / mL) not exceeding the recommended level, based on in vitro activity... Conversion and calculation of protein binding rate; To predict human clearance rate (unit: L / h); Dosing interval (in hours); To predict oral bioavailability (dimensionless), Caco-2 permeability and... Estimate; For the safety factor (dimensionless), the default value is 10. When forced to take 100. When the credibility label is medium, the system additionally... replace Recalculate the upper bound of the deviation Output in parallel with the main path results; when When the credibility label is low, force a switch to path B.
[0053] Path B from the target evidence quality report The field extracts the approved clinical starting dose of the L1 reference drug, and the starting dose of the candidate molecule is calculated using the following formula: ; in The reference drug's initial dose during Phase I; For reference drug in vitro activity value (units) ); In vitro activity value of candidate molecules (units) ); is the Tanimoto structural similarity coefficient between the candidate molecule and the reference drug, with a value range of [0,1]. The activity correction index; The similarity conservatism reduction index is α∈(0,1], β∈(0,1], and its specific values are determined by the system configuration. The more conservative (smaller) value is taken from the calculation results of path A and path B as the final recommended starting dose. .
[0054] S43, by route of administration and Based on the fundamental parameters, the escalation levels are determined according to the risk level and confidence labels of safety indicators in the composite prediction data structure, generating a modified Fibonacci sequence dose escalation protocol. The risk level classification rules for each safety indicator are as follows: hERG Predicted value greater than 30 Corresponding to green (low risk), in the range of 1 to 30 The interval corresponding to yellow (medium risk) is less than 1. Red corresponds to high risk; the predicted values for hepatotoxicity and nephrotoxicity are categorized into three levels: green, yellow, and red, based on preset thresholds. The dose escalation level is determined by the highest risk level among all safety indicators: when all are green, a standard escalation sequence is used. When a yellow indicator is present, the upper limit of the dose escalation factor is reduced to no more than 2 times; when a red indicator is present, the dose escalation factor is further reduced to no more than 1.5 times; when the confidence label of the corresponding indicator is low confidence, its risk level is treated as the worst case (i.e., it is raised one level).
[0055] S44. Generate corresponding safety monitoring plans based on the risk level and confidence labels of each safety and pharmacokinetic indicator in the composite prediction data structure. (For hERG) Corresponding to the yellow risk level, a cardiac safety monitoring protocol is generated: 12-lead electrocardiograms are performed before and after each dose, with sampling time points at 0 h before administration, 2 h after administration, and 8 h after administration, recording the QTcF interval; if the QTcF interval is prolonged by more than 30 ms from baseline, the monitoring frequency is increased to cover the entire period before and after each dose; if the confidence label of this indicator is moderately reliable, hERG is included in the worst-case analysis. After adjusting according to the upper limit of the deviation, the risk level is reassessed. If the adjusted level is upgraded to red (hERG)... If the predicted risk level is yellow, the monitoring protocol will be upgraded accordingly. For liver toxicity predictions corresponding to a yellow risk level, a liver function monitoring protocol will be generated: weekly testing of alanine aminotransferase (ALT), aspartate aminotransferase (AST), and total bilirubin; if ALT levels exceed three times the upper limit of normal, medication will be discontinued and the adverse event procedure will be initiated. For low-reliability safety indicators, regardless of their original predicted risk level, the monitoring frequency will be uniformly set to the highest level, and reassessed after actual data acquisition.
[0056] In this invention, The labels and confidence levels of each indicator are incorporated as explicit constraints into the reasoning process of each submodule, rather than directly passing the raw predicted values to the large language model for free interpretation. This ensures that the uncertainty information of the predicted data is structurally transmitted in the parameterization generation of clinical protocols.
[0057] In one embodiment of the present invention, step S5 specifically includes: S51. A multi-dimensional weighted scoring mechanism was used to comprehensively score the feasibility of the Phase I clinical trial protocol. The scoring dimensions included evidence of target druggability, molecular safety, pharmacokinetic characteristics and differential potential. The credibility label in the target evidence quality report and composite prediction data structure was introduced as a moderating factor for each dimension score.
[0058] Target dimension scoring The calculation formula is as follows: ; in The proportion of Level 1 clinical evidence in the total overall evidence quality score; The clinical advancement rate of drugs targeting the same target (derived from step S2); The value range is [0, 10]; for The weight, for The weights satisfy .
[0059] Molecular safety dimension score Calculated based on the combined risk level and credibility label of each security indicator. For each security indicator... (including hERG) (Prediction of hepatotoxicity and nephrotoxicity), defining a numerical risk level mapping: green (low risk) = 3 points, yellow (medium risk) = 1.5 points, red (high risk) = 0 points; confidence adjustment coefficient. Defined as: when highly reliable Medium confidence level When the reliability is low .but: ; in For the complete set of security indicators, Its number of elements, As an indicator Risk level numerical mapping score, The value range is [0, 10].
[0060] Pharmacokinetic dimension score Based on oral bioavailability prediction Predicting half-life Risk levels and confidence labels for pharmacokinetic-related indicators such as solubility (LogS), calculation methods and Similarly, the risk level values of each pharmacokinetic indicator are mapped to scores and corresponding confidence adjustment coefficients. After weighting and averaging according to the above formula, the result is mapped to the interval [0, 10].
[0061] Differentiation potential dimension scoring Quantifying the competitive advantage of candidate molecules relative to reference drugs targeting the same target by comparing the activity of candidate molecules. Median activity compared to reference drugs targeting the same target : ; in In vitro activity value of candidate molecules (units) ); The median (in units) of the activity values of reference drugs targeting the same target. ); As a safety differentiation bonus, 1 to 2 points are awarded when the overall safety score of the candidate molecule is higher than the average safety level of the reference drug for the same target; otherwise, 0 points are awarded. The final constraint is within [0,10].
[0062] Overall Feasibility Score Weighted summation of the four dimensions: ; in As the weight, satisfying , The value range is [0, 10].
[0063] S52, based on Based on the credibility tags and preset thresholds, three levels of hierarchical decision-making suggestions are generated: When there are no low-confidence red indicators, the system outputs "Recommended to proceed" and simultaneously generates a list of priority action items sorted by time urgency. If low-confidence yellow or red indicators exist, output "It is recommended to optimize and then proceed", clearly indicating the indicators that need to be optimized and the possible directions for structural optimization; If a high-confidence red safety indicator exists, the output will be "Not recommended to proceed," along with a detailed explanation of the fatal flaw. Each decision recommendation will be accompanied by a confidence statement, indicating the source of credibility for each dimension of evidence and the verification process following verification with low-confidence indicators. A quantitative description of the potential impact.
[0064] when In the critical range Furthermore, when the confidence label of the core indicator leading to the critical state is medium confidence, the system automatically triggers a feedback iteration: the decision-making agent sends a supplementary analysis request to the prediction data interpretation agent in step S3, specifying that the clinical risk level of the core indicator be recalculated using the upper and lower bounds of the deviation, respectively; step S4 regenerates the corresponding monitoring protocol and dose conservation parameters based on the updated risk level, and then passes them to step S5 for recalculation. Output two decision conclusions: the upper bound and the lower bound. The iteration runs at most 3 rounds. If after 3 rounds... If the system remains in the critical range and the decision level is still unclear, it will mark "Currently available data is insufficient to support a deterministic decision" in the report, terminate the iteration, and output the current optimal analysis result. The feedback iteration mechanism is compatible with the data accessibility grading rule in step S3: when an indicator obtains a moderately reliable label due to limited B-level data and triggers the critical range iteration, this must be stated in the supplementary analysis request. and Step S5 distinguishes between "medium confidence factor bias uncertainty" and "medium confidence factor data limited" in the confidence statement, providing users with accurate improvement directions.
[0065] S53, will Scoring by each dimension , , , Integrating with tiered decision-making recommendations, the system outputs a final intelligent drug clinical trial protocol, and performs structured integration and report rendering to generate a complete clinical development decision support report. This report includes an executive summary, target analysis section, predictive data credibility assessment section, comprehensive feasibility assessment section, decision recommendations section, and appendices. The report is output in HTML and PDF formats. The HTML version allows users to click on each scoring field to view the detailed calculation process and cited data sources. Intermediate calculation values for all formulas are saved in expanded form in a JSON attachment for independent review by users or regulatory agencies.
[0066] In the technical solution of this invention, a credibility adjustment coefficient is introduced into the scoring of each dimension. This approach directly reflects the uncertainty of molecular prediction data in the score values of each dimension, rather than merely stating it qualitatively in the decision text. The feedback iterative mechanism further transforms the uncertainty range of moderately reliable indicators into the quantitative range of the decision range, providing verifiable two-sided analysis for projects in a critical state.
[0067] This invention also provides an intelligent drug clinical trial protocol generation system for implementing the above methods, such as... Figure 5As shown, the system includes a data acquisition module, a data processing module, a knowledge retrieval module, a data evaluation module, a solution generation module, a decision evaluation module, a data storage module, and a report output module. These modules work together to complete the entire process from receiving multi-source data to outputting the final solution.
[0068] The data acquisition module is responsible for receiving candidate molecule SMILES structure strings, target identifiers, indication texts, and molecular prediction data output by third-party prediction tools submitted by users through the web interface or API interface, serving as the unified data entry point for the system.
[0069] The data processing module performs format validation, field mapping, unit unification, and standardization on the above data. It merges the candidate molecule data with the measured data of the reference drug for the same target cached by the knowledge retrieval module and outputs a standardized dataset for use by downstream modules.
[0070] The knowledge retrieval module uses target and indication information from the standardized dataset as search terms to search for relevant literature in public academic literature databases and clinical trial registration platforms, and to search for published in vitro experimental data of reference drugs with the same target in public molecular databases. The search results are cached and provided to the data evaluation module.
[0071] The data evaluation module undertakes two parallel evaluation responsibilities: First, it performs research type classification and weighted scoring on the retrieved related literature, calculates the overall evidence quality score, evidence structure distribution vector, and clinical progress rate, and generates a target evidence quality report; Second, based on the measured data of reference drugs for the same target, it quantifies the local bias of each molecular prediction indicator by weighting it with chemical neighborhood similarity, adds a credibility label and data accessibility level to each indicator, and generates a composite prediction data structure; The above two types of outputs serve as inputs to the scheme generation module.
[0072] The protocol generation module uses the credibility labels in the target evidence quality report and composite prediction data structure as explicit constraints, and determines the administration route, starting dose, dose escalation plan and safety monitoring plan in sequence through multi-agent collaborative reasoning, and outputs the preliminary Phase I clinical trial protocol.
[0073] The decision evaluation module assigns a weighted score to the preliminary proposal based on four dimensions: evidence of target drugability, molecular safety, pharmacokinetic characteristics, and differentiation potential. This results in a comprehensive feasibility score, and a three-tiered decision recommendation is generated based on the score and the credibility distribution of key indicators. When the score is in the critical range, the module automatically triggers a feedback iteration, sending supplementary analysis requests back to the data evaluation module and the proposal generation module. This process is repeated until the results converge or the iteration limit is reached.
[0074] The data storage module runs through the entire process described above, persistently storing the intermediate inputs and outputs of each module—including original submitted data, standardized datasets, literature retrieval cache, deviation estimation results, scoring calculation process, and generated scheme text—to support audit traceability and historical analysis.
[0075] The report output module reads all output fields from the data evaluation module, protocol generation module, and decision evaluation module, merges the fields and renders the format according to the predetermined report template, and generates a complete intelligent drug clinical trial protocol that includes an execution summary, target analysis section, prediction data credibility assessment section, comprehensive feasibility assessment section, decision recommendation section, and appendix. The protocol is output in HTML and PDF formats, and the HTML version supports expanding and tracing the source of each scoring field.
[0076] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for generating intelligent drug clinical trial protocols, characterized in that, Includes the following steps: S1. Receive multi-source data of candidate molecules and perform standardization processing to obtain a standardized dataset. The multi-source data includes structural data, target information, indication information, and molecular prediction data generated based on the structural data of the candidate molecules. S2. Based on the target and indication information in the standardized dataset, retrieve relevant literature, use the research type classification strategy to classify the quality of the retrieved literature evidence, and combine the time decay weighting mechanism to calculate the comprehensive evidence quality score of the target-indication pair, and obtain the target evidence quality report. S3. Based on the target information in the standardized dataset, retrieve the measured data of the reference drug for the same target, use the measured data as an error proxy, quantify the local deviation of each molecular prediction data, attach a confidence label to each prediction index, and obtain the composite prediction data structure. S4. Using the credibility labels in the target evidence quality report and composite prediction data structure as explicit constraints, the clinical trials of candidate molecules are parameterized and generated through a multi-agent collaborative reasoning mechanism to obtain a preliminary Phase I clinical trial protocol. S5. A multi-dimensional weighted scoring mechanism is used to conduct a comprehensive feasibility score on the Phase I clinical trial protocol, and a graded decision-making suggestion is generated based on the comprehensive feasibility score to output the final intelligent drug clinical trial protocol. Step S4 specifically includes: S41. The administration route of the candidate molecule is determined based on the combination of permeability prediction value, confidence label and liver extraction rate in the composite prediction data structure. The liver extraction rate is calculated from liver microsomal clearance rate based on the Well-Stirred model. S42. A dual-pathway approach is used to calculate the starting dose of the candidate molecule: Pathway A calculates the system clearance rate based on pharmacokinetic parameters using the Well-Stirred model, and then estimates the starting dose of Pathway A at the target steady-state plasma concentration; Pathway B extracts the approved clinical starting dose of the reference drug from the target evidence quality report, and calculates the starting dose of Pathway B by combining the activity differences and structural similarities between the candidate molecule and the reference drug; the dose calculation path is determined according to the target evidence credibility marker and credibility label to obtain the final recommended starting dose; S43. Based on the route of administration and the initial dose as basic parameters, determine the increment fold level and generate a dose escalation plan based on the risk level and confidence label of the safety indicators in the composite prediction data structure. S44. Generate corresponding security monitoring schemes based on the risk level and credibility labels of each security and PK index in the composite prediction data structure. S45. Validate the generated route of administration, final recommended starting dose, dose escalation scheme, and safety monitoring scheme against clinical guidelines, and output the validated parameter set as the preliminary Phase I clinical trial protocol.
2. The intelligent drug clinical trial protocol generation method as described in claim 1, characterized in that, The method for constructing the standardized dataset in step S1 includes: The format of the SMILES structure string of the candidate molecule is validated and converted, and the published experimental data of the reference drug with the same target are retrieved from public databases based on the target information. The retrieval results are cached and stored. The molecular prediction data output by the third-party molecular property prediction tool submitted by the user is formatted and converted. The molecular prediction data includes solubility, permeability, plasma protein binding rate, liver microsomal clearance rate, cardiotoxicity risk, CYP inhibitory activity and hepatotoxicity risk rating. The cached measured data of reference drugs targeting the same target are merged with the predicted data of candidate molecules to form a standardized dataset.
3. The intelligent drug clinical trial protocol generation method as described in claim 1, characterized in that, Step S2 includes the following steps: S21. Retrieve relevant literature based on target and indication information in a standardized dataset to form a literature collection; assign level labels to each article in the collection according to a research type hierarchical strategy. L1 corresponds to human clinical data, L2 corresponds to in vivo animal data, L3 corresponds to in vitro mechanism data, and L4 corresponds to computational prediction data. S22. Statistically count the number of successful cases of drugs progressing from the preclinical to the clinical stage reported in literature at all levels, and calculate the success rate. S23. Calculate the weighted contribution score for each document in the document collection, and sum the weighted contribution scores of all documents to obtain the overall evidence quality score; at the same time, calculate the proportion of the sum of the weighted contribution scores of each level of documents to the overall evidence quality score, and construct the evidence structure distribution vector. S24. The credibility marker of target evidence is determined by the combined value of the weighted contribution score of L1-level literature to the total score of comprehensive evidence quality and the clinical advancement rate. S25. Integrate the overall evidence quality score, evidence structure distribution vector, clinical progress rate, target evidence credibility markers, and L1 literature list to form a target evidence quality report.
4. The intelligent drug clinical trial protocol generation method as described in claim 3, characterized in that, Each document mentioned in step S23 Weighted contribution score The calculation formula is: ; in, Document level The corresponding basic weights are assigned according to the evidence pyramid principle of evidence-based medicine; This represents the number of years since the publication of the document to the current date. This is the time decay coefficient; For the first The semantic relevance score of each document to the current target-indication pair is calculated as follows: a target keyword set and an indication keyword set are constructed separately, the frequency of each keyword set in the document title and body text is counted, the results are weighted and merged according to the target dimension weight and the indication dimension weight, and the merged relevance score of all documents is normalized to obtain the semantic relevance score of the document.
5. The intelligent drug clinical trial protocol generation method as described in claim 1, characterized in that, Step S3 specifically includes S31. For candidate molecules and all reference drugs with the same target cached in the standardized dataset, calculate Morgan molecular fingerprints based on SMILES structure, and calculate the structural similarity between candidate molecules and each reference drug according to Tanimoto coefficient, forming a local chemical neighborhood set of candidate molecules. S32. For each molecule prediction index, for reference drugs with known measured values in the local chemical neighborhood set, calculate the prediction deviation multiple of each reference drug on the index. The prediction deviation multiple is the ratio of the predicted value to the known measured value. Using the structural similarity between the candidate molecule and each reference drug as the weight, calculate the weighted deviation estimate and weighted deviation standard deviation of the index in the local neighborhood of the candidate molecule. S33. Label each indicator with a confidence level based on the weighted bias estimate and the weighted bias standard deviation; S34. For each indicator, perform data accessibility classification. Based on the classification results, perform corresponding processing on the credibility label. Then, attach the credibility label, weighted deviation estimate, weighted deviation standard deviation, data accessibility level, and predicted value estimation interval to the original predicted value one by one to form a composite predicted data structure.
6. The intelligent drug clinical trial protocol generation method as described in claim 5, characterized in that, The data accessibility grading described in step S34 is performed according to the following rules: Grade A: For a certain prediction indicator, the number of valid reference drugs with published measured values in the local chemical neighborhood set is no less than 3, and the structural similarity of at least 1 reference drug is not lower than the structural similarity threshold. The confidence label is forcibly assigned to Grade A, indicating that the data is sufficient. Grade B: The number of effective reference drugs is 1 to 2, or the number is no less than 3 but the highest structural similarity in the local neighborhood is lower than the sufficient structural similarity threshold. The credibility label is forcibly assigned to Grade B, indicating that the data is limited. Grade C: The number of valid reference drugs is 0, or there are no reference drugs in the local neighborhood that meet the minimum threshold of structural similarity. In this case, bias estimation is not performed, and the confidence label is forcibly assigned to Grade C, indicating insufficient data.
7. The intelligent drug clinical trial protocol generation method as described in claim 1, characterized in that, The formula for calculating the dual-path initial dose in step S42 is as follows: Pathway A, based on the Well-Stirred model, calculates the hepatic system clearance rate using predicted values of hepatic microsomal clearance and plasma protein binding, and then extrapolates the starting dose at the target steady-state plasma concentration. ; in, To recommend a predicted steady-state peak concentration not exceeding a certain threshold, CL was calculated from the predicted liver microsomal clearance value using the Well-Stirred model. The dosing interval is denoted by F, the predicted oral bioavailability is denoted by SF, and the safety factor is SF. Pathway B extracts the approved clinical starting dose of the reference drug from the target evidence quality report and calculates it using the following formula: ; in, The reference drug's initial dose during Phase I; For reference drug in vitro activity values; The in vitro activity value of the candidate molecule; The Tanimoto structural similarity coefficient between the candidate molecule and the reference drug, with a value range of... ; The activity correction index; The similarity index is conservatively reduced.
8. The intelligent drug clinical trial protocol generation method as described in claim 1, characterized in that, Step S5 specifically includes: S51. A multi-dimensional weighted scoring mechanism was used to comprehensively score the feasibility of the Phase I clinical trial protocol. The scoring dimensions included evidence of target druggability, molecular safety, pharmacokinetic characteristics and differential potential. The credibility label in the target evidence quality report and composite prediction data structure was introduced as a moderating factor for each dimension score. S52. Based on the comprehensive feasibility score and credibility label, generate hierarchical decision-making suggestions according to preset thresholds; S53. Integrate the comprehensive feasibility score, the sub-scores of each dimension, and the hierarchical decision-making suggestions to output the final intelligent drug clinical trial plan.
9. An intelligent drug clinical trial protocol generation system, characterized in that, The system is used to implement the method as described in any one of claims 1-8, comprising: The data acquisition module is used to acquire structural data, target information, indication information, and molecular prediction data of candidate molecules. The data processing module is used to perform format conversion, field mapping, unit unification and standardization on the acquired data to obtain a standardized dataset; The knowledge retrieval module is used to retrieve related literature based on target information and indication information in a standardized dataset, and to retrieve measured data of reference drugs with the same target based on the target information. The data evaluation module is used to grade the quality of the retrieved related literature and calculate the comprehensive evidence quality score of the target-indication pair to generate a target evidence quality report; and to perform a deviation quantification evaluation of the molecular prediction data based on the measured data of the reference drug, and to attach a confidence label to each prediction indicator to generate a composite prediction data structure. The protocol generation module is used to generate Phase I clinical trial parameters for candidate molecules based on the target evidence quality report and composite prediction data structure, and obtain a preliminary Phase I clinical trial protocol. The decision evaluation module is used to comprehensively assess the feasibility of preliminary Phase I clinical trial protocols and generate tiered decision recommendations. The data storage module is used to store the acquired data, standardized datasets, search results, evaluation results, generated clinical trial protocols, and hierarchical decision-making recommendations. The report output module is used to integrate the preliminary Phase I clinical trial protocol and tiered decision-making recommendations to output an intelligent drug clinical trial protocol.