Self-adaptive clinical trial design method based on machine learning

Through machine learning adaptive clinical trial design, combined with reinforcement learning and data augmentation technology, the lack of real-time and safety in traditional clinical trials is solved, dynamically optimized treatment allocation and causal effect evaluation are achieved, and trial success rate and reliability of rare disease research and personalized medical treatment are improved.

CN120260969APending Publication Date: 2025-07-04CHONGQING HENGYUKANG PHARM TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510623544.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional clinical trials lack multi-dimensional considerations, preset rules lack real-timeness, neglecting patient safety, leading to high-risk allocation; the sample size in rare diseases or early drug trials is limited, and traditional designs are difficult to achieve statistical significance, and the results are unreliable.

Method used

Adaptive clinical trial design method based on machine learning is adopted, and reinforcement learning, causal inference and data augmentation technology is integrated, data volume is enhanced through virtual patient data, treatment allocation is dynamically optimized, treatment plans are adjusted in real time, and trials are terminated through causal effect evaluation and multi-objective optimization rules.

Benefits of technology

Significantly reduce the risk of improper allocation of high-risk groups, improve the statistical power of small sample scenarios, shorten the trial cycle, enhance patient benefits and ethical compliance, improve the reliability and success rate of trial results, and expand the application range to rare disease research and personalized medical care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260969A_ABST
    Figure CN120260969A_ABST
Patent Text Reader

Abstract

The invention provides a self-adaptive clinical trial design method based on machine learning, relates to the technical field of clinical trials, and adopts a method integrating reinforcement learning, causal inference and data enhancement technologies to predict an optimal treatment scheme and adjust a distribution proportion in real time through a dynamic optimization module, and limits high-risk treatment selection in combination with dynamic ethical constraints. The problems that a traditional clinical test preset rule lacks real-time performance and neglects patient safety are solved, a data collection and enhancement module generates virtual patient data conforming to a causal relationship, real and virtual data are combined to form a unified data set, a high-quality data set is constructed, the data size is effectively expanded, the statistical power of a small sample scene is improved, and the accuracy of a clinical test result is improved. The termination rule module of comprehensive statistics, ethics and resources shortens the test period through weighted evaluation, significantly reduces the risk of improper allocation of high-risk groups, enhances the benefit of patients, ethics compliance and test result reliability, and expands the application range in rare disease research and personalized medical treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of clinical trials, and specifically to an adaptive clinical trial design method based on machine learning. Background Art

[0002] According to a design method of a drug clinical trial protocol disclosed in Chinese Patent No. "CN112951351A", it includes the following steps: Step 1, convert the clinical trial protocol into a row-restricted covering array design problem; Step 2, construct a covering array with fewer rows to generate a row-restricted covering array; Step 3, convert the generated row-restricted covering array into an actual clinical trial protocol. The present invention constructs a covering array with row restrictions, and then obtains an actual clinical test protocol. Compared with ordinary heuristic methods, the present invention is based on the existing covering array, and then through mathematical operations, a row-restricted covering array can be obtained. Finally, the generated row-restricted covering array is converted into an actual clinical test protocol, with obvious timeliness advantages.

[0003] The above patent document and the prior art have the following technical problems when in use:

[0004] Problem 1: Traditional clinical trial adjustments usually refer to a single standard, lack multi-dimensional considerations, and are usually based on preset rules, lacking real-time performance, unable to continuously respond to new data, unable to adjust according to the real-time reactions of patients. At the same time, in traditional clinical trials, it is easy to ignore patient safety, resulting in high-risk treatments, and usually the safety differences are not considered, which may lead to some high-risk elderly patients being assigned to inappropriate treatment plans;

[0005] Problem 2: In rare diseases or early drug trials, the patient sample size is limited, and traditional designs rely on real data, making it difficult to achieve statistical significance, resulting in unreliable results or trial failures. Summary of the Invention

[0006] Technical Problems to be Solved

[0007] In view of the deficiencies of the prior art, the present invention provides an adaptive clinical trial design method based on machine learning, which solves the following problems:

[0008] 1. Traditional clinical trials lack multi-dimensional considerations due to a single statistical significance standard, preset rules lack real-time performance, and ignore patient safety, resulting in high-risk allocations;

[0009] 2. In rare diseases or early drug trials, it is difficult to achieve statistical significance due to limited patient sample size and traditional designs relying on real data, resulting in unreliable results or trial failures.

[0010] Technical Solutions

[0011] To achieve the above objectives, the present invention is implemented through the following technical solutions: An adaptive clinical trial design method based on machine learning, the adaptive clinical trial design method comprising the following steps:

[0012] Sp1: Design an experimental framework that integrates reinforcement learning, causal inference, and data augmentation techniques. Determine the treatment plan and patient allocation method of the trial by defining the initial treatment arms and allocation ratios. Initialize a virtual data generation system to generate virtual patient data during the trial to augment the data volume, and set an ethical constraint mechanism.

[0013] Sp2: Continuously collect the real data of patients and generate virtual patient data to augment the data volume during the trial. By obtaining the baseline data and treatment response data of patients, generate virtual data that conforms to the causal relationship between treatment and outcome, and merge the real data and virtual data into a unified analysis dataset to support subsequent optimization and evaluation.

[0014] Sp3: Perform dynamic optimization of treatment allocation on the merged dataset, specifically including defining states, actions, and rewards, predicting the best treatment plan based on real-time data, and dynamically adjusting the treatment allocation ratio to preferentially allocate more optimal treatments for patients.

[0015] Sp4: Evaluate the dynamic causal effect of the treatment to optimize the trial design. By analyzing longitudinal data to quantify the impact of the treatment on patient outcomes at different time points, combine propensity score and outcome prediction methods to ensure the robustness of the evaluation, and adjust the trial parameters according to the causal effect results.

[0016] Sp5: Implement a multi-objective optimization rule to decide the termination of the trial. By comprehensively evaluating the statistical significance of the treatment effect, the safety of each treatment arm, and the resource usage of the trial, design a weighted evaluation mechanism, and terminate the trial when the comprehensive result of this mechanism reaches a preset threshold.

[0017] Preferably, the treatment arms in Sp1 include at least two different treatment plans, namely a novel intervention treatment plan and a standard control treatment plan. The initial allocation ratio is determined based on historical data, and the historical data includes the results of previous trials, literature reports, and data from patient cohort studies. The virtual data generator based on the generative adversarial network is used to generate virtual patient data during the trial to augment the data volume. Setting the ethical constraint mechanism of the trial includes the criteria for identifying high-risk actions and dynamically adjusted ethical thresholds.

[0018] Preferably, the baseline data in Sp2 includes age, gender, and biomarkers. Collecting the treatment response data of patients includes short-term and long-term outcomes after treatment. Training the generative adversarial network model includes a generator and a discriminator, where the generator generates virtual patient data and the discriminator distinguishes real data from virtual data.

[0019] Preferably, in Sp2, by continuously acquiring the patient's baseline data and treatment response data, virtual patient data conforming to the causal relationship between treatment and outcome is generated using a generative adversarial network, and the real data and virtual data are merged into a unified data set. The merging steps are as follows: First, standardize the real data and the virtual data generated by the generative adversarial network to unify the format and distribution. Then, verify the causal consistency of the virtual data to ensure that it conforms to the causal relationship between treatment and outcome. Next, add source labels to the two types of data and integrate them into a comprehensive data set, and adjust the weights to balance the ratio of real data and virtual data to avoid bias. Finally, perform quality inspection and encrypted storage on the merged data set, and use blockchain technology to ensure data integrity and traceability, so as to provide a data set for dynamically optimizing treatment allocation and causal effect evaluation.

[0020] Preferably, the dynamic optimization of treatment allocation in Sp3 is achieved through deep ethical reinforcement learning, which specifically includes defining patient characteristics and trial progress as states, treatment plans as actions, and comprehensive short-term responses and long-term outcomes as rewards, combining a deep Q-network to predict the expected cumulative rewards of each action, and introducing dynamic ethical thresholds to constrain the selection of high-risk actions to ensure safety, and dynamically adjusting the treatment allocation ratio according to the prediction results to preferentially allocate more optimal treatments for patients.

[0021] Preferably, the causal effect evaluation in Sp4 uses time series causal inference technology combined with the double robust estimation of propensity scores and outcome prediction to improve the robustness of the evaluation, adjusts the trial parameters according to the causal effect results, reduces the allocation ratio of treatment arms with poor effects, and increases the sample size of treatment arms with significant effects.

[0022] Preferably, the multi-objective optimization rule in Sp5 is achieved by designing a weighted evaluation mechanism to comprehensively evaluate the statistical significance of treatment effects, the safety of each treatment arm, and the resource usage of the trial. Specifically, it includes defining the statistical significance index as the p-value of the causal effect evaluation, the safety index as the severe adverse event rate, and the resource efficiency index as the sample size and budget utilization rate, allocating dynamic weights to balance the priorities of each objective, using a weighted formula to generate a comprehensive evaluation score, and terminating the trial when the score reaches a preset threshold, and dynamically updating the weights and thresholds using real-time data.

[0023] Preferably, the adaptive clinical trial design method additionally includes a computer system. The composition architecture of the computer system includes a trial design module, a data collection and enhancement module, a dynamic optimization module, a causal effect evaluation module, a termination rule module, a data security module, and a model interpretation module. The trial design module scientifically sets trial parameters by defining the initial treatment plan and ethical constraints. The data collection and enhancement module constructs a high-quality analysis data set by merging real data and virtual data. The dynamic optimization module dynamically adjusts treatment allocation by predicting the best treatment plan. The causal effect evaluation module robustly quantifies the treatment effect by analyzing longitudinal data. The termination rule module makes a scientific decision on trial termination through comprehensive multi-objective evaluation. The data security module realizes data integrity and privacy security through blockchain and privacy protection technologies. The model interpretation module realizes the transparency of the optimization process through decision analysis and visualization tools.

[0024] Advantages

[0025] The present invention provides an adaptive clinical trial design method based on machine learning, having the following advantages:

[0026] 1. The present invention adopts an adaptive clinical trial design method that integrates reinforcement learning, causal inference, and data enhancement technologies, comprehensively considering multiple dimensions of statistics, ethics, and resources, which is superior to traditional single criteria. The dynamic optimization module predicts the best treatment plan in real time and adjusts the allocation ratio. Combining dynamic ethical constraints ensures that high-risk treatment options are restricted, solving the problems of the lack of real-time nature in the preset rules of traditional clinical trials and the high-risk allocation caused by ignoring patient safety. The dynamic optimization module continuously adjusts treatment allocation based on patient characteristics and real-time response data. The causal effect evaluation module quantifies the safety difference through longitudinal data analysis. The termination rule module comprehensively considers significance, safety, and resource efficiency through a weighted evaluation mechanism, significantly reducing the risk of improper allocation for high-risk groups, and shortening the trial cycle by about 20% by responding to new data in real time, improving patient benefit and ethical compliance.

[0027] 2. The present invention uses the data collection and enhancement module to generate virtual patient data that conforms to the causal relationship between treatment and outcome through a generative adversarial network, and merges real and virtual data to form a unified data set, solving the problems of limited patient sample size in rare diseases or early drug trials and the dependence of traditional designs on real data. This module constructs a high-quality data set through standardization, causal consistency verification, and weight adjustment, supports high-precision analysis for dynamic optimization and causal effect evaluation, effectively expands the data volume, significantly improves the statistical power in small sample scenarios, thereby enhancing the reliability and success rate of the trial results, and expanding the application scope of the method in rare disease research and early drug development. Brief Description of the Drawings

[0028] Figure 1 is the method step diagram of the present invention;

[0029] Figure 2 is the system architecture diagram of the present invention;

[0030] Figure 3 is the system hardware diagram of the present invention;

[0031] Figure 4 is the graph of the treatment allocation ratio of the present invention changing with time;

[0032] Figure 5 is the graph of the main observation indexes of the present invention changing with time;

[0033] Figure 6 is the graph of the comprehensive evaluation score of the present invention changing with time;

[0034] Figure 7 is the bar graph of the treatment effect comparison of the present invention. Detailed implementation manners

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Specific embodiment 1:

[0037] As Figures 1 to 7 shown, an adaptive clinical trial design method based on machine learning, the adaptive clinical trial design method includes the following steps:

[0038] Sp1: Design a trial framework that integrates reinforcement learning, causal inference, and data augmentation techniques. Determine the treatment plan and patient allocation method of the trial by defining the initial treatment arm and allocation ratio. Initialize a virtual data generation system to generate virtual patient data during the trial to enhance the data volume, and set an ethical constraint mechanism to ensure the safety and fairness of the trial. The treatment arm includes at least two different treatment plans, namely a new intervention treatment plan and a standard control treatment plan. The initial allocation ratio is determined based on historical data, and the historical data includes the results of previous trials, literature reports, and data from patient cohort studies. The virtual data generator based on the generative adversarial network (GAN) is used to generate virtual patient data during the trial to enhance the data volume. Setting the ethical constraint mechanism of the trial includes identifying the criteria for high-risk actions and dynamically adjusted ethical thresholds;

[0039] Sp2: Continuously collect the real data of patients during the trial and generate virtual patient data to enhance the data volume. By obtaining the baseline data and treatment response data of patients, generate virtual data that conforms to the causal relationship between treatment and outcome, and merge the real data and virtual data into a unified analysis dataset to support subsequent optimization and evaluation. The baseline data includes age, gender, and biomarkers. The collected treatment response data of patients includes short-term and long-term outcomes after treatment. Training the generative adversarial network model includes a generator and a discriminator, where the generator generates virtual patient data and the discriminator distinguishes between real data and virtual data;

[0040] By continuously obtaining the baseline data and treatment response data of patients, use the generative adversarial network to generate virtual patient data that conforms to the causal relationship between treatment and outcome, and merge the real data and virtual data into a unified dataset. The merging steps are as follows: First, standardize the real data and the virtual data generated by the generative adversarial network to unify the format and distribution. Subsequently, verify the causal consistency of the virtual data to ensure that it conforms to the causal relationship between treatment and outcome. Then, add source labels to the two types of data and integrate them into a comprehensive dataset, and adjust the weights to balance the ratio of real data and virtual data to avoid bias. Finally, conduct quality inspection and encrypted storage on the merged dataset, and use blockchain technology to ensure data integrity and traceability, so as to provide a dataset for dynamically optimizing treatment allocation and causal effect evaluation;

[0041] Sp3: Dynamically optimize the treatment allocation for the merged dataset, predict the best treatment plan based on real-time data, and dynamically adjust the treatment allocation ratio to preferentially allocate more optimal treatments for patients. The dynamic optimization of treatment allocation is achieved through deep ethical reinforcement learning, specifically including defining patient characteristics and trial progress as states, treatment plans as actions, and comprehensive short-term and long-term outcomes as rewards, combining the deep Q-network to predict the expected cumulative rewards of each action, and at the same time introducing dynamic ethical thresholds to constrain the selection of high-risk actions to ensure safety, and dynamically adjusting the treatment allocation ratio according to the prediction results to preferentially allocate more optimal treatments for patients;

[0042] Sp4: Evaluate the dynamic causal effect of treatment to optimize the trial design. Quantify the impact of treatment on patient outcomes at different time points by analyzing longitudinal data, combine propensity score and outcome prediction methods to ensure the robustness of the evaluation, and adjust the trial parameters according to the causal effect results. The causal effect evaluation uses time series causal inference technology combined with the double robust estimation of propensity score and outcome prediction to enhance the robustness of the evaluation, adjust the trial parameters according to the causal effect results, reduce the allocation ratio of treatment arms with poor effects, and increase the sample size of treatment arms with significant effects;

[0043] Sp5: Implement multi-objective optimization rules to determine trial termination. By comprehensively evaluating the statistical significance of treatment effects, the safety of each treatment arm, and the resource utilization of the trial, design a weighted evaluation mechanism and terminate the trial when the comprehensive result of this mechanism reaches a preset threshold. The multi-objective optimization rules design a weighted evaluation mechanism to comprehensively evaluate the statistical significance of treatment effects, the safety of each treatment arm, and the resource utilization of the trial. Specifically, define the statistical significance index as the p-value for causal effect evaluation, the safety index as the severe adverse event rate, and the resource efficiency index as the sample size and budget utilization rate. Allocate dynamic weights to balance the priorities of each objective, use a weighted formula to generate a comprehensive evaluation score, and terminate the trial when the score reaches the preset threshold. At the same time, use real-time data to dynamically update the weights and thresholds.

[0044] This method breaks through the limitations of traditional clinical trial fixed designs and single-standard adjustments by integrating reinforcement learning, causal inference, and data augmentation techniques, achieving dynamic optimization of treatment allocation, robust evaluation of causal effects, and scientific termination of multi-objective optimization. It significantly reduces the risk of improper allocation for high-risk groups, effectively expands the data volume, and significantly improves the statistical power in small-sample scenarios, thereby enhancing the reliability and success rate of trial results. It is applicable to rare disease research, early drug trials, and personalized medicine scenarios, and has the characteristics of high efficiency, ethics, safety, and wide applicability. Specific Embodiment 2:

[0046] As Figures 1 to 7 shown, based on the content in the above specific embodiments, the following content is further disclosed:

[0047] Based on the content of the above Specific Embodiment 1, it further includes the following content:

[0048] In Sp1, this trial framework is an adaptive clinical trial design framework that integrates reinforcement learning, causal inference, and data augmentation techniques, aiming to optimize treatment allocation and trial termination decisions through real-time data analysis and dynamic adjustment. Design method: First, define the initial treatment arms, including new interventions and standard controls, and define the allocation ratio, set initial parameters using historical data, then initialize the generative adversarial network as a virtual data generation system to enhance the data volume, and finally set an ethical constraint mechanism, including criteria for identifying high-risk treatments and dynamically adjusted ethical thresholds, to ensure patient safety and trial fairness;

[0049] Among them, the novel intervention treatment plan is a plan based on newly developed drugs, therapies, and combination treatments, aiming to provide innovative treatments for the biological mechanisms of specific diseases. The novel intervention treatment plan is usually based on the latest scientific research results or preclinical data. Historical data from early trials indicate its potential efficacy, making it suitable as the main treatment arm of the trial. The standard control treatment plan is an existing treatment method that has been widely accepted and is used as a control group to compare the efficacy and safety with the novel intervention treatment plan. The standard control treatment plan is determined based on published clinical trial results and has known efficacy and side effect data, serving as a benchmark for evaluating the effect of the novel treatment;

[0050] Based on historical data, including data sources such as previous trial results, literature reports, and patient cohort studies, combined with Bayesian prior statistical methods to analyze the treatment effect and safety. Among them, if historical data shows that the novel intervention treatment plan shows a higher initial response rate in a specific subgroup, the initial allocation ratio may tend to this treatment arm, that is, 60% is allocated to the novel treatment and 40% is allocated to the standard treatment. If the historical data is insufficient or the efficacy of the two plans is comparable, uniform allocation is adopted to ensure the exploratory nature at the initial stage of the trial.

[0051] In Sp2, real data collection continuously monitors the patient's baseline data and treatment response data through medical sensors, wearable devices, and electronic health records. Virtual data generation uses GAN to learn the real data distribution and causal relationships to generate virtual patient data that conforms to the treatment and outcome relationship. Real data comes directly from patients. Real data includes baseline characteristics and actual treatment responses. Virtual data is simulated baseline and response data. Virtual data is a simulation based on the real data distribution, used to increase the data volume, expand small sample data through virtual data, support analysis, merge the two to form a unified dataset, improve the robustness and accuracy of statistical analysis, and support dynamic optimization and causal effect evaluation;

[0052] The architecture of the generative adversarial network includes an input layer, a generator, a discriminator, a causal constraint module, a data processing unit, and an output interface. The input layer includes receiving real patient data, obtaining baseline and response data from data collection devices, importing historical data, and ensuring that the generated data conforms to the real distribution. The generator includes a deep neural network that inputs random noise and causal graph priors to generate virtual patient data, and outputs data containing baseline features and response data. The discriminator includes a convolutional neural network that evaluates the authenticity of the data and outputs the probability of being real or virtual, and guides the optimization of the generator through adversarial training to approximate the real data distribution. The causal constraint module includes an embedded generator that, based on causal inference tools, ensures that the virtual data conforms to the causal relationship between treatment and outcome, regularly verifies causal consistency, and adjusts the weights of the generator. The data processing unit includes standardizing real and virtual data, unifying the format, performing merging operations, integrating data sets using database tools, adding source labels, balancing the data ratio through weight adjustment, and performing quality checks to ensure no missing or abnormal data. The output interface includes transmitting the unified data set to the dynamic optimization module to support deep ethical reinforcement learning for optimizing the allocation ratio, providing data to the causal effect evaluation module for time series causal inference, and storing the data set in the data security module using blockchain technology to ensure integrity and traceability;

[0053] The operating logic of the generative adversarial network is as follows: Obtain real patient data, including baseline features and treatment response data, from the trial design module as the training basis for the GAN model. The generator generates virtual patient data from random noise, simulating the distribution and causal relationship of real data. The discriminator distinguishes between real data and virtual data, and optimizes the generator through adversarial training to make the virtual data realistic and conform to the causal structure of treatment and outcome. Ensure that the virtual data reflects the real relationship between treatment and outcome through causal graph priors. Then the generator outputs virtual patient data, including simulated baseline features and treatment responses, aligned with the real data format, expanding the data set to support dynamic optimization and causal effect evaluation. Subsequently, standardize the virtual data and real data to unify the format and distribution, verify the distribution consistency through statistical tests, integrate them into a unified data set after adding source labels, balance the ratio of real and virtual data through weight adjustment to ensure unbiased analysis. Finally, provide the unified data set to the dynamic optimization module for treatment allocation optimization and the causal effect evaluation module for effect analysis, enhancing the statistical power in small sample scenarios and improving the reliability of trial results. Adjust the parameters of the GAN generator according to the allocation results of the dynamic optimization module and the feedback of the causal effect evaluation module to optimize the quality of virtual data and continuously support the trial process.

[0054] In Sp3, the data is a combined and unified analysis dataset, including patient baseline characteristics, treatment response data, and dummy data. The dynamic allocation analyzes this data through a dynamic optimization module, predicts the best treatment plan, and adjusts the ratio. The dynamic adjustment utilizes Deep Ethical Reinforcement Learning (DERL) technology, which combines a Deep Q-Network and ethical constraints to predict the best treatment plan based on real-time data. The specific method is as follows: Define states, actions, and rewards, and adjust the allocation ratio through DERL. Among them, if a new intervention shows better efficacy and safety, the ratio increases from 50% to 70%. The allocation ratio changes dynamically according to the prediction, and the treatment arm with better efficacy is preferentially allocated.

[0055] In Sp4, the longitudinal data is structured data obtained by repeatedly measuring the same group of patients at multiple time points, recording the dynamic changes of patients during the trial, covering baseline characteristics, treatment allocation, and health outcomes. The time points are determined according to the trial design and disease characteristics. Using multi-time point data can reduce random errors and enhance the robustness of the estimation. Among them, the baseline characteristics are the information of patients at the beginning of the trial, including age, gender, biomarkers, and disease severity. The treatment allocation is the treatment plan received by the patient and its time point. The health outcomes include short-term outcomes and long-term outcomes. The short-term outcomes are the symptom remission rate and changes in biomarker levels 1 month after treatment. The long-term outcomes are the survival rate and quality of life score at 6 months or 1 year. The data is collected in a time series through measurements every week, month, or at each follow-up, forming a curve of the patient's condition over time.

[0056] The propensity score calculates the probability of a patient receiving a certain treatment, which is used to adjust the confounding factors of the patient's baseline differences. It is estimated through a logistic regression model to ensure a fair comparison of treatment allocations. By calculating the propensity score, which represents the probability of a patient receiving a certain treatment given the baseline characteristics, the propensity score is estimated using logistic regression. The input is the baseline characteristics, and the output is the treatment allocation probability. Propensity Score Matching (PSM) is applied to balance the characteristic distributions of the treatment group and the control group and reduce selection bias.

[0057] The outcome prediction method predicts the potential outcomes of patients with or without treatment, improving the estimation accuracy. A linear regression model is used to predict the health outcomes of patients. The input includes baseline characteristics, treatment allocation, and longitudinal data. Based on the potential outcome framework, the expected outcomes of the treatment group and the control group are predicted respectively. Combining the longitudinal data, a time series model is constructed to capture the changing trend of the outcomes over time.

[0058] The double-robust estimation integrates the propensity score and outcome prediction to form a double protection mechanism, ensuring that the estimation remains accurate even if one of the models is incorrect. The stability of the estimation is evaluated through 10-fold cross-validation to check whether the results are sensitive to data partitioning.

[0059] Adjusting trial parameters optimizes the trial design based on the results of causal effects, reduces the use of ineffective treatments, improves efficiency and patient benefit. The specific adjustment methods are as follows: Reduce the allocation ratio of the treatment arm with poorer effects: If the causal effect assessment shows that the effect of a certain treatment arm is significantly lower than that of another treatment arm, it indicates that its contribution to patient outcomes is smaller. Quantify the causal effect difference and adjust the allocation ratio. When the treatment effect of the standard control is significantly lower than that of the new intervention, increase the allocation ratio of the new intervention from 50% to 70%, and reduce the standard control from 50% to 30%. Use dynamic allocation rules to update the probability according to the real-time causal effect to ensure that more patients are allocated to the treatment arm with better effects; Increase the sample size of the treatment arm with significant effects: If a certain treatment arm shows a high causal effect, more samples may be needed to confirm its statistical significance. Increase the sample size of this treatment arm, ensure that the preset significance level is reached by recalculating the sample size requirement, adjust the recruitment plan, and preferentially include patients suitable for this treatment. In rare disease trials, if the causal effect of gene therapy is high, increase its sample size to verify the stability of the efficacy; Terminate the treatment arm with poor effects: If the causal effect of a certain treatment arm is close to zero or negative and shows no improvement trend, it indicates that it has no clinical value. According to the preset threshold, mark this treatment arm as ineffective, stop the patient allocation of this treatment arm, and reallocate resources to other treatment arms or newly designed treatment plans; Adjust the trial time points or follow-up frequencies: When longitudinal data may show that the treatment effect is more obvious at specific time points, adjust the follow-up plan, increase the measurement frequency at key time points, and extend or shorten the trial period to capture long-term effects or accelerate termination.

[0060] In Sp5, based on the causal effect estimation of Sp4, the p-value is calculated. The p-value is an indicator of statistical significance and is used to evaluate whether the treatment effect reaches a significant level. Determine whether the average treatment effect (ATE) is significantly non-zero through statistical tests. ATE is the average difference in outcomes between the treatment group and the control group. The p-value judges whether this difference is caused by the treatment rather than random fluctuations through statistical tests. Therefore, Sp4 first estimates ATE through time series causal inference and then obtains the p-value using ATE for testing. Next, use a t-test to verify whether ATE is significant, calculate the standard error (SE) of ATE, based on the sample variance and sample size, the t-statistic formula is: Look up the two-tailed p-value according to the t-distribution table. If the sample size is as large as 1000 people, a z-test can be used. ATE follows a normal distribution, and the formula is When the sample size is small or the data is non-normal, the bootstrap method is used. Resample 1000 times to construct the ATE distribution, and calculate the probability that ATE = 0 as the p-value. The calculation is implemented in MATLAB to ensure the robustness of the results. The p-value is used to quantify whether the causal effect of the treatment is statistically significant, and to evaluate the effect difference between the new intervention and the standard control. The p-value represents the probability of observing the current data or more extreme data under the null hypothesis. A small p-value indicates a significant treatment effect and rejects the null hypothesis, while a large p-value indicates an insignificant effect.

[0061] The evaluation operation of the weighted evaluation mechanism is as follows: Based on the response data of Sp2, the side effect rate of each treatment arm is statistically analyzed. Based on the progress of the trial, the sample size utilization rate, budget consumption rate, and time schedule are calculated.

[0062] S = ω1·S sig + ω2·S saf + ω3·S res

[0063] S: Comprehensive evaluation score;

[0064] S sig : Statistical significance score (1 - p-value, range [0,1]);

[0065] S saf : Safety score (1 - side effect rate, range [0,1]);

[0066] S res : Resource efficiency score (1 - budget utilization rate, range [0,1]);

[0067] ω1, ω2, ω3: Weights (0.4, 0.4, 0.2), adjusted according to the trial objectives;

[0068] Based on the preliminary trial results in Sp2, the ethical sensitivity trial increases ω2, and the resource-constrained trial increases ω3. Based on statistical criteria, ethical guidelines, and budget constraints, the threshold is set to 0.85, requiring a significant p < 0.05, a safety side effect rate < 5%, and a budget utilization rate < 90%. The threshold can vary with the trial stage. Based on the unified dataset of Sp2, S is updated every time new data is collected. If S ≥ the threshold, a termination signal is triggered. Before termination, it needs to be confirmed by the ethics committee. Terminating the trial when the comprehensive result of the weighted evaluation mechanism reaches the preset threshold is to ensure that the termination decision is based on clear and quantifiable criteria, avoiding both premature termination resulting in insufficient evidence and false negatives, which affect the success rate of new drug development, and late termination resulting in patients being exposed to ineffective or high-risk treatments, or wasting resources. Specific Example Three:

[0070] As Figures 1 to 7 shown, based on the content in the above specific examples, the following content is further disclosed:

[0071] The core formula content of each algorithm in the above system is as follows:

[0072] Deep Ethical Reinforcement Learning (DERL): Deep Ethical Reinforcement Learning combines traditional reinforcement learning with ethical constraints to ensure that while optimizing long-term rewards, patient safety and ethical requirements are met. It not only pursues the maximization of treatment effects but also avoids high-risk decisions through mathematical constraints;

[0073] Q(s,a)←Q(s,a)+α(r+γmax a' Q(s',a')-Q(s,a))

[0074] Q(s,a): The expected cumulative reward (Q-value) of choosing action a in state s;

[0075] s: The current state, such as the patient's health indicators or the trial phase;

[0076] a: The action taken, such as assigning a certain treatment plan;

[0077] r: The immediate reward, such as the short-term effect after treatment;

[0078] α: The learning rate (0 ≤ α ≤ 1), which controls the speed of Q-value update;

[0079] γ: The discount factor (0 ≤ γ < 1), which measures the importance of future rewards;

[0080] s': The next state after executing action a;

[0081] a': The possible action in the next state;

[0082] max a' Q(s',a'): The Q-value of the optimal action in the next state;

[0083] Ethical constraints: When choosing an action, DERL adds an ethical condition:

[0084] ifa∈high-risk actions,thenQ(s,a)≥θ ethics

[0085] a∈high-risk actions: If the action is classified as high-risk;

[0086] θ ethics : The dynamic ethical threshold, which is adjusted according to the patient's situation or the trial phase;

[0087] Only when the Q-value exceeds the ethical threshold will high-risk actions be considered, thus ensuring safety.

[0088] Time - series Causal Inference (TCI): Time - series causal inference combines longitudinal data and double - robust estimation to dynamically evaluate the causal effect of a treatment. It uses time - series data to improve the accuracy of causal inference and is particularly suitable for the real - time analysis of treatment effects in adaptive clinical trials.

[0089] The causal effect of double - robust estimation is:

[0090]

[0091] ATE: The average causal effect of the treatment, which is the difference in outcomes between the treatment group and the control group.

[0092] ΙE: Mathematical expectation, which calculates the average value of all samples.

[0093] Indicator function, which is 1 when the treatment is received (A = 1) and 0 otherwise.

[0094] Indicator function, which is 1 when the treatment is not received (A = 0) and 0 otherwise.

[0095] Propensity score, which represents the probability of receiving treatment under the covariate X.

[0096] X: Covariate, such as the age, gender, or disease severity of a patient.

[0097] Y: Observed outcome, such as the recovery rate or survival time.

[0098] The predicted outcome of the treatment group (A = 1) under the covariate X.

[0099] The predicted outcome of the control group (A = 0) under the covariate X.

[0100] This formula adjusts for confounding factors through the propensity score and combines the double - robustness of the prediction model to ensure that the causal effect estimate remains reliable even if the model is partially incorrect.

[0101] Generative Adversarial Network (GAN): Generative adversarial network introduces causal consistency constraints to ensure that the generated data is not only realistic but also conforms to the causal relationship of the treatment. It can be used to make up for data deficiencies and support causal analysis.

[0102]

[0103] G: Generator, which is responsible for generating realistic data from noise z.

[0104] D: Discriminator, which is responsible for distinguishing between real data x and generated data G(z).

[0105] V(D, G): The objective function of GAN, representing the adversarial loss between the discriminator and the generator;

[0106] The expectation of the true data distribution p data ;

[0107] x: True data, such as the actual treatment records of patients;

[0108] The expectation of the noise distribution p z ;

[0109] z: Random noise input for generating data;

[0110] D(x): The probability that the discriminator judges x as true data;

[0111] G(z): The data generated by the generator based on the noise z;

[0112] 1 - D(G(z)): The probability that the discriminator judges the generated data as fake;

[0113] A causal graph prior is added to the generator to ensure that the generated data satisfies the causal structure between treatment and outcome. Specific Embodiment Four:

[0115] As Figures 1 to 7 shown, based on the content in the above specific embodiments, the following content is further disclosed:

[0116] The composition architecture of the system includes an experimental design module, a data collection and enhancement module, a dynamic optimization module, a causal effect evaluation module, a termination rule module, a data security module, and a model interpretation module, specifically including the following:

[0117] Experimental design module: By defining at least two treatment plans and initial allocation ratios, it lays a scientific foundation for the experiment, initializes the virtual data generation system to enhance the data volume, solves the small - sample problem, sets ethical constraints to ensure that the selection of high - risk treatment plans meets patient safety standards, and integrates a design framework of dynamic optimization, causal analysis, and data enhancement technologies to support real - time adjustment and personalized treatment. It provides initial parameters for the data collection and enhancement module and treatment plan options for the dynamic optimization module. The experimental design module uses a high - performance server equipped with a multi - core CPU and GPU. The GPU accelerates the training of the generative adversarial network model to quickly generate realistic virtual data, and the multi - core CPU processes the calculation of experimental parameters and monitors ethical constraints in real time to ensure that the experimental design is scientific, reasonable, and meets ethical standards;

[0118] Data Collection and Enhancement Module: By obtaining patients' baseline data and treatment response data, combining with virtual data to generate data that conforms to the causal relationship between treatment and outcome, merging real and virtual data to form a unified dataset, providing high-quality input for subsequent dynamic optimization and causal assessment, breaking through the data volume limitation of traditional trials, enhancing the applicability of the method, supporting the state definition of the dynamic optimization module and the longitudinal data analysis of the causal effect assessment module. The data collection and enhancement module collects data through wearable devices and medical sensors to ensure data diversity and accuracy, improving the statistical analysis ability in small-sample scenarios. Edge computing nodes process data in real time to reduce latency, and the data storage server saves real and virtual data to support subsequent analysis. The collection device ensures data diversity and accuracy, improving the statistical analysis ability in small-sample scenarios;

[0119] Dynamic Optimization Module: By defining states, actions, and rewards, predicting the best treatment plan based on real-time data, dynamically adjusting the allocation ratio, optimizing patient benefit and trial efficiency. Using dynamic optimization technology to achieve personalized treatment allocation, transcending the limitations of traditional random allocation. Relying on the unified dataset of the data collection and enhancement module, feedback the adjustment results to the causal effect assessment module. The dynamic optimization module relies on a high-performance computing cluster equipped with GPUs and TPUs. GPUs and TPUs accelerate model training and inference to ensure real-time prediction and decision-making, quickly respond to data changes, and improve trial efficiency and treatment effect;

[0120] Causal Effect Assessment Module: By analyzing longitudinal data to quantify the impact of treatment on patients' outcomes at different time points, combining propensity score and outcome prediction methods to ensure the robustness of the assessment, adjusting trial parameters according to the causal effect results, optimizing the trial design. Using the data of the data collection and enhancement module to provide a significance basis for the termination rule module. The causal effect assessment module uses a big data analysis platform equipped with high-memory servers and a distributed computing framework to analyze longitudinal data such as post-treatment scores, quantify the causal effect, adjust trial parameters. The distributed framework processes large-scale data, and the high-memory server supports complex causal model calculations, improving the accuracy and reliability of causal inference, providing a scientific basis for trial optimization;

[0121] Termination Rule Module: By comprehensively evaluating the statistical significance of treatment effects, the safety of each treatment arm, and resource utilization, a weighted evaluation mechanism is designed to generate a comprehensive score and terminate the trial when the score reaches a preset threshold, ensuring a balance between scientificity and ethics. The multi-objective optimization rule dynamically responds to the trial progress, differentiating from single-criterion termination. It integrates the significance data of the causal effect evaluation module and the resource information of the data collection module. The termination rule module evaluates the statistical effect differences, adverse event rates, and budget usage through a decision support system equipped with a multi-core CPU and real-time monitoring tools to decide whether to terminate the trial. The multi-core CPU efficiently processes the multi-objective optimization rule, and the monitoring tools track the trial progress to ensure that the termination decision is timely and scientific, balancing ethics and efficiency;

[0122] Data Security Module: Stores and shares trial data through blockchain technology to ensure data integrity and immutability, and uses differential privacy technology to protect patient privacy, providing a secure data environment for all modules, meeting the regulatory requirements of clinical trials. It provides secure storage for the data collection and enhancement module and reliable data for other modules. The data security module uses a blockchain server and encryption hardware modules such as a hardware security module to store and share trial data, ensuring data integrity and privacy protection. The blockchain server records operation logs for transparent traceability, and the encryption module manages keys to protect sensitive data from unauthorized access, meeting the strict regulatory requirements of clinical trials;

[0123] Model Explanation Module: By analyzing the decision-making basis of dynamic optimization and causal evaluation, it provides visualization tools to display the adjustment of allocation ratios and causal effect results, enhancing the trust of clinicians and regulatory agencies, supporting the transparency of the method and regulatory approval, explaining the decisions of the dynamic optimization module, and verifying the results of the causal effect evaluation module. The model explanation module interprets model decisions and provides visual displays through a visualization server equipped with a high-resolution monitor and GPU. The GPU accelerates the calculation and rendering of the SHAP value model explanation method, and the monitor clearly presents the dynamic adjustments and results, enhancing model transparency and regulatory trust for easy understanding by doctors and institutions;

[0124] The trial design module serves as the starting point of the system, providing an initial dataset and parameters for the data collection module, treatment options for the dynamic optimization module, an analysis basis for the causal effect evaluation module, and setting evaluation criteria for the termination rule module, laying the foundation for the entire trial. The data collection and enhancement module receives the initial data from the trial design module, outputs a high-quality dataset to the dynamic optimization module to support treatment allocation adjustment, to the causal effect evaluation module to support effect analysis, and receives feedback data from the dynamic optimization module to update the dataset. The dynamic optimization module dynamically adjusts the allocation ratio based on the real-time data provided by the data collection module, and feeds back the optimization results to the data collection module to update the data distribution, and outputs data to the causal effect evaluation module for effect analysis, indirectly affecting the significance evaluation of the termination rule module. The causal effect evaluation module receives the allocation data from the dynamic optimization module, verifies the optimization effect, outputs the significance result to the termination rule module to support the termination decision, and feeds back adjustment suggestions to the data collection module to optimize data generation. The termination rule module comprehensively considers the significance of the causal effect evaluation module, the safety of the data collection module, and resource efficiency such as the sample size, generates a comprehensive score through weighted evaluation, determines whether the termination threshold is reached, integrates the results of each module, and decides to terminate the trial to ensure the balance between science and ethics. The data security module and the model interpretation module run through the entire process, ensuring data privacy and model transparency respectively. Specific Embodiment Five:

[0126] As Figures 1 to 7 shown, based on the content in the above specific embodiments, the following content is further disclosed:

[0127] To further verify the advantages of the proposed solution of this application compared with the prior art, a comparative experiment is designed by comparing the proposed solution of this application with the prior art. The specific experimental content is as follows:

[0128] Experiment 1: Comparison of the effects of a novel intervention vs. a standard control on cell proliferation and apoptosis in small cell lung cancer (SCLC):

[0129] Experimental objective: To compare the effects of a novel intervention (PD-1 immune checkpoint inhibitor combined with an EGFR-targeted drug) and a standard control (traditional chemotherapy) on the inhibition of cell proliferation and induction of apoptosis in small cell lung cancer cells, and to evaluate the dynamic impact of the treatment on the biological behavior of cancer cells;

[0130] Treatment plan: In the novel intervention, patients received treatment with the PD-1 immune checkpoint inhibitor nivolumab combined with the EGFR-targeted drug osimertinib. The dose of nivolumab was 240 mg, administered intravenously once every 3 weeks, and the dose of osimertinib was 80 mg, taken orally once a day. The treatment lasted for 6 months, during which the patients' responses and side effects were regularly monitored. In the standard control, patients received chemotherapy with etoposide combined with cisplatin. The dose of etoposide was 100mg / m 2, administered by intravenous injection on days 1 to 3 of each cycle, with a cisplatin dose of 75 mg / m 2 , administered by intravenous injection on day 1 of each cycle, with a treatment cycle of every 3 weeks for a total of 6 cycles, and the treatment lasted for 6 months;

[0131] Observation indicators: Cell proliferation inhibition rate: The percentage reduction in the proliferation activity of cancer cells was evaluated by Ki-67 immunohistochemical staining; Apoptosis rate: The percentage of cancer cell apoptosis was evaluated by TUNEL staining; Reduction rate of circulating tumor cells (CTC): The percentage change in the number of CTCs in the blood was measured by liquid biopsy; Tumor volume reduction rate: The percentage reduction in the maximum diameter of the tumor was measured by CT / MRI; Clinical remission rate: The proportion of patients with partial remission (PR, tumor diameter reduction ≥ 30%) or complete remission (CR, tumor disappearance) was evaluated according to the RECIST criteria; Severe adverse event rate: Defined as side effects requiring medical intervention;

[0132] Experimental design: 120 patients with confirmed extensive-stage SCLC, aged 18 - 75 years, with balanced baseline characteristics such as gender and smoking history, were selected. Based on the historical data of Sp1, the initial allocation ratio was set, that is, 60% for the new intervention and 40% for the standard control. Sp1 initialized the treatment arm and the generative adversarial network virtual data generation system. Sp2 collected real data and generated virtual data through medical sensors and wearable devices, which were combined into a unified dataset. Sp3 used deep ethical reinforcement learning to dynamically optimize the allocation ratio. Sp4 evaluated the causal effect through time series causal inference. Sp5 decided to terminate the trial by comprehensively considering significance, safety, and resource efficiency. The treatment lasted for 6 months, with regular follow-up, that is, monthly CT / MRI, biopsies at months 0, 3, and 6, and monthly liquid biopsies.

[0133] Experiment 2: Comparison of the effects of gene therapy vs hormone therapy on muscle fiber repair and function improvement in Duchenne muscular dystrophy (DMD):

[0134] Experimental objectives: To compare the effects of gene therapy (CRISPR-Cas9 gene editing) and hormone therapy (glucocorticoids) on muscle fiber repair and function improvement in patients with Duchenne muscular dystrophy, and to evaluate the impact of the treatment on muscle cell regeneration and strength recovery;

[0135] Treatment Plan: In gene therapy, patients receive CRISPR-Cas9 gene editing therapy delivered via an adeno-associated virus vector, administered as a single intravenous injection at a dose of 1×10^14 viral genomes per kilogram of body weight, targeting the repair of exon mutations in the DMD gene to promote dystrophin expression. After treatment, patients are followed up for 12 months to monitor muscle function and safety; in hormone therapy, patients receive glucocorticoid prednisone treatment at a dose of 0.75 mg / kg, orally once a day. If severe side effects occur, i.e., if weight gain exceeds 20%, the dose can be reduced to 0.5 mg / kg / day orally. The treatment lasts for 12 months, and patients' responses and side effects are regularly evaluated.

[0136] Observation Indicators: Muscle fiber repair rate: The percentage of repaired muscle fibers is evaluated through muscle biopsy and immunofluorescence staining, i.e., the proportion of normal muscle fibers; Muscle strength improvement rate: The percentage increase in walking distance is evaluated through the standard 6-minute walk test (6MWT); Reduction rate of muscle fiber necrosis: The percentage reduction in necrotic muscle fibers is evaluated through HE staining of tissue sections; Decrease rate of creatine kinase (CK) level: The percentage reduction in CK concentration is evaluated through serum testing, reflecting the degree of muscle damage; Improvement rate of quality of life (QoL): The percentage increase in the quality of life score is evaluated through the PedsQL scale; Severe adverse event rate: Defined as side effects requiring medical intervention.

[0137] Experimental Design: Eighty DMD patients, aged 5 - 15 years with a balanced distribution of gene mutation types, are selected. Based on Sp1, they are evenly allocated, i.e., 50% receive gene therapy and 50% receive hormone therapy. Sp1 sets the treatment arms and ethical constraints. Sp2 monitors exercise data through wearable devices and generates virtual data. Sp3 dynamically adjusts the allocation ratio. Sp4 evaluates the causal effect. Sp5 decides to terminate. The treatment lasts for 12 months, with regular follow-up, i.e., muscle biopsy and 6MWT every 3 months, serum CK monthly, and PedsQL at months 0, 6, and 12.

[0138] Results of Experiment 1: Novel Intervention vs Standard Control (SCLC):

[0139]

[0140]

[0141] Results of Experiment 2: Gene Therapy vs Hormone Therapy (DMD):

[0142]

[0143] Experiment 1: Novel Intervention vs Standard Control (SCLC):

[0144] Cell proliferation inhibition rate: The cancer cell proliferation activity in the novel intervention group decreased by 68%. The Ki-67 positive rate decreased from 60% before treatment to 20%. In the standard control group, it decreased by 42%, from 58% to 34%. This indicates that the novel intervention effectively inhibits cancer cell division by activating the immune system and blocking growth signals, while the cytotoxic effect of chemotherapy is weaker and some cancer cells may develop drug resistance.

[0145] Cell apoptosis rate: The cancer cell apoptosis rate in the novel intervention group reached 27%. TUNEL staining showed that a large number of cells entered the apoptotic state. In the standard control group, it was 16%. The novel intervention induces a stronger apoptotic response through immune killing and targeted signals, while the apoptotic effect of chemotherapy is limited by drug resistance.

[0146] Reduction rate of circulating tumor cells: The number of circulating tumor cells in the novel intervention group decreased by 55%, from 15 per 7.5 ml of blood to 7. In the standard control group, it decreased by 25%, from 14 to 11. The novel intervention significantly reduces the risk of metastasis through immune clearance and inhibition of tumor spread, while the effect of chemotherapy is weaker.

[0147] Tumor volume reduction rate: The average tumor diameter in the novel intervention group decreased by 48%. Computed tomography and magnetic resonance imaging showed a 40% reduction rate at the 4th month. In the standard control group, it was 32%, and only 25% at the 4th month. The synergistic mechanism of the novel intervention more effectively controls tumor growth.

[0148] Clinical remission rate: 45% of the patients in the novel intervention group achieved remission, including 27 with partial remission and 2 with complete remission. In the standard control group, it was 22%, including 13 with partial remission and 1 with complete remission. The novel intervention significantly improves the tumor control rate of patients and improves the prognosis.

[0149] Severe adverse event rate: 4% in the novel intervention group, that is, 5 cases of rash among 120 people. In the standard control group, it was 8%, that is, 10 people had nausea and vomiting among 120 people. The average for the whole group was 4.5%, lower than the traditional design.

[0150] The novel intervention is comprehensively superior to chemotherapy in terms of cell proliferation inhibition, apoptosis induction, reduction of circulating tumor cells, tumor shrinkage, and clinical remission, showing its powerful anti-cancer effect. The 45% remission rate meets the treatment expectations for small cell lung cancer, proving its clinical potential.

[0151] Experiment 2: Gene therapy vs hormone therapy (DMD):

[0152] Muscle fiber repair rate: The proportion of normal muscle fibers in the gene therapy group increased from 10% to 65%, with a repair rate of 55%. In the hormone therapy group, it increased from 12% to 47%, with a repair rate of 35%. Gene therapy restores dystrophin expression by repairing the DMD gene and promotes muscle fiber regeneration. Hormone therapy only indirectly protects muscles through anti-inflammatory effects, with limited effects.

[0153] Muscle strength improvement rate: In the gene therapy group, the distance of the 6-minute walk test increased from 200 meters to 296 meters, with an improvement rate of 48%. In the hormone therapy group, it increased from 205 meters to 262 meters, with an improvement rate of 28%. Gene therapy directly improves muscle structure and significantly enhances walking ability, while the protective effect of hormone therapy is weaker.

[0154] Reduction rate of muscle fiber necrosis: In the gene therapy group, the proportion of necrotic muscle fibers decreased from 30% to 12%, with a reduction rate of 60%. In the hormone therapy group, it decreased from 32% to 19%, with a reduction rate of 40%. Gene therapy stabilizes the muscle membrane and reduces muscle damage, while hormone therapy only partially alleviates inflammation-related necrosis.

[0155] Decrease rate of creatine kinase level: In the gene therapy group, creatine kinase decreased from 8000 units per liter to 2800 units per liter, with a decrease rate of 65%. In the hormone therapy group, it decreased from 8200 units per liter to 4510 units per liter, with a decrease rate of 45%. Gene therapy reduces muscle damage and lowers creatine kinase level, while the effect of hormone therapy is weaker.

[0156] Improvement rate of quality of life: In the gene therapy group, the score of the children's quality of life scale increased from 50 points to 75 points, with an improvement rate of 50%. In the hormone therapy group, it increased from 52 points to 68 points, with an improvement rate of 30%. Gene therapy significantly improves the patient's daily life quality by improving motor ability, while hormone therapy limits the improvement amplitude due to side effects.

[0157] Severe adverse event rate: 3% for gene therapy (2 fevers among 80 people), 6% for hormone therapy (5 people with weight gain among 80 people), and the overall group average is 4.5%, which is lower than the traditional design.

[0158] Gene therapy is comprehensively superior to hormone therapy in muscle repair, strength improvement, necrosis reduction, creatine kinase decrease, and quality of life improvement, showing its advantage in precisely repairing gene defects. The 50% quality of life improvement rate meets the treatment challenges of Duchenne muscular dystrophy, proving its potential.

[0159] Summary:

[0160] Dynamic optimization allocation is achieved through deep ethical reinforcement learning: The trial design module initially allocates 60% of Experiment 1 to the new intervention and 40% to the standard control, and Experiment 2 to 50% gene therapy and 50% hormone therapy. The dynamic optimization module uses deep ethical reinforcement learning to predict treatment effects based on real-time data and adjust the allocation ratio. In Experiment 1, it was found that the Ki-67 proliferation inhibition rate of the new intervention reached 60% at the 3rd month while that of chemotherapy was only 35%, so the proportion of the new intervention was increased to 75%. In Experiment 2, it was found that the 6-minute walk test strength improvement rate of gene therapy reached 40% at the 6th month while that of hormone therapy was only 20%, so the proportion of gene therapy was increased to 65%. This resulted in 90 patients in Experiment 1, that is, 90 out of 120, receiving the new intervention, and 52 patients in Experiment 2, that is, 52 out of 80, receiving gene therapy, significantly improving the 45% clinical remission rate in Experiment 1 and the 55% muscle fiber repair rate in Experiment 2. Traditional fixed allocation would have half of the patients receive the less effective chemotherapy or hormone therapy, thus reducing the overall success rate. Deep ethical reinforcement learning preferentially allocates more effective treatments by responding to real-time data changes, surpassing the rigid mode of traditional random allocation and greatly improving patient benefit and trial efficiency;

[0161] Data augmentation is achieved through generative adversarial networks: The data collection and augmentation module generates virtual data for the small samples of 120 people in Experiment 1 and 80 people in Experiment 2. Experiment 1 generates 60 people of virtual data to expand the dataset to 180 people, and Experiment 2 generates 50 people of virtual data to expand the dataset to 130 people. The generative adversarial network simulates the Ki-67 staining and circulating tumor cell data in Experiment 1 and the muscle biopsy and creatine kinase data in Experiment 2, and ensures that the virtual data accurately reflects the treatment effect differences through causal consistency verification. The augmented dataset significantly improves the accuracy of statistical analysis. The p-value in Experiment 1 drops from 0.08 to 0.01, thus confirming the difference in the proliferation inhibition rate of 68% compared to 42%, and the p-value in Experiment 2 drops from 0.06 to 0.01, thus confirming the difference in the muscle fiber repair rate of 55% compared to 35%. Traditional small-sample trials may not draw significant conclusions due to insufficient data, thus missing the true advantages of the new intervention. Generative adversarial networks effectively make up for the sample size limitations of complex diseases and rare diseases such as small cell lung cancer and Duchenne muscular dystrophy, greatly enhancing the reliability and statistical significance of the results;

[0162] Robust causal effect assessment is achieved through time series causal inference: The causal effect assessment module uses time series causal inference and combines propensity score matching to balance baseline characteristics, i.e., age and smoking history in Experiment 1, and gene mutation type and age in Experiment 2. By analyzing longitudinal data, i.e., monthly circulating tumor cell detection in Experiment 1 and muscle biopsy every three months in Experiment 2, the treatment effect is estimated. In Experiment 1, the average treatment effect of the new intervention is calculated to be 0.26 with a p-value of 0.01, and in Experiment 2, the average treatment effect of gene therapy is calculated to be 0.20 with a p-value of 0.01. Based on these results, the chemotherapy proportion is reduced to 25% and the hormone therapy proportion is reduced to 35%. This precise assessment confirms the significant advantages of the new intervention, i.e., a 48% tumor shrinkage rate in Experiment 1 and a 60% reduction in muscle fiber necrosis in Experiment 2, and optimizes the trial design to ensure that resources are concentrated on more effective treatments. Traditional methods may underestimate the treatment effect due to confounding bias, thus affecting the credibility of the results. The combination of time series causal inference and double-robust estimation improves the scientific nature and precision of the assessment, provides a reliable basis for dynamic allocation, and directly boosts the overall trial success rate;

[0163] Multi-objective optimization termination is achieved through weighted assessment: The termination rule module synthesizes significance (i.e., p-value of 0.01 in Experiment 1 and p-value of 0.03 in Experiment 2), safety (i.e., side effect rate of 4% for the new intervention in Experiment 1 compared to 8% for chemotherapy and side effect rate of 3% for gene therapy in Experiment 2 compared to 6% for hormone therapy), and resource efficiency (i.e., 120 people used in Experiment 1 compared to the planned 150 people and 80 people used in Experiment 2 compared to the planned 100 people) to calculate the weighted score. The score in Experiment 1 is 0.88 and the score in Experiment 2 is 0.86, both exceeding the threshold of 0.85, thus triggering the termination of the trial. This timely avoids patients from continuing to receive less effective chemotherapy (i.e., 22% remission rate) and hormone therapy (i.e., 30% quality of life improvement rate), saves 20% of the sample size, and focuses on the 45% remission rate of the new intervention and the 50% quality of life improvement rate of gene therapy. Traditional single-criterion termination may affect the reliability of the results due to premature or late stopping. Multi-objective optimization achieves a scientific termination decision by balancing scientific nature, ethics, and resource efficiency, accelerating the verification and clinical application of effective treatments.

[0164] Ethics and safety are achieved through dynamic ethical thresholds: The trial design module sets the side effect rate threshold at 5%. The dynamic optimization module restricts high-risk treatments through deep ethical reinforcement learning, i.e., patients with no decrease in circulating tumor cells in Experiment 1 are avoided from chemotherapy, and patients with high creatine kinase levels in Experiment 2 are avoided from hormone therapy. The termination rule module continuously monitors side effects, i.e., 4% rash in the new intervention in Experiment 1 compared to 8% nausea and vomiting in chemotherapy, and 3% mild fever in gene therapy in Experiment 2 compared to 6% weight gain in hormone therapy. This reduces the risk of patients being exposed to high-side-effect treatments and preferentially allocates low-risk PD-1+EGFR and gene therapies, thus improving patient compliance and trial participation, indirectly increasing the apoptosis rate by 27% in Experiment 1 and the muscle strength improvement rate by 48% in Experiment 2. Traditional methods may ignore real-time safety, thus increasing patient risk. The dynamic ethical threshold ensures patient safety through real-time monitoring and adjustment, enhances the ethics of the trial, and improves the quality and credibility of the results.

[0165] Through the above experimental design and result comparison, the solution of this application is superior to the prior art in terms of trial efficiency, patient benefit, safety, and resource utilization efficiency. Especially in the small-sample scenario, this solution significantly improves the trial performance through virtual data augmentation and dynamic optimization, providing a scientific basis for its application in actual clinical trials. Specific Embodiment Six:

[0167] As Figures 1 to 7 shown, based on the content in the above specific embodiments, the following content is further disclosed:

[0168] To further verify the feasibility of the technical solution of this application, it is further illustrated through the following case:

[0169] Case background: This case takes the rare disease retinitis pigmentosa (RP) as the research object to verify the application of the adaptive clinical trial design method in the small-sample and high-heterogeneity scenario. RP is a hereditary eye disease. Due to the gradual degeneration of retinal photoreceptor cells, patients suffer from night blindness, narrowed visual fields, and may eventually become completely blind. Because of the diverse gene mutations and large individual differences in disease progression among patients, it is difficult for traditional clinical trials to quickly and effectively evaluate new therapies. The adaptive design aims to find a better treatment plan for RP patients by dynamically adjusting trial parameters;

[0170] Application method: For the RP patients, two treatment arms were designed for the trial: gene therapy, a novel intervention targeting RPE65 gene mutations, and supportive care, a standard control without specific drugs. The initial allocation ratio was 50:50. The trial recruited 40 real patients and generated 20 virtual patients through simulation, expanding the sample to 60 people. Every three months, according to the ETDRS visual acuity chart score and the results of the visual field test, the treatment allocation ratio was dynamically adjusted, giving priority to the group with better effects. The trial set a comprehensive termination threshold of 0.75, considering the treatment effect and the incidence rate of side effects comprehensively. In the initial stage, 40 patients were randomly assigned to the novel intervention group and the standard control group, with 20 people in each group. After three months, the data showed that the visual acuity decline rate in the novel intervention group was slower and the degree of visual field shrinkage was also smaller. Based on this, the adaptive algorithm adjusted the proportion of the novel intervention group to 65%. At six months, the average treatment effect in the novel intervention group was significant, the side effect rate was low, and the comprehensive score reached 0.78, triggering the early termination of the trial.

[0171]

[0172] The adaptive clinical trial design performed excellently in RP patients. By dynamically optimizing the treatment allocation, the proportion of the novel intervention group was increased to 65%. The visual acuity decline rate within six months was only 10%, the degree of visual field shrinkage was 15%, and the side effect rate was as low as 3%. The trial time was shortened from 10 months in the traditional design to 7 months, saving 12% of resources. This method efficiently and safely accelerated the verification of new therapies, providing a new path for the treatment of complex and rare diseases such as RP.

[0173] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a reference structure" does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0174] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made therein without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An adaptive clinical trial design method based on machine learning, characterized in that: The described adaptive clinical trial design method includes the following steps: Sp1: Design a trial framework that integrates reinforcement learning, causal inference, and data augmentation techniques. Determine the treatment plan and patient allocation method of the trial by defining the initial treatment arms and allocation ratios. Initialize a virtual data generation system to generate virtual patient data during the trial to increase the data volume, and set up an ethical constraint mechanism; Sp2: Continuously collect the real data of patients and generate virtual patient data during the trial. By obtaining the baseline data and treatment response data of patients, generate virtual data that conforms to the causal relationship between treatment and outcome, and merge the real data and virtual data into a unified analysis dataset to support subsequent optimization and evaluation; Sp3: Dynamically optimize the treatment allocation for the merged dataset, predict the best treatment plan based on real-time data, and dynamically adjust the treatment allocation ratio to preferentially allocate patients to treatments with better effects; Sp4: Evaluate the dynamic causal effect of the treatment to optimize the trial design. By analyzing longitudinal data to quantify the impact of the treatment on patient outcomes at different time points, combining propensity score and outcome prediction methods, and adjusting the trial parameters according to the causal effect results; Sp5: Implement a multi-objective optimization rule to decide the termination of the trial. By comprehensively evaluating the statistical significance of the treatment effect, the safety of each treatment arm, and the resource usage of the trial, design a weighted evaluation mechanism, and terminate the trial when the comprehensive result of this mechanism reaches a preset threshold.

2. The adaptive clinical trial design method based on machine learning according to claim 1, wherein: The treatment arms in Sp1 include at least two different treatment plans, namely a novel intervention treatment plan and a standard control treatment plan. The initial allocation ratio is determined based on historical data, which includes the results of previous trials, literature reports, and data from patient cohort studies. The virtual data generator based on the generative adversarial network is used to generate virtual patient data during the trial to increase the data volume. Setting the ethical constraint mechanism of the trial includes the criteria for identifying high-risk actions and the dynamically adjusted ethical threshold.

3. The adaptive clinical trial design method based on machine learning according to claim 1, wherein: The baseline data in Sp2 includes age, gender, and biomarkers. Collecting the treatment response data of patients includes the short-term and long-term outcomes after treatment. Training the generative adversarial network model includes a generator and a discriminator, where the generator generates virtual patient data and the discriminator distinguishes between real data and virtual data.

4. The adaptive clinical trial design method based on machine learning according to claim 1, characterized in that: In Sp2, by continuously obtaining the baseline data and treatment response data of patients, using the generative adversarial network to generate virtual patient data that conforms to the causal relationship between treatment and outcome, and merging the real data and virtual data into a unified dataset. The merging steps are as follows: First, standardize the real data and the virtual data generated by the generative adversarial network to unify the format and distribution. Subsequently, verify the causal consistency of the virtual data to ensure that it conforms to the causal relationship between treatment and outcome. Then, add source labels to the two types of data and integrate them into a comprehensive dataset, and balance the ratio of real data and virtual data through weight adjustment to avoid bias. Finally, conduct quality inspection and encrypted storage on the merged dataset, and use blockchain technology to ensure data integrity and traceability, so as to provide a dataset for dynamically optimizing treatment allocation and causal effect evaluation.

5. The adaptive clinical trial design method based on machine learning according to claim 1, characterized in that: The dynamic optimization of treatment allocation in Sp3 is achieved through deep ethical reinforcement learning, specifically including defining patient characteristics and trial progress as states, treatment options as actions, and combining short-term responses and long-term outcomes as rewards. By integrating a deep Q-network to predict the expected cumulative rewards of each action, a dynamic ethical threshold is introduced to constrain the selection of high-risk actions, ensuring safety. Meanwhile, the treatment allocation ratio is dynamically adjusted according to the prediction results to preferentially allocate more optimal treatments for patients.

6. The adaptive clinical trial design method based on machine learning according to claim 1, wherein: The causal effect evaluation in Sp4 uses time series causal inference techniques combined with double robust estimation of propensity scores and outcome prediction to enhance the robustness of the evaluation. According to the causal effect results, the trial parameters are adjusted to reduce the allocation ratio of ineffective treatment arms and increase the sample size of effective treatment arms.

7. The adaptive clinical trial design method based on machine learning according to claim 1, characterized in that: The multi-objective optimization rule in Sp5 designs a weighted evaluation mechanism to comprehensively evaluate the statistical significance of treatment effects, the safety of each treatment arm, and the resource utilization of the trial. Specifically, it defines the statistical significance index as the p-value of causal effect evaluation, the safety index as the severe adverse event rate, and the resource efficiency index as the sample size and budget utilization rate. Dynamic weights are assigned to balance the priorities of each objective, and a weighted formula is used to generate a comprehensive evaluation score. The trial is terminated when the score reaches a preset threshold, and the weights and thresholds are dynamically updated using real-time data.

8. The adaptive clinical trial design method based on machine learning according to claim 1, wherein: The adaptive clinical trial design method additionally includes a computer system. The composition architecture of the computer system includes a trial design module, a data collection and enhancement module, a dynamic optimization module, a causal effect evaluation module, a termination rule module, a data security module, and a model interpretation module. The trial design module scientifically sets the trial parameters by defining the initial treatment plan and ethical constraints. The data collection and enhancement module constructs a high-quality analysis data set by merging real data and virtual data. The dynamic optimization module dynamically adjusts treatment allocation by predicting the best treatment plan. The causal effect evaluation module robustly quantifies treatment effects by analyzing longitudinal data. The termination rule module makes a scientific decision on trial termination through comprehensive multi-objective evaluation. The data security module ensures data integrity and privacy security through blockchain and privacy protection technologies. The model interpretation module achieves the transparency of the optimization process through decision analysis and visualization tools.

Citation Information

Cited By

  • Generation method and generation system of myopia operation scheme and storage medium

    CN120579069A

  • Clinical test subject management system and method based on block chain

    CN120673948A

  • Real-time monitoring and processing method and system for impact resistance test data of foam concrete and application

    CN121185800A