Clinical experiment data analysis method based on reinforcement learning

Through hierarchical feature extraction and reinforcement learning strategy analysis, combined with clinical risk constraints and efficacy evaluation, the treatment strategy is dynamically adjusted, which solves the problems of insufficient dynamic response analysis and global optimization in traditional methods, and achieves improved accuracy and safety of personalized treatment plans.

CN120600337APending Publication Date: 2025-09-05BEIJING SHUMANDE MEDICAL TECH DEV CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510689978.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional clinical data analysis methods have difficulty capturing the dynamic correlation characteristics of multi-dimensional physiological indicators, lack dynamic adaptability and global optimization capabilities, and separate risk assessment from efficacy evaluation. Existing reinforcement learning methods fail to fully utilize the hierarchical structure of time series data, and multi-stage models do not consider the dynamic changes of the treatment cycle.

Method used

By performing hierarchical feature extraction on temporal physiological parameters, building a dynamic response feature model and combining it with reinforcement learning strategy analysis, an initial treatment response evaluation model is generated. After loading clinical risk constraints and efficacy evaluation indicators, a comprehensive treatment response evaluation model is constructed. Strategy control simulation is performed and the treatment path is optimized to achieve dynamic strategy adjustment and multi-stage optimization.

Benefits of technology

It achieves accurate description of patients' individual response characteristics, dynamically adjusts treatment strategies, reduces safety risks, improves the adaptability and global optimization capabilities of treatment plans, and ensures real-time matching and optimality of treatment effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600337A_ABST
    Figure CN120600337A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of clinical experiment data processing, and discloses a clinical experiment data analysis method based on reinforcement learning, and the method comprises the steps: obtaining time sequence physiological parameters and treatment scheme execution state parameters; performing hierarchical feature extraction on the time sequence physiological parameters, constructing a dynamic response feature model, and generating an initial treatment response evaluation model in combination with the execution state parameters; loading risk constraints and curative effect indexes to generate a comprehensive evaluation model, and simulating to obtain strategy regulation and control data; constructing and training a reinforcement learning decision model, and generating a treatment path correction model and initial treatment parameters; obtaining dynamic strategy adjustment information, and constructing a multi-stage optimization model for iterative optimization to obtain optimal treatment path parameters; reasoning a curative effect standard reaching probability based on a real-time model, and adjusting a strategy to realize dynamic curative effect matching. According to the method, the clinical data dynamic analysis and treatment scheme optimization capability is improved, and the method is suitable for precise medical scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of clinical trial data processing, and in particular to a clinical trial data analysis method based on reinforcement learning. Background Art

[0002] In modern medical research, clinical trials are a key step in verifying the safety and effectiveness of treatment plans. Their core lies in the in-depth analysis of temporal physiological parameters and treatment execution status. Traditional clinical data analysis methods rely primarily on statistical models and expert experience, which have the following significant limitations:

[0003] Traditional methods for processing temporal physiological parameters remain at the level of simple descriptive statistics, making it difficult to capture the dynamic correlation characteristics of multi-dimensional physiological indicators over time. For example, parameters such as blood pressure, heart rate, and blood oxygen saturation exhibit complex temporal coupling relationships during treatment. Traditional linear regression or univariate analysis cannot effectively characterize this dynamic response pattern, resulting in insufficient recognition of individual patient differences.

[0004] Treatment optimization lacks dynamic adaptability. Existing methods typically adjust doses and plan monitoring nodes based on fixed, pre-set rules, failing to respond in real time to subtle changes in a patient's physiological state. For example, when a patient experiences pharmacokinetic abnormalities or unexpected complications, traditional approaches struggle to quickly adjust intervention strategies, potentially leading to poor treatment outcomes or increased safety risks.

[0005] The separation of risk assessment and efficacy evaluation leads to insufficient comprehensive decision-making capabilities. Traditional models often construct risk prediction and efficacy assessment models independently, failing to integrate both into a unified optimization framework. This makes it difficult to balance efficacy improvement and risk control when formulating treatment plans. This can lead to excessive pursuit of efficacy while neglecting safety, or inadequate efficacy due to conservative strategies.

[0006] Furthermore, the ability to globally optimize multi-stage treatment processes is lacking. Clinical treatment is typically divided into multiple phases, with the objectives and parameter settings for each phase having temporal dependencies. Traditional approaches often employ independent optimization strategies for each phase, lacking a global overview of the entire treatment cycle. This can lead to inconsistent strategies between phases, compromising overall treatment effectiveness.

[0007] With the development of artificial intelligence (AI), reinforcement learning (RL), an advanced method capable of handling dynamic decision-making problems, has gradually been introduced into the medical field. However, existing RL-based clinical data analysis methods still have the following problems: First, the feature extraction method is single and fails to fully utilize the hierarchical structure of time series data; second, the strategy optimization process lacks explicit constraints on clinical risk and efficacy indicators, which may result in the generated solutions being inconsistent with clinical practice; and third, the construction of multi-stage models does not fully consider the dynamic changes in the treatment cycle, making it difficult to achieve optimal control throughout the entire treatment cycle.

[0008] Therefore, there is an urgent need for a clinical trial data analysis method that can integrate dynamic feature modeling, reinforcement learning strategy optimization and multi-stage global planning to address the shortcomings of traditional methods in dynamic response analysis, real-time strategy adjustment and global optimization, and to improve the scientific nature of clinical experiments and the accuracy of treatment plans. Summary of the Invention

[0009] The purpose of the present invention is to provide a clinical trial data analysis method based on reinforcement learning to solve the problems raised in the above background technology.

[0010] To achieve the above objectives, the present invention provides the following technical solution: a clinical trial data analysis method based on reinforcement learning, the method comprising:

[0011] Obtaining the time series physiological parameter data set and treatment plan execution status parameters of clinical trial subjects;

[0012] Perform hierarchical feature extraction on temporal physiological parameters and construct a dynamic response feature model. Combined with the treatment plan execution state parameters, the dynamic response feature model is subjected to reinforcement learning strategy analysis to generate an initial treatment response evaluation model.

[0013] The initial treatment response evaluation model is loaded with clinical risk constraints and efficacy evaluation indicators to generate a comprehensive treatment response evaluation model. The treatment plan is dynamically simulated in combination with the execution state parameters to obtain strategy control simulation data.

[0014] Construct a reinforcement learning decision model and perform strategy optimization training through strategy-controlled simulation data to generate a treatment pathway correction model, thereby generating initial treatment parameters. The initial treatment parameters include at least the initial intervention time node, the initial dosing sequence, and the initial monitoring feedback threshold.

[0015] Based on the initial treatment parameters, the comprehensive treatment response evaluation model and the execution status parameters, dynamic strategy adjustment information is obtained;

[0016] Combining dynamic strategy adjustment information with treatment cycle planning parameters, a multi-stage optimization model is constructed and the planning parameters are iteratively optimized to obtain the optimal treatment path parameters;

[0017] Based on the real-time dynamic response characteristic model, the current treatment cycle parameters are inferred to generate the probability of achieving the efficacy target in the current stage. Combined with the optimal treatment path parameters in the current stage and the actual physiological parameter fluctuations, the treatment plan execution strategy is adjusted to achieve the dynamic efficacy matching goal.

[0018] Preferably, the strategy control simulation data includes at least patient response characteristic distribution, risk clustering area, safe treatment boundary set and strategy adjustment priority sequence;

[0019] The dynamic strategy adjustment information includes at least dose offset, treatment phase completion rate and efficacy stability evaluation index.

[0020] Preferably, the step of performing hierarchical feature extraction on temporal physiological parameters and constructing a dynamic response feature model, performing reinforcement learning strategy analysis on the dynamic response feature model in combination with treatment plan execution state parameters, and generating an initial treatment response evaluation model comprises the following steps:

[0021] Perform multi-dimensional signal decomposition on time-series physiological parameters to obtain data sequences with individualized response characteristics of patients;

[0022] Preprocessing the data sequence to obtain standardized feature fusion data, wherein the preprocessing includes one or more of outlier filtering, time series alignment, feature dimensionality reduction, stage division, and data normalization;

[0023] Based on standardized feature fusion data and combined with reinforcement learning strategy network, a dynamic response feature model is constructed;

[0024] A strategy exploration analysis is performed on the dynamic response characteristic model to generate an initial treatment response evaluation model, wherein the strategy exploration analysis at least includes state space definition and action space mapping, the therapeutic effect response level is divided by the state space definition, and corresponding strategy exploration weights are assigned to the parameters of each stage.

[0025] Preferably, the initial treatment response evaluation model is loaded with clinical risk constraints and efficacy evaluation indicators to generate a comprehensive treatment response evaluation model, and a treatment plan dynamic simulation is performed in combination with the execution state parameters to obtain strategy control simulation data, including the following steps:

[0026] Loading clinical risk constraints to the initial treatment response assessment model to simulate real-time risk changes, thereby generating a first treatment response assessment model;

[0027] Loading efficacy evaluation indicators into the first treatment response evaluation model to generate a comprehensive treatment response evaluation model, wherein the efficacy evaluation indicators include a maximum dose tolerance threshold, a minimum efficacy response threshold, and a monitoring parameter redundancy range;

[0028] Based on the comprehensive treatment response evaluation model and execution status parameters, dynamic simulation of treatment plans is performed. The specific process includes:

[0029] Constructing a multi-stage strategy planning equation, the multi-stage strategy planning equation includes at least a stage coverage integrity equation, a response consistency equation, and a safety equation. In combination with a comprehensive treatment response evaluation model, the multi-stage strategy planning equation is numerically solved using a Monte Carlo tree search method to obtain a patient response characteristic distribution, risk clustering areas, a safe treatment boundary set, and a strategy adjustment priority sequence;

[0030] The generation of the safe treatment boundary set includes the following steps:

[0031] Based on the risk clustering areas, the clinical risk coverage density of each treatment stage is calculated;

[0032] The stages in the comprehensive treatment response assessment model where the risk coverage density is no greater than the preset threshold are identified, and a set of safe treatment boundaries is generated.

[0033] Preferably, the construction of the reinforcement learning decision model and the strategy optimization training by strategy-controlled simulation data to generate a treatment path correction model and then generate initial treatment parameters include the following steps:

[0034] Construct a reinforcement learning decision model based on the dynamic response feature model;

[0035] Through the strategy control simulation data, the reinforcement learning decision model is trained and verified through strategy iteration to generate a treatment path correction model;

[0036] Input the real-time execution state parameters into the treatment path correction model to predict the set of safe treatment boundaries;

[0037] Based on the predicted safe treatment boundary set, initial treatment parameters are generated, wherein the initial treatment parameters at least include an initial intervention time node, an initial drug administration sequence, and an initial monitoring feedback threshold.

[0038] Preferably, the method further includes reconstructing the policy control simulation data, specifically:

[0039] Construct an initial multi-stage treatment input tensor based on the distribution of patient response characteristics, risk cluster regions, and a set of safe treatment boundaries;

[0040] Normalize and enhance the features of the initial multi-stage treatment input tensor to generate the final multi-stage treatment input tensor;

[0041] Adjust the priority sequence based on the strategy and construct the treatment path correction label tensor;

[0042] The final multi-stage treatment input tensor and the treatment path correction label tensor are combined to form a strategy training sample set.

[0043] Preferably, generating initial treatment parameters based on the predicted safe treatment boundary set comprises the following steps:

[0044] Extract the treatment stage with the lowest risk coverage density from the predicted safe treatment boundary set and generate the initial intervention time node;

[0045] According to the temporal correlation of the predicted safe treatment boundary set, the feasible connection structure of the initial administration dose sequence is fitted to generate the initial monitoring feedback threshold;

[0046] The patient response gradient direction of the predicted safe treatment boundary set is calculated and normalized to a dose adjustment reference vector, which is the direction of the initial administration dose sequence.

[0047] Preferably, obtaining dynamic strategy adjustment information based on the initial treatment parameters, the comprehensive treatment response evaluation model and the execution state parameters comprises the following steps:

[0048] Map the initial intervention time node to the comprehensive treatment response evaluation model, match the initial administration dose sequence with the initial monitoring feedback threshold, and update the strategy node density in the adjustment area. Update the comprehensive treatment response evaluation model, define dynamic strategy trigger conditions, dose reconstruction rules, and feedback threshold increments, and generate a dynamic strategy adjustment model.

[0049] Based on the dynamic policy adjustment model, the multi-stage policy planning equation is iteratively solved through the Monte Carlo tree search method to obtain dynamic policy adjustment information, including:

[0050] When the dynamic strategy triggering conditions are met, the dose offset, treatment stage completion rate, and efficacy stability evaluation indicators are updated, and the multi-stage strategy planning equation is re-solved until the simulation termination conditions are met;

[0051] The dynamic strategy triggering conditions include triggering the feedback threshold update when the completion rate of the current treatment stage is no greater than the preset completion rate threshold; the dose reconstruction rule includes adjusting the initial administration dose sequence based on the patient response gradient direction; the feedback threshold increment is in a piecewise linear relationship with the current efficacy stability evaluation index.

[0052] Preferably, the construction of a multi-stage optimization model and iterative optimization of planning parameters to obtain optimal treatment path parameters includes the following steps:

[0053] A multi-stage optimization model was constructed, in which the optimization variables included the strategy node distribution density, dose adjustment threshold, and monitoring sampling frequency. The optimization objectives included maximizing stage coverage efficiency and minimizing treatment risk increment. The constraints included the patient's physiological tolerance limit and the monitoring device accuracy threshold.

[0054] Perform a preliminary solution to the multi-stage optimization model using a dynamic programming algorithm to generate an initial set of optimization strategies;

[0055] Based on the initial optimization strategy set, the policy gradient optimization algorithm is used for global optimization to generate the optimal treatment path parameters.

[0056] Preferably, the construction of a multi-stage optimization model and iterative optimization of planning parameters to obtain optimal treatment path parameters further includes the following steps:

[0057] Construct a set of constraint relaxation parameters based on the patient's physiological tolerance upper limit and the monitoring equipment's accuracy threshold;

[0058] Dynamically match each strategy parameter in the initial optimization strategy set with the constraint relaxation parameter set to generate a feasible strategy subset;

[0059] The policy gradient optimization algorithm is used to prioritize the subset of feasible strategies and generate the optimal treatment path parameters.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] In terms of dynamic feature modeling, multi-dimensional signal decomposition and hierarchical feature extraction of time-series physiological parameters effectively capture the dynamic changes in individual patient response characteristics. Preprocessing operations such as outlier filtering, time-series alignment, and feature dimensionality reduction ensure the standardization and usability of the data, laying a solid foundation for subsequent model construction. The dynamic response feature model constructed using a reinforcement learning strategy network can dynamically characterize the mapping relationship between physiological parameters and treatment plans, achieving a precise description of the patient's status and addressing the problem of traditional methods' inadequate capture of dynamic correlation features in time-series data.

[0062] The reinforcement learning strategy analysis module, through state-space definition and action-space mapping, achieves a scientific classification of therapeutic response levels. It also assigns strategy exploration weights to parameters at each stage, enabling the model to dynamically adjust exploration strategies based on the characteristics of different treatment phases, improving the flexibility and targeted nature of strategy analysis. This hierarchical modeling and strategy exploration mechanism enhances the model's adaptability to individual patient differences and provides technical support for the development of personalized treatment plans.

[0063] The construction of a comprehensive treatment response assessment model achieves the organic unity of risk assessment and efficacy assessment by loading clinical risk constraints and efficacy assessment indicators (such as maximum dose tolerance threshold, minimum efficacy response threshold, etc.). The solution of the multi-stage strategy planning equation based on Monte Carlo tree search can generate key simulation data such as patient response characteristic distribution and risk clustering areas, providing comprehensive information support for the dynamic simulation and strategy optimization of treatment plans. The generation mechanism of the safe treatment boundary set quantifies the risk coverage density of each stage, clarifies the boundary range of safe treatment, and effectively reduces the safety risk during the treatment process.

[0064] The construction and training of the reinforcement learning decision-making model, through data reconstruction and iterative policy training of policy-controlled simulation data, generated a treatment pathway correction model capable of real-time prediction of safe treatment boundaries. This model dynamically generates initial treatment parameters (such as the initial intervention time point and initial dosing sequence) based on real-time execution state parameters, achieving initial optimization of the treatment plan. Standardization and feature enhancement during the data reconstruction process improved the quality of the training data and the generalization ability of the model, ensuring the scientific nature and reliability of the initial treatment parameters.

[0065] The dynamic strategy adjustment mechanism enables real-time dynamic adjustment of treatment plans by defining dynamic strategy trigger conditions, dose reconstruction rules, and feedback threshold increments. When trigger conditions are met, the model automatically updates parameters such as dose offset and treatment phase completion rate. It then iteratively solves multi-phase strategy planning equations to ensure the treatment plan remains optimal. This dynamic adjustment capability enables the treatment plan to respond to changes in the patient's physiological state in real time, improving the adaptability and effectiveness of treatment.

[0066] The construction of a multi-stage optimization model achieves global optimization of the entire treatment cycle by integrating optimization variables such as the distribution density of strategy nodes and the dose adjustment threshold, with the goal of maximizing stage coverage efficiency and minimizing the incremental treatment risk. The combination of a dynamic programming algorithm and a policy gradient optimization algorithm ensures that the model can find the optimal treatment path parameters under complex constraints (such as the patient's physiological tolerance limit and the monitoring device accuracy threshold), solving the problem of insufficient global optimality caused by traditional multi-stage independent optimization. The introduction of a constraint relaxation parameter set and the generation of a feasible strategy subset further improve the robustness and clinical feasibility of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 This is a working principle diagram of the clinical trial data analysis method based on reinforcement learning according to the present invention;

[0068] Figure 2 Flowchart for model generation for hierarchical feature extraction of temporal physiological parameters and initial treatment response assessment;

[0069] Figure 3 Flowchart for the construction of comprehensive treatment response assessment model and dynamic simulation of treatment regimen;

[0070] Figure 4 Flowchart for reconstruction of simulation data for policy control and generation of policy training sample sets;

[0071] Figure 5 Flowchart for generating dynamic policy adjustment information and solving multi-stage policy planning equations. DETAILED DESCRIPTION

[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0073] See also Figure 1-Figure 5 The present invention relates to a clinical trial data analysis method based on reinforcement learning, and the specific implementation steps are as follows:

[0074] Obtain a time-series physiological parameter dataset and treatment plan execution status parameters for clinical trial subjects. The time-series physiological parameter dataset contains multiple physiological indicator data of the subjects at different time points, and the treatment plan execution status parameters record the actual execution status of the treatment plan at each stage during the treatment process.

[0075] Hierarchical feature extraction is performed on temporal physiological parameters to construct a dynamic response feature model. This dynamic response feature model is then analyzed using reinforcement learning strategies in conjunction with treatment execution state parameters to generate an initial treatment response assessment model. Through in-depth analysis of temporal physiological parameters, features that reflect the subject's response to the treatment plan are extracted. Using reinforcement learning strategy analysis, an initial treatment response assessment model is then established.

[0076] Clinical risk constraints and efficacy evaluation indicators are loaded into the initial treatment response assessment model to generate a comprehensive treatment response assessment model. Dynamic simulation of treatment plans is then performed in conjunction with execution status parameters to generate simulation data for strategic control. Based on the initial model, the dynamic simulation of treatment plan execution, taking into account clinical risk factors and efficacy evaluation criteria, generates simulation data for strategic adjustment through dynamic simulation of treatment plan execution.

[0077] A reinforcement learning decision model is constructed and trained using simulated data for strategy optimization. This generates a treatment pathway correction model, which then generates initial treatment parameters. Initial treatment parameters include at least the initial intervention time point, initial dosing sequence, and initial monitoring feedback thresholds. The decision model is trained and optimized using simulated data to obtain a model capable of correcting the treatment pathway and generating initial treatment parameters.

[0078] Based on the initial treatment parameters, the comprehensive treatment response evaluation model, and the execution status parameters, dynamic strategy adjustment information is obtained. By combining the initial treatment parameters with the comprehensive evaluation model and the execution status parameters, the strategy information that needs to be dynamically adjusted during the treatment process is analyzed.

[0079] Combining dynamic strategy adjustment information with treatment cycle planning parameters, a multi-stage optimization model is constructed and the planning parameters are iteratively optimized to obtain the optimal treatment pathway parameters. Taking into account dynamic adjustment information and treatment cycle planning, the optimal treatment pathway parameters are determined through iterative calculations of the optimization model.

[0080] Based on a real-time dynamic response feature model, the system infers the parameters of the current treatment cycle and generates the probability of achieving the desired therapeutic effect at the current stage. Combining the optimal treatment pathway parameters with actual physiological parameter fluctuations, the system adjusts the treatment plan execution strategy to achieve dynamic efficacy matching. Leveraging the real-time feature model and optimal treatment pathway parameters, the system dynamically adjusts the treatment plan execution strategy based on changes in actual physiological parameters to ensure maximum therapeutic effect.

[0081] The present invention will be further described below in conjunction with Examples 1 to 5:

[0082] Example 1:

[0083] In the process of performing hierarchical feature extraction on temporal physiological parameters and constructing a dynamic response feature model, and combining the treatment plan execution status parameters to generate an initial treatment response evaluation model, key steps such as multi-dimensional signal decomposition, data sequence preprocessing, dynamic response feature model construction, and strategy exploration and analysis must be completed in sequence. Each step is closely connected and has clear technical logic and implementation details.

[0084] Multidimensional signal decomposition is performed to obtain data sequences that characterize individual patient responses. Time-series physiological parameters typically include multiple indicators with distinct physiological significance, such as electrocardiogram (ECG) signals, blood pressure fluctuations, blood oxygen saturation, temperature changes, and respiratory rate. These indicators reflect the patient's physiological status and treatment response from different dimensions. Multidimensional signal decomposition requires the use of signal processing methods tailored to the characteristics of each indicator. For example, ECG signals can be decomposed into different frequency components using methods such as Fourier transforms and wavelet transforms to extract features such as heart rate variability and ST segment deviation. For blood pressure data, a sliding window algorithm can be used to calculate the fluctuation amplitude of systolic and diastolic blood pressure, as well as the mean blood pressure trend. This decomposition operation breaks down the original mixed time series data into independent feature sequences corresponding to each indicator, ensuring that each sequence contains only information on changes in a single dimension of physiological parameters, facilitating subsequent in-depth analysis of the characteristics of each indicator. For example, when processing time series data containing heart rate and blood pressure, decomposing the heart rate sequence and blood pressure sequence separately can independently obtain the circadian rhythm characteristics of heart rate and the stress response characteristics of blood pressure, thereby more accurately capturing the individual patient response patterns across different physiological dimensions.

[0085] After completing multidimensional signal decomposition, each data series needs to be preprocessed to obtain standardized feature fusion data. This preprocessing process includes multiple operations, including outlier filtering, time series alignment, feature dimensionality reduction, stage partitioning, and data normalization. One or more combinations can be selected based on the data characteristics. Outlier filtering uses statistical thresholds (such as the Z-score method or the IQR method) or machine learning-based anomaly detection algorithms (such as the Isolation Forest algorithm or the Local Outlier Factor algorithm) to identify and remove abnormal data points that deviate significantly from the normal range, preventing them from interfering with model construction. For example, in a temperature data series, outliers above 42°C or below 32°C can be identified and removed using the IQR method. Time series alignment addresses the issue of inconsistent data collection frequencies for different indicators by using methods such as linear interpolation and spline interpolation to unify the series onto the same time grid, ensuring one-to-one correspondence in the temporal dimension during subsequent feature fusion. For example, if the ECG signal is sampled at 1000 Hz and the blood pressure data is collected every minute, interpolation is required to synchronize the blood pressure data to the same time points as the ECG signal.

[0086] Feature dimensionality reduction is used to reduce redundant features and computational complexity. Linear methods such as principal component analysis (PCA) and linear discriminant analysis (LDA) or nonlinear methods such as t-SNE and UMAP can be used. For example, when the dimensions of the extracted physiological features are as high as hundreds of dimensions, PCA can compress them to dozens of dimensions while retaining more than 95% of the information variance. Stage division is based on the implementation process of the treatment plan, dividing the entire time series data into different stages such as the pre-treatment baseline period, the initial treatment period, the treatment stabilization period, and the treatment adjustment period, so as to perform differentiated analysis based on the physiological response characteristics of each stage. For example, when evaluating the efficacy of a certain antihypertensive drug, the week before taking the drug can be set as the baseline period, the first 1-2 weeks after taking the drug can be set as the initial period, the third to eighth weeks can be set as the stabilization period, and the subsequent adjustment period can be set as the adjustment period if the dosage needs to be adjusted. Data normalization uses min-max scaling or standardization to map each feature value to a uniform range (such as [0, 1] or a distribution with a mean of 0 and a standard deviation of 1), eliminating the impact of dimensional differences in different indicators on model training. For example, systolic blood pressure (unit: mmHg) and heart rate (unit: beats / minute) are converted into dimensionless numerical sequences through standardization.

[0087] Based on preprocessed, standardized feature fusion data, a dynamic response feature model is constructed using a reinforcement learning policy network. Reinforcement learning policy networks typically employ deep neural network architectures, such as multi-layer perceptrons (MLPs), recurrent neural networks (RNNs), or transformers, to adapt to the sequential nature of time series data. For physiological parameters with long-term dependencies (e.g., multi-day blood sugar fluctuations), gated recurrent neural network structures such as LSTMs or GRUs can be used to capture long-term features through memory units. For signals with short-term, high-frequency variations (e.g., electrocardiogram waveforms), convolutional neural networks (CNNs) can be used to extract local features before connecting them to recurrent layers. The network input is a time series segment of the standardized feature fusion data, and the output is a policy probability distribution or value estimate corresponding to the current state. During model training, classic reinforcement learning algorithms, such as policy gradient methods and Q-learning, are employed. The treatment execution state parameters serve as environmental feedback signals. Through continuous interaction with the environment, network parameters are optimized, enabling the model to predict the patient's response probability to different treatment actions based on the input physiological parameter sequence. For example, taking a patient's physiological parameter sequence for three consecutive days as input, the model outputs the probability of adopting different drug dosage strategies under the current state, thereby constructing a feature model that can dynamically respond to the patient's physiological state.

[0088] After the dynamic response feature model is constructed, it is necessary to conduct strategy exploration analysis on it to generate an initial treatment response evaluation model. The core of strategy exploration analysis lies in the state space definition and action space mapping. The state space definition divides the patient's physiological parameter states into multiple discrete efficacy response levels, such as "excellent", "good", "moderate" and "poor", and each level corresponds to a combination range of a set of physiological indicators. For example, the "excellent" level can be defined as a heart rate of 60-100 beats / minute, a systolic blood pressure of 90-120 mmHg, a blood oxygen saturation of ≥95%, and the fluctuation range of each indicator is less than the preset threshold; the "poor" level corresponds to a state where the indicators are obviously abnormal or fluctuate violently. Through clear state division, the model can convert the continuous physiological parameter space into a finite discrete state, which is convenient for strategy search and evaluation.

[0089] Action space mapping defines the specific operations in the treatment plan (such as adjusting the dosage, changing the intervention time node, switching the monitoring frequency, etc.) as a set of executable actions, and establishes a mapping relationship between actions and states. For example, when the patient's condition is at the "poor" level, the action space may include actions such as "increase the dosage by 20%", "advance the intervention time by 2 hours", and "start high-frequency monitoring (once every 15 minutes)"; when the condition is at the "good" level, the action space may include actions such as "maintain the current dosage", "intervene as planned", and "maintain routine monitoring (once every hour)". Through this mapping, the model can clearly identify the treatment actions that can be taken in different states and their possible impact.

[0090] Based on the definition of state space and action space, corresponding strategy exploration weights are assigned to the parameters of each stage. The strategy exploration weight is used to adjust the model's tendency to explore new strategies and utilize known strategies at different treatment stages. For example, in the early stages of treatment (such as the baseline period or the dose escalation phase), a higher exploration weight can be assigned to encourage the model to actively try different treatment actions to explore the patient's response pattern; while in the stable treatment period, the exploration weight can be lowered, and proven effective strategies can be prioritized to maintain the stability of the treatment effect. The weight setting can be dynamically adjusted based on factors such as the time course of the treatment stage and the changing trend of the patient's status. For example, a linear attenuation strategy can be used to gradually reduce the exploration weight from an initial value of 0.8 to 0.2 as the treatment cycle progresses.

[0091] Through the aforementioned strategy exploration and analysis process, the dynamic response feature model is able to continuously explore the effects of different actions in the state space, accumulate experience, and develop the ability to evaluate treatment responses, ultimately generating an initial treatment response assessment model. This model can output expected efficacy assessment values ​​(such as efficacy level probability, risk score, etc.) for different actions under each state, given a sequence of physiological parameters and treatment plan execution status, providing a foundation for subsequent treatment plan optimization and adjustment. The entire process closely revolves around feature mining of time-series physiological parameters and strategy optimization through reinforcement learning. Through a multi-step technical operation, the transformation from raw data to a treatment response assessment model is achieved, embodying the clinical trial data analysis approach that combines data-driven and intelligent algorithms.

[0092] Example 2:

[0093] In the process of loading clinical risk constraints and efficacy evaluation indicators into the initial treatment response evaluation model to generate a comprehensive treatment response evaluation model, and dynamically simulating the treatment plan in combination with the execution state parameters to obtain strategy control simulation data, it is necessary to complete the key links such as clinical risk constraint loading, efficacy evaluation indicator integration, multi-stage strategy planning equation construction and Monte Carlo tree search solution in sequence. Each link forms a complete simulation analysis chain through logical association and technical connection.

[0094] Clinical risk constraints are loaded into the initial treatment response assessment model to simulate real-time risk changes, generating a first treatment response assessment model. Clinical risk constraints encompass various types, including drug toxicity risk, organ dysfunction risk, and metabolic abnormality risk. These constraints can be loaded using pre-set risk thresholds, risk probability distributions, or risk association rules. For example, to address the hepatotoxicity risk of a chemotherapy drug, a serum alanine aminotransferase (ALT) level exceeding 3 times the upper limit of normal (ULN) can be set as the risk trigger threshold. When the model-predicted ALT value reaches this threshold, a risk warning is automatically triggered and further increases in drug dosage are restricted. Risk constraints can be loaded using a penalty function. Specifically, a risk penalty term is introduced into the model's objective function. When the predicted result reaches the risk threshold, the corresponding strategy is prioritized by increasing the penalty value. Furthermore, risk constraints can be dynamically adjusted based on individual patient characteristics (such as age and baseline liver and kidney function). For example, for elderly patients or those with hepatic insufficiency, the ALT risk threshold can be set at 2 times the ULN to reflect personalized risk management needs.

[0095] After generating the first treatment response assessment model, efficacy evaluation indicators need to be loaded to generate a comprehensive treatment response assessment model. These efficacy evaluation indicators include core parameters such as the maximum dose tolerance threshold, the minimum efficacy response threshold, and the monitoring parameter redundancy range. The maximum dose tolerance threshold is determined based on drug clinical trial data or pharmacology studies and represents the maximum drug dose a patient can tolerate during treatment. Exceeding this threshold may lead to unacceptable toxicity. For example, the maximum dose tolerance threshold for a certain antihypertensive drug is set at 160 mg per day. When simulating dosing strategies, the model automatically excludes regimens with doses exceeding this threshold. The minimum efficacy response threshold measures whether a treatment has achieved its primary efficacy goal. For example, in anticancer treatment, a 10% tumor volume reduction might be set as the minimum efficacy response threshold. If the model predicts that a treatment regimen will not reduce tumor volume above this threshold, the regimen is deemed ineffective. The monitoring parameter redundancy range defines the allowable fluctuation range of the monitoring indicator. For example, the monitoring redundancy range for blood oxygen saturation is set at 90%-100%. If the model predicts values ​​outside this range, it indicates that the monitoring frequency or treatment regimen needs to be adjusted.

[0096] Efficacy evaluation indicators are loaded through model parameter configuration. Specifically, various indicators can be embedded into the model structure as constraints or optimization goals. For example, an efficacy evaluation branch is added to the model's output layer. This branch calculates the efficacy indicator value of the current plan based on the input treatment parameters and physiological parameters, compares it with the preset threshold, and outputs the probability of achieving the efficacy standard or the efficacy score. By gradually superimposing risk constraints and efficacy evaluation indicators, the initial model is gradually upgraded to a comprehensive treatment response evaluation model that can comprehensively evaluate the safety and effectiveness of treatment, providing an evaluation framework closer to clinical practice for subsequent dynamic simulations.

[0097] Based on the comprehensive treatment response evaluation model and execution state parameters, the core of dynamic simulation of treatment plans lies in constructing a multi-stage strategy planning equation and solving it using the Monte Carlo tree search method. The multi-stage strategy planning equation contains multiple sub-equations such as the stage coverage integrity equation, the response consistency equation, and the safety equation. Each sub-equation constrains the rationality of the treatment plan from different dimensions. The stage coverage integrity equation ensures that the treatment plan has clear strategic coverage for each stage on the timeline to avoid treatment gaps. For example, for a three-stage treatment plan divided into an induction phase, a consolidation phase, and a maintenance phase, the equation requires that each phase must define specific dosage, intervention time, and monitoring frequency, and the strategies of adjacent phases must be continuous.

[0098] The response consistency equation is used to measure whether the patient's physiological responses at different stages of treatment conform to the expected rules and avoid contradictory or abnormal response patterns. For example, in antihypertensive treatment, if an escalating dose strategy is used during the induction phase, the response consistency equation requires that the patient's blood pressure values ​​should show a gradual downward trend as the dose increases. If the model predicts that a certain regimen will increase blood pressure when the dose increases, then the regimen is judged to violate the response consistency principle and is excluded. The safety equation converts clinical risk constraints into mathematical expressions to ensure that all strategies in the simulation process do not exceed the preset risk threshold, such as the drug dose does not exceed the maximum tolerance threshold and the physiological indicators do not reach the dangerous range.

[0099] The Monte Carlo Tree Search (MCTS) method uses iterative simulation and statistical analysis to search for optimal treatment strategies within the constraints of multi-stage strategy planning equations. The specific implementation steps include: first, selecting an underexplored node in the current tree as the root node, representing the state of the current treatment stage; then, generating several candidate strategies (actions) through random sampling, and using a comprehensive treatment response evaluation model to predict the next stage state that each strategy may lead to; then, evaluating the prediction results using the stage coverage completeness equation, response consistency equation, and safety equation, calculating the score of each strategy (e.g., efficacy score minus risk score); finally, the highest-scoring strategy and its corresponding state node are added to the search tree, and the above process is repeated until a preset number of simulations or convergence criteria are reached. In this way, MCTS can efficiently search for treatment options with high comprehensive benefits in an exponentially growing strategy space, generating strategy control simulation data such as patient response feature distributions, risk clustering regions, safe treatment boundary sets, and strategy adjustment priority sequences.

[0100] The generation of a safe treatment boundary set is one of the key outputs of the dynamic simulation process. The specific steps are as follows: First, based on the risk clustering area (i.e., the stage or strategy combination with higher risk score during the simulation process), the clinical risk coverage density of each treatment stage is calculated. The risk coverage density is measured by the number of risk events occurring per unit time or the cumulative value of the risk score. For example, within a certain treatment stage, the average risk score per hour is the risk coverage density of that stage. Then, a preset threshold (such as an average risk score of 0.5) is set to identify the stages in the comprehensive treatment response assessment model where the risk coverage density is not greater than the threshold. These stages constitute the safe treatment boundary set due to their low risk. For example, when simulating a treatment regimen for a certain antibiotic, if the risk coverage density on days 3-5 of treatment is found to be lower than the preset threshold, while the risk on days 1-2 and days 6-7 is higher, then days 3-5 are designated as the safe treatment boundary, indicating that the risk of adjusting the treatment strategy within this stage is low and can be used as a priority strategy adjustment window.

[0101] The distribution of patient response characteristics in the strategy regulation simulation data is generated by statistically analyzing the change trajectories of patient physiological indicators corresponding to each strategy in the Monte Carlo simulation, reflecting the response patterns and their probability distributions that may be caused by different treatment actions. For example, in the simulation of the dosage adjustment strategy, the distribution of patient response characteristics can show the magnitude and probability of blood pressure reduction under different doses. The risk clustering area groups high-risk strategies through clustering algorithms (such as K-means and DBSCAN) to identify treatment stages or strategy combinations with similar risk characteristics, so that clinical personnel can formulate targeted risk prevention and control measures. The strategy adjustment priority sequence is sorted according to the comprehensive score of the strategy (the difference between the efficacy score and the risk score). The higher the score, the higher the priority of the strategy, indicating that it should be given priority in actual treatment.

[0102] The entire implementation process achieves dynamic simulation of treatment plans and generation of strategic control data by gradually loading constraints, constructing multidimensional planning equations, and employing intelligent search algorithms. This process not only considers the balance between treatment safety and effectiveness but also provides a concrete basis for strategic adjustments in clinical decision-making through quantitative analysis. This demonstrates the advantages of reinforcement learning in handling complex constraints and multi-objective optimization problems in clinical trial data analysis.

[0103] Example 3:

[0104] In the process of building a reinforcement learning decision model and performing strategy optimization training through strategy-controlled simulated data to generate a treatment path correction model and then generate initial treatment parameters, it is necessary to complete key steps such as model architecture design, data-driven training, real-time parameter prediction and parameter generation algorithm implementation in sequence. Each step forms a complete intelligent decision-making chain through technical integration and logical linkage.

[0105] First, a reinforcement learning decision model is constructed based on the dynamic response feature model. The dynamic response feature model establishes a mapping relationship between physiological state and treatment response through hierarchical feature extraction of temporal physiological parameters. The reinforcement learning decision model further incorporates decision logic, forming a complete closed loop from state perception to action selection. This model typically employs an "actor-critic" architecture: the "actor network" (ActorNetwork) is responsible for outputting a specific treatment strategy (such as dosage, intervention time, and other actions) based on the current state. Its input is the state feature vector output by the dynamic response feature model, and its output is a probability distribution over the action space. The "critic network" (CriticNetwork) evaluates the strategy selected by the "actor", taking as input the combination of state features and action vectors and outputting an estimated value for the strategy (such as an expected efficacy score or risk-adjusted return). This architecture enables a feedback loop between the "actor's" strategy exploration and the "critic's" value assessment, enabling continuous strategy optimization. For example, in a hypertension treatment scenario, the "actor" network outputs the probability distribution of increasing, decreasing, or maintaining the current dose based on the patient's current blood pressure fluctuation characteristics and treatment stage. The "critic" network scores the expected effect of each action based on risk constraints and efficacy indicators, guiding the "actor" network to adjust strategy parameters.

[0106] A reinforcement learning decision model is trained and validated through policy iteration using policy-controlled simulated data. This simulated data contains multi-dimensional information, including patient response feature distributions, risk clusters, safe treatment boundaries, and priority sequences for policy adjustments. This data is generated through dynamic simulations of previous treatment plans and accurately reflects the potential effectiveness of different treatment strategies in clinical scenarios. The training process utilizes an offline policy learning approach, which uses historical simulated data to train the model without requiring real-time interaction with the real environment, thereby reducing the risk and cost of clinical trials. The specific steps are as follows: First, state features, action vectors, next-state features, and reward signals (e.g., efficacy score minus risk score) from the simulated data are organized into training samples. Then, optimization algorithms such as stochastic gradient descent (SGD) are used to update the parameters of the actor and critic networks, ensuring that the policy output by the actor network achieves higher cumulative rewards under the critic's evaluation. Finally, cross-validation is used to partition the simulated data into training and validation sets. The validation set is periodically used during training to assess the model's generalization ability and avoid overfitting. For example, when training a decision-making model for a certain anti-tumor drug, iterative training is performed using a data set containing 1,000 simulated treatment pathways. After every 100 iterations, the model is verified using the reserved 200 pathways to ensure that the model can maintain stable evaluation capabilities in unseen strategy scenarios.

[0107] After training, real-time execution state parameters are input into the treatment pathway correction model to predict the safe treatment boundary set. These real-time execution state parameters include the actual dosage, intervention time points, measured values ​​of monitoring parameters, and treatment plan execution deviations during the current treatment phase. These parameters are collected in real time by the clinical monitoring system and synchronized to the model input. Based on a trained reinforcement learning decision model, the treatment pathway correction model analyzes the real-time state. First, a dynamic response feature model extracts features from the current physiological parameters to generate a feature vector for the current state. Then, the "actor" network outputs possible adjustment strategies based on this feature vector. The "critic" network evaluates each strategy and calculates its corresponding risk coverage density. Finally, treatment phases with a risk coverage density below a preset threshold are selected to form a predicted safe treatment boundary set. For example, if a patient's measured blood pressure on the fourth day of treatment is above the target range and the current dose is approaching the maximum tolerated threshold, the model, through real-time state analysis, predicts that if the current dose is maintained and the monitoring frequency is increased to every 30 minutes over the next 12 hours, the risk coverage density of the treatment phase will fall below the threshold, thus demarcating that period as a safe treatment boundary.

[0108] To generate initial treatment parameters based on the predicted safe treatment boundary set, a series of algorithms are needed to map boundary features to specific parameters. First, the treatment stage with the lowest risk coverage density in the safe treatment boundary set is extracted, and its starting time point is used as the initial intervention time node. The stage with the lowest risk coverage density means that the safety of treatment adjustment at this stage is the highest. For example, in the simulation data, the average risk score of a certain stage is only 0.2 (out of a full score of 1), which is significantly lower than the 0.6-0.8 of other stages. The starting time of this stage (such as 8 am on the fifth day of treatment) is determined as the initial intervention time node.

[0109] Next, a feasible connectivity structure for the initial dosing sequence is fitted based on the temporal correlation of the safe treatment boundary set. This temporal correlation is determined by analyzing the strategic dependencies between adjacent treatment phases. For example, if a dose-escalation strategy is employed in the phase prior to the safe treatment boundary, and the therapeutic response during that phase is stable, the feasible connectivity structure is likely to continue the escalating trend or transition to dose maintenance. A specific fitting method can employ an autoregressive model (AR) or long short-term memory (LSTM) network from time series analysis to predict the appropriate dose range for the next phase based on the historical dose sequence and physiological response data. For example, if the doses administered on the first three days are 50 mg, 75 mg, and 100 mg, respectively, and patient response characteristics show a linear increase in efficacy with increasing dose, the LSTM model fitting yields a feasible dose range for day 4 of 100-125 mg. Combined with the risk characteristics of the safe treatment boundary, the initial dosing sequence is ultimately determined to be 100 mg (day 4), 110 mg (day 5), and 120 mg (day 6).

[0110] Finally, the patient response gradient direction for the safe treatment boundary set is calculated and normalized to form a dose adjustment reference vector. The patient response gradient direction is determined by solving the partial derivative of the efficacy assessment indicator with respect to the administered dose, reflecting the sensitivity between dose changes and improved efficacy. For example, if the partial derivative of the efficacy assessment indicator (such as tumor shrinkage rate) with respect to dose is 0.02 / 10mg, the gradient direction is positive, indicating that increasing the dose can improve efficacy. If the partial derivative is -0.01 / 10mg, the gradient direction is negative, suggesting that the dose should be reduced to avoid a decrease in efficacy. Through normalization (e.g., scaling the modulus of the gradient vector to 1), the dose adjustment reference vector is obtained. The direction and magnitude of this vector directly determine the adjustment trend of the initial dose sequence. For example, if the normalized gradient vector is (0.8, 0.6), the dose adjustment direction is along the direction indicated by this vector, which simultaneously considers the combined effects of dose increases and improved efficacy, generating an initial dose sequence direction that conforms to the individual patient's response characteristics.

[0111] During data processing, the policy-controlled simulation data must be reconstructed to improve model training efficiency. Specifically, an initial multi-stage treatment input tensor is constructed based on the distribution of patient response characteristics, risk clustering regions, and a set of safe treatment boundaries. The dimensions of this tensor include time phase, physiological indicator dimension, and feature type (such as original value, rate of change, and volatility). For example, a tensor containing 10 treatment phases, 5 physiological indicators, and 3 feature types can be represented as a three-dimensional structure of [10, 5, 3]. The initial tensor is then normalized (e.g., using Z-score normalization) and feature augmented (e.g., adding time codes and phase labels) to generate the final multi-stage treatment input tensor. This enhances the model's ability to capture temporal features and phase differences. Furthermore, a treatment pathway correction label tensor is constructed based on the policy adjustment priority sequence. This tensor is a label matrix of the same dimensions as the input tensor, with each element corresponding to the desired policy adjustment direction or magnitude for the input feature. For example, higher-priority phases are assigned a larger adjustment weight in the label tensor, guiding the model to focus on key treatment nodes. Finally, the input tensor and the label tensor are combined to form a policy training sample set, which is used for supervised pre-training of the reinforcement learning decision model to accelerate the model convergence process.

[0112] The entire implementation process utilizes deep neural network architecture design, multi-dimensional data-driven training, and mathematical mapping algorithms to achieve intelligent decision-making from physiological state perception to treatment parameter generation. This process not only leverages the adaptive optimization capabilities of reinforcement learning to handle the complexity and uncertainty of clinical data, but also ensures the safety of treatment plans through real-time parameter feedback and safety margin prediction, providing a scientific and practical technical solution for generating personalized treatment pathways in clinical trials.

[0113] Example 4:

[0114] In the process of obtaining dynamic strategy adjustment information based on initial treatment parameters, a comprehensive treatment response evaluation model, and execution status parameters, key steps such as model mapping, parameter matching, strategy node updating, and dynamic solution are required to form a closed-loop control mechanism from static initial parameters to dynamic strategy adjustment. The following is a detailed description of the specific implementation method:

[0115] First, the initial intervention time node is mapped to the comprehensive treatment response evaluation model to achieve alignment between the treatment stage and the model evaluation space. As the key time anchor point of the treatment plan, the initial intervention time node needs to be accurately located in the model to determine its corresponding physiological state and strategy space. For example, if the initial intervention time node is set to 10 a.m. on the third day of treatment, the time point needs to be mapped to the time axis of the model through a temporal alignment algorithm, and the physiological parameter sequence 72 hours before this time point needs to be extracted as the model input to evaluate the treatment response baseline under the current state. The stage division of the treatment plan needs to be considered during the mapping process, such as dividing the intervention time node into specific sub-stages of the induction period, adjustment period, or maintenance period, so that the model can apply corresponding evaluation rules based on the stage characteristics.

[0116] After completing the time node mapping, it is necessary to match the initial dosing sequence with the initial monitoring feedback threshold. The initial dosing sequence is usually a preset multi-stage dosing regimen, such as 50 mg on the first day, 75 mg on the second day, and 100 mg on the third day. Each dose value in the sequence needs to be matched with the dose-response curve in the model to determine the expected range of changes in physiological indicators corresponding to each dose. For example, the comprehensive treatment response evaluation model predicts that a 100 mg dose may reduce the target physiological indicator (such as blood pressure) by 10-15 mmHg on the third day. If the actual monitoring value exceeds this range, the dose adjustment mechanism is triggered. The initial monitoring feedback threshold involves the warning range of each physiological indicator. For example, an alarm is triggered when the heart rate exceeds 120 beats / minute or is lower than 50 beats / minute. These threshold parameters need to be input into the safety constraint module of the model as one of the triggering conditions for dynamic strategy adjustment.

[0117] Updating the strategy node density in the adjustment area is a key step in optimizing the accuracy of model evaluation. The strategy node density reflects the distribution density of adjustable strategies in the treatment plan. A higher node density is usually set in the early stages of treatment or at higher risk stages to reserve more space for strategy adjustment. For example, in the first three days of treatment (adjustment area), a strategy node is set every 12 hours, while a node is set every 24 hours during the stable period. The updating process is achieved by adding or deleting nodes in the time-strategy two-dimensional space of the model. The node density is calculated as: number of nodes / unit time (such as number of nodes / day). By dynamically adjusting this density value, the model can more accurately capture changes in physiological responses within the adjustment area. For example, if physiological indicators are found to fluctuate more on the second day of treatment, the strategy node density on that day can be increased from 2 / day to 4 / day to evaluate the effectiveness of the strategy more frequently.

[0118] When updating the comprehensive treatment response evaluation model, the strategy node density, initial parameter matching results, and execution state parameters need to be integrated into the model's state space and action space. Specifically, the state space adds a new feature dimension of strategy node density, and the action space expands adjustment actions related to node density (such as adding nodes and deleting nodes). The state-action mapping relationship is adapted to the new parameter configuration through model training. For example, after receiving a signal that the strategy node density has increased, the model will automatically activate the action option of "evaluating the strategy every 6 hours" in the action space and adjust the corresponding value function to reflect the costs and benefits of high-frequency evaluation.

[0119] Defining dynamic strategy trigger conditions, dose reconstruction rules, and feedback threshold increments is the core of generating a dynamic strategy adjustment model. Dynamic strategy trigger conditions are based on preset logical rules. For example, when the completion rate of the current treatment phase (the number of treatment steps actually performed / the number of planned steps) is not greater than the preset completion rate threshold (such as 80%), the feedback threshold update is triggered; when the physiological indicators exceed the safe treatment boundary set for two consecutive monitoring cycles, the dose reconstruction is triggered. The dose reconstruction rules clearly define the direction and amplitude of the dose adjustment. For example, based on the patient response gradient direction (the dose-response sensitivity direction obtained through previous simulation calculations), the initial administration dose sequence is adjusted. If the gradient direction is positive (increase in dose corresponds to improved efficacy), the adjustment amplitude is 10% of the current dose each time until the maximum tolerance threshold is reached or the efficacy is achieved. The feedback threshold increment is in a piecewise linear relationship with the current efficacy stability evaluation index. For example, when the efficacy stability index (such as the standard deviation of physiological index fluctuation) is less than 0.5, the threshold increment is 0; when the index is between 0.5-1.0, the increment is linearly positively correlated with the index value, and the threshold increases by 5% for every increase of 0.1; when the index is greater than 1.0, the threshold reset mechanism is triggered, and the threshold is adjusted to 120% of the initial value to expand the monitoring range.

[0120] Based on the dynamic policy adjustment model, the multi-stage policy planning equation is iteratively solved using the Monte Carlo tree search method. The core process of the Monte Carlo tree search includes four stages: selection, expansion, simulation, and backpropagation. In the selection stage, the algorithm starts from the current policy node and selects underexplored or high-value child nodes according to the upper confidence bound (UCB) formula. The formula is:

[0121]

[0122] Where Q is the average node reward, c is the exploration coefficient, N is the number of parent node visits, and n is the number of child node visits. The search direction is determined by balancing exploration and utilization. In the expansion phase, a new strategy branch is generated under the selected node, corresponding to different dose adjustment or monitoring frequency change actions. In the simulation phase, a random strategy is used to quickly traverse the subsequent treatment phases and calculate the cumulative reward of each branch (such as the sum of the efficacy score minus the sum of the risk score). In the backtracking phase, the simulation results are fed back to each node in the tree structure, the number of node visits and the average reward value are updated, and the strategy selection is gradually optimized.

[0123] When the trigger conditions for a dynamic strategy are met, the following actions are performed: First, the dose offset is updated—the difference between the actual administered dose and the initial dose sequence. For example, if the initial dose is 100 mg and is updated to 110 mg due to a trigger condition, the offset is +10 mg. Second, the treatment stage completion rate is updated. If a stage plans to perform five actions and only four are completed, the completion rate is 80%. Finally, the efficacy stability assessment metric is updated, calculated by calculating the standard deviation of physiological indicators over the last three monitoring cycles. After the parameter update, the multi-stage strategy planning equation is re-solved to assess the impact of the new strategy on subsequent treatment stages. For example, after the dose offset is updated to +10 mg, the model recalculates the risk coverage density and efficacy target probability for each stage at that dose. If the risk of the next stage exceeds the threshold, a compensatory strategy of "increasing monitoring frequency" or "premature intervention" is automatically generated.

[0124] The simulation termination conditions usually include reaching the preset treatment cycle endpoint, the cumulative rewards of all strategy branches converging to a stable value, or key physiological indicators reaching a dangerous threshold (such as blood pressure below the shock threshold). When the termination conditions are met, the Monte Carlo tree search stops iterating and outputs dynamic strategy adjustment information, including: dose offset sequence (such as +10mg on the 3rd day, -5mg on the 5th day), treatment stage completion rate curve (a time series reflecting the execution status of each stage), and a trend chart of the efficacy stability evaluation indicator (showing changes in the degree of fluctuation). This information is output in the form of structured data and can be directly used in the clinical decision support system to prompt medical staff to the execution deviation and adjustment direction of the current treatment plan.

[0125] During the implementation process, special attention should be paid to the compatibility of the dynamic strategy adjustment model and the comprehensive treatment response evaluation model. After each parameter update, it is necessary to verify whether the output of the model complies with the clinical logic constraints. For example, the dose offset must not exceed ±20% of the maximum tolerance threshold, and the monitoring frequency adjustment range must be within the range supported by the device (such as a minimum of 15 minutes / time and a maximum of 12 hours / time). In addition, the strategy adjustment priority sequence must be consistent with the strategy control simulation data generated by the previous simulation. For example, strategies with higher priority in risk clustering areas (such as emergency dose adjustment) should be triggered first to ensure treatment safety.

[0126] The entire dynamic strategy adjustment process achieves adaptive adjustment of treatment plans through model-driven parameter mapping, conditional triggering of the rule engine, and search optimization of intelligent algorithms. This process not only promptly responds to real-time changes in the patient's physiological state but also ensures the scientific and rationality of the adjustments through a quantitative strategy evaluation system, providing technical support for achieving dynamic efficacy matching goals in clinical trials.

[0127] Example 5:

[0128] In the process of building a multi-stage optimization model and iteratively optimizing planning parameters to obtain the optimal treatment path parameters, key links such as model architecture definition, optimization algorithm selection, constraint processing, and strategy space search are required to form a transformation chain from multi-objective optimization to clinically feasible solutions. The following is a detailed description of the specific implementation method:

[0129] First, a multi-stage optimization model is constructed to clarify the optimization variables, optimization objectives, and constraints. The optimization variables include parameters directly related to treatment path planning, such as the strategy node distribution density, dose adjustment threshold, and monitoring sampling frequency. The strategy node distribution density represents the number of strategy evaluation points per unit time. For example, in the initial stage of treatment, it is set to one node every 12 hours, and in the stable period, it is set to one node every 24 hours. Its value range is [1,4] nodes / day; the dose adjustment threshold is defined as the maximum amplitude of a single dose adjustment (such as ±15% of the current dose), and the specific value is determined according to the drug characteristics and patient tolerance; the monitoring sampling frequency refers to the collection interval of physiological indicators (such as 15 minutes, 30 minutes, 1 hour, etc.), which needs to match the technical parameters of the monitoring equipment.

[0130] The optimization objectives include maximizing the stage coverage efficiency and minimizing the treatment risk increment. The stage coverage efficiency is evaluated by the strategy completeness of the treatment plan at each stage. The calculation formula is: the number of strategy nodes actually covered / the number of strategy nodes planned to be covered. The goal is to make this value close to 1 to ensure that no treatment stage is missed. The treatment risk increment is defined as the difference in risk scores between adjacent treatment stages. The goal is to make it as small as possible to avoid a sudden increase in risk due to strategy mutations. Constraints include the upper limit of the patient's physiological tolerance (such as the maximum dose of the drug, the indicator danger threshold) and the accuracy threshold of the monitoring equipment (such as the blood pressure monitoring error does not exceed ±2mmHg). The values ​​of all optimization variables must be within the range allowed by the constraints.

[0131] The multi-stage optimization model is preliminarily solved using a dynamic programming algorithm to generate an initial set of optimization strategies. The dynamic programming algorithm divides the entire treatment cycle into multiple stages, each corresponding to a state space (such as current dose, monitoring frequency), and recursively calculates the optimal strategy for each stage to construct the optimal path from the initial state to the terminal state. The specific steps are as follows: First, define the initial state (such as the baseline dose and routine monitoring frequency on the first day of treatment), then for each stage, traverse all possible state transitions (such as a 10% increase in dose and an increase in monitoring frequency to once every 30 minutes), calculate the stage coverage efficiency and risk increment after the transition, and select the state that optimizes the cumulative objective function value as the optimal solution for that stage. In this way, the algorithm can generate a series of initial optimization strategies that meet the constraints, such as a strategy set containing 10 different combinations of dose adjustment paths and monitoring frequencies.

[0132] Based on the initial set of optimized strategies, a policy gradient optimization algorithm is used for global optimization to generate the optimal treatment pathway parameters. The policy gradient algorithm directly optimizes the policy parameters, guiding the parameter updates in a direction that improves the objective function value by calculating the gradient of the objective function with respect to the policy parameters. In its implementation, each strategy is encoded as a parameter vector (e.g., [node density, dose threshold, sampling frequency]). A set of parameter samples is generated through random initialization. Monte Carlo simulation is then used to estimate the expected reward for each sample (i.e., stage coverage efficiency minus risk increment). The gradient of the reward with respect to the parameters is calculated, and the parameters are updated. This process is iterative until the objective function value converges or a preset maximum number of iterations (e.g., 1000) is reached. For example, when optimizing the insulin dosing regimen for a diabetic patient, a population of 200 policy parameter samples was iteratively optimized using the policy gradient algorithm, ultimately identifying the optimal strategy with a stage coverage efficiency of 95% and a risk increment of less than 0.1.

[0133] During the optimization process, constraints must also be addressed to ensure the clinical feasibility of the strategy, specifically by constructing a set of constrained relaxation parameters. The set of constrained relaxation parameters includes relaxation coefficients for the patient's physiological tolerance upper limit and the monitoring device's accuracy threshold. For example, the relaxation coefficient for the maximum drug dose is set to 0.9-1.1 (i.e., the dose is allowed to be fine-tuned within the range of 90%-110%), and the relaxation coefficient for the monitoring device's accuracy threshold is set to 1.0-1.2 (i.e., the error is allowed to fluctuate within the range of 10%-20%). The value of the relaxation coefficient needs to be determined through clinical expert evaluation to ensure that adjusting the parameters within the relaxation range will not significantly increase treatment risks or reduce monitoring quality.

[0134] The policy parameters in the initial optimization policy set are dynamically matched with the constraint relaxation parameter set to generate a subset of feasible policies. The matching process uses heuristic rules. For example, for the dose adjustment strategy, if the current dose exceeds 105% of the maximum tolerance threshold, it is adjusted to within the threshold by multiplying it by a relaxation coefficient of 0.95; for the monitoring sampling frequency, if the device cannot support a frequency lower than 15 minutes / time, the frequency is automatically adjusted to 15 minutes / time. In this way, strategies that violate hard constraints are filtered out, and feasible strategies that meet clinical reality are retained. For example, after matching the 30 strategies in the initial set, 25 feasible strategies that meet the requirements of physiological tolerance and device accuracy were screened out.

[0135] The policy gradient optimization algorithm is used to prioritize the subset of feasible strategies and generate the optimal treatment pathway parameters. The sorting is based on the comprehensive score, calculated as follows: comprehensive score = stage coverage efficiency × weight 1 - treatment risk increment × weight 2, where weight 1 and weight 2 are set according to clinical needs (such as weight 1 = 0.6, weight 2 = 0.4). The algorithm sorts the comprehensive scores of the feasible strategies in descending order, and the strategy with the highest score is determined as the optimal treatment pathway parameter. For example, in the subset of feasible strategies, the comprehensive score of strategy A is 0.85, strategy B is 0.82, and strategy C is 0.88. Then strategy C is selected as the optimal parameter, and its corresponding node density is 2 nodes / day, the dose adjustment threshold is ±12%, and the monitoring sampling frequency is 30 minutes / time.

[0136] The construction of a multi-stage optimization model also needs to consider the time dependence of the treatment cycle, for example, by applying different optimization weights to different treatment stages. During the induction phase, due to the need to rapidly explore effective doses, the weight of stage coverage efficiency can be increased to 0.7, while the weight of incremental risk can be reduced to 0.3. During the maintenance phase, to ensure treatment safety, the weight of incremental risk can be increased to 0.6, while the weight of stage coverage efficiency can be reduced to 0.4. By dynamically adjusting weights, the model can adapt to changing clinical goals at different stages of the treatment process.

[0137] When implementing the policy gradient optimization algorithm, it's important to pay attention to the learning rate setting and decay strategy. The initial learning rate is typically set to 0.1, decaying to 0.9 of the current value after every 100 iterations to balance rapid exploration in the early stages with refined convergence later. Furthermore, a momentum factor (e.g., 0.9) can be introduced to accelerate the gradient descent process and avoid getting stuck in local optima. For example, when optimizing a model with 10 treatment stages, by setting the learning rate decay and momentum factor, the algorithm converged to a stable solution within 500 iterations.

[0138] The entire multi-stage optimization process transforms clinical treatment objectives into computable optimization problems through mathematical modeling. Dynamic programming is combined with a policy gradient algorithm to efficiently search the complex policy space. Constraint relaxation and feasible policy screening are used to ensure the clinical applicability of the optimization results. This method not only handles multi-objective, multi-constraint optimization problems but also adapts to individual patient differences through iterative training, providing a systematic solution for developing personalized treatment pathways in clinical trials.

[0139] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0140] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A clinical experimental data analysis method based on reinforcement learning, characterized in that: The following steps are involved: Obtaining the time series physiological parameter data set and treatment plan execution status parameters of clinical trial subjects; Perform hierarchical feature extraction on temporal physiological parameters and construct a dynamic response feature model. Combined with the treatment plan execution state parameters, the dynamic response feature model is subjected to reinforcement learning strategy analysis to generate an initial treatment response evaluation model. The initial treatment response evaluation model is loaded with clinical risk constraints and efficacy evaluation indicators to generate a comprehensive treatment response evaluation model. The treatment plan is dynamically simulated in combination with the execution state parameters to obtain strategy control simulation data. Construct a reinforcement learning decision model and perform strategy optimization training through strategy-controlled simulation data to generate a treatment pathway correction model, thereby generating initial treatment parameters. The initial treatment parameters include at least the initial intervention time node, the initial dosing sequence, and the initial monitoring feedback threshold. Based on the initial treatment parameters, the comprehensive treatment response evaluation model and the execution status parameters, dynamic strategy adjustment information is obtained; Combining dynamic strategy adjustment information with treatment cycle planning parameters, a multi-stage optimization model is constructed and the planning parameters are iteratively optimized to obtain the optimal treatment path parameters; Based on the real-time dynamic response characteristic model, the current treatment cycle parameters are inferred to generate the probability of achieving the efficacy target in the current stage. Combined with the optimal treatment path parameters in the current stage and the actual physiological parameter fluctuations, the treatment plan execution strategy is adjusted to achieve the dynamic efficacy matching goal.

2. The clinical trial data analysis method based on reinforcement learning according to claim 1, characterized in that: The strategy control simulation data at least includes patient response characteristic distribution, risk clustering area, safe treatment boundary set and strategy adjustment priority sequence; The dynamic strategy adjustment information includes at least dose offset, treatment phase completion rate and efficacy stability evaluation index.

3. The clinical trial data analysis method based on reinforcement learning according to claim 1, characterized in that: The method of extracting hierarchical features of temporal physiological parameters and constructing a dynamic response feature model, performing reinforcement learning strategy analysis on the dynamic response feature model in combination with treatment plan execution state parameters, and generating an initial treatment response evaluation model includes the following steps: Perform multi-dimensional signal decomposition on time-series physiological parameters to obtain data sequences with individualized response characteristics of patients; Preprocessing the data sequence to obtain standardized feature fusion data, wherein the preprocessing includes one or more of outlier filtering, time series alignment, feature dimensionality reduction, stage division, and data normalization; Based on standardized feature fusion data and combined with reinforcement learning strategy network, a dynamic response feature model is constructed; A strategy exploration analysis is performed on the dynamic response characteristic model to generate an initial treatment response evaluation model, wherein the strategy exploration analysis at least includes state space definition and action space mapping, the therapeutic effect response level is divided by the state space definition, and corresponding strategy exploration weights are assigned to the parameters of each stage.

4. The clinical trial data analysis method based on reinforcement learning according to claim 1, characterized in that: The method of loading clinical risk constraints and efficacy evaluation indicators into the initial treatment response evaluation model to generate a comprehensive treatment response evaluation model, dynamically simulating the treatment plan in combination with the execution state parameters, and obtaining strategy control simulation data includes the following steps: Loading clinical risk constraints to the initial treatment response assessment model to simulate real-time risk changes, thereby generating a first treatment response assessment model; Loading efficacy evaluation indicators into the first treatment response evaluation model to generate a comprehensive treatment response evaluation model, wherein the efficacy evaluation indicators include a maximum dose tolerance threshold, a minimum efficacy response threshold, and a monitoring parameter redundancy range; Based on the comprehensive treatment response evaluation model and execution status parameters, dynamic simulation of treatment plans is performed. The specific process includes: Constructing a multi-stage strategy planning equation, the multi-stage strategy planning equation includes at least a stage coverage integrity equation, a response consistency equation, and a safety equation. In combination with a comprehensive treatment response evaluation model, the multi-stage strategy planning equation is numerically solved using a Monte Carlo tree search method to obtain a patient response characteristic distribution, risk clustering areas, a safe treatment boundary set, and a strategy adjustment priority sequence; The generation of the safe treatment boundary set includes the following steps: Based on the risk clustering areas, the clinical risk coverage density of each treatment stage is calculated; The stages in the comprehensive treatment response assessment model where the risk coverage density is no greater than the preset threshold are identified, and a set of safe treatment boundaries is generated.

5. The clinical trial data analysis method based on reinforcement learning according to claim 1, characterized in that: The construction of the reinforcement learning decision model and the strategy optimization training by using the strategy-controlled simulation data to generate the treatment path correction model and then generate the initial treatment parameters include the following steps: Construct a reinforcement learning decision model based on the dynamic response feature model; Through the strategy control simulation data, the reinforcement learning decision model is trained and verified through strategy iteration to generate a treatment path correction model; Input the real-time execution state parameters into the treatment path correction model to predict the set of safe treatment boundaries; Based on the predicted safe treatment boundary set, initial treatment parameters are generated, wherein the initial treatment parameters at least include an initial intervention time node, an initial drug administration sequence, and an initial monitoring feedback threshold.

6. The clinical trial data analysis method based on reinforcement learning according to claim 5, characterized in that: It also includes data reconstruction of strategy control simulation data, specifically: Construct an initial multi-stage treatment input tensor based on the distribution of patient response characteristics, risk cluster regions, and a set of safe treatment boundaries; Normalize and enhance the features of the initial multi-stage treatment input tensor to generate the final multi-stage treatment input tensor; Adjust the priority sequence based on the strategy and construct the treatment path correction label tensor; The final multi-stage treatment input tensor and the treatment path correction label tensor are combined to form a strategy training sample set.

7. The clinical trial data analysis method based on reinforcement learning according to claim 6, characterized in that: The method of generating initial treatment parameters based on the predicted safe treatment boundary set includes the following steps: Extract the treatment stage with the lowest risk coverage density from the predicted safe treatment boundary set and generate the initial intervention time node; According to the temporal correlation of the predicted safe treatment boundary set, the feasible connection structure of the initial administration dose sequence is fitted to generate the initial monitoring feedback threshold; The patient response gradient direction of the predicted safe treatment boundary set is calculated and normalized to a dose adjustment reference vector, which is the direction of the initial administration dose sequence.

8. The clinical trial data analysis method based on reinforcement learning according to claim 1, characterized in that: The method of obtaining dynamic strategy adjustment information based on the initial treatment parameters, the comprehensive treatment response evaluation model, and the execution status parameters includes the following steps: Map the initial intervention time node to the comprehensive treatment response evaluation model, match the initial administration dose sequence with the initial monitoring feedback threshold, and update the strategy node density in the adjustment area. Update the comprehensive treatment response evaluation model, define dynamic strategy trigger conditions, dose reconstruction rules, and feedback threshold increments, and generate a dynamic strategy adjustment model. Based on the dynamic policy adjustment model, the multi-stage policy planning equation is iteratively solved through the Monte Carlo tree search method to obtain dynamic policy adjustment information, including: When the dynamic strategy triggering conditions are met, the dose offset, treatment stage completion rate, and efficacy stability evaluation indicators are updated, and the multi-stage strategy planning equation is re-solved until the simulation termination conditions are met; The dynamic strategy triggering conditions include triggering the feedback threshold update when the completion rate of the current treatment stage is no greater than the preset completion rate threshold; the dose reconstruction rule includes adjusting the initial administration dose sequence based on the patient response gradient direction; the feedback threshold increment is in a piecewise linear relationship with the current efficacy stability evaluation index.

9. The clinical trial data analysis method based on reinforcement learning according to claim 1, characterized in that: The multi-stage optimization model is constructed and the planning parameters are iteratively optimized to obtain the optimal treatment path parameters, including the following steps: A multi-stage optimization model was constructed, in which the optimization variables included the strategy node distribution density, dose adjustment threshold, and monitoring sampling frequency. The optimization objectives included maximizing stage coverage efficiency and minimizing treatment risk increment. The constraints included the patient's physiological tolerance limit and the monitoring device accuracy threshold. Perform a preliminary solution to the multi-stage optimization model using a dynamic programming algorithm to generate an initial set of optimization strategies; Based on the initial optimization strategy set, the policy gradient optimization algorithm is used for global optimization to generate the optimal treatment path parameters.

10. The clinical trial data analysis method based on reinforcement learning according to claim 9, characterized in that: The multi-stage optimization model is constructed and the planning parameters are iteratively optimized to obtain the optimal treatment path parameters, and the following steps are also included: Construct a set of constraint relaxation parameters based on the patient's physiological tolerance upper limit and the monitoring equipment's accuracy threshold; Dynamically match each strategy parameter in the initial optimization strategy set with the constraint relaxation parameter set to generate a feasible strategy subset; The feasible strategy subsets are prioritized by the policy gradient optimization algorithm to generate the optimal treatment path parameters.

Citation Information

Cited By

  • Intelligent evaluation and optimization method and system for osteoarthritis treatment scheme

    CN121096664A

  • Intelligent evaluation and optimization method and system for osteoarthritis treatment plan

    CN121096664B

  • Gastric cancer immunotherapy curative effect prediction method based on multi-mode fusion

    CN121237451A

  • Automatic lung function examination data analysis and classification system based on big data

    CN121281864A

  • Intelligent decision-making system for CAP treatment of nephropathy skin complications based on multi-parameter dynamic monitoring

    CN121725984A