Historical data-based Saiyang capsule treatment effect prediction method

By constructing a method for predicting the treatment effect of Qiyang capsules based on historical data, and combining kernel function support vector regression and dynamic Bayesian network modules, the problem of accurately modeling the time-varying coupling relationship of multidimensional physiological parameters in the prediction of treatment effect of diabetic impotence patients was solved, thereby improving the prediction accuracy and model reliability.

CN121885186AInactive Publication Date: 2026-04-17BAODING SECOND CENT HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-06
Publication Date
2026-04-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies lack accurate modeling of the time-varying coupling relationship of multidimensional physiological parameters when predicting the treatment effect of diabetic erectile dysfunction patients, resulting in insufficient prediction accuracy.

Method used

A method for predicting the therapeutic effect of Qiyang capsules based on historical data was adopted. Baseline data of patients with impotence due to kidney deficiency and blood stasis in diabetes were obtained, and quality control and feature engineering were performed. A distributed drug efficacy synergistic prediction algorithm based on Lagrange dual decomposition was used, combined with a kernel function support vector regression module and a dynamic Bayesian network module, to construct a time-series coupled prediction model for therapeutic efficacy. An entropy increase monitoring model was used to continuously monitor the prediction error, thereby realizing the modeling of the nonlinear mapping and time-varying coupling relationship of physiological parameters.

Benefits of technology

It improves the predictive accuracy of treatment outcomes for diabetic erectile dysfunction patients, ensures the reliability and stability of prediction results, and avoids misdiagnosis or treatment errors caused by data drift in the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121885186A_ABST
    Figure CN121885186A_ABST
Patent Text Reader

Abstract

The invention provides a historical data-based Saiyang capsule treatment effect prediction method, and belongs to the technical field of Saiyang capsule treatment effect prediction.The method comprises the steps of establishing an original data set by collecting baseline data of a patient, and performing quality control and feature engineering processing on the data to obtain an optimized feature set; a Lagrange dual decomposition algorithm is adopted to carry out distributed processing, a CUDA flow is started on a GPU to execute a double-layer thread grid, a primary processing thread grid runs a kernel function support vector regression module to calculate a primary curative effect evaluation value, a deep processing thread grid runs a dynamic Bayesian network module to carry out time-varying coupling analysis, and a final curative effect prediction value is output. And after the model is deployed, an entropy increase monitoring algorithm is adopted to continuously monitor and predict error information entropy to trigger model retraining, so that the technical problem of insufficient prediction precision caused by lack of accurate modeling of a multi-dimensional physiological parameter time-varying coupling relationship when the treatment effect of the diabetic impotence patient is predicted is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of Qiyang Capsule therapeutic effect prediction technology, specifically, it relates to a method for predicting the therapeutic effect of Qiyang Capsule based on historical data. Background Technology

[0002] In evaluating the treatment efficacy of diabetic erectile dysfunction, traditional methods primarily rely on single-time-point physiological index measurements combined with statistical regression models for efficacy prediction. Existing prediction methods typically employ linear regression or simple machine learning algorithms, inputting indicators such as blood glucose concentration, hormone levels, and hemodynamic parameters as independent variables into the model, neglecting the complex nonlinear coupling relationships between these physiological parameters. Because the physiological state of diabetic patients exhibits dynamic changes, the interactions between different parameters evolve over time, making it difficult for static models to capture the time-varying dependence structure between parameters. In other words, existing technologies suffer from insufficient prediction accuracy due to a lack of accurate modeling of the time-varying coupling relationships of multidimensional physiological parameters when predicting the treatment efficacy of diabetic erectile dysfunction. Summary of the Invention

[0003] In view of this, the present invention provides a method for predicting the therapeutic effect of Qiyang capsules based on historical data, which can solve the technical problem in the prior art that the lack of accurate modeling of the time-varying coupling relationship of multidimensional physiological parameters in predicting the therapeutic effect of diabetic impotence patients leads to insufficient prediction accuracy.

[0004] This invention is implemented as follows: It provides a method for predicting the therapeutic effect of Qiyang Capsules based on historical data, comprising: acquiring baseline data of patients with kidney deficiency and blood stasis type diabetes and impotence to establish an original dataset; performing quality control processing on the original dataset to obtain a cleaned dataset; performing feature engineering processing on the cleaned dataset to obtain an optimized feature set; labeling the optimized feature set according to treatment groups and then distributing it to multiple computing nodes for distributed processing using a distributed pharmacodynamic synergistic prediction algorithm based on Lagrange dual decomposition to obtain node optimization results; using the CPU to read patient data from the node optimization results and transmit it to the global memory block of the GPU; starting CUDA streaming on the GPU to execute two-layer threaded grid processing; the primary processing threaded grid running the kernel function support vector regression module of the efficacy temporal coupling prediction model to calculate preliminary efficacy evaluation values ​​and store them in a shared memory block; the deep processing threaded grid running the dynamic Bayesian network module of the efficacy temporal coupling prediction model to perform time-varying coupling analysis and output the final efficacy prediction value; after deploying the efficacy temporal coupling prediction model, using an entropy increase monitoring model degradation detection algorithm based on the second law of thermodynamics to continuously monitor the prediction error information entropy and trigger model retraining when the entropy value monotonically increases beyond the degradation threshold.

[0005] The quality control process includes using an expectation-maximization multiple imputation framework combined with a random forest proximity matrix to fill missing values, and using a local anomaly detection method to identify outliers and then replacing them with robust regression.

[0006] Specifically, the expectation-maximization multiple imputation framework combines a random forest proximity matrix to fill missing values. It uses complete samples to train a random forest model to calculate the proximity matrix between samples. For missing values, it selects several complete samples with the highest proximity and fills them based on the weighted average of their corresponding feature values.

[0007] The local anomaly detection method involves determining the k nearest neighbors of each point in the dataset, calculating the reachability distance to the k nearest neighbor, calculating the local reachability density of the point as the inverse mean of the reachability distances of its k nearest neighbors, and the local anomaly factor being the average ratio of the local reachability density of each point in the neighborhood of the point to the local reachability density of the point itself.

[0008] The feature engineering process involves using the maximum correlation and minimum redundancy criterion of mutual information to screen a subset of features that are highly correlated with therapeutic efficacy and have low redundancy, and then projecting them onto a low-dimensional manifold through principal component analysis.

[0009] The maximum relevance and minimum redundancy criterion for mutual information requires that the sum of mutual information between the selected features and the therapeutic target variable be maximized and the sum of mutual information between the selected features be minimized.

[0010] The principal component analysis involves calculating the covariance matrix of the standardized feature matrix, solving for eigenvalues ​​and eigenvectors, and selecting principal components with a cumulative variance contribution rate of 85% to 95% based on the eigenvalues ​​sorted from largest to smallest.

[0011] The distributed pharmacodynamic synergistic prediction algorithm based on Lagrange dual decomposition distributes patient data to four computing nodes according to treatment groups. Each node independently solves local subproblems and then exchanges dual variables. The Lagrange multipliers are updated according to the subgradient method to adjust the local solution so that it satisfies global constraints.

[0012] The efficacy time-series coupling prediction model combines a kernel function support vector regression module and a dynamic Bayesian network module. The kernel function support vector regression module uses a radial basis kernel function to map low-dimensional nonlinear features to a high-dimensional space to establish a linear regression relationship. The dynamic Bayesian network module constructs a directed acyclic graph to represent the conditional dependency structure and time-varying coupling strength between parameters.

[0013] The steps for establishing the training dataset for the efficacy time-series coupling prediction model include: randomly selecting the training set and validation set at a ratio of 8 to 2; pairing the baseline data with the efficacy evaluation results after three treatment courses; generating time slices by sampling weekly within each treatment course for the time-series data according to the treatment course; and expanding the training sample size using small sample augmentation techniques.

[0014] The small sample augmentation technique involves using nonlinear least squares to fit the parameters of the differential equation describing the metabolic dynamics of physiological parameters to the time-series data of each training patient, thereby obtaining an individualized dynamic model. Gaussian perturbation is then performed in the parameter space to generate new parameter combinations, which are substituted into the differential equation to obtain new time-series trajectories as synthetic samples.

[0015] The training steps of the efficacy time-series coupled prediction model include: using a sequence minimum optimization algorithm to train the kernel function support vector regression module to obtain support vectors and regression coefficients by solving the Lagrange dual problem through quadratic programming; and using an expectation-maximization algorithm to train the dynamic Bayesian network module to alternately execute the expectation step and the maximization step to update the conditional probability parameters.

[0016] The conditional probability table update step of the dynamic Bayesian network module adjusts the learning rate of the time transition probability based on three parameters: blood glucose concentration fluctuation coefficient, hormone level variation coefficient, and standard deviation of TCM symptom scores. When the blood glucose concentration fluctuation coefficient exceeds 0.3 or the hormone level variation coefficient exceeds 0.4, the learning rate is reduced to 60% of the standard value.

[0017] In the primary processing thread grid, each thread independently processes a patient's blood glucose concentration, hormone levels, and hemodynamic parameters, and calculates preliminary efficacy assessment values ​​through a kernel function support vector regression module. In the deep processing thread grid, each thread processes all time-series data of a patient, and analyzes the time-varying coupling relationship between physiological parameters through a dynamic Bayesian network module.

[0018] The degradation detection algorithm of the entropy increase monitoring model based on the second law of thermodynamics calculates the prediction error of each batch of prediction results and discretizes it into several intervals. It then statistically analyzes the frequency of each interval to establish the error distribution probability, calculates the prediction error information entropy of the current batch and records the entropy value time sequence. When it is observed that the entropy value of 5 consecutive batches shows monotonic growth and the growth rate exceeds 10%, the model retraining mechanism is triggered.

[0019] The CUDA stream is a task execution queue on the GPU. Each CUDA stream independently manages the processing flow of a batch of patient data, and the pipelined processing of preliminary assessment and in-depth analysis is realized through a two-layer threaded grid within the stream.

[0020] This invention constructs a time-series coupled prediction model for treatment efficacy, combining a kernel function support vector regression module and a dynamic Bayesian network module to integrate the nonlinear mapping of physiological parameters with time-varying causal relationship modeling. The kernel function support vector regression module uses radial basis function to map low-dimensional nonlinear features to a high-dimensional space to establish regression relationships, capturing complex nonlinear correlations between parameters. The dynamic Bayesian network module constructs a directed acyclic graph to represent the conditional dependency structure between parameters and uses time transition probabilities to characterize the temporal evolution of parameters, accurately modeling the time-varying coupling relationship between physiological parameters. This solves the problem of insufficient prediction accuracy caused by traditional static models neglecting time-varying parameter dependencies. In summary, this invention addresses the technical problem mentioned in the background art of insufficient prediction accuracy in predicting the treatment effect of diabetic erectile dysfunction patients due to the lack of accurate modeling of the time-varying coupling relationship of multidimensional physiological parameters. Attached Figure Description

[0021] Figure 1 This is a flowchart of the method of the present invention.

[0022] Figure 2 This is a flowchart for data quality control.

[0023] Figure 3 This is a diagram illustrating the convergence process of the distributed drug efficacy synergistic prediction algorithm.

[0024] Figure 4 This is a schematic diagram of the feature mapping of the kernel function support vector regression module.

[0025] Figure 5 This is a graph showing the trend of information entropy changes in prediction error. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.

[0027] like Figure 1 The diagram shown is a flowchart of a method for predicting the therapeutic effect of Qiyang capsules based on historical data, provided by the present invention. This method includes the following steps:

[0028] S1. Collect baseline data of patients with impotence due to kidney deficiency and blood stasis in diabetes, including blood glucose concentration, hormone levels, hemodynamic parameters, routine blood indicators, biochemical indicators and TCM symptom scores, and establish the original dataset;

[0029] S2. The original dataset is subjected to quality control processing. The missing values ​​are filled by using the expectation-maximization multiple imputation framework combined with the random forest proximity matrix. The outliers are identified by the local outlier detection method and then replaced by robust regression to obtain the cleaned dataset.

[0030] S3. Perform feature engineering on the cleaned dataset, use the maximum correlation and minimum redundancy criterion of mutual information to screen the feature subset with strong correlation to efficacy and low redundancy, and then project it to a low-dimensional manifold through principal component analysis to obtain the optimized feature set.

[0031] S4. The optimized feature set is labeled according to four categories: Qiyang capsule group, acupuncture group, conventional treatment group, and placebo group. The distributed pharmacodynamic synergistic prediction algorithm based on Lagrange dual decomposition is used to distribute the feature set to multiple computing nodes for distributed processing to obtain the node optimization results.

[0032] S5. Use the CPU to read patient data from the node optimization results and transfer the patient data to the GPU's global memory block in batches according to the treatment group;

[0033] S6. Start the maximum number of CUDA streams on the GPU. Each CUDA stream starts two layers of thread grids, including a primary processing thread grid and a deep processing thread grid. The primary processing thread grid runs the kernel function support vector regression module of the efficacy time-series coupled prediction model, and the deep processing thread grid runs the dynamic Bayesian network module of the efficacy time-series coupled prediction model.

[0034] S7. The primary processing thread grid reads patient data from the global memory block, maps blood glucose concentration, hormone levels and hemodynamic parameters to a high-dimensional feature space through the kernel function support vector regression module, calculates the preliminary efficacy assessment value, and stores the preliminary efficacy assessment value in the shared memory block.

[0035] S8. The deep processing thread grid reads the preliminary efficacy assessment value from the shared memory block, extracts all time-series data of the patient corresponding to the preliminary efficacy assessment value, inputs it into the dynamic Bayesian network module for time-varying coupling analysis, outputs the final efficacy prediction value and confidence interval, and transmits the final efficacy prediction value back to the CPU-side main memory.

[0036] S9 and CPU read the final efficacy prediction value from main memory. After deploying the efficacy time-series coupled prediction model, they use an entropy increase monitoring model degradation detection algorithm based on the second law of thermodynamics to continuously monitor the prediction error information entropy. When the entropy value monotonically increases and exceeds the degradation threshold, the model is retrained to maintain prediction reliability.

[0037] The aforementioned expectation-maximization multiple imputation framework is an iterative optimization algorithm that estimates the conditional expectation of missing data by alternately executing expectation and maximization steps. The expectation step calculates the conditional distribution of the missing data based on the current parameter estimates, while the maximization step updates the parameter estimates based on the likelihood function of the complete data. This process iterates until the parameters converge. When combined with the random forest proximity matrix, the random forest model is first trained using complete samples, and the proximity matrix between samples is calculated. Proximity is defined as the frequency with which two samples fall into the same leaf node across all decision trees. For missing values, several complete samples with the highest proximity are selected, and imputation is performed based on a weighted average of their corresponding feature values, with the weights proportional to the proximity. This method can capture the nonlinear relationships between features and, compared to simple mean imputation, preserves the inherent structural features of the data, significantly improving the accuracy and stability of subsequent modeling.

[0038] The local anomaly detection method identifies outliers by calculating the degree of deviation of each data point from its neighborhood density. Specifically, for each point in the dataset, its... Calculate the nearest neighbor point of the given point to the th nearest neighbor point. The distance to the nearest neighbor is used as the reachability distance. Then, the local reachability density of the point is calculated, defined as its... The local outlier factor is the reciprocal mean of the reachability distances of the nearest neighbors. The local outlier factor is the average ratio of the local reachability density of all points in the neighborhood of the given point to the local reachability density of the point itself. When the local outlier factor is significantly greater than 1, it indicates that the density of the given point is much lower than that of its neighborhood, and it is determined to be an outlier. After detecting outliers, a robust regression method is used for replacement. Robust regression reduces the impact of outliers on the fitting results through iterative weighted least squares. In each iteration, the sample weights are adjusted according to the residual size; samples with larger residuals have lower weights, thus obtaining robust estimates that are insensitive to outliers.

[0039] The maximum relevance and minimum redundancy criterion for mutual information is a feature selection standard designed to select a subset of features that are most relevant to the target variable and have minimal redundancy among features. Mutual information measures the statistical dependence between two variables and is defined as the expected value of the ratio of the product of the joint probability distribution and the marginal probability distribution. The maximum relevance criterion requires that the sum of the mutual information between the selected features and the efficacy target variable be maximized, ensuring that the features contain rich predictive information. The minimum redundancy criterion requires that the sum of the mutual information among the selected features be minimized, avoiding computational redundancy and overfitting risks caused by overlapping information between features. By balancing relevance and redundancy, this criterion selects a subset of features that are informative and independent, effectively mitigating the curse of dimensionality and improving model training efficiency and generalization performance.

[0040] Principal component analysis (PCA) is a linear dimensionality reduction technique that transforms correlated variables into linearly independent principal components through orthogonal transformations. The covariance matrix is ​​calculated from the standardized feature matrix, and the eigenvalues ​​and eigenvectors of the covariance matrix are solved. The top few principal components are selected based on their eigenvalues, sorted from largest to smallest. Each principal component is a linear combination of the original features, and its eigenvalues ​​reflect the variance contribution rate explained by that principal component. Principal components with a cumulative variance contribution rate of 85% to 95% are selected, projecting the high-dimensional features onto a low-dimensional manifold. This significantly reduces the feature dimensionality while preserving key information, improving the computational efficiency of subsequent algorithms.

[0041] The distributed pharmacodynamic synergy prediction algorithm based on Lagrange dual decomposition decomposes the global optimization problem into multiple local subproblems that are solved in parallel. Based on Lagrange duality theory, the algorithm introduces dual variables to relax coupling constraints into a separable form. Specifically, patient data is distributed to four computing nodes according to treatment groups, with each node storing complete data for one treatment group. The global objective function is to minimize the sum of prediction errors for efficacy across all patients, with constraints based on the synergistic pharmacodynamic relationships between different groups. By introducing Lagrange multipliers, the global problem is transformed into a local optimization problem for each node, plus a penalty term. Each node independently solves its local subproblems to obtain the optimal prediction parameters for its own patient group. Subsequently, each node exchanges its dual variables, updates the Lagrange multipliers using the subgradient method, and adjusts the local solutions to gradually satisfy the global constraints. During the iteration process, the duality gap gradually decreases, eventually converging to the global optimum. The algorithm avoids the transmission of raw data between nodes, exchanging only low-dimensional dual variables, effectively protecting patient privacy. Simultaneously, utilizing a parallel computing architecture, each node processes its local data synchronously, significantly reducing the overall computation time. In multi-center collaborative research scenarios, data from different medical institutions can achieve globally optimal collaborative prediction without centralized aggregation, which not only ensures data security but also improves computational efficiency, providing a technical path for the efficient integration of large-scale clinical data.

[0042] The efficacy time-series coupling prediction model combines a kernel function support vector regression module and a dynamic Bayesian network module to capture the nonlinear time-varying coupling relationship between physiological parameters. The specific structure of the efficacy time-series coupling prediction model is as follows: the input layer receives multidimensional physiological parameter time-series data from an optimized feature set; the kernel function support vector regression module uses a radial basis function to map low-dimensional nonlinear features to a high-dimensional space, establishing a linear regression relationship in the high-dimensional space; the dynamic Bayesian network module constructs a directed acyclic graph to represent the conditional dependency structure between parameters, where nodes represent physiological parameters, directed edges represent causal relationships, and the conditional probabilities on the edges characterize the time-varying coupling strength between parameters; the fusion layer weightedly fuses the output of the kernel function support vector regression module and the inference results of the dynamic Bayesian network module, with the weights dynamically adjusted according to their respective prediction confidence levels; the output layer generates the final efficacy prediction value and confidence interval for each treatment group.

[0043] The steps for establishing the training dataset for the efficacy time-series coupling prediction model specifically include: randomly selecting samples from patients at a ratio of 8:2 as the training set and validation set; pairing the baseline data of the patients in the training set with the efficacy evaluation results after three treatment courses to establish input-output sample pairs; segmenting the time-series data according to the treatment courses, sampling weekly within each treatment course to generate time slices, forming a time-series sample sequence; expanding the training sample size using small sample augmentation techniques, and generating synthetic samples by solving a set of differential equations describing the evolution of physiological parameters.

[0044] The specific steps for training the efficacy time-series coupled prediction model include: initializing the kernel function parameters and regularization parameters of the kernel function support vector regression module, and initializing the conditional probability table of the dynamic Bayesian network module; training the kernel function support vector regression module using the sequence minimum optimization algorithm, and obtaining the support vectors and regression coefficients by solving the Lagrange dual problem through quadratic programming; training the dynamic Bayesian network module using the expectation-maximization algorithm, alternately executing the expectation step to infer the posterior distribution of latent variables and the maximization step to update the conditional probability parameters; evaluating the weights of the fusion layer on the validation set, and using grid search to determine the optimal weight combination that minimizes the prediction mean square error; iteratively training until the validation set loss function no longer decreases for 10 consecutive rounds, and then saving the model parameters.

[0045] The small-sample augmentation technique generates synthetic samples based on fitting the temporal evolution of physiological parameters using differential equations. The principle is as follows: the temporal changes of patients' physiological parameters in clinical data follow certain physiological laws, which are described by a system of ordinary differential equations. For parameters such as blood glucose concentration and hormone levels, differential equations describing their metabolic kinetic processes are established, and the equation parameters are estimated by fitting real patient data. In specific implementation, the parameters of the differential equations are fitted using a nonlinear least squares method for the temporal data of each training patient to obtain an individualized kinetic model. Gaussian perturbations are applied to the fitted parameters in the parameter space to generate new parameter combinations, which are then substituted into the differential equations to obtain new temporal trajectories as synthetic samples. This technique expands the training data under reasonable physiological constraints, increasing sample diversity while ensuring the physiological rationality of the generated data. In the efficacy temporal coupling prediction model structure, the augmented samples are mixed with the original samples as input, and the kernel function support vector regression module and the dynamic Bayesian network module learn on a larger training set, effectively alleviating the overfitting problem caused by small samples. The technology enables the model to learn more robust coupling patterns between parameters, improving its generalization prediction ability for patients not yet seen. By introducing differential equation constraints, the synthetic samples are not only sufficient in number but also maintain the intrinsic correlation between physiological parameters, avoiding non-physiologically reasonable samples generated by methods such as random interpolation. This ensures that the knowledge learned by the model conforms to actual clinical patterns, providing a theoretical basis and practical path for modeling small-sample medical data, and significantly enhancing the reliability and practicality of the prediction system in data-constrained scenarios.

[0046] The conditional probability table update step of the dynamic Bayesian network module in the efficacy time-series coupled prediction model is determined based on three parameters: blood glucose concentration fluctuation coefficient, hormone level coefficient of variation, and standard deviation of TCM symptom scores. The blood glucose concentration fluctuation coefficient reflects the stability of the patient's blood glucose control and is calculated as the ratio of the standard deviation to the mean of the blood glucose concentration time-series data. The hormone level coefficient of variation reflects the dynamic equilibrium of the endocrine system and is calculated as the ratio of the standard deviation to the mean of the hormone level time-series data. The standard deviation of TCM symptom scores reflects the degree of fluctuation in the patient's symptoms. During training, the learning rate of the time transition probabilities in the dynamic Bayesian network module is adjusted according to the numerical range of these three parameters. When all three parameters are within the normal range, the standard learning rate is used for parameter updates. When the blood glucose concentration fluctuation coefficient exceeds 0.3 or the hormone level coefficient of variation exceeds 0.4, the patient's physiological state is considered unstable, and the learning rate is reduced to 60% of the standard value to reduce the impact of abnormal samples on the model parameters. When the standard deviation of the TCM symptom score exceeds a threshold of 15 points, the update weight of the conditional probability table for the corresponding symptom node is increased, making the model more attentive to symptom change patterns. Through the aforementioned adaptive mechanism, the model can dynamically adjust its learning strategy based on individual patient differences, thereby improving the accuracy of predictions for patients in different physiological states.

[0047] The global memory block is the largest storage area on the GPU but has the highest access latency, used to store patient data and model parameters transferred from the CPU. The shared memory block is a high-speed cache area shared by all threads within the CUDA stream, located on the streaming multiprocessor chip, with access latency much lower than the global memory block, used to store preliminary efficacy assessment values ​​and inter-thread communication data.

[0048] Each thread in the primary processing thread grid independently processes a patient's blood glucose concentration, hormone levels, and hemodynamic parameters, calculating preliminary efficacy assessment values ​​using a kernel function support vector regression module. Each thread in the deep processing thread grid processes all time-series data for a patient, analyzing the time-varying coupling relationships between physiological parameters using a dynamic Bayesian network module.

[0049] The CUDA stream is a task execution queue on the GPU, which allows multiple computing tasks to be executed in parallel. Each CUDA stream independently manages the processing flow of a batch of patient data, and the pipelined processing of preliminary assessment and in-depth analysis is realized through a two-layer threaded grid within the stream.

[0050] The entropy increase monitoring model degradation detection algorithm based on the second law of thermodynamics borrows the principle of entropy increase in isolated systems to monitor model performance degradation. The algorithm defines the information entropy of the prediction error as a measure of system disorder, calculated as the negative logarithmic probability weighted sum of the prediction error distribution. In practice, after model deployment, the prediction error is calculated for each batch of prediction results, discretized into several intervals, and the frequency of each interval is statistically analyzed to establish the error distribution probability. The information entropy of the prediction error for the current batch is calculated, and the entropy value time series is recorded. A sliding window method is used to monitor the trend of entropy value changes. When five consecutive batches of entropy values ​​show a monotonically increasing trend with a growth rate exceeding 10%, the model is determined to have degraded due to data distribution drift. At this time, the system automatically triggers the model retraining mechanism, calling the latest collected clinical data to update the training set and re-execute the model training steps. If the data volume is insufficient to support complete retraining, an online incremental learning mode is activated, updating only the parameters in the model that have the greatest impact on prediction. The algorithm can detect model performance degradation without real labels because the increase in information entropy directly reflects the increase in prediction uncertainty, which conforms to the natural law of the system tending towards disorder described by the second law of thermodynamics. By continuously monitoring the entropy increase trend, the system can provide timely warnings before a significant decline in model performance, ensuring the long-term stable operation of the clinical decision-making system. This method provides a model maintenance solution for medical scenarios lacking real-time label feedback, ensuring the reliability of prediction results and avoiding misdiagnosis or treatment errors due to model degradation, thus providing technical assurance for patient safety and treatment effectiveness.

[0051] The kernel function support vector regression module transforms the nonlinear problem into a linear problem in a high-dimensional space through kernel mapping. The radial basis function calculates the similarity between two samples, with the similarity decreasing Gaussian as the sample distance increases. Support vectors are key samples in the training set that have a decisive influence on the position of the regression hyperplane, and are determined by solving a quadratic programming problem using the Lagrange multiplier method. The regularization parameter controls the balance between the model's difficulty and the fitting error; a larger parameter value tends to make the model simpler and smoother, while a smaller parameter value makes the model fit the training data better.

[0052] The dynamic Bayesian network module is a probabilistic graphical model that uses a directed acyclic graph to represent the conditional dependencies and causal structures between variables. Nodes represent physiological parameters or latent variables, directed edges represent causal relationships, and the conditional probabilities on the edges characterize the probability distribution of child nodes given a parent node's state. The dynamic Bayesian network module extends the static Bayesian network by introducing a time dimension to characterize the temporal evolution of parameters, describing the evolution of parameters from the previous time step to the current time step through time transition probabilities.

[0053] The described sequence minimum optimization algorithm is an efficient algorithm for solving the quadratic programming problem of support vector machines. In each iteration, the algorithm selects two Lagrange multipliers for optimization, fixing the other multipliers, and decomposes the multidimensional optimization problem into a series of two-dimensional subproblems. For the selected two multipliers, updated values ​​are obtained by analytically solving the quadratic programming problem, ensuring that the constraints are satisfied. During the iteration process, the multiplier pair that most severely violates the Kuhn-Tak condition is prioritized to accelerate the convergence speed.

[0054] The robust regression method reduces the impact of outliers using iterative weighted least squares. In the initial stage, ordinary least squares is used for fitting, and the residuals of each sample are calculated. Weights are assigned based on the residual magnitude, with samples having smaller residuals having weights close to 1 and samples having larger residuals having weights close to 0. A dual-weighting function is used to smooth the weight allocation and avoid abrupt weight changes. The regression coefficients and sample weights are iteratively updated until the change in the regression coefficients is less than the convergence threshold.

[0055] The degradation threshold is determined based on historical model performance data. When the growth rate of the prediction error information entropy exceeds 10% for five consecutive batches, the model retraining mechanism is triggered.

[0056] As an alternative, the present invention also provides a computer-implemented method for forming a Qiyang capsule treatment effect prediction system, wherein the computer is provided with a readable storage medium storing program instructions, and the program instructions execute the above-described method when the computer is run.

[0057] The specific implementation methods of the above steps are described in detail below.

[0058] The specific implementation of step S1 is as follows: Electronic medical record data of patients with kidney deficiency and blood stasis type diabetes and impotence are automatically collected through the hospital information system interface. The collected data includes fasting blood glucose concentration, 2-hour postprandial blood glucose concentration, testosterone level, luteinizing hormone level, peak velocity of penile corpora cavernosa, end-diastolic velocity, red blood cell count, white blood cell count, serum creatinine concentration, alanine aminotransferase activity, and results from a traditional Chinese medicine symptom scoring scale. The collected data is organized into a structured data table according to patient ID, with each row corresponding to one patient and each column corresponding to one physiological parameter, forming a raw dataset and storing it in a relational database. The purpose of this step is to establish a patient baseline database containing multidimensional physiological parameters, providing a complete source of raw data for subsequent data cleaning and feature engineering, ensuring that the predictive model can obtain sufficient input information.

[0059] The specific implementation of step S2 is as follows: First, missing value detection is performed on the original dataset, and the missing proportion of each physiological parameter is calculated. Parameters with a missing proportion exceeding 30% are removed. For parameters with a missing proportion below 30%, an expectation-maximization multiple imputation framework is used for imputation. The missing value is initialized as the mean of the parameter. The expectation step is then used to calculate the conditional expectation of the missing data, followed by the maximization step to update the parameter estimate. This process is iterated until convergence occurs when the parameter change is less than 0.001. During imputation, a random forest proximity matrix is ​​used. A random forest classifier is trained using complete samples, with 100 decision trees. The frequency of samples falling into the same leaf node is calculated as the proximity, and the top 5 samples in terms of proximity are selected for weighted imputation. After missing value imputation, a local anomaly detection method is used to identify outliers, and the number of nearest neighbors is set. The local reachability density and local outlier factor are calculated for each data point, with a local outlier factor greater than 1.5 considered an outlier. Robust regression is used to replace the detected outliers, employing iterative weighted least squares fitting. Sample weights are dynamically adjusted based on the residuals, and after 10 iterations, robust estimates are obtained to replace the outliers, forming a cleaned dataset. The purpose of these steps is to eliminate the interference of data quality issues on model training, preserving the statistical characteristics and physiological rationality of the data through scientific imputation and replacement methods, thereby improving the stability and unbiasedness of subsequent modeling.

[0060] The specific implementation of step S3 is as follows: All physiological parameters in the cleaned dataset are standardized by subtracting the mean from each parameter and dividing by the standard deviation, ensuring that parameters of different dimensions have the same numerical scale. The mutual information value between each physiological parameter and the therapeutic target variable is calculated. Mutual information reflects the statistical dependence strength between two variables; a larger value indicates a stronger correlation. Simultaneously, the mutual information values ​​between each pair of physiological parameters are calculated, constructing a feature redundancy matrix. Feature selection is performed using the maximum mutual information and minimum correlation criterion. First, parameters with the largest mutual information with the therapeutic target variable are added to the feature subset. Then, in each iteration, parameters with large mutual information with the target variable and small mutual information with the selected features are selected, balancing correlation and redundancy until 15 features are selected. Principal component analysis is performed on the selected feature subset to reduce dimensionality. The covariance matrix of the feature matrix is ​​calculated, and the eigenvalues ​​and eigenvectors are solved. The eigenvalues ​​are sorted from largest to smallest, and the top 8 principal components with a cumulative variance contribution rate of 90% are selected. The 15-dimensional features are projected onto an 8-dimensional low-dimensional manifold to obtain the optimized feature set. The purpose of these steps is to reduce feature dimensionality while retaining key prediction information, eliminate linear correlation and information redundancy between features, alleviate the computational complexity problem caused by the curse of dimensionality, and improve model training efficiency and generalization ability.

[0061] The specific implementation of step S4 is as follows: Based on the treatment plan received by the patients, the patient data in the optimized feature set is labeled into four categories: Qiyang Capsule group, acupuncture group, conventional treatment group, and placebo group. The patient data of the four categories are allocated to four independent computing nodes, each configured with the same computing resources and storage capacity. A distributed pharmacodynamic synergistic prediction algorithm based on Lagrange dual decomposition is used for processing. The global objective function is defined as minimizing the weighted sum of the efficacy prediction errors of all patients, with the constraint that the pharmacodynamic synergistic relationship between different treatment groups satisfies prior clinical knowledge. A Lagrange multiplier vector is introduced to decompose the global coupling problem into local optimization subproblems for each node. Each node independently solves for the optimal prediction parameters for its own group of patients. Nodes exchange dual variables through a message passing interface, and the Lagrange multipliers are updated using the subgradient method, with the step size set to an initial value of 0.1 and decreasing with the number of iterations. During the iteration process, the duality gap is monitored; convergence is determined when the duality gap is less than 0.01, and the node optimization results for each node are obtained. The purpose of these steps is to achieve collaborative optimization of multi-center data while protecting patient privacy, avoid centralized transmission of raw data, significantly shorten the overall processing time through distributed parallel computing, and provide globally optimal input data for subsequent prediction models.

[0062] The specific implementation of step S5 is as follows: The central processing unit (CPU) of the main control computer reads patient data from the node optimization results of the four computing nodes and organizes it into four data batches according to the treatment group. The patient data of each data batch is sequentially transferred to the global memory block of the graphics processor (GPU). The transfer process uses direct memory access technology to reduce the CPU's involvement overhead. A contiguous storage space is allocated for each patient in the global memory block, storing an 8-dimensional optimized feature vector and the patient's ID, ensuring that the GPU's computing core can quickly access the required data. The purpose of this step is to complete the migration of computing tasks from the CPU to the GPU, utilizing the GPU's massively parallel computing capabilities to accelerate the inference process of the prediction model, and preparing the data foundation for the subsequent initiation of CUDA streams and the execution of the threaded grid.

[0063] The specific implementation of step S6 is as follows: Based on the graphics processor's hardware specifications, query the maximum supported number of CUDA streams, and launch all CUDA streams within this range to achieve maximum parallelism. Allocate execution resources for two-layer thread grids within each CUDA stream. The primary processing thread grid contains a number of thread blocks equal to the number of patients in the current batch, with each thread block containing 32 threads. The deep processing thread grid is similarly configured with a corresponding number of thread blocks and threads. Pre-allocate storage space in shared memory blocks for both thread grids. The shared memory block for the primary processing thread grid stores preliminary efficacy assessment values, while the shared memory block for the deep processing thread grid stores intermediate calculation results and inter-thread communication data. Load the kernel function support vector regression module parameters of the efficacy temporal coupling prediction model into the constant memory of the primary processing thread grid, and load the dynamic Bayesian network module parameters into the constant memory of the deep processing thread grid. The purpose of these steps is to construct an efficient parallel computing pipeline, achieving phased processing of the prediction task through the division of labor and cooperation between the two thread grids, optimizing data access efficiency using the graphics processor's storage hierarchy, and providing hardware acceleration support for model inference.

[0064] The specific implementation of step S7 is as follows: Each thread in the primary processing thread grid reads an 8-dimensional optimized feature vector of a patient from the global memory block, extracting three key dimensions: blood glucose concentration, hormone level, and hemodynamic parameters. The thread inputs these three parameters into the kernel function support vector regression module to calculate the radial basis function kernel value, with the width parameter of the kernel function set to 1.0. Based on the pre-trained support vectors and regression coefficients, a weighted sum of the kernel function values ​​of the input sample and each support vector is calculated to obtain the regression prediction value in the high-dimensional feature space as a preliminary efficacy assessment value. The thread writes the calculated preliminary efficacy assessment value to the corresponding location in the shared memory block and sets a completion flag to notify the deep processing thread grid. The purpose of this step is to utilize the nonlinear mapping capability of the kernel function support vector regression module to quickly obtain preliminary efficacy predictions based on key physiological parameters, thereby reducing unnecessary computational overhead for subsequent in-depth analysis to screen patient samples requiring special attention.

[0065] The specific implementation of step S8 is as follows: Each thread in the deep processing thread grid reads the preliminary efficacy assessment value of the corresponding patient from the shared memory block, and simultaneously reads the patient's complete time-series data from the global memory block, including all physiological parameters at the baseline time and at each time point during treatment. The thread inputs the time-series data into the dynamic Bayesian network module, calculates the conditional probability of each node based on the directed acyclic graph structure, and uses a forward inference algorithm to extrapolate the parameter evolution probability distribution at each time point from the baseline time. It calculates the time-varying coupling strength between parameters, quantifies the influence of blood glucose concentration fluctuations on hormone level changes, and the regulatory effect of hormone levels on hemodynamic parameters. Based on the conditional probability and coupling strength, it comprehensively assesses the efficacy evolution trend and outputs the final efficacy prediction value and the lower and upper bounds of the confidence interval. The thread transmits the final efficacy prediction value and confidence interval back to the main memory on the central processing unit side through direct memory access technology, avoiding data transmission becoming a performance bottleneck. The purpose of these steps is to capture the causal relationships and temporal dependencies between physiological parameters through the dynamic Bayesian network module, perform in-depth analysis based on the preliminary assessment, generate accurate efficacy predictions considering parameter coupling effects, and provide reliable quantitative basis for clinical decision-making.

[0066] The specific implementation of step S9 is as follows: The central processing unit reads the final efficacy prediction values ​​of all patients from main memory and calculates the residuals between the predicted values ​​and the actual efficacy evaluation results. The residuals are discretized into 10 equal-width intervals according to their numerical values. The frequency of each interval is counted to establish a residual distribution histogram, and the residual probability distribution is obtained by normalization. The prediction error information entropy is calculated based on the probability distribution. The information entropy is the negative logarithmic weighted sum of the probabilities of each interval. The information entropy value of the current batch is appended to the entropy value time series. The entropy values ​​of the most recent 5 batches are extracted using a sliding window of length 5. The linear regression slope of the entropy value within the sliding window is calculated. When the slope is positive and the slope of 5 consecutive windows is greater than 0, the entropy value is determined to have a monotonically increasing trend. The entropy growth rate is calculated as the difference between the current entropy value and the initial entropy value of the window divided by the initial entropy value of the window. When the growth rate exceeds the degradation threshold of 10%, the model retraining mechanism is triggered. During model retraining, the training set is expanded using newly acquired clinical data. Feature engineering and model training steps are re-executed, and the parameters of the kernel function support vector regression module and dynamic Bayesian network module are updated. Finally, the old model parameters in the graphics processor's memory are replaced. The purpose of these steps is to establish a continuous monitoring mechanism for model performance, quantify the changing trend of prediction uncertainty using the information entropy index, and intervene promptly when the model degrades due to data distribution drift. This ensures the reliability and accuracy of the prediction system in long-term operation, guaranteeing consistently effective treatment guidance for patients.

[0067] It should be noted that the key technical ideas of this invention include the following aspects. The first key technical idea is to use a distributed drug efficacy collaborative prediction algorithm based on Lagrange dual decomposition to solve the contradiction between data privacy protection and global optimization in multi-center data. Traditional centralized processing requires aggregating patient data from different medical institutions to a central server, which poses a risk of data leakage and incurs huge transmission overhead. This invention decomposes the globally coupled problem into local subproblems that each node solves independently through Lagrange dual theory. Nodes only exchange low-dimensional dual variables rather than the original sensitive data, fundamentally protecting patient privacy. At the same time, each node processes local data in parallel, making full use of distributed computing resources, significantly shortening the overall computation time compared to serial centralized processing. The algorithm achieves convergence of the global optimal solution while ensuring data security, providing a technical path for multi-center collaborative research that balances privacy and performance optimization. The second key technical idea is to construct a two-layer threaded grid pipeline architecture based on CUDA streams to achieve hardware acceleration of model inference. Traditional central processing units face computational bottlenecks when executing prediction tasks serially, making it difficult to meet the real-time processing needs of large-scale patient data. This invention leverages the massively parallel computing capabilities of graphics processing units (GPUs) to achieve pipelined parallelism for preliminary evaluation and deep analysis through the division of labor between primary processing thread grids and deep processing thread grids. The primary thread grid quickly filters samples requiring deep analysis, preventing all samples from entering the computationally intensive deep network modules and significantly reducing average processing latency. Combined with hierarchical storage management of global and shared memory blocks, the data access pattern is optimized, reducing the number of high-latency global memory accesses and further improving inference throughput. This architecture fully utilizes the performance advantages of heterogeneous computing platforms, providing a hardware foundation for real-time response in clinical decision-making systems. The third key technical idea is to achieve unsupervised detection of model degradation based on entropy increase monitoring using the second law of thermodynamics. Traditional model maintenance relies on continuously acquiring real label feedback, but label acquisition in medical scenarios is often delayed and costly. This invention draws on the principle of entropy increase in isolated systems, using prediction error information entropy as a measure of system disorder, and judging model performance degradation by monitoring the monotonic growth trend of entropy values. This method can detect model degradation caused by data distribution drift without real labels, achieving autonomous assessment of model health status. When an abnormal increase in entropy is detected, a retraining mechanism is automatically triggered to ensure that the model always adapts to the current data distribution and avoids prediction inaccuracies due to performance degradation. This approach provides a guarantee mechanism for the long-term stable operation of models in scenarios lacking real-time labels. The synergistic effect of the above three key technical approaches constitutes a complete distributed privacy-preserving prediction system. The Lagrange dual decomposition algorithm ensures privacy and security at the data source and achieves global optima, providing high-quality input for subsequent models. The CUDA streaming two-layer threaded grid architecture accelerates the model inference process, enabling the system to process large-scale real-time data.An entropy monitoring mechanism continuously assesses the model's health status and proactively intervenes to maintain it before performance deteriorates, ensuring long-term operational reliability. These three elements work together to form a closed-loop technology system covering the entire lifecycle from data acquisition to model deployment and continuous maintenance. Compared to traditional centralized, single-threaded static model solutions, this invention achieves a qualitative breakthrough in privacy protection, computational efficiency, and long-term stability, providing a systematic solution for multi-center collaborative clinical decision support systems.

[0068] It should be noted that this invention also solves the following technical problem: overfitting easily occurs when training deep prediction models in scenarios with small sample clinical data, leading to insufficient generalization ability. This invention generates synthetic samples by employing a small-sample augmentation technique, based on the temporal evolution law of physiological parameters fitted by differential equations. For the temporal data of each training patient, the parameters of the differential equation are fitted using a nonlinear least squares method to obtain an individualized dynamic model. Gaussian perturbation is applied in the parameter space to generate new parameter combinations, which are then substituted into the differential equation to obtain new temporal trajectories as synthetic samples. By expanding the training data under reasonable physiological constraints, both sample diversity and physiological rationality of the generated data are increased, enabling the model to learn more robust parameter coupling patterns, effectively alleviating the overfitting problem caused by small samples, and improving the generalization prediction ability for patients not yet seen.

[0069] Specifically, the principle of this invention is as follows: This invention achieves accurate characterization of the time-varying coupling relationship of physiological parameters through a two-layer modeling architecture. The kernel function support vector regression module uses kernel mapping technology to transform the nonlinear relationship in the original feature space into a linear relationship in a high-dimensional feature space. It determines the support vectors and regression coefficients by solving the Lagrange dual problem, establishing a nonlinear mapping function between parameters. The dynamic Bayesian network module introduces a time dimension, representing the causal relationship between parameters through directed edges, using conditional probability to characterize the probability distribution of child nodes given a parent node state, and time transition probability to describe the evolution of parameters from the previous time step to the current time step. The fusion layer dynamically adjusts the weights based on the prediction confidence of the two modules, combining the nonlinear mapping results and the time-varying causal inference results to accurately capture the coupling pattern between physiological parameters over time, thereby improving prediction accuracy.

[0070] The following provides a specific embodiment 1 of the present invention. The specific implementation methods of steps S1 and S5-S8 in this embodiment 1 are the same as those described above, and will not be repeated in detail here. The specific implementation methods of other steps are described in detail below.

[0071] The specific implementation of the expected maximization multiple imputation framework in step S2 is as follows: For the original dataset containing missing values, firstly, a random forest model is trained using complete samples, and the proximity matrix between samples is calculated. Its expression is:

[0072] ;

[0073] In the formula, For the sample With sample The proximity between them is dimensionless; This represents the total number of decision trees in the random forest, with an empirical value of 500. As an indicator function, when the sample With sample In the The value is 1 when the decision tree falls into the same leaf node, and 0 otherwise. It is dimensionless. The sample number for the missing values ​​to be filled; For complete sample numbering; Number the decision tree. For samples with missing values... The Each feature is filled using a weighted average, and the filled value is... The calculation formula is:

[0074] ;

[0075] In the formula, For the sample No. Imputed values ​​for each feature; This is the number of complete samples with the highest proximity, typically set to 10. For complete samples The Each characteristic value has a unit determined based on the specific characteristic; for example, the unit for blood glucose concentration is [unit missing]. Hormone level units are ; This is used for feature numbering. In local anomaly detection methods, the sample is first calculated. Locally achievable density :

[0076] ;

[0077] In the formula, For the sample Locally achievable density; This is the number of nearest neighbors, which defaults to 20. For the sample of A set of nearest neighbors; For the sample to neighborhood samples The reachable distance is calculated as ,in For the sample and The Euclidean distance between them For the sample To its first The distance to the nearest neighbor. Sample Local anomalous factors The calculation formula is:

[0078] ;

[0079] In the formula, For the sample Local anomaly factors, dimensionless; for neighborhood samples The locally attainable density. When When an outlier is identified, robust regression is used for substitution. The substitution value is calculated using iterative weighted least squares, with the weight function as follows:

[0080] ;

[0081] In the formula, For the sample The weights in robust regression are dimensionless. For the sample The residual value; The standard deviation of the residuals, in units of same; The subscript is used to indicate robust regression. The regression coefficient update formula is:

[0082] ;

[0083] In the formula, For the first The regression coefficient vector of the next iteration; The characteristic matrix; For the first The weight diagonal matrix of the next iteration has diagonal elements of... Dimensionless; This is a vector of therapeutic efficacy values; superscript Indicates the number of iterations; superscript This indicates the matrix transpose.

[0084] The specific implementation of the maximum correlation and minimum redundancy criterion of mutual information in step S3 is as follows: Features With therapeutic target variables Mutual information between The calculation formula is:

[0085] ;

[0086] In the formula, Features With therapeutic target variables Mutual information between them, dimensionless; Features and target variable The joint probability distribution of is dimensionless; and Features and target variable The marginal probability distribution of is dimensionless; For the first Features; the summation symbol iterates through all possible feature values ​​and the target value. Feature selection criteria. The expression is:

[0087] ;

[0088] In the formula, The criterion value for selecting features is dimensionless; A subset of candidate features; The number of features contained in the feature subset; Features With features The mutual information between them is dimensionless. In principal component analysis, the original features are first standardized:

[0089] ;

[0090] In the formula, For the first A standardized feature, dimensionless; For the first One original feature value; For the first The mean of each feature, in units of same; For the first The standard deviation of each feature, in units of Same. (Number) Principal Components The calculation formula is:

[0091] ;

[0092] In the formula, For the first One principal component, dimensionless; For the first The th eigenvector of the th feature vector One component, dimensionless; This represents the total number of original features; Principal component numbering. Cumulative variance contribution rate. The calculation is as follows:

[0093] ;

[0094] In the formula, The contribution rate to the cumulative variance is dimensionless. The number of principal components selected; For the first Each eigenvalue is dimensionless.

[0095] The specific implementation of the distributed pharmacodynamic synergistic prediction algorithm based on Lagrange dual decomposition in step S4 is as follows: the objective function of the global optimization problem... Expressed as:

[0096] ;

[0097] In the formula, The global loss function is dimensionless. For the first Patients from each treatment group were assembled. Values ​​1 to 4 correspond to the Qiyang capsule group, acupuncture group, conventional treatment group, and placebo group, respectively. For patients The actual therapeutic effect value; For patients Predicted therapeutic efficacy value, in units of same; The standard deviation of the therapeutic effect value, in units of 1 / 2000 and 1 / 3000. same; For the first The group of Lagrange multipliers is dimensionless. For the first The local constraint function of the group is dimensionless. For the first The local parameter vector of the group. Local subproblem objective function of the group Expressed as:

[0098] ;

[0099] In the formula, For the first The local loss function of the group is dimensionless. The Lagrange multipliers are updated using the subgradient method, with the update formula as follows:

[0100] ;

[0101] In the formula, For the first After the nth iteration The group of Lagrange multipliers is dimensionless. For the first After the nth iteration The group of Lagrange multipliers is dimensionless. For the first The learning rate for each iteration is initially 0.01 and decreases with each iteration, and is dimensionless. For the first During the nth iteration The subgradient of the group constraint function is dimensionless. For the first During the nth iteration The local parameter vector of the group.

[0102] The specific implementation of the degradation detection algorithm for the entropy increase monitoring model based on the second law of thermodynamics in step S9 is as follows: First, calculate the frequency of the prediction error distribution. :

[0103] ;

[0104] In the formula, To fall into the first The sample frequency of each error interval, dimensionless; To fall into the first The number of samples in each error interval; This represents the total number of samples; This is the error identifier subscript. Prediction error information entropy. The calculation formula is:

[0105] ;

[0106] In the formula, The prediction error information entropy is dimensionless; The number of error discretization intervals, defaulting to 20. Entropy growth rate. The calculation is as follows:

[0107] ;

[0108] In the formula, The entropy growth rate is dimensionless. For the first The information entropy of a batch is dimensionless. For the first The information entropy of a batch is dimensionless. When 5 consecutive batches satisfy... The model is retrained at that time.

[0109] The specific implementation method of establishing a small-sample augmentation technique for the training dataset of the efficacy time-series coupling prediction model is as follows: The metabolic kinetic differential equation of blood glucose concentration is:

[0110] ;

[0111] In the formula, This refers to blood glucose concentration, in units of... ; For time, the unit is ; This is the rate constant of glucose consumption, in units of... The result was obtained by fitting using the nonlinear least squares method. The insulin-promoted glucose uptake coefficient, in units of... ; Hormone levels, in units of ; The input function for treatment is blood glucose, in units of... The kinetic equation for hormone levels is:

[0112] ;

[0113] In the formula, This is the hormone metabolic rate constant, in units of... ; The symptom-stimulating hormone secretion coefficient, in units of ; Score the symptoms according to Traditional Chinese Medicine (TCM), with the unit being points; The effect of treatment on hormones is a function, in units of The parameter perturbation uses a Gaussian distribution, and the perturbed parameters... The calculation is as follows:

[0114] ;

[0115] In the formula, The perturbed glucose consumption rate constant is expressed in units of 100 mg / L. ; The disturbance intensity coefficient has an empirical value of 0.15 and is dimensionless. Let be a random variable that follows a standard normal distribution and is dimensionless; Subscript for disturbance identifier.

[0116] The specific implementation of the conditional probability table update step in the dynamic Bayesian network module of the efficacy time-series coupled prediction model is as follows: blood glucose concentration fluctuation coefficient The calculation formula is:

[0117] ;

[0118] In the formula, is the blood glucose concentration fluctuation coefficient, which is dimensionless; The standard deviation of blood glucose concentration time series data is given in units of 1000 mg / L. ; This represents the mean of the time-series blood glucose concentration data, in units of... Hormone level variation coefficient The calculation formula is:

[0119] ;

[0120] In the formula, is the coefficient of variation of hormone levels, dimensionless; The standard deviation of hormone level time series data is given in units of 1000 m / s. ; This represents the mean of hormone level time series data, in units of... The TCM symptom scoring standard is poor. The calculation formula is:

[0121] ;

[0122] In the formula, The standard deviation of the TCM symptom scores is expressed in points. This represents the number of time-series sampling points; For the first Symptom scores at individual time points, in points; The mean of the symptom scores is expressed in points. This is a time-series identifier subscript. Adaptive learning rate. The adjustment formula is:

[0123] ;

[0124] In the formula, The adaptive learning rate is dimensionless. The standard learning rate is 0.01 by default and is dimensionless. Use adaptive identifier subscripts.

[0125] Kernel function: Radial basis function in support vector regression module The expression is:

[0126] ;

[0127] In the formula, For the sample and The kernel function values ​​between these values ​​are dimensionless. This is a parameter for the kernel function, with an empirical value of 0.1, and is dimensionless. For the sample and The Euclidean distance between them; is the standard deviation of the feature space, with the same units as the feature distance; and The first The and the first The feature vector of each sample. The Lagrange multiplier update formula in the sequence minimum optimization algorithm is:

[0128] ;

[0129] In the formula, For the updated number One Lagrange multiplier, dimensionless; For the previous version One Lagrange multiplier, dimensionless; For the sample The label value is dimensionless; For the sample The prediction error; For the sample The prediction error, in units of same; This is the normalization coefficient, with the same unit as the squared prediction error, calculated as follows: Dimensionless; Minimize the index of the sequence; and These are the indexes for the updated and unupdated versions, respectively.

[0130] To better understand and implement this invention, the following is an example 2 of a specific application scenario: Technicians collected clinical data from 120 patients with kidney deficiency and blood stasis type diabetes and impotence at a traditional Chinese medicine hospital. These included 30 patients in the Qiyang capsule group, 30 in the acupuncture group, 30 in the conventional treatment group, and 30 in the placebo group. Baseline data included 23 indicators such as fasting blood glucose concentration, 2-hour postprandial blood glucose concentration, testosterone level, luteinizing hormone level, peak velocity of penile corpora cavernosa, end-diastolic velocity, white blood cell count, hemoglobin concentration, fasting insulin level, glycated hemoglobin, kidney deficiency symptom score, and blood stasis symptom score. The original dataset had an 8.3% missing value rate and several outliers. Technicians first performed data quality control processing. Figure 2As shown, a multiple imputation framework combining expectation-maximization and random forest proximity matrix was used to impute missing values. A random forest model was trained on data from 120 patients, with 500 decision trees, and the proximity matrix between samples was calculated. For missing fasting blood glucose concentration data, the 15 complete samples with the highest proximity were selected, and imputation was performed based on the weighted average of their fasting blood glucose values. The weights were 0.089, 0.086, 0.083, 0.078, 0.075, 0.071, 0.068, 0.064, 0.061, 0.057, 0.054, 0.051, 0.048, 0.044, and 0.041, respectively. Local anomaly detection was used to identify outliers, with 10 nearest neighbors, and the local anomaly factor was calculated for each data point. A patient in the Qiyang capsule group was found to have a testosterone level of 26.8 nmol / L, with a local anomaly factor of 2.37, significantly higher than the threshold of 1.5, and was therefore identified as an outlier. A robust regression method was used as an alternative, and the iterative weighted least squares method converged after 8 iterations, adjusting the patient's testosterone level to 18.5 nmol / L. The cleaned dataset contained complete baseline data from 120 patients.

[0131] Technicians performed feature engineering on the cleaned dataset and used the maximum relevance and minimum redundancy criterion for mutual information to select a subset of features. The mutual information values ​​between the 23 original features and the therapeutic target variable were calculated: fasting blood glucose concentration (0.78), testosterone level (0.82), peak systolic velocity of penile cavernous body (0.76), glycated hemoglobin (0.74), and kidney deficiency symptom score (0.71). The mutual information matrix between features was calculated; the mutual information value between fasting blood glucose concentration and glycated hemoglobin was 0.68, indicating high redundancy. By balancing correlation and redundancy, a subset of 12 features was selected, including fasting blood glucose concentration, testosterone level, luteinizing hormone level, peak systolic velocity of penile cavernous body, end-diastolic velocity, fasting insulin level, kidney deficiency symptom score, and blood stasis symptom score. Principal component analysis was performed on the feature subset, and the eigenvalues ​​of the covariance matrix were calculated to be 4.86, 2.73, 1.94, 1.38, 0.92, 0.67, 0.54, 0.43, 0.31, 0.18, 0.09, and 0.05, respectively. The cumulative variance contribution rate of the first 5 principal components reached 89.4%. The 12-dimensional features were projected onto a 5-dimensional low-dimensional manifold to obtain the optimized feature set.

[0132] Technicians optimized the feature set by labeling it according to the four treatment groups and deployed a distributed pharmacodynamic synergy prediction algorithm based on Lagrange dual decomposition. Patient data was distributed across four computing nodes, with each node storing data from 30 patients in one treatment group. The global objective function was to minimize the sum of squared prediction errors for all patients, with constraints including a synergy coefficient greater than 0.15 between the Qiyang capsule group and the acupuncture group, and a synergy coefficient greater than 0.08 between the Qiyang capsule group and the conventional treatment group. Figure 3 As shown, the global problem is transformed into a local optimization problem for each node by introducing Lagrange multipliers, with the Lagrange multipliers initialized to 0.5. Each node independently solves its local subproblem: the nodes for the Qiyang capsule group, acupuncture group, conventional treatment group, and placebo group obtain the optimal predicted parameter vectors. Each node then swaps its dual variables and updates the Lagrange multipliers using the subgradient method. After the first iteration, the multipliers are updated to 0.53; after the second iteration, they are updated to 0.56; and after the third iteration, they are updated to 0.58. After 12 iterations, the dual gap converges to 0.003, yielding the node optimization results.

[0133] Technicians used the CPU to read patient data from the node optimization results, transferring data from 120 patients in batches according to treatment groups to the GPU's global memory block. Eight CUDA streams were launched on the NVIDIA Tesla V100 GPU, each processing 15 patient data points. Each CUDA stream ran two-layer thread grids: a primary processing thread grid containing 15 thread blocks (each containing 32 threads), and a deep processing thread grid containing 15 thread blocks (each containing 64 threads). The primary processing thread grid read data such as fasting blood glucose concentration, testosterone level, and peak penile corpus cavernosum blood flow velocity from the global memory block, mapping it to a high-dimensional feature space using a kernel function support vector regression module. Figure 4As shown, the bandwidth parameter of the radial basis function kernel was set to 0.35, the regularization parameter to 1.2, and the number of support vectors to 76. For a patient in the Qiyang capsule group, the fasting blood glucose concentration was 8.6 mmol / L, the testosterone level was 16.3 nmol / L, and the peak velocity of penile corpus cavernosum blood flow was 24.7 cm / s. The kernel function support vector regression module calculated a preliminary efficacy assessment value of 72.4 points, which was stored in a shared memory block. The deep processing thread grid read the preliminary efficacy assessment value from the shared memory block and extracted all time-series data of the patient for 3 treatment courses totaling 12 weeks, sampling once a week to generate 12 time slices. The dynamic Bayesian network module was input for time-varying coupling analysis. The network structure contained 8 observation nodes and 3 latent variable nodes, with 17 directed edges. The patient's blood glucose concentration fluctuation coefficient was 0.18, the hormone level coefficient of variation was 0.26, and the standard deviation of the TCM symptom score was 9.3 points, all within the normal range. The conditional probability table was updated using a standard learning rate of 0.01. The dynamic Bayesian network module inferred a time-varying coupling strength of 0.64 between blood glucose concentration and testosterone levels, 0.71 between testosterone levels and penile cavernous body blood flow, and 0.58 between blood flow parameters and symptom scores. The fusion layer weighted and fused the output score of 72.4 from the kernel function support vector regression module and the inferred score of 76.8 from the dynamic Bayesian network module. The weights were dynamically adjusted based on their respective prediction confidence levels: the kernel function support vector regression module had a weight of 0.42, and the dynamic Bayesian network module had a weight of 0.58. The final efficacy prediction value was 74.9, with a confidence interval of 71.3 to 78.5. The final efficacy prediction value was then transmitted back to the CPU's main memory.

[0134] Technicians compiled the predicted efficacy results for 120 patients, as shown in Table 1. The average predicted efficacy score was 74.2 points for the Qiyang Capsule group, 68.5 points for the acupuncture group, 63.8 points for the conventional treatment group, and 52.3 points for the placebo group. Comparing the predicted results with the actual clinical efficacy evaluation results, the average absolute error between the predicted and actual values ​​was 3.6 points for the Qiyang Capsule group, 4.2 points for the acupuncture group, 4.8 points for the conventional treatment group, and 3.9 points for the placebo group.

[0135] Table 1. Statistical table of predicted efficacy results for each treatment group

[0136]

[0137] After deploying the efficacy-time coupled prediction model, technicians continuously monitor the prediction performance using an entropy increase monitoring model degradation detection algorithm based on the second law of thermodynamics. For example... Figure 5As shown, the system processes prediction requests from 20 new patients per batch, calculates the prediction error and discretizes it into 8 intervals, and statistically analyzes the frequency of each interval to establish the error distribution probability. The prediction error information entropy of the initial batch was 1.86, the second batch was 1.91, the third batch was 1.95, the fourth batch was 2.03, and the fifth batch was 2.14. Using a sliding window method to monitor the entropy trend, it was observed that the entropy value of five consecutive batches showed a monotonically increasing trend: the growth rate of the second batch relative to the first batch was 2.7%, the third batch relative to the second batch was 2.1%, the fourth batch relative to the third batch was 4.1%, and the fifth batch relative to the fourth batch was 5.4%, with a cumulative growth rate reaching 15.1%, exceeding the degradation threshold of 10%. The system automatically triggered the model retraining mechanism, calling the latest 80 clinical data to update the training set and re-executing feature engineering and model training steps. After retraining, the model's prediction error information entropy decreased to 1.92, and the model performance returned to a stable state.

[0138] This invention represents a significant advancement over traditional methods for evaluating the efficacy of Qiyang Capsules. Traditional methods rely on subjective judgment by clinicians based on patient symptoms and examination results, which are influenced by physician experience and the accuracy of patient self-reports, lacking objective and quantitative predictive basis. This invention, through an expectation-maximization multiple imputation framework combined with a random forest proximity matrix to fill missing values, captures the nonlinear relationships between features, preserves the inherent structural features of the data, and avoids information loss caused by simple mean imputation. The local anomaly detection method identifies outliers by calculating the deviation of data points from their neighborhood density, and, combined with robust regression substitution, effectively reduces the interference of outlier data on model training. The mutual information maximum correlation minimum redundancy criterion filters out feature subsets that are information-rich and independent, and, combined with principal component analysis for dimensionality reduction, effectively alleviates the curse of dimensionality problem, improving model training efficiency and generalization performance. The distributed pharmacodynamic synergistic prediction algorithm based on Lagrange dual decomposition decomposes the global optimization problem into multiple local subproblems that are solved in parallel, avoiding the transmission of original data between nodes, protecting patient privacy, and significantly shortening computation time. The efficacy time-series coupled prediction model combines a kernel function support vector regression module and a dynamic Bayesian network module to capture the nonlinear time-varying coupling relationship between physiological parameters. Compared with traditional linear regression models, it can more accurately describe complex physiological evolution processes. The entropy increase monitoring model degradation detection algorithm based on the second law of thermodynamics detects model performance degradation by monitoring the changing trend of prediction error information entropy, without requiring real labels. This ensures the long-term stable operation of the prediction system and provides reliable technical support for clinical decision-making.

[0139] It should be noted that the variables involved in this invention are explained in detail in Tables 2 and 3.

[0140] Table 2. Variable Explanation Table (Part 1)

[0141]

[0142] Table 3. Variable Explanation Table (Part Two)

[0143]

[0144] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A historical data-based Qiyang capsule treatment effect prediction method, characterized by, include: Baseline data of patients with diabetic impotence of the kidney deficiency and blood stasis type were obtained to establish an original dataset. The original dataset was cleaned by quality control. The cleaned dataset was then processed by feature engineering to obtain an optimized feature set. The optimized feature set was labeled according to the treatment group and distributed to multiple computing nodes using a distributed pharmacodynamic synergistic prediction algorithm based on Lagrange dual decomposition for distributed processing to obtain node optimization results. The CPU reads patient data from the node optimization results and transmits it to the global memory block of the GPU. On the GPU, a CUDA stream is started to execute a two-layer threaded grid processing. The primary processing threaded grid runs the kernel function support vector regression module of the efficacy temporal coupling prediction model to calculate the preliminary efficacy assessment value and store it in the shared memory block. The deep processing threaded grid runs the dynamic Bayesian network module of the efficacy temporal coupling prediction model to perform time-varying coupling analysis and output the final efficacy prediction value. After deploying the efficacy temporal coupling prediction model, an entropy increase monitoring model degradation detection algorithm based on the second law of thermodynamics is used to continuously monitor the prediction error information entropy and trigger model retraining when the entropy value monotonically increases and exceeds the degradation threshold.

2. The method of claim 1, wherein, The quality control process involves using an expectation-maximization multiple imputation framework combined with a random forest proximity matrix to fill in missing values, and using a local outlier detection method to identify outliers before replacing them with robust regression.

3. The method according to claim 2, characterized in that, The expected value maximization multiple imputation framework combines a random forest proximity matrix to fill missing values. It uses complete samples to train a random forest model to calculate the proximity matrix between samples. For missing values, it selects several complete samples with the highest proximity and fills them with a weighted average of their corresponding feature values.

4. The method according to claim 3, characterized in that, The local anomaly detection method determines the reachability distance to the k nearest neighbors for each point in the dataset, calculates the local reachability density of the point as the inverse mean of the reachability distances of its k nearest neighbors, and the local anomaly factor is the average ratio of the local reachability density of each point in the neighborhood of the point to the local reachability density of the point itself.

5. The method according to claim 4, characterized in that, The feature engineering process involves using the maximum correlation and minimum redundancy criterion of mutual information to screen feature subsets that are highly correlated with therapeutic efficacy and have low redundancy, and then projecting them onto a low-dimensional manifold through principal component analysis.

6. The method according to claim 5, characterized in that, The maximum relevance and minimum redundancy criterion for mutual information requires that the sum of mutual information between the selected features and the target efficacy variable be maximized and the sum of mutual information between the selected features be minimized.

7. The method according to claim 6, characterized in that, The principal component analysis involves calculating the covariance matrix of the standardized feature matrix, solving for eigenvalues ​​and eigenvectors, and selecting principal components with a cumulative variance contribution rate of 85% to 95% based on the eigenvalues ​​sorted from largest to smallest.

8. The method according to claim 7, characterized in that, The distributed pharmacodynamic synergistic prediction algorithm based on Lagrange dual decomposition distributes patient data to four computing nodes according to treatment groups. Each node independently solves local subproblems and then exchanges dual variables. The Lagrange multipliers are updated according to the subgradient method to adjust the local solutions so that they satisfy global constraints.

9. The method according to claim 8, characterized in that, The efficacy time-series coupling prediction model combines a kernel function support vector regression module and a dynamic Bayesian network module. The kernel function support vector regression module uses radial basis kernel functions to map low-dimensional nonlinear features to a high-dimensional space to establish a linear regression relationship. The dynamic Bayesian network module constructs a directed acyclic graph to represent the conditional dependency structure and time-varying coupling strength between parameters.

10. The method according to claim 9, characterized in that, The steps for establishing the training dataset for the efficacy time-series coupled prediction model are as follows: the training set and validation set are randomly selected in an 8:2 ratio; the baseline data and the efficacy evaluation results after three treatment courses are paired; the time-series data are segmented according to the treatment course, and time slices are generated by sampling weekly within each treatment course; and the training sample size is expanded using small sample augmentation technology.