Multi-modal data analysis method for drug screening safety evaluation
By constructing multimodal standardized states and state evolution simulations, the toxicity index in the drug screening process is quantified, solving the problem that the dynamic response of drug screening safety evaluation is difficult to quantify in existing technologies, and realizing efficient and reliable drug screening safety evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-26
AI Technical Summary
Existing drug screening methods cannot fully reflect the dynamic molecular response process of cells under the action of drugs in safety evaluation, and lack a systematic and quantifiable measurement mechanism, making it difficult to form a universal standard for drug screening safety evaluation.
By constructing a multimodal standardized state, using discrete-time difference and integral iteration, a state evolution simulation is established to quantify the toxicity index, and a safety score is given by the accumulation of model prediction error and time-state deviation.
It enables global dynamics modeling of cell states under drug action, provides reliable evaluation of drug screening safety, improves the comparability and repeatability of results, reduces sources of error, and enhances the sensitivity and reliability of early screening.
Smart Images

Figure CN122087707A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, specifically to a multimodal data analysis method for drug screening and safety evaluation. Background Technology
[0002] The evaluation of safety and toxicity during drug screening is a crucial step in new drug development and preclinical research. Currently, drug safety evaluation primarily relies on traditional in vitro cell activity assays, animal experiments, and limited molecular biology detection methods. These traditional techniques are often limited to point-to-point analysis of single-modality data (such as cell viability, expression of specific biomarkers, etc.), failing to fully reflect the dynamic molecular response processes of cells under drug influence. In recent years, with the development of high-throughput sequencing, metabolomics, and other multi-omics detection methods, the acquisition of multi-modal and multi-timepoint biological data has gradually become possible, providing richer reference information for drug safety evaluation.
[0003] However, existing multimodal data analysis methods still have many shortcomings. On the one hand, commonly used analytical techniques such as principal component analysis (PCA), cluster analysis, and classical time series statistical methods often fail to model the global dynamic characteristics of cell states under drug influence, only capturing local or low-dimensional features and struggling to reveal the synergistic changes among multi-omics data. On the other hand, some recent literature has proposed using machine learning or deep learning models to predict and classify multimodal time series data, but these methods heavily rely on large amounts of labeled data, complex model structures, or readily available computational frameworks, lacking kinetic modeling oriented towards mechanism explanation, and offering limited biological interpretability of the results. Furthermore, existing methods often lack systematic and quantifiable measurement mechanisms for issues such as abnormal fluctuations in experimental data, inter-group differences, and overall deviations from dynamic processes, making it difficult to establish universal standards for evaluating drug screening safety.
[0004] Therefore, this case aims to propose a multimodal data analysis method for drug screening safety evaluation. Using a drug-free control group as a benchmark, a multimodal standardized state is constructed. A concise and efficient state evolution simulation is established through discrete time difference and integral iteration. Finally, the toxicity index is quantified by two methods: model prediction error and time-state deviation accumulation, and a relative safety score is established with the baseline control. Summary of the Invention
[0005] This invention provides a multimodal data analysis method for drug screening and safety evaluation, which helps to solve the problems mentioned in the background art.
[0006] This invention provides the following technical solution: a multimodal data analysis method for drug screening safety evaluation, comprising: The total duration of the experiment and the sample collection time interval were set, and gene expression level and metabolite concentration data were obtained at multiple sampling time points for the drug-treated group and the untreated control group. The statistical characteristics of the target gene and target metabolites were calculated based on the data of the untreated control group at each sampling time point. The dimensions of the two sets of data were standardized, and a comprehensive standardized state vector was constructed at each sampling time point. Between adjacent sampling time points, the changes in the standardized state vector of the drug treatment group are used to form discrete-time differential components, and an integral iteration of the state evolution function is established. Based on the deviation of the comprehensive standardized state vector from the initial state and the magnitude of its components, a state evolution function vector composed of component change rates is formed. Using the combined standardized state vectors of the two initial sampling time points as the initial predicted state, the state evolution function is called iteratively according to the sample collection time interval to obtain the theoretical state trajectories of the two sets at each sampling time point; At each sampling time point, the actual comprehensive standardized state vector of the drug treatment group is compared with the corresponding predicted vector to generate a multidimensional difference vector, and the average modeling error is obtained by summarizing the results. At each sampling time point, the absolute deviation of the two groups of actual and predicted vectors is taken dimension by dimension and accumulated in the time and state dimensions to obtain the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group. The safety score is calculated based on the relationship between the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group. The candidate drugs are then classified, ranked, and output.
[0007] Optionally, the setting of the total experimental duration and sample collection time interval, and the acquisition of gene expression levels and metabolite concentration data at multiple sampling time points for the drug-treated group and the untreated control group, specifically includes: Set the total duration of a single in vitro cell drug treatment experiment and the sample collection time interval between two adjacent sample collections. By dividing the total experimental duration by the sample collection time interval and adding one, obtain the total number of discrete sampling time points that are actually collected during the entire experimental cycle, so that the total number of discrete sampling time points is not less than two, and assign the corresponding physical time to each sampling time point in chronological order. The total number of target gene variables to be analyzed is preset. At each sampling time point, the expression level of all target genes in the drug treatment group is obtained. The expression level of each target gene at the current sampling time point is arranged into a data vector in a preset order, and the index number corresponding to each target gene and the total number of target gene variables are recorded. At each sampling time point, the expression levels of the same set of target genes as the drug-treated group were obtained from the untreated control group. The expression levels of each target gene at the current sampling time point were arranged into a data vector in the same order as the drug-treated group, and the index number corresponding to each target gene was recorded. The total number of target metabolite variables to be analyzed is preset. At each sampling time point, the concentration of all target metabolites in the drug treatment group is obtained. The concentration of each target metabolite at the current sampling time point is arranged into a data vector in a preset order, and the index number of each target metabolite and the total number of target metabolite variables are recorded. At each sampling time point, the concentrations of the same set of target metabolites as those in the drug-treated group were obtained for the untreated control group. The concentrations of each target metabolite at the current sampling time point were arranged into a data vector in the same order as those in the drug-treated group.
[0008] Optionally, the step of calculating the statistical characteristics of the target gene and target metabolite based on data from the untreated control group at each sampling time point, standardizing the dimensions of the two sets of data, and constructing a comprehensive standardized state vector at each sampling time point specifically includes: For the control group without medication, the arithmetic mean and standard deviation of the expression data of each target gene were calculated at all discrete sampling time points to obtain the average expression level and standard deviation of each target gene throughout the entire experimental period. For the untreated control group, the arithmetic mean and standard deviation of the concentration data of each target metabolite were calculated at all discrete sampling time points. The average concentration and concentration standard deviation of each target metabolite were obtained throughout the entire experimental period. When the standard deviation of a target gene or a target metabolite was equal to zero, the corresponding standard deviation value was replaced with one. For the drug-treated group, at each sampling time point, the original expression level of each target gene was subtracted from the average expression level of each target gene in the untreated control group, and then divided by the corresponding expression level standard deviation to obtain the standardized expression level of each target gene at the current sampling time point; the original concentration of each target metabolite was subtracted from the average concentration of each target metabolite in the untreated control group, and then divided by the corresponding concentration standard deviation to obtain the standardized concentration of each target metabolite at the current sampling time point. At each sampling time point, the standardized expression levels of all target genes and the standardized concentrations of all target metabolites in the drug treatment group are arranged in a preset order to construct a comprehensive standardized state vector of the drug treatment group at the current sampling time point. The dimension of the comprehensive standardized state vector is equal to the sum of the total number of target gene variables and the total number of target metabolite variables. At each sampling time point, the target gene expression level and target metabolite concentration of the untreated control group were standardized by subtracting the mean and dividing by the standard deviation in the same way as the drug-treated group. The standardized values were then arranged in the same order as the drug-treated group to construct the comprehensive standardized state vector of the untreated control group at the current sampling time point.
[0009] Optionally, the step of forming discrete-time differential components from the changes in the standardized state vector of the drug treatment group between adjacent sampling time points, and establishing an integral iteration of the state evolution function, specifically includes: In the comprehensive normalized state vector sequence of the drug treatment group, for each pair of adjacent sampling time points, the comprehensive normalized state vector of the next sampling time point is subtracted from the comprehensive normalized state vector of the previous sampling time point, and then divided by the sample collection time interval to obtain the average rate of change of the comprehensive normalized state vector in the corresponding discrete time interval. The average rates of change of all time intervals are then arranged in chronological order to form a discrete time differential component sequence. Based on the given state evolution function, during state prediction calculation, for each sampling time point, the predicted comprehensive standardized state vector of the previous sampling time point is used as the input of the state evolution function. The state evolution function outputs the rate of change vector of the comprehensive standardized state vector at the current sampling time point. The product of the output rate of change vector and the sample collection time interval is added to the predicted comprehensive standardized state vector of the previous sampling time point to obtain the predicted comprehensive standardized state vector at the current sampling time point, thus forming an integral iterative form of state update relationship.
[0010] Optionally, the step of forming a state evolution function vector composed of component change rates based on the deviation of the comprehensive standardized state vector from the initial state and the component amplitudes specifically includes: For any given comprehensive normalized state vector, record each component of the given comprehensive normalized state vector in all state dimensions and each component of the comprehensive normalized state vector at the initial sampling time point. Subtract the corresponding initial state component from each component of the current state, accumulate the difference in all state dimensions, and scale it according to the sample collection time interval to obtain the first cumulative amount that represents the overall deviation of the current state from the initial state. For a given comprehensive standardized state vector, the absolute value of the current component is calculated in all state dimensions, and the absolute values are accumulated in all state dimensions to obtain the second cumulative quantity that represents the overall magnitude level of the current state. At each state dimension, the result of subtracting the second cumulative amount from the first cumulative amount is taken as the rate of change value of the corresponding state dimension. All rate of change values are combined in the order of each state dimension to form the vector output of the state evolution function, which is used as the rate of change input of the comprehensive standardized state vector in the state update relationship in the form of integral iteration.
[0011] Optionally, the step of using the combined standardized state vectors of the two initial sampling time points as the initial predicted state, and iteratively calculating the state evolution function according to the sample collection time interval to obtain the two sets of theoretical state trajectories at each sampling time point specifically includes: In the drug treatment group, the comprehensive standardized state vector at the initial sampling time point is set as the predicted initial state, denoted as the predicted comprehensive standardized state vector of the drug treatment group at the initial sampling time point. In the drug treatment group, for each subsequent sampling time point other than the initial sampling time point, the predicted comprehensive normalized state vector of the previous sampling time point is used as the input of the state evolution function. The rate of change vector corresponding to the input predicted comprehensive normalized state vector is obtained through the state evolution function. The product of the output rate of change vector and the sample collection time interval is added to the predicted comprehensive normalized state vector of the previous sampling time point. The predicted comprehensive normalized state vector sequence of the drug treatment group at all sampling time points is generated in sequence. The generated sequence constitutes the theoretical state trajectory of the drug treatment group. In the untreated control group, the comprehensive standardized state vector at the initial sampling time point is set as the predicted initial state. Using the same integral iteration method as the drug-treated group, the state evolution function is calculated and the state is updated one by one at each sampling time point other than the initial sampling time point. This generates the predicted comprehensive standardized state vector sequence of the untreated control group at all sampling time points. The generated sequence constitutes the theoretical state trajectory of the untreated control group.
[0012] Optionally, the step of comparing the actual comprehensive standardized state vector of the drug treatment group with the corresponding predicted vector at each sampling time point to generate a multidimensional difference vector, and summing them to obtain the average modeling error, specifically includes: In the drug treatment group, for each sampling time point, the actual comprehensive standardized state vector of the current sampling time point is subtracted from the corresponding predicted comprehensive standardized state vector in each state dimension to obtain a difference vector covering all state dimensions. Each component of the difference vector represents the deviation between the predicted value and the actual value in the current state dimension. For each sampling time point, the difference vector is squared and accumulated for all state dimensions. Then, the square root operation is performed on the accumulated result to obtain the multidimensional error magnitude of the current sampling time point. The corresponding multidimensional error magnitude is used as the modeling error metric of the current sampling time point. Throughout the entire experimental period, the arithmetic mean of the multidimensional error magnitudes at all sampling time points other than the initial sampling time point is calculated to obtain the average modeling error for the entire experimental period.
[0013] Optionally, the step of taking the absolute deviation of the two sets of actual and predicted vectors dimension by dimension at each sampling time point and accumulating them in the time and state dimensions to obtain the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group specifically includes: In the drug treatment group, for each sampling time point other than the initial sampling time point, the actual comprehensive standardized state vector of the current sampling time point is subtracted from the corresponding predicted comprehensive standardized state vector in each state dimension, and the absolute value of the difference in each state dimension is taken. The absolute values are accumulated in all state dimensions to form the multidimensional absolute deviation sum of the current sampling time point. In the drug treatment group, the sum of the multidimensional absolute deviations of all sampling time points other than the initial sampling time point is accumulated and normalized by the product of the total number of discrete sampling time points and the dimension of the comprehensive standardized state vector to obtain the toxicity index of the drug treatment group. The toxicity index reflects the average absolute deviation level between the predicted comprehensive standardized state and the actual comprehensive standardized state of the drug treatment group under the combined effect of the time dimension and the state dimension. In the untreated control group, the same method as in the drug-treated group was used to subtract the actual and predicted comprehensive standardized state vectors from each sampling time point dimension by dimension and take the absolute value. The absolute deviations of each sampling time point were accumulated in the state dimension to obtain the sum of the multidimensional absolute deviations. Then, the deviations were accumulated in the time dimension and normalized by the product of the total number of discrete sampling time points and the dimension of the comprehensive standardized state vector to obtain the baseline toxicity index of the untreated control group.
[0014] Optionally, the step of calculating a safety score based on the relationship between the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group, completing the classification and ranking of candidate drugs, and outputting the results specifically includes: For each candidate drug, after obtaining the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group, if the baseline toxicity index of the untreated control group is greater than zero, the ratio of the toxicity index of the drug-treated group to the baseline toxicity index of the untreated control group is first calculated, and the ratio is used as the dimensionless toxicity ratio. Then, the dimensionless toxicity ratio is subtracted from one to obtain a safety score between zero and one. If the baseline toxicity index of the untreated control group is equal to zero and the toxicity index of the drug-treated group is equal to zero, the safety score of the corresponding candidate drug is set to one. If the baseline toxicity index of the untreated control group is equal to zero and the toxicity index of the drug-treated group is greater than zero, the safety score of the corresponding candidate drug is set to zero. Based on the safety score of each candidate drug, the drug safety is classified. When the safety score is not less than the first threshold, the candidate drug is classified as high safety level; when the safety score is between the second threshold and the first threshold, the candidate drug is classified as medium safety level; when the safety score is lower than the second threshold, the candidate drug is classified as low safety level. The first threshold is set to 0.8 and the second threshold is set to 0.5. The safety scores of all candidate drugs are compared and ranked from highest to lowest to form a sorted list. The sorting results are used as the output of the drug screening safety evaluation.
[0015] The present invention has the following beneficial effects: Simultaneously, two sets of data were collected: gene expression and metabolite concentration. A unified, standardized state vector was constructed by statistically analyzing the periodic mean and variance of the control group. Unlike existing approaches that focus only on a single biomarker or rely on manually set thresholds based on experience, this method achieves comparison and fusion of different types of data in the same space through dimensional unification and vectorization. This maximizes the preservation of the dynamic information of each indicator over time, avoiding biases caused by inconsistent data dimensions, and providing a reliable foundation for subsequent modeling. In reality, different laboratories often have significant differences in the absolute value range of the same indicator. This method eliminates batch effects through unified normalization, improving the comparability and repeatability of the results.
[0016] Between adjacent sampling time points, this scheme proposes to divide the difference of the standardized vectors by the time interval to obtain the average rate of change, and then use a simple iterative accumulation to predict subsequent states, rather than employing complex numerical solutions of differential equations. Unlike traditional continuous-time dynamics models that require solving high-order differential equations or using black-box machine learning for prediction, this innovative approach has the advantages of simple algorithm implementation and high computational efficiency, making it suitable for rapid simulation of high-dimensional vectors. In practical applications, this method can quickly generate theoretical state trajectories for the entire cycle with limited samples and computational resources, providing real-time prediction results for subsequent error assessment and toxicity index calculation.
[0017] By combining the cumulative deviation of the current state from the initial state with the sum of the amplitudes of each component, the rate of change for each dimension is calculated, avoiding the tedious parameter calibration or black-box neural network learning required by traditional models. The entire calculation is based on standardized experimental observation data, eliminating the need for additional hyperparameters or training data, thus ensuring the model's consistency and interpretability under different drug conditions.
[0018] By fully aligning the prediction initialization and iteration processes of the control and treatment groups, and generating two sets of trajectories in parallel using the same integral iteration strategy, this method effectively avoids the offset problem caused by the difference in model paths under control and treatment conditions. Compared with existing techniques that often employ separate simulations or subtract batch effects from the control group, this method completes parallel computation within the same iterative framework, ensuring the consistency of the two sets of trajectories on the same numerical and temporal scales. Its advantages include simplifying the preprocessing work for comparing the control and treatment groups, reducing error sources, and providing highly comparable data support for the subsequent direct calculation of error vectors and toxicity indices.
[0019] By subtracting the predicted and actual state vectors dimension-wise at each sampling time point and calculating the error magnitude, this scheme accurately characterizes the multidimensional deviation of the model prediction from reality, ultimately obtaining the average modeling error. Unlike traditional methods that only focus on the maximum deviation or mean square error, this method uses Euclidean norm to quantify the total vector error, providing a more holistic error metric while considering the contribution of different components. It can comprehensively reflect the degree of perturbation of the overall state of the cellular system by the drug, providing an objective numerical indicator for toxicity assessment and overcoming the risk of misjudgment or omission by a single indicator.
[0020] First, the state dimension of the multidimensional absolute deviations at each time point is accumulated. Then, the summation is performed across all time points and normalized to obtain a toxicity index that reflects the cumulative effect of both time and state. Compared with existing methods that judge toxicity based on instantaneous peak deviations or single endpoint indicators, the toxicity index constructed in this scheme has deep cumulative characteristics, reflecting both the cumulative effect of small deviations at multiple time points and the combined effect of various components. When drugs cause long-term cumulative perturbations to cellular systems, this index can detect potential toxicity risks in advance, improving the sensitivity and reliability of early screening.
[0021] Ultimately, this approach constructs a dimensionless safety score by calculating the ratio or difference between the toxicity index of the treatment group and the baseline index of the control group, and sets multi-level thresholds to classify candidate drugs into high, medium, and low safety levels. Unlike traditional empirical or expert scoring, this method provides clear numerical classification standards, reducing subjective bias; at the same time, the dimensionless safety score facilitates cross-sectional comparisons between different experimental systems and models. It transforms complex multimodal dynamic simulation results into intuitive and easy-to-understand classification results, meeting the comprehensive needs of efficiency, interpretability, and decision support in the drug screening process. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Example, refer to Figure 1 A multimodal data analysis method for drug screening safety evaluation, comprising: The total duration of the experiment and the sample collection time interval were set, and gene expression level and metabolite concentration data were obtained at multiple sampling time points for the drug-treated group and the untreated control group. The statistical characteristics of the target gene and target metabolites were calculated based on the data of the untreated control group at each sampling time point. The dimensions of the two sets of data were standardized, and a comprehensive standardized state vector was constructed at each sampling time point. Between adjacent sampling time points, the changes in the standardized state vector of the drug treatment group are used to form discrete-time differential components, and an integral iteration of the state evolution function is established. Based on the deviation of the comprehensive standardized state vector from the initial state and the magnitude of its components, a state evolution function vector composed of component change rates is formed. Using the combined standardized state vectors of the two initial sampling time points as the initial predicted state, the state evolution function is called iteratively according to the sample collection time interval to obtain the theoretical state trajectories of the two sets at each sampling time point; At each sampling time point, the actual comprehensive standardized state vector of the drug treatment group is compared with the corresponding predicted vector to generate a multidimensional difference vector, and the average modeling error is obtained by summarizing the results. At each sampling time point, the absolute deviation of the two groups of actual and predicted vectors is taken dimension by dimension and accumulated in the time and state dimensions to obtain the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group. The safety score is calculated based on the relationship between the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group. The candidate drugs are then classified, ranked, and output.
[0025] First, the proposed method simultaneously acquires gene expression and metabolite concentration data from both the drug-treated and control groups at multiple time points, avoiding the limitations of existing technologies that focus only on a single indicator or time segment. Second, the statistical characteristics of the control group data are incorporated into the standardization process, and a comprehensive standardized state vector is constructed simultaneously at each time point. This fundamentally solves the technical bottleneck of different biological indicators having different dimensions and being difficult to integrate, improving the feasibility and stability of multi-level data fusion. Subsequently, the state evolution function is defined using discrete-time differential components and state deviation amplitudes, preserving the temporal changes of the dynamic system while eliminating the high-cost solution process of traditional continuous differential equations, significantly improving computational efficiency. Next, the initial states of the two groups are iteratively generated in parallel to produce the theoretical state trajectory of the entire cycle, naturally aligning the control and treatment results and avoiding the tedious steps of subsequent correction or batch effect correction. Based on the average error of the multi-dimensional difference vector between prediction and reality, the model fitting deviation is quantified as a whole, solving the incomplete problem of existing methods that only take a single dimension or peak value for error. Based on the cumulative effect of time and state dimensions, a toxicity index is constructed, which can simultaneously reflect the cumulative effect over time and the interactive deviation of various dimensions of the system, making it more sensitive and comprehensive than single-point or single-dimensional indicators. Finally, the toxicity index is ratioed to the baseline index of the control group to give a dimensionless safety score, which is output in stages according to a set threshold, ensuring that the score is intuitive and easy to understand, as well as comparability and repeatability under different experimental conditions.
[0026] The process involves setting the total experimental duration and sample collection time intervals, and obtaining gene expression levels and metabolite concentration data for both the drug-treated group and the untreated control group at multiple sampling time points. Specifically, this includes: Set the total duration of a single in vitro cell drug treatment experiment and the sample collection time interval between two adjacent sample collections. By dividing the total experimental duration by the sample collection time interval and adding one, obtain the total number of discrete sampling time points that are actually collected during the entire experimental cycle, so that the total number of discrete sampling time points is not less than two, and assign the corresponding physical time to each sampling time point in chronological order. The total number of target gene variables to be analyzed is preset. At each sampling time point, the expression level of all target genes in the drug treatment group is obtained. The expression level of each target gene at the current sampling time point is arranged into a data vector in a preset order, and the index number corresponding to each target gene and the total number of target gene variables are recorded. At each sampling time point, the expression levels of the same set of target genes as the drug-treated group were obtained from the untreated control group. The expression levels of each target gene at the current sampling time point were arranged into a data vector in the same order as the drug-treated group, and the index number corresponding to each target gene was recorded. The total number of target metabolite variables to be analyzed is preset. At each sampling time point, the concentration of all target metabolites in the drug treatment group is obtained. The concentration of each target metabolite at the current sampling time point is arranged into a data vector in a preset order, and the index number of each target metabolite and the total number of target metabolite variables are recorded. At each sampling time point, the concentrations of the same set of target metabolites as those in the drug-treated group were obtained for the untreated control group. The concentrations of each target metabolite at the current sampling time point were arranged into a data vector in the same order as those in the drug-treated group.
[0027] The total duration of a single in vitro cell drug treatment experiment was set to... Set a fixed time interval between two adjacent sample collections. Therefore, the total number of discrete time points where data was actually collected during the entire experimental period is and satisfy ; Set the corresponding time index The physical time is Then we have: , ;in, Indexes for discrete time points; At every point in time The following two types of data were collected from the drug-treated group and the untreated control group: S11, Gene expression level data vector of the drug-treated group: ;in, For the drug treatment group at time All collected A column vector consisting of the expression levels of each gene; The first in the drug treatment group Each gene at time The amount of expression; For gene indexing; The total number of genetic variables analyzed; S12, Gene expression data vector of the untreated control group: ;in, For the untreated control group at time [time] All collected A column vector consisting of gene expression levels; The first in the untreated control group Each gene at time The amount of expression; S21, Metabolite Concentration Data Vector for Drug-Treated Groups: ;in, For the drug treatment group at time All collected A column vector consisting of the concentrations of each metabolite; The first in the drug treatment group A metabolite at time The concentration; The total number of metabolite variables analyzed; For metabolite indexing; S22, Metabolite Concentration Data Vector for the Untreated Control Group: ;in, For the untreated control group at time [time] All collected A column vector consisting of the concentrations of each metabolite; The first in the untreated control group A metabolite at time The concentration.
[0028] The statistical characteristics of the target gene and target metabolite are calculated based on data from the untreated control group at each sampling time point. The dimensions of the two sets of data are standardized, and a comprehensive standardized state vector is constructed at each sampling time point. Specifically, this includes: For the control group without medication, the arithmetic mean and standard deviation of the expression data of each target gene were calculated at all discrete sampling time points to obtain the average expression level and standard deviation of each target gene throughout the entire experimental period. For the untreated control group, the arithmetic mean and standard deviation of the concentration data of each target metabolite were calculated at all discrete sampling time points. The average concentration and concentration standard deviation of each target metabolite were obtained throughout the entire experimental period. When the standard deviation of a target gene or a target metabolite was equal to zero, the corresponding standard deviation value was replaced with one. For the drug-treated group, at each sampling time point, the original expression level of each target gene was subtracted from the average expression level of each target gene in the untreated control group, and then divided by the corresponding expression level standard deviation to obtain the standardized expression level of each target gene at the current sampling time point; the original concentration of each target metabolite was subtracted from the average concentration of each target metabolite in the untreated control group, and then divided by the corresponding concentration standard deviation to obtain the standardized concentration of each target metabolite at the current sampling time point. At each sampling time point, the standardized expression levels of all target genes and the standardized concentrations of all target metabolites in the drug treatment group are arranged in a preset order to construct a comprehensive standardized state vector of the drug treatment group at the current sampling time point. The dimension of the comprehensive standardized state vector is equal to the sum of the total number of target gene variables and the total number of target metabolite variables. At each sampling time point, the target gene expression level and target metabolite concentration of the untreated control group were standardized by subtracting the mean and dividing by the standard deviation in the same way as the drug-treated group. The standardized values were then arranged in the same order as the drug-treated group to construct the comprehensive standardized state vector of the untreated control group at the current sampling time point.
[0029] Steps S201 and S202 were performed to analyze gene expression levels and metabolite concentrations in the untreated control group, and the mean and standard deviation of each variable were calculated at all time points. S201. Calculate the mean and standard deviation of gene expression variables: , ;in, The first in the untreated control group Each gene in all Average expression level at each sampling time point; The first in the untreated control group Each gene in all Standard deviation at each sampling time; S202. Calculate the mean and standard deviation of metabolite concentration variables: , ;in, The first in the untreated control group Each metabolite in all The average concentration at each sampling time; The first in the untreated control group Each metabolite in all The standard deviation of concentration at each sampling time; when season ; when season ; The data from the drug treatment groups were standardized to obtain dimensionless variables: , ;in, The first in the drug treatment group Each gene at time Standardized expression level; The first in the drug treatment group A metabolite at time Standardized concentration; The standardized state variable vector of the drug treatment group is constructed as follows: ;in, For the drug treatment group at time The comprehensive state vector; The total dimension of the state vector; For dimension The real vector space; The data from the untreated control group were standardized to construct a state variable vector for the control group: , ;in, The first in the untreated control group Each gene at time Standardized expression level; The first in the untreated control group A metabolite at time Standardized concentration; ;in, For the untreated control group at time [time] The comprehensive state vector.
[0030] The process of forming discrete-time differential components from the changes in the standardized state vector of the drug treatment group between adjacent sampling time points, and establishing an integral iteration of the state evolution function, specifically includes: In the comprehensive normalized state vector sequence of the drug treatment group, for each pair of adjacent sampling time points, the comprehensive normalized state vector of the next sampling time point is subtracted from the comprehensive normalized state vector of the previous sampling time point, and then divided by the sample collection time interval to obtain the average rate of change of the comprehensive normalized state vector in the corresponding discrete time interval. The average rates of change of all time intervals are then arranged in chronological order to form a discrete time differential component sequence. Based on the given state evolution function, during state prediction calculation, for each sampling time point, the predicted comprehensive standardized state vector of the previous sampling time point is used as the input of the state evolution function. The state evolution function outputs the rate of change vector of the comprehensive standardized state vector at the current sampling time point. The product of the output rate of change vector and the sample collection time interval is added to the predicted comprehensive standardized state vector of the previous sampling time point to obtain the predicted comprehensive standardized state vector at the current sampling time point, thus forming an integral iterative form of state update relationship.
[0031] Calculation in discrete time interval Average rate of change of internal integrated state vector Specifically: , ; Calculate the predicted drug treatment group at time t under the action of the kinetic model. State vector estimate Specifically: ;in, The state evolution function takes a single input. A dimensionless state vector is output as a state change rate vector of the same dimension.
[0032] The process of forming a state evolution function vector composed of component change rates based on the deviation of the comprehensive standardized state vector from the initial state and the component amplitudes specifically includes: For any given comprehensive normalized state vector, record each component of the given comprehensive normalized state vector in all state dimensions and each component of the comprehensive normalized state vector at the initial sampling time point. Subtract the corresponding initial state component from each component of the current state, accumulate the difference in all state dimensions, and scale it according to the sample collection time interval to obtain the first cumulative amount that represents the overall deviation of the current state from the initial state. For a given comprehensive standardized state vector, the absolute value of the current component is calculated in all state dimensions, and the absolute values are accumulated in all state dimensions to obtain the second cumulative quantity that represents the overall magnitude level of the current state. At each state dimension, the result of subtracting the second cumulative amount from the first cumulative amount is taken as the rate of change value of the corresponding state dimension. All rate of change values are combined in the order of each state dimension to form the vector output of the state evolution function, which is used as the rate of change input of the comprehensive standardized state vector in the state update relationship in the form of integral iteration.
[0033] For each dimension of the state variable vector, let its index be denoted as . The components of the state evolution function are constructed as follows: ;in, For the component index in the state vector dimension; is the independent variable of the state evolution function; For the independent variable in a given state Next, the The rate of change of each state component; These are the indices of the state components when summing within the state evolution function; State vector The One component; For the first Each state component at the initial sampling time The corresponding standardized value; All The combined state evolution function vector is: ;in, For vector functions, by component functions The overall state evolution function.
[0034] The process involves using a standardized state vector from two initial sampling time points as the initial predicted state, and iteratively calculating the state evolution function according to the sample collection time interval to obtain two theoretical state trajectories at each sampling time point. Specifically, this includes: In the drug treatment group, the comprehensive standardized state vector at the initial sampling time point is set as the predicted initial state, denoted as the predicted comprehensive standardized state vector of the drug treatment group at the initial sampling time point. In the drug treatment group, for each subsequent sampling time point other than the initial sampling time point, the predicted comprehensive normalized state vector of the previous sampling time point is used as the input of the state evolution function. The rate of change vector corresponding to the input predicted comprehensive normalized state vector is obtained through the state evolution function. The product of the output rate of change vector and the sample collection time interval is added to the predicted comprehensive normalized state vector of the previous sampling time point. The predicted comprehensive normalized state vector sequence of the drug treatment group at all sampling time points is generated in sequence. The generated sequence constitutes the theoretical state trajectory of the drug treatment group. In the untreated control group, the comprehensive standardized state vector at the initial sampling time point is set as the predicted initial state. Using the same integral iteration method as the drug-treated group, the state evolution function is calculated and the state is updated one by one at each sampling time point other than the initial sampling time point. This generates the predicted comprehensive standardized state vector sequence of the untreated control group at all sampling time points. The generated sequence constitutes the theoretical state trajectory of the untreated control group.
[0035] The actual initial state of the drug treatment group is taken as the predicted initial state. ;in, For the drug treatment group at the initial sampling time The actual integrated state vector; This represents the predicted state vector of the dynamic model at the initial moment. for to The predicted state sequence of the drug treatment group is generated through integral iteration: ; The actual initial state of the control group without medication was taken as the predicted initial state. ;in, The control group without medication was sampled at the initial sampling time. The actual integrated state vector; This represents the predicted state vector of the untreated control group at the initial time. for to The predicted state sequence of the untreated control group is generated through integral iteration: ;in, The untreated control group obtained after integral iteration at time 1000 The predicted state vector.
[0036] The step involves comparing the actual comprehensive standardized state vector of the drug treatment group with the corresponding predicted vector at each sampling time point to generate a multidimensional difference vector, which is then summarized to obtain the average modeling error. Specifically, this includes: In the drug treatment group, for each sampling time point, the actual comprehensive standardized state vector of the current sampling time point is subtracted from the corresponding predicted comprehensive standardized state vector in each state dimension to obtain a difference vector covering all state dimensions. Each component of the difference vector represents the deviation between the predicted value and the actual value in the current state dimension. For each sampling time point, the difference vector is squared and accumulated for all state dimensions. Then, the square root operation is performed on the accumulated result to obtain the multidimensional error magnitude of the current sampling time point. The corresponding multidimensional error magnitude is used as the modeling error metric of the current sampling time point. Throughout the entire experimental period, the arithmetic mean of the multidimensional error magnitudes at all sampling time points other than the initial sampling time point is calculated to obtain the average modeling error for the entire experimental period.
[0037] For each time point Calculate the difference vector between the actual and predicted states of the drug-treated group: ;in, For the drug treatment group at time The difference vector between the predicted state and the actual state; Calculate the error modulus at this time point: ;in, Error vector The One component; For at any time Euclidean norm of multidimensional state error; Calculate the average modeling error over the entire experimental period: ;in, In the time interval Error modulus at each time point The average value.
[0038] The absolute deviation between the actual and predicted vectors at each sampling time point is calculated dimension-wise and accumulated along the time and state dimensions to obtain the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group. Specifically, this includes: In the drug treatment group, for each sampling time point other than the initial sampling time point, the actual comprehensive standardized state vector of the current sampling time point is subtracted from the corresponding predicted comprehensive standardized state vector in each state dimension, and the absolute value of the difference in each state dimension is taken. The absolute values are accumulated in all state dimensions to form the multidimensional absolute deviation sum of the current sampling time point. In the drug treatment group, the sum of the multidimensional absolute deviations of all sampling time points other than the initial sampling time point is accumulated and normalized by the product of the total number of discrete sampling time points and the dimension of the comprehensive standardized state vector to obtain the toxicity index of the drug treatment group. The toxicity index reflects the average absolute deviation level between the predicted comprehensive standardized state and the actual comprehensive standardized state of the drug treatment group under the combined effect of the time dimension and the state dimension. In the untreated control group, the same method as in the drug-treated group was used to subtract the actual and predicted comprehensive standardized state vectors from each sampling time point dimension by dimension and take the absolute value. The absolute deviations of each sampling time point were accumulated in the state dimension to obtain the sum of the multidimensional absolute deviations. Then, the deviations were accumulated in the time dimension and normalized by the product of the total number of discrete sampling time points and the dimension of the comprehensive standardized state vector to obtain the baseline toxicity index of the untreated control group.
[0039] The toxicity index of the drug treatment group was constructed as follows: ;in, The toxicity index of the drug-treated group; For the drug treatment group at time state vector The One component; For the drug treatment group at time Predicted state vector The One component; The toxicity index of the drug-treated group; In the control group without medication, the toxicity index of the control group was constructed as follows: ;in, State vector of the untreated control group The One component; Predicted state vector for the untreated control group The One component; The baseline toxicity index is for the control group without medication.
[0040] The process of calculating a safety score based on the relationship between the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group, completing the classification and ranking of candidate drugs, and outputting the results specifically includes: For each candidate drug, after obtaining the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group, if the baseline toxicity index of the untreated control group is greater than zero, the ratio of the toxicity index of the drug-treated group to the baseline toxicity index of the untreated control group is first calculated, and the ratio is used as the dimensionless toxicity ratio. Then, the dimensionless toxicity ratio is subtracted from one to obtain a safety score between zero and one. If the baseline toxicity index of the untreated control group is equal to zero and the toxicity index of the drug-treated group is equal to zero, the safety score of the corresponding candidate drug is set to one. If the baseline toxicity index of the untreated control group is equal to zero and the toxicity index of the drug-treated group is greater than zero, the safety score of the corresponding candidate drug is set to zero. Based on the safety score of each candidate drug, the drug safety is classified. When the safety score is not less than the first threshold, the candidate drug is classified as high safety level; when the safety score is between the second threshold and the first threshold, the candidate drug is classified as medium safety level; when the safety score is lower than the second threshold, the candidate drug is classified as low safety level. The first threshold is set to 0.8 and the second threshold is set to 0.5. The safety scores of all candidate drugs are compared and ranked from highest to lowest to form a sorted list. The sorting results are used as the output of the drug screening safety evaluation.
[0041] Construct a safety score for the drug under this kinetic model. Specifically: ; According to safety rating Drug classification: when At that time, it was rated as high safety; when At that time, it was assessed as medium safety; when At that time, it was rated as low safety; Calculate separately for all candidate drugs ,according to The drugs are sorted from largest to smallest, and the sorting results are used as the output of the drug screening safety evaluation.
[0042] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0043] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A multimodal data analysis method for drug screening and safety evaluation, characterized in that, include: The total duration of the experiment and the sample collection time interval were set, and gene expression level and metabolite concentration data were obtained at multiple sampling time points for the drug-treated group and the untreated control group. The statistical characteristics of the target gene and target metabolites were calculated based on the data of the untreated control group at each sampling time point. The dimensions of the two sets of data were standardized, and a comprehensive standardized state vector was constructed at each sampling time point. Between adjacent sampling time points, the changes in the standardized state vector of the drug treatment group are used to form discrete-time differential components, and an integral iteration of the state evolution function is established. Based on the deviation of the comprehensive standardized state vector from the initial state and the magnitude of its components, a state evolution function vector composed of component change rates is formed. Using the combined standardized state vectors of the two initial sampling time points as the initial predicted state, the state evolution function is called iteratively according to the sample collection time interval to obtain the theoretical state trajectories of the two sets at each sampling time point; At each sampling time point, the actual comprehensive standardized state vector of the drug treatment group is compared with the corresponding predicted vector to generate a multidimensional difference vector, and the average modeling error is obtained by summarizing the results. At each sampling time point, the absolute deviation of the two groups of actual and predicted vectors is taken dimension by dimension and accumulated in the time and state dimensions to obtain the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group. The safety score is calculated based on the relationship between the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group. Candidate drugs are then classified, ranked, and output.
2. The multimodal data analysis method for drug screening safety evaluation according to claim 1, characterized in that, The process involves setting the total experimental duration and sample collection time intervals, and obtaining gene expression levels and metabolite concentration data for both the drug-treated group and the untreated control group at multiple sampling time points. Specifically, this includes: Set the total duration of a single in vitro cell drug treatment experiment and the sample collection time interval between two adjacent sample collections. By dividing the total experimental duration by the sample collection time interval and adding one, obtain the total number of discrete sampling time points that are actually collected during the entire experimental cycle, so that the total number of discrete sampling time points is not less than two, and assign the corresponding physical time to each sampling time point in chronological order. The total number of target gene variables to be analyzed is preset. At each sampling time point, the expression level of all target genes in the drug treatment group is obtained. The expression level of each target gene at the current sampling time point is arranged into a data vector in a preset order, and the index number corresponding to each target gene and the total number of target gene variables are recorded. At each sampling time point, the expression levels of the same set of target genes as the drug-treated group were obtained from the untreated control group. The expression levels of each target gene at the current sampling time point were arranged into a data vector in the same order as the drug-treated group, and the index number corresponding to each target gene was recorded. The total number of target metabolite variables to be analyzed is preset. At each sampling time point, the concentration of all target metabolites in the drug treatment group is obtained. The concentration of each target metabolite at the current sampling time point is arranged into a data vector in a preset order, and the index number of each target metabolite and the total number of target metabolite variables are recorded. At each sampling time point, the concentrations of the same set of target metabolites as those in the drug-treated group were obtained for the untreated control group. The concentrations of each target metabolite at the current sampling time point were arranged into a data vector in the same order as those in the drug-treated group.
3. The multimodal data analysis method for drug screening safety evaluation according to claim 2, characterized in that, The statistical characteristics of the target gene and target metabolite are calculated based on data from the untreated control group at each sampling time point. The dimensions of the two sets of data are standardized, and a comprehensive standardized state vector is constructed at each sampling time point. Specifically, this includes: For the control group without medication, the arithmetic mean and standard deviation of the expression data of each target gene were calculated at all discrete sampling time points to obtain the average expression level and standard deviation of each target gene throughout the entire experimental period. For the untreated control group, the arithmetic mean and standard deviation of the concentration data of each target metabolite were calculated at all discrete sampling time points. The average concentration and concentration standard deviation of each target metabolite were obtained throughout the entire experimental period. When the standard deviation of a target gene or a target metabolite was equal to zero, the corresponding standard deviation value was replaced with one. For the drug-treated group, at each sampling time point, the original expression level of each target gene was subtracted from the average expression level of each target gene in the untreated control group, and then divided by the corresponding expression level standard deviation to obtain the standardized expression level of each target gene at the current sampling time point; the original concentration of each target metabolite was subtracted from the average concentration of each target metabolite in the untreated control group, and then divided by the corresponding concentration standard deviation to obtain the standardized concentration of each target metabolite at the current sampling time point. At each sampling time point, the standardized expression levels of all target genes and the standardized concentrations of all target metabolites in the drug treatment group are arranged in a preset order to construct a comprehensive standardized state vector of the drug treatment group at the current sampling time point. The dimension of the comprehensive standardized state vector is equal to the sum of the total number of target gene variables and the total number of target metabolite variables. At each sampling time point, the target gene expression level and target metabolite concentration of the untreated control group were standardized by subtracting the mean and dividing by the standard deviation in the same way as the drug-treated group. The standardized values were then arranged in the same order as the drug-treated group to construct the comprehensive standardized state vector of the untreated control group at the current sampling time point.
4. The multimodal data analysis method for drug screening safety evaluation according to claim 3, characterized in that, The process of forming discrete-time differential components from the changes in the standardized state vector of the drug treatment group between adjacent sampling time points, and establishing an integral iteration of the state evolution function, specifically includes: In the comprehensive normalized state vector sequence of the drug treatment group, for each pair of adjacent sampling time points, the comprehensive normalized state vector of the next sampling time point is subtracted from the comprehensive normalized state vector of the previous sampling time point, and then divided by the sample collection time interval to obtain the average rate of change of the comprehensive normalized state vector in the corresponding discrete time interval. The average rates of change of all time intervals are then arranged in chronological order to form a discrete time differential component sequence. Based on the given state evolution function, during state prediction calculation, for each sampling time point, the predicted comprehensive standardized state vector of the previous sampling time point is used as the input of the state evolution function. The state evolution function outputs the rate of change vector of the comprehensive standardized state vector at the current sampling time point. The product of the output rate of change vector and the sample collection time interval is added to the predicted comprehensive standardized state vector of the previous sampling time point to obtain the predicted comprehensive standardized state vector at the current sampling time point, thus forming an integral iterative form of state update relationship.
5. The multimodal data analysis method for drug screening safety evaluation according to claim 4, characterized in that, The process of forming a state evolution function vector composed of component change rates based on the deviation of the comprehensive standardized state vector from the initial state and the component amplitudes specifically includes: For any given comprehensive normalized state vector, record each component of the given comprehensive normalized state vector in all state dimensions and each component of the comprehensive normalized state vector at the initial sampling time point. Subtract the corresponding initial state component from each component of the current state, accumulate the difference in all state dimensions, and scale it according to the sample collection time interval to obtain the first cumulative amount that represents the overall deviation of the current state from the initial state. For a given comprehensive standardized state vector, the absolute value of the current component is calculated in all state dimensions, and the absolute values are accumulated in all state dimensions to obtain the second cumulative quantity that represents the overall magnitude level of the current state. At each state dimension, the result of subtracting the second cumulative amount from the first cumulative amount is taken as the rate of change value of the corresponding state dimension. All rate of change values are combined in the order of each state dimension to form the vector output of the state evolution function, which is used as the rate of change input of the comprehensive standardized state vector in the state update relationship in the form of integral iteration.
6. The multimodal data analysis method for drug screening safety evaluation according to claim 5, characterized in that, The process involves using a standardized state vector from two initial sampling time points as the initial predicted state, and iteratively calculating the state evolution function according to the sample collection time interval to obtain two theoretical state trajectories at each sampling time point. Specifically, this includes: In the drug treatment group, the comprehensive standardized state vector at the initial sampling time point is set as the predicted initial state, denoted as the predicted comprehensive standardized state vector of the drug treatment group at the initial sampling time point. In the drug treatment group, for each subsequent sampling time point other than the initial sampling time point, the predicted comprehensive normalized state vector of the previous sampling time point is used as the input of the state evolution function. The rate of change vector corresponding to the input predicted comprehensive normalized state vector is obtained through the state evolution function. The product of the output rate of change vector and the sample collection time interval is added to the predicted comprehensive normalized state vector of the previous sampling time point. The predicted comprehensive normalized state vector sequence of the drug treatment group at all sampling time points is generated in sequence. The generated sequence constitutes the theoretical state trajectory of the drug treatment group. In the untreated control group, the comprehensive standardized state vector at the initial sampling time point is set as the predicted initial state. Using the same integral iteration method as the drug-treated group, the state evolution function is calculated and the state is updated one by one at each sampling time point other than the initial sampling time point. This generates the predicted comprehensive standardized state vector sequence of the untreated control group at all sampling time points. The generated sequence constitutes the theoretical state trajectory of the untreated control group.
7. The multimodal data analysis method for drug screening safety evaluation according to claim 6, characterized in that, The step involves comparing the actual comprehensive standardized state vector of the drug treatment group with the corresponding predicted vector at each sampling time point to generate a multidimensional difference vector, which is then summarized to obtain the average modeling error. Specifically, this includes: In the drug treatment group, for each sampling time point, the actual comprehensive standardized state vector of the current sampling time point is subtracted from the corresponding predicted comprehensive standardized state vector in each state dimension to obtain a difference vector covering all state dimensions. Each component of the difference vector represents the deviation between the predicted value and the actual value in the current state dimension. For each sampling time point, the difference vector is squared and accumulated for all state dimensions. Then, the square root operation is performed on the accumulated result to obtain the multidimensional error magnitude of the current sampling time point. The corresponding multidimensional error magnitude is used as the modeling error metric of the current sampling time point. Throughout the entire experimental period, the arithmetic mean of the multidimensional error magnitudes at all sampling time points other than the initial sampling time point is calculated to obtain the average modeling error for the entire experimental period.
8. The multimodal data analysis method for drug screening safety evaluation according to claim 7, characterized in that, The absolute deviation between the actual and predicted vectors at each sampling time point is calculated dimension-wise and accumulated along the time and state dimensions to obtain the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group. Specifically, this includes: In the drug treatment group, for each sampling time point other than the initial sampling time point, the actual comprehensive standardized state vector of the current sampling time point is subtracted from the corresponding predicted comprehensive standardized state vector in each state dimension, and the absolute value of the difference in each state dimension is taken. The absolute values are accumulated in all state dimensions to form the multidimensional absolute deviation sum of the current sampling time point. In the drug treatment group, the sum of the multidimensional absolute deviations of all sampling time points other than the initial sampling time point is accumulated and normalized by the product of the total number of discrete sampling time points and the dimension of the comprehensive standardized state vector to obtain the toxicity index of the drug treatment group. The toxicity index reflects the average absolute deviation level between the predicted comprehensive standardized state and the actual comprehensive standardized state of the drug treatment group under the combined effect of the time dimension and the state dimension. In the untreated control group, the same method as in the drug-treated group was used to subtract the actual and predicted comprehensive standardized state vectors from each sampling time point dimension by dimension and take the absolute value. The absolute deviations of each sampling time point were accumulated in the state dimension to obtain the sum of the multidimensional absolute deviations. Then, the deviations were accumulated in the time dimension and normalized by the product of the total number of discrete sampling time points and the dimension of the comprehensive standardized state vector to obtain the baseline toxicity index of the untreated control group.
9. The multimodal data analysis method for drug screening safety evaluation according to claim 8, characterized in that, The process of calculating a safety score based on the relationship between the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group, completing the classification and ranking of candidate drugs, and outputting the results specifically includes: For each candidate drug, after obtaining the toxicity index of the drug-treated group and the baseline toxicity index of the untreated control group, if the baseline toxicity index of the untreated control group is greater than zero, the ratio of the toxicity index of the drug-treated group to the baseline toxicity index of the untreated control group is first calculated, and the ratio is used as the dimensionless toxicity ratio. Then, the dimensionless toxicity ratio is subtracted from one to obtain a safety score between zero and one. If the baseline toxicity index of the untreated control group is equal to zero and the toxicity index of the drug-treated group is equal to zero, the safety score of the corresponding candidate drug is set to one. If the baseline toxicity index of the untreated control group is equal to zero and the toxicity index of the drug-treated group is greater than zero, the safety score of the corresponding candidate drug is set to zero. Based on the safety score of each candidate drug, the drug safety is classified. When the safety score is not less than the first threshold, the candidate drug is classified as high safety level; when the safety score is between the second threshold and the first threshold, the candidate drug is classified as medium safety level; when the safety score is lower than the second threshold, the candidate drug is classified as low safety level. The first threshold is set to 0.8 and the second threshold is set to 0.
5. The safety scores of all candidate drugs are compared and ranked from highest to lowest to form a sorted list. The sorting results are used as the output of the drug screening safety evaluation.