Method and system for predicting lung injury repairing effect of isoliquiritigenin based on machine learning
By acquiring molecular structure and biomarker data of isoliquiritigenin, calculating interaction forces and permeability indices, and combining multiple regression and neural network models, the problem of insufficient prediction accuracy of isoliquiritigenin in repairing lung injury in existing technologies has been solved, achieving accurate effect prediction and optimization guidance.
Patent Information
- Application Number
- CN202511708331.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies lack comprehensive analysis of molecular mechanisms and biomarkers, resulting in insufficient predictive accuracy and overly simplistic predictions of isoliquiritigenin in lung injury repair.
By acquiring molecular structure data of isoliquiritigenin and biomarker data related to lung injury, molecular interaction forces and cell permeability indicators are calculated. Combined with historical experimental repair effect data, a multivariate regression analyzer and neural network model are used to construct an effect prediction model and output the predicted value of repair effect.
It enables accurate prediction of the lung injury repair effect of isoliquiritigenin, provides decision support for drug development, reduces prediction bias caused by data bias, and improves prediction accuracy and reliability.
Smart Images

Figure CN121545783A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of drug effect prediction, in particular to a glycyrrhizin repair lung injury effect prediction method and system based on machine learning. BACKGROUND
[0002] Acute lung injury (ALI) is a common respiratory critical illness with complex pathogenesis involving inflammation, oxidative stress and other pathological processes. Glycyrrhizin, as a natural flavonoid compound, has anti-inflammatory and antioxidant properties and shows potential therapeutic value in lung injury repair.
[0003] The existing drug effect prediction relies on in vitro experiments or simple statistical models, lacks comprehensive analysis of molecular mechanisms and biomarkers, and cannot accurately predict the actual repair effect in complex biological environments, resulting in insufficient prediction accuracy and overly one-sided prediction.
[0004] In summary, there is a need for a high-precision prediction method based on machine learning to evaluate the repair effect of glycyrrhizin and achieve precise evaluation and optimized guidance of lung injury repair effect. SUMMARY
[0005] The present application provides a glycyrrhizin repair lung injury effect prediction method and system based on machine learning, aiming to solve the problem of lack of comprehensive analysis of molecular mechanisms and biomarkers in the prior art, insufficient prediction accuracy and overly one-sided prediction.
[0006] In view of the above problems, the present application provides a glycyrrhizin repair lung injury effect prediction method and system based on machine learning.
[0007] In a first aspect, the present application provides a glycyrrhizin repair lung injury effect prediction method based on machine learning, comprising:
[0008] Obtaining molecular structure data of glycyrrhizin, lung injury related biomarker data and historical experimental repair effect data;
[0009] Based on the molecular structure data and biomarker data, calculating the molecular interaction force index and cell permeability index of glycyrrhizin, combining the historical experimental repair effect data, and obtaining the repair effect coefficient by multivariate regression analysis;
[0010] According to the repair effect coefficient, configuring the effect confidence interval, and optimizing the range of the molecular interaction force index and cell permeability index to obtain the optimized biological activity index;
[0011] The effect prediction model based on the neural network is constructed, the optimized bioactivity index of the target isoliquiritigenin is input into the effect prediction model, and a prediction value of a lung injury repair effect is output.
[0012] In a second aspect, the application provides an isoliquiritigenin lung injury repair effect prediction system based on machine learning, comprising:
[0013] A data acquisition module is configured to acquire molecular structure data of isoliquiritigenin, lung injury related biomarker data, and historical experimental repair effect data.
[0014] A repair effect coefficient analysis module is configured to calculate a molecular interaction force index and a cell permeability index of isoliquiritigenin based on the molecular structure data and the biomarker data, combine the historical experimental repair effect data, and analyze a repair effect coefficient through a multiple regression analyzer.
[0015] An optimized bioactivity index calculation module is configured to configure an effect confidence interval according to the repair effect coefficient, and perform range optimization on the molecular interaction force index and the cell permeability index to obtain an optimized bioactivity index.
[0016] An effect prediction model construction module is configured to construct an effect prediction model based on a neural network, input the optimized bioactivity index of the target isoliquiritigenin into the effect prediction model, and output a prediction value of a lung injury repair effect.
[0017] One or more technical solutions provided in the application have at least the following technical effects or advantages:
[0018] The application first integrates multi-source data to provide comprehensive input for subsequent analysis, reduces prediction deviation caused by one-sided data, calculates a molecular interaction force index and a cell permeability index based on molecular structure data and biomarker data, and improves the representativeness of the indexes. A multiple regression analyzer is constructed combined with historical experimental data, an effect confidence interval is configured, and range optimization is performed on the molecular interaction force index and the cell permeability index. Meanwhile, a neural network model is constructed to efficiently process the molecular structure data and the biomarker data of isoliquiritigenin, and output a quantitative repair effect prediction value, and provide decision support for drug development to realize accurate prediction and optimization guidance of the repair effect. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0020] Figure 1 A flowchart of a method for predicting the effect of glycyrrhizin on repairing lung injury based on machine learning is shown in the figure.
[0021] Figure 2 A structural diagram of a system for predicting the effect of glycyrrhizin on repairing lung injury based on machine learning is shown in the figure.
[0022] In the figure, the explanations of the respective reference signs are as follows.
[0023] Data acquisition module 11; repair effect coefficient analysis module 12; optimized biological activity index calculation module 13; effect prediction model construction module 14. DETAILED DESCRIPTION
[0024] The present application provides a method and system for predicting the effect of glycyrrhizin on repairing lung injury based on machine learning, which is used to solve the problems of lack of comprehensive analysis of molecular mechanisms and biomarkers, insufficient prediction accuracy, and one-sided prediction in the prior art.
[0025] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0026] It should be noted that the terms "include" and "have" are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units need not be limited to only those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to these processes, methods, products or devices.
[0027] Embodiment one, as shown in the figure, the present application provides a method for predicting the effect of glycyrrhizin on repairing lung injury based on machine learning, which comprises: Figure 1
[0028] S10: Obtain the molecular structure data of glycyrrhizin, the biomarker data related to lung injury, and the historical experimental repair effect data;
[0029] In the embodiments of the present application, first, the molecular structure related information of glycyrrhizin is obtained from the chemical database; the biomarker data of animal lung injury experiment is obtained from the biological experiment database. Then, the historical experimental result data of glycyrrhizin repairing lung injury is collected. After data acquisition, preprocessing is performed, including normalization, missing value filling and outlier removal, to ensure data quality and consistency.
[0030] The step S10 in the method provided in the embodiments of the present application comprises:
[0031] The molecular structure data of glycyrrhizin is obtained from a chemical database, wherein the molecular structure data comprises molecular weight, lipid-water partition coefficient and hydrogen bond donor-acceptor number;
[0032] The biomarker data of the lung injury animal model is obtained from a biological experiment database, wherein the biomarker data comprises pro-inflammatory cytokine concentration and oxidative stress index;
[0033] The historical experimental repair effect data of glycyrrhizin in repairing lung injury in historical literature is collected, wherein the historical experimental repair effect data comprises repair rate and side effect incidence.
[0034] In the embodiments of the present application, first, the molecular structure data of glycyrrhizin is obtained from a chemical database, wherein the chemical database is usually PubChem, which can provide molecular structure information of a compound.
[0035] The molecular weight is the relative molecular mass of a molecule, which can reflect the size of the molecule; the lipid-water partition coefficient is the concentration ratio when the molecule is distributed and balanced in a lipid-soluble environment and a water-soluble environment, and too strong or too weak lipid solubility can affect the drug to enter the cell to play a role; the hydrogen bond donor-acceptor number is the number of atoms / groups in the molecule that can provide hydrogen bonds.
[0036] For example, the molecular weight of glycyrrhizin obtained from PubChem is 432.4 Da, the lipid-water partition coefficient is 2.5, and the hydrogen bond donor-acceptor number is 8.
[0037] Secondly, the biomarker data is obtained from a biological experiment database. The biomarker data of the lung injury animal includes pro-inflammatory cytokine concentration and oxidative stress index.
[0038] The biological experiment database is a database for storing biomarker experimental data, including GEO or ArrayExpress; the pro-inflammatory cytokine concentration is an important pro-inflammatory factor in acute lung injury, which can be represented by the concentration of tumor necrosis factor-α (TNF-α); the oxidative stress index is a reflection of the degree of oxidative damage to lung tissue, which can be represented by malondialdehyde (MDA).
[0039] For example, the TNF-α concentration of the mouse lung injury model obtained from the GEO database is 50 pg / mL, and the MDA concentration is 10 μM.
[0040] Finally, the historical experimental repair effect data of glycyrrhizin in repairing lung injury in historical literature is collected, wherein the historical experimental repair effect data includes repair rate and side effect incidence. After obtaining the historical experimental repair effect data, standardization processing is performed to ensure the unit consistency, which is convenient for subsequent calculation.
[0041] Exemplary, from the literature to collect 10 experiments average repair rate of 70% and the incidence of side effects 5%.
[0042] In the embodiments of the application, the data reliability and comparability are improved by standardizing the data acquisition method. By obtaining data from authoritative databases, errors can be reduced, and by using animal model data for simulation, the actual biological environment is approached, and finally by using historical experimental repair effect data, basic data for real effect prediction are provided.
[0043] On the basis of data acquisition, key indicators also need to be calculated. To solve the above problems, based on the molecular structure data and biomarker data, the molecular interaction force index and cell permeability index of isofraxidin are calculated, and the repair effect coefficient is obtained by multivariate regression analyzer analysis by combining the historical experimental repair effect data.
[0044] S20: Based on the molecular structure data and biomarker data, the molecular interaction force index and cell permeability index of isofraxidin are calculated, and the repair effect coefficient is obtained by multivariate regression analyzer analysis by combining the historical experimental repair effect data;
[0045] In the embodiments of the application, first, based on the molecular structure data, the molecular interaction force index is calculated by weighted summation. Second, the concentration of pro-inflammatory cytokines is calculated to obtain the cell permeability index. Finally, the molecular interaction force index, the cell permeability index and the historical experimental repair effect data are input into the multivariate regression analyzer. Redundant information is eliminated by multivariate regression analysis to improve the representativeness of the index.
[0046] The method provided in the embodiments of the application comprises the following steps:
[0047] Based on the molecular weight, the lipid-water partition coefficient and the hydrogen bond donor-acceptor number of isofraxidin, the molecular interaction force index is calculated by weighted summation;
[0048] Based on the concentration of pro-inflammatory cytokines and oxidative stress index in the biomarker data, the ratio of the two is calculated and logarithm is taken to obtain the cell permeability index;
[0049] The molecular interaction force index, the cell permeability index and the historical experimental repair effect data are input into the multivariate regression analyzer, and the repair effect coefficient is output.
[0050] In the embodiments of the present application, firstly, the influence degree of the molecular interaction force index on the molecular weight, the lipid-water partition coefficient and the hydrogen bond donor-acceptor number is respectively set according to the molecular weight, the lipid-water partition coefficient and the hydrogen bond donor-acceptor number of glycyrrhizin, and the corresponding weight values are respectively set by referring to historical data. Subsequently, the molecular interaction force index is calculated by weighting and fusing the molecular interaction force index using the weights, and the molecular interaction force index = w1 x molecular weight + w2 x lipid-water partition coefficient + w3 x hydrogen bond donor-acceptor number, wherein w1, w2 and w3 are weights.
[0051] For example, the weight w1 of the molecular weight is 0.3, the weight w2 of the lipid-water partition coefficient is 0.5, and the weight w3 of the hydrogen bond donor-acceptor number is 0.2, and the molecular interaction force index = (0.3 x 432.4 + 0.5 x 2.5 + 0.2 x 8) = 132.57.
[0052] Secondly, based on the biomarker data, the ratio of the pro-inflammatory cytokine concentration to the oxidative stress index is calculated, and the natural logarithm is taken to obtain the cell permeability index, and the cell permeability index = log (促炎细胞因子浓度 / 氧化应激指标) .
[0053] For example, the pro-inflammatory cytokine concentration is 50 pg / mL, the oxidative stress index is 10 μM, the ratio is 5, and the cell permeability index after taking the logarithm is 1.61.
[0054] Thirdly, the molecular interaction force index and the cell permeability index are input into a multiple regression analyzer together with historical experimental repair effect data to output a repair effect coefficient. The repair effect coefficient is a comprehensive score obtained by weighting calculation based on the repair rate and the incidence of side effects, and the range is 0-1. The repair effect prediction value is obtained by comprehensive evaluation and labeling of the repair rate and the incidence of side effects in the historical experimental repair effect data.
[0055] For example, after the molecular interaction force index and the cell permeability index are input into the multiple regression analyzer, the output repair effect coefficient is 0.75, indicating that the repair effect is moderately high.
[0056] After obtaining the repair effect coefficient, the confidence interval needs to be further configured and the index needs to be optimized to improve the prediction accuracy. The optimization process will be described in detail below.
[0057] The step S20 in the method provided in the embodiments of the present application, the construction step of the multiple regression analyzer includes:
[0058] A sample molecular interaction force index set, a sample cell permeability index set and a sample historical experimental repair effect data set are collected, and the repair rate and the incidence of side effects corresponding to each sample are labeled to obtain a sample repair effect coefficient set, wherein the sample repair effect coefficient is a comprehensive score obtained by weighting calculation based on the repair rate and the incidence of side effects;
[0059] Based on machine learning, a network architecture of a multiple regression analyzer is constructed.
[0060] The sample molecular interaction force indicator set, the sample cell permeability indicator set, and the sample historical experiment repair effect data set are used as input, and the sample repair effect coefficient set is used as supervision. The multiple regression analyzer is supervised trained until the verification accuracy converges, and the construction of the multiple regression analyzer is completed.
[0061] In the embodiments of the application, first, a sample molecular interaction force indicator set, a sample cell permeability indicator set, and a sample historical experiment repair effect data set are collected, and a labeling tool is used to label the repair rate and the side effect occurrence rate corresponding to each sample to obtain a sample repair effect coefficient set.
[0062] Data labeling refers to processing raw data, adding structured labels or annotations, so that it can be understood and used by machine learning models. Data labeling is the basis of supervised learning. The labeled data is usually used to train and verify the model, which can help the model learn to extract meaningful patterns and information from the original data.
[0063] For example, a BP neural network is used to construct a multiple regression analyzer. The BP neural network model is a feedforward neural network that is trained by error backpropagation and is commonly used for continuous value prediction.
[0064] Based on the BP neural network, the steps of constructing the multiple regression analyzer are as follows:
[0065] First, data preparation. The sample molecular interaction force indicator set, the sample cell permeability indicator set, and the sample historical experiment repair effect data set are used as input, and the sample repair effect coefficient set is used as supervision. The training set, the validation set, and the test set are divided according to the ratio of 7:2:1.
[0066] Second, model building. It mainly consists of an input layer, a hidden layer, and an output layer, which contains 5 input nodes, 5 hidden layer nodes, and 1 output layer node. The input layer weights and biases the sample molecular interaction force indicator set, the sample cell permeability indicator set, and the sample historical experiment repair effect data set. The hidden layer performs nonlinear transformation on the sample molecular interaction force indicator set, the sample cell permeability indicator set, and the sample historical experiment repair effect data set through the activation function ReLU with the sample repair effect coefficient set as the supervision label. The output layer again weights and biases the sum, and limits the predicted value to the range of 0-1 through the activation function Sigmoid, and outputs the predicted value.
[0067] Next, model training is performed. Using the sample repair effect coefficient set as the supervision label, an initial learning rate and weights are set, and weights are assigned. The mean squared error function is used to calculate the error between the predicted results and the sample molecular interaction force index set, sample cell permeability index set, and sample historical experimental repair effect data set. Weights are adjusted and calculated iteratively until the error is minimized. Parameters are updated through forward and backpropagation. Performance is evaluated using a validation set after each training epoch to avoid overfitting. When the MSE loss decreases by less than 1% over five consecutive training epochs, the model is considered successful. e-5 When the MSE loss on the validation set stabilizes below 0.01, the model is considered converged, and the multivariate regression analyzer is obtained. Finally, the model is tested by selecting a device, moving the model to the device, and performing 10 training cycles, with a test performed after each cycle.
[0068] In this embodiment, a molecular interaction force index is calculated by weighted summation based on molecular structure data. Then, based on biomarker data, the ratio of pro-inflammatory cytokine concentration to oxidative stress index is calculated, and the natural logarithm is taken to obtain a cell permeability index. Finally, the molecular interaction force index, cell permeability index, and historical experimental repair effect data are input into a multiple regression analyzer, which outputs a repair effect coefficient through a trained regression model. This transforms multidimensional data into quantifiable indicators, and eliminates redundant information through multiple regression analysis, improving the representativeness of the indicators.
[0069] S30: Based on the repair effect coefficient, configure the effect confidence interval, and optimize the range of the molecular interaction force index and cell permeability index to obtain optimized bioactivity index;
[0070] In this embodiment, the effect confidence interval is a statistical interval that represents the credible range of the repair effect coefficient. It is usually calculated based on the mean and standard deviation and is used to screen reliable data points.
[0071] Based on the repair effect coefficients output by the multivariate regression analyzer, corresponding effect confidence intervals are configured, the ranges of molecular interaction force indicators and cell permeability indicators are optimized, and the mean and standard deviation are calculated to obtain optimized bioactivity indicators, which are convenient for subsequent screening.
[0072] Step S30 in the method provided in this application embodiment includes:
[0073] Calculate the mean and standard deviation of the repair effect coefficient, and configure the effect confidence interval as [μ-2σ, μ+2σ], where μ is the mean and σ is the standard deviation;
[0074] The molecular interaction force index and cell permeability index of the current sample are compared with the effect confidence interval, and the confidence weight coefficient is calculated based on the relative position of each index value within the effect confidence interval.
[0075] Based on feature importance analysis, the basic weight coefficients of interaction force indicators and cell permeability indicators were determined;
[0076] The standardized molecular interaction force index and cell permeability index are weighted and fused together using the product of the corresponding basic weight coefficient and confidence weight coefficient as the final weight to obtain the optimized bioactivity index.
[0077] In this embodiment, based on the repair effect coefficient output by the multiple regression analyzer, a corresponding effect confidence interval is configured. The effect confidence interval is configured as [μ-2σ, μ+2σ], where μ is the mean and σ is the standard deviation. The mean represents the average value; the standard deviation represents the degree of data dispersion.
[0078] For example, if the mean repair effect coefficient is 132.5 and the standard deviation is 0.3, then the confidence interval for the molecular interaction force index is [131.9, 133.1], the mean repair effect coefficient is 1.6 and the standard deviation is 0.2, and the confidence interval for the cell permeability index is [1.2, 2.0].
[0079] Secondly, the molecular interaction force and cell permeability indices of the current sample are compared with the confidence intervals, and confidence weight coefficients are calculated based on the relative positions of each index value within the effect confidence interval. The index values are adjusted through this comparison to eliminate outlier data.
[0080] Among them, the confidence weight coefficient represents the weight calculated based on the position of the index value within the confidence interval, reflecting the reliability of the data.
[0081] For example, the molecular interaction force index is 132.57, the cell permeability index is 1.61, and the confidence weight coefficients are 0.83 and 1, respectively.
[0082] Finally, the basic weight coefficients were determined through feature importance analysis. The standardized indicators were then weighted and fused according to the product of the basic weight and the confidence weight to obtain the optimized bioactivity index. Subsequently, the optimized bioactivity index can be calculated as follows: Optimized bioactivity index = Standardized molecular interaction force index × Basic weight × Confidence weight + Standardized cell permeability index × Basic weight × Confidence weight.
[0083] Feature importance analysis is a machine learning method used to determine the degree of influence of different indicators on prediction.
[0084] For example, the standardized molecular interaction force index and cell permeability index are 0.8 and 0.6, respectively. After feature importance analysis, the basic weight coefficients are 0.6 and 0.4, and the confidence weight coefficients are 0.83 and 1, respectively. The optimized bioactivity index of the molecular interaction force index = 0.8 × 0.6 × 0.83 + 0.6 × 0.4 × 1 ≈ 0.64.
[0085] Step S30 of the method provided in this application embodiment compares the molecular interaction force index and cell permeability index of the current sample with the effect confidence interval, and calculates the confidence weight coefficient based on the relative position of each index value within the effect confidence interval, including:
[0086] Determine whether the current indicator value is within the effect confidence interval. If the current indicator value is within the effect confidence interval, calculate the relative position coefficient based on the distance between the current indicator value and the center of the interval. If the current indicator value exceeds the effect confidence interval, use the least squares method to adjust the corresponding indicator to the boundary value of the effect confidence interval.
[0087] Using a preset weight mapping rule, the relative position coefficient is mapped to a confidence weight coefficient between 0 and 1. The weight mapping rule satisfies the following condition: when the current index value is equal to the center of the interval, the confidence weight coefficient is set to 1, and it monotonically decreases to approach 0 as the index value approaches the boundary of the interval.
[0088] In this embodiment, firstly, it is determined whether the index value is within the confidence interval. If it is within the interval, the relative position coefficient is calculated; if it exceeds the interval, the least squares method is used to adjust it to the boundary value; if the index value is greater than the upper limit, it is set as the upper limit value.
[0089] The relative position coefficient represents the position of the index value within the confidence interval, and the calculation formula is 1 - |index value - interval center| / (interval width / 2); the weight mapping rule is a function that maps the relative position coefficient to the confidence weight coefficient, which is usually a linearly decreasing function; the least squares method is an optimization method that can adjust the index that is outside the interval to the boundary value.
[0090] For example, the confidence interval is [131.9, 133.1], with the center at 132.5. If the molecular interaction force index is 0.8, the relative position coefficient is 0.5. If the index is 1.0, it is outside the interval, and the relative position weight coefficient is 0.
[0091] The confidence interval for the molecular interaction force index is [131.9, 133.1], the molecular interaction force index is 132.57, and the relative position weighting coefficient is 1 - |132.57 - 132.5| / (1.2 / 2) ≈ 0.83. The confidence interval for the cell permeability index is [1.2, 2.0], the cell permeability index is 1.61, and the relative position weighting coefficient is 1 - |1.6 - 1.6| / (0.8 / 2) = 1.0.
[0092] Secondly, the relative position coefficient is converted into a confidence weight coefficient using a weight mapping rule. This relative position coefficient is mapped to a confidence weight coefficient between 0 and 1. Specifically, if the current indicator value equals the center of the interval, the confidence weight coefficient is set to 1, and it monotonically decreases towards 0 as the relative position coefficient approaches the interval boundary. The closer the indicator value is to the center of the interval, the higher the confidence weight coefficient, indicating more reliable data.
[0093] For example, the confidence interval for the molecular interaction force index is [131.9, 133.1], and the relative position weighting coefficient is 0.83. The confidence interval for the cell permeability index is [1.2, 2.0], and the relative position weighting coefficient is 1.0. The resulting confidence weighting coefficients are 0.83 and 1, respectively.
[0094] In this embodiment, confidence intervals are configured based on the repair effect coefficient. The molecular interaction force and cell permeability indices of the current sample are compared with the confidence intervals to calculate confidence weight coefficients. Finally, the base weight coefficients are determined, and weighted fusion is performed to obtain optimized bioactivity indices. Outliers are filtered out using confidence intervals, and important features are emphasized through weight optimization to improve index stability. Dynamic weight adjustments ensure that the optimized indices are more focused on high-confidence regions. This concentrates the optimized bioactivity indices in high-probability effective regions, improving the accuracy of subsequent predictions.
[0095] After optimizing the indicators, it is necessary to construct an effect prediction model. To address the above issues, this application constructs an effect prediction model based on a neural network. The optimized bioactivity indicators of the target isoglycyrrhizin are input into the effect prediction model, and the predicted value of the lung injury repair effect is output.
[0096] S40: Construct an effect prediction model based on a neural network, input the optimized bioactivity index of the target isoglycyrrhizin into the effect prediction model, and output the predicted value of lung injury repair effect.
[0097] In this embodiment, an efficacy prediction model is obtained by constructing and supervising the training of a neural network model. The optimized bioactivity index of the target isoglycyrrhizin is input, and the efficacy prediction model is used to capture nonlinear relationships to obtain the predicted value of the lung injury repair effect of the target isoglycyrrhizin. Predicting the lung injury repair effect is beneficial for drug evaluation.
[0098] Step S40 in the method provided in this application embodiment includes:
[0099] A set of optimized bioactivity indicators was collected from samples, and the lung injury repair effect corresponding to each sample was labeled to obtain a set of predicted values for sample repair effect.
[0100] Based on neural networks, construct the network architecture for an effect prediction model;
[0101] The optimized bioactivity index set of the samples is used as input, and the predicted value set of the sample repair effect is used as supervision to supervise the training of the effect prediction model until the mean square error on the validation set converges, thus completing the construction of the effect prediction model.
[0102] For example, the steps to construct an effect prediction model based on a BP neural network are as follows:
[0103] First, data preparation. The optimized set of bioactivity indicators of the samples was used as input, and the predicted value set of sample repair effect was used as supervision label. The samples were divided into training set, validation set, and test set in a 7:2:1 ratio.
[0104] Secondly, the model is constructed. It mainly consists of an input layer, a hidden layer, and an output layer, containing 5 input nodes, 5 hidden layer nodes, and 1 output layer node. The input layer performs a weighted summation of the weights and biases of the sample's optimized bioactivity index set. The hidden layer uses the ReLU activation function and the predicted value set of sample repair effect as the supervision label to perform a nonlinear transformation on the sample's optimized bioactivity index set. The output layer again performs a weighted summation of the weights and biases and uses the Sigmoid activation function to restrict the predicted value to the range of 0-1 before outputting the predicted value.
[0105] Next, model training. The set of predicted repair effects from samples is used as the supervision label. An initial learning rate and weights are set, and weights are assigned. The mean squared error function is used to calculate the error between the predicted results and the optimized bioactivity index set of samples. Weight adjustments are made and the calculation is repeated iteratively until the error is minimized. Parameters are updated through forward and backpropagation. Performance is evaluated using a validation set after each training epoch to avoid overfitting. When the MSE loss decreases by less than 1% over five consecutive training epochs, the model is considered successful. e-5 When the MSE loss on the validation set stabilizes below 0.01, the model is considered converged, and the predicted performance model is obtained. Finally, the model is tested by selecting a device, moving the model to the device, and performing 10 training cycles, with a test performed after each cycle.
[0106] In step S40 of the method provided in this application embodiment, the sample repair effect prediction value set is obtained based on the repair rate and side effect incidence rate annotation in the historical experimental repair effect data. The sample repair effect prediction value set includes a repair effect score that comprehensively considers the degree of improvement in repair rate and the degree of control over side effect incidence rate.
[0107] In this embodiment of the application, the repair rate and the incidence of side effects are extracted from historical experimental repair effect data, and a repair effect score is obtained by weighted calculation. The repair effect score, also known as the repair effect coefficient, is a comprehensive score obtained by weighted calculation of the repair rate and the incidence of side effects.
[0108] Repair effectiveness score = Repair rate × Repair rate weight - Side effect incidence rate × Side effect incidence rate weight. The weights can be configured and set according to the actual scenario. The repair effectiveness score obtained by weighted fusion is normalized to the range of 0-1 and labeled, which can be used as a supervision label for model training.
[0109] The predicted repair effect value is obtained by comprehensively evaluating and labeling the repair rate and side effect incidence rate in the historical experimental repair effect data; the historical experimental repair effect data is the original data of repair rate and side effect incidence rate.
[0110] For example, with a repair rate weight of 0.8 and a side effect incidence rate weight of 0.2, and a repair rate of 70% and a side effect incidence rate of 5% in historical data, the repair effect score = 0.7 × 0.8 - 0.05 × 0.2 = 0.56 - 0.01 = 0.55, which is 0.55 after normalization.
[0111] Step S40 of the method provided in this application embodiment, which involves constructing an effect prediction model based on a neural network, inputting the optimized bioactivity index of the target isoglycyrrhizin into the effect prediction model, and outputting the predicted value of lung injury repair effect, further includes:
[0112] If the predicted value of the lung injury repair effect is higher than the preset first threshold, the target isoglycyrrhizin is determined to have a significant repair effect, and a recommendation to advance to preclinical trials is output.
[0113] If the predicted value of the lung injury repair effect is between the preset first threshold and the second threshold, then the target isoglycyrrhizin is determined to have a moderate repair effect, and a suggestion to optimize the molecular structure is output.
[0114] If the predicted value of the lung injury repair effect is lower than the preset second threshold, it is determined that the repair effect of the target isoglycyrrhizin is insufficient, and a suggestion to adjust the dosing regimen is output.
[0115] In this embodiment, firstly, based on the historical repair effect scores of lung injury repair effect values, a first threshold and a second threshold are set as preset critical values for classifying and predicting results.
[0116] Secondly, the predicted lung injury repair effects output by the effect prediction model are distinguished based on the significance of the first and second thresholds. If the predicted lung injury repair effect value is higher than the first threshold, it is considered a high score for lung injury repair effect. In this case, the target isoglycyrrhizin is determined to have a significant repair effect, and a recommendation to proceed to preclinical trials is generated. Preclinical trials are drug trials in the drug development stage, including animal experiments and toxicity tests.
[0117] If the predicted lung injury repair effect falls between the first and second thresholds, it is considered a moderate score for lung injury repair effect. Therefore, the target isoglycyrrhizin is determined to have a moderate repair effect, and a suggestion for molecular structure optimization is output. Further optimization using chemical modification to improve drug properties is recommended.
[0118] If the predicted lung injury repair effect is lower than the second threshold, it is considered a low score for lung injury repair effect. Therefore, the target isoglycyrrhizin is deemed to have insufficient repair effect, and a recommendation to adjust the dosing regimen is generated. It is recommended to adjust the regimen and repeat the experiment.
[0119] For example, a first threshold of 0.7 and a second threshold of 0.4 are set. If the predicted value 0.8 > 0.7, the output recommends proceeding to preclinical trials; if the predicted value is 0.5, 0.4 < 0.5 < 0.7, the output suggests molecular structure optimization; if the predicted value is 0.3 < 0.4, the output suggests adjusting the dosing regimen.
[0120] In this embodiment, a set of predicted lung injury repair effects is obtained through data collection and annotation. Supervised training is then performed using a backpropagation (BP) neural network until the mean squared error on the validation set converges, thus constructing an effect prediction model. Furthermore, a preset threshold is used to differentiate the predicted lung injury repair effects based on significance, providing a tiered evaluation output to guide drug development decisions and reduce blind experimentation. This ensures that drugs with high predictive values are prioritized, saving resources.
[0121] The embodiments of this application, through the specific implementation methods described above, achieve the following technical effects:
[0122] In this embodiment, the reliability and comparability of the data are improved by standardizing data sources. Obtaining data from authoritative databases reduces errors, while animal model data is used for simulation to closely resemble the actual biological environment. Finally, historical experimental data on the repair effects are used to provide a reference for the real-world results.
[0123] Secondly, based on molecular structure data, a weighted summation of molecular interaction forces is used to calculate the molecular interaction force index. Then, based on biomarker data, the ratio of pro-inflammatory cytokine concentration to oxidative stress index is calculated, and the natural logarithm is taken to obtain the cell permeability index. Finally, the molecular interaction force index, cell permeability index, and historical experimental repair effect data are input into a multiple regression analyzer. This analyzer outputs the repair effect coefficient through a trained regression model. This process transforms multidimensional data into quantifiable indicators and eliminates redundant information through multiple regression analysis, improving the representativeness of the indicators.
[0124] Furthermore, based on the confidence interval configured for the repair effect coefficient, the molecular interaction force and cell permeability indices of the current sample are compared with the confidence interval to calculate the confidence weight coefficients. Finally, the base weight coefficients are determined, and weighted fusion is performed to obtain the optimized bioactivity indices. Outliers are filtered out using confidence intervals, and important features are emphasized through weight optimization to improve index stability. Dynamic weight adjustments ensure that the optimized indices are more focused on high-confidence regions. This concentrates the optimized bioactivity indices in high-probability effective regions, improving the accuracy of subsequent predictions.
[0125] Finally, through data collection and annotation, a set of predicted values for lung injury repair effects was obtained. Based on a backpropagation (BP) neural network, supervised training was performed until the mean squared error on the validation set converged, thus constructing an effect prediction model. Furthermore, by using preset thresholds, the predicted values for lung injury repair effects were differentiated by significance, providing stratified evaluation outputs to guide drug development decisions and reduce blind experiments. This ensures that drugs with high predictive values are prioritized, saving resources.
[0126] Example 2, as Figure 2 As shown, this application provides a machine learning-based system for predicting the effect of isoliquiritigenin on lung injury repair, including:
[0127] The data acquisition module 11 is used to acquire molecular structure data of isoliquiritigenin, biomarker data related to lung injury, and historical experimental repair effect data.
[0128] The repair effect coefficient analysis module 12 is used to calculate the molecular interaction force index and cell permeability index of isoliquiritigenin based on the molecular structure data and biomarker data, and to obtain the repair effect coefficient by combining the historical experimental repair effect data through a multivariate regression analyzer.
[0129] The optimized bioactivity index calculation module 13 is used to configure the effect confidence interval according to the repair effect coefficient, and to optimize the range of the molecular interaction force index and cell permeability index to obtain the optimized bioactivity index.
[0130] The effect prediction model construction module 14 is used to construct an effect prediction model based on a neural network. The optimized bioactivity index of the target isoglycyrrhizin is input into the effect prediction model, and the predicted value of lung injury repair effect is output.
[0131] In one embodiment, the data acquisition module 11 is configured to:
[0132] Molecular structure data of isoliquiritigenin were obtained from a chemical database, wherein the molecular structure data included molecular weight, lipid-water partition coefficient, and number of hydrogen bond donors and acceptors.
[0133] Biomarker data of lung injury animal models were obtained from a biological experimental database, wherein the biomarker data included the concentration of pro-inflammatory cytokines and oxidative stress indicators;
[0134] Historical experimental data on the repair effects of isoliquiritigenin in lung injury were collected from historical literature, including repair rate and incidence of side effects.
[0135] In one embodiment, the repair effect coefficient analysis module 12 is used for:
[0136] Based on the molecular weight, lipid-water partition coefficient, and number of hydrogen bond donors and acceptors of isoliquiritigenin, the molecular interaction force index was obtained by weighted summation calculation.
[0137] Based on the concentrations of pro-inflammatory cytokines and oxidative stress indicators in the biomarker data, the ratio between the two is calculated and the logarithm is taken to obtain the cell permeability index.
[0138] The molecular interaction force index, cell permeability index, and historical experimental repair effect data are input into a multiple regression analyzer, which outputs the repair effect coefficient.
[0139] The construction steps of the multivariate regression analyzer include:
[0140] A set of sample molecular interaction force indicators, a set of sample cell permeability indicators, and a set of historical experimental repair effect data are collected. The repair rate and side effect incidence rate of each sample are labeled to obtain a set of sample repair effect coefficients. The sample repair effect coefficients are comprehensive scores calculated based on the repair rate and the side effect incidence rate.
[0141] A network architecture for a multivariate regression analyzer is constructed based on machine learning.
[0142] Using the sample molecular interaction force index set, sample cell permeability index set, and sample historical experimental repair effect dataset as inputs, and the sample repair effect coefficient set as supervision, the multivariate regression analyzer is trained under supervision until the verification accuracy converges, thus completing the construction of the multivariate regression analyzer.
[0143] In one embodiment, the optimized bioactivity index calculation module 13 is used for:
[0144] Calculate the mean and standard deviation of the repair effect coefficient, and configure the effect confidence interval as [μ-2σ, μ+2σ], where μ is the mean and σ is the standard deviation;
[0145] The molecular interaction force index and cell permeability index of the current sample are compared with the effect confidence interval, and the confidence weight coefficient is calculated based on the relative position of each index value within the effect confidence interval.
[0146] Based on feature importance analysis, the basic weight coefficients of interaction force indicators and cell permeability indicators were determined;
[0147] The standardized molecular interaction force index and cell permeability index are weighted and fused together using the product of the corresponding basic weight coefficient and confidence weight coefficient as the final weight to obtain the optimized bioactivity index.
[0148] The step of comparing the molecular interaction force index and cell permeability index of the current sample with the effect confidence interval, and calculating the confidence weight coefficient based on the relative position of each index value within the effect confidence interval, includes:
[0149] Determine whether the current indicator value is within the effect confidence interval. If the current indicator value is within the effect confidence interval, calculate the relative position coefficient based on the distance between the current indicator value and the center of the interval. If the current indicator value exceeds the effect confidence interval, use the least squares method to adjust the corresponding indicator to the boundary value of the effect confidence interval.
[0150] Using a preset weight mapping rule, the relative position coefficient is mapped to a confidence weight coefficient between 0 and 1. The weight mapping rule satisfies the following condition: when the current index value is equal to the center of the interval, the confidence weight coefficient is set to 1, and it monotonically decreases to approach 0 as the index value approaches the boundary of the interval.
[0151] In one embodiment, the effect prediction model building module 14 is used for:
[0152] A set of optimized bioactivity indicators was collected from samples, and the lung injury repair effect corresponding to each sample was labeled to obtain a set of predicted values for sample repair effect.
[0153] Based on neural networks, construct the network architecture for an effect prediction model;
[0154] The optimized bioactivity index set of the samples is used as input, and the predicted value set of the sample repair effect is used as supervision to supervise the training of the effect prediction model until the mean square error on the validation set converges, thus completing the construction of the effect prediction model.
[0155] The process of using the optimized bioactivity index set of the samples as input and the predicted value set of sample repair effects as supervision to supervise the training of the effect prediction model until the mean squared error on the validation set converges, thereby completing the construction of the effect prediction model, includes:
[0156] The sample repair effect prediction value set is obtained based on the repair rate and side effect incidence rate in historical experimental repair effect data. The sample repair effect prediction value set includes a repair effect score that comprehensively considers the degree of improvement in repair rate and the degree of control over side effect incidence.
[0157] In one embodiment, the effect prediction model building module 14 is further configured to:
[0158] If the predicted value of the lung injury repair effect is higher than the preset first threshold, the target isoglycyrrhizin is determined to have a significant repair effect, and a recommendation to advance to preclinical trials is output.
[0159] If the predicted value of the lung injury repair effect is between the preset first threshold and the second threshold, then the target isoglycyrrhizin is determined to have a moderate repair effect, and a suggestion to optimize the molecular structure is output.
[0160] If the predicted value of the lung injury repair effect is lower than the preset second threshold, it is determined that the repair effect of the target isoglycyrrhizin is insufficient, and a suggestion to adjust the dosing regimen is output.
[0161] The embodiments of this application, through the specific implementation methods described above, achieve the following technical effects:
[0162] In this embodiment, the data acquisition module 11 first standardizes data sources, improving data reliability and comparability. Obtaining data from authoritative databases reduces errors, while animal model data is used for simulation to closely approximate actual biological environments. Finally, historical experimental data on repair effects provides a reference for real-world results.
[0163] Secondly, the molecular interaction force index is calculated by weighted summation based on molecular structure data using the repair effect coefficient analysis module 12. Then, based on biomarker data, the ratio of pro-inflammatory cytokine concentration to oxidative stress index is calculated, and the natural logarithm is taken to obtain the cell permeability index. Finally, the molecular interaction force index, cell permeability index, and historical experimental repair effect data are input into a multiple regression analyzer, which outputs the repair effect coefficient through a trained regression model. This process transforms multidimensional data into quantifiable indicators and eliminates redundant information through multiple regression analysis, improving the representativeness of the indicators.
[0164] Furthermore, by optimizing the bioactivity index calculation module 13, confidence intervals are configured based on the repair effect coefficients. The molecular interaction force index and cell permeability index of the current sample are compared with the confidence intervals to calculate confidence weight coefficients. Finally, the base weight coefficients are determined, and weighted fusion is performed to obtain the optimized bioactivity index. Outliers are filtered out using confidence intervals, and important features are emphasized through weight optimization to improve index stability. Dynamic weight adjustments ensure that the optimized index is more focused on high-confidence regions. This concentrates the optimized bioactivity index in high-probability effective regions, improving the accuracy of subsequent predictions.
[0165] Finally, through the effect prediction model construction module 14, data is collected and labeled to obtain a set of predicted values for sample repair effects. Based on a BP neural network, supervised training is performed until the mean squared error on the validation set converges, thus constructing the effect prediction model. Furthermore, by using preset thresholds, the predicted values for lung injury repair effects are distinguished by significance, providing stratified evaluation outputs to guide drug development decisions and reduce blind experiments. This ensures that drugs with high predictive values are prioritized, saving resources.
[0166] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0167] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A machine learning-based method for predicting the effect of isoliquiritigenin on lung injury repair, characterized in that, The method includes: To obtain molecular structure data of isoliquiritigenin, biomarker data related to lung injury, and historical experimental data on repair effects; Based on the molecular structure data and biomarker data, the molecular interaction force index and cell permeability index of isoliquiritigenin were calculated. Combined with the historical experimental repair effect data, the repair effect coefficient was obtained by analyzing the data through a multivariate regression analyzer. Based on the repair effect coefficient, an effect confidence interval is configured, and the range of the molecular interaction force index and cell permeability index is optimized to obtain optimized bioactivity index; A neural network-based effect prediction model is constructed. The optimized bioactivity index of the target isoglycyrrhizin is input into the effect prediction model, and the predicted value of lung injury repair effect is output.
2. The method for predicting the effect of isoliquiritigenin on lung injury repair based on machine learning according to claim 1, characterized in that, Obtain molecular structure data of isoliquiritigenin, lung injury-related biomarker data, and historical experimental repair efficacy data, including: Molecular structure data of isoliquiritigenin were obtained from a chemical database, wherein the molecular structure data included molecular weight, lipid-water partition coefficient, and number of hydrogen bond donors and acceptors. Biomarker data of lung injury animal models were obtained from a biological experimental database, wherein the biomarker data included the concentration of pro-inflammatory cytokines and oxidative stress indicators; Historical experimental data on the repair effects of isoliquiritigenin in lung injury were collected from historical literature, including repair rate and incidence of side effects.
3. The method for predicting the effect of isoliquiritigenin on lung injury repair based on machine learning according to claim 1, characterized in that, Based on the aforementioned molecular structure and biomarker data, the molecular interaction force and cell permeability indices of isoliquiritigenin were calculated. Combined with the historical experimental repair effect data, the repair effect coefficient was analyzed and obtained, including: Based on the molecular weight, lipid-water partition coefficient, and number of hydrogen bond donors and acceptors of isoliquiritigenin, the molecular interaction force index was obtained by weighted summation calculation. Based on the concentrations of pro-inflammatory cytokines and oxidative stress indicators in the biomarker data, the ratio between the two is calculated and the logarithm is taken to obtain the cell permeability index. The molecular interaction force index, cell permeability index, and historical experimental repair effect data are input into a multiple regression analyzer, which outputs the repair effect coefficient.
4. The method for predicting the effect of isoliquiritigenin on lung injury repair based on machine learning according to claim 3, characterized in that, The steps for constructing the multivariate regression analyzer include: A set of sample molecular interaction force indicators, a set of sample cell permeability indicators, and a set of historical experimental repair effect data are collected. The repair rate and side effect incidence rate of each sample are labeled to obtain a set of sample repair effect coefficients. The sample repair effect coefficients are comprehensive scores calculated based on the repair rate and the side effect incidence rate. A network architecture for a multivariate regression analyzer is constructed based on machine learning. Using the sample molecular interaction force index set, sample cell permeability index set, and sample historical experimental repair effect dataset as inputs, and the sample repair effect coefficient set as supervision, the multivariate regression analyzer is trained under supervision until the verification accuracy converges, thus completing the construction of the multivariate regression analyzer.
5. The method for predicting the effect of isoliquiritigenin on lung injury repair based on machine learning according to claim 1, characterized in that, Based on the repair effect coefficient, an effect confidence interval is configured, and the ranges of the molecular interaction force index and cell permeability index are optimized to obtain optimized bioactivity indicators, including: Calculate the mean and standard deviation of the repair effect coefficient, and configure the effect confidence interval as [μ-2σ, μ+2σ], where μ is the mean and σ is the standard deviation; The molecular interaction force index and cell permeability index of the current sample are compared with the effect confidence interval, and the confidence weight coefficient is calculated based on the relative position of each index value within the effect confidence interval. Based on feature importance analysis, the basic weight coefficients of interaction force indicators and cell permeability indicators were determined; The standardized molecular interaction force index and cell permeability index are weighted and fused together using the product of the corresponding basic weight coefficient and confidence weight coefficient as the final weight to obtain the optimized bioactivity index.
6. The method for predicting the effect of isoliquiritigenin on lung injury repair based on machine learning according to claim 5, characterized in that, The molecular interaction force index and cell permeability index of the current sample are compared with the effect confidence interval, and the confidence weight coefficients are calculated based on the relative positions of each index value within the effect confidence interval, including: Determine whether the current indicator value is within the effect confidence interval. If the current indicator value is within the effect confidence interval, calculate the relative position coefficient based on the distance between the current indicator value and the center of the interval. If the current indicator value exceeds the effect confidence interval, use the least squares method to adjust the corresponding indicator to the boundary value of the effect confidence interval. Using a preset weight mapping rule, the relative position coefficient is mapped to a confidence weight coefficient between 0 and 1. The weight mapping rule satisfies the following condition: when the current index value is equal to the center of the interval, the confidence weight coefficient is set to 1, and it monotonically decreases to approach 0 as the index value approaches the boundary of the interval.
7. The method for predicting the effect of isoliquiritigenin on lung injury repair based on machine learning according to claim 1, characterized in that, The steps of a neural network-based performance prediction model include: A set of optimized bioactivity indicators was collected from samples, and the lung injury repair effect corresponding to each sample was labeled to obtain a set of predicted values for sample repair effect. Based on neural networks, construct the network architecture for an effect prediction model; The optimized bioactivity index set of the samples is used as input, and the predicted value set of the sample repair effect is used as supervision to supervise the training of the effect prediction model until the mean square error on the validation set converges, thus completing the construction of the effect prediction model.
8. The method for predicting the effect of isoliquiritigenin on lung injury repair based on machine learning according to claim 7, characterized in that, The sample repair effect prediction value set is obtained based on the repair rate and side effect incidence rate in historical experimental repair effect data. The sample repair effect prediction value set includes a repair effect score that comprehensively considers the degree of improvement in repair rate and the degree of control over side effect incidence.
9. The method for predicting the effect of isoliquiritigenin on lung injury repair based on machine learning according to claim 1, comprising constructing an effect prediction model based on a neural network, inputting the optimized bioactivity index of the target isoliquiritigenin into the effect prediction model, and outputting the predicted value of lung injury repair effect, further comprising: If the predicted value of the lung injury repair effect is higher than the preset first threshold, the target isoglycyrrhizin is determined to have a significant repair effect, and a recommendation to advance to preclinical trials is output. If the predicted value of the lung injury repair effect is between the preset first threshold and the second threshold, then the target isoglycyrrhizin is determined to have a moderate repair effect, and a suggestion to optimize the molecular structure is output. If the predicted value of the lung injury repair effect is lower than the preset second threshold, it is determined that the repair effect of the target isoglycyrrhizin is insufficient, and a suggestion to adjust the dosing regimen is output.
10. A machine learning-based predictive system for the effect of isoliquiritigenin on lung injury repair, characterized in that, The system for implementing the method according to any one of claims 1-9 comprises: The data acquisition module is used to acquire molecular structure data of isoliquiritigenin, biomarker data related to lung injury, and historical experimental repair effect data. The repair effect coefficient analysis module is used to calculate the molecular interaction force index and cell permeability index of isoliquiritigenin based on the molecular structure data and biomarker data, and to obtain the repair effect coefficient by combining the historical experimental repair effect data with a multivariate regression analyzer. The optimized bioactivity index calculation module is used to configure the effect confidence interval based on the repair effect coefficient, and to optimize the range of the molecular interaction force index and cell permeability index to obtain the optimized bioactivity index. The effect prediction model construction module is used to construct an effect prediction model based on a neural network. The optimized bioactivity index of the target isoglycyrrhizin is input into the effect prediction model, and the predicted value of lung injury repair effect is output.