A Safety Prediction Method for Dendrobium officinale Extract Based on Multi-Indicator Joint Modeling

By employing a multi-indicator joint modeling and multi-dimensional prediction verification strategy, the uncertainty caused by batch differences and nonlinear relationships in the safety prediction of Dendrobium officinale extract was resolved, achieving accurate, reliable, and efficient prediction of safe concentrations, while reducing experimental costs and time.

CN122489976APending Publication Date: 2026-07-31DOCTOR PLANT GUANGDONG BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-09
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for predicting the safety of Dendrobium officinale extract suffer from drawbacks such as batch-to-batch variations and nonlinear dose-response relationships, making it difficult to determine the specific working concentration. This results in high experimental costs, long cycles, and unstable results, and the reliance on empirical rules leads to uncertain safety boundaries.

Method used

A multi-index joint modeling method is adopted. By constructing a comprehensive activity quantification value sequence, monotonicity, slope constraint, noise robustness constraint and boundary consistency constraint are applied to train the correlation model. The weighted loss function is used to optimize the fitting process, and loss convergence judgment and fitting adaptive optimization mechanism are introduced. Combined with multi-dimensional prediction and verification strategy, the interpretability and consistency of safe concentration are ensured.

Benefits of technology

It significantly reduces redundant verification and blind trial and error, improves the accuracy and batch consistency of safe concentration prediction, narrows the safe concentration window, reduces experimental costs, and ensures the reliability and reproducibility of safe prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489976A_ABST
    Figure CN122489976A_ABST
Patent Text Reader

Abstract

This invention discloses a safety prediction method for Dendrobium officinale extract based on multi-index joint modeling, belonging to the field of multi-index joint modeling. The method includes the following steps: In the safety prediction scenario of Dendrobium officinale extract, a training dataset is constructed based on sample batches. Each sample batch contains multiple concentration gradients and corresponds to a complete set of experimental data. The safe working concentration and experimental process information of each sample batch are recorded. The experimental data of each sample batch are converted into a corresponding comprehensive activity quantification value sequence. Monotonicity constraints, slope constraints, noise robustness constraints, and boundary consistency constraints are applied to fit the comprehensive activity quantification curve of each sample batch. This method ensures safety constraints while narrowing the safe concentration window, improving cross-batch judgment consistency and reproducibility, significantly reducing blind trial and error and repeated verification, and ensuring the accuracy of experimental data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-indicator joint modeling technology, and in particular to a method for predicting the safety of Dendrobium officinale extract based on multi-indicator joint modeling. Background Technology

[0002] Conducting safety predictions during the research, formulation, and product release of Dendrobium officinale extract aims to identify potential safety risks in advance and establish actionable dosage boundaries, thereby avoiding insufficient safety judgments based solely on natural origin or compliance with a single indicator. Dendrobium officinale extract is a multi-component mixture system, and its component and impurity profiles fluctuate batch-to-batch due to factors such as raw material origin, harvesting season, processing and storage, and extraction and purification processes. Such fluctuations may lead to the enrichment or transformation of certain accompanying components, degradation products, or process residues under specific conditions, resulting in changes in cytotoxicity, irritation, or sensitivity risks. Therefore, safety predictions are necessary for early quantitative assessment and tiered management of risk trends.

[0003] In existing technologies, the safety prediction of Dendrobium officinale extract typically employs an empirical process based on standardized indicators and graded verification: first, routine quality and risk factor screening of raw materials and extracts is conducted, such as for heavy metals, pesticide residues, microorganisms, solvent residues, and consistency of major components, to exclude obviously substandard batches; then, basic safety concentration assessments are performed in in vitro models, commonly including cytotoxicity (e.g., metabolic activity or membrane integrity methods), irritation or corrosivity (e.g., epidermal reconstruction or replacement tests), initial screening indicators for sensitization risk, and, if necessary, phototoxicity or genotoxicity, to determine an acceptable concentration range based on the principle of not triggering significant adverse reactions under preset exposure conditions; in the application stage, more reliance is placed on existing industry experience and conservative margins, such as calculating the exposure level based on the maximum level without significant adverse reactions and setting a safety margin to provide recommended addition amounts or maximum allowable dosages in formulations.

[0004] The existing technology has the following technical problems: This type of method is characterized by its reliance on controlled experiments and threshold determination, and its use of conservative assumptions to ensure manageable risk. However, it often relies on empirical rules and repeated validation to handle batch variations, nonlinear dose-response relationships, and the determination of specific working concentrations between adjacent discrete concentration points. Typically, an intermediate concentration is manually selected between adjacent concentration points for subsequent in vitro tissue experiments. However, because the dose-response relationship may exhibit inflection points or nonlinear changes between adjacent concentration points, and because the criteria for cell viability and morphological damage are subject to fluctuations and uncertainties, the selection of the intermediate concentration is prone to two unfavorable situations: one is that the intermediate concentration... The first risk factor has already passed the risk inflection point, and potential safety or stimulation risks have emerged. Secondly, due to caution, the values ​​are set too low, making it difficult to detect efficacy signals in tissue models. This forces experiments to repeatedly increase or decrease concentrations and conduct repeated verifications within the range before barely achieving convergence. This not only increases the cost of cells, tissue models, and reagents, but also lengthens the research and evaluation cycle, causing delays in formulation iteration, process change verification, and batch release decisions. At the same time, with the time window lengthened, factors such as changes in the stability of raw materials and samples and differences in storage conditions will further introduce new deviations, forming a chain reaction where the more working concentrations are tested, the less convergent the results become. Summary of the Invention

[0005] To address the technical problems existing in the prior art, this invention provides a safety prediction method for Dendrobium officinale extract based on multi-indicator joint modeling. The technical solution is as follows: A method for predicting the safety of Dendrobium officinale extract based on multi-indicator joint modeling is provided. This method includes: Step 1: In the Dendrobium officinale extract safety prediction scenario, a training dataset is constructed based on sample batches. Each sample batch contains multiple concentration gradients and corresponds to a complete set of experimental data, recording the safe working concentration and experimental process information for each sample batch; Step 2: The experimental data of each sample batch is converted into a corresponding comprehensive activity quantification value sequence, and monotonicity constraints, slope constraints, noise robustness constraints, and boundary consistency constraints are applied to fit the comprehensive activity quantification curve for each sample batch; Step 3: Based on the comprehensive activity quantification curves of each sample batch, a correlation model jointly established by multiple indicators is trained. A weighted loss function based on curve correlation is used for joint training, and the convergence trend of the loss value sequence during training is used to determine whether to optimize the comprehensive activity quantification curve fitting process; Step 4: After the correlation model is trained, the comprehensive activity quantification curve is fitted to the target sample and input into the correlation model to obtain the predicted safe activity quantification value sequence. A multi-dimensional prediction verification strategy is used to determine whether to output the safe concentration of the extract for the target sample.

[0006] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. This invention addresses the technical problems in safety prediction of Dendrobium officinale extract, such as the wide safety boundary caused by discrete concentration gradients, the difficulty in uniformly characterizing batch differences, the reliance on experience in selecting intermediate working concentrations, and the high cost of retesting. It solves these problems by weighting and aggregating cell activity values ​​and morphological quantitative indicators to form a comprehensive activity quantification value sequence. Furthermore, it introduces monotonicity, slope range, robust noise, and boundary consistency constraints in the curve fitting process, ensuring that the dose-response relationship remains interpretable, comparable, and convergent even under small sample and fluctuating conditions. Finally, it uses multiple indicators such as local curve slope, confidence interval width, and residual scale to jointly train the correlation model. Based on the curve correlation degree and a loss function weighted by the curve uncertainty, the model learns the stable landing point pattern of the safe working concentration under the constraints of curve shape and uncertainty. When the training fails to converge, it drives the adaptive optimization of the fitting parameters, reducing the bias caused by outliers and excessive smoothing from the source. During the inference stage, it outputs the predicted safety activity quantification value sequence and combines the coverage probability and consistency check to realize the adaptive output of single value, interval or supplementary test suggestions. This ensures that the safety constraint is maintained while narrowing the safe concentration window, improving the consistency and reproducibility of cross-batch judgment, significantly reducing blind trial and repeated verification, and ensuring the accuracy of experimental data.

[0007] 2. This invention addresses the technical problems in the safety prediction of Dendrobium officinale extract, such as the susceptibility of curve fitting to outliers and batch fluctuations, leading to unstable safety boundaries and difficulties in convergence of the correlation model. It achieves closed-loop correction by introducing a loss convergence determination and adaptive optimization mechanism during the training phase. Specifically, the convergence trend is determined based on the weighted loss value sequence during the correlation model training process. When the loss does not show a convergence trend, curve fitting optimization is triggered. In the fitting optimization, the fitting residual sequence for each sample batch is calculated and the residual discrete values ​​are extracted. If the residual discrete value exceeds a first threshold, a significant outlier is identified, and the degrees of freedom of the heavy-tailed error for that batch are reduced to enhance robustness and weaken the impact of abnormal observations. If the residual discrete value is below a second threshold, the curve is deemed overly smoothed and may lose effective features, and the degrees of freedom are increased to improve sensitivity to changes in the actual response. Through the above adaptive adjustment, the comprehensive activity quantification curve remains comparable and convergent across different batches, thereby improving the consistency and generalization ability of the correlation model in fitting the safety activity quantification value, reducing repeated testing, and stabilizing the output of safe concentration conclusions.

[0008] 3. This method addresses the problems in the safety prediction of Dendrobium officinale extract, such as unstable safety point positioning, excessively wide boundaries, and reliance on empirical point selection due to concentration gradient dispersion and large curve uncertainty. It proposes a multi-dimensional prediction and verification strategy to achieve convergent safety concentration output. First, fluctuation verification is performed. Within a preset slope range, multiple candidate curves are generated from the comprehensive activity quantification value sequence and input into a correlation model to obtain the predicted safety activity quantification value sequence. The maximum and minimum differences are calculated as the predicted fluctuation value and compared with a threshold. If the fluctuation is too large, supplementary experimental feedback is triggered to reduce uncertainty. When the fluctuation is controllable, deviation verification is performed. The target safety activity quantification value is screened from the predicted sequence, and the deviation ratio between it and the minimum difference between the predicted and candidate curves is calculated and compared with a threshold. If the deviation ratio meets the condition, the safety concentration is calculated and output. If the deviation ratio does not meet the condition, constrained refitting is initiated. The target reference curve is retrieved by similarity to the training curve, and the degree of freedom interval is mapped based on the similarity and the degree of freedom of the reference curve. Constraints are applied to the target curve for refitting to redetermine the safety point. This forms a closed loop of fluctuation screening, deviation calibration, and constraint refitting, transforming the safety point location from empirical trial-and-error to a verifiable and correctable calculation process, reducing repeated retesting and improving cross-batch consistency. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 Flowchart of a method for predicting the safety of Dendrobium officinale extract based on multi-index joint modeling provided in this application embodiment; Figure 2 These are comparative images of cell morphology under different concentrations of Dendrobium officinale extract provided in the embodiments of this application. Figure 3 The image shows the results of Collagen I immunofluorescence staining of ex vivo skin tissue treated with 2.0% Dendrobium officinale extract in the embodiments of this application. Figure 4 The results of Collagen III immunofluorescence replicates of isolated skin tissue treated with 2.0% Dendrobium officinale extract in the embodiments of this application are shown. Detailed Implementation

[0011] Embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of the present disclosure are shown in the drawings, it should be understood that embodiments of the present disclosure may be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure.

[0012] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0013] Example 1, as Figure 1 The diagram shows a flowchart of a method for predicting the safety of Dendrobium officinale extract based on multi-indicator joint modeling, provided in an embodiment of this application. The method includes the following steps: Sample batches are used as the basic units of the training dataset. A sample batch can be understood as a Dendrobium officinale extract sample with the same production batch number or the same submitted batch, corresponding to a complete cell safety screening record. Taking batch B2024-01 as an example, a batch-level file is established for this batch, recording at least the batch number, sample source and preparation information, and simultaneously recording the final safe working concentration used for this batch and experimental process information. The experimental process information is used to describe the key contextual conditions of this experiment, such as the experimental plate number being Plate-07, the experimental date being February 1, 2024, the cell passage number being P12, and the plate reading time being 18:30, which is used for subsequent tracing and correction of systematic biases caused by different batches and different experimental days.

[0014] At the concentration level, the system sets multiple preset concentration points for each batch and collects experimental data. Taking this embodiment as an example, the concentration gradient adopts a serial dilution sequence, with volume fractions set to 0.0781%, 0.1563%, 0.3125%, 0.625%, 1.25%, 2.5%, 5%, and 10%, and three replicate wells are set for each concentration.

[0015] Experimental data includes cell viability values, quantitative indicators of cell morphology, and information from the control group experiments. The control group experimental information includes at least the readings from the zeroing well, solvent control well, and positive control well, used to standardize readings across different well locations to the same scale. For example, the zeroing well reading is 0.05, the average solvent control well reading is 0.80, and the average positive control well reading is 0.10; the readings of the three replicate wells at a 1.25% concentration are 0.73, 0.71, and 0.72, respectively. The system normalizes the data by using the zeroing well background and setting the solvent control to 100%, resulting in a cell viability value of approximately (0.72−0.05) / (0.80−0.05)≈89% at that concentration. Similarly, if the average reading at a 2.5% concentration is 0.66, the cell viability value is approximately (0.66−0.05) / (0.80−0.05)≈81%, thus forming a concentration-relative activity sequence for this batch, sorted by concentration, for subsequent curve fitting and safety point localization.

[0016] The system transforms morphological observations from descriptive conclusions into calculable indicators. This embodiment employs three types of indicators: adhesion rate, rounding ratio, and confluence. The adhesion rate characterizes the proportion of adherent cells, the rounding ratio characterizes the proportion of rounded or stressed cells, and the confluence characterizes the proportion of cell-covered area. For example, at a 1.25% concentration, three fields of view were statistically analyzed: Field 1: 120 total cells, 110 adherent, 8 rounded, confluence 78%; Field 2: 115 total cells, 104 adherent, 10 rounded, confluence 76%; Field 3: 118 total cells, 108 adherent, 9 rounded, confluence 77%. The overall adhesion rate is approximately 90%, the rounding ratio is approximately 8%, and the confluence is approximately 77%.

[0017] The system converts the above indicators into morphological activity coefficients according to preset mapping rules. This implementation uses a weighted aggregation mapping rule for the conversion. The morphological activity coefficient ranges from 0 to 1, with a larger value indicating a more normal morphology. It is then weighted and aggregated with the relative activity according to preset weights to obtain the comprehensive activity quantification value for that concentration point. For example, with an activity weight of 0.8 and a morphological weight of 0.2, the comprehensive activity quantification value can be expressed as 0.8 × cell activity value + 0.2 × morphological activity coefficient. Thus, each concentration point simultaneously possesses activity value, morphological quantification result, and normalized control information, forming a trainable point-level sample.

[0018] It should be explained that the adhesion rate and confluence are positively correlated with the morphological activity coefficient. That is, the higher the adhesion rate and the higher the confluence, the more stable the cell adhesion and the more complete the coverage, and the closer the morphology is to normal, thus increasing the morphological activity coefficient. The rounding ratio is negatively correlated with the morphological activity coefficient. That is, the higher the rounding ratio, the more obvious the cell stress or damage morphology, thus decreasing the morphological activity coefficient. Therefore, in this embodiment, the weights of the adhesion rate, rounding ratio, and confluence are set to 0.4, 0.3, and 0.3. The morphological activity coefficient is calculated as 0.4 × 90% + 0.3 × 8% + 0.3 × (1 / 77%) = 0.77, where the weights represent the degree of influence of different parameters on the morphological activity coefficient.

[0019] Figure 2 The accompanying images show the cell morphology comparison of different concentrations of Dendrobium officinale extract in an in vitro cell model. The images are labeled as 5%, 2.5%, 1.25%, 0.625%, and 0.3125% treatment groups, and a solvent control group (SolventControl) indicating the control condition where only solvent was added and no extract was added. By comparing the cell adhesion, cell density, and morphological changes such as rounding / shrinkage between the concentration groups and the solvent control group, the effects of different concentrations of extract on cell morphological stability can be assessed. This provides visual evidence for extracting morphological quantitative indicators such as adhesion rate, rounding / shrinkage ratio, and confluence.

[0020] At the batch level, the system associates and stores the final selected safe operating concentration for that batch with the aforementioned point-level data. Taking this embodiment as an example, if the batch ultimately adopts 2.0% as the safe operating concentration in subsequent verifications, then 2.0% is written as a batch-level label into the batch file, and its corresponding experimental process information and point-level evidence chain are retained. This allows the model to learn not only the concentration-response curve shape during training, but also the relationship between the safe operating concentration and the curve shape, fluctuation level, and experimental conditions, thereby supporting cross-batch safety prediction and traceable output.

[0021] Taking this embodiment as an example, the concentration points are 0.625%, 1.25%, 2.5%, 5% and 10%, and the corresponding comprehensive activity quantification values ​​are 92.0%, 88.5%, 80.8%, 67.4% and 52.0%, respectively.

[0022] During the fitting process, the system uses a monotonic curve to fit the above data and applies a monotonicity constraint. The monotonicity constraint is used to constrain the overall activity quantification curve obtained from the fitting process from showing an overall rebound as the concentration increases. That is, when the concentration increases from 1.25% to 2.5%, the predicted value of the curve should not rise from approximately 88.5 to 90 or higher. Taking the data in this embodiment as an example, the system requires the fitted curve to remain unchanged within the range of 0.625% to 1.25% to 2.5%, so that the curve output meets the overall trend of 92.0% ≥ 88.5% ≥ 80.8%, thereby avoiding misleading subsequent safety boundary positioning due to false rebounds in the high concentration range caused by a small number of noise points.

[0023] The comprehensive activity quantification curve shows the concentration point on the horizontal axis (in percentage) and the comprehensive activity quantification value on the vertical axis (in percentage), demonstrating the relationship between the comprehensive activity quantification value and the concentration.

[0024] A slope constraint is applied to control the steepness of the curve. The slope constraint is used to control the steepness of the comprehensive activity quantification curve within a preset slope range; in this embodiment, the slope range is 0.8 to 4. This means that the curve is allowed to decrease with increasing concentration, but abrupt drops are not permitted. For example, without a slope constraint, the curve might still predict a comprehensive activity quantification value of approximately 80.8% at the 2.5% concentration point, but suddenly drop to 50.0% near the 3.0% concentration point. This is inconsistent with the measured 67.4% at the 5% concentration point, indicating an overfitting drop induced by a small number of points. With the slope constraint, the system constrains the curve to have a smooth transition in the 2.5% to 5% range, keeping the predicted comprehensive activity quantification value around 75% near the 3.0% concentration point and gradually decreasing to around 67%, thereby stabilizing the inflection point and improving the curve's interpretability.

[0025] In this embodiment, to reduce the interference of abnormal wells and batch fluctuations on the shape of the comprehensive activity quantification curve, the system introduces both noise robustness constraint and fitting weight constraint when fitting the comprehensive activity quantification curve for each sample batch. The comprehensive activity quantification value is normalized to the 0-1 range. Taking batch B2024-01 as an example, at a concentration of 2.5%, the comprehensive activity quantification values ​​for the three wells are 0.808, 0.812, and 0.804, with a calculated standard deviation of approximately 0.004. At a concentration of 10%, the comprehensive activity quantification values ​​for the three wells are 0.520, 0.600, and 0.480, with a calculated standard deviation of approximately 0.060, indicating that the fluctuation at the 10% point is greater and there is an abnormally high reading of 0.600. The system first uses the standard deviation of each concentration point as the volatility. It then retrieves the fitting weights from a pre-set volatility-fitting weight mapping table in the database. The mapping rule adopts a form that is inversely proportional to the square of the standard deviation and is cropped to avoid extreme weights. For example, the weight values ​​are limited to between 0.2 and 5. Based on this, 2.5% of the points can be given a larger fitting weight, while 10% of the points can be given a smaller fitting weight. This allows the residual term of the concentration point to be scaled by the weighting coefficient in the objective function, so that even if outliers occur at highly volatile points, they will only have a limited impact. Meanwhile, the system uniformly sets the degrees of freedom of the t-distribution error model for the entire curve of this batch to 5. The degrees of freedom are determined by a preset degree of freedom mapping rule and fixed for residual modeling of all concentration points in this batch. This means that the residual penalty is set to a heavy-tailed form, so that observations with large residuals will not be excessively amplified like normal errors under the same fitting weights, thereby further weakening the pull of outliers on the curve parameter estimation. The degree of freedom mapping rule is to mean the fitting weights of all concentration points on the curve, and then look up the corresponding degrees of freedom by the mean processing result and the fitting weight average-degrees of freedom mapping table stored in the database.

[0026] In summary, the fitting weights are responsible for quantifying and allocating the residual contribution at the concentration point level, while the t-distribution error with 5 degrees of freedom is responsible for robustly handling large residuals at the error distribution level. Together, they work to fit the comprehensive activity quantification curve, so that while maintaining the overall dose response trend of not rebounding with increasing concentration, the curve shape is reduced by the distortion of high-fluctuation concentration points of 10%, and the stability and reproducibility of subsequent safety activity quantification value prediction and safety concentration inverse calculation process are improved.

[0027] Finally, boundary consistency constraints are applied to ensure that both ends of the curve remain consistent with the control results. Boundary consistency constraints require that the low-concentration end of the curve not exceed the control baseline level, and the high-concentration end of the curve not exceed the low activity level corresponding to the positive control. The control baseline level refers to the baseline value of the quantified overall cell activity measured under solvent control conditions, representing the reference level at which cells are in a normal state without the addition of the extract. This reference value serves as the upper limit constraint for the low-concentration end of the curve to avoid extrapolation exceeding the control. The low activity level refers to the lower limit reference value of the quantified overall cell activity measured under positive control conditions, representing the reference level at which cells exhibit significantly low activity under known damage or strong inhibitory treatment. This reference value serves as the upper limit constraint for the high-concentration end of the curve to avoid unreasonable rebounds in extrapolation. For example, in this embodiment, the solvent control corresponds to a comprehensive activity quantification value of 100%. The system anchors the comprehensive activity quantification value of the curve at 100% at the low concentration end, preventing the curve from drifting to 105% near the 0% concentration point due to extrapolation. The positive control corresponds to a comprehensive activity quantification value of 15%. The system constrains the comprehensive activity quantification value of the curve at the high concentration end to not exceed 15%, thereby avoiding the unreasonable phenomenon of the curve rising back to 30% when extrapolating after the comprehensive activity quantification value of 10%. Through the combined effects of monotonicity constraints, slope constraints, noise robustness constraints, and boundary consistency constraints, this embodiment obtains the comprehensive activity quantification curve for this batch of samples.

[0028] This embodiment details the training process of the correlation model for predicting the safety of Dendrobium officinale extract, as well as the fitting and optimization process of the comprehensive activity quantification curve during training. The correlation model used is a machine learning model based on the joint construction of multiple indicators. It takes the comprehensive activity quantification curve and curve descriptor of each sample batch as input features and the safety activity quantification value corresponding to each sample batch as output label. The core is to learn the correlation between the shape of the comprehensive activity quantification curve and the safety activity quantification value of Dendrobium officinale extract, so as to provide model support for the prediction of the safety activity quantification value of subsequent target samples.

[0029] The model training process is as follows: The safe working concentration of each sample batch on the corresponding comprehensive activity quantification curve is marked as the safe activity quantification value of each sample batch, i.e., the true label of the model training, Q. s,i =Q ’ (c i ), where Q s,i Let Q be the quantification value of the safety activity of the i-th sample batch. ’ (·) represents the comprehensive activity quantification curve function for the i-th sample batch, c i Let be the safe working concentration of the i-th sample batch, where i represents the batch number, i = 1, 2, 3, ..., m, and m represents the total number of sample batches.

[0030] The input to the correlation model is the comprehensive activity quantification curve and curve descriptor of each sample batch. The curve descriptor is a set of parameters that characterize the shape of the comprehensive activity quantification curve, including the comprehensive activity quantification value of each concentration point on the curve, the slope between adjacent concentration points, the curvature and other key parameters. The input features of each sample batch are uniformly integrated into a feature vector in a specific order.

[0031] The association model outputs the safety activity quantification value for each batch of samples after labeling.

[0032] The training objective is to minimize the weighted loss value. A weighted loss function based on curve correlation is used for training. Curve correlation characterizes the consistency between the model's predicted safety activity quantification value and the actual safety activity quantification value on the curve descriptor. The weighted loss function is constructed based on the correlation between the curve descriptors corresponding to the predicted safety activity quantification value and the actual safety activity quantification value. The correlation is represented by the cosine similarity of the curve descriptors, which is used to quantify the correlation between the curve descriptors corresponding to the predicted safety activity quantification value and the actual safety activity quantification value. The higher the cosine similarity value, the higher the fit between the curve descriptors and the stronger the correlation.

[0033] In this embodiment, the weighted loss function is as follows: Where L(θ) is the weighted loss value of the model; θ is the set of parameters to be trained for the association model, such as the weights and biases of the neural network, the coefficients of the regression model, etc.; a is a local constant to avoid the denominator being 0; This is the quantified value of the predicted safety activity of the i-th sample batch based on the parameter θ. p represents the true safety activity quantification value for the i-th sample batch. i This represents the correlation between the predicted safety activity quantification value and the actual safety activity quantification value of the i-th sample batch and the corresponding curve descriptor.

[0034] In this method, the weighted loss function used for training the association model is based on the cosine similarity of curve descriptors as the quantification of association. Weight coefficients are constructed based on this association to weight the loss for different sample batches. The principle is to calculate the cosine similarity between the curve descriptors corresponding to the predicted safety activity quantification value and the actual safety activity quantification value. This similarity value characterizes the degree of association between the two in terms of curve morphology; the higher the similarity, the greater the association, and the higher the corresponding loss weight, and vice versa. The loss function is based on the mean squared error as the fundamental loss calculation form, and the weight coefficients of each sample batch are compared with the predicted... The average is calculated by multiplying the squared error terms of the true and false values. This allows the model training process to focus more on sample batches with high matching and strong correlation between the curve shape and the safety activity quantification value, reducing the interference of samples with abnormal curve shapes and low correlation on model training. At the same time, it ensures that the loss value can accurately reflect the model's learning effect on the core rules, guaranteeing the training accuracy of the model when learning the correlation between the comprehensive activity quantification curve and the safety activity quantification value of Dendrobium officinale extract. This makes the trained model more closely match the characteristics of actual experimental data, improving the prediction accuracy of the safety activity quantification value of the target sample and the cross-batch generalization ability.

[0035] To determine whether optimization of the fitting process for the comprehensive activity quantification curve is needed, the system records the weighted loss value for each iteration during the training of the correlation model, forming a loss value sequence. The weighted loss value sequence refers to the set of loss values ​​indexed by the iteration round; for example, the loss values ​​for rounds 1 to 60 are 0.420, 0.395, 0.362, and up to 0.210, respectively. The convergent decreasing trend refers to the situation where, within several consecutive windows, the loss value not only decreases overall but also reaches a preset threshold, while the fluctuation range remains controllable. To quantify this trend, this embodiment uses a sliding window method, treating 10 consecutive rounds as a window. The mean decreasing rate between adjacent windows is calculated, which is equal to the mean loss of the previous window minus the mean loss of the next window, then divided by the mean loss of the previous window. The volatility of the loss value within the next window is also calculated as a stability index, where volatility is the standard deviation of the loss value within that window. When the mean decline rate is not less than 3% and the volatility is not greater than 0.01, and this condition is met continuously for 3 windows, the loss value sequence is determined to show a convergent downward trend, and the system continues to train the correlation model. Conversely, if the mean decline rate is less than 3% or the mean increases, or the volatility is greater than 0.01 and lasts for more than 2 windows, it is determined that there is no convergent downward trend, triggering the fitting process optimization of the comprehensive activity quantification curve and retraining the correlation model after optimization. For example, if the mean of the first window is 0.300, the mean of the second window is 0.285, and the mean of the third window is 0.271, the mean decline rate is approximately 5.0% and 4.9% respectively, and the corresponding volatility is 0.006 and 0.005 respectively, which can be determined as convergent decline. If the mean of the first window is 0.280, the mean of the second window is 0.279, and the mean of the third window is 0.281, the mean decline rate is insufficient and rebounds, which can be determined as non-convergence and trigger optimization.

[0036] In this embodiment, the fitting process of the comprehensive activity quantification curve adopts adaptive adjustment of degrees of freedom driven by residual dispersion. The fitted residual sequence refers to the sequence obtained by subtracting the comprehensive activity quantification value from the predicted value of the fitted curve at each preset concentration point within each sample batch. For example, the residual equals the comprehensive activity quantification value at the concentration point minus the predicted value of the curve.

[0037] The discrete value of the fitted residual refers to the quantification result of the dispersion of the residual sequence. In this embodiment, the standard deviation of the fitted residual sequence is preferably used as the discrete value. The discrete value of the residual is calculated for each sample batch and compared with a preset first threshold and a preset second threshold. The first threshold is used to determine whether the dispersion of the fitted residual of a sample batch is abnormally large. If it exceeds the threshold, it is determined that the batch has obvious outliers or heavy-tailed fluctuations, thereby triggering a reduction in the degrees of freedom of the t-distribution error to enhance robustness. The second threshold is used to determine whether the dispersion of the fitted residual of a sample batch is abnormally small. If it is below the threshold, it is determined that the batch may have over-smoothed and lost effective features, thereby triggering an increase in the degrees of freedom of the t-distribution error to improve the sensitivity to effective changes. Both the first threshold and the second threshold are stored in the database.

[0038] When the residual discrete value of a sample batch exceeds the first threshold, it is determined that the batch has obvious outliers or heavy-tailed fluctuations. The observation error of the batch is then modeled using a Student's t-distribution with reduced degrees of freedom to enhance heavy-tailed robustness. When the residual discrete value of a sample batch is below the second threshold, it is determined that the batch's fit is overly suppressing noise, potentially resulting in over-smoothing and loss of effective features. The degrees of freedom of the Student's t-distribution error of the batch are increased to improve sensitivity to effective changes. For ease of implementation, this embodiment limits the degrees of freedom to the range of 3 to 10 and adopts a piecewise mapping rule: when the residual discrete value is greater than 2.0, the degrees of freedom are set to 3; when the residual discrete value is between 0.8 and 2.0, the degrees of freedom are set to 5; and when the residual discrete value is less than 0.8, the degrees of freedom are set to 8. The first threshold is 2.0, and the second threshold is 0.8.

[0039] After the association model is trained, the system receives a batch of unknown target samples from the experiment as the objects to be predicted. These unknown target samples refer to samples that have undergone multi-concentration gradient cell experiments consistent with the training data, generating point-level experimental data, but for which a final safe working concentration label has not yet been provided; these are samples for which the method needs to automatically output a safe concentration. The system first converts the cell activity values ​​and morphological quantification indicators of each concentration point of the unknown target sample into a comprehensive activity quantification value sequence, and then fits the comprehensive activity quantification curve of the unknown target sample under constraints of monotonicity, slope range, noise robustness, and boundary consistency.

[0040] To avoid instability in the safety point output due to slope values ​​or local fluctuations in single-fit results, this embodiment introduces a multi-dimensional prediction verification strategy, which includes a fluctuation verification strategy and a bias verification strategy. The fluctuation verification strategy involves generating several candidate comprehensive activity quantification curves within a preset slope range using different connection methods, without altering the original comprehensive activity quantification value sequence of the unknown target sample. This assesses the model's sensitivity to curve morphology perturbations. The preset slope range can be consistent with the training phase, for example, set to 0.8 to 4, and several candidate slope values ​​are selected according to a preset step size to generate multiple candidate curves. The system inputs each candidate comprehensive activity quantification curve into the correlation model to obtain the corresponding predicted safety activity quantification value, thereby forming the predicted safety activity quantification value sequence for the unknown target sample.

[0041] In the fluctuation verification strategy, the system performs difference processing on the maximum and minimum values ​​of the predicted safety activity quantification value sequence to obtain the predicted fluctuation value, which is used to quantify the sensitivity of the model output to the candidate curve set. For example, if the predicted safety activity quantification values ​​output by the correlation model under four candidate curves are 0.86, 0.84, 0.85, and 0.83 respectively, then the predicted fluctuation value is 0.86 minus 0.83 equals 0.03. The system compares this predicted fluctuation value with a preset threshold predicted fluctuation value, which can be preset to 0.05, to distinguish whether the output is stable, and stores it in the database. When the predicted fluctuation value is higher than 0.05, it is determined that the safety point prediction of the unknown target sample is too sensitive to curve morphology perturbations, indicating that the current experimental information is insufficient to support stable output. The system triggers experimental data supplementation feedback and outputs a supplementary test suggestion with the uncertainty of the convergence curve. When the predicted fluctuation value is not higher than 0.05, it is determined that the predicted output is stable. The system enters the deviation verification strategy to further verify the candidate safety points before deciding whether to output the safe concentration of the extract.

[0042] After completing the fluctuation verification and confirming that the predicted fluctuation value is not higher than the defined predicted fluctuation value, the system enters the deviation verification strategy to determine whether the predicted safety activity quantification value of the unknown target sample can stably land on the comprehensive activity quantification curve of the target sample. Here, the predicted safety activity quantification value sequence refers to a set of predicted outputs obtained by inputting the correlation model into multiple candidate comprehensive activity quantification curves in the fluctuation verification strategy; the target safety activity quantification value refers to a representative predicted value selected from this predicted safety activity quantification value sequence, used as the benchmark value for this round of safety point positioning. In this embodiment, the median of the sequence is preferably used as the target safety activity quantification value to reduce the impact of extreme predicted values ​​on subsequent positioning. For example, if the predicted safety activity quantification value sequence is 0.86, 0.84, 0.85, 0.83, then the median is 0.845, and the system marks 0.845 as the target safety activity quantification value.

[0043] The system calculates a prediction deviation value to measure the degree of fit between the target safety activity quantification value and the target sample's comprehensive activity quantification curve. The prediction deviation value refers to the absolute value of the difference between the target sample's comprehensive activity quantification curve and the closest curve value to the target safety activity quantification value. For ease of implementation, this embodiment searches on a preset candidate concentration grid with a grid step size of 0.1%. For example, if the target sample's comprehensive activity quantification curve values ​​at 1.8%, 1.9%, 2.0%, and 2.1% are 0.84, 0.85, 0.87, and 0.88 respectively, then the curve closest to the target safety activity quantification value of 0.845 is either 0.84 or 0.85, corresponding to a minimum absolute difference of 0.005. The system marks 0.005 as the prediction deviation value. To avoid the impact of different sample scales on threshold uniformity, the system further processes the prediction deviation value and the defined prediction deviation value to obtain the prediction deviation ratio. The defined prediction deviation value is a preset tolerance benchmark in the database, which is set to 0.01 in this embodiment. Therefore, the prediction deviation ratio is 0.005 divided by 0.01, equaling 0.5. The system compares the prediction deviation ratio with the defined prediction deviation ratio. In this embodiment, the defined prediction deviation ratio is set to 1.0. When the prediction deviation ratio is lower than 1.0, it is determined that the target safety activity quantification value can form a stable landing point on the curve.

[0044] When it is determined that the target safety activity quantification value can form a stable landing point on the curve, the safe concentration is obtained by inversely locating the target safety activity quantification value on the comprehensive activity quantification curve. Specifically, the concentration points are traversed on a preset candidate concentration grid, and the curve prediction value at each concentration point is calculated. The absolute value of the difference between the predicted value and the target safety activity quantification value is used as the matching error. The concentration point with the smallest matching error and that meets the safety constraint is selected as the safe concentration, where the candidate concentration grid step size is preferably 0.1%. The safety constraint means that the curve prediction value at the concentration point is not lower than the minimum value allowed by the preset comprehensive activity quantification, thereby ensuring that the selected concentration still meets the safety requirements under uncertainty conditions. For example, if the target safety activity quantification value is 0.845, and the predicted values ​​of the curve at 1.8%, 1.9%, 2.0%, and 2.1% are 0.84, 0.85, 0.87, and 0.88 respectively, then the matching errors at each point are 0.005, 0.005, 0.025, and 0.035 respectively. Under the condition of simultaneously satisfying safety constraints, the system prioritizes the concentration point with the smallest error, and selects the point with the higher concentration when the error is the same to improve the detection capability of subsequent verification. Therefore, 1.9% is finally mapped to a clean concentration of 2.0% and 2.0% is output as the safe concentration of the extract of the unknown target sample.

[0045] Mapping 1.9% to a clean concentration of 2.0% and outputting 2.0% as the safe concentration of the extract for the unknown target sample is because this method, during reverse localization, first identifies the candidate concentration point in the candidate concentration grid that has the smallest matching error with the target safety activity quantification value and meets the safety constraints. When multiple candidate points have the same or similar matching errors and all meet the safety constraints, to ensure the output results have engineering feasibility and cross-batch reproducibility, the system further executes the "clean concentration alignment" rule, that is, mapping the candidate concentration to a preset standard concentration scale to avoid label drift caused by experimental preparation errors, differences in measurement accuracy, or inconsistencies in concentration scales between different laboratories. At the same time, under the premise of meeting the safety constraints, the standard concentration closer to the upper limit is preferentially selected to improve the detectability probability of efficacy signals in subsequent ex vivo tissue or formulation validation and reduce the number of retests. Based on the above principles, when 1.9% and 2.0% are in the same standard scale neighborhood and both safety constraints are met, the system aligns 1.9% to 2.0% and outputs 2.0% as the final safe concentration.

[0046] When the prediction bias ratio is not lower than the defined prediction bias ratio, the system determines that the target safety activity quantification value does not fit the current comprehensive activity quantification curve sufficiently. The system then executes a constrained refitting strategy to reduce curve fitting uncertainty and re-select the target safety activity quantification value. Constrained refitting refers to not directly altering the original experimental data, but rather introducing a noise robustness parameter constraint from the reference curve to achieve a more reasonable balance between robustness and sensitivity in the curve fitting of the target sample. Specifically, the system compares the comprehensive activity quantification curve of the target sample with multiple comprehensive activity quantification curves in the training dataset. The similarity can be calculated from the differences in curve values ​​at several standard concentration points; a higher value indicates greater similarity. The system sorts the similarities from highest to lowest and selects the training batch curve with the highest similarity as the target reference curve. For example, if the similarity between the target sample and training curve T3 is 0.92, with T7 it is 0.88, and with T9 it is 0.84, then the system selects T3 as the target reference curve.

[0047] After determining the target reference curve, the system reads the t-distribution error degrees of freedom parameter used in fitting the reference curve. Assuming this degree of freedom is 5, the system obtains a degree of freedom interval based on the similarity and this degree of freedom, which is used to constrain the refitting process of the target sample. This embodiment provides an easy-to-implement mapping method: when the similarity is not lower than 0.90, the degree of freedom interval is set to [4, 6]; when the similarity is between 0.80 and 0.90, the degree of freedom interval is set to [3, 7]; when the similarity is lower than 0.80, the degree of freedom interval is set to [3, 10]. Accordingly, when the similarity is 0.92 and the reference degree of freedom is 5, the degree of freedom interval is [4, 6]. The system then uses this degree of freedom interval as a constraint to refit the target sample, that is, only allowing the t-distribution error degrees of freedom to take values ​​in the range of [4, 6] and optimize them, so that the curve maintains its robustness to outliers while avoiding excessive smoothing due to excessive robustness. After refitting, the system regenerates the comprehensive activity quantification curve of the target sample and recalculates the predicted safety activity quantification value sequence. Then, it re-screens the target safety activity quantification values ​​according to the above deviation verification process until the prediction deviation ratio is lower than the defined prediction deviation ratio and outputs the safe concentration of the extract, or triggers further experimental data supplementation feedback.

[0048] When the fluctuation verification strategy determines that the predicted fluctuation value is higher than the defined predicted fluctuation value, or the deviation verification strategy determines that the predicted deviation ratio does not meet the threshold requirement, the system enters the experimental data supplementation feedback process. This process is used to determine whether there is a convergent safe point coverage interval for the current unknown target sample, and accordingly decides whether to trigger supplementary testing or output an anomaly warning. The predicted safety activity quantification value sequence refers to a set of predicted output values ​​obtained by generating multiple candidate comprehensive activity quantification curves within a preset slope range and inputting them into the correlation model. The coverage interval refers to the concentration range on the comprehensive activity quantification curve of the target sample that can accommodate the above-mentioned predicted safety activity quantification values, used to characterize the possible landing point range of the predicted safe point on the concentration axis.

[0049] In terms of coverage interval generation, the system first obtains the comprehensive activity quantification curve and its confidence interval of the target sample, and then discretizes the curve using a preset concentration grid, with a preferred grid step size of 0.1%. For each predicted value in the predicted safety activity quantification value sequence, the system searches for a candidate concentration point on the concentration grid that makes the curve prediction value closest to that predicted value, and records the position of the candidate concentration point on the concentration axis as the projection position of the predicted value. When the same predicted value can achieve a similar matching error at multiple concentration points, the point that satisfies the safety constraint and has a higher concentration is preferred as the projection position. Thus, the system projects the predicted safety activity quantification value sequence into a set of concentration positions, and uses the minimum envelope interval of this set on the concentration axis as the coverage interval. For example, if the predicted safety activity quantification value sequence has five values: 0.83, 0.84, 0.85, 0.86, and 0.88, and the concentration positions after projection onto the curve are 1.8%, 1.9%, 2.0%, 2.0%, and 2.2%, respectively, then the coverage interval can be [1.8%, 2.2%].

[0050] Regarding coverage determination, the system presets a coverage determination probability threshold to quantify the concentration of predicted safety activity quantification values ​​within a candidate coverage interval. In this embodiment, the coverage determination probability threshold is set to 0.8. The system further slides within the coverage interval to generate several candidate sub-intervals, for example, generating candidate sub-intervals with lengths of 0.2%, 0.3%, 0.4%, etc., with a step size of 0.1%, and calculates the coverage probability for each candidate sub-interval. The coverage probability refers to the proportion of the number of predicted safety activity quantification values ​​falling into the candidate sub-interval to the total number of predicted safety activity quantification value sequences. For example, if the total number of predicted safety activity quantification value sequences is 5, and the candidate sub-interval [1.9%, 2.1%] contains three projection positions (1.9%, 2.0%, 2.0%), the coverage probability is 3 / 5 = 0.6; if the candidate sub-interval [1.8%, 2.2%] contains all five projection positions, the coverage probability is 5 / 5 = 1.0.

[0051] In selecting the minimum coverage interval, the system uses the coverage determination probability threshold as a constraint and the minimum interval width as the objective to select and determine the minimum coverage interval. The coverage determination probability threshold, representing the minimum allowable coverage probability, is stored in the database. Continuing the example above, if the threshold is 0.8, then at least 4 / 5 predicted values ​​need to be covered. The system finds that the candidate sub-interval [1.8%, 2.2%] has a coverage probability of 1.0 and meets the threshold, while the narrower interval [1.8%, 2.1%] only covers 4 predicted values, with a coverage probability of 0.8, which also meets the threshold, and its interval width is smaller. Therefore, [1.8%, 2.1%] is determined as the minimum coverage interval. After finding the minimum coverage interval, the system triggers the experimental data supplementation feedback process, outputting supplementation suggestions to converge the uncertainty of the interval. For example, it suggests prioritizing supplementation of the concentration point within the minimum coverage interval that maximizes the curve confidence interval width or minimizes the predicted value matching error, thereby achieving safe point convergence with the fewest supplementation measurements.

[0052] If the minimum coverage interval that meets the coverage determination probability threshold cannot be found after retrieval, for example, if the predicted safety activity quantification value projection position is scattered in [1.0%, 3.0%] and the coverage probability of any candidate sub-interval not exceeding 0.5% is less than 0.8, then the system determines that the predicted safety point of the unknown target sample cannot converge under the current experimental information and outputs an abnormal warning message. The abnormal warning message includes at least a coverage failure indicator and a suggestion to review the experimental conditions or supplement key concentration points to avoid giving misleading safety concentration conclusions under unstable prediction.

[0053] Figure 3 This image shows the results of Collagen I immunofluorescence staining of ex vivo skin tissue treated with 2.0% Dendrobium officinale extract according to the embodiments of this application. In the image, green fluorescence represents Collagen I positive signal, and blue fluorescence represents cell nuclear signal. The scale bar in the lower right corner is used to calibrate the imaging scale. From left to right, the images correspond to the representative fields of replicate 1, replicate 2, and replicate 3, which are used to demonstrate the consistency of results between different replicates under the same concentration and experimental conditions. The advantage of accurately determining the safe concentration is that it can maintain the stability of tissue structure and staining background without causing tissue stimulation or decreased cell activity, reduce false differences caused by tissue damage, shedding, or non-specific background increase due to excessively high concentration, and avoid weak signals due to excessively low concentration, making it difficult to detect. This makes the Collagen I fluorescence intensity and distribution of the three replicates more reproducible and comparable, and improves the reliability of immunofluorescence results when used to determine safety and efficacy.

[0054] Figure 4The image shows the results of the Collagen III immunofluorescence replicates of ex vivo skin tissue treated with 2.0% Dendrobium officinale extract in this embodiment of the application. In the image, green fluorescence represents a positive Collagen III signal (reflecting the distribution and relative abundance of type III collagen), and blue fluorescence represents a cell nuclear signal. The scale bar in the lower right corner is used to calibrate the imaging scale. From left to right, these represent the representative fields of replicate 1, replicate 2, and replicate 3, demonstrating the consistency of results between different replicates under the same concentration and experimental conditions. Accurately determining the safe concentration is beneficial because when the concentration is within the safe window, tissue activity and structural integrity are more easily maintained, and non-specific background, tissue necrosis and shedding, and stress-induced fluorescence interference are significantly reduced. This allows the Collagen III signal intensity and fibrous distribution to more accurately reflect the effect of the extract rather than the artifacts of damage caused by excessively high concentrations. Simultaneously, it avoids weak signals and large replicate differences due to excessively low concentrations, which makes judgment difficult. This improves the reproducibility, comparability, and statistical reliability of the three-replica results, supporting stable conclusions regarding safety and tissue response.

[0055] While maintaining the training dataset structure, point-level experimental data fields, and safe working concentration recording method described in Example 1, this Example 2 is proposed when observable overall drift occurs when the same batch of Dendrobium officinale extract samples is repeatedly tested under different experimental plate numbers or different reading times, leading to unstable fitting results of the comprehensive activity quantification curve or significant differences in safe point positioning. Here, overall drift refers to a systematic shift in the same direction across multiple concentration points, such as the overall activity quantification obtained from the same batch on Plate-07. The initial values ​​were 0.94, 0.92, 0.88, 0.67, and 0.52, while the values ​​for the same concentration on Plate-09 were 0.91, 0.89, 0.85, 0.64, and 0.49, a decrease of approximately 0.03. Furthermore, when the reading time for the same plate number was delayed from 18:30 to 19:10, the values ​​decreased from 0.92, 0.88, 0.81, 0.67, and 0.52 to 0.90, 0.86, 0.79, 0.65, and 0.50, a decrease of approximately 0.02. If this type of drift is not addressed, it will cause the position of the subsequent fitted curve to shift downwards, resulting in a deviation in the predicted safety activity quantification value's landing point on the curve and causing inconsistencies in the output safety concentration.

[0056] In Example 2, the system uses the experimental plate number, cell passage number, and plate reading time from the experimental process information as correction factors in the comprehensive activity quantification curve fitting. Specifically, the system first establishes plate number bias and time bias for each sample batch: using the low-concentration baseline segment corresponding to the solvent control as the anchor point, the mean values ​​of the baseline segments at different plate numbers or different times are aligned. Taking the above example, the average low-concentration baseline value of Plate-07 is 0.94, and the average low-concentration baseline value of Plate-09 is 0.91. The plate offset is set to 0.03. The system adds 0.03 to the overall comprehensive activity quantification value of each concentration point in this batch under Plate-09 to obtain 0.94, 0.92, 0.88, 0.67, and 0.52, aligning it with Plate-07. If the plate reading time is delayed, causing the baseline average value to drop from 0.92 to 0.90, the time offset is set to 0.02. The system adds 0.02 to all concentration points at that time to compensate, thereby eliminating the systematic downward shift caused by the difference in plate reading time. To address the state differences caused by cell passage number, this embodiment employs a segmented correction method: when a higher passage number leads to an overall decrease in the low concentration segment but a larger decrease in the high concentration segment, the system calculates the bias amount separately for the low concentration segment and the medium-to-high concentration segment, and applies different correction biases to different segments to avoid distortion of the curve shape caused by simple overall translation.

[0057] After completing the above corrections, the system uses the corrected comprehensive activity quantification value sequence to perform monotonic curve fitting consistent with Example 1, and continues to apply monotonicity constraints, slope constraints, noise robustness constraints, and boundary consistency constraints to obtain the corrected comprehensive activity quantification curve. Since the plate number bias and time bias have absorbed systematic errors before fitting, this embodiment can ensure that the curve shape and position obtained by the same batch under different plate numbers and different reading times are consistent. This makes the inverse positioning result of the predicted safety activity quantification value output by the correlation model on the curve more stable, reduces inconsistencies such as the safety concentration changing from 1.9% to 2.2% in multiple experiments of the same batch, and improves the reproducibility across experimental conditions and the reliability of subsequent verification.

[0058] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A safety prediction method for Dendrobium officinale extract based on multi-indicator joint modeling, characterized in that, Includes the following steps: Step 1: In the safety prediction scenario of Dendrobium officinale extract, construct a training dataset based on sample batches. Each sample batch contains multiple concentration gradients and corresponds to a complete set of experimental data. Record the safe working concentration and experimental process information for each sample batch. Step 2: Convert the experimental data of each sample batch into the corresponding comprehensive activity quantification value sequence, and apply monotonicity constraint, slope constraint, noise robustness constraint and boundary consistency constraint to fit the comprehensive activity quantification curve of each sample batch. Step 3: Based on the comprehensive activity quantification curves of each sample batch, train the correlation model jointly established by multiple indicators, use a weighted loss function based on the correlation degree of the curve to carry out joint training, and determine whether to optimize the fitting process of the comprehensive activity quantification curve based on the convergence trend of the loss value sequence during the training process. Step 4: After the association model is trained, fit the comprehensive activity quantification curve to the target sample and input it into the association model to obtain the predicted safety activity quantification value sequence. Use a multi-dimensional prediction verification strategy to determine whether to output the safe concentration of the extract of the target sample.

2. The method for predicting the safety of Dendrobium officinale extract based on multi-index joint modeling as described in claim 1, characterized in that: The specific process for constructing the training dataset in batches is as follows: The training dataset is organized in batches, with each batch corresponding to a complete set of experimental data for Dendrobium officinale extract, and each batch containing multiple concentration points of a preset concentration gradient. For each concentration point, corresponding experimental data were collected and recorded. The experimental data specifically included cell viability values, cell morphology quantitative indicators, and control group experimental information used to normalize cell viability values ​​at that concentration point. The cell morphology quantitative indicators included adhesion rate, rounding ratio, and confluence. At each sample batch level, the final determined safe working concentration for that sample batch, as well as the experimental process information for that sample batch, are recorded. The experimental process information includes the experimental plate number, experimental date, number of cell passages, and plate reading time.

3. The method for predicting the safety of Dendrobium officinale extract based on multi-index joint modeling as described in claim 1, characterized in that: The specific fitting process for the comprehensive activity quantification curves of each sample batch is as follows: Cell morphology quantitative indicators are converted into morphological activity coefficients according to a preset mapping rule; The cell activity value and morphological activity coefficient are weighted and aggregated according to preset weights to obtain a comprehensive activity quantification value; For each sample batch, a monotonic curve is used to fit the comprehensive activity quantification value at each corresponding concentration point. Monotonicity constraints, slope constraints, noise robustness constraints, and boundary consistency constraints are applied to the fitting process to obtain the comprehensive activity quantification curve for each sample batch. The monotonicity constraint means that the overall activity quantification curve obtained by the constraint fitting must not show an overall rebound as the concentration increases; The slope limitation refers to controlling the slope of the comprehensive activity quantification curve within a preset slope range; The noise robustness limit refers to using a t-distribution error and preset degrees of freedom to make the fitting insensitive to occasional abnormal holes, batch fluctuations, or outliers. The boundary consistency constraint refers to constraining the curve at the low concentration end to not be higher than the control baseline level, and constraining the curve at the high concentration end to not be higher than the low activity level corresponding to the positive control, so that the comprehensive activity quantification curve is consistent with the experimental control results at both ends.

4. The method for predicting the safety of Dendrobium officinale extract based on multi-index joint modeling as described in claim 1, characterized in that: The training process involves establishing a correlation model using multiple indicators. The safe working concentration of each sample batch on the corresponding comprehensive activity quantification curve is marked as the safe activity quantification value of each sample batch. The correlation model takes the comprehensive activity quantification curve of each sample batch and the corresponding curve descriptor as input, and outputs the corresponding safety activity quantification value. The curve descriptor refers to the set of parameters that describe each point on the comprehensive activity quantification curve; The correlation model is trained using a weighted loss function, which is constructed based on the correlation between the predicted safety activity quantification value and the actual safety activity quantification value on the comprehensive activity quantification curve. The correlation degree is used to characterize the consistency between the predicted safety activity quantification value and the actual safety activity quantification value on the curve descriptor.

5. The method for predicting the safety of Dendrobium officinale extract based on multi-index joint modeling as described in claim 1, characterized in that: The specific process for determining whether to optimize the fitting of the comprehensive activity quantification curve is as follows: Obtain the loss value sequence of the weighted loss function during the training process of the association model, and determine whether the loss value sequence shows a convergent decreasing trend; If the loss value sequence shows a converging decreasing trend, then continue training the correlation model; If the loss value sequence does not show a convergent decreasing trend, it is determined that the fitting process of the comprehensive activity quantification curve should be optimized, and the correlation model should be retrained based on the optimized fitting results.

6. The method for predicting the safety of Dendrobium officinale extract based on multi-index joint modeling as described in claim 5, characterized in that: The fitting process of the comprehensive activity quantification curve is optimized, and the specific optimization process is as follows: Obtain the fitting residual sequence for each sample batch and analyze the discrete values ​​of the fitting residual for each sample batch. If the discrete value of the fitting residual of a certain sample batch exceeds the preset first threshold, then the degrees of freedom in the t-distribution error of that sample batch are reduced. If the discrete value of the fitting residual of a certain sample batch is lower than the preset second threshold, then the degrees of freedom in the t-distribution error of that sample batch are increased.

7. The method for predicting the safety of Dendrobium officinale extract based on multi-index joint modeling as described in claim 1, characterized in that: The multi-dimensional prediction verification strategy specifically includes a volatility verification strategy and a deviation verification strategy. The fluctuation verification strategy refers to obtaining the comprehensive activity quantification value sequence of the target sample, connecting them according to a preset slope range to obtain several comprehensive activity quantification curves of the target sample, and inputting them into the correlation model to obtain the predicted safety activity quantification value sequence of the target sample. The maximum and minimum values ​​in the predicted safety activity quantification value sequence are processed by difference, and the result is marked as the predicted fluctuation value. The predicted volatility value is compared with the preset defined predicted volatility value; If the predicted volatility value is higher than the defined predicted volatility value, then supplementary experimental data feedback will be provided. If the predicted volatility is not higher than the defined predicted volatility, then the deviation verification strategy is executed.

8. The method for predicting the safety of Dendrobium officinale extract based on multi-index joint modeling as described in claim 7, characterized in that: The deviation verification strategy specifically refers to: The median is selected from the predicted safety activity quantification value sequence of the target sample and marked as the target safety activity quantification value; The minimum absolute value of the difference between the target safety activity quantification value and the corresponding comprehensive activity quantification curve of the target sample is marked as the prediction deviation value. The prediction deviation value is compared with the preset defined prediction deviation value, and the result is marked as the prediction deviation ratio. The prediction deviation ratio is compared with the preset defined prediction deviation ratio; If the prediction bias ratio is lower than the defined prediction bias ratio, the safe concentration of the extract of the target sample is analyzed based on the current predicted safe activity quantification value, and the safe concentration of the extract of the target sample is output. If the prediction bias ratio is not lower than the defined prediction bias ratio, a constrained refitting strategy is applied to the comprehensive activity quantification curve of the target sample to reanalyze the predicted safety activity quantification value of the target sample.

9. The method for predicting the safety of Dendrobium officinale extract based on multi-index joint modeling as described in claim 8, characterized in that: The constrained refitting strategy specifically refers to: The similarity of the comprehensive activity quantification curve of the target sample with several comprehensive activity quantification curves in the training dataset is compared. The similarity scores are sorted in descending order, and the comprehensive activity quantification curve corresponding to the training dataset with the highest similarity score is extracted and marked as the target reference curve. Obtain the degrees of freedom in the t-distribution error of the target reference curve, and map the degree of freedom interval based on the similarity and the degrees of freedom in the t-distribution error of the target reference curve; The comprehensive activity quantification curve of the target sample is constrained based on the degree of freedom interval and then refitted to re-select the target safety activity quantification value.

10. The method for predicting the safety of Dendrobium officinale extract based on multi-index joint modeling as described in claim 1, characterized in that: The process of supplementing and feeding back experimental data is as follows: Obtain the coverage range of the predicted safety activity quantification value sequence on the comprehensive activity quantification curve of the target sample; A preset coverage determination probability threshold is set, and the ratio of the total number of predicted safety activity quantification values ​​in a certain coverage interval to the total number of predicted safety activity quantification values ​​in the predicted safety activity quantification value sequence is processed, and the processing result is marked as the coverage probability. If the coverage probability of a certain coverage interval is greater than the coverage determination probability threshold, then it is marked as satisfying the coverage determination probability threshold; Filter and determine the smallest coverage interval that meets the coverage determination probability threshold. If the smallest coverage interval can be found, trigger the experimental data supplementation feedback process. If the minimum coverage interval that meets the coverage determination probability threshold cannot be found after the search, an abnormal warning message will be fed back.