A Data-Driven Modeling and Predictive Control Method for Respiratory Virus Inactivation
By combining molecular dynamics simulations and machine learning models, we can identify viral structural changes and immune responses, optimize viral inactivation conditions, solve the problem of uncertain inactivation effects in existing technologies, and achieve precise control and safety assessment of the viral inactivation process.
Patent Information
- Application Number
- CN202610476090.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-03
- Estimated Expiration
- 2046-04-13
AI Technical Summary
Existing virus inactivation methods rely on experience or single parameter adjustment, making it difficult to fully characterize the structural changes of viruses under multivariate inactivation processes. This leads to uncertainty in inactivation efficacy and immune safety, and lacks data-driven system modeling and prediction capabilities, making it difficult to identify potential unexpected interaction risks.
The viral structural change matrix was obtained through molecular dynamics simulation, potential unexpected interaction sites were identified using a support vector machine classifier, and the impact of inactivation process variables was predicted using a random forest regression model. By combining immune response simulation and gradient descent optimization algorithm, inactivation conditions were optimized to achieve precise control.
It enables forward-looking assessment and precise control of the virus inactivation process, reduces the risk of unintended interactions, improves the scientific validity and stability of inactivation process parameters, shortens the process optimization cycle, and reduces costs.
Smart Images

Figure CN122024814B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, specifically to a data-driven modeling and predictive control method for respiratory virus inactivation. Background Technology
[0002] In the biomedical field, respiratory viruses, due to their diverse transmission routes and tendency to cause cluster infections, pose a persistent threat to public health and safety, and have always been a key research subject in virology research, vaccine development, and biosafety control. Especially in vaccine development and virus control, virus inactivation processes are crucial for ensuring safety and immunogenicity; their scientific validity and controllability directly affect the quality of the final product and the effectiveness of disease control. Therefore, conducting systematic research on respiratory virus inactivation processes has become an important development direction in this technological field.
[0003] In existing technologies, virus inactivation typically relies on physical or chemical methods, using specific process conditions to render the virus infective. However, due to the complex surface structure and conformational dependence of viruses on environmental conditions, different inactivation conditions often cause significant changes in viral structure and immunological properties. In practical applications, existing methods often rely on experience or single-parameter adjustments to determine inactivation conditions, making it difficult to comprehensively characterize the structural changes of viruses under multivariate inactivation processes, leading to uncertainties in inactivation efficacy and immune safety. Furthermore, changes in key viral surface structures during inactivation may affect the interaction between the virus and antibodies. Existing technologies lack systematic analytical methods for the correlation between viral structural changes and antibody binding behavior, making it difficult to identify potential unintended interaction risks in a timely manner. In some cases, inappropriate inactivation conditions may induce non-target immune responses or immune dysregulation, adversely affecting vaccine safety and efficacy; however, these risks are often difficult to accurately predict and assess in the current process design stage. Simultaneously, with the continuous development of viral data scale and computational methods, how to effectively utilize multi-source data for modeling and analyzing the virus inactivation process remains a weak link in existing technologies. Existing solutions generally lack data-driven system modeling and prediction capabilities, making it difficult to achieve forward-looking assessment and control decisions for the inactivation process under complex process variables, resulting in long optimization cycles and high trial-and-error costs for the inactivation process. Summary of the Invention
[0004] The purpose of this invention is to provide a data-driven modeling and predictive control method for respiratory virus inactivation, thereby solving the problems existing in the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a data-driven modeling and predictive control method for respiratory virus inactivation, comprising the following steps:
[0006] S1. Initial viral characteristic data were obtained from the viral surface structure database through molecular dynamics simulation, and the structural change matrix of the virus under different inactivation conditions was obtained after processing.
[0007] S2. Based on the structural change matrix, a support vector machine classifier is used to extract features of viral characteristics and antibody binding sites to determine the set of potential unexpected interaction sites.
[0008] S3. If the set of potential unexpected interaction sites exceeds the preset threshold, the influence distribution of inactivation process variables on interaction intensity is predicted by random forest regression model to obtain the range of optimized inactivation conditions.
[0009] S4. Obtain simulated immune response data from the optimized inactivation condition range, determine the probability of immune disorder response, and obtain the precise adjustment vector for regulation;
[0010] S5. For the precise adjustment vector of regulation, the gradient descent optimization algorithm is used to iterate the virus characteristic model to determine the final inactivation process parameter set;
[0011] S6. Simulate the vaccine development process using the final inactivation process parameter set to obtain a set of safety and efficacy evaluation indicators;
[0012] S7. If the set of safety and effectiveness assessment indicators meets the preset standards, then output a complete description of the control scheme.
[0013] As can be seen from the above technical solution, the present invention has the following beneficial effects:
[0014] This invention systematically integrates viral surface structure data, inactivation process variables, and immune response data, and introduces molecular dynamics simulations and various data-driven modeling methods to predict and analyze the structural changes of the virus under different inactivation conditions and their immunological impacts. This allows for the early identification of potential unexpected interactions and probabilities of disordered immune responses during the inactivation process design phase, enabling proactive assessment and precise control of the inactivation process. Instead of relying on single experiences or trial-and-error methods to determine inactivation conditions, this invention develops process control schemes based on prediction results to regulate the inactivation of respiratory viruses. This effectively improves the scientific rigor and stability of inactivation process parameter determination, reduces safety risks associated with inappropriate inactivation conditions, shortens the process optimization cycle, reduces experimental costs, and helps maintain good immunogenicity while ensuring the safety of virus inactivation, thus better meeting the practical application needs in vaccine development and virus control. Attached Figure Description
[0015] Figure 1 This is a flowchart of the respiratory virus inactivation data-driven modeling, prediction, and control method of the present invention;
[0016] Figure 2 The control chart is optimized for this invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] like Figure 1 and Figure 2 As shown, the present invention provides a technical solution: a data-driven modeling and predictive control method for respiratory virus inactivation, comprising the following steps:
[0019] S1. Initial viral characteristic data were obtained from the viral surface structure database through molecular dynamics simulation, and the structural change matrix of the virus under different inactivation conditions was obtained after processing.
[0020] S2. Based on the structural change matrix, a support vector machine classifier is used to extract features of viral characteristics and antibody binding sites to determine the set of potential unexpected interaction sites.
[0021] S3. If the set of potential unexpected interaction sites exceeds the preset threshold, the influence distribution of inactivation process variables on interaction intensity is predicted by random forest regression model to obtain the range of optimized inactivation conditions.
[0022] S4. Obtain simulated immune response data from the optimized inactivation condition range, determine the probability of immune disorder response, and obtain the precise adjustment vector for regulation;
[0023] S5. For the precise adjustment vector of regulation, the gradient descent optimization algorithm is used to iterate the virus characteristic model to determine the final inactivation process parameter set;
[0024] S6. Simulate the vaccine development process using the final inactivation process parameter set to obtain a set of safety and efficacy evaluation indicators;
[0025] S7. If the set of safety and effectiveness assessment indicators meets the preset standards, then output a complete description of the control scheme.
[0026] In the above embodiments, this method focuses on viral structural changes and immune response regulation, using data-driven modeling to predict and optimize the inactivation process. First, molecular dynamics simulations are used to perform high-temporal-resolution calculations of conformational changes in viral surface proteins under different inactivation conditions, including parameters such as temperature, chemical inactivating agent concentration, pH value, and inactivation time. By comparing the structures of key viral surface proteins before and after simulation, a structural change matrix is constructed to characterize the response of viral properties to changes in inactivation conditions.
[0027] Building upon this, a support vector machine classifier is introduced to extract and classify high-dimensional features related to antibody binding sites in the structural change matrix, identifying a set of potential unexpected interaction sites that may trigger abnormal immune responses. When the number or overall weight of this set exceeds a preset threshold, it indicates that the current inactivation conditions may lead to adverse changes in viral antigenic epitopes.
[0028] Subsequently, a random forest regression model was used to model the nonlinear relationship between different inactivation process variables and the intensity of unexpected interactions, predicting the distribution of the impact of each process variable on the interaction intensity, thereby defining the range of optimized inactivation conditions that can reduce the risk of abnormal interactions. Within this range, an immune response simulation model was further introduced to generate corresponding immune response data to assess the probability of immune dysregulation, and based on this, a precise adjustment vector for regulation was formed.
[0029] Based on this, a gradient descent optimization algorithm is used to iteratively update the virus characteristic model, gradually bringing the model output closer to the optimal state combining immune safety and inactivation effectiveness, ultimately obtaining a stable and convergent set of inactivation process parameters. This parameter set is then used to simulate the entire vaccine development process, obtaining a set of safety and efficacy evaluation indicators, including antigen integrity, safety, immunogenicity, and stability. Based on preset standards, it is determined whether a complete control plan should be output.
[0030] Compared with existing virus inactivation methods that rely on experience or single-parameter optimization, the above implementation method has at least the following advantages: Through the synergistic application of molecular dynamics and machine learning models, a refined characterization of viral structural changes is achieved, improving the controllability of the inactivation process in protecting key antigen structures; by using support vector machines and random forest models to identify and quantify potential unexpected interaction sites, the risk of inducing abnormal immune responses during inactivation is reduced; data-driven predictive control methods are used to globally optimize inactivation process parameters, reducing the cost of repeated experiments and improving the efficiency of inactivation process design; by introducing immune response probability assessment and gradient descent iterative mechanisms, the final determined inactivation process parameters achieve a balance between safety and efficacy, thereby improving the reliability and reproducibility of the vaccine development process.
[0031] Specifically, S1 includes obtaining initial viral characteristic data from a viral surface structure database through molecular dynamics simulations, processing it to obtain a preliminary structural model; for the preliminary structural model, inactivation condition variables, including temperature and pH, are introduced to simulate protein folding changes under different inactivating agent concentrations, resulting in a conditional influence dataset; based on the conditional influence dataset, a structural change matrix is constructed where row vectors represent inactivation condition variables and column vectors represent viral surface structural sites, and the matrix element values are determined to be site displacement differences, which are obtained by comparing the Euclidean distance between the coordinates of viral surface structural sites before and after inactivation; if the matrix element values exceed a preset threshold, the simulation parameters are adjusted to optimize the protein folding change simulation, resulting in a refined structural change matrix; trend tracking features are extracted from the refined structural change matrix, and the accuracy of the matrix is confirmed by combining it with control data obtained from the conditional influence dataset, thus obtaining the structural change matrix of the virus under different inactivation conditions.
[0032] In this embodiment, initial viral characteristic data is first obtained from a viral surface structure database. This initial viral characteristic data includes three-dimensional coordinate information of viral surface proteins, resolution information from structural analysis, missing fragment annotation information, and annotation information related to surface exposure. The three-dimensional coordinate information originates from the master record version of the same viral strain in the database, which is determined by deduplicating and merging structural entries of the same viral strain within the database. Resolution information is used to filter out entries with excessive structural uncertainty. The resolution screening threshold is determined based on the correspondence between structural fitting errors and subsequent conformational drift errors in historical modeling tasks, ensuring that the conformational drift background noise of the selected entries remains stable within a preset tolerance range. Missing fragment annotation information is used for subsequent structural completion. The completion boundary is jointly determined by the continuity of the main chain and the consistency of the secondary structure on both sides of the missing fragment, avoiding the introduction of discontinuous conformations. After data screening, preprocessing was performed on the target structure. This preprocessing included removing ligand residues irrelevant to inactivation, completing missing side chains, correcting abnormal bond lengths and angles, adding necessary hydrogen atoms, and performing preliminary charge homogenization. The hydrogen atom addition rule was determined by the protonation state corresponding to the pH value. The pH value was derived from the target environment of the inactivation process, set according to the buffer system of the production system, and determined based on the median value of online monitoring records to avoid frequent switching of the protonation state caused by instantaneous fluctuations. After preprocessing, a preliminary structural model was obtained. This model was used to obtain the conformation after local stress release through energy minimization. The termination condition for energy minimization was determined based on the magnitude of the total energy change in adjacent iterations. Iteration was stopped once the total energy change entered a stable region, thus providing a numerically stable starting point for subsequent molecular dynamics simulations.
[0033] For the preliminary structural model, inactivation conditional variables were introduced, and protein folding changes were simulated. These inactivation conditional variables included temperature and pH values. The temperature value was determined based on the target temperature window of the proposed inactivation process, with the upper and lower limits corresponding to the process control capability serving as boundaries. The upper and lower limits of the temperature window were determined through the statistical distribution of equipment temperature control errors, ensuring that temperature deviation remained within a controllable range under most operating conditions. The pH value was determined based on the buffer system formulation target, with the stable range of online detection after process scale-up serving as boundaries. The stable range was determined through the quantile range of continuous sampling, avoiding the inclusion of short-term spikes in the modeling. Simultaneously, an inactivating agent concentration variable was introduced. The inactivating agent concentration was determined based on the process-operable concentration range, which was jointly limited by the allowable concentration of raw materials, mixing uniformity constraints, and safe disposal constraints, with the intersection between the minimum effective inactivation concentration and the maximum acceptable antigen retention concentration serving as the boundary. For each combination of temperature, pH, and inactivator concentration, a corresponding simulation system was constructed. System construction included adding a solvent environment to the protein, setting the ionic strength, establishing a periodic boundary, and assigning initial velocities. The ionic strength was determined based on the salt concentration in the actual formulation and corrected using the intersection of osmotic pressure and the protein stability window. The boundary size of the periodic boundary was determined based on the protein's circumscribed envelope size and the minimum solvent layer thickness. The minimum solvent layer thickness was determined by the distance requirement to avoid protein self-interactions, ensuring sufficient spacing between the protein surface and its periodic mirror image. The initial velocity assignment was determined based on the statistical law of thermal equilibrium at the target temperature, and multiple sets of initial velocities generated using different random seeds were used to reduce randomness. The number of random seeds was determined based on the variance convergence of the site displacement differences, ensuring the distribution of displacement differences in repeated simulations tended to stabilize. After system construction was completed, an equilibrium phase was executed to bring temperature, pressure, and density into a stable range. The stable range was determined by the mean drift and fluctuation amplitude statistically analyzed using a sliding window. Once the determination criteria were met, the system entered the production phase and its trajectory was recorded. The trajectory recording frequency is determined based on the characteristic timescale of protein folding changes. This frequency is set after pre-running to estimate the relevant time of conformational changes, ensuring that key conformational transitions can be distinguished between adjacent recording frames while avoiding excessive redundant data that could impact subsequent computational efficiency. This yields a conditional influence dataset, which includes a set of surface site coordinates evolving over time under each condition, conformational cluster labels, and statistics related to folding changes. The conformational cluster labels are obtained by clustering trajectory conformations based on similarity measurements. The threshold for similarity measurement is determined by combining gyroscopic stability within the same conformational cluster with inter-cluster separation, enabling the clustering results to distinguish the dominant folding change paths.
[0034] After obtaining the conditional influence dataset, a structural change matrix is constructed. The row vectors of the structural change matrix represent combinations of inactivation condition variables, and their order is determined by the lexicographical order of temperature, pH, and inactivator concentration to ensure consistent condition indexes and facilitate traceability. The column vectors of the structural change matrix represent viral surface structural sites. The determination of surface structural sites begins with a surface exposure assessment of viral surface proteins, based on solvent accessibility calculations. The solvent probe size is determined and fixed based on the effective scale of water molecules to avoid inconsistencies in the site set caused by probe size variations. Further structural sites are identified within the exposed residues, characterized by residue representative points. These representative points are selected based on the geometric center of the residue backbone to ensure comparability of displacements between different residues. The matrix element values are defined as site displacement differences, which are obtained by comparing the coordinate differences of the same structural site before and after inactivation. The specific calculation process is as follows: First, the coordinates of the site in three-dimensional space are read in the baseline state. Then, the three-dimensional coordinates of the same site are read at the end of the stable phase of the corresponding inactivation condition. Subsequently, the differences in the three coordinate directions are calculated respectively. Then, the differences in the three directions are squared and summed. Finally, the square root of the summation result is taken to obtain the linear displacement distance of the site in three-dimensional space. This linear displacement distance is written as the site displacement difference into the corresponding element of the matrix. To reduce the influence of instantaneous fluctuations, the coordinate reading after inactivation uses the average coordinates of the time window at the end of the stable phase. The length of the time window is determined according to the autocorrelation decay rate of the site displacement difference, so that the average window can cover the main thermal fluctuation cycle and retain the net effect of folding changes. The coordinate reading in the baseline state also uses the average coordinates of the time window after equilibrium is completed to ensure the consistency of the comparison before and after. This results in a structural change matrix with conditions as rows and sites as columns, where each element has a clear geometric displacement meaning and corresponds to a specific condition and a specific site.
[0035] When any element in the structure change matrix exceeds a preset threshold, the simulation parameters are adjusted to optimize the protein folding simulation and generate a refined structure change matrix. The preset threshold is used to distinguish between normal thermal fluctuation displacements and inactivation-induced abnormal conformational migrations. The threshold is determined as follows: First, under reference conditions without introducing changes in inactivator concentration, the system is repeatedly simulated using the same temperature and pH values to obtain the distribution of reference site displacement differences; then, the mean and standard deviation of this distribution are calculated; finally, the threshold is set to the mean plus three times the standard deviation, ensuring that the vast majority of displacements under reference conditions fall within the threshold, and displacements exceeding the threshold are judged as abnormal migrations. The factor of three times the standard deviation is determined based on a trade-off between false alarm rate and false negative rate, ensuring that the abnormal displacement detection rate meets the preset target. After triggering the threshold, the adjusted simulation parameters include at least the time step, constraint strategy, and sampling length. The time step value is determined based on energy drift monitoring, and the final time step is determined by gradually reducing the time step until the energy drift enters a stable range. The constraint strategy is used to limit numerical instability caused by high-frequency vibrations. The constraint strength is determined based on the balance between maintaining bond length stability and preserving conformational degrees of freedom, so that main chain conformational changes can still occur sufficiently. The sampling length is used to improve the coverage of folding change paths. The sampling length is determined based on the convergence of the site displacement difference distribution, and is determined by extending the simulation until the variation of key statistics within a continuous window is less than the preset tolerance. The tolerance value is determined based on the median level of the statistical differences between different repeated simulations. After completing the parameter adjustment, the system construction, equilibration, and production stage calculations are re-executed for the combination of triggering anomalies. The matrix elements are refilled according to the aforementioned coordinate reading and displacement difference calculation process to form a refined structural change matrix. The refined structural change matrix meets the reliability requirements of the threshold determination in terms of numerical stability and statistical convergence.
[0036] After obtaining the refined structural change matrix, trend tracking features were extracted and their accuracy confirmed by comparing them with control data. The extraction of trend tracking features focused on the displacement changes of the same structural site across different conditional sequences. Specifically, the displacement difference sequence of the same site after row vector sorting was smoothed. The smoothing window length was determined based on the local fluctuation scale of the displacement difference sequence, ensuring that smoothing suppressed random noise while preserving the overall trend driven by the conditions. Subsequently, the increment of the displacement difference between adjacent conditional points was calculated, and the sign consistency of the increments was statistically analyzed to determine the monotonicity with changing conditions. The number and location of extreme points in the displacement difference sequence were then statistically analyzed to characterize the sensitive intervals of folding changes. Finally, the dispersion of the displacement difference sequence was calculated to reflect the sensitivity of the site to conditional perturbations. The control data came from pre-set control conditions in the conditional influence dataset. These control conditions referred to simulated data where temperature and pH remained constant and the inactivator concentration was a baseline value. The baseline value was determined based on the actual concentration state when no inactivator was introduced in the process, ensuring that the control conditions reflected the intrinsic fluctuation level of the system. The accuracy verification process includes: comparing the distribution of displacement differences in the control rows of the refined structural change matrix with the independently calculated distribution of displacement differences in the control data. The comparison index is determined based on both the distribution position and the distribution width. The determination is completed by checking whether the median deviation and quantile span deviation of the two fall within the preset tolerance range. The preset tolerance range is determined based on the statistical range of natural differences between repeated simulations, so that the tolerance can cover both randomness and systematic bias. Through the joint constraints of trend tracking features and control consistency comparison, the structural change matrix of the virus under different inactivation conditions is finally obtained.
[0037] Specifically, S2 includes obtaining protein folding change data corresponding to inactivation condition variables from the structural change matrix, using a support vector machine (SVM) classifier to initially classify viral characteristics, and obtaining a feature vector set; for the feature vector set, introducing the inactivator type variable combined with site displacement difference, extracting antibody binding site-related attributes, and determining a preliminary list of interaction sites; based on the preliminary list of interaction sites, obtaining change trend tracking information from the conditional influence dataset, and determining that if the site displacement difference exceeds a preset threshold, it is marked as a potential unexpected site, thus obtaining a screening site group; through the screening site group, integrating viral surface structure data to optimize the SVM classifier, where the SVM classifier takes the site displacement difference as input, constructs a decision boundary by maximizing the classification margin, outputs site classification results, constructs a site stability evaluation index based on the site classification results, and obtains a refined interaction site matrix; aggregating unexpected interaction attributes from the refined interaction site matrix, and verifying them with change trend tracking, to determine a set of potential unexpected interaction sites.
[0038] In this embodiment, protein folding change data is first extracted using the structural change matrix as the input source. Specifically, each row of the structural change matrix is read sequentially according to the combination order of the inactivation condition variables, and the site displacement differences of all viral surface structural sites corresponding to that row are combined to form a folding change data sequence under the same conditions. To ensure the consistency of subsequent calculations, site index alignment is performed first. The site index alignment is implemented as follows: using the master sequence of the surface site determined in the viral surface structure database as the unique reference sequence, the column index of the structural change matrix is checked against the reference sequence number one by one. When a missing site is found, missing site repair is performed. The missing site displacement difference is obtained by retrieving the average coordinates of the stable window of adjacent frames. The length of the stable window is determined by the autocorrelation decay rate of the displacement difference time series, so that the window covers the main thermal fluctuation cycle without crossing the conformational transition segment. After missing data repair is completed, scale unification is performed. The scale unification is achieved as follows: under control conditions, all locus displacement differences are collected to form a baseline set. The baseline set is sorted from smallest to largest, and the 25th and 75th percentiles are taken as the baseline span. Then, the locus displacement differences under any condition are translated and scaled so that their deviation from the baseline median falls into a uniform dimension. The values of the 25th and 75th percentiles are determined by evaluating the span stability of multiple batches of control simulations. The combination of quantiles with the smallest inter-batch fluctuations is selected to avoid a small number of extreme values affecting the scale unification results.
[0039] After preparing the folding change data, a support vector machine classifier is used to initially classify the virus characteristics and form a set of feature vectors. Specifically, the core input is the "site displacement difference under specific inactivation conditions," while local neighborhood statistics are introduced to reflect the folding coupling relationship. The determination of the local neighborhood begins by constructing the surface site adjacency relationship. This is done by calculating the spatial linear distance between any site and the remaining sites in the baseline structure, sorting the distances from smallest to largest, and taking 10% of the distance distribution as the neighborhood radius threshold. This threshold is determined based on the stable interval of neighborhood connectivity obtained statistically across different viral structural entries. A quantile is selected that ensures the average number of neighbors falls within the center of the stable interval, thus ensuring that the neighborhood covers the same local structural segment while avoiding cross-segment misconnections. After obtaining the neighborhood, the average level and discrete level of the displacement differences within the neighborhood are calculated for each point. The average level is obtained by summing all displacement differences within the neighborhood and dividing by the number of neighborhood points. The discrete level is obtained by first calculating the difference between each neighborhood displacement difference and the average level, then squaring the difference, summing the results, dividing by the number of neighborhood points, and taking the square root. When the number of neighborhood points is insufficient, neighborhood expansion is triggered. The expansion rule is to increase the neighborhood radius threshold to 15% until the number of neighborhood points reaches a preset lower limit. The preset lower limit is set to 8, which is determined by comparing the statistical stability of neighborhoods for various surface resolution structures, so that the discrete level calculation is not affected by insufficient sample size. This forms the input description for each point under each condition. The input description consists of the point's own displacement difference, the neighborhood average level, the neighborhood discrete level, and the increment of the point's change in the condition sequence. The increment of change is obtained by subtracting the displacement difference of the same point under adjacent conditions. The adjacency relationship of the condition sequence is determined according to the sorting rules of the inactivation condition variables to avoid increment distortion caused by crossing discontinuous conditions.
[0040] The classification labels are obtained using a priori partitioning method. The specific process of priori partitioning is as follows: Under control conditions, all site displacement differences are collected and sorted, and the 95th percentile is taken as the upper bound threshold for control stability. This threshold is determined based on the fact that the probability of the displacement difference exceeding this percentile under control conditions is controlled within 0.05, so that the natural fluctuations under the steady state are limited within the threshold. Subsequently, under non-control conditions, the displacement difference of the same site is calculated to see if it exceeds the upper bound threshold for control stability. Those exceeding the threshold are marked as unstable candidates, and those not exceeding the threshold are marked as stable candidates. Then, a consistency constraint for repeated simulation is introduced. For each site, the consistency ratio of its candidate labels in repeated simulation is calculated. Sites with a consistency ratio lower than 0.7 are removed from the training set to avoid noisy labels. The value of 0.7 is based on the trade-off between the training set size and the reliability of the labels, and the minimum consistency ratio is selected that makes the training set size decrease by no more than 20% and the validation error significantly reduced. After prior partitioning, a support vector machine (SVM) classifier is trained. The training process aims to maximize the classification margin between the two classes. The specific steps are as follows: First, a similarity matrix is calculated between the training samples. The similarity is calculated using a radial basis function kernel. The width parameter of the similarity is determined through cross-validation. Cross-validation is implemented by stratifying the training samples according to conditional combinations and using an alternating evaluation method to assess the classification stability score under different width parameters. The classification stability score is composed of the "recall rate of the unstable class in the validation set" and the "false positive rate of the stable class," with the weight of the unstable class recall set to 0.7 and the weight of the stable class false positive rate set to 0.3. The weight values are determined based on the principle that the cost of missing potential risk sites for subsequent process control is higher than the cost of false positives. The penalty strength parameter is also determined using an alternating evaluation method. The candidate range of the penalty strength parameter increases from small to large, with the increment step determined based on the dense interval of the inflection point of the validation error curve, ensuring that the search focuses on the region with the greatest impact on performance. Training is achieved through an iterative update approach. After each iteration, the classification boundary is recalculated to increase the distance from the boundary to the nearest sample while controlling for a small number of samples with acceptable error. The stopping condition is "the validation stability score improvement over three consecutive iterations does not exceed 0.001." The value of 0.001 is determined based on the natural noise level of score fluctuations across multiple batches of data, selecting a threshold slightly higher than the noise level to avoid premature stopping. After training, the set of samples closest to the classification boundary is determined, corresponding to support vectors. Support vector identification is achieved by calculating the distance from each sample to the classification boundary. Distance calculation involves determining the sign of the sample's boundary response value in kernel space and using its absolute value as the proximity scale; a smaller distance indicates closer proximity to the boundary.
[0041] After obtaining the preliminary classification results, a feature vector set is formed. The construction process of the feature vector set is as follows: using the locus as an index, the classification labels of the same locus across all conditions are summarized and statistically analyzed. The statistics include the proportion of unstable labels, the proportion of stable labels, the median boundary distance, the lower quartile of the boundary distance, and the upper quartile of the boundary distance. The proportion is calculated by dividing the corresponding label count by the number of conditions participating in the statistics for that locus. The number of conditions is the number of valid conditions. Loci with more than 0.2 of the total number of conditions for missing data repair are removed. The value of 0.2 is based on the fact that excessive missing data repair will significantly reduce the confidence of the shift difference. The quartile of the boundary distance is obtained by taking the 25th and 75th percentiles after sorting the distance sequence. It is used to describe the classification confidence interval and avoid the influence of extreme values when using only the mean. The above statistics, together with the maximum increment of the locus in the condition sequence and the condition index where the maximum increment occurred, constitute the feature vector of that locus, so as to extract interaction attributes in combination with the inactivator type variable.
[0042] For the feature vector set, an inactivator type variable is introduced, and antibody binding site-related attributes are extracted by combining site shift differences, thus forming a preliminary list of interaction sites. The inactivator type variable is determined based on the list of inactivator types in the process formulation. The list is constrained by both the raw material batch record and the formulation version, with the formulation version as the primary key to determine a unique type identifier, avoiding confusion caused by different formulations with the same name. After associating the type identifier with each condition combination, type grouping statistics are performed for each site. The statistical method is to calculate the unstable tag percentage and median boundary distance of the site under each inactivator type, and calculate the difference magnitude between different types. The difference magnitude is obtained by taking the difference between the maximum and minimum medians, which is used to identify inactivator-dependent responses. The determination of antibody binding site-related attributes includes two categories: direct attributes and proximity attributes. Direct attributes are derived from annotations of antibody-binding regions in a viral surface structure database. Mapping the annotation residue numbers to surface site indices yields the set of direct binding sites. Proximity attributes are achieved through spatial proximity inference. This involves calculating the spatial distance from any surface site to the nearest site in the direct binding site set, sorting the distance distribution, and using the 5th percentile as the proximity threshold. This threshold is determined based on a stable lower bound of antibody-interface contact distances across different structural entries, ensuring the proximity threshold covers tightly contacted areas without spreading to distant regions. Sites meeting either the criteria of "belonging to the direct binding site set" or "distance to the nearest direct binding site not exceeding the proximity threshold" are labeled as binding-related sites. Subsequently, sites with an unstable label percentage exceeding 0.3 and a median boundary distance lower than the median boundary distance of all sites are selected as initial interaction sites. The value of 0.3 is based on the statistical correlation between the percentage of unstable labels in the training set and the subsequent trend validation pass rate, selecting the minimum percentage threshold that achieves the target pass rate, thereby reducing the number of invalid candidate sites.
[0043] Based on the preliminary list of interactive sites, trend tracking information was extracted from the conditional influence dataset and threshold determination was performed to form a screening site group. The method for extracting trend tracking information was as follows: for each candidate site, the displacement difference time series under various conditions was retrieved, and window smoothing was first performed. The window length was 20 times the time series sampling interval. The value of 20 was based on the statistical number of sampling points corresponding to the main cycle of thermal fluctuations under the control condition, so that the smoothing covered one main cycle and suppressed high-frequency noise. Then, the incremental sequence of displacement difference between adjacent sampling points was calculated, and the median and upper quartile of the absolute increment of the incremental sequence were calculated to characterize the intensity of the mutation. A preset threshold for the locus shift difference is used to identify potential unexpected loci. The determination process is as follows: Under control conditions, the shift differences of all loci in the preliminary interactive locus list are collected to form a control shift set. After sorting the control shift set, the 99th percentile is taken as the upper threshold for anomalies. This threshold is determined based on controlling the false alarm probability under control conditions to within 0.01. To enhance robustness to batch differences, the maximum value of the upper threshold for anomalies from different batches is taken as the final threshold to avoid false alarms caused by excessively low thresholds due to batch fluctuations. For each candidate locus, if its shift difference exceeds this final threshold during the stable phase under any non-control condition, and the upper quartile of its increment sequence exceeds twice the upper quartile of the control increment, then the locus is marked as a potential unexpected locus and included in the screening locus group. The value of twice is based on a balance assessment between the scale expansion of the control increment distribution and the mutation identification sensitivity, selecting a fold that significantly improves the mutation detection rate while controlling the increase in false alarms.
[0044] The support vector machine classifier was optimized using a selected site group to generate a refined interaction site matrix. The optimization training data construction process involved using samples marked as potentially unexpected sites in the selected site group as key samples. Simultaneously, sites that did not trigger a threshold under the same conditions were extracted from the spatial neighborhood of each key sample as control samples. The number of control samples was three times the number of key samples; the value of 3 was based on an assessment of the minimum control size required to maintain stable classification boundaries under class imbalance conditions. Subsequently, viral surface structure data was introduced as auxiliary constraint features, including site surface exposure levels and neighborhood coupling strength. Surface exposure levels were determined by ranking solvent accessibility and then dividing the data into ternary segments, with the 33rd and 66th percentiles as the dividing thresholds. These thresholds were based on stable segmentation points in the exposure distribution across different structural entries. Neighborhood coupling strength was calculated by assessing the degree of coordinated change in the displacement differences between the site and its neighbors. Specifically, the median of the displacement difference time series between the site and each neighboring site was subtracted from each of the neighboring sites. Then, the two decentralized series were multiplied point-by-point and averaged. Finally, the average of the average values from all neighboring sites was calculated to obtain the coupling strength index for that site. After feature expansion, the classifier was retrained. During training, a risk weighting was applied to the penalty strength parameter. The risk weighting was determined by the magnitude by which the displacement difference at the site exceeded the threshold. The magnitudes were ranked and divided into four levels, with the highest level weighted four times the lowest level weight. The four levels and the 4-fold factor were chosen based on the hierarchical separability assessment of the magnitude distribution, selecting a grading scheme that provides stronger boundary constraints for high-risk sites without degrading the overall error. After training, new site classification results were obtained, and a site stability assessment index was constructed based on these results. The calculation process for the site stability assessment index is as follows: First, calculate the proportion of stable labels for the site under all conditions. Then, calculate the median boundary distance and the lower quartile of the boundary distance. Map the stable label proportion and boundary distance statistics segment by segment to obtain the stability score. The mapping segment points are determined by the intersection of the score distributions of stable and unstable classes in the training set, ensuring a clear boundary between the two classes. Simultaneously, calculate the difference in inactivator type as a risk sensitivity index. The sensitivity index is the maximum difference between the median boundary distances of different types, combined with the threshold trigger count, and ranked accordingly. The threshold trigger count is the number of conditions under which the site exceeds the final threshold in non-control conditions. The stability score, sensitivity index, combined with relevant attribute markers and threshold trigger count, together constitute the core content of the refined interactive site matrix. The matrix organizes information using site indices as the main thread, ensuring that subsequent aggregation can be completed within the same index system.
[0045] Based on the refined interaction site matrix, unexpected interaction attributes are aggregated and verified by tracking changes in trends, ultimately determining the set of potential unexpected interaction sites. The aggregation process first filters sites based on their association with relevant attributes, then sorts them by stability score from low to high, selecting sites with stability scores below a cutoff point as risk candidates. The cutoff point is the score value corresponding to the intersection of the two score distributions in the training set; this cutoff point minimizes the error between the stable and unstable classes. Subsequently, trend verification was performed on each risk candidate, using three consistency criteria: First, the overlap ratio between the condition index where the displacement difference exceeds the final threshold and the condition index where the unstable label appears in the classification result is not less than 0.6. The value of 0.6 is based on the statistical relationship between the overlap ratio in historical data and the explanatory power of subsequent random forest regression. Second, the duration of the mutation segment is not less than twice the upper bound of the duration of the mutation segment in the control condition. The upper bound of the duration is obtained by statistically analyzing the length of continuous sampling points in the control condition increment sequence that exceed the upper quartile of the control increment. The value of twice is used to distinguish between short burst fluctuations in the control and inactivation-induced persistent mutations. Third, the decline in the recovery segment is insufficient to bring the displacement difference back within the final threshold. The decline amplitude is obtained by the difference between the mutation peak value and the end value of the stable segment. Sites that meet the three consistency criteria are identified as potential unexpected interaction sites and are compiled into a set of potential unexpected interaction sites.
[0046] S3 involves obtaining inactivation process variable data from a set of potential unexpected interaction sites. If the number of sites exceeds a preset threshold, a random forest regression model is used to process the influence of the inactivation process variable data as input to obtain the distribution of interaction intensity. For the distribution of interaction intensity, a condition range optimization attribute is introduced as the basis for adjusting the distribution boundary. The distribution data is integrated with the site set threshold to determine the variable influence prediction result. Based on the variable influence prediction result, the distribution attribute of the regression model is used as the basis for calculating the intensity deviation. Information obtained from the intensity distribution is extracted. If the distribution deviation exceeds the range, the process variable distribution is adjusted to obtain the preliminary optimization conditions. Through the preliminary optimization conditions, the unexpected interaction sites and the preset threshold judgment attribute are aggregated as verification inputs to verify the range of optimization conditions and obtain the final inactivation condition range.
[0047] In this embodiment, the set of potential unintended interaction sites serves as the entry point for risk constraints, transforming the statistical correlation between inactivation process variables and interaction intensity characterization into an optimizable condition range. The initial step involves extracting inactivation process variable data associated with each site from the set of potential unintended interaction sites. This inactivation process variable data originates from the condition index corresponding to the preceding structural changes and site classification results. Temperature, pH value, inactivator concentration, inactivator type identifier, and action time history index are obtained by backtracking using the same condition index. The action time history index is determined by the end window position of the stable phase of the simulation trajectory, ensuring a comparable time baseline between different conditions. To ensure consistency across data sources, the process variables were first standardized. Temperature was defined using the target value as the primary key, along with the control fluctuation range. This range was determined by the median deviation and upper quartile deviation of the online control records, avoiding the use of instantaneous extreme values. pH values were defined using the median value of the buffer system's stability plateau as the primary key, along with the plateau width. The plateau width was determined based on the quantile span of continuous sampling, ensuring the input reflects the achievable process stability. Inactivator concentration was defined using the formulation target value as the primary key, along with the mixing uniformity deviation. This deviation was determined by the upper quantile statistics of the concentration differences from multiple sampling points, avoiding misclassification of localized incomplete mixing as a conditional effect. Inactivator type identification was determined by the joint key of the formulation version and raw material batch record, preventing misclassification due to different specifications with the same name. After standardization, the process variable values under different conditions were aggregated by site index, forming a set of associated samples from sites to process variables. At the sample level, site displacement difference, site stability score, threshold trigger count, and trend change persistence index were added for subsequent characterization of interaction intensity and regression modeling.
[0048] When the number of sites exceeds a preset threshold, variable impact processing of the random forest regression model is initiated. This preset threshold ensures the statistical stability of the regression model's estimation of the multivariate impact distribution. The threshold is determined as follows: first, the median size of historically similar viral surface sites is used as a baseline; then, the effective sample density is calculated by combining the coverage of the current site set in the conditional space. The effective sample density is obtained by dividing the number of different conditional combinations by the number of sites. Subsequently, the error convergence curve of the regression model under cross-validation is used as a criterion to find the inflection point where the marginal improvement of error slows down with increasing site count. The threshold is set to the number of sites corresponding to this inflection point, ensuring that the model error enters a stable range after the threshold is reached, thus avoiding unstable impact distributions when samples are insufficient. After triggering the threshold, regression training samples are constructed. The input of the samples is inactivation process variable data, and the output is the interaction intensity characterization value. The interaction intensity characterization value is generated jointly by the site shift difference and antibody binding-related attributes. The generation process is as follows: First, the site shift difference is scaled uniformly according to the control baseline span to ensure that different sites have comparable variation amplitudes. Then, binding-related sites are assigned higher risk weights. The risk weights are determined based on the correlation strength ranking of binding-related sites with the immune confusion probability index in the training data. Sites with correlation strength in the upper quantile are weighted twice that in the lower quantile. The value of 2 times comes from the validation results of achieving an optimal balance between recall risk sites and control of false positives. Finally, the scaled shift amplitude and the risk weight are weighted and synthesized to obtain the interaction intensity characterization value, which reflects both the structural perturbation amplitude and the immune correlation.
[0049] The training process of the random forest regression model is implemented using split sampling and ensemble averaging as its core. First, the number of trees is determined based on the convergence behavior of the validation error as the number of trees increases. Specifically, multiple models are trained using an incremental sequence, and the median error of each model under stratified cross-validation is recorded. When the error improvement from three consecutive increments falls below a preset convergence threshold, the increment is stopped, and the number of trees in that group is taken as the final value. The convergence threshold is set to 0.001, which is determined based on the upper bound of the natural fluctuation of the validation error in the stable phase, avoiding misjudging noise improvement as genuine improvement. Next, the number of variables participating in each split is determined. This parameter is determined by the stability of the importance ranking of variables in the candidate set. Ranking stability is measured by the consistency ratio of the importance ranking obtained from multiple resampling training sessions. The parameter value with the highest consistency ratio is selected, ensuring interpretable stability of the output influence distribution. The training samples for each tree are obtained by resampling the original sample set, with a resampling ratio of 0.632. This ratio is derived from the desired proportion of resampling samples covering the original samples, allowing the model to balance fitting ability and generalization ability. The criterion for splitting tree nodes aims to minimize the regression residuals. The residuals are calculated based on the dispersion of the current node's sample output, determined by the median absolute deviation of the output from its median. This makes the splitting process insensitive to extreme values. After training, the predicted outputs of all trees for the same input condition combination are integrated. The median of the predicted outputs is used as the predictive interaction strength for that condition to reduce the impact of outlier predictions from a few trees. By traversing the range of input variables and statistically analyzing the distribution of the predicted outputs, the interaction strength influence distribution is obtained. This distribution describes the quantile structure of the output strength with the process variable as the axis, forming the basis for subsequent condition range optimization.
[0050] To address the impact of interaction intensity on distribution, a conditional range optimization attribute is introduced as the basis for adjusting the distribution boundary, and the site set threshold is integrated to determine the variable's influence on the prediction results. The conditional range optimization attribute is jointly composed of the process achievable boundary, the quality target boundary, and the risk tolerance boundary. The process achievable boundary is jointly determined by the equipment control precision, raw material specification tolerance, and mixing uniformity deviation. The quantile interval of historical batch control data is used as the boundary source, and the 5th and 95th percentiles are taken to form a stable achievable interval, avoiding the misinterpretation of occasional extreme capabilities as normal capabilities. The quality target boundary is jointly determined by the consistency requirements of the structural change matrix and the target level of the site stability assessment index. The target level is defined by raising the median of the stability score distribution to the preset target quantile, which is set at 75%. This value is based on a comprehensive trade-off of maintaining a sufficient process window while ensuring that the key structure of the antigen does not drift significantly. The risk tolerance boundary is jointly determined by the threshold of the number of potential unexpected interaction sites and the threshold trigger strength. The threshold of the number of sites is taken as the lower limit of the threshold for triggering regression modeling in the previous step to maintain consistency in risk control logic. The threshold trigger strength is determined by the stratification of the magnitude of the site displacement difference exceeding the threshold, and the risk of the magnitude in the upper quartile interval is regarded as an unacceptable interval. The process of adjusting the distribution boundary is as follows: First, align the distribution of interaction intensity impact with the process achievable boundary, and delete variable value ranges that exceed the achievable boundary; then, introduce a quality target boundary within the remaining range, and filter variable combinations whose predicted interaction intensity is below the target quantile interval; subsequently, integrate the site set threshold into the distribution data integration, and count the number of predicted high-risk sites for each variable combination. The calculation of the number of predicted high-risk sites is based on the fact that sites whose predicted interaction intensity under that variable combination exceeds the risk tolerance quantile are considered high-risk sites. The risk tolerance quantile is taken as the 90th percentile, which is used to isolate extreme high-intensity cases outside the optimization range and maintain sensitivity to risk peaks. After completing the above boundary adjustment and threshold integration, the variable impact prediction results are obtained. These results include the ranking of the contribution of each process variable to the interaction intensity distribution within the feasible range and the candidate condition regions that satisfy multiple boundary constraints.
[0051] Based on the predicted results of variable influence, the intensity distribution information is extracted using the distribution attributes of the regression model as the basis for calculating the intensity deviation. When the distribution deviation exceeds the range, the distribution of process variables is adjusted to form preliminary optimization conditions. The extraction of intensity distribution information includes three types of statistical descriptions: distribution center, distribution dispersion, and distribution skewness. The distribution center is the median of the predicted interaction intensity, reflecting the typical intensity level; the distribution dispersion is the difference between the 75th and 25th percentiles of the predicted interaction intensity, reflecting the uncertainty span; the distribution skewness is determined by comparing the difference between the 90th percentile and the median with the difference between the median and the 10th percentile, used to identify high-intensity long-tail risks. The intensity deviation is calculated based on the target distribution, which is obtained by extrapolating the interaction intensity characterization values under control conditions using a safety margin. The magnitude of the safety margin extrapolation is determined by the maximum median drift of the control distribution across different batches, ensuring that the target distribution remains conservatively safe under batch variations. The preset deviation range is derived from the risk tolerance boundary, specifically using the 90th percentile of the target distribution as the upper limit of the deviation. This upper limit is used to ensure that the interaction intensity does not enter the high-risk tail in most cases. If the 90th percentile of the current predicted distribution exceeds the deviation limit, the distribution deviation is deemed to be out of range, triggering an adjustment to the process variable distribution. The adjustment prioritizes variable importance, shrinking the value range of the variable that contributes most to the intensity deviation. The shrinkage magnitude is determined by the length of the sensitive segment that causes the intensity upper tail to rise in the distribution. The sensitive segment is obtained by statistically analyzing the 90th percentile rate of change using a sliding window on the variable axis. The rate of change threshold is twice the upper bound of the rate of change under the control condition; this double value is used to distinguish between natural fluctuations and significantly sensitive segments. After shrinkage, collaborative correction is performed on the remaining variables. Collaborative correction is determined based on the interaction importance of variables. Interaction importance is obtained by comparing the difference between the magnitude of the upper tail change in the predicted intensity when both variables are changed simultaneously and the sum of the magnitudes when they are changed separately. A larger difference indicates a stronger interaction, which is then prioritized. Through the above adjustments, preliminary optimization conditions are obtained, expressed in the form of variable value ranges and satisfying the deviation range constraints.
[0052] The unexpected interaction sites and preset threshold judgment attributes are aggregated through preliminary optimization conditions as validation inputs to verify the range of optimized conditions and obtain the final inactivation condition range. The validation input is constructed with the site set as the core. The predicted interaction intensity of each site under the condition combinations covered by the preliminary optimization conditions is recalculated to form a site risk profile. The risk profile uses the tail quantile of the site's intensity within the condition range as the representative value, with the representative value being the 90th percentile, used to focus on the most unfavorable scenario. The preset threshold judgment attributes include two items: a site set quantity threshold and a risk site proportion threshold. The site set quantity threshold follows the trigger threshold starting from S3 to maintain consistency in risk scale judgment; the risk site proportion threshold is set to 0.1. The determination of 0.1 is based on controlling high-risk sites within a small proportion range while maintaining a sufficient process window and ensuring a manageable input scale for subsequent immunization disorder probability assessment. The verification process first counts the number of sites whose risk profiles exceed the upper limit of the target distribution deviation within the initial optimization conditions. If the number exceeds the threshold for the number of sites in the site set, the range is deemed too wide, triggering further narrowing. Next, the proportion of sites exceeding the upper limit to the total number of sites is counted. If this proportion exceeds the risk site proportion threshold, the risk concentration is deemed too high, triggering adjustment. Adjustment prioritizes narrowing the sensitive segments of variables that cause risk site aggregation, and the verification indicators are recalculated until both threshold constraints are met. Once the constraints are met, the final inactivation condition range is formed. This range consists of a joint value interval of a set of process variables. The interval boundaries simultaneously satisfy multiple constraints of process achievability, quality objectives, and risk tolerance, and maintain consistency with the risk profile of the potential unexpected interaction site set.
[0053] S4 includes obtaining simulated immune response data from the optimized inactivation condition range, aggregating and processing the site interaction intensity through response data aggregation, whereby the aggregation processing includes classifying and summarizing the response data by site, calculating the average interaction intensity, and determining the probability of immune dysregulation; for the probability of immune dysregulation, a probability threshold calibration is introduced as the adjustment basis, whereby the threshold calibration compares the probability with a preset standard value, and uses distribution boundary adjustment to integrate the influence of process variables, whereby the distribution boundary adjustment includes expanding the variable influence interval and determining the preliminary form of the precise adjustment vector; based on the preliminary form of the precise adjustment vector, aggregating the simulated response data and calculating the intensity deviation, whereby the intensity deviation is calculated by subtracting the interaction intensity value from the average response data to obtain the immune response simulation verification attribute, and determining if the verification attribute exceeds the preset threshold, then adjusting the vector distribution; through adjusting the vector distribution, extracting the influence distribution prediction for the final inactivation range, whereby the influence distribution prediction includes extracting peak points from the distribution to obtain an optimized version of the precise adjustment vector; obtaining response data aggregation from the optimized version of the precise adjustment vector, whereby the response data aggregation integrates the previous response data through the optimized version vector to determine the final precise adjustment vector.
[0054] In this embodiment, the optimized inactivation condition range is first used as the input boundary for the immune response simulation. Simulated immune response data is obtained from this range, and a unified index system for subsequent calculations is established. The optimized inactivation condition range consists of the joint value range of multiple process variables. The variables include at least temperature, pH value, inactivator concentration, action time, and inactivator type identifier. The inactivator type identifier is derived from the joint unique key of the formulation version and raw material batch record to avoid classification bias caused by different specifications with the same name. To form a uniformly covered and traceable set of condition sample points within this range, the value range of each process variable is hierarchically segmented. The number of segments is set to 5. This value is based on obtaining sufficient boundary and center coverage within the joint range of multiple variables, while controlling the number of condition combinations within the scale where the probability estimation variance enters the stable range. When segmenting, the interval length is calculated first, and then the interval length is divided into 5 equal parts to determine the endpoint values of each segment. The endpoint values are obtained by accumulating segment by segment from the lower bound of the interval upwards. After univariate segmentation, a set of joint conditional sample points is constructed. The construction rules are as follows: lower boundary endpoints, middle boundary endpoints, and upper boundary endpoints of each variable are prioritized for combination to ensure simultaneous coverage of the boundary and central regions. For inactivator type identifiers, a type-by-type rotation method is used to combine them with continuous variables. The rotation order is arranged from high to low historical usage frequency, derived from production batch record counts, to avoid missing less common types in the samples. An immune response simulation is initiated for each joint conditional sample point. The simulation output forms simulated immune response data, organized by site index, and includes at least the corresponding interaction intensity output, immune response intensity output, abnormal response marker output, and response time series. To ensure consistency in the comparison benchmark between sample points under different conditions, response time window alignment is first performed. The alignment process uses the starting point of the stable response phase as a unified zero point. The starting point of the stable response phase is determined by the slope change of the response curve. The determination method is as follows: calculate the difference in response values of adjacent sampling points in the time series, and then divide it by the time interval between adjacent sampling points to obtain the slope sequence. The absolute values of the slope sequences under the control conditions are taken and sorted, and the median is taken as the control slope benchmark. The alignment threshold is set to 0.1 times the control slope benchmark. The basis for determining 0.1 is that after the control conditions enter the plateau period, the slope fluctuation amplitude shows a significant contraction relative to the control slope benchmark. Taking one-tenth of it can distinguish between plateau fluctuations and continuous upward segments. When the absolute value of the slope is lower than this threshold for several consecutive sampling points, the first point that meets the threshold is taken as the starting point of the stable response phase, and a fixed-length window after this point is used as the statistical window. The window length is taken as the median value of the number of sampling points corresponding to the main fluctuation cycle of the plateau period under the control conditions. The main fluctuation cycle is obtained by counting the peak-valley intervals of the control response during the plateau period.
[0055] After obtaining and aligning the simulated immune response data, response data aggregation is performed to process site interaction strength and calculate the probability of disordered immune responses. The aggregation process begins with site-based categorization, using a site index system derived from the previous site screening process and maintained unchanged in this step. This ensures that the outputs of the same site under different conditional sample points are included in the same aggregation unit. Subsequently, the average site interaction strength is calculated by summing the interaction strength outputs of the same site under all valid conditional sample points and then dividing by the number of valid conditional sample points. The number of valid conditional sample points is obtained by removing sample points that do not meet the time window alignment criteria. The upper limit for the removal ratio is set to 0.2. This 0.2 is determined because when the removal ratio exceeds this value, the sensitivity of the average value to the remaining samples increases significantly, leading to unstable probability estimation. After the average value of the loci is completed, a chaotic label is generated for each conditional sample point. The chaotic label is triggered by two types of evidence: one is the abnormal reaction label in the immune response simulation output, which is generated by the discrimination logic inside the simulation module and is consistent with the locus index; the other is the label of the locus interaction intensity exceeding the limit. The threshold for the label of exceeding the limit is taken as 99% of the interaction intensity output distribution under the control condition. The basis for determining this threshold is to control the false alarm probability of the control condition to within 0.01, while maintaining sensitivity to extreme high intensity.
[0056] Specifically, the calculation process for exceeding the limit marker is as follows: In the control condition sample point set, the interaction intensity outputs of all loci are collected, merged, and sorted. The value of the sorted position multiplied by the total number is taken as the threshold. For any non-control condition sample point, if the interaction intensity output of any locus exceeds this threshold, the condition sample point is considered to have exceeded the limit marker. The synthesis rule for the confusion marker is: a confusion marker is determined to be true if either the abnormal reaction marker or the exceeding limit marker is true. The calculation process for the probability of an immune confusion reaction is as follows: The number of condition sample points with true confusion markers is counted, and then divided by the number of valid condition sample points to obtain the probability value. This probability represents the frequency level of a confusion reaction within the optimized inactivation condition range.
[0057] To determine the probability of immune dysregulation, a probability threshold calibration is performed, generating a preliminary form of a precise adjustment vector. A preset standard value of 0.05 is used, determined to keep the frequency of dysregulation within a low-probability range while maintaining a sufficient process window to support subsequent parameter iterations. This standard value is derived from a risk tolerance strategy, constrained by a safety and effectiveness assessment objective and referencing the false alarm level of dysregulation labeling under historical batch control conditions as a lower limit, ensuring the standard value is higher than the false alarm level but significantly lower than the high-risk range. The calibration is performed by comparing the currently calculated immune dysregulation probability with 0.05. When the probability exceeds 0.05, a distribution boundary adjustment is triggered to integrate the influence of process variables. The distribution boundary adjustment is based on the variable influence distribution output in step 3. First, the direction of sensitive variables is determined. The process for determining the direction of sensitive variables is as follows: For each process variable, its value interval is divided into the aforementioned 5 segments, and the segment to which each conditional sample point belongs is identified; the number of cases marked as true in each segment is counted and the total number of samples in that segment is divided to obtain the segment confusion rate; the difference between the highest and lowest segment confusion rates is calculated as the monotonic correlation strength of the variable. The larger the difference, the more significant the directional influence of the variable on the confusion probability. The sign of the difference is used to determine the direction of risk increase and risk decrease. After the direction determination is completed, the variable influence interval is expanded. The expansion range is 0.1 times the original interval length. The basis for determining 0.1 is to provide sufficient boundary movement to reduce the confusion probability without exceeding the process achievable boundary, while avoiding excessive expansion in a single step that could lead to failure of subsequent deviation verification; during expansion, only the risk decrease direction is expanded, while the risk increase direction is simultaneously contracted. The contraction range is consistent with the expansion range, so that the center position of the variable interval shifts towards the low-risk direction while the interval length remains unchanged. The expanded and contracted boundary values are obtained by adding or subtracting one-tenth of the original boundary value from the interval length. The effectiveness of the boundary shift is verified by the process, which compares the shifted boundary with the 5% to 95% range of the equipment's stable control interval. If it exceeds this range, the boundary is truncated to the corresponding quantile boundary, which is derived from the statistical distribution of the control records. Based on the boundary shift amount and direction for each variable, a preliminary form of the precise adjustment vector is formed. Each component of the vector is determined by the difference between the new and old boundaries of that variable, with a directional indicator indicating an offset towards the low-probability direction.
[0058] After obtaining the preliminary form of the precise adjustment vector, intensity deviation calculation and immune response simulation verification attribute extraction are performed to determine whether the vector adjustment introduces instability in the response distribution. The intensity deviation calculation strictly follows the definition of "mean response data minus interaction intensity value," where the mean response data refers to the average interaction intensity of the aforementioned sites, and the interaction intensity value refers to the interaction intensity output of that site under a certain conditional sample point. The specific calculation process is as follows: for each site and each effective conditional sample point, the average interaction intensity of that site is taken, and then the interaction intensity output of that site under that conditional sample point is subtracted to obtain the deviation value of that site under that conditional sample point; the absolute value of the deviation value is taken to reflect the deviation magnitude; the absolute deviation values of all sites under the same conditional sample point are summed, and then divided by the number of sites to obtain the overall deviation level of that conditional sample point; the overall deviation levels of all effective conditional sample points are sorted, and the 90th percentile is taken as the immune response simulation verification attribute. This value is used to focus on the high deviation tail and avoid a small number of extreme values dominating the judgment. A preset threshold is used to determine whether a validation attribute exceeds the limit. The threshold is set to the 95th percentile of the validation attribute under the control condition. This threshold is determined based on keeping the false trigger probability below 0.05 under the control condition, while maintaining sensitivity to tail expansion introduced by the non-control condition. The threshold is obtained by calculating the validation attribute in the control condition sample point set using the same process and taking the 95th percentile value. If the current validation attribute is higher than the threshold, a vector distribution adjustment is triggered. The execution rules for the vector distribution adjustment are as follows: First, priority variables are determined based on the importance of the variables' influence on the distribution. The importance ranking comes from the statistical results of the tail contribution of the prior regression model to the intensity. The boundary shift of the priority variable is reduced by a reduction ratio of 0.5. The basis for 0.5 is to significantly compress the tail of the bias in a single adjustment without causing the boundary shift to fall back into the invalid interval. The reduction operation is performed by multiplying the boundary shift of the variable by 0.5 to obtain a new shift, while keeping the shift direction unchanged. After the reduction is completed, the joint condition sample point set is regenerated, and the calculation of the immune disorder response probability and validation attribute is repeated until the validation attribute falls within the threshold and the disorder response probability converges towards the standard value.
[0059] After the vector distribution is verified to be stable, the influence distribution prediction is extracted for the final inactivation range, and the peak point is extracted from the distribution to form an optimized version of the control vector for precise adjustment. The final inactivation range is determined by the joint interval of each process variable after boundary adjustment and vector reduction. Within this range, stratified coverage sampling is performed again, and immune response simulation is run to obtain an updated set of disordered labels and a sequence of disordered responses. The formation process of the influence distribution prediction is as follows: For each conditional sample point, its contribution to the disordered response probability is recorded. The contribution is defined as the disordered label value of the sample point, where a disordered label of true is recorded as 1, and a disordered label of false is recorded as 0; all sample points are grouped according to a certain statistical index, the statistical index being selected as "overall deviation level of conditional sample points", which can simultaneously reflect the site response deviation and risk triggering trend; the range of values of the overall deviation level is divided into 10 segments, the determination of 10 is based on forming sufficient resolution to identify risk-dense areas without introducing excessive segment noise; the number of disordered labels of true in each segment is counted and divided by the number of sample points in that segment to obtain the disorder rate within the segment, and the segment with the highest disorder rate within the segment is regarded as the peak segment. The extraction of peak points includes two parts: the probabilistic peak representative value and the conditional peak representative point. The probabilistic peak representative value is the disorder rate within the peak segment. The conditional peak representative point is obtained by backtracking the sample point set that falls into the peak segment and taking the median of the values in the set for each process variable. The median is selected based on the fact that the median is not sensitive to extreme values and can represent the central location of the risk-intensive area.
[0060] Furthermore, the construction of the optimized version of the precise adjustment vector is based on the principle of staying away from the peak representative point. The specific process is as follows: For each variable, calculate the direction of the difference between the median value of its value in the peak set and the center value of the current interval. If the median value of the peak is located at the high end of the current interval, the optimization direction is to reduce the variable; if the median value of the peak is located at the low end of the current interval, the optimization direction is to increase the variable. The shift amplitude is allocated according to the sensitivity of the variable. The sensitivity is determined by arranging the variable values in the peak set in ascending order, calculating the change amplitude within the segment marked by the difference between adjacent values, and the larger the change amplitude, the stronger the sensitivity. The sensitivity is normalized and used as the weight for amplitude allocation, so that the total shift amount is controlled and highly sensitive variables receive a larger adjustment amount.
[0061] Finally, based on the optimized version of the precise adjustment vector, the existing response data is aggregated and integrated to determine the final precise adjustment vector. The aggregation process adopts a vector direction consistency weighting mechanism. Specifically, for each existing conditional sample point, its relative position to the optimization direction is compared. If the sample point is located on the safe side indicated by the optimization direction in terms of most variables, it is assigned a higher weight; if it is located on the risk side, it is assigned a lower weight. The weight ratio is 3. This value is determined to significantly emphasize the contribution of samples on the safe side during aggregation, while retaining a small amount of information from samples on the risk side to maintain the model's ability to identify boundary risks. The weight allocation adopts a piecewise linear method, so that the weight increases monotonically with the distance from the sample point to the peak representative point. The distance is calculated by summing the absolute values of the normalized deviations of each variable. The normalized deviations are scaled uniformly based on the interval length of each variable to avoid a variable with a large dimension dominating the distance. After weighting, a weighted average interaction intensity is calculated for each locus. This is achieved by multiplying the interaction intensity output for that locus across all sample points by its corresponding weight, summing the results, and then dividing by the total weights to obtain the weighted average. Based on this weighted average, the disordered labeling and disordered response probabilities are calculated repeatedly, along with the validation attribute calculations. If the disordered response probability is not higher than 0.05 and the validation attribute is not higher than the 95th percentile threshold of the control, the vector is confirmed to meet the probability threshold calibration and bias stability constraints. The components, direction, sensitivity weights, and constraint satisfaction markers of the optimized vector are then combined to form the final precise adjustment vector.
[0062] S5 includes obtaining immune response data from virus inactivation conditions, introducing a gradient descent optimization algorithm for precise adjustment vector control, where the gradient descent optimization algorithm takes the vector as input and iteratively updates the model parameters by minimizing the loss function, iterating the virus characteristic model, where the virus characteristic model takes the immune response data as input and outputs optimized characteristic parameters to obtain a preliminary set of process parameters; aggregating response data through the preliminary process parameter set, calculating the average interaction intensity, where the average interaction intensity is obtained by summing the response data and dividing by the number of data points, determining the probability threshold calibration, where the probability threshold calibration compares the average value with a preset threshold to determine the distribution boundary adjustment; integrating the intensity deviation calculation based on the distribution boundary adjustment, where the intensity deviation calculation is obtained by subtracting the interaction intensity value from the average value to obtain the influence distribution prediction, optimizing the form of the precise adjustment vector control; extracting the inactivation temperature range for the optimized precise adjustment vector control, introducing a support vector machine algorithm to process the range data, where the support vector machine algorithm takes the range data as input and outputs the classification result by maximizing the classification boundary to obtain the temperature influence distribution; extracting peak points from the temperature influence distribution, integrating previous response data, and determining the final inactivation process parameter set.
[0063] In this embodiment, a unified data input is first obtained from the immune response data corresponding to the virus inactivation conditions. The source of the immune response data is limited to the simulation results of conditional sample points within the final inactivation range output by S3, and the same site indexing system and the same time window alignment rules are used to avoid parameter update direction drift caused by inconsistencies in the scope of different stages. Before being incorporated into the virus characteristic model, the immune response data underwent a standardization process, which included sample deduplication, missing data removal, and scaling. Sample deduplication was achieved by comparing the discrete identifiers of the process variable combinations. Only the record with higher quality was retained for duplicate identifiers. The quality judgment was based on the number of valid sampling points after passing the time window alignment judgment, with priority given to those with more valid sampling points. Missing data removal directly excluded samples that did not meet the steady-state window alignment, with an upper limit of 0.2. The basis for determining 0.2 is that when the exclusion ratio exceeds this value, the variance of subsequent average and gradient estimates increases significantly, weakening the optimization convergence stability. Scaling was based on the quantile span under the control condition, and the immune intensity-related output and the site interaction intensity-related output were normalized separately to ensure that different dimensional indicators participate in error measurement within the same numerical scale. The quantile span was the difference between the 25th and 75th percentiles, a value chosen because it is insensitive to extreme values and stable across multiple batches of data.
[0064] Subsequently, a gradient descent optimization algorithm was introduced to optimize the precise adjustment vector and iterated on the virus characteristic model to obtain a preliminary set of process parameters. The precise adjustment vector as optimization input is implemented by expanding each component of the vector into four categories of information corresponding to the process variables: "boundary movement direction, boundary movement amplitude, sensitivity weight, and constraint satisfaction mark." The sensitivity weight is used to construct a weighting rule for parameter updates, ensuring that highly sensitive variables receive a higher priority in adjustment under the same error improvement objective. The virus characteristic model takes immune response data as input and outputs optimized characteristic parameters. The characteristic parameters are defined as a set of numerical parameter vectors characterizing "site interaction intensity distribution pattern, immune disorder probability level, and deviation tail strength." This vector has a differentiable or numerically approximate response relationship with the process variables, satisfying the requirements of subsequent iterative updates. The objective of gradient descent optimization is to minimize the comprehensive error index, which consists of three parts: Part 1 is the excess between the immune disorder response probability and a preset standard value. This excess is summarized according to the rule of "including the portion above the standard value and excluding the portion below the standard value," thus focusing the optimization on risk suppression. Part 2 is the tail statistic of the intensity deviation, which is taken as 90% of the overall deviation level. The 90% threshold is chosen to focus on high-risk tails and avoid the dominance of a few extreme values. Part 3 is the boundary constraint penalty term, triggered when a variable goes out of bounds. The penalty strength is monotonically correlated with the magnitude of the out-of-bounds movement, ensuring that the update process does not push the parameters away from the process's achievable boundary. The weights of these three parts are determined through sensitivity analysis on the validation set. Specifically, the rate of decrease of the comprehensive error index and the number of constraint triggers are calculated alternately within the candidate weight set. The weight combination with "faster decrease rate and fewer triggers" is selected, and the weight of the immune disorder probability component is set to the highest. This highest weight is chosen because missed detections or amplified disorder risks bring unacceptable uncertainty to subsequent safety and effectiveness assessments.
[0065] The gradient is calculated using a numerical difference path to avoid dependence on specific analytical forms. The numerical difference process is as follows: for each parameter to be updated, a small perturbation is applied along its corresponding process variable direction, and then the virus characteristic model output and comprehensive error index calculation are re-executed; the change in the comprehensive error index before and after the perturbation is divided by the perturbation amplitude to obtain the approximate gradient value of the parameter. The perturbation amplitude is determined based on the joint constraints of parameter scaling and numerical stability. Initially, 0.01 of the current value of the parameter is used as the initial perturbation ratio. The determination of 0.01 is based on achieving a balance between "distinguishing the differential signal" and "not destroying the local linear approximation" in most numerical optimization tasks. Subsequently, the sign stability of the difference result is checked. If the sign flips during repeated calculations, the perturbation ratio is reduced to 0.005 until the sign stabilizes. The value of 0.005 is based on suppressing the differential noise to an acceptable level. The step size is determined using a backtracking search rule. An initial step size of 1 is set as a preset baseline value, determined by the median error reduction per unit step size in the previous iteration, ensuring the step size matches the current error terrain. If the overall error index does not decrease after the update, or the number of constraint penalty triggers increases, the step size is reduced by 0.5 and updated again. The 0.5 is determined as a trade-off between backtracking efficiency and fine-tuning the search. The stopping condition uses three parallel criteria: the overall error index decreases by less than 0.001 in three consecutive iterations; the parameter update magnitude is less than 0.001 of the current parameter magnitude in three consecutive iterations; and the number of constraint penalty triggers remains unchanged in three consecutive iterations. The 0.001 threshold is based on a slightly higher upper bound of error fluctuation noise on the validation set, used to avoid false stoppages caused by noise. After the stopping condition is met, the corresponding preliminary process parameter set is output. The preliminary process parameter set is characterized by "the center value of the candidate value interval of each process variable, the interval width, and the direction mark consistent with the control vector". The center value is the midpoint of the interval after the current optimization iteration, the interval width is the current boundary difference, and the direction mark is oriented towards the low-risk direction.
[0066] The response data is aggregated using a preliminary set of process parameters, and the average interaction intensity is calculated. Probability threshold calibration is then performed, and distribution boundary adjustments are determined. The aggregation of response data is achieved by selecting a set of sample points from the immune response data whose process variable combinations fall within the interval defined by the preliminary set of process parameters, and summing the interaction intensity output for each site. The average interaction intensity is obtained by summing and dividing by the number of data points according to the aforementioned rules. The summation object is the interaction intensity output of that site within the sample point set, and the number of data points is the number of valid records for that site within the sample point set. Records that fail to pass the steady-state window alignment determination are excluded from the valid record count. In this step, probability threshold calibration uses the risk representation of the average interaction intensity instead of the direct probability output for calibration. The calibration process involves comparing the average interaction intensity with a preset threshold; if it exceeds the threshold, it is considered a risk exceeding the limit, triggering distribution boundary adjustments. The preset threshold is determined using a high-quantile strategy based on the control baseline. Specifically, the average distribution of interaction intensity is calculated in the control condition sample set, and the 95th percentile is taken as the threshold. The 95th percentile is determined to keep the probability of false triggering in the control below 0.05 and to maintain sensitivity to upper tail expansion. If multiple batches of control data exist, the maximum value of the 95th percentile from each batch is taken as the final threshold to avoid low thresholds due to batch differences. The rules for adjusting the distribution boundary are as follows: identify the sensitive direction of the variable that causes the exceedance. The sensitive direction is determined by comparing the median difference of each variable between the "exceedance sample point set" and the "non-exceedance sample point set," with the variable having a larger median difference taking priority. For the priority variable, interval shrinkage is performed with a shrinkage ratio of 0.1. The 0.1 is determined based on the fact that a single adjustment has an observable effect on the boundary without causing the interval to collapse rapidly. For low-sensitivity variables, the interval remains unchanged to preserve the feasible region width. After shrinkage, a new distribution boundary adjustment result is formed, which serves as the boundary input for subsequent bias calculations and influence distribution prediction.
[0067] The intensity deviation is calculated and the impact distribution is predicted based on the distribution boundary adjustment, thereby optimizing the form of the precise adjustment vector for regulation. The intensity deviation is calculated by subtracting the interaction intensity value from the average value according to the aforementioned rules. The average value refers to the average of the aforementioned interaction intensity, and the interaction intensity value refers to the interaction intensity output of a single sample point. The specific calculation process is as follows: the deviation value is calculated for each site and each sample point, and the absolute value of the deviation value is taken to form the deviation amplitude; the deviation amplitudes of all sites under the same sample point are summed and divided by the number of sites to obtain the overall deviation level of the sample point; the overall deviation levels are sorted from smallest to largest, and the 90th percentile is extracted as the deviation tail index to characterize the risk tail. The impact distribution prediction uses "changes in the tail index of deviation caused by changes in process variable values" as the output. The process is as follows: Within the interval after adjusting the distribution boundary, multiple representative values are selected for each variable and combined with other variables at the center of the interval to generate a set of prediction points. The number of representative values is set to 5, determined to balance resolution and computational scale in univariate scanning. For each set of prediction points, a virus characteristic model is executed to infer and calculate the tail index of deviation, resulting in a mapping sequence from variable to tail index. The mapping sequence is sorted by the tail index of deviation from smallest to largest, and the neighborhood of the minimum value is used as the low-risk segment. Based on this low-risk segment, the precise adjustment vector is updated. The update rule is: the adjustment direction of each variable's components points to the center of the low-risk segment, and the component amplitude is allocated according to the variable's sensitivity. Sensitivity is determined by the range of changes in the tail index of deviation in the variable's scan sequence; a larger range indicates higher sensitivity and a larger amplitude share. Simultaneously, the constraint satisfaction flag is updated to "threshold calibration passed" or "threshold calibration failed," used for weight adjustment in subsequent iterations.
[0068] To optimize the control and precisely adjust the vector form, the inactivation temperature range is extracted, and a support vector machine (SVM) algorithm is introduced to process the range data to obtain the temperature influence distribution. The temperature range extraction is implemented as follows: the lower and upper bounds of the temperature range after boundary shift are read from the temperature-related components of the vector, and this range is used as the temperature candidate domain. Within the candidate domain, the temperature value sequence is generated by equal-interval segmentation, with 20 segmentation points. The number 20 is determined to provide sufficiently fine resolution in the one-dimensional temperature space to support classification boundary learning, while avoiding redundancy due to excessively dense sample points. The training labels for the SVM algorithm are derived from the risk assessment results. The label construction process is as follows: for each temperature value, combined with the condition that other variables are fixed at the center of the current range, the corresponding comprehensive error index or deviation tail index is calculated and compared with the aforementioned preset threshold. Values below the threshold are marked as safe, and values above the threshold are marked as risky. The threshold is the 95th percentile threshold used in the control group to maintain consistency. The core objective of Support Vector Machines (SVMs) is to maximize the classification boundary. Parameter determination employs hierarchical cross-validation. Cross-validation is implemented by sampling in segments according to the temperature sequence to form training and validation subsets. The number of segments is set to 5, and validation is performed alternately. The 5 is chosen to balance training stability and validation reliability in scenarios with limited sample sizes. The penalty strength parameter is selected with the "minimum false negative rate for risk classes" as the priority criterion. When false negative rates are the same, a parameter combination with a "larger boundary interval" is selected. The priority of false negative rate is based on minimizing the impact of missed risks on subsequent process determination. After training, the temperature influence distribution is obtained, which is presented as a mapping from temperature values to classification results. Simultaneously, the distance from each temperature point to the classification boundary is output as a confidence scale; a larger distance indicates a more robust classification.
[0069] The distribution of temperature influence is standardized, and the sorting direction and granularity of temperature values are unified. Multiple records corresponding to repeated temperature values in the distribution are merged according to the same temperature key. The merging rule adopts a robust statistical caliber. The median level, upper tail level and sample density are calculated for the risk characterization output under each temperature key. The sample density is represented by the number of valid records under that temperature key. The number of valid records is determined by the count of records that have passed the time window alignment and missing consistency verification to avoid false peaks introduced by missing data.
[0070] After data processing, peak point extraction is performed. The temperature axis is first divided into continuous segments. The number of segments is not set by a fixed constant but determined by the stability of peak positions. Multiple rounds of resampling are performed on the same dataset. In each round, a subset is proportionally extracted from the original records, and the segment risk density sequence is recalculated. The maximum segment position fluctuation of each round's risk density sequence is compared, and the smallest segment number that minimizes the maximum segment position fluctuation without multi-peak jumps is selected as the final number of segments. After segment division, a segment risk density index is calculated for each temperature segment. The segment risk density index consists of two parts: one is the proportion of high-risk markers within the segment, which are determined by the risk category output from the classification; the other is the risk intensity level within the segment, represented by the upper tail quantile value output from the risk characterization within the segment. The quantile of the upper tail quantile is taken from the corresponding quantile in the control distribution when the false alarm probability is controlled. The quantile is determined based on the upper tail stability bound of the control data across multiple batches. The proportion of high-risk individuals and the level of risk intensity are weighted and synthesized according to a risk priority principle. The weights are determined by a verification rule that prioritizes prioritizing missed detection constraints, ensuring that the contribution of the proportion of high-risk individuals to the segment risk density is higher than that of the level of risk intensity. The segment risk density sequence is smoothed, with the smoothing window length determined by the number of segments and the fluctuation amplitude. The shortest window length is selected to ensure that the peak segment boundary remains consistent across resampling rounds, avoiding peak shift caused by over-smoothing. After smoothing, the segment with the highest segment risk density is taken as the peak segment, and the median temperature value within the peak segment is defined as the representative value of the peak point. The median is used to reduce the pull of sparse samples at the segment endpoints on the representative value.
[0071] Subsequently, the previous response data is integrated. Using the peak segment as a filter, all records whose temperatures fall within the peak segment are extracted from the previous response data, forming a peak-related data subset. Simultaneously, records in the adjacent segments to the left and right of the peak segment are extracted, forming a peak neighborhood data subset. The number of neighborhood segments is set to 1, based on the fact that the marginal contribution of adjacent segments to peak gradient discrimination is the most stable across multiple rounds of resampling. The peak-related data subset is aggregated in two layers by site index and condition index. First, the interaction intensity-related output and the confusion label-related output are summarized at the site level, and then summarized at the condition level to form a conditional risk profile. The interaction intensity summary uses the mean within the steady-state window, and the denominator of the mean is the number of valid data points at that site under that condition. The number of valid data points is counted according to the number of data points that have passed the steady-state alignment judgment. The confusion label summary uses the trigger frequency, which is obtained by dividing the number of confusion labels as true by the number of valid records under that condition. The peak neighborhood data subset is aggregated again using the same caliber to obtain the neighborhood risk profile. Subsequently, the peak risk profile was compared with the neighborhood risk profile to form a peak gradient discrimination result. The comparison indicators were the difference between the trigger frequency of the peak segment and the trigger frequency of the neighborhood segment, and the difference between the upper tail level of the interaction intensity of the peak segment and the upper tail level of the interaction intensity of the neighborhood segment. The difference discrimination threshold was determined by the upper bound of the difference of the same indicator in the control data. The upper bound of the control was taken as the maximum upper quantile value of the difference of multiple batches of controls to avoid false triggering of peak judgment by control noise. This discrimination result is used to distinguish between peak-type high risk and plateau-type high risk, thereby providing a basis for the convergence direction of the parameter set.
[0072] After completing peak point location and response data integration, the final inactivation process parameter set is determined. First, the temperature parameter value strategy is determined: if the discrimination result shows a peak-type high risk, the center of the temperature candidate interval is shifted from the peak point representative value to the side where the risk gradient decreases faster. The shift direction is determined by the side with the larger decrease in trigger frequency among the neighboring segments to the left and right of the peak. The shift magnitude is determined by the constraint that "the effective sample density covered by the candidate interval after shift is not lower than the lower limit." The lower limit of the effective sample density is derived from the minimum sample size required for previous model training and validation, ensuring that subsequent evaluations still have statistical stability. If the discrimination result shows a plateau-type high risk, the entire temperature candidate interval is moved out of the plateau boundary. The plateau boundary is determined by the outer edge of the set of segments where the trigger frequency of continuous segments is higher than the upper limit of the control. After moving out, boundary truncation verification is performed on the candidate interval. The truncation boundary is determined by the process-realizable boundary, which originates from the statistical quantile range of the existing control stability interval.
[0073] After the candidate temperature intervals are determined, the remaining process variables are jointly defined, based on the principle of "shrinking the range of variable values in the same direction as the peak risk." Specifically, in the peak-related data subset, each process variable is segmented according to its current candidate interval, with the number of segments following the minimum stable segment number determined by the aforementioned segment stability criterion. The frequency of disturbance triggering and the upper tail level of interaction intensity for each segment are calculated to form a variable segment risk sequence. The variable segment risk sequence is then compared with the risk profile of the temperature peak interval. The consistency comparison uses a proportional change in the same direction, obtained by the percentage of segments in each segment whose relative median risk level rises or falls in the same direction as the temperature risk rises or falls. Variables with high consistency comparison results are identified as dominant linkage variables, and their candidate intervals are shrunken. The shrunkenness target is to retain continuous intervals of low-risk segments. Low-risk segments are defined by having a triggering frequency lower than a preset standard value and an upper tail level lower than the control upper bound. Both the preset standard value and the control upper bound follow the previous calibration logic to avoid caliber drift. The contracted joint interval forms a draft parameter set, which includes a combination of temperature candidate intervals and other variable candidate intervals, and is accompanied by an evidence label of associated risk. The evidence label consists of peak segment location, gradient discrimination results, and consistency comparison results of linked variables.
[0074] Finally, a consistency review was performed on the draft parameter set, and the final inactivation process parameter set was output. The review employed a dual-threshold constraint: the first threshold was the probability threshold for disordered reactions, which was set to a preset standard value and kept consistent with previous calibration; the second threshold was the deviation stability threshold, which was set to the upper bound of the control distribution and kept consistent with previous deviation thresholds. During the review process, the disorder trigger frequency and deviation tail statistics were recalculated within the subset of conditions covered by the draft parameter set. The denominator for the disorder trigger frequency was the number of valid records, and the deviation tail statistics used a high quantile and the upper bound of the control as the comparison benchmark. When both indicators simultaneously met the thresholds, the draft parameter set was confirmed and used as the final inactivation process parameter set; if either indicator did not meet the threshold, the corresponding variable interval was shrunk according to the priority order of the dominant linked variables, with the shrinkage range limited by the "sample density lower limit constraint," and the temperature candidate interval was kept from retreating to the peak range until the review passed.
[0075] Specifically, S6 includes obtaining simulated input data from the final inactivation process parameter set, introducing dose response simulation for vaccine process modeling, where dose response simulation maps the input data to the response curve to obtain a preliminary safety index set; aggregating immune response predictions through the preliminary safety index set, calculating the average effectiveness measure, where the average effectiveness measure is obtained by summing the response predictions and dividing by the quantity, determining threshold calibration, and determining distribution boundary adjustment; adjusting and integrating intensity deviation calculation based on the distribution boundary, where intensity deviation is obtained by subtracting the intensity value from the average value, extracting temperature range data to obtain peak point extraction results, where the peak point extraction results are obtained by selecting the maximum value point from the range data; optimizing the process iteration form based on the peak point extraction results, where the process iteration form is processed by updating parameter vectors, introducing a support vector machine algorithm to process the response data, where the support vector machine algorithm takes data as input, outputs the classification boundary by maximizing the interval, and obtains the optimized parameter set; extracting evaluation indicators from the optimized parameter set, integrating the previous simulation data, and obtaining a safety and effectiveness evaluation index set.
[0076] In this embodiment, firstly, the simulation input data required for vaccine process modeling is compiled from the final inactivation process parameter set. The final inactivation process parameter set consists of the combined value ranges of multiple process variables. To ensure the representativeness and stability of the modeling input, a representative value is determined for each process variable within its corresponding range. The representative value is determined using the median principle, that is, the midpoint of the variable range is taken as the main input value. This method is based on the fact that the median level is not sensitive to extreme values in multiple batches of simulations and can reflect the typical state within the range. Simultaneously, the upper and lower bounds of the range are recorded as a perturbation reference for subsequent verification of the completeness of the response curve. After completing the variable-level compilation, the representative values of each variable are combined to form a single input vector, which is then used as the basic input for vaccine process modeling.
[0077] After obtaining the simulated input data, dose response simulation is introduced to construct a vaccine process model. The implementation logic of dose response simulation is as follows: the simulated input data is mapped to the immune response output space, establishing a continuous correspondence between input intensity and immune response level. This correspondence is determined by jointly fitting existing simulated samples with simulation results under the current parameter set. During joint fitting, the monotonic trend and smoothness within the interval of response change are maintained as constraints, ensuring that the response curve does not have non-physical jumps within the input interval. Subsequently, several representative nodes are selected on the input intensity axis, with the number of nodes set to 7. The determination of 7 is based on controlling computational complexity while covering low, medium, and high intensity segments, so that the overall shape of the response curve is fully characterized without introducing over-dense sampling. For each node, the corresponding safety-related output index is calculated, and the safety-related outputs of all nodes are summarized to form a preliminary safety index set. The preliminary safety index set includes at least the abnormal response trigger frequency, the width of the response stability interval, and the response fluctuation amplitude index. Each index is based on statistical results under the same time alignment window.
[0078] After obtaining the preliminary safety indicator set, the aggregation and efficacy measurement calculation stage for immune response prediction begins. First, the preliminary safety indicator set is categorized and summarized by indicator type, and the prediction results for the same indicator at all dose nodes are compiled. Then, the average efficacy measurement is calculated. The average efficacy measurement is calculated by summing the predicted values of a specific efficacy-related indicator at all nodes, then dividing by the number of nodes (the previously determined 7 nodes). This average reflects the overall immune efficacy performance across the entire dose range. After calculating the average, threshold calibration is performed. The average efficacy measurement is compared with a preset standard value to determine whether distribution boundary adjustment is triggered. The preset standard value is determined as follows: under controlled simulation conditions, the average distribution of similar efficacy indicators is collected, and a stable high percentile is selected from multiple batches of control results as a unified standard value. This high percentile is set at 95%, and its selection is based on controlling the false positive probability under control conditions to within 0.05, while maintaining sensitivity to insufficient efficacy. When the average efficacy measurement is lower than this standard value, the current distribution is deemed to have efficacy risk, and distribution boundary adjustment is initiated.
[0079] During the distribution boundary adjustment process, intensity deviation calculation is introduced to locate critical sections with significant impact. The intensity deviation calculation follows the definition of "average value minus single-point intensity value," specifically implemented as follows: for each dose node, the corresponding effectiveness prediction value is taken, and the difference between this value and the aforementioned effectiveness metric average is calculated to obtain the deviation value. The absolute magnitude of the deviation value is then used to characterize the degree of deviation. Subsequently, the deviation magnitudes of all nodes are sorted and statistically analyzed to form a deviation distribution. Based on this deviation distribution, the process variable values corresponding to nodes with larger deviation magnitudes are traced back, with a focus on extracting the corresponding interval data for temperature variables. The temperature interval data consists of the temperature values corresponding to nodes with deviation magnitudes located in the upper quantile interval. The upper quantile interval is the 90th percentile, which is based on focusing on the few sections that have the most significant impact on overall effectiveness. The extracted temperature interval data is sorted, and the value with the largest deviation magnitude is selected as the peak point extraction result. This peak point is used to identify the temperature position that has the most prominent impact on safety or effectiveness under the current parameter combination.
[0080] After obtaining the peak point extraction results, the process iteration method is optimized. Specifically, the peak point is first used as a risk reference to update the current parameter vector. The update direction follows the principle of moving away from the high-deviation segment corresponding to the peak point, causing the parameter vector to move towards a low-deviation, stable segment. The update magnitude is constrained by the remaining width of the parameter interval and the model's convergence stability to avoid excessive one-time movement that could lead to evaluation distortion. Subsequently, a support vector machine (SVM) algorithm is introduced to reprocess the response data, which includes safety and effectiveness index outputs obtained under different parameter vectors. The SVM algorithm uses the response data as input and constructs a classification boundary by maximizing the classification margin between samples of different risk levels. The classification label is jointly defined by "whether effectiveness threshold calibration is triggered" and "whether it falls within the upper quantile interval of deviation," allowing the classification boundary to directly reflect the comprehensive judgment result of safety and effectiveness. During model training, the penalty weight parameter is determined through hierarchical cross-validation. The validation objective prioritizes constraining the omission rate of high-risk samples, and then, under the premise of satisfying this constraint, selects the parameter combination with the lowest overall misclassification rate. Based on the classification boundary after training, the parameter vectors are filtered and corrected, retaining the parameter combinations that are located on the low-risk side and maintain a sufficient distance from the classification boundary, forming the optimized parameter set.
[0081] Finally, evaluation indicators were extracted from the optimized parameter set and integrated with previous simulation data to form the final safety and effectiveness evaluation indicator set. The extraction of evaluation indicators followed a unified standard, summarizing the safety and effectiveness indicators recalculated under the optimized parameter set and comparing them with the corresponding indicators in the previous simulation data. The consistency comparison was based on whether the indicator changes fell within the natural fluctuation range under control conditions, which was determined by the quantile span of the control simulation results. Abnormal indicators that did not meet the consistency requirements were removed, based on their unstable shifts in multiple simulations. The safety and effectiveness evaluation indicator set formed after the above integration can simultaneously reflect the comprehensive performance of the final inactivation process parameter set in terms of safety boundaries and immunogenicity.
[0082] Specifically, S7 includes: acquiring data that meets preset standards from a set of safety and effectiveness assessment indicators; extracting key control point data from the prediction results for indicators that meet the standards; matching the key control point data with virus inactivation control conditions using a data mapping method to obtain preliminary process control boundary delineation results; based on the preliminary process control boundary delineation results, acquiring environmental constraint data in the inactivation condition setting to meet the needs of process parameter adjustment; determining the parameter range range in the process control scheme by integrating the content of safety indicator verification; acquiring the boundary conditions required for control process optimization for the parameter range range; combining the parameter range range with the inactivation environmental constraint setting using a pre-established mapping table to determine the feasibility of the control scheme and obtain an optimized process control scheme framework; and, based on the optimized process control scheme framework, integrating the results of effectiveness data acquired from the set of safety and effectiveness assessment indicators to meet the actual needs of virus inactivation control, outputting a complete process control scheme description and determining the final control process used for respiratory virus inactivation.
[0083] In this implementation, firstly, a subset of data that meets preset standards is selected from the set of safety and effectiveness assessment indicators. The set of safety and effectiveness assessment indicators consists of multiple indicators, each with a clearly defined statistical scope and evaluation direction. The preset standards are not fixed constants but are derived based on control conditions and historical stable operating intervals. Specifically, for each type of safety indicator, the indicator distribution under control conditions and multiple rounds of stable simulations is collected, the upper quantile boundary of this distribution is calculated, and this upper quantile boundary is used as the lower limit for safety judgment. The principle for its value is to cover natural fluctuations under control conditions and exclude abnormal tails. For each type of effectiveness indicator, the data distribution from the same source is collected, the quantile intervals that are above the median level and have long-term stability are calculated, and the lower limit of this interval is used as the effectiveness judgment standard, so that the screening results simultaneously meet the dual requirements of "no abnormal risk" and "achieving the target effect." After the standards are determined, the assessment indicator set is compared item by item, and only data records that simultaneously meet all corresponding standards are retained, forming a subset of data that meets the standards.
[0084] After obtaining a subset of data that meets the criteria, a contribution analysis is performed on the prediction results to extract key control point (CCP) data. CCPs are not directly determined by a single variable, but are obtained by ranking the degree of influence of changes in each control variable on the indicator changes in the prediction results. Specifically, the control variables in the criterion-compliant subset are segmented according to their value ranges. The statistical change magnitude of the corresponding evaluation indicator is calculated for each segment. The change magnitudes between different segments are then compared; variables with larger change magnitudes are considered more sensitive to the results. To avoid random fluctuations interfering with the ranking results, the calculation of the change magnitude is based on stable statistical results after multiple resampling, with consistency in the direction of change as an additional constraint. Only when a variable shows a consistent influence trend in most resamplings is it included in the candidate set of CCPs. Finally, the top-ranked control variables and their corresponding value ranges are selected from the candidate set to form the CCP dataset.
[0085] Subsequently, a data mapping method was used to match the critical control point data with the virus inactivation control conditions, resulting in preliminary process control boundary delineation. The mapping relationship was constructed based on the correspondence between the value ranges of control variables and the states of evaluation indicators. This was achieved by: for each critical control point variable, reading its value distribution within a compliant subset of data, selecting the concentrated segment of this distribution as the safe and effective segment, and mapping this segment as the preliminary control boundary for that variable. The determination of the concentrated segment adopted robust statistical principles, excluding small proportions of extreme values at both ends of the distribution and retaining the continuous intermediate range as the boundary range. This ensured that the control boundary covered the main safe and effective areas while avoiding interference from unstable points at the edges. The combination of boundary segments from multiple critical control point variables formed the preliminary process control boundary delineation.
[0086] Based on this, and addressing the actual needs of process parameter adjustment, environmental constraint data from the inactivation condition setting is introduced to further converge the initial boundary. The environmental constraint data originates from statistical results of existing control systems and operational experience, including the stable operating range of the equipment, the upper and lower limits of the process flow, and the range of external condition constraints. These environmental constraints are not expressed as single values but as stable ranges formed by long-term operational data, determined by selecting the middle high-density segment from the historical data distribution. The initial process control boundary is superimposed on the environmental constraint range; portions exceeding the environmental constraints are eliminated, and the remaining segments are verified for consistency with safety indicator validation results, ensuring that safety indicators remain within the standard range under environmental constraints. After the above processing, the final feasible range range for each parameter in the process control scheme is obtained.
[0087] For the aforementioned parameter range, the boundary conditions required for control process optimization are further obtained to complete the feasibility assessment at the scheme level. The determination of boundary conditions is based on the tolerance of key process nodes to parameter changes. The process involves: retrospectively analyzing existing process data and statistically analyzing the response stability of different nodes under parameter changes; identifying nodes with high sensitivity to parameter changes; and using the parameter value ranges corresponding to these nodes as process-level boundary constraints. Subsequently, a pre-established mapping table is invoked to combine and match the parameter range with inactivation environment constraints. This mapping table, based on historical verification results, records the adaptation status of different parameter ranges under different environmental constraint combinations, with its judgment logic centered on "whether there are known conflicts or unsustainable states." Each parameter combination is verified item by item through this mapping table, selecting combinations that have no conflicts in terms of process nodes, environmental constraints, and safety verification, thus forming an optimized process control scheme framework.
[0088] Finally, through an optimized process control framework, the integration and output of solutions tailored to actual inactivation control needs are completed. Specifically, the effectiveness data results corresponding to the safety and effectiveness assessment indicators are matched one-to-one with the parameter combinations in the solution framework. This verifies whether the effectiveness indicators remain within the target range at each control stage, and checks the continuity of parameter connections between different control stages to avoid parameter gaps or boundary conflicts during stage transitions. After the above integration, the control scheme is organized using a structured process description, clarifying the control objectives, parameter ranges, judgment conditions, and adjustment directions for each stage, ultimately outputting a complete description of the process control scheme.
[0089] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A data-driven modeling and predictive control method for respiratory virus inactivation, characterized in that, Includes the following steps: S1. Initial viral characteristic data were obtained from the viral surface structure database through molecular dynamics simulation, and the structural change matrix of the virus under different inactivation conditions was obtained after processing. S2. Based on the structural change matrix, a support vector machine classifier is used to extract features of viral characteristics and antibody binding sites to determine the set of potential unexpected interaction sites. S3. If the set of potential unexpected interaction sites exceeds the preset threshold, the influence distribution of inactivation process variables on interaction intensity is predicted by random forest regression model to obtain the range of optimized inactivation conditions. S4. Obtain simulated immune response data from the optimized inactivation condition range, determine the probability of immune disorder response, and obtain the precise adjustment vector for regulation; S5. For the precise adjustment vector of regulation, the gradient descent optimization algorithm is used to iterate the virus characteristic model to determine the final inactivation process parameter set; S6. Simulate the vaccine development process using the final inactivation process parameter set to obtain a set of safety and efficacy evaluation indicators; S7. If the set of safety and effectiveness assessment indicators meets the preset standards, output a complete description of the control scheme.
2. The respiratory virus inactivation data-driven modeling and predictive control method according to claim 1, characterized in that: S1 includes: Initial viral characteristic data were obtained from a viral surface structure database through molecular dynamics simulations, and a preliminary structural model was obtained after processing. Based on the preliminary structural model, inactivation condition variables including temperature and pH were introduced to simulate protein folding changes under different inactivating agent concentrations, resulting in a conditional influence dataset. Based on the conditional influence dataset, the row vectors of the structural change matrix represent the inactivation condition variables, and the column vectors represent the viral surface structural sites. The matrix element values are determined to be the site displacement differences, which are obtained by comparing the Euclidean distance between the coordinates of the viral surface structural sites before and after inactivation. If the matrix element values exceed the preset threshold, the simulation parameters are adjusted to optimize the simulation of protein folding changes, and a refined structure change matrix is obtained. By extracting trend tracking features from the refined structural change matrix and combining them with control data obtained from the conditional influence dataset to confirm the accuracy of the matrix, the structural change matrix of the virus under different inactivation conditions was obtained.
3. The respiratory virus inactivation data-driven modeling and predictive control method according to claim 1, characterized in that: S2 includes: Protein folding change data corresponding to inactivation condition variables were obtained from the structural change matrix. A support vector machine classifier was used to preliminarily classify the virus characteristics and obtain a set of feature vectors. For the feature vector set, the inactivator type variable and the site shift difference are introduced to extract antibody binding site related attributes and determine a preliminary list of interaction sites. Based on the preliminary list of interactive sites, change trend tracking information is obtained from the conditional influence dataset. If the difference in site displacement exceeds a preset threshold, it is marked as a potential unexpected site, thus obtaining a screening site group. By screening site groups, integrating viral surface structure data, and optimizing the support vector machine classifier, the support vector machine classifier takes the site displacement difference as input, constructs the decision boundary by maximizing the classification margin, outputs the site classification result, constructs the site stability evaluation index based on the site classification result, and obtains a refined interactive site matrix. Unexpected interaction attributes are aggregated from the refined interaction site matrix, and verified by tracking changes in trends to determine the set of potential unexpected interaction sites.
4. The respiratory virus inactivation data-driven modeling and predictive control method according to claim 1, characterized in that: S3 includes: Data on inactivation process variables are obtained from the set of potential unintended interaction sites. If the number of sites exceeds a preset threshold, the influence of the inactivation process variables is processed by a random forest regression model with the data on inactivation process variables as input, and the distribution of the influence of interaction intensity is obtained. To address the impact of interaction intensity on distribution, a conditional range optimization attribute is introduced as the basis for adjusting the distribution boundary. The threshold of the site set is used to integrate the distribution data to determine the influence of variables on the prediction results. Based on the prediction results of the variable impact, and combined with the distribution attributes of the regression model as the basis for the intensity deviation calculation, the intensity distribution information is extracted. If the distribution deviation exceeds the range, the distribution of process variables is adjusted to obtain the preliminary optimization conditions. By initially optimizing the conditions, we aggregate unexpected interaction sites and preset threshold judgment attributes as verification inputs to verify the range of optimized conditions and obtain the final inactivation condition range.
5. The respiratory virus inactivation data-driven modeling and predictive control method according to claim 1, characterized in that: S4 includes: Simulated immune response data are obtained from the optimized inactivation conditions range. The site interaction intensity is processed by aggregation of the response data. The aggregation process includes classifying and summarizing the response data by site, calculating the average interaction intensity, and determining the probability of immune disorder. To address the probability of immune disorder, a probability threshold calibration is introduced as the basis for adjustment. The threshold calibration compares the probability with a preset standard value and integrates the influence of process variables by adjusting the distribution boundary. The distribution boundary adjustment includes expanding the influence range of variables and determining the preliminary form of the precise adjustment vector. Based on the preliminary form of the vector for precise adjustment of regulation, the simulated response data and intensity deviation calculation are aggregated. The intensity deviation calculation is obtained by subtracting the interaction intensity value from the average response data to obtain the immune response simulation verification attribute. If the verification attribute exceeds the preset threshold, the vector distribution is adjusted. By adjusting the vector distribution, the influence distribution prediction is extracted for the final inactivation range. The influence distribution prediction includes extracting peak points from the distribution to obtain an optimized version of the regulation and adjustment vector. The response data is aggregated from the optimized version of the control precision adjustment vector. The response data aggregation integrates previous response data by optimizing the version vector to determine the final control precision adjustment vector.
6. The respiratory virus inactivation data-driven modeling and predictive control method according to claim 1, characterized in that: S5 includes: Immune response data is obtained from virus inactivation conditions. A gradient descent optimization algorithm is introduced to precisely adjust the vector for regulation. The gradient descent optimization algorithm takes the vector as input and iteratively updates the model parameters by minimizing the loss function. The virus characteristic model takes the immune response data as input and outputs the optimized characteristic parameters to obtain the preliminary process parameter set. By aggregating response data using preliminary process parameter sets, the average interaction intensity is calculated. The average interaction intensity is obtained by summing the response data and dividing by the number of data points. The probability threshold calibration status is then assessed by comparing the average value with a preset threshold to determine the distribution boundary adjustment. Intensity deviation is calculated based on the distribution boundary adjustment. The intensity deviation is obtained by subtracting the interaction intensity value from the average value to obtain the influence distribution prediction results and optimize the form of the precise adjustment vector for regulation. To optimize and precisely adjust the vector form of inactivation temperature range, a support vector machine algorithm is introduced to process the range data. The support vector machine algorithm takes the range data as input and outputs the classification result by maximizing the classification boundary to obtain the temperature influence distribution. Peak points were extracted from the temperature effect distribution, and previous response data were integrated to determine the final inactivation process parameter set.
7. The respiratory virus inactivation data-driven modeling and predictive control method according to claim 1, characterized in that: S6 includes: Simulation input data is obtained from the final inactivation process parameter set, and dose response simulation is introduced for vaccine process modeling. The dose response simulation maps the input data to the response curve to obtain a preliminary safety index set. By using a preliminary set of safety indicators to predict immune response, the average effectiveness metric is calculated. The average effectiveness metric is obtained by summing the response predictions and dividing by the number of response prediction sample points. Threshold calibration is then performed to determine the distribution boundary adjustment.
8. The respiratory virus inactivation data-driven modeling and predictive control method according to claim 7, characterized in that: S6 further includes: Intensity deviation is calculated based on the distribution boundary adjustment. The intensity deviation is obtained by subtracting the intensity value from the average value of the effectiveness index at all dose response nodes. Temperature interval data is extracted to obtain peak point extraction results, where the peak point extraction results are obtained by selecting the maximum value point from the interval data. The process iteration form is optimized based on the peak point extraction results. The process iteration form is processed by updating the parameter vector and introducing the support vector machine algorithm to process the response data. The support vector machine algorithm takes the response data as input, and the response data includes the safety index output and effectiveness index output obtained under different parameter vectors. The optimized parameter set is obtained by maximizing the interval to output the classification boundary; Evaluation indicators are extracted from the optimized parameter set and integrated with previous simulation data to obtain a set of safe and effective evaluation indicators.
9. The respiratory virus inactivation data-driven modeling and predictive control method according to claim 1, characterized in that: S7 includes: Data that meets preset standards is obtained from the set of safety and effectiveness assessment indicators. For indicators that meet the standards, key control point data in the prediction results are extracted. Using data mapping relationship, the key control point data is matched with the conditions for virus inactivation control to obtain the preliminary process control boundary division results. Based on the preliminary process control boundary delineation results, and in response to the need for process parameter adjustment, environmental constraint data in the inactivation condition setting are obtained. By integrating the content of safety indicator verification, the parameter range in the process control scheme is determined.
10. The respiratory virus inactivation data-driven modeling and predictive control method according to claim 9, characterized in that: The S7 also includes: For the parameter range, the boundary conditions required for control process optimization are obtained. A pre-established mapping table is used to combine the parameter range with the inactivation environment constraints to determine the feasibility of the control scheme and obtain the optimized process control scheme framework. By optimizing the process control framework and integrating the results of effectiveness data obtained from the set of safety and effectiveness assessment indicators to meet the actual needs of virus inactivation control, a complete description of the process control scheme is output, and the final control process used for respiratory virus inactivation is determined.
Citation Information
Patent Citations
Method for detecting virus inactivation effect by ELISA (enzyme-linked immuno sorbent assay) method
CN120577532A
Vaccine target screening system based on calculation model simulation
CN121483364A