Medical peptide quantitative addition parameter optimization method and system based on machine learning
By constructing a combination of chemical rules knowledge base and data-driven model, identifying and correcting parameter combinations in medical peptide synthesis, the problem of inefficient parameter optimization in traditional methods is solved, and the reliability and continuous improvement of parameter optimization results are achieved.
Patent Information
- Application Number
- CN202510733686.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
In the existing medical peptide synthesis process, traditional parameter optimization methods rely on experience or a large number of experiments, resulting in long R&D cycles, large material consumption, and difficulty in systematically revealing the deep correlation between parameters and synthesis results, and being unable to quickly respond to the development needs of complex structural peptides. At the same time, the data analysis system lacks a mechanism to integrate chemical rules, which leads to model recommendations contrary to the experience of synthetic chemists.
Build a knowledge base of chemical rules, combine data-driven models, quantitatively add parameter combinations through checksum correction, identify chemical rules inconsistencies or potential side reactions, generate warnings, and update the knowledge base or model parameters based on feedback, integrate chemical rules for parameter optimization.
Improve the reliability and efficiency of parameter optimization results, support iterative improvement of system knowledge and models, ensure that parameter combinations comply with chemical rules, avoid violations of basic principles, and continuously improve optimization accuracy.
Smart Images

Figure CN120260731A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a parameter optimization decision-making technology in the field of pharmaceutical peptide synthesis, and particularly relates to a method and system for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning. Background Art
[0002] In the research and development and production process of pharmaceutical peptides, solid-phase synthesis or liquid-phase synthesis is the core process for constructing the target peptide chain. These synthesis methods highly rely on the precise and sequential addition of various materials, including different types of amino acid derivatives, activators, coupling agents, protecting group removal reagents, and various reaction solvents. The quantitative addition parameters of each material, such as the molar ratio between each component, the concentration of the solution, the added volume, the rate of material addition, and the precise time node, together with process conditions such as the temperature of the reaction system and the reaction duration, jointly determine the efficiency of the peptide chain elongation reaction, the accuracy of the target sequence construction, the purity level of the final product, and the types and relative contents of by-products.
[0003] Traditional parameter optimization approaches mainly rely on the experience accumulation of R & D personnel or explore by performing a large number of orthogonal experiments and single-factor variable screening experiments. This method may be applicable when dealing with peptides with simple structures and mature synthesis routes, but when the target is a new type of pharmaceutical peptide with a high degree of structural complexity and a wide variety, it shows a long R & D cycle, a huge consumption of raw materials, and it is difficult to systematically reveal the deep relationship between parameters and synthesis results, resulting in low overall efficiency of the optimization work and inability to quickly respond to the development needs of new or complex-structured peptides.
[0004] To improve the optimization efficiency, pharmaceutical R & D institutions have introduced a data analysis system trained based on historical synthesis data to assist in predicting and optimizing the quantitative addition parameters in pharmaceutical peptide synthesis. This system constructs a prediction model by analyzing the polypeptide structure information, detailed process parameters, and corresponding synthesis yield, purity and other result data recorded in the database. However, the quantitative addition parameter combinations recommended by the data analysis model trained solely based on the statistical laws of historical data sometimes deviate significantly from the empirical judgments formed by senior synthetic chemists based on their understanding of the chemical behaviors of these special structures, or from the specific synthesis strategies reported in the literature for such special structural units.
[0005] Researchers have attempted to apply these known chemical rules related to specific peptide structures, or the relationships between structure and reactivity. However, existing data analysis systems are often designed without an effective mechanism to smoothly integrate this type of chemically symbolic knowledge. When the quantitative addition parameter combinations recommended by the model conflict with the professional intuition of synthetic chemists or their long-accumulated experimental experience, since the internal decision-making processes of most current data analysis models are not transparent enough to users, it is difficult for researchers to clearly and intuitively understand which input features the model is based on and how these features interact, ultimately resulting in the system recommending this specific set of parameters that may even go against conventional knowledge. Summary of the Invention
[0006] The object of the present invention is to provide a method and system for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning, which has the functions of combining chemical rule knowledge and data-driven models for parameter verification, identifying situations where parameter combinations do not conform to chemical rules or cause side reactions, correcting or generating warnings for parameter combinations, updating the chemical rule knowledge base or adjusting the preset parameters of the data-driven model based on the decision results, improving the reliability of parameter optimization results, and supporting iterative improvement of system knowledge and models.
[0007] To achieve the above object, the technical solution of the present invention is as follows: On the one hand, the present application provides a method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning, including the following steps: S1. Obtain raw data related to the field of pharmaceutical peptide synthesis from multiple specific channels, and use a data fusion algorithm to fuse the obtained raw data to obtain a chemical rule knowledge base; S2. Obtain the structural characteristics of the target peptide, and based on the structural characteristics of the target peptide, combine the parameters input by the data-driven model to obtain a quantitative addition parameter combination. According to the chemical rule knowledge base, verify the obtained quantitative addition parameter combination to identify the verification results that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions, and correct or generate a warning for the quantitative addition parameter combination in combination with the verification results to obtain a corrected or warning-attached quantitative addition parameter combination; S3. Generate a decision result based on the corrected or warning-attached quantitative addition parameter combination; S4. Obtain feedback information based on the decision result and, based on the obtained feedback information, update the chemical rule knowledge base or adjust the preset parameters in the data-driven model, where the feedback information includes parameter adjustments suggested by the user or the chemical principles supporting their suggestions.
[0008] Further, the step S1 includes the following steps: S11. Obtain chemical application knowledge information related to amino acid sequences, protecting groups, and coupling reagents in the field of pharmaceutical peptide synthesis; S12. Based on the chemical application knowledge information, formulate rules according to peptide structure characteristics, process conditions, and potential risks or recommended operations corresponding to peptide structure characteristics and process conditions, and store them in a structured manner according to the data set. Combine the obtained data sets to obtain a chemical rule knowledge base.
[0009] Further, the step S2 specifically includes the following steps: S21. Obtain a quantitative addition parameter combination combined with the structure characteristics of the target peptide and the parameters output by the data-driven model; S22. According to the chemical rule knowledge base, conduct a primary verification on the obtained quantitative addition parameter combination to identify a primary verification result that does not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions; S23. Based on the primary verification result, conduct a primary correction on the current quantitative addition parameter combination or generate a warning to obtain a corrected or warning-attached quantitative addition parameter combination obtained after the primary verification; S24. Perform iterative verification processing on the target peptide information set with the corrected or warning-attached quantitative addition parameter combination obtained after the primary verification, verify the current quantitative addition parameter combination to obtain a final verification result associated with the current quantitative addition parameter combination, and obtain a corrected or warning-attached quantitative addition parameter combination based on the final verification result.
[0010] Further, the step S22 specifically includes: S221. Obtain one or more parameter subsets in which there are synergistic effects or antagonistic effects among the member parameters in the current quantitative addition parameter combination, and for each identified parameter subset, retrieve the parameter interaction rules corresponding to the parameter subset from the chemical rule knowledge base or the parameter interaction knowledge database between parameters. The parameter interaction rules indicate the influence on the process and efficiency of the target chemical reaction or the occurrence probability and degree of side reactions when the member parameters within the parameter subset act together; S222. Based on each retrieved parameter interaction rule, execute an adjustment on the compliance verification conclusion of a single parameter or the entire parameter combination within the parameter subset in the single rule verification result or evaluate the risk level or occurrence probability of side reactions caused by the joint action of the member parameters within the parameter subset to obtain parameter interaction adjustment information; S223. According to the chemical rule knowledge base, conduct a single rule verification on the obtained quantitative addition parameter combination to identify parameters that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions, and obtain a single rule verification result; S224. Combine the single rule verification result and the parameter interaction adjustment information to obtain the initial verification result.
[0011] Further, the step S223 specifically includes: S2231. Retrieve chemical rules related to the current quantitative addition parameter combination and the structural characteristics of the target peptide from the chemical rule knowledge base, and identify the rules with their own confidence scores or probability descriptions of their applicable conditions attached in the relevant chemical rules, so as to obtain the chemical rules with uncertainty information attached. S2232. For each obtained chemical rule with uncertainty information attached, calculate the compliance probability of the current quantitative addition parameter combination for the chemical rule with uncertainty information attached based on the confidence score of the chemical rule itself or the probability description of its applicable conditions, and the matching situation between the quantitative addition parameter combination and the preconditions of the chemical rule. S2233. For one or more preset side reactions, based on one or more chemical rules related to the side reaction, combine the compliance probability of the current quantitative addition parameter combination for the chemical rule with uncertainty information attached, and when there are multiple relevant chemical rules contributing to the risk of the same side reaction, use the probability synthesis method to calculate the comprehensive risk probability of the current quantitative addition parameter combination triggering the side reaction. S2234. Use the compliance probability of the current quantitative addition parameter combination for each chemical rule calculated and the comprehensive risk probability of the current quantitative addition parameter combination triggering each preset side reaction calculated to identify the parameters that do not conform to the chemical rules in the chemical rule knowledge base or may trigger side reactions, and obtain the single rule verification result.
[0012] Further, the step S2233 specifically includes: S22331. For multiple chemical rules related to the side reaction, retrieve and obtain the dependency relationship descriptions between these chemical rules from the rule dependency data source, and the dependency relationship descriptions indicate the interaction modes of these chemical rules when jointly affecting the occurrence probability of the side reaction. S22332. Before performing probability synthesis operations, according to the obtained dependency description, correct the individual risk contribution of at least one of the multiple chemical rules to the side reaction generated by the rule to obtain the corrected individual risk contribution, and / or construct a probability synthesis structure that integrates the dependency description, and obtain the individual risk contributions of the multiple chemical rules related to the side reaction generated by each rule as the input of the probability synthesis structure, where the individual risk contribution indicates the compliance probability of the current quantitative addition parameter combination calculated for the corresponding rule for the chemical rule, or the probability information directly included in the corresponding rule for describing the probability of triggering the side reaction; S22333. Perform probability synthesis operations using the obtained corrected individual risk contribution, and / or perform probability synthesis operations using the probability synthesis structure constructed in the previous step and the obtained individual risk contribution as its input, so as to calculate, through the probability synthesis operation, the comprehensive risk probability of triggering the side reaction by combining the dependency relationship between chemical rules and the current quantitative addition parameter combination.
[0013] Further, step S24 specifically includes: S241. After obtaining the corrected or warning-attached quantitative addition parameter combination after the initial verification, start the iterative verification process. In each iteration of the iterative verification process, re-verify the quantitative addition parameter combination of the current iteration round according to the chemical rule knowledge base; S242. Determine whether the iterative verification process meets the preset iterative termination condition. The iterative termination condition indicates that the updated quantitative addition parameter combination reaches a stable state without further correction after a complete re-verification, or the number of iterations of the iterative process reaches the preset upper limit. If the iterative termination condition is met, use the updated quantitative addition parameter combination as the corrected or warning-attached quantitative addition parameter combination. If the preset iterative termination condition is not met, use the updated quantitative addition parameter combination as the input quantitative addition parameter combination for the next iteration round, and return to perform the re-verification in each iteration of the iterative process to finally obtain the re-verification result; S242. Based on the re-verification result, perform re-correction on the quantitative addition parameter combination of the current iteration round to obtain the final corrected or warning-attached quantitative addition parameter combination.
[0014] Further, step S3 specifically includes: S31. Obtain the relevance result between the corrected or warning - attached quantitative addition parameter combination and the chemical rule knowledge base; S32. Evaluate the impact of the corrected or warning - attached quantitative addition parameter combination on at least one preset optimization goal based on the relevance result, and obtain a decision result.
[0015] Further, the step S32 specifically includes: S321. Obtain the selected corrected or warning - attached quantitative addition parameter combination and the user's preference information for multiple conflicting pharmaceutical peptide optimization goals; S322. For each pharmaceutical peptide optimization goal among multiple conflicting pharmaceutical peptide optimization goals, calculate the quantitative achievement degree of this pharmaceutical peptide optimization goal based on the obtained corrected or warning - attached quantitative addition parameter combination; based on the obtained corrected or warning - attached quantitative addition parameter combination and the multiple conflicting pharmaceutical peptide optimization goals, identify the conflict relationship between the multiple conflicting pharmaceutical peptide optimization goals, and quantify the conflict relationship to obtain a conflict relationship quantification result; S323. According to the calculated quantitative achievement degree of each pharmaceutical peptide optimization goal, the obtained conflict relationship quantification result, and the obtained user preference information, adopt a comprehensive evaluation algorithm to calculate the comprehensive impact evaluation of the selected parameter combination on the multiple conflicting pharmaceutical peptide optimization goals; S324. Obtain a decision result based on the comprehensive impact evaluation.
[0016] In one aspect of the present application, a method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning is provided. By constructing a chemical rule knowledge base and verifying the parameter combination output by the data - driven model, identify the situations where the chemical rules in the chemical rule knowledge base do not conform or cause side reactions, and make corrections or generate warnings, so as to obtain a decision result. And obtain feedback according to the decision result, update the knowledge base or adjust the model parameters, thereby constructing an optimization decision scenario for the quantitative addition parameters of pharmaceutical peptides, which can integrate chemical rules to correct the parameter verification, provide a decision explanation with chemical logic, and can update the chemical rule knowledge base or adjust the data - driven model parameters based on the decision result and accompanied by expert participation in the decision - making process during the decision result feedback process, improve the reliability of the parameter optimization result, and support the iterative improvement of the system knowledge and model.
[0017] As the second aspect of the present application, an optimization system for the quantitative addition parameters of pharmaceutical peptides based on machine learning is proposed, including: A chemical rule knowledge base construction module, which is used to obtain the original data associated with the field of pharmaceutical peptide synthesis from multiple specific channels, and fuse the obtained original data using a data fusion algorithm to obtain a chemical rule knowledge base; A data acquisition and parameter correction warning module, which is used to acquire the structural characteristics of the target peptide, obtain a quantitative addition parameter combination based on the structural characteristics of the target peptide and the parameters input by the data-driven model, and verify the obtained quantitative addition parameter combination according to the chemical rule knowledge base to identify verification results that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions, and correct or generate a warning for the quantitative addition parameter combination in combination with the verification results, so as to obtain a corrected or warning-attached quantitative addition parameter combination; A decision basis generation module, which is used to generate a decision result based on the corrected or warning-attached quantitative addition parameter combination; A feedback and calibration module, which is used to obtain feedback information based on the decision result and update the chemical rule knowledge base or adjust the preset parameters in the data-driven model based on the obtained feedback information, where the feedback information includes parameter adjustments suggested by the user or the chemical principles supporting their suggestions.
[0018] In the second aspect of the present application, a machine learning-based optimization system for quantitative addition parameters of pharmaceutical peptides is provided, which is constructed based on an optimization decision method for quantitative parameters of pharmaceutical peptides. The system framework is completed by configuring a chemical rule knowledge base construction module, a data acquisition and parameter correction warning module, a decision basis generation module, and a feedback and calibration module, so as to complete the verification by constructing a chemical rule knowledge base and combining the parameter combination output by the data-driven model, identify situations that do not conform to the chemical rules in the chemical rule knowledge base or cause side reactions and make corrections or generate warnings, thereby obtaining a decision result, and obtaining feedback according to the decision result and updating the knowledge base or adjusting the model parameters, and then constructing an optimization decision scenario for quantitative addition parameters of pharmaceutical peptides, which can integrate chemical rules to correct parameters under verification, provide a decision explanation with chemical logic, and can update the chemical rule knowledge base or adjust the parameters of the data-driven model based on the decision result and with the participation of experts in the decision-making process during the feedback of the decision result, improve the reliability of the parameter optimization result, and support the iterative improvement of the system knowledge and model.
[0019] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Description of the Drawings
[0020] Figure 1 It is a schematic flowchart of a machine learning-based optimization method for quantitative addition parameters of pharmaceutical peptides in this embodiment; Figure 2 It is a schematic flowchart of the step S1 in a machine learning-based optimization method for quantitative addition parameters of pharmaceutical peptides in this embodiment; Figure 3 It is a schematic flow chart indicating step S2 in a method for optimizing parameters of quantitative addition of pharmaceutical peptides based on machine learning in this embodiment; Figure 4 It is a schematic flow chart indicating step S22 in a method for optimizing parameters of quantitative addition of pharmaceutical peptides based on machine learning in this embodiment; Figure 5 It is a schematic flow chart indicating step S223 in a method for optimizing parameters of quantitative addition of pharmaceutical peptides based on machine learning in this embodiment; Figure 6 It is a schematic flow chart indicating step S2233 in a method for optimizing parameters of quantitative addition of pharmaceutical peptides based on machine learning in this embodiment; Figure 7 It is a schematic flow chart indicating step S24 in a method for optimizing parameters of quantitative addition of pharmaceutical peptides based on machine learning in this embodiment; Figure 8 It is a schematic flow chart indicating step S3 in a method for optimizing parameters of quantitative addition of pharmaceutical peptides based on machine learning in this embodiment; Figure 9 It is a schematic flow chart indicating step S32 in a method for optimizing parameters of quantitative addition of pharmaceutical peptides based on machine learning in this embodiment; Figure 10 It is a system block diagram of a system for optimizing parameters of quantitative addition of pharmaceutical peptides based on machine learning in this embodiment. Detailed implementation manners
[0021] To better elaborate the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0022] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope protected by the embodiments of the present application.
[0023] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present application. The singular forms of "a", "the" and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0024] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0025] In addition, in the description of the present application, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0026] In the prior art, based on the content of the background art, it can be known that for a data analysis model trained solely relying on the statistical laws of historical data, the recommended quantitative addition parameter combinations sometimes deviate significantly from the empirical judgments formed by senior synthetic chemists based on their understanding of the chemical behaviors of these special structures, or from the specific synthesis strategies reported in the literature for such special structural units.
[0027] Therefore, the technical problem actually solved by the present invention is how to improve the reliability of parameter optimization results and support the iterative improvement of system knowledge and models in the technical field of quantitative addition parameters of pharmaceutical peptides. Based on the above, a preferred example is provided in the present application for illustration as follows: As one aspect of the embodiment of the present invention, as Figure 1 shown, a method for optimizing quantitative addition parameters of pharmaceutical peptides based on machine learning is provided, which includes the following steps: S1. Obtain raw data associated with the field of pharmaceutical peptide synthesis from multiple specific channels, and use a data fusion algorithm to fuse the obtained raw data to obtain a chemical rule knowledge base; S2. Obtain the structural characteristics of the target peptide, and based on the structural characteristics of the target peptide, combine the parameters input by the data-driven model to obtain a quantitative addition parameter combination. According to the chemical rule knowledge base, verify the obtained quantitative addition parameter combination to identify verification results that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions, and correct the quantitative addition parameter combination or generate a warning in combination with the verification results to obtain a corrected or warning-attached quantitative addition parameter combination; S3. Generate a decision result based on the corrected or caution - attached quantitative addition parameter combination; S4. Obtain feedback information based on the decision result and, based on the obtained feedback information, update the chemical rule knowledge base or adjust the preset parameters in the data - driven model, where the feedback information includes parameter adjustments suggested by the user or the chemical principles supporting their suggestions.
[0028] Among them, the construction of the chemical rule knowledge base can collect data from channels such as literature databases, experimental records, and expert experience interviews; data fusion algorithms can adopt techniques such as ontology matching and rule extraction to integrate data from different sources and in different formats to form a structured set of chemical rules, providing a unified and reliable basis for subsequent parameter verification; the structural characteristics of the target peptide can be achieved by inputting amino acid sequences, modification information, etc.; the data - driven model can be a machine - learning model, such as a neural network or a decision tree, which predicts parameter combinations based on historical data; the verification process can be a rule engine that compares the parameter combinations output by the model with the rules in the knowledge base; the verification result can be a boolean value (compliant / non - compliant) or a risk level, and based on the verification result, a risk level can be generated or the parameter values can be automatically adjusted to generate a warning; the decision result can be a recommended parameter solution, a list of alternative solutions, or a parameter combination with a risk assessment. The generation process can be directly outputting the corrected parameter combination or sorting and presenting it in combination with other factors; the acquisition of feedback information can be through user - interface input or experimental result analysis. Updating the chemical rule knowledge base can be converting new experimental findings or expert experience into rules and adding them to the knowledge base, and adjusting the parameters of the data - driven model can be through retraining or parameter fine - tuning to make the model better adapt to new data and rules.
[0029] In a specific example application, after knowledge extraction, each rule in the chemical rule knowledge base can be stored in a structured / vectorized form as "IF [peptide structure feature A AND / OR process condition B] THEN [potential risk C OR recommended operation D]", and experts are allowed to add, delete, query, and modify it. After inputting the target peptide structure and a set of preliminary parameters corresponding to the target peptide structure, the target peptide structure is analyzed, and rules related to these features and input parameters are retrieved in the rule base. If the parameters trigger the "potential risk C" condition of a certain rule, the system marks this risk. If for the identified risk, the corresponding "recommended operation D" is included in the chemical rule knowledge base, the system generates specific parameter adjustment suggestion warnings. Or if there is no correction strategy, a risk warning is given to alert R & D personnel to pay attention and presented in the form of a report, etc., to form a decision result. The decision result can be a recommended parameter solution presented in a list state, a list of alternative solutions, or a parameter group and with a risk assessment.
[0030] Based on the results of the decision, extraction and analysis are carried out. The successful experiences verified by chemical principles or experiments submitted by experts can be formalized into new chemical rule entries and added to the knowledge base after being reviewed and confirmed by the field expert group for their universality and accuracy, or used to correct the applicable conditions or parameter ranges of existing rules.
[0031] Combined with the above, this method structures and stores the chemical principles and expert experiences in the field of pharmaceutical peptide synthesis by establishing a chemical rule knowledge base. After the data-driven model generates a preliminary quantitative addition parameter combination, the chemical rule knowledge base is used to verify the parameter combination. The verification process identifies the parts of the parameter combination that conflict with known chemical rules or may cause side reactions. Based on the verification results, the parameter combination is corrected, such as adjusting the reagent ratio or reaction conditions, or generating a warning message to remind the user of potential risk points. Thus, it ensures that the recommended parameter scheme is chemically reasonable and avoids recommendations that may violate basic chemical principles generated by the model. Further, this method introduces a feedback mechanism. The feedback information generated by the user or the actual synthesis results, including suggestions for parameter adjustment or the chemical principles supporting these suggestions, is used to update the chemical rule knowledge base or adjust the internal parameters of the data-driven model. This feedback loop enables the system to continuously learn and adapt, continuously improving the accuracy and reliability of parameter optimization, and solving the problems of the existing system's difficulty in integrating expert knowledge and lack of continuous improvement ability.
[0032] As a specific example in this embodiment, it is assumed that the target peptide structure contains a sequence known to be prone to β-elimination under alkaline conditions.
[0033] Extract and integrate rules from literature and experimental data to form a chemical rule knowledge base, which contains a rule: the risk of β-elimination is high for a specific sequence when the pH is higher than X; the data-driven model recommends a parameter combination based on historical data, which includes using a buffer with a pH of Y (Y > X). The system obtains the target peptide structure characteristics and this parameter combination. According to the chemical rule knowledge base, the verification module identifies that the buffer pH in this parameter combination conflicts with the β-elimination rule in the knowledge base. The verification result indicates a high risk. The system combines the verification result and automatically corrects the buffer pH to Z (Z < X) to obtain a corrected parameter combination, or generates a warning message: "Attention: The current buffer pH may cause β-elimination side reactions"; generates a parameter scheme containing the corrected pH value as the decision result: if the user adopts the corrected scheme and obtains good results in the experiment, this result can be used as feedback for further model fine-tuning.
[0034] As Figure 2 shown, step S1 in this embodiment includes the following steps: S11. Obtain chemical application knowledge information related to amino acid sequences, protecting groups, and coupling reagents in the field of pharmaceutical peptide synthesis; S12. Based on the chemical application knowledge information, formulate rules according to the peptide structure characteristics, process conditions, and potential risks or recommended operations corresponding to the peptide structure characteristics and process conditions, and store them in a structured manner according to the data sets. Combine the obtained data sets to obtain a chemical rule knowledge base.
[0035] Among them, in step S11, specific chemical application knowledge directly related to pharmaceutical peptide synthesis is obtained. This knowledge can be collected from public chemical literature, patent databases, chemical reaction databases, expert experience bases, or internal experimental data. The acquisition process can adopt methods such as information extraction, text mining, or manual collation. The obtained information focuses on the behavior and influence of specific amino acid sequences (e.g., amino acids prone to racemization, having steric hindrance or specific reaction activities), common protecting groups (e.g., their removal conditions, stability, possible side reactions), and coupling reagents (e.g., coupling efficiency, side reaction tendency, applicable scope) under different synthesis conditions.
[0036] Furthermore, in step S12, the obtained chemical application knowledge is transformed into structured rules. The rule formulation can be based on a predefined rule template, mapping the knowledge information to fields such as "peptide structure characteristics", "process conditions", "potential risks", and "recommended operations". For example: A rule can be described as "When the peptide chain contains a specific sequence fragment and a specific coupling reagent is used, there is a certain risk of side reactions". These rules are organized into data sets and stored in a structured manner, such as stored in a table of a relational database. Each rule is used as a record, containing the values of each field, and finally a chemical rule knowledge base is formed, such as the reference structure sequence "[Peptide structure characteristic A AND / OR Process condition B] THEN [Potential risk C OR Recommended operation D]".
[0037] In a specific example, consider building a chemical rule knowledge base for racemizable amino acids under different coupling conditions. Obtain racemization rate data and related literature reports on common racemizable amino acids (e.g., His, Cys, Asp) in solid-phase peptide synthesis when using different coupling reagents (e.g., HATU, HBTU, PyBOP) and additives (e.g., HOBt, HOAt). For example, obtain the knowledge that when coupling the C-terminal His with HATU, the racemization risk is high. Step S12 formulates a rule based on this knowledge: "If the C-terminal of the peptide chain contains His and the coupling reagent is HATU, there is a high racemization risk, and it is recommended to use an alternative coupling reagent or add HOAt". This rule is structurally stored in a table of a relational database, with the field "peptide structure feature" having a value of "C-terminal His", "process condition" having a value of "coupling reagent HATU", "potential risk" having a value of "high racemization risk", and "recommended operation" having a value of "use an alternative coupling reagent / add HOAt".
[0038] As Figure 3 shown, step S2 in this embodiment includes the following steps: S21. Obtain a quantitative addition parameter combination obtained by combining the structural features of the target peptide and the parameters output by the data-driven model; S22. According to the chemical rule knowledge base, conduct a primary verification on the obtained quantitative addition parameter combination to identify a primary verification result that does not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions; S23. Based on the primary verification result, conduct a primary correction on the current quantitative addition parameter combination or generate a warning to obtain a corrected or warning-attached quantitative addition parameter combination obtained after the primary verification; S24. Perform an iterative verification process on the target peptide information set with the corrected or warning-attached quantitative addition parameter combination obtained after the primary verification, verify the current quantitative addition parameter combination to obtain a final verification result associated with the current quantitative addition parameter combination, and obtain a corrected or warning-attached quantitative addition parameter combination based on the final verification result.
[0039] Among them, step 21 is responsible for receiving or generating a quantitative addition parameter combination to be processed, and this combination is a parameter set calculated based on the structural information of the target peptide and the data-driven model.
[0040] Step S22 is to perform the first rule check, compare with the chemical rule knowledge base, and identify the part of the parameter combination that conflicts with the known chemical rules or has a potential side reaction risk. Thus, the primary verification result is obtained.
[0041] Step S23 is to adjust the parameter combination or mark a warning according to the primary verification result, which is a direct response to the initially discovered problem.
[0042] Step S24 and step S23 are iterative verification processes. Using the corrected or warning-attached parameter combination obtained in step S23 as the input, start the loop verification. In each loop, recheck the current parameter combination according to the chemical rule knowledge base.
[0043] Specifically, in this process step, by introducing an iterative verification process, the verification and correction capabilities for the quantitative addition parameter combination are enhanced, solving the problem that a single verification and correction process may introduce new situations inconsistent with chemical rules or trigger the risk of new side reactions. In the iterative process, use the initially corrected parameter combination as the input and perform verification again according to the chemical rule knowledge base. If new inconsistencies or risks are found, perform re-correction. This verification and correction loop continues until the parameter combination no longer requires correction after a complete verification, indicating that it has reached a stable state, or until the preset maximum number of iterations is reached. The finally output parameter combination is the result of multiple iterations of verification and correction, reducing the risk of introducing new problems due to single correction and improving the reliability of the parameter combination.
[0044] In a specific example, assume that the target peptide sequence contains an amino acid residue prone to racemization. Step S21 obtains the quantitative addition parameter combination output by the data-driven model, which may include a relatively high reaction temperature and a certain general coupling agent. Step S22 performs an initial verification according to the chemical rule knowledge base. There is a rule in the knowledge base indicating that "easily racemizable residue + high temperature + general coupling agent -> high racemization risk". The initial verification result identifies that this parameter combination has a high racemization risk. Based on this result, step S23 makes an initial correction to the parameter combination, for example, reducing the reaction temperature to 10 °C and suggesting the use of a coupling agent known to inhibit racemization. Step S24 starts the iterative verification. Use the corrected parameter combination (low temperature, specific coupling agent) as the input for re-verification. In the re-verification of the first iteration, check this new combination according to the chemical rule knowledge base. It may be found that the low temperature will significantly extend the reaction time, conflicting with another rule "too long reaction time + certain protecting groups -> risk of protecting group detachment". The re-verification result identifies a new risk. Based on this result, perform re-correction, for example, fine-tuning the temperature to 15 °C and adding a small amount of catalyst to the reaction system to shorten the reaction time. Enter the second iteration and perform re-verification on the parameter combination (15 °C, specific coupling agent, small amount of catalyst) after the second correction. If no new inconsistencies with chemical rules or side reactions are found in this verification, it is determined that the stable state has been reached and the iteration terminates. Finally, output this parameter combination that has been corrected and verified multiple times.
[0045] As Figure 4 shown, step S22 in this embodiment includes the following steps: S221. Obtain one or more parameter subsets in the current quantitative addition parameter combination where there are synergistic or antagonistic effects among the member parameters. For each identified parameter subset, retrieve from the chemical rule knowledge base or the parameter interaction knowledge database the parameter interaction rules corresponding to this parameter subset. The parameter interaction rules indicate the impact on the process and efficiency of the target chemical reaction, or on the occurrence probability and degree of side reactions, when the member parameters within this parameter subset act together. S222. Based on each retrieved parameter interaction rule, perform an adjustment to verify the compliance conclusion of a single parameter or the entire parameter combination within the parameter subset in the single rule verification result, or evaluate the risk level or occurrence probability of side reactions caused by the joint action of the member parameters within the parameter subset, to obtain parameter interaction adjustment information. S224. According to the chemical rule knowledge base, perform a single rule verification on the obtained quantitative addition parameter combination to identify the parameters that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions, to obtain a single rule verification result. S224. Combine the single rule verification result and the parameter interaction adjustment information to obtain a preliminary verification result.
[0046] Specifically, during the synthesis of pharmaceutical peptides, the accuracy of the quantitative addition parameter combination directly affects the synthesis result. Traditional verification methods may only check whether each parameter conforms to independent rules, ignoring the additional risks or impacts that may be brought about by the interaction between parameters. By introducing the analysis of the interaction between parameters in step S22, the comprehensiveness of the verification is improved.
[0047] First, identify the parameter groups in the quantitative addition parameter combination that may have synergistic or antagonistic effects. For example, the concentration of the coupling agent and the reaction temperature may interact with each other. Then, query the knowledge base to obtain the rules describing how these parameter combinations affect the reaction or side reactions. These rules may indicate that even though the concentration of the coupling agent and the temperature are individually within the acceptable range, their specific combination may significantly increase the occurrence probability of a certain side reaction.
[0048] Meanwhile, perform a basic single rule verification to check whether each parameter or parameter combination conforms to independent chemical rules. Finally, combine the result of the single rule verification with the adjustment information obtained from the parameter interaction analysis. If the parameter interaction rule indicates that a certain parameter combination will significantly increase the risk of side reactions, even if no problems are found in the single rule verification, this risk information will be integrated into the final preliminary verification result.
[0049] Thus, the initial verification result not only reflects the compliance of the parameters with the independent rules, but also takes into account the overall effect of the parameter combination, thereby more accurately identifying potential problems and reducing the synthesis risk caused by parameter interactions.
[0050] In a specific example, for a target peptide structure, the system obtains its quantitative addition parameter combination. For example, the molar equivalent of coupling agent A is 1.5, the molar equivalent of base B is 2.0, and the reaction temperature is 25°C. The system identifies that there may be a synergistic effect between the molar equivalent of coupling agent A and the molar equivalent of base B, and retrieves the rule from the knowledge base: when the molar equivalent of coupling agent A is higher than 1.2 and the molar equivalent of base B is higher than 1.8, the occurrence probability of a specific side reaction C increases significantly. At the same time, the system performs a single rule verification and finds that the 1.5 molar equivalent of coupling agent A complies with its independent use rule (for example, the recommended range is 0.8 - 2.0), the 2.0 molar equivalent of base B also complies with its independent use rule (for example, the recommended range is 1.0 - 2.5), and the reaction temperature of 25°C complies with its independent use rule (for example, the recommended range is 20 - 30°C). The single rule verification result shows that these parameters individually comply with the rules. However, based on the retrieved parameter interaction rule, the system evaluates that the combination of the 1.5 molar equivalent of coupling agent A and the 2.0 molar equivalent of base B significantly increases the risk of side reaction C, generates parameter interaction adjustment information indicating a high risk of side reaction C. Finally, the system combines the single rule verification result with the parameter interaction adjustment information to obtain the initial verification result, indicating that although this parameter combination individually complies with the rules, there is a high risk of side reaction C due to parameter interactions and may generate corresponding warning information.
[0051] As Figure 5 shown, step S223 of this embodiment includes the following steps: S2231. Retrieve chemical rules related to the current quantitative addition parameter combination and the structural characteristics of the target peptide from the chemical rule knowledge base, and identify the rules with their own confidence scores or probability descriptions of their applicable conditions in the relevant chemical rules, to obtain the identified chemical rules with uncertainty information attached; S2232. For each obtained identified chemical rule with uncertainty information attached, calculate the compliance probability of the current quantitative addition parameter combination for this chemical rule with uncertainty information attached based on the confidence score of this chemical rule itself or the probability description of its applicable conditions, and the matching situation between the quantitative addition parameter combination and the preconditions of this chemical rule; S2233. For one or more preset side reactions, based on one or more chemical rules related to the side reaction, combine the compliance probability of the current quantitative addition parameter combination with the chemical rule with uncertainty information, and when multiple relevant chemical rules contribute to the risk of the same side reaction, use the probability synthesis method to calculate the comprehensive risk probability of the current quantitative addition parameter combination triggering the side reaction; S2234. Use the compliance probability of the calculated current quantitative addition parameter combination for each chemical rule and the comprehensive risk probability of the calculated current quantitative addition parameter combination triggering each preset side reaction to identify the parameters that do not conform to the chemical rules in the chemical rule knowledge base or may trigger side reactions, and obtain a single rule verification result.
[0052] Among them, retrieve the chemical rules related to the current quantitative addition parameter combination and the structural characteristics of the target peptide from the chemical rule knowledge base, and identify the rules with probability descriptions of their own confidence scores or applicable conditions attached to the relevant chemical rules to obtain the chemical rules with uncertainty information attached; for each obtained chemical rule with uncertainty information attached, based on the confidence score of the chemical rule itself or the probability description of the applicable conditions, and the matching situation between the quantitative addition parameter combination and the prerequisite conditions of the chemical rule, calculate the compliance probability of the current quantitative addition parameter combination with the chemical rule with uncertainty information attached; for one or more preset side reactions, based on one or more chemical rules related to the side reaction, combine the compliance probability of the current quantitative addition parameter combination with the chemical rule with uncertainty information attached, and when multiple relevant chemical rules contribute to the risk of the same side reaction, use the probability synthesis method to calculate the comprehensive risk probability of the current quantitative addition parameter combination triggering the side reaction; use the compliance probability of the calculated current quantitative addition parameter combination for each chemical rule and the comprehensive risk probability of the calculated current quantitative addition parameter combination triggering each preset side reaction to identify the parameters that do not conform to the chemical rules in the chemical rule knowledge base or may trigger side reactions, and obtain a single rule verification result.
[0053] In addition, chemical rules with attached uncertainty information can be identified by means of metadata tags or fields preset for each rule in the chemical rule knowledge base, such as a field representing a confidence value or a field describing an applicable probability distribution. The retrieval process distinguishes between deterministic rules and uncertain rules based on the presence or absence of these fields. Calculating the compliance probability of the current quantitative addition parameter combination for a chemical rule with attached uncertainty information can involve combining the confidence or probability information of the rule itself with the degree to which the parameter combination meets the prerequisite conditions of the rule. For example, if the prerequisite conditions of the rule are fully met, the compliance probability is equal to the confidence of the rule itself; if partially met, the confidence is weighted according to the degree of match. When calculating the comprehensive risk probability of side reactions, the individual risk contributions of multiple rules related to a specific side reaction are used as inputs and integrated using a probability synthesis method. The probability synthesis method can be implemented using Bayesian networks, fuzzy logic reasoning, or a simple probability aggregation model, thereby obtaining a quantified probability of side reaction occurrence.
[0054] In a specific example, when synthesizing a peptide segment containing the Asp-Gly sequence, this sequence is known to be prone to the side reaction of imidization under specific conditions. A data-driven model recommends a set of parameters, including using HATU as the coupling agent, DIPEA as the base, DMF as the solvent, and setting the reaction temperature at 25°C. The system performs a single rule check. First, retrieve the rules related to the Asp-Gly sequence, HATU, DIPEA, DMF, and 25°C temperature from the chemical rule knowledge base. The system identifies a rule: "Using HATU / DIPEA to couple the Asp-Gly sequence in DMF increases the risk of imidization at 25°C (confidence: 0.7)". This rule is identified as carrying uncertainty information. The system calculates the matching situation between the current parameter combination (HATU, DIPEA, DMF, 25°C) and the preconditions of this rule and finds a perfect match. Based on the 0.7 confidence of the rule, the compliance probability of the current parameter combination for this rule is calculated to be 0.7. The system also identifies another rule: "The Asp-Gly sequence itself is prone to imidization (probability under typical conditions: 0.8)". Both of these rules are related to the side reaction of imidization. The system adopts a probability synthesis method, such as considering the independent contributions of the two rules for probability aggregation, and calculates the comprehensive risk probability of the current parameter combination causing the side reaction of imidization. For example, the calculation result is 0.65. Based on the calculated rule compliance probability of 0.7 and the imidization risk probability of 0.65, the system identifies that the parameters HATU, DIPEA, DMF, and 25°C are related to the potential imidization risk and outputs this information as the result of the single rule check. For example, these parameters are marked as "parameters highly associated with the imidization risk" and accompanied by the risk probability value. Thus, the R & D personnel obtain a quantitative risk assessment rather than just a simple rule violation prompt.
[0055] As Figure 6 shown, in this embodiment, step S2233 includes the following steps: S22331. For multiple chemical rules related to the side reaction, retrieve and obtain the dependency descriptions between these chemical rules from the rule dependency data source. The dependency description indicates the interaction mode between these chemical rules when jointly affecting the occurrence probability of the side reaction; S22332. Before performing the probability synthesis operation, according to the obtained dependency description, correct the individual risk contribution to the side reaction generated by at least one of the multiple chemical rules, to obtain the corrected individual risk contribution, and / or construct a probability synthesis structure that integrates the dependency description, and obtain the individual risk contributions to the side reaction generated by each of the multiple chemical rules related to the side reaction as the input of the probability synthesis structure, where the individual risk contribution indicates the compliance probability of the current quantitative addition parameter combination for the chemical rule calculated for the corresponding rule, or the probability information directly included in the corresponding rule for describing the initiation of the side reaction; S22333. Perform the probability synthesis operation using the obtained corrected individual risk contribution, and / or perform the probability synthesis operation using the probability synthesis structure constructed in the previous step and the obtained individual risk contribution as its input, so as to calculate, through the probability synthesis operation, the comprehensive risk probability of the combination of the dependency relationship between chemical rules and the current quantitative addition parameter combination triggering the side reaction.
[0056] Among them, for multiple chemical rules related to the side reaction, retrieve and obtain the dependency description between these chemical rules from the rule dependency data source. The dependency description indicates the interaction mode between these chemical rules when jointly affecting the occurrence probability of the side reaction; before performing the probability synthesis operation, according to the obtained dependency description, correct the individual risk contribution to the side reaction generated by at least one of the multiple chemical rules, to obtain the corrected individual risk contribution, and / or construct a probability synthesis structure that integrates the dependency description, and obtain the individual risk contributions to the side reaction generated by each of the multiple chemical rules related to the side reaction as the input of the probability synthesis structure, where the individual risk contribution indicates the compliance probability of the current quantitative addition parameter combination for the chemical rule calculated for the corresponding rule, or the probability information directly included in the corresponding rule for describing the initiation of the side reaction; perform the probability synthesis operation using the obtained corrected individual risk contribution, and / or perform the probability synthesis operation using the probability synthesis structure constructed in the previous step and the obtained individual risk contribution as its input, so as to calculate, through the probability synthesis operation, the comprehensive risk probability of the combination of the dependency relationship between chemical rules and the current quantitative addition parameter combination triggering the side reaction, and the rule dependency data source can store the interaction mode between chemical rules, such as indicating an enhancement, inhibition, or mutual exclusion relationship between rules.
[0057] In a specific example, consider that a target peptide may undergo a β-elimination side reaction under specific conditions. Suppose there are two chemical rules related to this: Rule R1 indicates that a specific sequence fragment is prone to β-elimination under alkaline conditions, and its individual risk contribution is P1; Rule R2 indicates that a specific protecting group is prone to trigger β-elimination under the action of a specific deprotecting reagent, and its individual risk contribution is P2. It is retrieved from the rule-dependency data source that there is a synergistic enhancement relationship between Rule R1 and Rule R2, that is, when the sequence fragment and the protecting group both meet the conditions, the risk of β-elimination is significantly higher than the simple superposition when they act alone. According to the obtained description of the synergistic enhancement dependency, before performing the probability synthesis operation, P1 and P2 can be corrected. For example, a non-linear combination function f(P1, P2) = P1 + P2 + α * P1 * P2 (where α > 0 represents the enhancement factor) is used to calculate the corrected joint contribution. Alternatively, a simple Bayesian network can be constructed, in which there are two parent nodes representing the satisfaction of the conditions of Rule R1 and Rule R2 respectively, and a child node representing the occurrence of the β-elimination side reaction. The conditional probability table of the child node is defined according to the dependency between R1 and R2. Taking P1 and P2 as the input probabilities of the parent nodes, the comprehensive risk probability of the β-elimination side reaction is calculated through network inference. For example, if P1 = 0.4, P2 = 0.3, and the dependency indicates synergistic enhancement, simple addition may result in 0.7, but considering the enhancement effect, the corrected joint contribution or the result of Bayesian network inference may be 0.85, which more accurately reflects the actual risk.
[0058] As Figure 7 shown, in this embodiment, step S24 includes the following steps: S241. After obtaining the corrected or warning-attached quantitative addition parameter combination obtained after the initial verification, start the iterative verification process. In each iteration of the iterative verification process, re-verify the quantitative addition parameter combination of the current iteration round according to the chemical rule knowledge base; S242. Determine whether the iterative verification process meets the preset iterative termination condition. The iterative termination condition indicates that the updated quantitative addition parameter combination reaches a stable state without further correction after a complete re-verification or the number of iterations of the iterative process reaches the preset upper limit; if the iterative termination condition is met, use the updated quantitative addition parameter combination as the corrected or warning-attached quantitative addition parameter combination; if the preset iterative termination condition is not met, use the updated quantitative addition parameter combination as the input quantitative addition parameter combination for the next iteration round, and return to perform the re-verification in each iteration of the iterative process, and finally obtain the re-verification result; S242. Based on the re-verification result, perform re-correction on the quantitative addition parameter combination of the current iteration round to obtain the finally corrected or warning-attached quantitative addition parameter combination.
[0059] Among them, after presenting the corrected or warning-attached quantitative addition parameter combination obtained after the initial verification, start the iterative verification process. In each iteration of the iterative verification process, perform re-verification on the quantitative addition parameter combination of the current iteration round according to the chemical rule knowledge base; determine whether the iterative verification process meets the preset iterative termination condition, where the iterative termination condition indicates that the updated quantitative addition parameter combination reaches a stable state without further re-correction after a complete re-verification or the number of iterations of the iterative process reaches the preset upper limit; if the iterative termination condition is met, use the updated quantitative addition parameter combination as the corrected or warning-attached quantitative addition parameter combination; if the preset iterative termination condition is not met, use the updated quantitative addition parameter combination as the input quantitative addition parameter combination for the next iteration round, return to perform the re-verification in each iteration of the iterative process, and finally obtain the re-verification result; based on the re-verification result, perform re-correction on the quantitative addition parameter combination of the current iteration round to obtain the finally corrected or warning-attached quantitative addition parameter combination.
[0060] In a specific example, for instance, an initial parameter combination P = {p1: 1.0, p2: 2.0} becomes P' = {p1: 0.8, p2: 2.2} after the initial verification and correction, and a warning message W1 is attached indicating that the value of p1 may not conform to rule R_A. Start the iterative verification process. In the first iteration, perform re-verification on P' (S241). The re-verification finds that although p1 = 0.8 no longer violates R_A, p2 = 2.2 does not conform to rule R_B, and the combination of p1 and p2 may trigger side reaction F1. By judging the termination condition, the parameter combination has changed and the upper limit of the number of iterations has not been reached. The termination condition is not met. Based on the re-verification result, perform re-correction on P'. The system corrects p2 to 2.0 and adds a warning W2 indicating the combination risk of p1 and p2. Obtain a new parameter combination P'' = {p1: 0.8, p2: 2.0}, with warnings W1 and W2 attached. Use P'' as the input for the next iteration. In the second iteration, perform re-verification on P''. The re-verification finds that P'' conforms to all rules and there are no new risks. By judging the termination condition (S242), the parameter combination P'' does not require further correction after this re-verification and reaches a stable state. The termination condition is met. The iteration terminates, and P'' = {p1: 0.8, p2: 2.0} with warnings W1 and W2 attached is output as the finally corrected or warning-attached quantitative addition parameter combination.
[0061] As Figure 8 shown, in this embodiment, step S3 includes the following steps: S31. Obtain the relevance result between the corrected or warning - attached quantitative addition parameter combination and the chemical rule knowledge base; S32. Evaluate the impact of the corrected or warning - attached quantitative addition parameter combination on at least one preset optimization goal based on the relevance result, and obtain a decision result.
[0062] Among them, step S31 is realized by querying the chemical rule knowledge base for the corrected or warning - attached quantitative addition parameter combination. By comparing each parameter in the parameter combination with the chemical rules stored in the knowledge base, it is identified whether the parameter combination conforms to the rules, whether there are potential conflicts or risks, and the relevance result is presented in the form of a relevance report or score.
[0063] Based on the relevance result obtained in step S32, step S32 evaluates the impact of the parameter combination on the preset optimization goal. The preset optimization goals can include yield, purity, impurity content, cost, reaction time, etc. And the evaluation process receives the parameter combination, the relevance result, and the preset optimization goal as inputs, outputs the predicted impact on each goal or a comprehensive evaluation result, and finally generates a decision result, which reflects the comprehensive performance of the parameter combination on multiple optimization goals.
[0064] As Figure 9 shown, in this embodiment, step S321 includes the following steps: S321. Obtain the selected corrected or warning - attached quantitative addition parameter combination and the user's preference information for multiple conflicting pharmaceutical peptide optimization goals; S322. For each pharmaceutical peptide optimization goal among multiple conflicting pharmaceutical peptide optimization goals, calculate the quantitative achievement degree of this pharmaceutical peptide optimization goal based on the obtained corrected or warning - attached quantitative addition parameter combination; based on the obtained corrected or warning - attached quantitative addition parameter combination and multiple conflicting pharmaceutical peptide optimization goals, identify the conflict relationship between multiple conflicting pharmaceutical peptide optimization goals, and quantify the conflict relationship to obtain the conflict relationship quantification result; S323. According to the calculated quantitative achievement degree of each pharmaceutical peptide optimization goal, the obtained conflict relationship quantification result, and the obtained user preference information, use a comprehensive evaluation algorithm to calculate the comprehensive impact evaluation of the selected parameter combination on multiple conflicting pharmaceutical peptide optimization goals; S324. Obtain a decision result based on the comprehensive impact evaluation.
[0065] Among them, obtain the quantitative addition parameter combinations to be evaluated and the optimization target weights or priority information set by the user. For each optimization target, calculate the predicted performance value of the current parameter combination on this target. At the same time, analyze the mutual influence patterns between different optimization targets. For example, an improvement in one target may lead to a decline in another target, and convert this degree of mutual influence into a numerical representation. Use the obtained individual target performance values, the numerical values of the mutual influence between targets, and the obtained user preference information as inputs, and calculate through a preset algorithm model to output a numerical value or level reflecting the overall performance of the parameter combination. Step S324 performs a decision generation operation, and determines the final decision output based on the obtained overall performance numerical value or level.
[0066] This step is used to solve the problem of how to evaluate parameter combinations and make decisions when there are multiple interrelated and possibly mutually restrictive optimization targets (such as yield, purity, cost, time) during the optimization of the quantitative addition parameters of pharmaceutical peptides. First, obtain a parameter combination that has been preliminarily corrected or with a warning, and receive the user's input on the importance of different optimization targets. Then, the system calculates the expected performance values of this parameter combination on each single optimization target, analyzes the degree of mutual influence between these targets, and quantifies this influence. Subsequently, combine this information with the user's preferences and calculate through an algorithm the overall evaluation of this parameter combination considering all targets, the relationships between targets, and the user's requirements. Finally, based on this overall evaluation, generate a decision result regarding this parameter combination.
[0067] In summary, in one aspect described in this embodiment, a method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning is provided. By constructing a chemical rule knowledge base and verifying the parameter combinations output by the data-driven model, identify situations where the chemical rules in the chemical rule knowledge base do not conform or cause side reactions, and make corrections or generate warnings, so as to obtain a decision result. And obtain feedback based on the decision result and update the knowledge base or adjust the model parameters, thereby constructing a decision-making scenario for optimizing the quantitative addition parameters of pharmaceutical peptides, which can integrate chemical rules to correct the parameter verification, provide a decision explanation of chemical logic, and can update the chemical rule knowledge base or adjust the data-driven model parameters based on the decision result and accompanied by expert participation in the decision-making during the decision result feedback process, improve the reliability of the parameter optimization result, and support the iterative improvement of the system knowledge and model.
[0068] As a second aspect of the embodiment of the present invention, as Figure 10 shown, a system 100 for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning is provided, including: The Chemical Rule Knowledge Base Construction Module 101 is used to obtain the original data related to the field of pharmaceutical peptide synthesis from multiple specific channels, and fuse the obtained original data using a data fusion algorithm to obtain the chemical rule knowledge base; The Data Acquisition and Parameter Correction Warning Module 102 is used to obtain the structural characteristics of the target peptide, obtain the quantitative addition parameter combination based on the structural characteristics of the target peptide and the parameters input by the data-driven model, and verify the obtained quantitative addition parameter combination according to the chemical rule knowledge base to identify the verification results that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions, and correct or generate warnings for the quantitative addition parameter combination based on the verification results to obtain the corrected or warning-attached quantitative addition parameter combination; The Decision Basis Generation Module 103 is used to generate a decision result based on the corrected or warning-attached quantitative addition parameter combination; The Feedback and Calibration Module 104 is used to obtain feedback information based on the decision result and update the chemical rule knowledge base or adjust the preset parameters in the data-driven model based on the obtained feedback information, where the feedback information includes the parameter adjustment suggested by the user or the chemical principle supporting the suggestion.
[0069] In the second aspect of this embodiment, a machine learning-based pharmaceutical peptide quantitative addition parameter optimization system is constructed based on an optimization decision method for pharmaceutical peptide quantitative parameters. The system framework is completed by configuring the chemical rule knowledge base construction module, the data acquisition and parameter correction warning module, the decision basis generation module, and the feedback and calibration module. By constructing the chemical rule knowledge base and verifying the parameter combination output by the data-driven model, situations that do not conform to the chemical rules in the chemical rule knowledge base or cause side reactions are identified and corrected or warnings are generated, so as to obtain a decision result. And feedback is obtained according to the decision result and the knowledge base is updated or the model parameters are adjusted, thereby constructing an optimization decision scenario for pharmaceutical peptide quantitative addition parameters, which can integrate chemical rules to correct parameter verification, provide a decision explanation of chemical logic, and can update the chemical rule knowledge base or adjust the data-driven model parameters based on the decision result and with the participation of experts in the decision-making process during the decision result feedback process, improve the reliability of the parameter optimization result, and support the iterative improvement of the system knowledge and model.
[0070] Based on the disclosure and teachings of the above specification, those skilled in the art to which the present invention pertains can also make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the protection scope of the claims of the present invention. In addition, although some specific terms are used in this specification, these terms are only for convenience of description and do not constitute any limitation to the present invention.
Claims
1. A method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning, characterized in that, It includes the following steps: S1. Obtain the original data related to the field of pharmaceutical peptide synthesis from multiple specific channels, and use a data fusion algorithm to fuse the obtained original data to obtain a chemical rule knowledge base; S2. Obtain the structural characteristics of the target peptide, and based on the structural characteristics of the target peptide, combine the parameters input by the data-driven model to obtain a quantitative addition parameter combination. According to the chemical rule knowledge base, verify the obtained quantitative addition parameter combination to identify the verification results that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions, and correct the quantitative addition parameter combination or generate a warning in combination with the verification results to obtain a corrected or warning-attached quantitative addition parameter combination; S3. Generate a decision result based on the corrected or warning-attached quantitative addition parameter combination; S4. Obtain feedback information based on the decision result and, based on the obtained feedback information, update the chemical rule knowledge base or adjust the preset parameters in the data-driven model, where the feedback information includes the parameter adjustment suggested by the user or the chemical principle supporting the suggestion.
2. The method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning according to claim 1, wherein, The step S1 includes the following steps: S11. Obtain the chemical application knowledge information related to amino acid sequences, protecting groups, and coupling reagents in the field of pharmaceutical peptide synthesis; S12. Based on the chemical application knowledge information, formulate rules according to the peptide structure characteristics and process conditions and the potential risks or recommended operations corresponding to the peptide structure characteristics and process conditions, and store them structurally according to the data set. Combine the obtained data sets to obtain a chemical rule knowledge base.
3. A method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning according to claim 1, characterized in that, The step S2 specifically includes the following steps: S21. Obtain a quantitative addition parameter combination by combining the structural characteristics of the target peptide and the parameters output by the data-driven model; S22. According to the chemical rule knowledge base, conduct a primary verification on the obtained quantitative addition parameter combination to identify the primary verification results that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions; S23. Based on the primary verification results, conduct a primary correction on the current quantitative addition parameter combination or generate a warning to obtain a corrected or warning-attached quantitative addition parameter combination obtained after the primary verification; S24. Conduct an iterative verification process on the target peptide information set with the corrected or warning-attached quantitative addition parameter combination obtained after the primary verification, verify the current quantitative addition parameter combination to obtain the final verification result associated with the current quantitative addition parameter combination, and obtain a corrected or warning-attached quantitative addition parameter combination based on the final verification result.
4. A method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning according to claim 3, characterized in that The step S22 specifically includes: S221. Obtain one or more parameter subsets in which there are synergistic or antagonistic effects among the member parameters in the current quantitative addition parameter combination, and for each identified parameter subset, retrieve the parameter interaction rules corresponding to the parameter subset from the chemical rule knowledge base or the parameter interaction knowledge database, where the parameter interaction rules indicate the influence on the process and efficiency of the target chemical reaction or the probability and degree of side reaction occurrence when the member parameters within the parameter subset act together; S222. Based on each retrieved parameter interaction rule, perform adjustments to verify the compliance conclusion of a single parameter or the entire parameter combination within the parameter subset in the single-rule verification result, or evaluate the risk level or occurrence probability of side reactions caused by the combined action of member parameters within the parameter subset, to obtain parameter interaction adjustment information; S223. According to the chemical rule knowledge base, perform a single-rule verification on the obtained quantitative addition parameter combination to identify parameters that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions, to obtain a single-rule verification result; S224. Combine the single-rule verification result and the parameter interaction adjustment information to obtain a primary verification result.
5. A method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning according to claim 4, characterized in that The specific steps of step S223 include: S2231. Retrieve chemical rules related to the current quantitative addition parameter combination and the structural characteristics of the target peptide from the chemical rule knowledge base, and identify rules with their own confidence scores or probability descriptions of their applicable conditions attached to them, to obtain chemical rules with attached uncertainty information; S2232. For each obtained chemical rule with attached uncertainty information, calculate the compliance probability of the current quantitative addition parameter combination for this chemical rule with attached uncertainty information based on the confidence score of the chemical rule itself or the probability description of its applicable conditions, and the matching situation between the quantitative addition parameter combination and the preconditions of this chemical rule; S2233. For one or more preset side reactions, based on one or more chemical rules related to this side reaction, combine the compliance probability of the current quantitative addition parameter combination for this chemical rule with attached uncertainty information, and when multiple relevant chemical rules contribute to the risk of the same side reaction, use a probability synthesis method to calculate the comprehensive risk probability of the current quantitative addition parameter combination causing this side reaction; S2234. Use the calculated compliance probability of the current quantitative addition parameter combination for each chemical rule and the calculated comprehensive risk probability of the current quantitative addition parameter combination causing each preset side reaction to identify parameters that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions, and obtain a single-rule verification result.
6. A method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning according to claim 5, characterized in that, The specific steps of step S2233 include: S22331. For multiple chemical rules related to the side reaction, retrieve and obtain the dependency relationship descriptions between these chemical rules from the rule dependency data source, and the dependency relationship descriptions indicate the interaction patterns of these chemical rules when jointly affecting the occurrence probability of the side reaction; S22332. According to the obtained dependency description, before performing probability synthesis operations, correct the individual risk contribution of at least one of the multiple chemical rules to the side reaction generated by this rule to obtain the corrected individual risk contribution, and / or construct a probability synthesis structure that integrates the dependency description, and obtain the individual risk contributions of the multiple chemical rules related to the side reaction generated by each rule as the input of the probability synthesis structure, where the individual risk contribution indicates the compliance probability of the current quantitative addition parameter combination calculated for the corresponding rule for this chemical rule, or the probability information directly included in the corresponding rule for describing the probability of triggering this side reaction; S22333. Perform probability synthesis operations using the obtained corrected individual risk contribution, and / or perform probability synthesis operations using the probability synthesis structure constructed in the previous step and the obtained individual risk contribution as its input, so as to calculate, through the probability synthesis operations, the comprehensive risk probability of triggering this side reaction by combining the dependency relationship between chemical rules and the current quantitative addition parameter combination.
7. A method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning according to claim 3, characterized in that, The specific steps of step S24 include: S241. After obtaining the corrected or warning-attached quantitative addition parameter combination obtained after the initial verification, start iterative verification processing. In each iteration of the iterative verification processing, re-verify the quantitative addition parameter combination of the current iteration round according to the chemical rule knowledge base; S242. Determine whether the iterative verification processing meets the preset iterative termination conditions. The iterative termination conditions indicate that the updated quantitative addition parameter combination reaches a stable state without further correction after a complete re-verification, or the number of iterations of the iterative processing reaches the preset upper limit; if the iterative termination conditions are met, use the updated quantitative addition parameter combination as the corrected or warning-attached quantitative addition parameter combination; if the preset iterative termination conditions are not met, use the updated quantitative addition parameter combination as the input quantitative addition parameter combination for the next iteration round, and return to perform the re-verification in each iteration of the iterative processing, and finally obtain the re-verification result; S243. Based on the re-verification result, perform re-correction on the quantitative addition parameter combination of the current iteration round to obtain the final corrected or warning-attached quantitative addition parameter combination.
8. A method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning according to claim 1, characterized in that, The specific steps of step S3 include: S31. Obtain the relevance result between the corrected or warning-attached quantitative addition parameter combination and the chemical rule knowledge base; S32. Evaluate the impact of the corrected or warning-attached quantitative addition parameter combination on at least one preset optimization goal based on the relevance result, and obtain a decision result.
9. A method for optimizing the quantitative addition parameters of pharmaceutical peptides based on machine learning according to claim 8, characterized in that, The specific steps of step S32 include: S321. Obtain the selected corrected or warning-attached quantitative addition parameter combination and the user's preference information for multiple conflicting pharmaceutical peptide optimization goals; S322. For each of the multiple conflicting pharmaceutical peptide optimization objectives, based on the obtained corrected or caution-attached quantitative addition parameter combinations, calculate the quantitative achievement degree of this pharmaceutical peptide optimization objective; based on the obtained corrected or caution-attached quantitative addition parameter combinations and the multiple conflicting pharmaceutical peptide optimization objectives, identify the conflict relationships between the multiple conflicting pharmaceutical peptide optimization objectives, and quantify the conflict relationships to obtain the conflict relationship quantification results; S323. According to the quantitative achievement degree of each calculated pharmaceutical peptide optimization objective, the obtained conflict relationship quantification results, and the obtained user preference information, use a comprehensive evaluation algorithm to calculate the comprehensive impact evaluation of the selected parameter combination on the multiple conflicting pharmaceutical peptide optimization objectives; S324. Obtain the decision result based on the comprehensive impact evaluation.
10. A parameter optimization system for quantitatively adding pharmaceutical peptides based on machine learning, characterized in that, Including: A chemical rule knowledge base construction module, which is used to obtain the original data associated with the pharmaceutical peptide synthesis field from multiple specific channels, and use a data fusion algorithm to fuse the obtained original data to obtain a chemical rule knowledge base; A data acquisition and parameter correction warning module, which is used to obtain the structural characteristics of the target peptide, obtain the quantitative addition parameter combination based on the structural characteristics of the target peptide and the parameters input by the data-driven model, and verify the obtained quantitative addition parameter combination according to the chemical rule knowledge base to identify the verification results that do not conform to the chemical rules in the chemical rule knowledge base or may cause side reactions, and correct or generate a warning for the quantitative addition parameter combination in combination with the verification results to obtain the corrected or caution-attached quantitative addition parameter combination; A decision basis generation module, which is used to generate a decision result based on the corrected or caution-attached quantitative addition parameter combination; A feedback and calibration module, which is used to obtain feedback information based on the decision result and, based on the obtained feedback information, update the chemical rule knowledge base or adjust the preset parameters in the data-driven model, where the feedback information includes the parameter adjustment suggested by the user or the chemical principle supporting their suggestion.
Citation Information
Patent Citations
Simulation decision-making method combined with dynamic knowledge graph
CN118820487A
Drug combination risk management system based on knowledge graph
CN118888078A
Method for designing biochemical experiment based on artificial intelligence and man-machine interaction system
CN119127986A
Medical knowledge base verification method and system, medium and equipment
CN119271651A
Intelligent assistant system and method based on AI technology
CN119443292A