A method, device and storage medium for optimizing a drug synthesis experiment
By establishing a compound database and optimizing reaction condition parameters, the optimal drug synthesis route was obtained, and the problem of difficulty in optimizing the process route for traditional Chinese medicine synthesis in drug research and development was solved, which improved R&D efficiency and reduced costs.
Patent Information
- Application Number
- CN202111673458.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-31
AI Technical Summary
There are problems such as long R&D cycle, low result rate and high cost in the drug development process. Especially when looking for the best drug synthesis process route, a large number of repeated experiments are required, resulting in wasting time and cost.
By establishing a compound database, including basic data of chemical structural formulas and basic reaction formulas, selectable intermediate reaction step information can be obtained, drug synthesis paths can be generated, and optimal drug synthesis routes can be obtained by optimizing reaction condition parameters and catalytic condition parameters.
This method effectively reduces the number of repeated times of drug synthesis experiments, improves work efficiency, shortens the R&D cycle, reduces the consumption of reagents and raw materials, and reduces R&D costs.
Smart Images

Figure CN114388070B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pharmaceutical informatization, and particularly to a method, device and storage medium for optimizing drug synthesis experiments. Background Art
[0002] As is well known, drug R & D is a long process, facing the dilemmas of long R & D cycle, low R & D success rate and high R & D cost. The synthetic process route of drugs is the most important link in the drug R & D process. In the process of researching generic drugs, in order to obtain samples with the same parameters as the reference substance, multiple chemical reactions are required. To obtain the same drug efficacy as the reference preparation, it is often necessary to continuously repeat synthesis and analysis verification to find the final appropriate drug synthesis process route, laying a foundation for subsequent process scale-up. In the existing R & D work of pharmaceutical enterprises, to obtain the optimal drug synthesis process route, it still relies on researchers to conduct multiple repeated experiments to find the optimal solution of the process route. It is difficult to guarantee the R & D time and sample quality, and a large amount of reagents and reactant costs are required. Summary of the Invention
[0003] In view of the deficiencies in the prior art, the present invention provides a method for optimizing drug synthesis experiments, including the following steps:
[0004] S1, establish a compound database, where the compound database includes basic chemical structure data and basic reaction formula data, and the basic reaction formula data includes, but is not limited to, main and side reactants, main and side products, reaction condition parameters and catalytic condition parameters;
[0005] S2, obtain optional intermediate reaction step information from the compound database according to the initial reactant and final product information, and generate a drug synthesis path composed of multiple chemically related reactions before and after according to the selected intermediate step information, other reactant information, reaction condition parameters, and catalytic condition parameters input;
[0006] S3, change the reaction condition parameters or catalytic condition parameters of each step in each drug synthesis path, and sequentially select the yields of the drug synthesis path under different parameters to obtain the drug synthesis path with the optimal yield and the corresponding reaction condition parameters and catalytic condition parameters to form the optimal drug synthesis route;
[0007] S4, conduct a physical experiment according to the obtained optimal drug synthesis route, compare the output and yield of the physical experiment with the theoretical output and theoretical yield in the optimal synthesis route, and update the theoretical yield of the relevant reaction formula in the compound database according to the deviation between the actual yield and the theoretical yield.
[0008] Preferably, step S3 specifically includes: dividing multiple groups of experiments with the main and secondary reactants, main and secondary products, reaction condition parameters, and catalytic condition parameters as condition factors, setting the minimum granularity to 1 equivalent, the number of samples to N, dividing each factor into N parts, and combining the conditions through the enumeration method. Then, taking one condition factor as a fixed value, sorting and combining the other condition factors to form samples.
[0009] Preferably, step S3 further includes: based on a preset factor range, optimization objective function, and compound database, using a genetic algorithm to calculate through population iteration to obtain a local optimal solution, specifically including:
[0010] Taking the product yield N and the yield d as the representation forms of the genetic algorithm, using a {0,1} binary string to represent the population individuals, and establishing a mapping relationship with the gene form;
[0011] Obtaining the binary string encoded in the above steps and randomly generating an initial population;
[0012] Using the fitness function Fit[f(N,d)] = -f(N,d), randomly selecting a group of individuals from the population, taking the best among them as the parent individual, and repeating the above operation multiple times until the selection is completed;
[0013] Crossing individual individuals, with the crossing range being [1,num], where num is the number of variables, and exchanging variables with each other until the crossing point does not exceed the limit.
[0014] Preferably, step S3 further includes: designing a corresponding experimental plan and conducting a physical experiment according to the obtained optimal synthesis route, comparing the output and yield of the physical experiment with the theoretical output and theoretical yield in the optimal synthesis route. If the deviation is greater than the preset value, separately conducting physical experiments on each level of chemical reaction in the optimal synthesis route and replacing the theoretical yield of the corresponding reaction formula in the compound database with the actual yield of this level of reaction. If the deviation is less than the preset value, entering the actual yield as the theoretical yield corresponding to the synthesis route into the corresponding synthesis route data group in the compound database.
[0015] Preferably, the step S2 specifically includes: retrieving a corresponding reaction formula from the compound database according to the initial reactant information, reaction condition parameters, and catalytic condition parameters, and obtaining the main product and by-products of this stage reaction; retrieving each reaction formula in the reactant database according to the main product information of the previous stage reaction, using this main product as the main reactant, traversing each reaction formula starting from this main reactant, and querying whether there is a reaction formula whose product is the chemical formula of the required synthetic drug. If so, enter the basic reaction formula data of this reaction formula into the alternative reaction formula data for selection; if not, continue to traverse all the second-stage reaction formulas with the main products of all the current-stage reaction formulas as reactants, and query whether there is a reaction formula whose product is the chemical formula of the synthetic drug. If so, enter the basic reaction formula data of this multi-stage reaction formula into the alternative reaction formula data for selection; otherwise, continue to traverse the third stage until there are no subsequent reaction formulas or the system preset level is reached.
[0016] The present invention also discloses a drug synthesis experiment optimization system, including: a database setting module for establishing a compound database, where the compound database includes chemical structural formula basic data and basic reaction formula data, and the basic reaction formula data includes but is not limited to main and by-reactants, main and by-products, reaction condition parameters, and catalytic condition parameters; a synthesis path acquisition module for obtaining optional intermediate reaction step information from the compound database according to the initial reactant and final product information, and generating a drug synthesis path composed of multiple chemically related reaction steps before and after according to the selected intermediate step information and the input other reactant information, reaction condition parameters, and catalytic condition parameters; a path optimization module for changing the reaction condition parameters or catalytic condition parameters of each step in each drug synthesis path, sequentially selecting the yields of the drug synthesis path under different parameters, obtaining the drug synthesis path with the optimal yield and the corresponding reaction condition parameters and catalytic condition parameters to form an optimal drug synthesis route; a verification module for conducting a physical experiment according to the obtained optimal drug synthesis route, comparing the output and yield of the physical experiment with the theoretical output and theoretical yield in the optimal synthesis route, and updating the theoretical yield of the relevant reaction formula in the compound database according to the deviation between the actual yield and the theoretical yield.
[0017] Preferably, the path optimization module is configured to divide multiple groups of experiments with the main and by-reactants, main and by-products, reaction condition parameters, and catalytic condition parameters as condition factors, set the minimum granularity to 1 equivalent, the number of samples to N, divide each factor into N parts, and combine each condition by the enumeration method, and form samples by sorting and combining other condition factors with one condition factor as a fixed value.
[0018] Preferably, the path optimization module is further configured to obtain a local optimal solution by iterative calculation of a population using a genetic algorithm based on a preset factor range, an optimization objective function, and a compound database, specifically including: using the product yield N and the yield d as the representation forms of the genetic algorithm, using a {0,1} binary string to represent an individual in the population, and establishing a mapping relationship with the gene form; obtaining the binary string encoded in the above steps, and randomly generating an initial population; using a fitness function Fit[f(N,d)] = -f(N,d), randomly selecting a group of individuals from the population, and taking the best of them as the parent individual, repeating the above operation multiple times until the selection is completed; crossing individual individuals, with the crossing range being [1,num], where num is the number of variables, and exchanging variables with each other until the crossing point does not exceed the boundary.
[0019] The present invention also discloses a drug synthesis experiment optimization device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, each step of the above method is implemented.
[0020] The present invention also discloses a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, each step of the above method is implemented.
[0021] The drug synthesis experiment optimization method, device, and storage medium disclosed in the present invention establish a compound database containing basic data of each chemical structure formula and basic reaction formula data by integrating a chemical structure formula editor, a docking electronic experiment record system, and a third-party compound database. Based on the compound and reaction formula data obtained previously in the compound database, a synthetic drug process route is constructed, and multiple groups of optimizations are performed by setting reaction factor parameter ranges, optimization objects, boundary conditions, etc., thereby completing the modeling and optimization of the drug synthesis process route. At the same time, through physical experiments on the generated optimal drug synthesis route, the yield and productivity of the physical experiment are compared with the theoretical yield and theoretical productivity in the optimal synthesis route, and the theoretical productivity of the relevant reaction formula in the compound database is updated according to the deviation between the actual productivity and the theoretical productivity, thereby continuously optimizing the accuracy of the system. It solves the problem that drug researchers need to spend a lot of time in a large number of repeated experiments during the research process, and consume a large amount of reagents and raw materials during process improvement, resulting in high R & D costs, provides a strong guarantee for process scale-up and optimization, and effectively improves work efficiency and scientific research output.
[0022] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0024] Figure 1 It is a schematic flowchart of an optimized method for a drug synthesis experiment disclosed in an embodiment.
[0025] Figure 2 It is a schematic diagram of the principle of a compound database disclosed in an embodiment.
[0026] Figure 3 It is a specific flowchart of step S2 disclosed in an embodiment.
[0027] Figure 4 It is a specific flowchart of step S22 disclosed in an embodiment.
[0028] Figure 5 It is another specific flowchart of step S22 disclosed in an embodiment.
[0029] Figure 6 It is another specific flowchart of step S22 disclosed in an embodiment.
[0030] Figure 7 It is another specific flowchart of step S22 disclosed in an embodiment.
[0031] Figure 8 It is a flowchart of the genetic algorithm of step S3 disclosed in an embodiment. Detailed implementation manners
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0033] In the present invention, unless otherwise clearly defined and limited, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention pertains. The terms "first", "second", and similar terms used in the specification and claims of this patent application for the present invention do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, terms such as "a" or "one" do not denote a quantity limitation, but mean that there is at least one.
[0034] In the actual drug R & D process, a lot of detection equipment is used in the analysis process, generating a large amount of data. For example, when an experimenter conducts chromatographic analysis on a drug, there will be original data records on the equipment. However, equipment such as chromatographs works continuously, and samples of many projects need to be detected later. Therefore, the data needs to be exported to Excel or PDF files to the experimenter's own computer for data processing and calculation. This will generate many PDF files, and the experimenter needs to distinguish them according to the numbers of projects, equipment, samples, etc., and process each file one by one. Due to compliance requirements, the PDF format will ultimately be selected for transmission, and multiple files need to be copied separately by the experimenter, often resulting in problems such as transcription errors and missed copying. This patent discloses a method for automatically analyzing and generating data records, which can automatically obtain PDF files, upload them in batches, parse them in batches, automatically extract the table content in the PDF files, and copy and paste it into other systems for experimental record processing as needed, such as an electronic experimental record system. After extracting and editable processing of the table content in the PDF files, it participates in subsequent experimental data processing, which can greatly improve the data record efficiency, accuracy, reduce the workload of transcription, and avoid manual transcription errors.
[0035] The design of the drug synthesis process model includes information such as the experimental equipment, materials, methods, reference standards, and experimental steps required for synthesis. During the research of generic drugs, in order to obtain samples with the same parameters as the reference substance, there will be multiple-step reactions. The product of this reaction will be used as the reactant for the next reaction, and the most suitable synthesis process model can be obtained through research on multiple process routes. The general experimental process of drug R & D is divided into four major steps, including experimental scheme design, experimental execution, experimental result analysis, and summary and reporting. According to the relevant information obtained from preliminary research and queries, such as pharmacopoeias, literature, historical data, and external relevant databases, the experimental scheme design is output, and experiments are carried out according to the experimental scheme design, including experimental conditions and experimental process records. Subsequently, experimental result analysis is carried out based on the original data, calculation results, review records, comments, and recheck results of the experimental records, and the final experimental report is output. Based on the above experimental process, the present invention discloses a method for optimizing drug synthesis experiments, as shown in the appendix Figure 1 shown, including the following steps:
[0036] Step S1, establish a compound database, the compound database includes basic chemical structure data and basic reaction formula data, and the basic reaction formula data includes but is not limited to main and side reactants, main and side products, reaction condition parameters, and catalytic condition parameters.
[0037] In this embodiment, as shown in the appendix Figure 2As shown, the compound database of this system is established mainly through three ways: docking with an external compound database, self-building by MarvinJS, and importing ELN experimental record data. This step S1 may specifically include the following content.
[0038] Step S11, docking with an external compound database, and obtaining the basic data of the chemical structure of each compound by searching by CAS number and / or structural formula. Dock with an external compound database, for example, PubChem: a free database that stores small molecule chemical structures and their biological activity information. Similar data sources include PharmaPendium, Animal TFDB, etc. After the interface is docked, the corresponding compound is retrieved by CAS number, structural formula, etc.
[0039] Step S12, creating a required compound that does not exist in the database through an editor, assigning it the properties of reactants, catalysts and / or products, generating a corresponding CAS number and Smiles string, and saving it to the compound database.
[0040] Specifically, you can use the MarvinJS structural editor to manually create compound structures and generate corresponding CAS numbers and Smiles strings. Smile is a simplified molecular language used to input and represent linear symbols for molecular reactions. It is an ASCII code, and the Smiles structure can be used to store chemical information and chemical information. At the same time, you can draw a variety of chemical reaction formulas, including reactants, catalysts, and products. In addition, you can reference the system's compound data in the editor, and save the manually created structural formulas, products, etc. to the database for storage, which is convenient for subsequent optimization and analysis.
[0041] Chemical formulas created using the MarvinJS structural formula editor tool do not require additional format conversion. Structural formulas imported from other external editors can be imported through XML, MOL and other format files, and need to be converted to formats, generally smiles format strings, or unique identification content such as Inch key. During this process, the structure-activity relationship will be checked, such as the composition of common rings, functional groups, groups, and repeating groups. It is very helpful to correct errors such as loss of reaction conditions and reaction relationships in reaction formulas due to import format conversion. The system's own compound database can be directly formed through the import function, compound registration, and reaction formula registration.
[0042] Step S13, obtaining the experimental record data in the electronic experimental record system, performing data structuring processing on the experimental record data, splitting and classifying the data, and classifying and reorganizing the data according to the same yield range, the same reaction type, the same reactant and / or the same product according to preset rules.
[0043] The electronic laboratory notebook (ELN) system is used to facilitate researchers in recording experimental data in a timely manner. Its experimental record module includes forms such as rich text, structural formula, spreadsheet, calculation formula, attachment, experimental instrument, and experimental material. The data formats are basically in standard string, number, etc. formats. In this step, various ELN systems can be connected through the background interface. After obtaining the required electronic experimental record data, it is subjected to data structuring, and then the front end performs splitting, classification, and recombination display. The splitting operation can be marked according to data types such as strings and texts. The classification operation can be classified and archived according to the same type of data and the same type of tags. The recombination operation can automatically recombine various data types according to the preset rules of this system to form a module that meets the requirements of this patent. The preset rules include the following categories: reaction structure, reaction yield, reaction solvent, reagent catalyst, reaction time, reaction temperature, reaction pressure, reaction type, etc. Grouping reactions with the same yield range, the same reaction type, the same substrate, the same product, etc. together is beneficial for pharmaceutical laboratory technicians to evaluate.
[0044] In this embodiment, step S1 may further include: classifying the types of compounds in the compound database in different ways. Classify according to the physical and chemical properties of the compounds, and the physical and chemical properties of the compounds include but are not limited to high boiling point, medium boiling point, and low boiling point. Classify according to the catalytic type, and the catalytic type includes but is not limited to heterogeneous catalysis and biocatalysis. At the same time, generate an identity number for the compounds in the compound database, and the identity number includes a classification code, time information, and a serial number. Specifically, classifying according to the types of compounds includes classifying according to the physical and chemical properties of the substances, such as high boiling point, medium boiling point, and low boiling point; classifying according to the physical state of the substances, such as attributes like liquid state and solid state; classifying according to the catalytic type including heterogeneous catalysis, biocatalysis, etc. For each type, assign a unique code, and automatically generate a unique number for the compounds under the classification in the system. The numbering rule can be classification code + year, month, day + number of digits of the serial number, and reset the serial number by day to avoid overflow of the serial number. Subsequently, when forming a process route in combination, each reactant and catalyst has a unique code that can be associated, facilitating reference pairing and query.
[0045] In this embodiment, the compound database aggregates compound data, structural formulas, reaction formula data, etc. Through the above steps, the basic part in the drug synthesis experiment optimization method, that is, the construction of the compound database, is realized, providing data support for the development of subsequent steps.
[0046] Step S14, perform a material balance calculation on the chemical reaction process in the compound database, and enter the product yield or production rate in each chemical reaction into the corresponding reaction formula in the compound database.
[0047] In this embodiment, when compound feeding is carried out, the calculation logic formula for automatic material balance is as follows:
[0048] Field calculation formula for each material: Theoretical feeding amount = MW * molar amount; Actual feeding amount = Theoretical feeding amount / purity; Actual feeding amount = VOL * density; Molarity = molar amount / volume.
[0049] Logic between materials: Molar amount of B / Molar amount of A = Eq of B / Eq of A; The logic for the molar amount of the product is the same.
[0050] This step S14 may further include: The compound database judges the product yield, yield and catalytic condition parameters corresponding to each chemical reaction formula obtained from the electronic experiment record system. If it is greater than the preset range, an alarm prompt is given; if it is within the preset range, the input data is compared with the theoretical value of this data, and the parameters are adjusted according to a preset ratio until the relative average deviation of the adjusted parameter relative to the previous adjusted value is lower than the preset value.
[0051] Specifically, the specific processes for input value standard threshold setting, early warning and automatic calibration are as follows: In the basic setting center, the nominal units, upper and lower limits of the design range of parameters such as reactants and catalysts can be configured. When the experimenter fills in the actual process, or the value calculated according to the chemical reaction formula exceeds the design range, there will be prompts such as high - light early warning. For the calculated results, this system is designed with a set of numerical feedback compensation mechanism. For the deviation F between the result and the theoretical design, the design range of the parameter is adjusted accordingly in proportion. Each time it is adjusted, the relative average deviation F compared with the previous time is calculated until it approaches a stable value, or within the actually specified deviation range, that is, the range acceptable to the designer. In the subsequent continuously enriched reactions and massive data, enough samples are provided for this compensation mechanism to make it reach an almost perfect level
[0052] This step S14 further includes judging the input reactant attribute information. If the input reactant attribute information is the first attribute, and the first attribute is molar amount, actual feeding amount, theoretical feeding amount or purity, then enter the first - attribute acquisition mode and abort the input of the second - attribute information. The first - attribute acquisition mode includes:
[0053] If the input is molar amount, the theoretical feeding amount is automatically calculated, and when there is a purity value, the actual feeding amount is automatically calculated, and when there is no purity value, the actual feeding amount is calculated according to a purity of 100%.
[0054] If the input is theoretical feeding amount, the molar amount is automatically calculated, and when there is a purity value, the actual feeding amount is automatically calculated, and when there is no purity value, the actual feeding amount is calculated according to a purity of 100%.
[0055] If the input is purity, the actual feeding amount is automatically calculated when the theoretical feeding amount exists, or the theoretical feeding amount is automatically calculated when the actual feeding amount exists. If both the theoretical feeding amount and the actual feeding amount exist, the theoretical feeding amount is changed according to the purity.
[0056] If the input is the actual feeding amount, the theoretical feeding amount is calculated when the purity exists, or the purity is automatically calculated when the theoretical feeding amount exists. If both the purity and the theoretical feeding amount exist, the theoretical feeding amount is changed according to the purity. If neither the purity nor the theoretical feeding amount exists, the theoretical feeding amount is calculated according to 100% purity.
[0057] If the input is the theoretical feeding amount, the actual feeding amount is calculated when the purity exists, or the purity is automatically calculated when the actual feeding amount exists. If both the purity and the actual feeding amount exist, the actual feeding amount is changed according to the purity. If neither the purity nor the actual feeding amount exists, the actual feeding amount is calculated according to 100% purity.
[0058] Specifically, when the manually input fields are molar amount, actual feeding amount, theoretical feeding amount, and purity, it is the mass mode. In this mode, if there is density information in the volume field, it is automatically calculated; if there is no density, it is cleared. Specifically, when the molar amount is input, the theoretical feeding amount is automatically calculated because the MW remains unchanged, and the actual feeding amount is automatically calculated according to the purity. If the purity information column is empty, it is calculated according to 100% purity. When the theoretical feeding amount is input, the molar amount is automatically calculated, and the actual feeding amount is automatically calculated according to the purity. When the purity is input, if there is a theoretical feeding amount, the actual feeding amount is automatically calculated; if there is an actual feeding amount, the theoretical feeding amount is automatically calculated; if both exist, the actual feeding amount remains unchanged; if both are empty, no processing is performed. When the actual feeding amount is input, if there is purity, the theoretical feeding amount is calculated; if there is a theoretical feeding amount, the purity is automatically calculated; if both exist, the purity remains unchanged; if both are empty, the theoretical feeding amount is calculated according to 100% purity. When the theoretical feeding amount is input, if there is an actual feeding amount, the purity is automatically calculated; if there is purity, the actual feeding amount is automatically calculated; if both exist, the purity remains unchanged; if both are empty, the actual feeding amount is calculated according to 100% purity.
[0059] In this embodiment, step S2 further includes judging the input reactant attribute information. If the input reactant attribute information is the second attribute, enter the second attribute acquisition mode and abort the input of the first attribute information. The second attribute is volume, volume molar concentration, or density. The second attribute acquisition mode includes: if the input is volume, the volume molar concentration is automatically calculated, and when there is a density value, the actual feeding amount is automatically calculated, and the theoretical feeding amount and molar amount are calculated according to the purity value. When there is no density value, the stored actual feeding amount value is cleared; if the input is density, the actual feeding amount is automatically calculated when there is a volume value, and the theoretical feeding amount and molar amount are calculated according to the purity value.
[0060] Specifically, when the input fields are volume, molar concentration, or density, it is in volume mode. In this mode, the theoretical feeding amount is cleared, and the actual feeding amount must be input with density for calculation. When volume is input, the molar concentration is automatically calculated. If there is density, the actual feeding amount is automatically calculated, and then the theoretical feeding amount and molar amount are calculated according to purity subsequently. If there is no density, the actual feeding amount is cleared. When density is input, if the volume is not empty, the actual feeding amount is automatically calculated, and then the theoretical feeding amount and molar amount are calculated according to purity subsequently. If the volume is empty, when the actual feeding amount is not empty, the volume is automatically calculated, and if the actual feeding amount is also empty, no processing is performed.
[0061] Step S2: Obtain the information of each optional intermediate reaction step from the compound database according to the initial reactant and final product information. Generate a drug synthesis route composed of a series of chemically related reactions before and after according to the selected intermediate step information, the input information of other reactants, reaction condition parameters, and catalytic condition parameters.
[0062] In the previous steps, a basic compound database has been established, which already contains basic data of various compound structural formulas and basic reaction formulas, that is, it contains all objects participating in chemical reactions: main and side reactants, reaction condition parameters, catalytic condition parameters, main and side products, etc. Next, the next reaction operation can be quickly created, and the product of the previous step is directly used as the reactant of the next step, and they are associated through the compound identity number. At the same time, the system can directly recommend a variety of suitable reactants, reaction conditions, and catalytic conditions according to the reaction rules set in the library. Of course, manual intervention can still be performed at this time, and some condition parameters or reactants can be added or deleted as needed. For example, to meet the set goal, it can be automatically constructed by the system according to certain rules. For example, according to factors such as the physical and chemical properties, thermodynamic characteristics, crystal structure, separation and purification methods or conditions of the reactants, the reactants can also be selected manually and then the route can be reorganized.
[0063] In step S2, the system can provide a table such as a drug synthesis process feeding table to display the specific information of each step in the drug synthesis process route in tabular form. For example, the main and side reactants, main and side products, catalytic condition parameters, and reaction condition parameters of the primary reactions in the synthesis process route are entered and presented in units of rows or columns. The arrangement order of each row or column can be used as the sequence before and after the reactions at all levels of the synthesis process route. In this embodiment, as shown in the appendix Figure 3 shown, this step can specifically include the following content.
[0064] S21. Obtain the main and side reactants and various reaction parameter information entered in the cell corresponding to the initial reaction in the drug synthesis process feeding table, retrieve the corresponding reaction formula from the compound database, and input the main product and by-products of this step into the corresponding product cells of this step. The reaction parameter information includes reaction condition parameters and catalytic condition parameters.
[0065] In this embodiment, the drug synthesis process feeding table is configured to arrange the chemical reaction information at all levels before and after the drug synthesis process in sequence from top to bottom in each row of the table according to the sequence before and after the reaction. The main and side reactant information, main and side product information, and reaction parameter information of each level of reaction are correspondingly arranged in the cells of the same row.
[0066] S22. Obtain the main product or by-product in the previous reaction as the reactant in this reaction and enter it into the reactant cell of this reaction, and generate the main and side products of this reaction according to the selected other reactants and reaction parameter information and input them into the product cells of this level.
[0067] In the previous steps, a basic compound database has been established, which already contains various compound structural formula basic data and basic reaction formulas, that is, it contains all objects participating in chemical reactions: main and side reactants, reaction condition parameters, catalytic condition parameters, main and side products, etc. Next, the next reaction operation can be quickly created, directly using the product of the previous step as the reactant of the next step, and associating through the compound identity number. At the same time, the system can directly recommend a variety of suitable reactants, reaction conditions, and catalytic conditions according to the reaction rules set in the library. Of course, manual intervention can still be performed at this time to appropriately add or delete some condition parameters or reactants as needed. For example, to meet the set goals, it can be automatically constructed by the system according to certain rules. For example, according to factors such as the physical and chemical properties, thermodynamic characteristics, crystal structure, separation and purification methods or conditions of the reactants, the reactants can also be selected manually and then the route can be reorganized.
[0068] In the synthesis reaction, the product generated in the current experiment can be continued for the next reaction, so the experiment of the next reaction can be quickly created to reduce manual operation. Through the product of the synthesis reaction formula, quickly create the experimental record of the next reaction, create the experimental record under the same experiment, automatically add the structural formula module and the structural formula automatically includes the product generated in the previous experiment, for example, it can include information such as reactants, molecular weight mw, molecular formula, cas number, batch number, source, etc., and realize the total synthesis route through the binding of each step reaction formula relationship.
[0069] In this embodiment, as shown in the appendix Figure 4 It is shown that step S22 specifically includes:
[0070] Step S221: Obtain at least one product information from the previous - level reaction as the main reactant of this level reaction and enter it into the reactant cell of this level.
[0071] Step S222: After detecting an action instruction on the remaining empty reactant cells, obtain the main reactant information in the other already - entered reactant cells of this level reaction, and retrieve the compound database based on the main reactant. Traverse the compound database to obtain various co - reactants that can undergo chemical reactions with the main reactant, and generate an alternative co - reactant group corresponding to the empty reactant cell. Specifically, when detecting that the operator needs to input a co - reactant, obtain the compound information in the other already - entered reactant cells in this row, retrieve the reaction formulas in the compound database that can form reactions with these already - entered compounds, and use the reactants not entered in these reaction formulas as the alternative co - reactant group. These alternative co - reactant groups can be used as the selectable reactants in the drop - down menu of the reactant cell for the operator to choose. In this step, it may also include: if after selecting the compound to be entered into the empty reactant cell, there are still remaining compounds in its corresponding reaction formula that need to participate in the reaction but are not entered in the reactant cells of this row, then insert a new empty reactant cell after the reactant cells of this row, and enter the remaining un - entered reactants into this cell.
[0072] In this embodiment, as shown in the appendix Figure 5 it can also specifically include the following content.
[0073] Step S101: After detecting a whole - row data translation instruction, determine whether there is a compound in the reactant cell of the row to be moved that is the same as at least one compound in the product cell of the row above the target position.
[0074] Step S102: If there is, insert the moved row into the target position of the feeding table and move the original target - position data and the following rows of data down one row; otherwise, reject moving the data of this row.
[0075] Step S103: After inserting the moved row into the target position of the feeding table, determine whether there is a compound in the product cell of this moved row that is the same as at least one compound in the reactant cell of the row below the target position. If there is, complete the response to this row - data translation; otherwise, adjust the information in the reactant cell of the next row, and readjust the reactants and products of the subsequent rows. Thus, drug researchers can quickly adjust the feeding table, conveniently adjust the order of each reaction step in the already - edited compound synthesis path, and automatically detect the association between the adjusted step and the reactions of the previous and subsequent steps, and readjust the reaction steps that cannot be associated.
[0076] In another embodiment, as shown in the appendix Figure 6As shown, step S22 may specifically include the following content.
[0077] Step S201, after detecting the insert row instruction, move the original target position data and the data of the following rows down one row.
[0078] Step S202, obtain the information in each product cell in the row above the target position as the first compound group, and obtain the information in the reactant cell at the original target position as the second compound group.
[0079] Step S203, retrieve the reaction formula information in the compound database, and obtain the associated chemical formulas with at least one compound in the first compound group as the reactant and at least one compound in the second compound group as the product.
[0080] Step S204, use the reactant information corresponding to each retrieved associated chemical formula as the alternative reactant group to be entered into the reactant cell of the inserted row, use the corresponding product information as the alternative product group to be input into the product cell of the inserted row, and after confirming the reactant in the reactant cell, obtain the corresponding product information in the alternative product group and fill it into the product cell of the inserted row.
[0081] The insert instruction for the above feeding table can conveniently insert a new reaction step in the edited compound synthesis path, automatically detect the association between the inserted new reaction step and the front and rear stage steps, and alarm and automatically adjust the reactions that cannot be associated, greatly facilitating the efficiency of the experimenter in adjusting the compound synthesis path.
[0082] In another specific embodiment, as shown in the appendix Figure 7 Step S22 may specifically include:
[0083] Step S301, retrieve the corresponding reaction formula from the compound database according to the initial reactant information, reaction condition parameters, and catalytic condition parameters, and obtain the main product and by-products of this stage of the reaction;
[0084] Step S302, according to the main product information of the previous level reaction, each reaction formula in the reactant database is retrieved, and the main product is used as the main reactant to traverse each reaction formula starting from the main reactant, and it is inquired whether there is a reaction formula whose product is the desired synthetic drug chemical formula. If so, the basic reaction formula data of the reaction formula is entered into the alternative reaction formula data for selection; if not, all the second-level reaction formulas are traversed with the main product of all the above-mentioned current-level reaction formulas as reactants, and it is inquired whether there is a reaction formula whose product is the synthetic drug chemical formula. If so, the basic reaction formula data of the multi-level reaction formula is entered into the alternative reaction formula data for selection; otherwise, the third level is continued to be traversed until there is no subsequent reaction formula or the system preset level is reached. Through the above steps, the experimenter only needs to input the initial reactant and the final drug compound that he hopes to synthesize, and the system can retrieve all possible compound synthesis paths that use the reactant as a reactant in the synthesis path and finally produce the desired drug compound from the chemical database, and provide the experimenter with selection, which effectively improves the work efficiency of the experimenter.
[0085] In this embodiment, the step S22 may further specifically include the following contents:
[0086] Verify the newly added row data in the feed table to obtain the information in the reactant cell and the product cell of the newly added row in the feed table.
[0087] It is determined whether there is a compound in the newly added reactant cell that is identical to at least one product in the product cell in the previous row. If not, the newly added reactant cell and the product cell in the previous row are marked.
[0088] It is determined whether there is a compound identical to at least one reactant in the reactant cell of the next row in the newly added product cell. If not, the newly added product cell and the reactant cell of the next row are marked. In this embodiment, the system can turn off the automatic association function of the compound input of the feeding table, and the reactants and products of all rows can be input by themselves, and the reaction conditions and catalytic conditions can be input by themselves. At this time, it can be verified at the end, and there is a verification link for intermediate reactants and products during verification. If the product of the previous step is not the reactant of the next step, the relevant cells that fail the verification are highlighted. Specifically, these verification rules can also be set, and the selected reactants are the same, and the duplicates are automatically removed, such as the CAS number of the compound is consistent, the structural formula is consistent, and sometimes the same ring system may have the same name but the CAS number and structural formula are inconsistent. The physical and chemical parameters of the selected reactants do not match, and error prompts such as exceeding the standard are given, for example, the boiling point, hydrophilicity, water solubility, solubility, etc. of the reactants. The separation and purification method does not meet the requirements, etc., and prompts are given, which are automatically eliminated or manually intervened. When the substructure retrieval matches, the protection prompts the experimenter to confirm the ring system.
[0089] In the above step S2, the drug synthesis path is intelligently edited by setting the drug synthesis process feeding table. Through the input initial reactant information, the system can automatically query the corresponding matching reaction formulas that the reactants can participate in the compound database, and automatically enter the products of the selected reaction formulas into the product cells corresponding to this level of reaction. At the same time, for subsequent levels of reactions, the initial reactant information can be used to automatically query the compound database and provide optional reactants for this level for the experimenter to select. And it can automatically generate the corresponding adapted reaction condition parameters, catalytic condition parameters, and the products of this level of reaction formula, and automatically complete the entry of the reaction information of this level in the feeding table. Thus, it greatly shortens the configuration time of the feeding table for the drug compound synthesis path, and at the same time can recommend the optimal reaction information for each level based on the process data in the compound database to the synthesis path feeding table, effectively improving work efficiency and scientific research output.
[0090] In step S3, the reaction condition parameters or catalytic condition parameters of each step in each drug synthesis path are changed, and the yields of the drug synthesis paths under different parameters are sequentially selected to obtain the drug synthesis path with the optimal yield and the corresponding reaction condition parameters and catalytic condition parameters to form the optimal drug synthesis route.
[0091] In this embodiment, this step constructs a model of the synthesis process route based on the above compound database with complete chemical structural formula basic data and basic reaction formula data. According to conditions such as the main reactant and catalyst, it automatically calculates situations such as products and yields, and conducts DOE experimental optimization design. Among them, a large number of sample calculation results provided by historical data can be used to inversely obtain the associated rules, obtain the optimal rules based on the calculated optimal results, classify them according to types and form a rule library. Subsequently, for different types of synthesis plans, the corresponding optimal rules can be recommended. In addition, because the range covered by the sample quantity has limitations, it may lead to the problem of local optimal solutions, so it is necessary to calibrate and repair the corresponding rules and standards. The rules and standards are continuously improved through the results. With the continuous enrichment of massive data such as reaction formula data in the future, it provides enough samples for this compensation mechanism to make it approach perfection. Among them, the Elasticsearch professional search engine can be used, and its powerful search function can be used, combined with methods such as scrolling loading of search results, to improve the search efficiency as much as possible and avoid the impact of massive data search on efficiency.
[0092] Furthermore, in this step, an asynchronous iterative calculation method can be adopted, decomposing the calculation process into several sub-units. Each unit participates in one calculation, and all iterations fluctuate within a certain upper and lower limit range, not exceeding the theoretical calibration range. For example, the deviation of any 30 iterative calculations does not exceed 5%. At the same time, by adjusting the number of iterations and the density granularity of the units, the calculation efficiency can also be optimized.
[0093] Based on the initially obtained drug synthesis path model in the previous step S3, the design variables are changed, and the models with a large number of different parameters under different conditions are solved. Through multiple combination methods such as the physical and chemical properties, catalyst type, ring system, main and side reaction types, crystal structure, and thermodynamic properties described above, different parameters are selected in sequence. This can effectively assist researchers in exploring a wide design space and lay a foundation for SO (Sequential Optimization, automatic sequential optimization). By performing gradient modeling and solving for multiple variables of the original model, the optimal solution can be accurately obtained. The sequential optimization method can automatically select the optimal solution according to the constraint conditions.
[0094] The specific steps of step S3 include: dividing multiple groups of experiments with the main and side reactants, main and side products, reaction condition parameters, and catalytic condition parameters as condition factors, setting the minimum granularity to 1 equivalent, the number of samples to N, dividing each factor into N parts, and combining the conditions through the enumeration method. Each condition factor is set as a fixed value, and the other condition factors are sorted and combined to form samples.
[0095] Step S3 also includes obtaining a local optimal solution by using the genetic algorithm through population iterative calculation based on a preset factor range, optimization objective function, and compound database. As shown in the appendix Figure 8 shown, specifically including:
[0096] Taking the product yield N and the recovery rate d as the manifestation forms of the genetic algorithm, using a {0,1} binary string to represent the population individuals, and establishing a mapping relationship with the gene form.
[0097] Obtain the binary string encoded in the above steps and randomly generate an initial population.
[0098] Adopt the fitness function Fit[f(N,d)] = -f(N,d), randomly select a group of individuals from the population, and take the best among them as the parent individual. Repeat the above operation multiple times until the selection is completed.
[0099] Cross a single individual, with the cross range being [1, num], where num is the number of variables, and exchange variables with each other until the crossover point does not exceed the boundary.
[0100] In this embodiment, the entire optimization step process mainly includes the following four parts:
[0101] Step S41, input setting: variable definition of the synthesis experiment plan, including factor definitions and range settings such as reactants, catalysts, temperature and humidity, and main reactions.
[0102] Step S42, Output Setting: The reactive product includes the output factor type, judgment baseline setting, theoretical yield, yield, and determination of the optimized objective function.
[0103] Step S43, Multiple Sets of Experimental Schemes: According to the reaction conditions and the set range of factors, multiple sets of experiments are divided. First, set the basic conditions of the synthesis plan, factors such as reactants, catalysts, environmental conditions, main and side reactions, etc. The minimum granularity is set to 1 equivalent, and the number of samples is N. Then each factor is divided into N parts.
[0104] The enumeration method combines various conditions and matches them quickly. First, sort the reactants, conditions, catalysts, etc. according to the formula conditions. Later, some actually required factors can be artificially intervened and added. Before starting the calculation, the content of the subdivided combination units can be understood.
[0105] Step S44, Determination of the Optimal Solution of the Reaction Output and the Total Synthesis Route: Based on the previously set conditions, factor range, optimized objective function, and basic database, the final step is carried out next. Utilizing the characteristics of the genetic algorithm, through population iterative calculation, local optimal solution optimization can be performed. The process of the genetic algorithm optimization method for the specific objective function is as follows:
[0106] Multivariable Encoding: Taking the product yield N and the yield d as the representation form of the genetic algorithm, establishing the mapping relationship with the gene form is encoding. In this embodiment, a {0,1} binary string is used to represent the population individuals.
[0107] Planning the Initial Population: From the binary strings encoded in the multivariable encoding step, a random initial population is generated. In this embodiment, the initial population size is initially set to 20.
[0108] Fitness: The fitness function is formed by converting the objective function. The fitness function adopted in this embodiment is: Fit[f(N,d)] = -f(N,d).
[0109] Individual Selection of the Best: Randomly select a group of individuals from the population, and take the best among them as the parent individual. Repeat this operation multiple times until the selection is completed.
[0110] Crossing Individuals: In this embodiment, single individuals are crossed. The crossing range is [1, num], where num is the number of variables. Exchange variables with each other until the crossing point, without exceeding the boundary.
[0111] Taking two parent individuals with 6 - bit variables as an example, as follows:
[0112] Parent Individual 1: 0 1 1 1 0 1
[0113] Parent Individual 2: 0 1 1 0 1 1
[0114] Assume that the crossover point is at position 3. Then the two offspring individuals generated after the crossover operation are as follows:
[0115] Offspring individual 1: 0 1 1 0 1 1
[0116] Offspring individual 2: 1 0 1 1 0 1
[0117] Mutation operation: The crossover operation of the population causes mutations in the offspring individuals. Assume that the mutation rate is 0.01, and randomly change the variable values of the individuals. The fourth position of an individual mutates as shown below.
[0118] Before mutation: 1 1 0 0 0 0
[0119] After mutation: 1 1 0 1 0 0
[0120] Convergence judgment: In this embodiment, the convergence criterion is that the best and the worst individuals reach consistency. At this time, the optimization ends. Otherwise, it transfers to the aforementioned fitness step to start the next round of loop. During the iteration process, use some excellent offspring individuals to replace some parent individuals in the new population, which can improve the optimization performance of the genetic algorithm.
[0121] By using the genetic algorithm to simulate natural selection, the problem can be solved through parallel computing, only needing to evaluate a small number of possibilities simultaneously.
[0122] In this embodiment, step S3 further includes: designing a corresponding experimental plan and conducting a physical experiment according to the obtained optimal synthesis route, comparing the yield and productivity of the physical experiment with the theoretical yield and theoretical productivity in the optimal synthesis route. If the deviation is greater than the preset value, conduct individual physical experiments on each level of chemical reaction in the optimal synthesis route and replace the theoretical productivity of the corresponding reaction formula in the compound database with the actual productivity obtained for this level of reaction. If the deviation is less than the preset value, record the actual productivity as the theoretical productivity corresponding to this synthesis route in the corresponding synthesis route data group in the compound database.
[0123] Step S4, conduct a physical experiment according to the obtained optimal drug synthesis route, compare the yield and productivity of the physical experiment with the theoretical yield and theoretical productivity in the optimal synthesis route, and update the theoretical productivity of the relevant reaction formula in the compound database according to the deviation between the actual productivity and the theoretical productivity.
[0124] Specifically, compare the optimization result with the actual test: According to the obtained optimal synthesis reaction route, design an experimental plan, compare the yield and productivity of the actual experiment with the system. The system corrects the subsequent reactions of the same type according to the actual situation by fitting according to the deviation, so as to continuously optimize the accuracy of the system.
[0125] Deviation Analysis: Simulation can be carried out before the implementation of the experimental design plan, during the implementation process, and even after the actual experiment to achieve the functions of design evaluation and actual verification. However, since it is based on theoretical calculations under given conditions and optimization within a local range, there will be deviations from the results of the actual research experiment. The deviation can be reduced by increasing the number of experimental samples and the population size of the genetic algorithm in the system, or even by adding variable factor types. However, this will increase the time required to obtain the optimal solution, and a solution needs to be formulated based on the actual situation.
[0126] The drug synthesis experiment optimization method, device, and storage medium disclosed in the present invention establish a compound database containing basic chemical structure data and basic reaction formula data by integrating a chemical structure editor, a docking electronic experimental record system, and a third-party compound database. Based on the compound and reaction formula data obtained previously in the compound database, a synthetic drug process route is constructed, and multiple groups of optimizations are carried out by setting the parameter range of reaction factors, optimization objects, boundary conditions, etc., thereby completing the modeling and optimization of the drug synthesis process route. At the same time, physical experiments are carried out on the generated optimal drug synthesis route, and the output and yield of the physical experiment are compared with the theoretical output and theoretical yield in the optimal synthesis route. According to the deviation between the actual yield and the theoretical yield, the theoretical yield of the relevant reaction formula in the compound database is updated, thereby continuously optimizing the accuracy of the system. This solves the problem that drug researchers need to spend a lot of time on a large number of repeated experiments during the research process, and consume a large amount of reagents and raw materials during process improvement, resulting in high R & D costs. It provides a strong guarantee for process scale-up and optimization, effectively improving work efficiency and scientific research output. The optimal solution of the synthesis reaction route can be found through this method before formulating the physical experiment plan, providing more possibilities and foresight for drug R & D; reducing the time of actual experiments relying on manpower and materials, improving experimental efficiency, reducing repeated experiments, and greatly shortening the R & D cycle.
[0127] In another embodiment, a drug synthesis experiment optimization system is also disclosed, including: a database setting module for establishing a compound database, where the compound database includes basic chemical structural formula data and basic reaction formula data, and the basic reaction formula data includes, but is not limited to, main and side reactants, main and side products, reaction condition parameters, and catalytic condition parameters; a synthesis path acquisition module for obtaining optional intermediate reaction step information from the compound database according to the initial reactant and final product information, and generating a drug synthesis path composed of multiple chemically related reaction steps before and after according to the selected intermediate step information, the input other reactant information, reaction condition parameters, and catalytic condition parameters; a path optimization module for changing the reaction condition parameters or catalytic condition parameters of each step in each drug synthesis path, successively selecting the yields of the drug synthesis paths under different parameters, and obtaining the drug synthesis path with the optimal yield and the corresponding reaction condition parameters and catalytic condition parameters to form an optimal drug synthesis route; a verification module for conducting a physical experiment according to the obtained optimal drug synthesis route, comparing the output and yield of the physical experiment with the theoretical output and theoretical yield in the optimal synthesis route, and updating the theoretical yield of the relevant reaction formula in the compound database according to the deviation between the actual yield and the theoretical yield.
[0128] In this embodiment, the path optimization module is configured to divide multiple groups of experiments with the main and side reactants, main and side products, reaction condition parameters, and catalytic condition parameters as condition factors, set the minimum granularity to 1 equivalent, the number of samples to N, divide each factor into N parts, and combine the conditions by enumeration method, and respectively form samples by sorting and combining other condition factors with one condition factor as a fixed value.
[0129] In this embodiment, the path optimization module is further configured to obtain a local optimal solution by population iteration calculation using a genetic algorithm based on a preset factor range, optimization objective function, and compound database, specifically including: using the product output N and yield d as the representation forms of the genetic algorithm, using a {0,1} binary string to represent the population individuals, and establishing a mapping relationship with the gene form; obtaining the binary string encoded in the above steps and randomly generating an initial population; using a fitness function, randomly selecting a group of individuals from the population, and taking the best of them as the parent individual, repeating the above operations multiple times until the selection is completed; crossing individual individuals, with the crossing range being [1,num], where num is the number of variables, and exchanging variables with each other until the crossing point does not exceed the limit.
[0130] The specific functions of the above drug synthesis experiment optimization system correspond one by one to the drug synthesis experiment optimization methods disclosed in the previous embodiments, so they will not be described in detail here. For details, reference can be made to the respective embodiments of the drug synthesis experiment optimization methods disclosed above. It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0131] In some other embodiments, a drug synthesis experiment optimization device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements each step of the drug synthesis experiment optimization method described in the above embodiments.
[0132] The drug synthesis experiment optimization device may include, but is not limited to, a processor and a memory. The server may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of the server and does not constitute a limitation on the server device. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the server device may also include input / output devices, network access devices, buses, etc.
[0133] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the server device, connecting various parts of the entire server device through various interfaces and lines.
[0134] The memory can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by invoking the data stored in the memory, the processor implements various functions of the server device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disks, memory, plug-in hard disks, smart media cards (SMC), secure digital (SD) cards, flash cards, at least one magnetic disk storage device, flash memory devices, or other volatile solid-state storage devices.
[0135] If the method for optimizing drug synthesis experiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
[0137] In summary, the above are only the preferred embodiments of the present invention, and all equivalent changes and modifications made in accordance with the scope of the patent application of the present invention shall fall within the scope of the patent of the present invention.
Claims
1. A method for optimizing drug synthesis experiments, characterized in that, it includes the following steps: S1. Establish a compound database, which includes basic chemical structural formula data and basic reaction formula data. The basic reaction formula data includes, but is not limited to, main and side reactants, main and side products, reaction condition parameters, and catalytic condition parameters; S2. Obtain optional intermediate reaction step information from the compound database according to the initial reactant and final product information. According to the selected intermediate step information, other reactant information, reaction condition parameters, and catalytic condition parameters input, generate a drug synthesis path composed of multiple chemically related reactions before and after; S3. Change the reaction condition parameters or catalytic condition parameters of each step in each drug synthesis path, and sequentially select the yields of the drug synthesis paths under different parameters to obtain the drug synthesis path with the optimal yield and the corresponding reaction condition parameters and catalytic condition parameters to form the optimal drug synthesis route; Divide multiple groups of experiments with main and side reactants, main and side products, reaction condition parameters, and catalytic condition parameters as condition factors. Set the minimum granularity to 1 equivalent and the number of samples to N. Divide each factor into N parts, and combine the conditions by enumeration method. Respectively, take one condition factor as a fixed value and sort and combine other condition factors to form samples; The step S3 also includes obtaining a local optimal solution by using a genetic algorithm through population iteration calculation based on a preset factor range, optimization objective function, and compound database, specifically including: Taking the product yield N and the yield d as the manifestation forms of the genetic algorithm, using a {0,1} binary string to represent the population individuals, and establishing a mapping relationship with the gene form; Obtain the binary string encoded in the above steps and randomly generate an initial population; Adopt the fitness function Fit[f(N,d)] = -f(N,d), randomly select a group of individuals from the population, and take the best of them as the parent individual. Repeat the above operation multiple times until the selection is completed; Cross individual individuals, and the cross range is [1,num], where num is the number of variables, and exchange variables with each other until the crossover point does not exceed the boundary; S4. Conduct physical experiments according to the obtained optimal drug synthesis route, compare the output and yield of the physical experiments with the theoretical output and theoretical yield in the optimal synthesis route, and update the theoretical yield of the relevant reaction formula in the compound database according to the deviation between the actual yield and the theoretical yield.
2. The method for optimizing drug synthesis experiments according to claim 1, characterized in that, the step S3 also includes: Design and conduct corresponding experimental plans according to the obtained optimal synthesis route and conduct physical experiments. Compare the output and yield of the physical experiments with the theoretical output and theoretical yield in the optimal synthesis route. If the deviation is greater than the preset value, conduct separate physical experiments on each level of chemical reactions in the optimal synthesis route and replace the theoretical yield of the corresponding reaction formula in the compound database with the actual yield obtained at this level of reaction. If the deviation is less than the preset value, record the actual yield as the theoretical yield corresponding to the synthesis route in the corresponding synthesis route data group in the compound database.
3. The method for optimizing a drug synthesis experiment according to claim 2, wherein, the step S2 specifically includes: retrieving a corresponding reaction formula from the compound database according to the initial reactant information, reaction condition parameters, and catalytic condition parameters, and obtaining the main product and by-products of this level of reaction; retrieving each reaction formula in the reactant database according to the main product information of the previous level of reaction, using the main product as the main reactant, traversing each reaction formula starting from the main reactant, and querying whether there is a reaction formula whose product is the chemical formula of the required synthetic drug. If so, enter the basic reaction formula data of this reaction formula into the alternative reaction formula data for selection; if not, continue to traverse all the second-level reaction formulas with the main products of all the current-level reaction formulas as reactants, and query whether there is a reaction formula whose product is the chemical formula of the synthetic drug. If so, enter the basic reaction formula data of each multi-level reaction formula into the alternative reaction formula data for selection; otherwise, continue to traverse the third level until there are no subsequent reaction formulas or the system preset level is reached.
4. A drug synthesis experiment optimization system, wherein, it includes: a database setting module for establishing a compound database, where the compound database includes basic chemical structure data and basic reaction formula data, and the basic reaction formula data includes but is not limited to main and by-reactants, main and by-products, reaction condition parameters, and catalytic condition parameters; a synthesis path obtaining module for obtaining optional intermediate reaction step information from the compound database according to the initial reactant and final product information, and generating a drug synthesis path composed of multiple chemically related reaction steps before and after according to the selected intermediate step information and the input other reactant information, reaction condition parameters, and catalytic condition parameters; A path optimization module is used to change the reaction condition parameters or catalytic condition parameters of each step in each drug synthesis path, and sequentially select the yields of the drug synthesis paths under different parameters to obtain the drug synthesis path with the optimal yield and the corresponding reaction condition parameters and catalytic condition parameters to form an optimal drug synthesis route; the path optimization module is configured to divide multiple groups of experiments with the main and side reactants, main and side products, reaction condition parameters and catalytic condition parameters as condition factors, set the minimum granularity to 1 equivalent, the number of samples to N, divide each factor into N parts, and combine the conditions by enumeration method, and respectively sort and combine other condition factors with one condition factor as a fixed value to form samples; it is also configured to obtain a local optimal solution by using a genetic algorithm through population iteration calculation based on a preset factor range, optimization objective function and compound database, specifically including: using the product yield N and the yield d as the representation forms of the genetic algorithm, using a {0,1} binary string to represent the population individuals, and establishing a mapping relationship with the gene form; obtaining the binary string encoded in the above steps and randomly generating an initial population; using the fitness function Fit[f(N,d)]=-f(N,d), randomly selecting a group of individuals from the population, and taking the best of them as the parent individual, repeating the above operations multiple times until the selection is completed; crossing individual individuals, with the crossing range being [1,num], where num is the number of variables, and exchanging variables with each other until the crossing point does not exceed the limit. A verification module is used to conduct a physical experiment according to the obtained optimal drug synthesis route, compare the output and yield of the physical experiment with the theoretical output and theoretical yield in the optimal synthesis route, and update the theoretical yield of the relevant reaction formula in the compound database according to the deviation between the actual yield and the theoretical yield.
5. A drug synthesis experiment optimization device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: when the processor executes the computer program, the steps of the method according to any one of claims 1-3 are implemented.
6. A computer-readable storage medium storing a computer program, characterized in that: when the computer program is executed by the processor, the steps of the method according to any one of claims 1-3 are implemented.
Citation Information
Patent Citations
Lead compound discovery and synthesis method based on quantum group intelligent optimization
CN109559786A
Document creation assistance server and document creation assistance method
CN111406273A