A Greedy Algorithm-Based Strain Construction Scheduling Design Method

By optimizing the strain construction schedule using a greedy algorithm and determining the common ancestor strain using integer programming and clustering algorithms, the problem of repetitive gene operations in biofoundries is solved, and efficient and low-cost strain library construction is achieved.

CN116524994BActive Publication Date: 2026-01-30TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310143252.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-20
Publication Date
2026-01-30
Estimated Expiration
2043-02-20

AI Technical Summary

Technical Problem

Existing biofoundries suffer from repetitive operations on common genes and waste of resources when constructing high-breadth and high-depth strain libraries, leading to increased production pressure and costs.

Method used

A multi-strain construction scheduling design method based on a greedy algorithm is adopted. Through iterative clustering and integer programming optimization algorithms, the common ancestor strain and strain parentage are determined, the number of intermediate strains constructed is reduced, and the gene operation sequence is optimized.

Benefits of technology

It significantly reduced the workload and cost of strain construction, reduced the repetitive modification of the same gene targets, and improved the efficiency of strain construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524994B_ABST
    Figure CN116524994B_ABST
Patent Text Reader

Abstract

This invention relates to the field of biotechnology, specifically to optimizing genetic manipulation scheduling, and discloses a multi-strain construction scheduling design method based on a greedy algorithm. This method uses integer programming to iteratively cluster target strains sharing the maximum number of gene target sites, then constructs a common ancestor strain in each cluster subset, and builds individual target strains from the common ancestor strain, forming a strain modification scheduling tree, ultimately resulting in a strain modification scheduling tree. This method significantly reduces redundant modification of the same gene target sites among different strains, thereby reducing the workload and cost of large-scale strain construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, specifically to optimizing genetic operation scheduling, and discloses a multi-strain construction scheduling design method based on a greedy algorithm. Background Technology

[0002] Manual methods of strain development are typically labor-intensive and time-consuming. It was reported that in 2018, developing commercial strains required 100 person-years and $50 million [1]. The scale and efficiency of manual methods limit strain development to a low level, such as sequentially adding a dozen gene modifications to strains to overproduce a specific product. Establishing a strain library that is 1 to 2 orders of magnitude larger than manual methods (high breadth) and successfully conducting large-scale screening is considered an effective approach to strain development [2,3]. Biocasting technology is a rapidly developing technology for large-scale strain development. Biocasting workshops use software to coordinate plate / liquid handling robots and high-throughput analysis equipment to automate the processes required for strain development, including PCR, clone selection, plasmid isolation and insertion, protein expression, and strain performance characterization [7,8].

[0003] Biofoundries support parallel, standardized, and automated genetic operations that accelerate strain construction while improving reproducibility[9], reducing labor costs, and enabling the development of a wide range of strains. For example, Enghiad et al. developed PlasmidMaker, an automated plasmid construction platform that can rapidly and automatically construct 101 plasmids from six different species, ranging in size from 5 to 18 kb, with up to 11 DNA fragments

[10] . Wang et al. developed MACBETH, an automated base editing platform for Corynebacterium glutamicum that can construct genome-scale inactivated libraries of 94 transcription factors at once with 100% success rate and support the construction of thousands of Corynebacterium glutamicum strains per month

[11] . Biofoundries accelerate the design-build-test-learn (DBTL) cycle in synthetic biology[1]

[12] .

[0004] However, the depth of physical processing in the current biofoundry construction process is still relatively low (usually a single gene mutation for each strain). Combining the high-depth strain construction of traditional strains with the high-breadth advantage of biofoundries is the future trend of biofoundry development. Currently, advanced mechanistic models [13–15] and mixed integer algorithms [16–18] can design microbial strains for overproduction of hundreds of products

[19] , providing multiple metabolic engineering strategies for each of them. This requires a biofoundry to be able to construct one hundred (or hundreds) target strains, each of which may carry a dozen or so gene manipulations. This is a high-breadth and high-depth task, thus putting enormous production pressure on biofoundry workshops.

[0005] Different target strains often share common genes with identical operations, especially in central metabolic pathways. Constructing these strains separately requires repeatedly manipulating these common genes, necessitating the creation of numerous intermediate strains and resulting in a significant waste of materials and manpower.

[0006] References:

[0007] [1] E. Marcellin, LK Nielsen, Advances in analytical tools for high throughput strain engineering, Curr. Opin. Biotechnol. 54 (2018) 33–40. https: / / doi.org / 10.1016 / j.copbio.2018.01.027.

[0008] [2] J. Zhang, SD Petersen, T. Radivojevic, A. Ramirez, A. Pérez-Manríquez, E. Abeliuk, BJ Sánchez, Z. Costello, Y. Chen, MJ Fero, HG Martin, J. Nielsen, JD Keasling, MK Jensen, Combining mechanistic and machine learning models for predictive engineering and optimization of tryptophan metabolism, Nat. Commun. 11(2020). https: / / doi.org / 10.1038 / s41467-020-17910-1.

[0009] [3] Y. Li, EOMensah, E. Fordjour, J. Bai, Y. Yang, Z. Bai, Recent advances in high-throughput metabolic engineering: Generation of oligonucleotide-mediated genetic libraries., Biotechnol. Adv. 59 (2022) 107970. https: / / doi.org / 10.1016 / j.biotechadv.2022.107970.

[0010] [4]T.Farzaneh,P.S.Freemont,Biofoundries are a nucleating hub forindustrial translation,Synth.Biol.(2021)1–6.https: / / doi.org / 10.1093 / synbio / ysab013.

[0011] [5]I.Holland,J.A.Davies,Automation in the Life Science ResearchLaboratory,Front.Bioeng.Biotechnol.8(2020)1–18.https: / / doi.org / 10.3389 / fbioe.2020.571777.

[0012] [6]M.HamediRad,R.Chao,S.Weisberg,J.Lian,S.Sinha,H.Zhao,Towards afully automated algorithm driven platform for biosystems design,Nat.Commun.10(2019)1–10.https: / / doi.org / 10.1038 / s41467-019-13189-z.

[0013] [7]S.R.Hughes,T.R.Butt,S.Bartolett,S.B.Riedmuller,P.Farrelly,Designand construction of a first-generation high-throughput integrated roboticmolecular biology platform for bioenergy applications,J.Lab.Autom.16(2011)292–307.https: / / doi.org / 10.1016 / j.jala.2011.04.004.

[0014] [8]J.Zhang,Y.Chen,L.Fu,E.Guo,B.Wang,L.Dai,T.Si,Accelerating strainengineering in biofuel research via build and test automation of syntheticbiology,Curr.Opin.Biotechnol.67(2021)88–98.https: / / doi.org / 10.1016 / j.copbio.2021.01.010.

[0015] [9]M.M.Jessop-Fabre,N.Sonnenschein,Improving reproducibility insynthetic biology,Front.Bioeng.Biotechnol.7(2019)1–6.https: / / doi.org / 10.3389 / fbioe.2019.00018.

[0016]

[10] B.Enghiad,P.Xue,N.Singh,A.G.Boob,C.Shi,V.A.Petrov,R.Liu,S.S.Peri,S.T.Lane,E.D.Gaither,H.Zhao,PlasmidMaker is a versatile,automated,and highthroughput end-to-end platform for plasmid construction,Nat.Commun.13(2022).https: / / doi.org / 10.1038 / s41467-022-30355-y.

[0017]

[11] Y.Wang,Y.Liu,J.Liu,Y.Guo,L.Fan,X.Ni,X.Zheng,M.Wang,P.Zheng,J.Sun,Y.Ma,MACBETH:Multiplex automated Corynebacterium glutamicum base editingmethod,Elsevier Inc.,2018.https: / / doi.org / 10.1016 / j.ymben.2018.02.016.

[0018]

[12] N.Hillson, M.Caddick, Y.Cai, J.A.Carrasco, M.W.Chang, N.C.Curach, D.J.Bell, R.Le Feuvre, D.C.Friedman, X.Fu, N.D.Gold, M.J. MBHolowko,JRJohnson,RAJohnson,JDKeasling,RIKitney,A.Kondo,C.Liu,VJJMartin,F.Menolascina,C.Ogino,NJPatron,M.Pavan,C LPoh,ISPretorius,SJRosser,NSScrutton,M.Storch,H.Tekotte,E.Travnik,CEVickers,WSYew,Y.Yuan,H.Zhao,PSFreemont,Building a global alliance of biofoundries,Nat.Commun.10(2019)1038–1041.https: / / doi.org / 10.1038 / s41467-019-10079-2.

[0019]

[13] JMMonk,CJLloyd,E.Brunk,N.Mih,A.Sastry,Z.King,R.Takeuchi,W.Nomura,Z.Zhang,H.Mori,AMFeist,BOPalsson,iML1515,a knowledgebase thatcomputes Escherichia coli traits,Nat.Biotechnol.35(2017)904–908.https: / / doi.org / 10.1038 / nbt.3956.

[0020]

[14] H.Lu,F.Li,BJSánchez,Z.Zhu,G.Li,I.Domenzain,S. P.M.Anton,D.Lappa,C.Lieven,M.E.Beber,N.Sonnenschein,E.J.Kerkhoven,J.Nielsen,Aconsensus S.cerevisiae metabolicmodel Yeast8 and its ecosystem forcomprehensively probing cellular metabolism,Nat.Commun.10(2019).https: / / doi.org / 10.1038 / s41467-019-11581-3.

[0021]

[15] Y.Zhang,J.Cai,X.Shang,B.Wang,S.Liu,X.Chai,T.Tan,Y.Zhang,T.Wen,Anew genome-scale metabolic model of Corynebacterium glutamicum and itsapplication,Biotechnol.Biofuels.10(2017).https: / / doi.org / 10.1186 / s13068-017-0856-3.

[0022]

[16] K.Jensen,V.Broeken,A.S.L.Hansen,N.Sonnenschein,M.J. OptCouple:Joint simulation of gene knockouts,insertions and mediummodifications for prediction of growth-coupled strain designs,Metab.Eng.Commun.8(2019).https: / / doi.org / 10.1016 / j.mec.2019.e00087.

[0023]

[17] S. Jiang, I. Otero-muras, JRBanga, Y. Wang, N. Krasnogor, OptDesign: Identifying Optimum Design Strategies in Strain Engineering for Biochemical Production, (2021) 1–29. https: / / doi.org / 10.1021 / acssynbio.1c00610.

[0024]

[18] S.Klamt,R.Mahadevan,A.von Kamp,Speeding up the core algorithm for the dual calculation of minimal cut sets in large metabolic networks.,BMCBioinformatics.21(2020)510.https: / / doi.org / 10.1186 / s12859-020-03837-3.

[0025]

[19] A. von Kamp, S. Klamt, Growth-coupled overproduction is feasible for almost all metabolites in five major production organisms., Nat. Commun. 8 (2017) 15956. https: / / doi.org / 10.1038 / ncomms15956. Summary of the Invention

[0026] In the process of constructing a series of strains, this invention considers using intermediate strains that share a common ancestor with different target strains during the construction process. This reduces the total number of intermediate strains required, thereby lowering the cost of strain construction. It requires an optimization algorithm to determine the optimal common ancestor strain and the parentage of the strains for iterative construction in each round. An intuitive strategy for determining the "common ancestor strain" is to cluster the target strains into several subsets based on their shared genetic operations.

[0027] This invention provides a multi-strain construction scheduling design method based on a greedy algorithm, comprising the following steps:

[0028] Step 1: Obtain the genotypes of all starting strains: Obtain the genotypes of each target strain, including all genetic modifications made based on the starting strains;

[0029] Step 2: Use integer programming to iteratively cluster the target strains that share the maximum number of gene target points until all strains are assigned to cluster subsets. This step assigns the target strains to P subsets and determines the common ancestor strain corresponding to each subset.

[0030] Step 3: The common ancestor strains obtained in Step 2 are constructed and scheduled, and the parent-child inheritance order of the modified target points is arranged in descending order of the difficulty of the gene target modification.

[0031] Step 4: Starting from the common ancestral strains, add each target strain and the target modification difference set of the corresponding ancestral strain in descending order of gene target modification difficulty to obtain the construction schedule of the target strains in each subset.

[0032] Preferably, the algorithm used for iterative clustering calculation in the second step is:

[0033] maximize

[0034]

[0035]

[0036] Among them, I s =1 indicates that the target strain s is included in the subS{p} set; otherwise, I s =0;

[0037]

[0038] T g =1 indicates that the gene operation g is included in the gene operation of the common ancestor of the subset; otherwise, T g =0;

[0039]

[0040] W g,s =1 indicates that the common ancestral strain carries the gene. The purpose of operation g is to construct strain s; otherwise, W g,s =0; if the target strain s is not included in S c In the set, i.e., I s =0, then W g,s =0, if the target strain s is contained in S c In the set, i.e., I s =1, W g,s The value of is determined by whether the gene operation g is included in the gene operation of the common ancestor of the subset:

[0041]

[0042] If the operator gene g is not included in the target strain s, then the common ancestor cannot carry the operator g:

[0043]

[0044] This is determined by identifying whether a common ancestor contains the gene manipulation used to construct the target strain. c In the set,

[0045]

[0046] The subset must contain more than 2 strains in total, meaning at least two target strains must share a common ancestor.

[0047]

[0048] Under the constraints described above, the optimization objective is set to maximize the number of shared gene target modifications, resulting in the strain set S under this objective. c .

[0049] Preferably, the third step also includes constructing a scheduling design for each common ancestral strain.

[0050] Furthermore, the fourth step also includes drawing the scheduling tree.

[0051] The present invention also provides a system for parallel construction of target strains based on the method to minimize the total number of gene operations, which includes a data input module, a data processing module and a result output module.

[0052] Specifically, the data processing module executes the algorithm of the second step using a computer; the result output module includes a scheduling tree drawing submodule.

[0053] Optionally, the results may be displayed via a computer screen, a remote terminal, or a mobile terminal.

[0054] The present invention further provides an apparatus for the system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program, the computer program being encoded to implement the system, and preferably further comprising a data input device such as a keyboard, an image recognition device, etc., and a display device such as a screen or a remote result display.

[0055] Preferably, the device exists in the form of a computer package or an Internet platform.

[0056] The method of this invention uses integer programming to find the subset of strains with the most shared genes in the unclustered strain set. This subset is then moved from the unclustered set to the clustered set, and this process is iterated to guide all target strains to be assigned to various subsets. Then, based on the shared genes in each subset, a common ancestor strain is constructed. Finally, starting from each ancestor strain, each target strain is constructed, forming a strain modification scheduling tree, ultimately resulting in a scheduled tree for strain modification. This method significantly reduces the duplication of modification of the same gene targets among different strains, thereby reducing the workload and cost of large-scale strain construction. Attached Figure Description

[0057] Figure 1 A schematic diagram of the scheduling constructed using the method of the present invention for the target strain.

[0058] Figure 2 For the design tasks in Table 1, a scheduling tree is constructed using the strains designed by the method of this invention. Detailed Implementation

[0059] The following describes specific embodiments of the present invention in detail, but this does not constitute a limitation on the present invention.

[0060] Example 1:

[0061] This invention provides a multi-strain construction scheduling design method based on a greedy algorithm, comprising the following steps:

[0062] Step 1: Obtain the genotypes of all starting strains: Obtain the genotypes of each target strain, including all genetic modifications made based on the starting strains.

[0063] Step 2: Iterative clustering calculation is performed on the target strains sharing the maximum number of gene target sites using integer programming. In each iteration, strains not assigned to a subset are clustered, and then it is checked whether all strains have been assigned to a subset. If the result is true, proceed to the next step; otherwise, continue to the next iteration. This step assigns the target strains to P subsets and determines the common ancestor strain corresponding to each subset.

[0064] The algorithm used in the iterative clustering calculation in the second step is as follows:

[0065] maximize

[0066]

[0067] In the above method, I s =1 indicates that the target strain s is included in the subS{p} set; otherwise, I s=0. Obviously, if strain s has already been assigned to a subset, it is impossible for it to be assigned to the current subset p.

[0068]

[0069] T g =1 indicates that the gene operation g is included in the gene operation of the common ancestor of the subset; otherwise, T g =0.

[0070]

[0071] W g,s =1 indicates that the common ancestral strain carries the gene. The purpose of operation g is to construct strain s; otherwise, W g,s =0. Therefore, if the target strain s is not included in S c In the set (i.e., I) s =0), then W g,s =0, if the target strain s is contained in S c In the set (i.e., I) s =1), W g,s The value of is determined by whether the gene operation g is included in the gene operation of the common ancestor of the subset:

[0072]

[0073] Clearly, if the operator gene g is not included in the target strain s, then the common ancestor cannot carry the operator g:

[0074]

[0075] This is determined by identifying whether a common ancestor contains the gene manipulation used to construct the target strain. c In the set,

[0076]

[0077] The subset must contain more than 2 strains in total, because at least two target strains can share a common ancestor:

[0078]

[0079] Under the constraints described above, the optimization objective is set to maximize the number of shared gene target modifications, resulting in the strain set S under this objective. c And the scheduling design for constructing various common ancestral strains.

[0080] Step 3: The common ancestor strains obtained in Step 2 are constructed and scheduled, and the parent-child inheritance order of the modified target points is arranged in descending order of the difficulty of the gene target modification.

[0081] Step 4: Starting from the common ancestral strains, the construction schedule of the target strains in each subset is obtained. The parent-child inheritance order of the newly added modification targets is arranged in descending order of the difficulty of gene target modification, and finally the scheduling tree is drawn.

[0082] Application example:

[0083] Table 1 lists 30 strains to be constructed and the gene targets to be knocked out for each strain. Using a separate construction approach, meaning constructing the first strain entirely independently, requires 335 gene knockouts based on the tasks in Table 1.

[0084] Table 1 shows the strain construction tasks used for testing (this task includes 30 strains, all of which have had their genes knocked out, represented by R followed by the gene abbreviation).

[0085]

[0086]

[0087]

[0088] The method of this invention is used to optimize the scheduling. Each strain in Table 1 and its gene manipulation requirements (i.e., the genotypes in the table) are used as the raw data for calculation (the starting strain is a wild-type strain unless otherwise specified). That is, the set of strains S to be constructed given in Table 1 and the genotypes Gs of each strain are used as input. The GSCAS algorithm is run according to the aforementioned steps, and finally the strain construction scheduling diagram is obtained. Figure 2 ).

[0089] Thus, using the scheduling algorithm of this invention, the number of gene operations obtained is 217. In total, only 217 genes need to be knocked out (e.g., ...). Figure 2 As shown in the figure, where dots represent strains and edges represent gene target modifications, the total number of gene knockouts was reduced by 35.2% compared to constructing these strains individually, thus obtaining all the strains listed in Table 1. Figure 2 The numbers in the table correspond to the target strain numbers in Table 1.

[0090] Examples of common ancestral strains: Figure 2 Taking the leftmost branch as an example, this branch contains two subbranches. The common ancestor strain refers to the strain represented by the nearest node above the two gene modifications GND and DRPA. This strain is the common ancestor strain of the strains represented by the four nodes below it.

Claims

1. A method for scheduling the construction of multiple strains based on greedy algorithm, comprising the following steps: Step 1: Obtain the genotypes of all starting strains: obtain the genotypes of each target strain, which includes all genetic modifications made on the basis of the starting strain; Step 2: Use integer programming to iteratively cluster target strains that share the maximum number of genetic targets until all strains are assigned to a cluster subset, this step assigns target strains to P subsets and determines the corresponding common ancestor strain for each subset; Step 3: Schedule the construction of each common ancestor strain obtained in Step 2, and arrange the parent-child inheritance order of target modification in descending order of genetic target modification difficulty; Step 4: Starting from each common ancestor strain, add each target modification in the difference set of target modifications of each target strain and the corresponding ancestor strain in descending order of genetic target modification difficulty to obtain the construction schedule of target strains in each subset; wherein The algorithm used in the iterative clustering calculation in Step 2 is: maximizing where I s = 1 indicates that the target strain s is contained in the set subS{p}, otherwise I s = 0. T g = 1 indicates that the gene operation g is contained in the gene operations of the common ancestor of the subset, otherwise T g = 0; W g,s = 1 means that the common ancestor strain carries the gene manipulation g for the purpose of constructing strain s, otherwise W g,s = 0; if the target strain s is not contained in the set S c , i.e. I s = 0, then W g,s = 0, if the target strain s is contained in the set S c , i.e. I s = 1, the value of W g,s is determined by whether the gene manipulation g is contained in the gene manipulations of the common ancestor of the subset: If the operon g is not included in the target strain s, the common ancestor cannot carry the operon g: By identifying whether the common ancestor contains the genetic manipulation that built the target strain s, it is determined whether this strain is in S c the set, The total number of strains included in the subset must be greater than 2, and at least two target strains have a common ancestor: Under the constraints of the above conditions, the optimization objective is set to maximize the number of shared genetic target modifications, resulting in a set of strains S under this objective c .

2. The multi-strain construction scheduling design method of claim 1, wherein, Step 3 also includes scheduling the construction of each common ancestor strain.

3. The multi-strain construction scheduling method of claim 1 or 2, wherein, Step 4 also includes drawing a schedule tree.

4. A system for designing target strains for parallel construction that minimizes the total number of genetic operations based on the method of any one of claims 1 to 3, comprising a data entry module, a data processing module, and a result output module.

5. The system of claim 4, wherein, The data processing module is executed by a computer to perform the algorithm of Step 2; the result output module includes a schedule tree drawing submodule.

6. The system of claim 5, wherein, The results are displayed on the computer screen, remote terminal and mobile terminal.

7. Apparatus comprising a system as claimed in any one of claims 4 to 6, characterized in that It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executes the computer program, and the computer program code implements the system of any one of claims 4 to 6.

8. The apparatus of claim 7, wherein, It also includes data entry devices such as keyboards or image recognition devices, and display devices such as screens or remote result displays.

9. The apparatus of claim 7 or 8, wherein, It exists in the form of a computer program package or an internet platform.

Citation Information

Patent Citations

  • Weighted microbial clustering analysis method based on microbial mass spectrometer

    CN109859799A

  • HIV vaccines comprising one or more population episensus antigens

    WO2016054654A1