A method and use for the development of cell membrane and / or serum albumin lipidated analogs using computer aided design

By using computer-aided design and genetic code expansion methods, suitable lipidation analogs were screened, solving the problems of site specificity and uniformity of lipidation modification in protein drugs. This resulted in extended drug half-life and improved cell delivery efficiency, providing a new strategy for studying the function of protein lipidation modification.

CN117292741BActive Publication Date: 2026-07-10SHAOXING RES INST OF ZHEJIANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHAOXING RES INST OF ZHEJIANG UNIV
Filing Date
2023-01-18
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies make it difficult to design and evaluate lipidation-modified analogs, which limits the application of lipidation modifications in protein drugs, especially in terms of site specificity and uniformity.

Method used

Using a computer-aided design approach, suitable lipid analogs were screened by evaluating their hydrophobicity, affinity for serum albumin, and likelihood of recognition by orthogonalyl-tRNA synthetase. Then, their site-specific introduction into biomolecules was achieved using a genetic code extension method.

Benefits of technology

This approach enables site-specific introduction of lipid analogs onto biomolecules, prolonging drug half-life, improving cell delivery efficiency, and providing a new strategy for studying the function of protein lipid modification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117292741B_ABST
    Figure CN117292741B_ABST
Patent Text Reader

Abstract

The application discloses a method and application for developing a cell membrane and / or serum albumin lipid analogue combined method by using computer aided design, introduces a kind of bifunctional group molecular containing aromatic ring and linear aliphatic side chain to biological macromolecule, can give biological macromolecule membrane binding and serum albumin binding ability, the half-life of biological macromolecule drug is prolonged by using this bifunctional lipid molecule, and the drug efficacy is enhanced, the function of protein lipid modification is studied in the mode of site-specific functional mutation, and the cell delivery efficiency of biological macromolecule is improved by using site-specific lipid modification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of genetic engineering technology, and in particular to a method and application for developing lipid-modified analogs of cell membranes and / or serum albumin using computation-aided design. Background Technology

[0002] Lipidification of proteins covalently links highly hydrophobic lipid molecules to specific protein sites, reshaping protein-membrane and protein-protein interactions, thus significantly influencing protein structure, localization, and migration. Lipidification modifications include isopentenylation of cysteine, palmitoylation of cysteine, myristoylation of N-terminal glycine, and fatty acylation of serine and lysine. Organisms have evolved hundreds of enzymes and regulators involved in the proper addition and removal of these modifications, playing a central role in membrane-related biological processes, including cell signal transduction, programmed cell death, cell secretion, and cellular immunity. However, the highly variable, reversible nature of lipidification, and its crosstalk with other post-translational modifications make elucidating its biological and physiological functions challenging. Inactivation mutations at modification sites are a commonly used method in lipidification research, and this method has demonstrated the important role of lipidification. However, it is difficult to elucidate the functions of specific lipidification modifications and the dynamic transformation of lipidification modifications using this method.

[0003] With the introduction of chemical biology research methods and approaches, site-specific gain-of-function mutations have become a novel strategy for elucidating the function and importance of lipidation modifications. This strategy elucidates the function of lipidation modifications at the biochemical and biophysical levels by preparing lipidated proteins with functional mutations through the following three methods: 1) direct chemical modification, using lipid molecules with reactive groups such as maleamide to specifically couple to cysteine ​​residues of proteins; 2) semi-synthesis of lipid-modified proteins, using solid-phase peptide synthesis to synthesize lipid-modified peptides, and then using natural chemical linkages to link them to specific proteins; 3) genetic code expansion coupled to bioorthogonal reactions, using genetic code expansion to introduce orthogonal reactive groups at specific sites, and then using bioorthogonal reactions to link lipid groups to proteins.

[0004] Lipidification is a clinically approved post-translational modification widely used in many best-selling drugs, including liraglutide, semaglutide, insulin degludec, and telposide. Lipidification of peptides or proteins can prolong their half-life, reduce their immunogenicity, and improve their absorption efficiency. The increased half-life after lipidification is mainly due to the ability of lipidation to bind to serum albumin, thereby resisting protease degradation and reducing renal clearance through albumin reabsorption. The production methods for lipid-modified peptides and proteins are similar to those used in research on obtaining functional mutations through lipidation modifications, and can be accomplished using the three methods mentioned above. However, site specificity and drug homogeneity are bottlenecks in the production of lipid-modified protein drugs. Method 1 requires multiple mutations in the protein to retain only the reactive groups at the target sites, and these mutations may affect protein function. Furthermore, due to the highly hydrophobic nature of lipid molecules, both Method 1 and Method 2 require the addition of organic solvents, making it difficult to obtain modified proteins. While method 3 can overcome the above problems to some extent, it is still difficult to obtain homogeneous lipid-modified proteins due to the efficiency of bioorthogonal reactions, and the additional groups introduced by orthogonal reactions may also affect the function and efficacy of proteins.

[0005] To address the above issues, there is an urgent need to develop a class of lipid-modified analogs. These lipid-modified analogs should meet the following requirements: 1) They should possess strong hydrophobicity, enabling them to bind to cell membranes; 2) They should also have the ability to bind to serum albumin, which, when introduced into protein drugs, can prolong the drug's half-life and improve its pharmacokinetics; 3) They should be site-specifically introduced into biomolecules through genetic coding or in vitro conjugation. However, there are no reports on how to design such molecules, how to evaluate them, or how to design libraries of aminoacyl-tRNA synthetases to recognize these molecules and introduce them through genetic coding. To address these issues, a solution is proposed below. Summary of the Invention

[0006] To address the problems existing in current technologies, the present invention aims to develop a computer-aided screening method. This method designs and evaluates lipid analogs that can mimic protein lipidation modifications, and uses genetic code extension to specifically introduce the genetically encoded lipid analogs obtained through virtual screening onto biomolecules. Simultaneously, this system extends the half-life of biomolecule drugs, allows for the study of protein lipid modification function through site-specific gain-of-function mutations, and improves the cellular delivery efficiency of biomolecules.

[0007] The above-mentioned technical objective of the present invention is achieved through the following technical solution:

[0008] Part 1: Establishment of a computational-aided design method for lipid-modified analogs

[0009] A computer-aided virtual screening method was developed to design and screen lipid analogs that can mimic protein lipidation modifications. This method comprises three aspects:

[0010] 1) Evaluate the hydrophobicity of the designed lipotropic analogs;

[0011] 2) Evaluate the affinity of the designed lipid analogue for serum albumin;

[0012] 3) Assess the likelihood that the designed lipid analogues will be recognized by orthogonal aminoacyl-tRNA synthetases.

[0013] The hydrophobicity of lipid analogs was evaluated by predicting cLogP; a higher cLogP value indicates stronger hydrophobicity and membrane binding. The affinity between lipid analogs and serum albumin was measured by the Gibbs free energy (ΔG) of their binding; a smaller ΔG indicates higher affinity. Using AutoDocking Vina, the designed lipid analogs were hooked to the seven major fatty acid binding sites of serum albumin. The appropriate conformation after docking was analyzed, and the Gibbs free energy of this conformation was obtained. The likelihood of lipid analogs being recognized by orthogonal aminoacyl-tRNA synthetase was measured by calculating the root mean square difference (RMSD) between the lipid analog and a reference amino acid; a smaller RMSD value indicates a higher likelihood of recognition. Further thresholds were set by calculating the RMSD between identified lipid analogs and reference amino acids to assess the likelihood of recognition. The RMSD calculation was performed using the following steps:

[0014] 1) By hooking the designed lipid analogs into the substrate binding pockets of the corresponding orthogonal aminoacyl-tRNA synthetases, the energy-optimal conformation (pose) is obtained;

[0015] 2) Extract the conformation (pose) of the reference amino acid from the reference amino acid and aminoacyl-tRNA synthetase structure;

[0016] 3) The RMSD of the designed non-natural amino acid and the reference amino acid was calculated using the LigRMSD web server.

[0017] Furthermore, this invention designs hundreds of lipid analogs, including phenylalanine analogs, tryptophan analogs, lysine analogs, and non-natural amino acids with aliphatic side chains. The structures of these amino acids are shown in Table 1. The designed amino acids were evaluated using a developed virtual screening method. Analysis revealed that the hydrophobicity and HSA binding affinity of the lipid analogs increased with carbon chain elongation and decreased with the introduction of hydrophilic atoms. Simultaneously, most of the designed lipid analogs exhibited higher hydrophobicity and HSA affinity than previously reported HepoK, indicating that this virtual screening method can obtain lipid analogs that mimic protein lipidation modifications. Further analysis showed that amino acids containing aromatic rings possessed stronger HSA affinity, particularly lipid analogs with benzene rings and linear aliphatic side chains, and lipid analogs containing two aromatic rings.

[0018] Specifically, the previously invented chimeric phenylalanine-tRNA synthetase can recognize a series of phenylalanine derivatives. The RMSD of these recognized amino acids is less than 2.2. Therefore, further analysis was conducted on lipid analogs with RSMD < 2.2, -ΔG > 8.2 (Kd < 1 μM), and cLogP > 4.0. Further screening of these lipid analogs showed that para-phenylalanine derivatives with linear aliphatic side chains containing 6, 7, 8, and 9 carbons are ideal lipid analogs. Regardless of whether the linker between the linear aliphatic side chain and the benzene ring is a single, double, or triple bond, the RMSD increases sharply when the number of carbons exceeds 9, making recognition by the aminoacyl-tRNA synthetase extremely unlikely. This indicates that the limit for recognizing lipid analogs is between 9 and 10 carbons in the linear aliphatic chain. Further analysis revealed a significant increase in serum albumin affinity for non-natural amino acids with two benzene rings.

[0019] Preferably, this invention further compares the properties of phenylalanine derivatives with triple-bonded linear aliphatic side chains with lipid-modified lipids and other analogs reported in the literature. These phenylalanine derivatives include 4HexyF, 4HepyF, 4OctyF, 4NonyF, and 4DecyF. The cLogP of these five genetically encoded lipid analogs increases with carbon chain elongation, significantly higher than that of HepoK reported in the literature. Their ability to bind serum albumin is higher than that of fatty acids such as myristic acid and palmitic acid. Molecular docking analysis shows that the high serum albumin affinity of these five genetically encoded lipid analogs is conferred by benzene ring-mediated π-π interactions.

[0020] II: Synthesis and Validation of Lipid-Based Analogs

[0021] To obtain ideal lipid-modified analogs and verify the reliability of computer-aided virtual screening, the non-natural amino acid synthesized in this invention has the following structural formula:

[0022]

[0023] As preferred, 4HexyF and 4OctyF were selected to determine their affinity for serum albumin, with HepoK as a control. Surface plasmon resonance (SPR) measurements showed that the Kd values ​​of 4HexyF and 4OctyF binding to serum albumin were 103 μM and 23.9 μM, respectively, which are 15.3 and 67 times higher than those of HepoK (1.6 mM). This preliminarily demonstrates that lipid-modified analogs containing aromatic rings and aliphatic side chains can interact with serum albumin through hydrophobic and π-π interactions, making them ideal lipid-modified analogs.

[0024] 3: Screening for identification of lipoylation analogue chimeric phenylalanine-tRNA synthetase mutants

[0025] Lipid-modified analogs such as 4HexyF and 4OctyF were coupled to the substrate-binding pocket of a chimeric phenylalanine-tRNA synthetase, and a saturated mutagenesis library (Q356NNK, L360NNK, E391GAN, V393NNK, M490NNK, L494NNK, T467G, and A507G) was constructed by selecting surrounding amino acid residues. Positive and negative selection yielded a positive mutant only in the 4HexyF experimental group, named LipRS-1 (Q356G, E391D, T467G, M490G, and A507G), with its nucleotide sequence shown in SeqID No. 1. Further experiments revealed that this mutant could also recognize 4HepyF, 4OctyF, 4NonyF, 4HexeF, 4OcteF, 4FbutF, 4FproF, and 4FpenF, but not 4DecyF, consistent with previous virtual screening results. LC-MS analysis showed that, except for 4NonyF which had a fidelity of only 50%, the other genetically encoded lipid analogs had high fidelity.

[0026] To further improve the efficiency of recognizing genetically encoded lipid analogs, the 4HexyF molecule was inserted into LipRS-1. The results showed that the Q356G and M490G mutations provided sufficient space to accommodate these genetically encoded lipid analogs, but their interaction with the aliphatic side chain was weak. Further rational mutations at these two positions are needed to improve the efficiency of recognizing genetically encoded lipid analogs. Experiments showed that the 490A mutation significantly improved the efficiency of recognizing 4HexyF and 4OctyF. This mutant was named LipRS-2, and the corresponding nucleotide sequence is shown in Seq ID No. 2. Conversely, mutating the 490 position to an amino acid larger than the alanine side chain resulted in the loss of the ability to recognize genetically encoded lipid analogs. Next, through rational design, the fidelity of 4NonyF recognition was further improved, and a mutant with high fidelity in recognizing 4NonyF (L225V, Q356G, E391D, F464I, T467G, M490G, and A507G) was finally obtained, named LipRS-3, and the corresponding nucleotide sequence is shown in SeqID No.3.

[0027] The genetically encoded lipid analogs of this invention have high cLogP values, resulting in low solubility. Specifically, 4HexyF has a solubility of 224 μM, and 4DecyF has a solubility of 18.6 μM. Further, using 4OctyF as an example, the effect of dosage on recognition activity was investigated. Experimental results showed that even at a working concentration as low as 4 μM, LipRS-2 could still efficiently recognize and introduce the protein.

[0028] IV: Engineering Design of Therapeutic Drug Candidates Based on Precision Lipolysis

[0029] Lipid modification has been successfully applied to various FDA-approved peptides, significantly increasing their half-life, reducing patient medication costs, and alleviating patient suffering. However, progress on lipid modification strategies for large protein molecules has been slow, for the following reasons:

[0030] 1) There is a lack of site-specific lipidation modification tools. The target protein may have multiple reactive groups, and direct chemical modification is difficult to achieve specific coupling.

[0031] 2) Because lipids are highly hydrophobic, direct chemical modification must be carried out in an organic solution, which will damage the structure and function of the protein.

[0032] The genetically encoded lipid analogs of the present invention can be efficiently and site-specifically introduced onto target proteins through a chimeric phenylalanine translation system, providing a new platform for novel lipid repair protein drugs.

[0033] Lipidification modification prolongs drug half-life primarily because it endows drugs with the ability to bind to serum albumin; the stronger the binding, the longer the half-life. This invention first assesses the feasibility of creating precise lipolysis-based therapeutic drug candidates by measuring the affinity of the introduced lipid analogue's biomolecules for serum albumin.

[0034] Specifically, 4HexyF and 4OctyF were first introduced into the K20 site of GLP-1, and their affinity for human serum albumin (HSA) was measured using a microthermophoresis apparatus. The results showed that the Kd values ​​of GLP1-20-4HexyF and GLP1-20-4OctyF for binding to HSA were 2.31 μM and 0.58 μM, respectively, which are 6.5 and 25.9 times that of GLP1-20-HepoK. Wild-type GLP-1 did not bind to HSA, further demonstrating that the genetically encoded lipid analogs of this invention can confer strong HSA binding ability to the target protein.

[0035] Furthermore, a universal method was developed to introduce lipid analogs to prolong the half-life of the target protein. Specifically, lipid analogs were introduced at the N-terminus of the protein to minimize the impact on its function. GFP and Neo-2 / 15 were selected as model proteins, with Neo-2 / 15 being a de novo synthesized, biased IL-2 analog reported in 2019, exhibiting antitumor effects. The affinity of GFP and IL-2 mutants for HSA was determined using isothermal titration calorimetry, showing that the Kd values ​​of GFP-4HexyF, GFP-4OctyF, Neo-2 / 15-4HexyF, and Neo-2 / 15-4OctyF for HSA binding were 260 nM, 160 nM, 474 nM, and 370 nM, respectively. Further analysis revealed that one HSA molecule binds to 6-7 genetically encoded lipid-encoded proteins. Next, the binding ability of the introduced 4OcteF, 4FbutF, 4FproF and 4FpenF proteins to HSA was determined, and the results showed that the Kd values ​​of these mutants were all in the nM range.

[0036] Furthermore, the precisely lipid-modified Neo-2 / 15 mutant, Neo-2 / 15-4OctyF, was applied to the treatment of a mouse model of colon cancer. Results showed that compared to wild-type Neo-2 / 15, the Neo-2 / 15-4OctyF treatment group resulted in a 32% smaller tumor volume after 15 days, and a median survival time of 2.2 days in mice. These experimental results further demonstrate that the site-specific introduction of genetically encoded lipid analogs can serve as a novel strategy and platform for producing novel drug candidates, particularly protein drugs that are difficult to manipulate using traditional chemical modification methods. Therefore, the strategy of this patent is not limited to the proteins GFP and Neo-2 / 15, but also includes all other large drug molecules.

[0037] 5. Genetically encoded lipid analogs mimic the lipidation modification of proteins.

[0038] Lipidification of proteins, including myristoylation of N-terminal glycine, palmitoylation of cysteine, and isopentenylation, plays a crucial role in cellular life processes. However, the lack of genetic coding tools applicable to living cells limits the elucidation of lipidation modification functions. Lipids are highly hydrophobic molecules, and the vast majority of lipidation modification functions are membrane-related. Previous virtual screening results indicate that the cLogP of 4OctyF developed in this invention is similar to that of myristic acid, potentially possessing the ability to mimic membrane binding through lipidation modifications.

[0039] Specifically, the chimeric phenylalanine genetic code extension system for recognizing lipidation analogs was first introduced into mammalian cells. Flow cytometry results showed that it could efficiently introduce genetically encoded lipidation analogs into mammalian cells. Further, myristoylation modification of proteins was selected, with LCK, Xrp1, SVIN, and Gαi1 being preferred. 4OctyF was introduced at their myristoylation modification site (G2), and live-cell observation was performed to observe the localization of the mutant proteins. Live-cell imaging showed that the four proteins with introduced 4OctyF were mostly distributed on the plasma membrane, similar to the distribution of wild-type proteins, while the control group mutated to phenylalanine did not accumulate on the plasma membrane. Next, two palmitoylated proteins, R7BP and STERX, were selected, and 4OctyF was introduced at their palmitoylation modification sites. Imaging results showed that their cellular distribution was similar to that of wild-type proteins, also distributed on the plasma membrane. Finally, KRas4B, which underwent isovalerylation modification, was selected, and 4OctyF was introduced at its isovalerylation modification site; the mutant protein was also distributed on the plasma membrane.

[0040] The above results demonstrate that the genetically encoded lipidation analog 4OctyF of this invention can endow proteins with membrane-binding ability and can mimic protein myristylation, palmitoylation, and isovalerylation modifications to study protein lipidation in a gain-of-function manner. The lipidation modification mimicking strategy of this invention is not limited to LCK, Xrp1, SVIN, Gαi1, R7BP, STERX, and KRas4B mentioned above, and can be applied to the functional study of all proteins undergoing lipidation modifications.

[0041] 6. The ability to regulate protein membrane binding through length-tunable genetically encoded lipid analogs.

[0042] Protein lipidation exhibits heterogeneity, with different lengths or unsaturated fatty acids linked at the same modification site. The function of this heterogeneity remains to be elucidated. This invention develops lipidation analogs of different lengths to investigate the function of proteins linking lipids of varying lengths. Specifically, three genetically encoded lipidation analogs—4HexyF, 4HepyF, and 4OctyF—with 13, 14, and 15 carbon chains, were selected and introduced into the C185 site of Kras4B, with Hepok and C185F mutants serving as controls. Experimental results show that Kras4B-4OctyF protein exhibits significant plasma membrane enrichment, while Kras4B-4OctyF is mainly distributed in the cytoplasm and nucleus, similar to the control group introduced with Hepok. Further quantitative analysis of the ratio of fluorescence intensity on the membrane to that in the cytoplasm (P / C) revealed that Kras4B-4OctyF has a P / C value of 12.4, which is 5 and 10 times higher than that of Kras4B-4HpeyF and 4HexyF, respectively. To further demonstrate the advantages of this invention in terms of membrane binding tunability, a GC dipeptide motif was selected. Introducing lipidation at the G position induced palmitoylation in the Golgi apparatus, thereby transporting the corresponding protein to the plasma membrane. The group introducing 4HexyF showed that the protein was primarily enriched in the Golgi apparatus, while the groups introducing 4HepyF and 4OctyF, in addition to enrichment in the Golgi apparatus, were also enriched on the plasma membrane. The P / C ratios for the two groups were 1.3 and 3.3, respectively.

[0043] The above results indicate that the membrane binding strength changes with the length of the carbon chain. Further analysis using liposome co-precipitation determined the membrane affinity of two genetically encoded lipid analogs, 4HexyF and 4OctyF. The results showed that the membrane affinity of PolyK-4OctyF-EGFP was 105 μM, 18 times that of PolyK-4HexyF-EGFP. Further thermodynamic analysis revealed that the difference in the two carbon chains resulted in a 1.6 kcal / mol difference in the Gibbs free energy of binding. These results suggest the existence of a cLogP threshold; a given chemical with a cLogP higher than this threshold tends to bind to the cell membrane. The cLogP values ​​of 4HexyF and 4OctyF were 4.4 and 5.3, respectively. Previous literature reported that lauric acid has a weaker cell membrane binding ability, with a cLogP of 5.07. Therefore, this invention hypothesizes that this threshold is 5.07-5.3.

[0044] 7. Genetically encoded lipid analogs promote protein delivery in cells.

[0045] Cell-penetrating peptides (CPPs) are short peptides that facilitate the entry of macromolecules such as proteins into cells. They are generally composed of positively charged amino acids and hydrophobic amino acids. PCP-mediated delivery primarily occurs through endocytosis and direct transduction, with biomolecules like proteins typically entering cells via endocytosis. However, these PCPs require high protein concentrations for protein delivery and tend to accumulate in endosomes, leading to low delivery efficiency to the cytoplasm or nucleus. Previous literature has reported that lipidation modification increases the direct transduction pathway during PCP delivery and enhances the chance of delivery cargo escaping from endosomes through lipid-mediated interactions with the endosome membrane. The genetically encoded lipidation analogs developed in this invention can increase protein hydrophobicity and enhance membrane interactions, showing great potential to improve protein delivery efficiency in cells.

[0046] Specifically, PolyK-4OctyF-EGFP, PolyK-4HexyF-EGFP, and PolyK-EGFP were added to HeLa cells and incubated at 37°C for 2 hours. Protein delivery was observed using confocal fluorescence microscopy, and delivery efficiency was analyzed by flow cytometry. The results showed that the fluorescence signals of PolyK-4HexyF-EGFP and PolyK-EGFP were mostly distributed in the endosomes, while the fluorescence signal of PolyK-4OctyF-EGFP was primarily enriched in the plasma membrane. This further confirms the strong affinity of PolyK-4OctyF for the plasma membrane. Flow cytometry results showed that the cell delivery efficiencies of PolyK-4OctyF-EGFP and PolyK-4HexyF-EGFP were 3.1 and 2.3 times that of PolyK-EGFP, respectively. These results indicate that the genetically encoded lipid analogs developed in this invention can enhance intracellular protein delivery efficiency. The introduction of genetically encoded lipid analogs is a common strategy that can be applied to a variety of proteins.

[0047] The beneficial effects of this invention are as follows:

[0048] (1) This invention develops a virtual screening strategy that can rapidly evaluate the hydrophobicity, serum albumin binding ability, and recognition probability of orthogonalyl-tRNA synthetase in the design of lipid analogs. This strategy can greatly improve the usability of designing lipid analogs and reduce operating costs. Furthermore, through this virtual screening, this invention discovered a class of lipid analogs with bifunctional groups consisting of linear aliphatic side chains and aromatic rings, which can effectively mimic the properties of lipid binding to serum albumin and membranes. More importantly, the virtual screening strategy developed in this invention can also be extended to the design and recognition of other non-natural amino acids, such as fluorescent non-natural amino acids, non-natural amino acids that bind to a specific receptor, and non-natural amino acids with proximity reaction characteristics.

[0049] (2) This invention utilizes a broad-spectrum orthogonal chimeric phenylalanine genetic code extension system to identify designed lipid analogs, specifically introducing the lipid analogs into proteins to endow the target proteins with the ability to bind serum albumin (Kd approximately 400 nM), thereby enhancing the therapeutic efficacy of the drug protein. Lipids are strongly hydrophobic, and their introduction into proteins may interfere with protein function, requiring meticulous experimental screening of the introduction sites. However, traditional in vitro chemical modifications involving lipidation make it difficult to screen lipid introduction sites, hindering the application of lipidation strategies in protein drugs. This site-specific genetically encoded lipid analog introduction system can not only rapidly screen introduction sites for genetically encoded lipid analogs but also quickly evaluate the efficacy of the target protein with the introduced genetically encoded lipid analogs. This invention further proposes a strategy for introducing genetically encoded lipid analogs at the N-terminus of proteins, which can minimize interference with protein function and can be applied to the modification of all proteins of interest.

[0050] (3) The 4OctyF developed in this invention is the first genetically encoded lipid analog that mimics the membrane-binding properties of protein lipidation modifications in mammalian cells. It can mimic not only myristoylation and palmitoylation modifications of proteins, but also isovalerylation modifications, possessing the potential to be applied to the functional study of proteins with all lipidation modifications. It can utilize a gain-of-function strategy to decode the function of lipid-modified proteins and the role of lipidation modifications in this function. More importantly, this invention develops genetically encoded lipid analogs with linear aliphatic side chains of different chain lengths, which can not only be used to decode the function of protein lipidation modification heterogeneity, but also applied to synthetic biology to construct novel signaling pathways through tunable membrane binding. Attached Figure Description

[0051] Figure 1 This is a computer-aided design diagram of genetically encoded lipid analogs for an example. Figure A: Overview of the development and application of genetically encoded lipid analogs; Figure B: Screening diagram of ΔG and cLogP for genetically encoded lipid analogs; Figure C: Distribution diagram of genetically encoded lipid analogs with -ΔG>8.2 and cLogP>4, where gray represents non-natural amino acids that do not meet the requirements and whose RMSD is greater than the threshold; Figure D: Chemical structures of candidate genetically encoded lipid analogs and myristic acid.

[0052] Figure 2For the example, calculations were performed to assist in the design of functionalized genetically encoded lipid analogs; Figure A: Distribution of genetically encoded lipid analogs with -ΔG>8.2 and cLogP>4, except for gray non-natural amino acids that meet the RMSD threshold requirements; Figure B: Chemical structure of candidate genetically encoded lipid analogs; Figure C: Relationship between the length of the linear aliphatic side chain of the phenylalanine analog in B and RMSD;

[0053] Figure 3 This diagram illustrates the computer-aided design and screening of functionalized genetically encoded lipid analogs. Figure A shows the calculation flow of the affinity ΔG between the genetically encoded lipid analog and serum albumin. Figure B shows the calculation flow of the RMSD of the genetically encoded lipid analog. Figure C shows the chemical structures of candidate genetically encoded lipid analogs and some control compounds. Figure D shows the distribution of cLogP in compound C. Figure E shows the affinity distribution of compound C with human serum albumin. Figure F shows the structure of 4HexyF molecules attached to the seven binding pockets of human serum albumin, indicating the presence of Pi-Pi interactions. Figure G shows the SPR determination of the affinity of some genetically encoded lipid analogs with human serum albumin.

[0054] Figure 4 An overview diagram of the synthesis of genetically encoded lipid analogs for this example;

[0055] Figure 5 This is a screening and identification diagram for functional genetically encoded lipid analogs, as shown in the example. Figure A: The 4HexyF molecule is coupled to the substrate binding pocket of the chimeric phenylalanine-tRNA synthetase, and the side chains of amino acids surrounding 4HexyF are shown. These amino acids are selected to construct a mutant library for screening genetically encoded lipid analogs. Figure B: Green fluorescent protein is used to determine the amber codon inhibition efficiency of genetically encoded lipid analogs. The mutations of LipRS-1 and 2 are Q356G / E391D / T467G / M490G / A507G and Q356G / E391D / T467G / M490A / A507G, respectively. Figure C: Mass spectra of the fidelity of LipRS-2 in recognizing 4HexyF, 4OctyF, and 4NonyF. Figure D: Recognition efficiency of LipRS-2 at different concentrations of 4OctyF. Figure E: Simulated complex structure of LipRS-1 binding to 4HexyF. Figure F: The large side chain mutation at position 490 hinders the recognition of genetically encoded lipid analogs.

[0056] Figure 6This is a recognition diagram of functional genetically encoded lipid analogs for an example. A: LipRS-2 recognition of 4HepyF; B: Mass spectrum of 4HepyF recognition; C: Efficiency diagram of LipRS-2 recognition of 4HexeF and 4OcteF; D: Chemical structure of candidate genetically encoded lipid analogs; E: Efficiency diagram of LipRS-2 recognition of 4FproF, 4FbutF, and 4FpenF; F: Mass spectra of 4FproF, 4FbutF, and 4FpenF.

[0057] Figure 7 This is a graph illustrating the serum albumin binding ability of the genetically encoded lipid analog protein used in this example. Figure A: Coq staining of the GLP1 mutant; Figure B: Mass spectrum of the GLP1 mutant; Figure C: Kd value of the GLP1 mutant and serum albumin determined by MST; Figure D: Kd value of the GFP mutant and serum albumin determined by ITC; Figure E: Kd value of the Neo-2 / 15 mutant and serum albumin determined by ITC.

[0058] Figure 8 This example illustrates a candidate therapeutic drug with precisely modified lipids. Figure A: Overview of precisely modified lipids therapeutic drugs; Figure B: LC-MS of the GFP mutant; Figure C: Kd of the GFP mutant and serum albumin; Figure D: LC-MS of the Neo-2 / 15 mutant; Figure E: Kd of the Neo-2 / 15 mutant and serum albumin; Figure F: Schematic diagram of colon cancer model construction; Figure G: Curve showing the change in colon cancer volume in mice over treatment time; Figure H: Survival curve of colon cancer mice.

[0059] Figure 9 For this example, ITC was used to determine the Kd of Neo-2 / 15 mutant and serum albumin;

[0060] Figure 10 As an example, the introduction of genetically encoded lipidation analogs can mimic the membrane-binding ability of protein lipidation modifications. Figure A: FACS-assay efficiency of LipRS-2 recognizing genetically encoded lipidation analogs in mammalian cells; Figure B: Overview of protein lipidation modifications mimicked by genetically encoded lipidation analogs; Figure C: The introduction of 4OctyF can mimic myristoylation of proteins in mammalian cells; Figure D: The introduction of 4OctyF can mimic palmitoylation of proteins in mammalian cells.

[0061] Figure 11This example utilizes genetically encoded lipid analogs of adjustable length to investigate the ability of protein lipidation modifications to bind to the cell membrane. Figure A: Subcellular fate of proteins after introducing genetically encoded lipid analogs of different lengths; Figure B: Introduction of 4OctyF into Kras4B, enabling its effective distribution on the plasma membrane; Figure C: Subcellular localization of 4OctyF and 4HexyF after introduction into the GC motif; Figure D: Fluorescence distribution analysis curves from Figures B and C; Figure E: Interaction curve between the GFP mutant and liposomes.

[0062] Figure 12 As an example, the genetically encoded lipidation analogs mimic the membrane-binding ability of protein lipidation modifications. Figure A: Subcellular distribution of different gene G2F mutants; Figure B: Subcellular distribution of XRP2-introduced 4OctyF and wild-type; Figure C: Fluorescence intensity analysis curve of B; Figure D: Subcellular distribution of Lck-introduced 4OctyF and wild-type; Figure E: Fluorescence intensity analysis curve of D; Figure F: Subcellular localization of XRP2 and Lck-introduced 4HexyF.

[0063] Figure 13 For the example, subcellular localization of the fluorescence of the genetically encoded lipid analogs 4FbutF, 4FproF, and 4FpenF was introduced;

[0064] Figure 14 This example demonstrates the ability of a length-tunable genetically encoded lipid analog to bind to the cell membrane. Figure A: Subcellular fluorescence localization of 4OctyF and 4HexyF introduced to their respective sites; Figure B: Fluorescence intensity analysis curve from Figure A; Figure C: LC-MS image of the EGFP mutant; Figure D: Graphical description of the liposome co-precipitation experiment; Figure E: Threshold for assigning the protein effective membrane-binding compound cLogP.

[0065] Figure 15 This example illustrates how genetically encoded lipidation analogs promote protein delivery in cells. Figure A: Microscopic image of genetically encoded lipidation analogs promoting protein delivery; Figure B: FACS quantification of protein delivery efficiency in cells;

[0066] Figure 16 The structure used for virtual screening is shown in the example. Detailed Implementation

[0067] The following description is only a preferred embodiment of the present invention, and the scope of protection is not limited to this embodiment. Any technical solution that falls within the scope of the present invention should be protected by the present invention.

[0068] Example 1: Establishment of a computer-aided virtual screening method for designing lipid analogs

[0069] The computer-aided design method for virtual screening of lipid analogs consists of four parts:

[0070] 1) Design of lipid analogs and format conversion before virtual screening;

[0071] 2) Prediction of the lipid analog cLogP;

[0072] 3) Prediction of the affinity of genetically encoded lipid analogs for HSA;

[0073] 4) Calculation of RMSD of lipid-modified analogs.

[0074] 1. Design and format conversion of lipid-modified analogs

[0075] (1) Using ChemDraw, we generated the two-dimensional structures and SMILES of four lipid analogs (phenylalanine derivative, tryptophan derivative, lysine derivative, and non-natural amino group with aliphatic side chain).

[0076] (2) The three-dimensional structure of the lipid analog was generated using Chem3D and converted into a pdbqt format file using OpenBabel software.

[0077] 2. Prediction of the lipid-modified analog cLogP

[0078] In this invention, the cLogP value of the lipid analog is the cLogP of its side chain. The cLogP of the lipid analog is calculated using the RDKIT software package by the SMILES number of the lipid analog, and then the cLogP value of the main chain (-3.56) is subtracted.

[0079] 3. Prediction of affinity between lipotropic analogs and HSA

[0080] The affinity of genetically encoded lipid analogs for HSA was determined by using AutoDock Vina software to attach the molecules to the seven major fatty acid binding pockets of HSA, selecting the highest -ΔG in the appropriate conformation.

[0081] (1) Use AutoDockTools 1.5.6 to convert the HSA PDB file (1E7H) into a pdbqt file and set it as the acceptor molecule in molecular docking.

[0082] (2) The parameters for the docking design and the parameter settings for the seven docking pockets are shown in the table below:

[0083]

[0084]

[0085] (3) Run AutoDock Vina and select the highest -ΔG in the appropriate conformation.

[0086] 4. Calculation of RMSD of lipid-modified analogues

[0087] The RMSD of lipid analogs is calculated as the average distance between the conformation generated by the genetically encoded lipid analog when ligated to the corresponding orthogonal aminoacyl-tRNA synthetase and the conformation of the reference molecule in the corresponding crystal structure. The smaller the RMSD value, the greater the probability that the genetically encoded lipid analog will be recognized. To measure this recognition probability, this invention calculates the RMSD of already identified non-natural amino acids and sets this threshold to 1.2 times the maximum value of the RMSD of these already identified non-natural amino acids. Taking phenylalanine derivatives as an example, the calculation steps of RMSD are as follows: 1. Design a phenylalanyl-tRNA synthetase mutant with a substrate binding pocket (T467G / M490G / A507G) and predict the three-dimensional structure of this mutant using SWISS-MODEL.

[0088] 2. Using AutoDock Vina software, the corresponding phenylalanine derivative molecules were docked to the substrate binding pocket. The four conformations with the best energy were selected and converted into MOL files using Open Babel software.

[0089] 3. The three-dimensional conformation of phenylalanine was extracted from the crystal structure of human mitochondrial phenylalanine-tRNA synthetase (PDB: 5MGW), converted into an MOL file, and set as the reference for calculating RMSD.

[0090] 4. The RMSD between the docking conformation of the phenylalanine derivative molecule and the reference molecule conformation was calculated using the LigRMSD webserver, and the smallest RMSD among the four docking conformations was set as the RMSD of this genetically encoded lipid analog.

[0091] Example 2: Synthesis of lipid analogs

[0092]

[0093] N-Boc-4-iodo-L-phenylalanine (10.0 mmol, 3.92 g) was dissolved in 40 mL of a dichloromethane-methanol mixture. Trimethylsilyldiazomethane (12 mmol, 12 mL) was added dropwise to the reaction system at 0 °C. The reaction was then carried out at room temperature for 2 hours. The solvent was removed by rotary evaporation to obtain A (Al yield 99%), a white solid.

[0094] The product A1 was analyzed, and the results are as follows: (A1) 1H NMR (500MHz, CDCl3) δ7.64–7.56(m,2H),6.87(d,J=7.9Hz,2H),4.99(d,J=8.3Hz,1H),4.62–4.4 8(m,1H),3.70(s,3H),3.06(dd,J=13.8,5.8Hz,1H),2.97(dd,J=13.8,6.2Hz,1H),1.41(s,9H). 13 C NMR(125MHz, CDCl3)δ172.14,155.09,137.68,135.84,131.42,92.61,80.17,54.30,52.43,38.00,28.39.HRMS(ESI)m / z calcd.For C 10 H 13 INO2 + (M-Boc) + 305.9985, found 306.0001.

[0095]

[0096] The solid A (4.0 mmol, 1.62 g) obtained in the first step was dissolved in 15 mL of anhydrous tetrahydrofuran. Pd(PPh3)2Cl2 (0.04 mmol, 28 mg), CuI (0.08 mmol, 15 mg), and 2 mL of piperidine were added to the solution at 0 °C, and the reaction was allowed to proceed for 10 minutes. Then, B1 (4.8 mmol, 410 mg) was added dropwise. The reaction was carried out under nitrogen protection and stirred at 40 °C for 4 hours. After the reaction was complete, the mixture was poured into water, extracted with ethyl acetate, and the upper organic layer was washed with brine, dried over anhydrous sodium sulfate, and concentrated under reduced pressure. The resulting compound, D1 (83% yield), was purified by column chromatography with petroleum ether and ethyl acetate as a pale yellow oil.

[0097] The analysis of product D1 yielded the following results: (D1) 1 H NMR (500MHz, CDCl3) δ7.33–7.26(m,2H),7.04(d,J=7.9Hz,2H),5.22(d,J=8.3Hz,1H),4.59–4.48(m,1H),3.66(s,3H),3.04(ddd,J=45 .9,13.8,6.2Hz,2H),2.38(t,J=7.0Hz,2H),1.60–1.52(m,2H),1.46(ddd,J=9.9,7.6,5.9Hz,2H),1.40(s,9H),0.93(t,J=7.3Hz,3H). 13C NMR (125MHz, CDCl3) δ172.00,154.90,135.47,131.41,128.96,122.64,90.18,80. 17,79.52,54.23,51.89,37.84,30.66,28.06,21.79,18.86,13.43.HRMS(ESI)m / z calcd.For C 16 H 22 NO2 + (M-Boc) + 260.1645, found 260.1668.

[0098]

[0099] Compound D1 (2.0 mmol, 720 mg) was dissolved in 10 mL of methanol, and NaOH (6.0 mmol, 240 mg) was dissolved in 10 mL of H2O. The mixture was stirred at room temperature for 4 h, and then the methanol was evaporated under reduced pressure to approximately half the reaction volume. The solution was acidified with ice-cold 2 M dilute hydrochloric acid to adjust the pH to 3. The aqueous solution was extracted with cold ethyl acetate, the upper organic layer was washed with saturated brine, dried over anhydrous sodium sulfate, and then evaporated under vacuum to give a colorless oil, which yielded carbamate E1 without further purification. E1 was then dissolved in dichloromethane, and trifluoroacetic acid (5.0 mmol, 570 mg) was slowly added to remove the protecting group of the amino group, yielding the target compound 4HexyF (74% yield), a pale yellow solid.

[0100] The product 4HexyF was analyzed, and the results are as follows: 4HexyF) 1 H NMR(500MHz,D2O)δ7.17(d,J=7.9Hz,2H),6.97(d,J=7.9Hz,2H),3.22(dd,J=9.5,4.1Hz,1H),2.90(dd,J=13.6,4.1Hz,1 H),2.38(dd,J=13.5,9.6Hz,1H),2.18(t,J=6.9Hz,2H),1.36(dtd,J=33.8,9.6,8.5,4.7Hz,4H),0.82(t,J=7.2Hz,3H). 13 C NMR(125MHz,D2O)δ181.25,138.64,131.52,129.26,121.44,89.51,80.96,57.39,41.51,30.78,21.92,18.79,13.37.HRMS(ESI)m / z calcd.ForC 15 H20 NO2 + (M+H) + 246.1489, found 246.1549.

[0101]

[0102] The solid A (4.0 mmol, 1.62 g) obtained in the first step was dissolved in 15 mL of anhydrous toluene. Pd(OAc)₂ (0.8 mmol, 180 mg), PPh₃ (1.6 mmol, 419 mg), and 1 mL of triethylamine were added to the solution at 0 °C, and the reaction was allowed to proceed for 10 minutes. C₁ (4.8 mmol, 410 mg) was then added dropwise. The reaction was carried out overnight at 100 °C under nitrogen protection with stirring. After the reaction was complete, the mixture was poured into water, extracted with ethyl acetate, and the upper organic layer was washed with brine, dried over anhydrous sodium sulfate, and concentrated under reduced pressure. The resulting product was purified by column chromatography with petroleum ether and ethyl acetate to give compound F₁ (67% yield), a pale yellow oil.

[0103] The product F1 was analyzed, and the results are as follows: F1) 1 H NMR (500MHz, CDCl3) δ7.38–7.23(m,2H),7.07(ddd,J=27.9,7.9,4.6Hz,3H),6.39–6.14(m,1H),4.98(d,J=8.5Hz,1H),4.63–4.45(m,1H) ,3.71(d,J=5.6Hz,3H),3.06(ddd,J=24.0,13.5,7.1Hz,2H),2.19(ddd,J=16.2,8.1,6.5Hz,2H),1.53–1.27(m,14H),0.98–0.87(m,3H). 13 C NMR (125MHz, CDCl3) δ172.26,172.23,155.03,134.45,130.85,129.31,129.27,128.98,125.95,125.57,79 .63,54.38,52.00,51.96,32.62,31.45,30.75,28.20,22.72,22.16,13.86,13.81.HRMS(ESI)m / zcalcd.For C 16 H 24 NO2 + (M-Boc) + 262.1802, found 262.1764.

[0104]

[0105] Compound F1 (2.0 mmol, 720 mg) was dissolved in 10 mL of methanol, and NaOH (6.0 mmol, 240 mg) was dissolved in 10 mL of H2O. The mixture was stirred at room temperature for 4 h, and then the methanol was evaporated under reduced pressure to approximately half the reaction volume. The solution was acidified with ice-cold 2 M dilute hydrochloric acid to adjust the pH to 3. The aqueous solution was extracted with cold ethyl acetate, the upper organic layer was washed with saturated brine, dried over anhydrous sodium sulfate, and then evaporated under vacuum to give a colorless oil, which yielded carbamate G1 without further purification. G1 was then dissolved in dichloromethane, and trifluoroacetic acid (5.0 mmol, 570 mg) was slowly added to remove the protecting group of the amino group, yielding the target compound 4HexeF (65% yield), a pale yellow solid.

[0106] The product 4HexeF was analyzed, and the results are as follows: 4HexeF) 1 H NMR (500MHz, CD3OD) δ7.40–7.25(m,2H),7.24–7.14(m,2H),6.42–6.16(m,1H),4.20(ddd,J=7.5,5.5,3.2Hz, 1H),3.31–3.10(m,2H),2.19(qd,J=7.0,1.4Hz,1H),2.05–1.90(m,1H),1.56–1.18(m,4H),1.00–0.78(m,3H). 13 C NMR (125MHz, Methanol-d4) δ156.85,136.67,136.32,129.10,123.13,119.25,112.76,107.80,70.22,69.96,68.68,54.42,30.79,9.95.HRMS (ESI) m / z calcd.ForC 15 H 22 NO2 + (M+H) + 248.1645, found 248.1631.

[0107] Example 3: Surface plasmon resonance (SPR) determination of the affinity between genetically encoded lipid analogs and serum albumin

[0108] The affinity between the genetically encoded lipid analog and HSA was determined using the Biacore T200 system. The specific steps are as follows:

[0109] 1. HSA was coupled to the CM5 chip using an amino coupling method. The CM5 chip was activated using a mixture of 1M NHS and 1M EDC in equal proportions. HSA (20 μg / ml) dissolved in 10 mM sodium acetate buffer was injected into the chip to achieve a response of 10000 RU. Finally, the chip was blocked with 1M ethanolamine.

[0110] 2. Dissolve 4OctyF, 4HexyF, and HepoK in running buffer (20 mM phosphate, 150 mM sodium chloride, 5% DMSO, pH 7.5) to prepare a 1 mM solution. After incubating at room temperature for 2 hours, centrifuge at 12000 rpm for 20 minutes to remove the precipitate.

[0111] 3. These non-natural amino acids were sequentially injected into a CM5 chip immobilized with HSA, from low to high concentrations. The binding time was set to 120 seconds and the dissociation time to 300 seconds. Finally, the Kd value of the non-natural amino acids bound to HSA was obtained using Biacore S200 analysis software.

[0112] Example 4: Screening for chimeric phenylalanine-tRNA synthetase mutations that identify lipid analogs

[0113] 1. The 4HexyF molecule was coupled into the substrate binding pocket of chPheRS. The amino acids surrounding the 4HexyF molecule were selected, and the mutant library was set as follows: Q356NNK, L360NNK, E391GAN, V393NNK, M490NNK, L494NNK, T467G, and A507G.

[0114] 2. The fragment was amplified by PCR, and then ligated into the pBK vector using the Gibson assembly method to construct the chPheRS library.

[0115] 3. The constructed library was electroporated into DH10B electrocompetent cells containing the negative screening plasmid pNEG-Barnase-Q3TAG-D45TAG-chPheT. The bacterial culture was then plated onto LB agar plates containing 50 μg / ml kanamycin, 100 μg / ml laminpicillin, and 0.2% L-arabinose and incubated at 37°C for 12 h.

[0116] 4. Collect bacteria from the plates, extract plasmids, and transform the extracted plasmids into DH10B electroporated competent cells containing the positive screening plasmid pNEG-CAT112TAG-chPheT-GFP190TAG. Distribute the plasmids onto positive screening plates containing the corresponding non-natural amino acids (50 μg / ml kanamycin, 100 μg / ml ampicillin, 10 μg / ml chloramphenicol, and 0.2% L-arabinose), and incubate at 37°C for 12 h, followed by incubation at 30°C for another 48 h. Select clones with strong green fluorescence on the plates and determine the effect of introducing the genetically encoded lipid analog using the method in Example 5. Sequencing analysis is performed on clones that show strong fluorescence in the presence of the genetically encoded lipid analog and low fluorescence in the absence of the lipid analog. Ultimately, two mutants, LipRS-1 (Q356G, E391D, T467G, M490G, and A507G) and LipRS-2 (Q356G, E391D, T467G, M490A, and A507G), were obtained that efficiently recognize 4HexyF, 4HepyF, 4OctyF, 4HepeF, 4OcteF, and 4prhF. Further mutation analysis yielded a mutant that recognizes 4NonyF with high fidelity, named LipRS-3 (L225V, Q356G, E391D, F464I, T467G, M490G, and A507G).

[0117] Example 5: Determining the efficiency of chimeric phenylalanine-tRNA synthetase mutants in recognizing lipid analogs using a GFP fluorescent reporter system.

[0118] 1. The pBK plasmid containing orthocyanin-tRNA synthetase and the pNEG plasmid containing GFP190TAG were co-transformed into Escherichia coli DH10B, plated on a plate containing the corresponding antibiotic, and incubated overnight at 37°C.

[0119] 2. Select three single clones from the transformed plate and transfer them to LB medium containing the corresponding antibiotic. Incubate with shaking until the OD600 reaches approximately 0.8. Add the corresponding genetically encoded lipid analog and then add arabinose to a final concentration of 0.2% to induce expression. Set up a control without the genetically encoded lipid analog. Continue incubation at 30°C with shaking for 16 hours.

[0120] 3. After expression, centrifuge 750 μL of cell culture, remove the culture medium, and lyse with 150 μL of BugBuster protein extraction reagent (Millipore). After lysis, centrifuge at 12000 rpm for 1 min, and transfer 100 μL of supernatant to a 96-well culture plate (COSTAR). Record the GFP signal in the supernatant using a BioTek Synergy NEO2 and normalize the OD600 value of the culture.

[0121] Example 6: Construction of a series of plasmids for introducing genetically encoded lipid analogs

[0122] Unless otherwise specified, all plasmids were constructed using the Gibson assembly system. The constructed plasmids and corresponding steps are as follows:

[0123] 1. Plasmid construction for E. coli expression systems

[0124] ⑴pNEG-2*chPheT-SUMO-GLP1 and pNEG-2*chPheT-SUMO-GLP1-K20TAG

[0125] The nucleotide sequence of SUMO-GLP1-K20TAG is shown in Seq ID No. 4. The SUMO and GLP1 fragments or the GLP1-K29TAG fragment were constructed into the pNEG-2*chPheT vector using Gibson recombination. The construction methods for other prokaryotic expression plasmids are consistent. Specifically, when constructing GFP and Neo-2 / 15 fragments containing an amber codon at the N-terminus, a synthetic sequence (Syn: ATGGAGTACGAA) was introduced at the N-terminus. TAG GAATACGAGGCCGAAGCGGCTGCAAAAGAGGCCGCTGCAAAGGAAGCTGCAGCGAAGGCT), where the underlined part is the position of the amber codon.

[0126] The nucleotide sequences of Syn-GFP and Syn-Neo-2 / 15 are shown in Seq ID Nos. 5 and 7. The constructed plasmids are pNEG-2*chPheT-Syn-GFP and pNEG-2*chPheT-Syn-Neo-2 / 15, respectively. Furthermore, plasmids expressing the polyK-EGFP mutant were constructed, and fragments of polyK-EGFP and polyK-TAG-EGFP were synthesized and ligated into the pNEG-2*chPheT vector using Gibson. The nucleotide sequence of polyK-TAG-EGFP is shown in Seq ID No. 7.

[0127] 2. Plasmids for mammalian cell expression system construction

[0128] (1) Construction of eukaryotic plasmid expressing LipRS-2

[0129] The LipRS-2 fragment was amplified using primers F8 and R8 and ligated into the pCMV-chPheT vector to construct the plasmid pCMV-LipRS-2-chPheT.

[0130] (2) Construction of a cardamom-simulated plasmid

[0131] Four myristylated proteins were selected: XRP (NM_006915.3), SVIP (NM_001391937), Lck (NM_005356.5), and Gαi1 (NM_002069.6). These four genes were synthesized and ligated into the pEGFP vector to construct the wild-type plasmid pEGFP-POIs-EGFP. The Ub gene was further amplified, and the second codon of each of the four POIs was mutated to an amber codon, thus constructing the plasmid pEGFP-Ub-POIs-G2TAG-EGFP, which introduces non-natural amino acids. The nucleotide sequences of Ub-XRP-G2TAG-EGFP, Ub-SVIP-G2TAG-EGFP, Ub-Lck-G2TAG-EGFP, and Ub-Gαi1-G2TAG-EGFP are shown in Seq ID No. 8-11.

[0132] (3) Construction of simulated isopentenylated plasmids

[0133] The synthesized fragment Kras4B (165-185) was ligated into the pEGFP-mCherry-T2A-EGFP plasmid to construct pEGFP-mCherry-T2A-Kras4B-EGFP. Further, the C185 site of Kras4B was mutated to an amber codon to construct the plasmid pEGFP-mCherry-T2A-Kras4B-C185TAG-EGFP, which incorporates non-natural amino acids.

[0134] The nucleotide sequence of mCherry-T2A-Kras4B-C185TAG-EGFP is shown in Seq ID No. 12.

[0135] (4) Plasmid construction containing two fatty acid modification sites

[0136] A nucleotide sequence encoding XC (where X is the amber codon used for the introduction of non-natural amino acids) was synthesized, ligated into the corresponding vector, and the plasmid pEGFP-mCherry-TAG-C-EGFP was constructed. The nucleotide sequence encoding mCherry-TAG-C-EGFP is shown in Seq ID No. 13.

[0137] (5) Construction of simulated palmitoylated plasmids

[0138] The palmitoylated proteins R7BP (NM_001029875.3) and the STREX domain of KCNMA1 were selected and synthesized into the pEGFP-mCherry-T2A-EGFP plasmid. Simultaneously, amber codons were introduced at the corresponding palmitoylation sites, resulting in plasmids pEGFP-mCherry-T2A-R7BP-C253TAG-EGFP and pEGFP-mCherry-T2A-STREX-C13TAG-EGFP, respectively. The nucleotides encoding mCherry-T2A-R7BP-C253TAG-EGFP and mCherry-T2A-STREX-C13TAG-EGFP are shown in Seq ID No. 14-15.

[0139] Example 7: Expression, purification, and mass spectrometry of the genetically encoded lipid analog protein. In this example, the recombinant protein was expressed using an *E. coli* expression system, purified by chelate metal ion affinity chromatography, and the insertion of the genetically encoded lipid analog was further confirmed by LC-MS. The specific steps are as follows:

[0140] (1) Protein expression and purification

[0141] Overnight-cultured DH10B cells were seeded at a 1:100 ratio into 100 ml of fresh LB medium with the required antibiotics added, and cultured with shaking until the OD600 reached 0.8. The expression of the target protein was induced by adding 100 μM of the corresponding genetically encoded lipid analog and 0.2% arabinose. After induction, the cells were centrifuged at 4000 rpm for 5 minutes at 4°C. The resulting cell pellet was resuspended in pre-chilled NTA-0 buffer (25 mM Tris, 250 mM NaCl, pH 8.0) and lysed by sonication. The lysis buffer was centrifuged at 4°C for 60 minutes at 12000 rpm, and the supernatant was loaded onto a nickel affinity chromatography plate pre-equilibrated with NTA-0 buffer. The cells were then washed with 6 volumes of NTA-0 buffer containing 50 mM imidazole. Finally, the protein was eluted with NTA-0 buffer containing 500 mM imidazole.

[0142] (2) Identification of the target protein by LC-MS

[0143] The purified proteins were analyzed using a SCIEX Triple TOF 6600MS mass spectrometer with electrospray ionization and SCIEX AnalystTF software. A PHENOMENEX AERIS wide-pore C4 column was used. Separation and desalting were performed using a 2.1 x 50 mm, 3.6 μm lens. Mobile phase A was 0.1% formic acid aqueous solution, and mobile phase B was 0.1% formic acid acetonitrile. A constant flow rate of 0.2 mL / min was set. Mass spectrometry data were analyzed using SCIEX OS-Q software (version 2.0, SCIEX Corporation) through deconvolution. The molecular weight of the protein was predicted using the ExPASy Compute pI / Mw tool.

[0144] Example 8: Determination of the affinity between selected proteins and serum albumin using a micro-thermal surge test (MST)

[0145] Taking the SUMO-GLP1 mutant as an example, the detailed method for determining the affinity of the protein encoding the lipid analog introduced into it for HSA is as follows:

[0146] (1) As described in Example 7, SUMO-GLP1-K20-4OctyF, SUMO-GLP1-K20-4HexyF,

[0147] LC-MS analysis of SUMO-GLP1-K20-HepoK and SUMO-GLP1-WT proteins showed that the corresponding non-natural amino acids had been introduced with high fidelity. The four proteins were dialyzed into MST buffer (PBS buffer containing 0.01% Tween 20) and then concentrated to the appropriate concentrations using ultrafiltration concentrators. The concentrated proteins were centrifuged at 10,000g for 20 min to remove the precipitate, and the concentrations were determined for later use.

[0148] (2) Dissolve HSA powder in PBS buffer, then label it with the Cy5 fluorescent group using the Monolith RED-NHS second-generation protein labeling kit (NanoTemper Technologies). Displace the labeled protein into MST buffer.

[0149] (3) Using an NT.115 Monolith instrument (Nano Temper Technologies, Munich, Germany), the Kd values ​​of SUMO-GLP1 mutant and HSA were determined under constant temperature conditions of 25℃. Refer to the instrument manual for specific procedures.

[0150] (4) Data processing. Using NT analysis software, the data were fitted according to a model in which the target protein and fluorescent peptide bind in a 1:1 ratio to obtain the dissociation constant Kd of the target protein, and the fitting curve was plotted using Origin software.

[0151] Experimental results show that the affinities of SUMO-GLP1-K20-4OctyF and SUMO-GLP1-K20-4HexyF with HSA are 0.58 μM and 2.31 μM, respectively.

[0152] The concentrations of SUMO-GLP1-K20-HepoK (15 μM) were 25.9 and 6.3 times higher, respectively. Simultaneously, experiments showed that SUMO-GLP1-WT did not interact with HSA. These results demonstrate that the genetically encoded lipid analogs developed in this invention can confer strong HSA-binding ability to proteins, possessing significant potential to extend the half-life of the corresponding protein or peptide drugs.

[0153] Example 9: Determination of the affinity between protein and serum albumin using an isothermal titration calorimeter (ITC)

[0154] The Kd values ​​of GFP mutants and Neo-2 / 15 mutants with HSA were determined using an isothermal titration calorimeter. Taking the GFP mutant as an example, the detailed method for determining the affinity of the protein encoding the genetically modified lipid analog with HSA is as follows:

[0155] (1) Following the steps described in Example 7, GFP-4OctyF, GFP-4HexyF, and GFP-F proteins were purified. LC-MS analysis showed that the corresponding non-natural amino acids had been introduced with high fidelity. The three proteins were dialyzed into PBS buffer and then concentrated to the appropriate concentrations using an ultrafiltration concentrator. The concentrated proteins were centrifuged at 10,000g for 20 min to remove the precipitate, and the concentration was determined for later use.

[0156] (2) Dissolve HSA powder in PBS to a final concentration of 80 μM and add it to the ITC syringe. Add 40 μM of GFP mutation to the titration cell and perform titration experiments according to the ITC manual.

[0157] (3) After the titration experiment, the kinetics were fitted using the Origin 7.0 software provided by ITC, and the Kd of each mutant and HSA was calculated.

[0158] The method for determining the Kd of the Neo-2 / 15 mutant with HSA was the same as that for the GFP mutant. Experimental results showed that the introduction of 4OctyF and 4HexyF conferred the ability of both GFP and Neo-2 / 15 proteins to bind to HSA. The Kd values ​​of GFP-4OctyF and Neo-2 / 15-4OctyF with HSA were 160 nM and 370 nM, respectively. Neo-2 / 15 is a de novo synthesized IL-2 and IL-15 biased analogue that can promote CD8 binding. +The expansion of T cells results in excellent anti-tumor activity. Because Neo-2 / 15-4OctyF binds to HSA at the nM level, it may have a longer serum half-life after injection into the bloodstream, thus exhibiting a more potent anti-tumor effect.

[0159] Example 10: Construction and Treatment of Mouse Colon Cancer Model

[0160] To further explore how the introduction of genetically encoded lipid analogs can enhance the efficacy of therapeutic proteins or peptides, this study investigates the effects of introducing genetically encoded lipid analogs in a mouse colon cancer model, using Neo-2 / 15 as an example.

[0161] 1. Construction of a mouse colon cancer model

[0162] (1) Mouse CT26 cell line cultured in 10 cm culture dishes was digested with trypsin and washed twice with PBS to remove serum.

[0163] (2) 1*10 6 CT26 cells were injected subcutaneously into the left or right side of mice, and the mice were cultured to allow the tumors to grow to 100 mm. 3 .

[0164] 2. Treatment of mouse colon cancer model

[0165] The established tumor model mice were randomly divided into three groups. Endotoxin-free Neo-2 / 15, Neo-2 / 15-4OctyF, and PBS were administered intraperitoneally to the mice once daily at a dose of 1 mg / kg, respectively. Tumor volume was measured daily. Tumor volume (V) was calculated using the formula V = length * width. 2 *0.5 is used for calculation. When the tumor volume in mice exceeds 1000 mm². 3 Then they would euthanize him.

[0166] The results showed that the introduction of 4OctyF enhanced the antitumor effect of Neo-2 / 15. Compared with wild-type Neo-2 / 15, the Neo-2 / 15-4OctyF treatment group showed a 32% smaller tumor volume after 15 days, and the median survival of mice was improved by 2.2 days.

[0167] Example 11: Flow cytometry analysis (FACS) of the insertion efficiency of lipid-modified analogs in mammalian cells

[0168] (1) Cell transfection. Cells were transfected according to the standard transient plasmid transfection procedure. The experimental group consisted of cells co-expressing the chimeric phenylalanine translation system plasmid pCMV-LipRS-2-chPheT and the fluorescent reporter plasmid pEGFP-mCherry-T2A-EGFP-190TAG, while the control group consisted of cells infected with pEGFP-mCherry and pEGFP-EGFP alone.

[0169] (2) Six hours after cell transfection, the culture medium was replaced with fresh culture medium, and the corresponding genetically encoded lipid analog was added at a final concentration of 200 μM and cultured for a longer period.

[0170] (3) After 48 hours, remove the culture medium and add 1x PBS to wash away any remaining culture medium. Remove the PBS solution, add trypsin to digest the cells, add 1 mL of DMEM culture medium to resuspend the cells, and transfer the cells to a 1.5 mL centrifuge tube.

[0171] (4) Use HEK 293T cells to set the forward scattering and side scattering gates of the flow cytometer, use mCherry-expressing cells to set the parameters and gates of the PE channel, and use EGFP-expressing cells to set the parameters and gates of the FITC channel.

[0172] (5) Measure the cells in the experimental group, setting up a sample size of 50,000 cells. Analyze the data using the software FlowJo.

[0173] Experimental results show that the screened LipRS-2 can efficiently introduce the genetically encoded lipid analogs 4OctyF and 4HexyF into proteins in mammalian cells.

[0174] Example 12: Live-cell imaging analysis of the binding of genetically encoded lipid analogs to the cell membrane system

[0175] Protein lipidation endows proteins with the ability to bind to membranes, and the exercise of this function mainly depends on membrane interaction. The genetically encoded lipidation analogs developed in this invention possess the potential to study the membrane interaction function of lipidation modifications in a gain-of-function manner. Protein-membrane interactions can be studied by analyzing protein localization through live-cell imaging. Taking the mimicking farnesyl-acylated Kras4B(65-185) as an example, the steps for live-cell imaging analysis of the genetically encoded lipidation analogs simulating the membrane-binding function of protein lipidation modifications are as follows:

[0176] 1. Cell Transfection. Cultured HEK 293T cells were seeded into 20mm glass-bottomed cell culture dishes. When the cells reached 70% confluence, pCMV-LipRS-2-chPheT and pEGFP-mCherry-T2A-Kras4B-C185TAG-EGFP plasmids were co-transfected using liposome transfection reagent. Six hours after transfection, the culture medium was replaced with fresh medium and 200 μM of the corresponding genetically encoded lipid analogs 4OctyF, 4HepyF, and 4HexyF were added. A control without the genetically encoded lipid analogs, a C185F control, and a control incorporating Hepok were also prepared.

[0177] 2. Confocal live-cell imaging. Twenty-four days after transfection, observation was performed using a Zeiss LSM 900 confocal microscope at 40x oil immersion. Images were captured using the mCherry and EGFP channels selected in ZEN software. After imaging, the subcellular localization of the fluorescence signal and the degree of signal enrichment on the membrane were analyzed using ZEN software.

[0178] Imaging experiments simulating other lipidation modifications were also conducted following the above steps. The experimental results are as follows:

[0179] 1) Genetically encoded lipidation analog 4OctyF can mimic post-translational modifications of proteins such as myristylation, palmitoylation, and isopentenylation, and simulate the ability of the corresponding modifications to interact with the cell membrane. The function of protein lipidation modifications can be studied in a gain-of-function manner.

[0180] 2) The number of carbon atoms in the fatty acid chain of genetically encoded lipidation analogs determines the strength of their membrane interaction. Imaging results show that the P / C ratio of Kras4B-4OctyF is 12.4, which is 5 and 10 times higher than that of Kras4B-4HpeyF and 4HexyF, respectively. Genetically encoded lipidation analogs with different carbon chain lengths can be used to study the function of lipidation modification heterogeneity and to construct tunable signal transduction transducers for application in synthetic biology.

[0181] Example 13: Determining the effect of introducing genetically encoded lipid analogs on protein-bound lipids using liposome co-precipitation assay

[0182] 1. Preparation of liposomes. Liposomes were prepared using the thin-film hydration method. 1-Palmitoyl-2-oleoyl-sn-glycerol-3-phosphocholine and 1-palmitoyl-2-oleoyl-sn-glycerol-3-(phospho-L-serine)) were dissolved in chloroform at a ratio of 2:1, and the chloroform was evaporated to dryness using a rotary evaporator, followed by vacuum drying for 2 h. The evaporated lipid film was resuspended in buffer A (25 mM Tris, 250 mM NaCl, pH 8.2) to prepare a 5 mM liposome reservoir.

[0183] 2. Liposome Co-precipitation Assay. Three proteins, polyK-EGFP, polyK-4HexyF-EGFP, and polyK-4OcytF-EGFP, were purified according to the method in Example 7. The purified proteins were centrifuged at 16000g for 30 min at 4°C to remove the precipitate and diluted to a final concentration of 2 μM. 100 μL of the diluted protein was added to 100 μL of liposome solutions of different concentrations, incubated at room temperature for 15 min, and then centrifuged at 16000g for 15 min at 22°C. The supernatant containing unbound liposomes was transferred to a 96-well plate, and the precipitate containing bound liposomes was resuspended in 200 μL of buffer A and also transferred to a 96-well plate. The corresponding fluorescence intensity was measured.

[0184] 3. Calculation of protein-liposome Kd. The ratio of protein bound to liposomes (f...) Binding Using formula f Binding =Fpel / (Fpel+Fsp) is calculated. Where Fpel is the fluorescence of the precipitate, and Fsp is the fluorescence of the supernatant. The calculated f... Binding Plot a curve with the corresponding liposome concentration and use equation f Binding =Ka[c] / (1+Ka[c] fitting. Ka represents the binding constant, and then Kd is calculated according to the formula Kd=1 / Ka.

[0185] The results showed that genetically encoded lipid analogs can confer the ability to bind proteins to cell membranes, and the affinity for the membrane is positively correlated with the length of its carbon chain. The Kd of polyK-4OcytF-EGFP with liposomes was 105 μM, which is 18 times that of polyK-4HexyF-EGFP.

[0186] Example 14: Determination of the effect of lipid-modified analogs on protein cell delivery efficiency

[0187] 1. Culture HeLa cells and passage them into 24-well plates until they reach 70% confluence. Remove the culture medium and wash three times with PBS to remove residual serum.

[0188] 2. PolyK-EGFP, polyK-4HexyF-EGFP, and polyK-4OcytF-EGFP proteins were diluted to a final concentration of 2 μM in serum-free Opti-MEM medium and added to cells. The cells were incubated at 37°C for 2 hours. The cells were washed three times with PBS containing 0.5 mg / ml heparin to remove membrane-bound proteins. The distribution of the proteins within the cells was observed using a confocal microscope, following the setup in Example XII. Furthermore, the efficiency of different protein mutants entering the cells was quantified using flow cytometry, following the setup in Example XI.

[0189] Microscopic images showed that polyK-4HexyF-EGFP, similar to polyK-EGFP, exhibited fluorescence signals distributed within the endosomes, indicating that these protein mutants still enter the cell via endocytosis. The fluorescence signal of polyK-4OcytF-EGFP was primarily concentrated on the membrane, further demonstrating that 4OcytF binding to PolyK is a strong membrane-binding signal. Flow cytometry analysis showed that the delivery efficiencies of PolyK-4OcytF-EGFP and polyK-4HexyF-EGFP into cells were 3.1 and 2.4 times that of PolyK-EGFP, respectively.

[0190] The following is the sequence list:

[0191] Seq ID 1: LipRS-1

[0192] ATGGATAAGAAGCCGCTGGATGTTCTGATCTCTGCGACCGGTCTGTGGATG

[0193] TCCCGTACCGGCACGCTGCACAAGATCAAGCACTATGAGATTTCTCGTTCT

[0194] AAAATCTACATCGAAATGGCGTGTGGTGACCATCTGGTTGTGAACAACTCT

[0195] CGTTCTTGTCGTCCCGCACGTGCATTCCGTTATCATAAATACCGTAAAACCT

[0196] GCAAACGTTGTCGTGTTTCTGACGAAGATATCAACAACTTCCTGACCCGTT

[0197] CTACCGAAGGCAAAACCTCTGTTAAAGTTAAAGTTGTTTCTGAGCCGAAAG

[0198] TGAAAAAAGCGATGCCGAAATCTGTTTCTCGTGCGCCGAAACCGCTGGAA

[0199] AATCCGGTTTCTGCGAAAGCGTCTACCGACACCTCTCGTTCTGTTCCGTCTC

[0200] CGGCGAAATCTACCCCGAACTCTCCGGTTCCGACCTCTGCAAGCGCCCCAG

[0201] CTCTGACtaaatcccagacggaccgtctggaggtgctgctgaacccaaaggatgaaatctctctgaacagcggcaa

[0202] gcctttccgtgagctggaaagcgagctgctgtctcgtcgtaaaaaggatctgcaacagatctacgctgaggaacgcgagg

[0203] gtggcggaagcggcggcggaagccaggcctggggatcgaggcctcctgcagcagagtgtgccacccaaagagctcca

[0204] ggcagtgtggtggagctgctgggcaaatcctaccctcaggacgaccacagcaacctcacccggaaggtcctcaccagag

[0205] ttggcaggaacctgcacaaccagcagcatcaccctctgtggctgatcaaggagagggtgttggagcacttcaacaagcag

[0206] tatgtgggcagctctgggaccccgttgttctcggtctatgacaacctttcgccagtggtcacgacctggcagaactttgacag

[0207] cctgctcatcccagctgatcacccctgcaggaagaagggggacaactattacctgaatcggactcacatgctgagagcgc

[0208] acacgtccgcacacGGTtgggacttgctgcacgcgggactggatgccttcctggtggtgggtgatgtctacaggcgtga

[0209] ccagatcgactcccagcactaccctattttccaccagctgGATgccgtgcggctcttcaccaagcatgagttatttgctggt

[0210] ataaaggatggggaaagcctgcagctctttgaacaaagttctcgctctgcgcataaacaagagacacacaccatggaggc

[0211] cgtgaagcttgttgagtttgatcttaagcaaacgcttaccaggctcatggcacatctttttggagatgagccggagataaggt

[0212] gggtagactgctacttcccttttggacatccttcctttgagatggagatcaactttcatggagaatggctggaagttcttggctg

[0213] cggggtgGGTgaacaacaactggtcaattcagctggtgctcaagaccgaatcggctggggatttggcctagggttagaa

[0214] aggctagccatgatcctctacgacatccctgatatccgtctcttctggtgtgaggacgagcgcttcctgaagcagttctgtgta

[0215] tccaacattaatcagaaggtgaagtttcagcctcttagcaaa

[0216] Seq ID 2:LipRS-2

[0217] ATGGATAAGAAGCCGCTGGATGTTCTGATCTCTGCGACCGGTCTGTGGATG

[0218] TCCCGTACCGGCACGCTGCACAAGATCAAGCACTATGAGATTTCTCGTTCT

[0219] AAAATCTACATCGAAATGGCGTGTGGTGACCATCTGGTTGTGAACAACTCT

[0220] CGTTCTTGTCGTCCCGCACGTGCATTCCGTTATCATAAATACCGTAAAACCT

[0221] GCAAACGTTGTCGTGTTTCTGACGAAGATATCAACAACTTCCTGACCCGTT

[0222] CTACCGAAGGCAAAACCTCTGTTAAAGTTAAAGTTGTTTCTGAGCCGAAAG

[0223] TGAAAAAAGCGATGCCGAAATCTGTTTCTCGTGCGCCGAAACCGCTGGAA

[0224] AATCCGGTTTCTGCGAAAGCGTCTACCGACACCTCTCGTTCTGTTCCGTCTC

[0225] CGGCGAAATCTACCCCGAACTCTCCGGTTCCGACCTCTGCAAGCGCCCCAG

[0226] CTCTGACtaaatcccagacggaccgtctggaggtgctgctgaacccaaaggatgaaatctctctgaacagcggcaa

[0227] gcctttccgtgagctggaaagcgagctgctgtctcgtcgtaaaaaggatctgcaacagatctacgctgaggaacgcgagg

[0228] gtggcggaagcggcggcggaagccaggcctggggatcgaggcctcctgcagcagagtgtgccacccaaagagctcca

[0229] ggcagtgtggtggagctgctgggcaaatcctaccctcaggacgaccacagcaacctcacccggaaggtcctcaccagag

[0230] ttggcaggaacctgcacaaccagcagcatcaccctctgtggctgatcaaggagagggtgttggagcacttcaacaagcag

[0231] tatgtgggcagctctgggaccccgttgttctcggtctatgacaacctttcgccagtggtcacgacctggcagaactttgacag

[0232] cctgctcatcccagctgatcacccctgcaggaagaagggggacaactattacctgaatcggactcacatgctgagagcgc

[0233] acacgtctgcacacGGTtgggacttgctgcacgcgggactggatgccttcctggtggtgggtgatgtctacaggcgtga

[0234] ccagatcgactcccagcactaccctattttccaccagctgGAcgccgtgcggctcttcaccaagcatgagttatttgctggt

[0235] ataaaggatggggaaagcctgcagctctttgaacaaagttctcgctctgcgcataaacaagagacacacaccatggaggc

[0236] cgtgaagcttgttgagtttgatcttaagcaaacgcttaccaggctcatggcacatctttttggagatgagccggagataaggt

[0237] gggtagactgctacttcccttttggacatccttcctttgagatggagatcaactttcatggagaatggctggaagttcttggctg

[0238] cggggtgGCTgaacaacaactggtcaattcagctggtgctcaagaccgaatcggctggggatttggcctagggttagaa

[0239] aggctagccatgatcctctacgacatccctgatatccgtctcttctggtgtgaggacgagcgcttcctgaagcagttctgtgta

[0240] tccaacattaatcagaaggtgaagtttcagcctcttagcaaa

[0241] Seq ID 3:LipRS-3

[0242] ATGGATAAGAAGCCGCTGGATGTTCTGATCTCTGCGACCGGTCTGTGGATG

[0243] TCCCGTACCGGCACGCTGCACAAGATCAAGCACTATGAGATTTCTCGTTCT

[0244] AAAATCTACATCGAAATGGCGTGTGGTGACCATCTGGTTGTGAACAACTCT

[0245] CGTTCTTGTCGTCCCGCACGTGCATTCCGTTATCATAAATACCGTAAAACCT

[0246] GCAAACGTTGTCGTGTTTCTGACGAAGATATCAACAACTTCCTGACCCGTT

[0247] CTACCGAAGGCAAAACCTCTGTTAAAGTTAAAGTTGTTTCTGAGCCGAAAG

[0248] TGAAAAAAGCGATGCCGAAATCTGTTTCTCGTGCGCCGAAACCGCTGGAA

[0249] AATCCGGTTTCTGCGAAAGCGTCTACCGACACCTCTCGTTCTGTTCCGTCTC

[0250] CGGCGAAATCTACCCCGAACTCTCCGGTTCCGACCTCTGCAAGCGCCCCAG

[0251] CTCTGACtaaatcccagacggaccgtctggaggtgctgctgaacccaaaggatgaaatctctctgaacagcggcaa

[0252] gcctttccgtgagctggaaagcgagctgctgtctcgtcgtaaaaaggatctgcaacagatctacgctgaggaacgcgagg

[0253] gtggcggaagcggcggcggaagccaggcctggggatcgaggcctcctgcagcagagtgtgccacccaaagagctcca

[0254] ggcagtgtggtggagctgctgggcaaatcctaccctcaggacgaccacagcaacctcacccggaaggtcctcaccagag

[0255] ttggcaggaacctgcacaaccagcagcatcaccctctgtggctgatcaaggagagggtgttggagcacttcaacaagcag

[0256] tatgtgggcagctctgggaccccgttgttctcggtctatgacaacctttcgccagtggtcacgacctggcagaactttgacag

[0257] cgtgctcatcccagctgatcacccctgcaggaagaagggggacaactattacctgaatcggactcacatgctgagagcgc

[0258] acacgtccgcacacGGTtgggacttgctgcacgcgggactggatgccttcctggtggtgggtgatgtctacaggcgtga

[0259] ccagatcgactcccagcactaccctattttccaccagctggacgccgtgcggctcttcaccaagcatgagttatttgctggtat

[0260] aaaggatggggaaagcctgcagctctttgaacaaagttctcgctctgcgcataaacaagagacacacaccatggaggccg

[0261] tgaagcttgttgagtttgatcttaagcaaacgcttaccaggctcatggcacatctttttggagatgagccggagataaggtgg

[0262] gtagactgctacataccttttggacatccttcctttgagatggagatcaactttcatggagaatggctggaagttcttggctgcg

[0263] gggtggctgaacaacaactggtcaattcagctggtgctcaagaccgaatcggctggggatttggcctagggttagaaagg

[0264] ctagccatgatcctctacgacatccctgatatccgtctcttctggtgtgaggacgagcgcttcctgaagcagttctgtgtatcc

[0265] aacattaatcagaaggtgaagtttcagcctcttagcaaa

[0266] Seq ID 4:SUMO-GLP1-K20TAG

[0267] ATGTCGGACTCAGAAGTCAATCAAGAAGCTAAGCCAGAGGTCAAGCCAGA

[0268] AGTCAAGCCTGAGACTCACATCAATTTAAAGGTGTCCGATGGATCTTCAGA

[0269] GATCTTCTTCAAGATCAAAAAGACCACTCCTTTAAGAAGGCTGATGGAAGC

[0270] GTTCGCTAAAAGACAGGGTAAGGAAATGGACTCCTTAAGATTCTTGTACGA

[0271] CGGTATTAGAATTCAAGCTGATCAGACCCCTGAAGATTTGGACATGGAGGA

[0272] TAACGATATTATTGAGGCTCACAGAGAACAGATTGGTGGATCCCATGGCGA

[0273] AGGCACCTTTACCAGCGATGTGAGCAGCTATCTGGAAGGCCAGGCGGCGta

[0274] gGAATTTATTGCGTGGCTGGTGAAA

[0275] Seq ID 5:Syn-GFP

[0276] atgGAGTACGAAtagGAATACGAGgccgaagcggctgcaaaagaggccgctgcaaaggaagctgcag

[0277] cgaaggctggtAAAggagaagaacttttcactggagttgtcccaattcttgttgaattagatggtgatgttaatgggcacaa

[0278] attttctgtcagtggagagggtgaaggtgatgcaacatacggaaaaacttacccttaaatttatttgcactactggaaaaactacct

[0279] gttccatggccaacacttgtcactactttctcttatggtgttcaatgcttttcccgttatccgGACcacatgaaacggcatgac

[0280] tttttcaagagtgccatgcccgaaggttatgtacaggaacgcactatatctttcaaagatgacgggaactacaagacgcgtg

[0281] ctgaagtcaagtttgaaggtgatacccttgttaatcgtatcgagttaaaaggtattgattttaaagaagatggaaaacattctcgg

[0282] acacaaactcgagtacaactataactcacacaacgtatacatcacggcagacaaacaaaaaatggaatcaaagctaactt

[0283] caaaattcgccacaacattgaagatggatccgttcaactagcagaccattatcaacaaaatactccaattggcgatggccct

[0284] gtccttttaccagacaaccattacctgtcgacacaatctgccctttcgaaagatcccaacgaaaagcgtgaccacatggtcct

[0285] tcttgagtttgtaactgctgctgggattacacatggcatggatgaactctacaaa

[0286] Seq ID 6:Syn-Neo-2 / 15

[0287] ATGGAGTACGAAtagGAATACGAGgccgaagcggctgcaaaagaggccgctgcaaaggaagctgca

[0288] gcgaaggctCCTAAAAAGAAAATCCAGCTGCACGCTGAACATGCACTGTATGAT

[0289] GCACTGATGATCCTGAATATCGTCAAAACCAACAGCCCGCCGGCAGAAGA

[0290] AAAACTGGAAGATTATGCATTTAACTTTGAACTGATCCTGGAAGAAATTGC

[0291] ACGTCTGTTTGAAAGCGGTGATCAGAAAGATGAAGCAGAAAAAGCAAAA

[0292] CGTATGAAAGAATGGATGAAACGCATTAAAACCACCGCAAGCGAAGATGA

[0293] ACAGGAAGAAATGGCAAATGCAATTATTACCATTCTGCAGAGCTGGATTTTT

[0294] AGT

[0295] Seq ID 7:polyK-TAG-EGFP

[0296] ATGgggaaaaagaagaaaaagaagtcaaagacaaagtagGGCGGAAGCGGCGGCAGCGTGAG

[0297] CAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGG

[0298] ACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGG

[0299] CGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAA

[0300] GCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCA

[0301] GTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTC

[0302] CGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACG

[0303] ACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTG

[0304] GTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACAT

[0305] CCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCAT

[0306] GGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACA

[0307] ACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACC

[0308] CCCATCGGCgacGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCC

[0309] AGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTG

[0310] CTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTAC

[0311] AAG

[0312] Seq ID 8:Ub-XRP-G2TAG-EGFP

[0313] atgcagatcttcgtgaagactctgactggtaagaccatcaccctcgaggtggagcccagtgacaccatcgagaatgtcaag

[0314] gcaaagatccaagataaggaaggcattcctcctgatcagcagaggttgatctttgccggaaaacagctggaagatggtcgt

[0315] accctgtctgactacaacatccagaaagagtccaccttgcacctggtgctccgtctcagaggtggctagtgcttcttctccaa

[0316] gagacggaaggctgacaaggagtcgcggcccgagaacgaggaggagcggccaaagcagtacagctgggatcagcg

[0317] cgagaaggttgatccaaaagactacatgttcagtggactgaaggatgaaacagtaggtcgcttacctgggacggtagcag

[0318] gacaacagtttctcattcaagactgtgagaactgtaacatctatatttttgatcactctgctacagttaccattgatgactgtacta

[0319] actgcataatttttctgggacccgtgaaaggcagcgtgtttttccggaattgcagagattgcaagtgcacattagcctgccaa

[0320] caatttcgtgtgcgagattgtagaaagctggaagtctttttgtgttgtgccactcaacccatcattgagtcttcctcaaatatcaa

[0321] atttggatgttttcaatggtactatcctgaattagctttccagttcaaagatgcagggctaagtatcttcaacaatacatggagta

[0322] acattcatgactttacacctgtgtcaggagaactcaactggagccttcttccagaagatgctgtggttcaggactatgttcctat

[0323] acctactaccgaagagctcaaagctgttcgtgtttccacagaagccaatagaagcattgttccaatatcccggggtcagaga

[0324] cagaagagcagcgatgaatcatgcttagtggtattatttgctggtgattacactattgcaaatgccagaaaactaattgatgag

[0325] atggttggtaaaggctttttcctagttcagacaaaggaagtgtccatgaaagctgaggatgctcaaagggtttttcgggaaaa

[0326] agcacctgacttccttcctcttctgaacaaaggtcctgttattgccttggagtttaatggggatggtgctgtagaagtatgtcaa

[0327] cttattgtaaacgagatattcaatgggaccaagatgtttgtatctgaaagcaaggagacggcatctggagatgtagacagctt

[0328] ctacaactttgctgatatacagatgggaataGGAAGCGGCGGCAGCGTGAGCAAGGGCGAGG

[0329] AGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTA

[0330] AACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTA

[0331] CGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGC

[0332] CCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCC

[0333] GCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCG

[0334] AAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTAC

[0335] AAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCAT

[0336] CGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACA

[0337] AGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGC

[0338] AGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGAC

[0339] GGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCgac

[0340] GGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCT

[0341] GAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCG

[0342] TGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAG

[0343] Seq ID 9:Ub-SVIP-G2TAG-EGFP

[0344] atgcagatcttcgtgaagactctgactggtaagaccatcaccctcgaggtggagcccagtgacaccatcgagaatgtcaag

[0345] gcaaagatccaagataaggaaggcattcctcctgatcagcagaggttgatctttgccggaaaacagctggaagatggtcgt

[0346] accctgtctgactacaacatccagaaagagtccaccttgcacctggtgctccgtctcagaggtggctagctgtgttttccttgt

[0347] cccggggagtccgcgcctcccacgccggacctggaagagaaaagagcaaagcttgcagaggctgcagagagaagac

[0348] aaaaagaggctgcatctcggggaattttagatgttcaatctgtgcaagaaaagagaaagaaaaaggaaaaaatagaaaaa

[0349] caaattgctacatccgggcccccaccagaaggtggacttaggtggacagtttcaGGAAGCGGCGGCAGCG

[0350] TGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAG

[0351] CTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGA

[0352] GGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCG

[0353] GCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGC

[0354] GTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTC

[0355] AAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAG

[0356] GACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACA

[0357] CCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGC

[0358] AACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTAT

[0359] ATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCG

[0360] CCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGA

[0361] ACACCCCCATCGGCgacGGCCCCGTGCTGCTGCCCGACAACCACTACCTGA

[0362] GCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATG

[0363] GTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGA

[0364] GCTGTACAAG

[0365] Seq ID 10:Ub-Lck-G2TAG-EGFP

[0366] atgcagatcttcgtgaagactctgactggtaagaccatcaccctcgaggtggagcccagtgacaccatcgagaatgtcaag

[0367] gcaaagatccaagataaggaaggcattcctcctgatcagcagaggttgatctttgccggaaaacagctggaagatggtcgt

[0368] accctgtctgactacaacatccagaaagagtccaccttgcacctggtgctccgtctcagaggtggctagtgtggctgcagct

[0369] cacacccggaagatgactggatggaaaacatcgatgtgtgtgagaactgccattatcccatagtcccactggatggcaagg

[0370] gcacgctgctcatccgaaatggctctgaggtgcgggacccactggttacctacgaaggctccaatccgccggcttcccca

[0371] ctgcaagacaacctggttatcgctctgcacagctatgagccctctcacgacggagatctgggctttgagaagggggaaca

[0372] gctccgcatcctggagcagagcggcgagtggtggaaggcgcagtccctgaccacgggccaggaaggcttcatccccttc

[0373] aattttgtggccaaagcgaacagcctggagcccgaaccctggttcttcaagaacctgagccgcaaggacgcggagcggc

[0374] agctcctggcgcccgggaacactcacggctccttcctcatccgggagagcgagagcaccgcgggatcgttttcactgtcg

[0375] gtccgggacttcgaccagaaccagggagaggtggtgaaacattacaagatccgtaatctggacaacggtggcttctacatc

[0376] tcccctcgaatcacttttcccggcctgcatgaactggtccgccattacaccaatgcttcagatgggctgtgcacacggttgag

[0377] ccgcccctgccagacccagaagccccagaagccgtggtgggaggacgagtgggaggttcccagggagacgctgaag

[0378] ctggtggagcggctgggggctggacagttcggggaggtgtggatggggtactacaacgggcacacgaaggtggcggt

[0379] gaagagcctgaagcagggcagcatgtccccggacgccttcctggccgaggccaacctcatgaagcagctgcaacacca

[0380] gcggctggttcggctctacgctgtggtcacccaggagcccatctacatcatcactgaatacatggagaatgggagtctagtg

[0381] gattttctcaagaccccttcaggcatcaagttgaccatcaacaaactcctggacatggcagcccaaattgcagaaggcatgg

[0382] cattcattgaagagcggaattatattcatcgtgaccttcgggctgccaacattctggtgtctgacaccctgagctgcaagattg

[0383] cagactttggcctagcacgcctcattgaggacaacgagtacacagccagggagggggccaagtttcccattaagtggaca

[0384] gcgccagaagccattaactacgggacattcaccatcaagtcagatgtgtggtcttttgggatcctgctgacggaaattgtcac

[0385] ccacggccgcatcccttacccagggatgaccaacccggaggtgattcagaacctggagcgaggctaccgcatggtgcgc

[0386] cctgacaactgtccagaggagctgtaccaactcatgaggctgtgctggaaggagcgcccagaggaccggcccacctttg

[0387] actacctgcgcagtgtgctggaggacttcttcacggccacagagggccagtaccagcctcagcctGGAAGCGGC

[0388] GGCAGCGTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCT

[0389] GGTCGAGCTGGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCG

[0390] AGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGC

[0391] ACCACCGGCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGAC

[0392] CTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACG

[0393] ACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCT

[0394] TCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAG

[0395] GGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGA

[0396] GGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACA

[0397] ACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCA

[0398] AGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTAC

[0399] CAGCAGAACACCCCCATCGGCgacGGCCCCGTGCTGCTGCCCGACAACCAC

[0400] TACCTGAGCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGA

[0401] TCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCAT

[0402] GGACGAGCTGTACAAG

[0403] Seq ID 11:Ub-Gαi1-G2TAG-EGFP

[0404] atgcagatcttcgtgaagactctgactggtaagaccatcaccctcgaggtggagcccagtgacaccatcgagaatgtcaag

[0405] gcaaagatccaagataaggaaggcattcctcctgatcagcagaggttgatctttgccggaaaacagctggaagatggtcgt

[0406] accctgtctgactacaacatccagaaagagtccaccttgcacctggtgctccgtctcagaggtggctagtgcacgctgagc

[0407] gccgaggacaaggcggcggtggagcggagtaagatgatcgaccgcaacctccgtgaggacggcgagaaggcggcgc

[0408] gcgaggtcaagctgctgctgctcggtgctggtgaatctggtaaaagtacaattgtgaagcagatgaaaattatccatgaagc

[0409] tggttattcagaagaggagtgtaaacaatacaaagcagtggtctacagtaacaccatccagtcaattattgctatcattaggg

[0410] ctatggggaggttgaagatagactttggtgactcagcccgggcggatgatgcacgccaactctttgtgctagctggagctg

[0411] ctgaagaaggctttatgactgcagaacttgctggagttataaagagattgtggaaagatagtggtgtacaagcctgtttcaac

[0412] agatcccgagagtaccagcttaatgattctgcagcatactatttgaatgacttggacagaatagctcaaccaaattacatccc

[0413] gactcaacaagatgttctcagaactagagtgaaaactacaggaattgttgaaacccattttactttcaaagatcttcattttaaaa

[0414] tgtttgatgtgggaggtcagagatctgagcggaagaagtggattcattgcttcgaaggagtgacggcgatcatcttctgtgta

[0415] gcactgagtgactacgacctggttctagctgaagatgaagaaatgaaccgaatgcatgaaagcatgaaattgtttgacagca

[0416] tatgtaacaacaagtggtttacagatacatccattatactttttctaaacaagaaggatctctttgaagaaaaaatcaaaaagag

[0417] ccctctcactatatgctatccagaatatgcaggatcaaacacatatgaagaggcagctgcatatattcaatgtcagtttgaaga

[0418] cctcaataaaagaaaggacacaaaggaaatatacacccacttcacatgtgccacagatactaagaatgtgcagtttgtttttg

[0419] atgctgtaacagatgtcatcataaaaaataatctaaaagattgtggtctctttGGAAGCGGCGGCAGCGTGA

[0420] GCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTG

[0421] GACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGG

[0422] GCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCA

[0423] AGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGC

[0424] AGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGT

[0425] CCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACG

[0426] ACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTG

[0427] GTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACAT

[0428] CCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCAT

[0429] GGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACA

[0430] ACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACC

[0431] CCCATCGGCgacGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCC

[0432] AGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTG

[0433] CTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTAC

[0434] AAG

[0435] Seq ID 12:mCherry-T2A-Kras4B-C185TAG-EGFP

[0436] atgGTGAGCAAGGGCGAGGAGGATAACATGATGGCCATCATCAAGGAGTTCA

[0437] TGCGCTTCAAGGTGCACATGGAGGGCTCCGTGAACGGCCACGAGTTCGAG

[0438] ATCGAGGGCGAGGGCGAGGGCCGCCCCTACGAGGGCACCCAGACCGCCA

[0439] AGCTGAAGGTGACCAAGGGTGGCCCCCTGCCCTTCGCCTGGGACATCCTG

[0440] TCCCCTCAGTTCATGTACGGCTCCAAGGCCTACGTGAAGCACCCCGCCGAC

[0441] ATCCCCGACTACTTGAAGCTGTCCTTCCCCGAGGGCTTCAAGTGGGAGCGC

[0442] GTGATGAACTTCGAGGACGGCGGCGTGGTGACCGTGACCCAGGACTCCTC

[0443] CCTGCAGGACGGCGAGTTCATCTACAAGGTGAAGCTGCGCGGCACCAACT

[0444] TCCCCTCCGACGGCCCCGTAATGCAGAAGAAGACCATGGGCTGGGAGGCC

[0445] TCCTCCGAGCGGATGTACCCCGAGGACGGCGCCCTGAAGGGCGAGATCAA

[0446] GCAGAGGCTGAAGCTGAAGGACGGCGGCCACTACGACGCTGAGGTCAAG

[0447] ACCACCTACAAGGCCAAGAAGCCCGTGCAGCTGCCCGGCGCCTACAACGT

[0448] CAACATCAAGTTGGACATCACCTCCCACAACGAGGACTACACCATCGTGGA

[0449] ACAGTACGAACGCGCCGAGGGCCGCCACTCCACCGGCGGCATGGACGAGC

[0450] TGTACAAGggaagcggaGAGGGGAGAGGAAGTCTGCTAACATGCGGTGACGTC

[0451] GAGGAGAATCCTGGCCCAaaacataaagaaaagatgagcaaagatgggaaaaagaagaaaaagaagtc

[0452] aaagacaaagTAGGGCGGAAGCGGCGGCAGCGTGAGCAAGGGCGAGGAGCTG

[0453] TTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACGG

[0454] CCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTACGGCA

[0455] AGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTGG

[0456] CCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTAC

[0457] CCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGG

[0458] CTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTACAAGA

[0459] CCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGAG

[0460] CTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGCT

[0461] GGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGA

[0462] AGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGC

[0463] AGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCgacGGC

[0464] CCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCTGAG

[0465] CAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGA

[0466] CCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAG

[0467] Seq ID 13:mCherry-TAG-C-EGFP

[0468] atgGTGAGCAAGGGCGAGGAGGATAACATGATGGCCATCATCAAGGAGTTCA

[0469] TGCGCTTCAAGGTGCACATGGAGGGCTCCGTGAACGGCCACGAGTTCGAG

[0470] ATCGAGGGCGAGGGCGAGGGCCGCCCCTACGAGGGCACCCAGACCGCCA

[0471] AGCTGAAGGTGACCAAGGGTGGCCCCCTGCCCTTCGCCTGGGACATCCTG

[0472] TCCCCTCAGTTCATGTACGGCTCCAAGGCCTACGTGAAGCACCCCGCCGAC

[0473] ATCCCCGACTACTTGAAGCTGTCCTTCCCCGAGGGCTTCAAGTGGGAGCGC

[0474] GTGATGAACTTCGAGGACGGCGGCGTGGTGACCGTGACCCAGGACTCCTC

[0475] CCTGCAGGACGGCGAGTTCATCTACAAGGTGAAGCTGCGCGGCACCAACT

[0476] TCCCCTCCGACGGCCCCGTAATGCAGAAGAAGACCATGGGCTGGGAGGCC

[0477] TCCTCCGAGCGGATGTACCCCGAGGACGGCGCCCTGAAGGGCGAGATCAA

[0478] GCAGAGGCTGAAGCTGAAGGACGGCGGCCACTACGACGCTGAGGTCAAG

[0479] ACCACCTACAAGGCCAAGAAGCCCGTGCAGCTGCCCGGCGCCTACAACGT

[0480] CAACATCAAGTTGGACATCACCTCCCACAACGAGGACTACACCATCGTGGA

[0481] ACAGTACGAACGCGCCGAGGGCCGCCACTCCACCGGCGGCATGGACGAGC

[0482] TGTACAAGggaagcggaGAGGGGAGAGGAAGTCTGCTAACATGCGGTGACGTC

[0483] GAGGAGAATCCTGGCCCAGGCTAGTGCGGAAGCGGCGGCAGCGTGAGCA

[0484] AGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGAC

[0485] GGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCG

[0486] ATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGC

[0487] TGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGT

[0488] GCTTCAGCCGCTACCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCG

[0489] CCATGCCCGAAGGCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGAC

[0490] GGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGT

[0491] GAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCC

[0492] TGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGG

[0493] CCGACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAAC

[0494] ATCGAGGACGGCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCC

[0495] CATCGGCgacGGCCCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCA

[0496] GTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGC

[0497] TGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACA

[0498] AG

[0499] Seq ID 14:mCherry-T2A-R7BP-C253TAG-EGFP

[0500] atgGTGAGCAAGGGCGAGGAGGATAACATGATGGCCATCATCAAGGAGTTCA

[0501] TGCGCTTCAAGGTGCACATGGAGGGCTCCGTGAACGGCCACGAGTTCGAG

[0502] ATCGAGGGCGAGGGCGAGGGCCGCCCCTACGAGGGCACCCAGACCGCCA

[0503] AGCTGAAGGTGACCAAGGGTGGCCCCCTGCCCTTCGCCTGGGACATCCTG

[0504] TCCCCTCAGTTCATGTACGGCTCCAAGGCCTACGTGAAGCACCCCGCCGAC

[0505] ATCCCCGACTACTTGAAGCTGTCCTTCCCCGAGGGCTTCAAGTGGGAGCGC

[0506] GTGATGAACTTCGAGGACGGCGGCGTGGTGACCGTGACCCAGGACTCCTC

[0507] CCTGCAGGACGGCGAGTTCATCTACAAGGTGAAGCTGCGCGGCACCAACT

[0508] TCCCCTCCGACGGCCCCGTAATGCAGAAGAAGACCATGGGCTGGGAGGCC

[0509] TCCTCCGAGCGGATGTACCCCGAGGACGGCGCCCTGAAGGGCGAGATCAA

[0510] GCAGAGGCTGAAGCTGAAGGACGGCGGCCACTACGACGCTGAGGTCAAG

[0511] ACCACCTACAAGGCCAAGAAGCCCGTGCAGCTGCCCGGCGCCTACAACGT

[0512] CAACATCAAGTTGGACATCACCTCCCACAACGAGGACTACACCATCGTGGA

[0513] ACAGTACGAACGCGCCGAGGGCCGCCACTCCACCGGCGGCATGGACGAGC

[0514] TGTACAAGggaagcggaGAGGGGAGAGGAAGTCTGCTAACATGCGGTGACGTC

[0515] GAGGAGAATCCTGGCCCAagttctgcaccgaatgggcgcaaaaagcgccccagccggtccacccgctc

[0516] ctcgatcttccagatcagcaagcccccgctgcagagcggagattgggagcgcaggggcagcggctccgagagcgccc

[0517] acaaaacccaacgagccctggacgactgcaagatgcttgtccaagagttcaacacacaagtggccctgtaccgagagctg

[0518] gtcatttctattggggatgtctcggtcagctgcccctcactccgggcgggaaatgcacaagaaagaaaccaaaggctgtgaa

[0519] atggcccgtcaggcacaccaaaaattggctgccatctcaggcccggaagatggtgagatccatccagaaatctgtcggctt

[0520] tacatccagctgcagtgctgcttagaaatgtataccacagagatgctaaatccatatgtctgctggggtctcttcagtttcat

[0521] gaaaaaaggaacctggcgggggaaccaagagtttggattgcaaaattgaggagagtgctgaaacacctgccctaga

[0522] agactcctcatcatcccccgtagatagtcagcaacattcctggcaggtttccacagacattgagaacactgaaagagacatg

[0523] agagaaatgaaaaaccttttaagcaaactcagggaaactatgcctttaccattgaaaaatcaagatgacagcagccttctga

[0524] atctaactccctacccctggtgaagaagacggaagaagaggttctttgggctgtgttagctcatctcaagcGGAAGCG

[0525] GCGGCAGCGTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATC

[0526] CTGGTCGAGCTGGACGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGG

[0527] CGAGGGCGAGGGCGATGCCACCTACGGCAAGCTGACCCTGAAGTTCATCT

[0528] GCACCACCGGCAAGCTGCCCGTGCCCTGGCCCACCCTCGTGACCACCCTG

[0529] ACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCACATGAAGCAGCA

[0530] CGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCACCAT

[0531] CTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCG

[0532] AGGGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAG

[0533] GAGGACGGCAACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCA

[0534] CAACGTCTATATCATGGCCGACAAGCAGAAGAACGGCATCAAGGTGAACTT

[0535] CAAGATCCGCCACAACATCGAGGACGGCAGCGTGCAGCTCGCCGACCACT

[0536] ACCAGCAGAACACCCCCATCGGCgacGGCCCCGTGCTGCTGCCCGACAACC

[0537] ACTACCTGAGCACCCAGTCCGCCCTGAGCAAAGACCCCAACGAGAAGCGC

[0538] GATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGC

[0539] ATGGACGAGCTGTACAAG

[0540] Seq ID 15:mCherry-T2A-STREX-C13TAG-EGFP

[0541] atgGTGAGCAAGGGCGAGGAGGATAACATGATGGCCATCATCAAGGAGTTCATGCGCTTCAAGGTGCACATGGAGGGCTCCGTGAACGGCCACGAGTT

[0542] CGAGATCGAGGGCGAGGGCGAGGGCCGCCCCTACGAGGGCACCCAGACC

[0543] GCCAAGCTGAAGGTGACCAAGGGTGGCCCCCTGCCCTTCGCCTGGGACAT

[0544] CCTGTCCCCTCAGTTCATGTACGGCTCCAAGGCCTACGTGAAGCACCCCGC

[0545] CGACATCCCCGACTACTTGAAGCTGTCCTTCCCCGAGGGCTTCAAGTGGGA

[0546] GCGCGTGATGAACTTCGAGGACGGCGGCGTGGTGACCGTGACCCAGGACT

[0547] CCTCCCTGCAGGACGGCGAGTTCATCTACAAGGTGAAGCTGCGCGGCACC

[0548] AACTTCCCCTCCGACGGCCCCGTAATGCAGAAGAAGACCATGGGCTGGGA

[0549] GGCCTCCTCCGAGCGGATGTACCCCGAGGACGGCGCCCTGAAGGGCGAGA

[0550] TCAAGCAGAGGCTGAAGCTGAAGGACGGCGGCCACTACGACGCTGAGGT

[0551] CAAGACCACCTACAAGGCCAAGAAGCCCGTGCAGCTGCCCGGCGCCTACA

[0552] ACGTCAACATCAAGTTGGACATCACCTCCCACAACGAGGACTACACCATCG

[0553] TGGAACAGTACGAACGCGCCGAGGGCCGCCACTCCACCGGCGGCATGGAC

[0554] GAGCTGTACAAGggaagcggaGAGGGGAGAGGAAGTCTGCTAACATGCGGTG

[0555] ACGTCGAGGAGAATCCTGGCCCAaaggcctgtcatgatgacatcacagatcccaaaagaataaaaaa

[0556] atgtggctgcaaacggcccaagatgtccatctacaagagaatgagacgggcatgttagtttgattgcggacgttctgagcgt

[0557] gactgctcatgcatgtcaggccgtgtgcgtggtaacgtggacacccttgagagagccttcccactttcttctgtctctgttaat

[0558] gattgctccaccagtttccgtgccGGAAGCGGCGGCAGCGTGAGCAAGGGCGAGGAGCT

[0559] GTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGGACGGCGACGTAAACG

[0560] GCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCACCTACGGC

[0561] AAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAGCTGCCCGTGCCCTG

[0562] GCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTA

[0563] CCCCGACCACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAG

[0564] GCTACGTCCAGGAGCGCACCATCTTCTTCAAGGACGACGGCAACTACAAG

[0565] ACCCGCGCCGAGGTGAAGTTCGAGGGCGACACCCTGGTGAACCGCATCGA

[0566] GCTGAAGGGCATCGACTTCAAGGAGGACGGCAACATCCTGGGGCACAAGC

[0567] TGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGACAAGCAGA

[0568] AGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGC

[0569] AGCGTGCAGCTCGCCGACCACTACCAGCAGAACAACCCCATCGGCgacGGC

[0570] CCCGTGCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCGCCCTGAG

[0571] CAAAGACCCCAACGAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGA CCGCCGCCGGGATCACTCTCGGCATGGACGAGCTGTACAAG.

[0572] The specific embodiments described above further illustrate the technical problems, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A lipid-modified non-natural amino acid that can bind to cell membranes and / or serum albumin, characterized in that, The esterified non-natural amino acid is an L-phenylalanine derivative with a linear aliphatic side chain, and its specific structure is as follows: 。 2. An orthogonal translation system for site-specific introduction of the esterified non-natural amino acid of claim 1 into biological macromolecules, characterized in that, The system comprises: (1) the esterified non-natural amino acid of claim 1; (2) a chimeric phenylalanine-tRNA synthetase (chPheRS) mutant that specifically recognizes the esterified non-natural amino acid; wherein the coding sequence of the amino acid sequence of the chimeric phenylalanine-tRNA synthetase mutant is selected from SEQ ID NO: 1 or SEQ ID NO: 2, and when the esterified non-natural amino acid is 4HexyF, the coding sequence of the amino acid sequence of the chimeric phenylalanine-tRNA synthetase mutant is SEQ ID NO: 1 or SEQ ID NO: 2; and when the esterified non-natural amino acid is 4OctyF, the coding sequence of the amino acid sequence of the chimeric phenylalanine-tRNA synthetase mutant is SEQ ID NO:

2.

3. The application of the lipid-modified non-natural amino acid bound to cell membranes and / or serum albumin as described in claim 1, characterized in that, The esterified non-natural amino acids are introduced into biological macromolecules through genetic encoding, including the following steps: S1: An orthogonal translation system encoding the lipid-containing non-natural amino acid and a target gene containing a TAG codon are introduced into a host cell, including *E. coli* or mammalian cells; S2: The lipid-containing non-natural amino acid is added to a culture medium and the target protein is induced to express, thereby specifically introducing the lipid-containing non-natural amino acid site onto a biomolecule to obtain a biomolecule containing the lipid-containing non-natural amino acid; the biomolecule containing the lipid-containing non-natural amino acid is glucagon-like peptide-1 (GLP-1), green fluorescent protein (GFP), interleukin-2 analog, KRas4B, LCK, Xrp1, SVIN, Gαi1, R7BP, or STREX, and the nucleotide sequences of the biomolecule containing the lipid-containing non-natural amino acid are shown in SEQ ID NO: 4, 5, 6, 12, 10, 8, 9, 11, 14, and 15, respectively; the coding sequence of the amino acid sequence of the chimeric phenylalanine-tRNA synthetase mutant that specifically recognizes the lipid-containing non-natural amino acid in the orthogonal translation system is selected from SEQ ID NO: 1 or SEQ ID NO:2.