Method for constructing multivalent protein drugs and vaccines via nucleic acid multimerization and applications thereof
The use of complementary nucleic acid backbones for multimerization addresses issues of specificity and stability in forming multivalent proteins, enhancing drug half-life and vaccine immunogenicity.
Patent Information
- Application Number
- JP2023532213
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-25
- Filing Date
- 2021-11-25
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2041-11-25
AI Technical Summary
Existing methods for forming multivalent proteins for drug and vaccine applications suffer from low binding specificity, heterogeneous multimers, and immune system activation risks, with challenges in extending half-life and improving activity.
A method involving complementary nucleic acid backbones to form stable, uniform multimers through nucleic acid multimerization, allowing for efficient assembly of multivalent protein drugs and vaccines without fusion proteins, using nucleic acid sequences that self-assemble via base complementarity.
The method enhances the half-life and activity of protein drugs and improves immunogenicity of vaccines by forming stable, specific multivalent complexes rapidly and efficiently, reducing immune system activation risks.
Smart Images

Figure 0007705672000070 
Figure 0007705672000071 
Figure 0007705672000072
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology medicine, and specifically to a method for constructing and applying multivalent protein drugs and vaccines through nucleic acid multimerization.
Background Art
[0002] In the case of many biopolymers, their aggregation or multivalent state directly affects their in vivo activity and half-life. For example, the activation of most immune receptors is accompanied by the aggregation of receptors on the cell membrane, thereby activating the downstream signal transduction pathway inside the cell. Therefore, when the natural ligands or antibodies of these receptors form multivalents or high valences, they often significantly improve the ability to activate the receptors. In addition, some protein drugs with low molecular weight (MW < 40 kDa), such as cytokines, growth hormones, and synthetic polypeptides, have high renal clearance and short half-lives in vivo. These protein drugs can also increase their molecular weight by forming high valences and extend their half-lives.
[0003] Therefore, the multivalentization of proteins is a process that has attracted attention in the field of biomedicine. There are many existing methods, but in the case of many chemical cross-linking methods, there are problems such as low binding specificity and heterogeneous binding multimers. Currently, the most widely applied method is to express and produce proteins in the form of multivalent fusion proteins by cells, and to fuse the protein region with effective functions with a protein that can form oligomers to form chimeras such as Fc bivalent fusion proteins and GCN4 trivalent fusion proteins. Although these fusion proteins can all form uniform oligomers, the cell expression level and activity of the fusion proteins are lower than those of the original protein drugs, and the presence of the Fc region poses a risk of activating the immune system, inducing the release of cytokines, and causing cytotoxicity. Therefore, in this field, there is an urgent need to develop a simple, flexible, and efficient method that can form non-fusion protein-carried, uniform, and highly specific multivalent proteins from verified protein drugs.
[0004] Furthermore, multimerization of proteins is also important for vaccine development. First, in B cell-based vaccine design, activation of the B cell receptor (BCR) with viral or bacterial proteins as antigens is a crucial step. The BCR is the same as the immune receptor described above, and for its effective activation, the receptors need to aggregate on the cell membrane. Therefore, compared to monomeric antigens, multimeric antigens have an absolute advantage in activating B cells. Next, the multimeric antigen does not have to be a single antigen for forming oligomers. The multimeric antigen can include different proteins within a virus or mutant forms and subtypes of the same protein from different viral strains. Thus, such diverse antigen presentation can theoretically induce a polyclonal B cell response of the host immune system and generate a broader range of neutralizing antibodies. Summary of the Invention Problems to be Solved by the Invention
[0005] The first object of the present invention is to provide an efficiently and stably assembled n-stage nucleic acid oligomer assembly backbone design suitable for the efficient and stable assembly of nucleic acid coupling protein drugs for forming multivalent drugs or vaccines.
[0006] The second object of the present invention is to provide a simple and efficient method for forming a protein drug into a multivalent polymer to extend the half-life of the drug and increase the activity of the drug.
[0007] The third object of the present invention is to provide a simple, flexible, efficient and modular method for forming a multivalent polymer complex with the same or different protein antigens to activate immune cells and improve the immunogenicity of vaccines. Means for Solving the Problems
[0008] The first aspect of the present invention provides a multimer complex based on complementary nucleic acid backbones, wherein the complex is a multimer formed by complexing n monomers having complementary nucleic acid backbones, where each monomer is a polypeptide having the single-stranded nucleic acid, and n is a positive integer from 3 to 6. In the multimer, the single-stranded nucleic acid of each monomer and the single-stranded nucleic acids of two other monomers form complementary double-strands by base complementarity to form a complementary nucleic acid backbone structure.
[0009] In another preferred example, n is 3, 4, 5 or 6. In another preferred example, the complex is a trimer, tetramer, or pentamer, and preferably the structure of the complex is as shown in Figure 1.
[0010] In another preferred example, the monomer has the structure of Formula I, Z1-W(I) In the formula, Z1 is a polypeptide moiety, W is a single-stranded nucleic acid sequence, "-" is a linker or a bond. In another preferred example, "-" is a covalent bond.
[0011] In another preferred example, the nucleic acid sequence is selected from the group consisting of left-handed nucleic acids, peptide nucleic acids, locked nucleic acids, sulfur-modified nucleic acids, 2'-fluoro-modified nucleic acids, 5-hydroxymethylcytosine nucleic acids, phosphorodiamidate morpholino nucleic acids, or combinations thereof.
[0012] In another preferred example, in the multimer, Z1 of each monomer is the same or different. In another preferred example, in the multimer, W of each monomer is different.
[0013] In another preferred example, the monomer has the structure of Formula II, D-[L-W]m(II) Here, D is the drug element part of the protein, W is a nucleic acid sequence, L is none or a linker, “-” is a covalent bond, m is 1, 2, or 3. In another preferred example, m is 1.
[0014] In another preferred example, the monomer has the structure of formula III, A-[L-W]m(III) Here, A is the polypeptide antigen element part, W is a nucleic acid sequence, L is none or a linker, “-” is a covalent bond, m is 1, 2, or 3. In another preferred example, m is 1.
[0015] In another preferred example, the nucleic acid sequence W has the structure shown in formula 1, X1-R1-X2-R2-X3(1) Here, R1 is the base complementary pairing region 1, R2 is the base complementary pairing region 2, X1, X2, and X3 are each independently none or redundant nucleic acids, “-” is a bond.
[0016] In another preferred example, the lengths of R1 and R2 are each independently 10 to 20 bases, preferably 14 to 16 bases. In another preferred example, the length of X1 is 0 to 5 bases. In another preferred example, the length of X3 is 0 to 5 bases. In another preferred example, the length of X2 is 0 to 3 bases. In another preferred example, the sequence of X2 is selected from the group consisting of A, AA, AGA, or AAA.
[0017] In another preferred example, R1 of each monomer forms a base complementary pairing structure with R2 of the left adjacent (or left side) monomer, and R2 forms a base complementary pairing structure with R1 of the right adjacent (or right side) monomer.
[0018] In another preferred example, the monomer sequence is selected from any one of the single-stranded nucleic acid sequences SEQ ID No: 1 to 60 (refer to Table 9-1) or a sequence set thereof that form a complementary nucleic acid backbone-based trimer complex.
[0019] In another preferred example, the monomer sequence is selected from any one of the single-stranded nucleic acid sequences SEQ ID No: 61 to 140 (refer to Table 9-2) or a sequence set thereof that form a complementary nucleic acid backbone-based tetramer complex.
[0020] In another preferred example, the monomer sequence is selected from any one of the single-stranded nucleic acid sequences SEQ ID No: 141 to 240 (refer to Table 9-3) or a sequence set thereof that form a complementary nucleic acid backbone-based pentamer complex.
[0021] In another preferred example, the monomer sequence is a phosphorodiamidate morpholino nucleic acid.
[0022] In another preferred example, the monomer sequence is selected from any one of the single-stranded nucleic acid sequences SEQ ID No: 275 to 278 or a sequence set thereof that form a complementary nucleic acid backbone-based tetramer complex.
[0023] The second aspect of the present invention provides a pharmaceutical composition, and the pharmaceutical composition comprises (a) the complementary nucleic acid backbone-based multimer complex described in the first aspect, and (b) It includes a pharmaceutically acceptable carrier. In another preferred example, the pharmaceutical composition includes a vaccine composition. In another preferred example, the pharmaceutical composition includes a therapeutic and / or prophylactic pharmaceutical composition. In another preferred example, the multimer complex includes a trimer complex, a tetramer complex, and a pentamer complex.
[0024] The third aspect of the present invention provides a nucleic acid sequence library, and the nucleic acid sequence library includes a nucleic acid sequence for forming a complementary nucleic acid backbone-based multimer complex described in the first aspect.
[0025] In another preferred example, the nucleic acid sequence (a) a nucleic acid sequence for forming a complementary nucleic acid backbone-based trimer complex, and (b) a nucleic acid sequence for forming a complementary nucleic acid backbone-based tetramer complex, and / or (c) a nucleic acid sequence for forming a complementary nucleic acid backbone-based pentamer complex.
[0026] In another preferred example, the nucleic acid sequence W has a structure shown in Formula 1, X1-R1-X2-R2-X3(1) Here, R1 is a base complementary pairing region 1, R2 is a base complementary pairing region 2, X1, X2, and X3 are each independently none or redundant nucleic acids, “-” is a bond.
[0027] The fourth aspect of the present invention provides the use of the nucleic acid sequence library described in the third aspect in the preparation of the multimer complex described in the first aspect or a pharmaceutical composition containing the multimer complex.
[0028] The fifth aspect of the present invention provides a method for determining a single-stranded nucleic acid sequence for forming a complementary nucleic acid backbone-based multimer complex, including the following steps: (a) Setting the parameters of simulated annealing: Setting the initial annealing temperature, the final annealing temperature, and the annealing temperature decay coefficient ΔT, setting the optimization constraint parameters, (1) The number n of single-stranded nucleic acids, preferably a positive integer from 3 to 6, (2) The paired sequence length L, preferably L is 12 to 16 bases, (3) The dissociation temperature threshold T of the paired region m 、 (4) The free energy threshold ΔG of the specific paired region sequence S °、 (5) The free energy threshold ΔG of non-specific pairing NS °、 (6) The connector X2, preferably A, AA, and AAA, (7) The dissociation temperature threshold T of the secondary structure (hairpin) m-H 、 (8) The CG ratio P within the paired sequence CG 、preferably, the range of P CG is [0.4, 0.6), (9) Optionally, when n = 4, use a symmetric sequence and initialize the sequence set
Number
Number
[0029] In another preferred example, the setting of the simulated annealing parameters in step (a) is For example, the initial annealing temperature T o = 50°C ± 2°C, the annealing end temperature T fSet to 0.12°C ± 0.02°C, and the annealing temperature decay coefficient ΔT is generally 0.98 ± 0.01 depending on the situation. Set the optimization constraint parameters. (1) The number n of single-stranded nucleic acids is a positive integer from 3 to 6. (2) The paired sequence length L is determined according to the situation (preferably, L is 12 to 16 bases). (3) The dissociation temperature threshold T m of the paired region is determined according to the paired sequence length (for example, when L = 14 bases, T m > 50°C, and when L = 16 bases, T m > 52°C). (4) The free energy threshold ΔG S ° of the specific paired region sequence is determined according to the paired sequence length (preferably, when L = 14 bases, ΔG S ° < -27 kcal / mol, and when L = 16 bases, ΔG S ° < -29 kcal / mol). (5) The free energy threshold ΔG NS ° of non-specific pairing is determined according to the sequence length (preferably, ΔG NS ° > -7 kcal / mol). (6) The connector X2 is determined according to the situation (it can be A, AA, AAA, etc.). (7) The dissociation temperature threshold T m-H of the secondary structure (hairpin) is determined according to the situation (preferably, T m-H < 40°C ± 2°C). (8) The range of the CG ratio P CG in the paired sequence is [0.4, 0.6). (9) In particular, when n = 4, a symmetric sequence is used, and according to the above parameters, the sequence set
Number
[0030] In another preferred example, each single-stranded nucleic acid sequence W has a structure shown in Formula 1. X1-R1-X2-R2-X3(1) Here, R1 is the base complementary pairing region 1, R2 is the base complementary pairing region 2, X1, X2, and X3 are each independently none or redundant nucleic acids, "-" is a bond.
[0031] In another preferred example, in step (d), the optimization set is a set that satisfies the following conditions: (C1) In a complementary nucleic acid backbone structure, the free energy (ΔG S °) of the DNA double-stranded structure formed by target pairing is small or minimized, and (C2) In a complementary nucleic acid backbone structure, the ΔG NS ° of non-target pairing is large or maximized.
[0032] In another preferred example, in step (d), the benefit set further satisfies the following conditions: (C3) The dissociation temperature T m > 50 °C for the R1 and R2 regions (when L = 14 bases).
[0033] In another preferred example, in step (c), the free energy (ΔG S °) of the DNA oligomer (i.e., the complementary nucleic acid backbone structure) is calculated by the nearest neighbor method.
[0034] In another preferred example, in step (c), the DNA oligomer (i.e., the complementary nucleic acid backbone structure) is decomposed into 10 different nearest neighbor pair interactions, and these pair interactions are AA / TT; AT / TA; TA / AT; CA / GT; GT / CA; CT / GA; GA / CT; CG / GC; GC / CG and GG / CC. Based on the enthalpy (ΔH°) and entropy (ΔS°) of these pair interactions, the corresponding ΔG° values are calculated respectively, and then the free energies of the pair interactions included in the complementary nucleic acid backbone structure are combined (or summed) to obtain the free energy of the complementary nucleic acid backbone structure.
[0035] In another preferred example, the method repeatedly performs steps (b), (c) and (d) (i.e., executes n1 times of iteration) to obtain an overall optimal solution for the iterative process.
[0036] In another preferred example, in the iterative process, according to the Metropolis criterion, worse solutions are accepted within a limited range so that the algorithm can find an overall optimal solution as much as possible when it ends, and the probability that a worse solution is accepted gradually tends to approach 0.
[0037] In another preferred example, the free energy of the non-target pairing region is optimized by simulating the iteration of simulated annealing using the following optimization objective function,
Number
Number
[0038] The sixth aspect of the present invention provides a set of single-stranded nucleic acid sequences for forming a multimer complex based on a complementary nucleic acid backbone, and the set of single-stranded nucleic acid sequences is determined by the method described in the fifth aspect.
[0039] In another preferred example, the set is selected from the group consisting of the following. (S1) Single-stranded nucleic acid sequences for forming a trimer complex based on a complementary nucleic acid backbone: [Table 1-1] [Table 1-2] [Table 1-3] (S2) Single-stranded nucleic acid sequences for forming a tetramer complex based on a complementary nucleic acid backbone: [Table 2-1] [Table 2-2] [Table 2-3]
[0040] (S3) Single-stranded nucleic acid sequences for forming a pentamer complex based on a complementary nucleic acid backbone: [Table 3-1] [Table 3-2] [Table 3-3]
[0041] The seventh aspect of the present invention provides an apparatus for determining a single-stranded nucleic acid sequence for forming a multimer complex based on a complementary nucleic acid backbone, the apparatus comprising: (M1) An input module used to input simulated annealing parameters, optimization constraint parameters, and an optionally optimized nucleic acid sequence, wherein the setting of the simulated annealing parameters includes an initial annealing temperature, an end annealing temperature, and an annealing temperature decay coefficient ΔT, The optimization constraint parameters are (1) The number n of single-stranded nucleic acids, which is preferably a positive integer from 3 to 6, (2) The paired sequence length L, which is preferably 12 to 16 bases, (3) The dissociation temperature threshold T of the paired region m and (4) The free energy threshold ΔG of the specific paired region sequence S ° and (5) The free energy threshold ΔG of non-specific pairing NS ° and (6) The connector X2, which is preferably A, AA, and AAA, (7) The secondary structure (hairpin) dissociation temperature threshold T m-H and (8) The CG ratio P within the paired sequence, where the range of P CG is preferably in the range of [0.4, 0.6), CG including (M2) An optimization operation module configured to obtain an optimized single-stranded nucleic acid sequence or a set thereof by performing the following sub-steps: (z1) Calculate the objective function value E0 of the initial set S, that is, calculate the sum of the free energies (ΔG NS °) of non-specific pairing between sequences and within the sequences themselves, and at the same time obtain the free energy matrix C n×n of non-specific pairing, search for S i and S j (1 ≦ i ≦ n, 1 ≦ j ≦ n) corresponding to the minimum value in its upper triangular matrix, and S i and S j The free energy ΔG of non-specific pairingNS °(S i, S j ) according to, S i or S j is randomly selected to perform an update operation to obtain a new nucleic acid sequence, and an updated sequence set S’ is obtained. (z2) Determine whether the sequences in the set S’ described in the previous step meet the set optimization constraint parameter conditions, and the dissociation temperature T of the specific pairing region m , the free energy ΔG of the specific pairing region sequence S °, the dissociation temperature T of the secondary structure m-H and the CG ratio P CG Verify the parameters including. If the above parameters meet the constraint conditions, execute step (z3). Otherwise, repeat step (z2). If S’ that meets the conditions cannot be obtained after continuously executing the steps 15 times under a specific annealing temperature, to prevent infinite loop, set set S to set S’, and execute the next step. (z3) Calculate the objective function value E1 of the set S’ described in the previous step, compare E0 and E1. If E1≧E0, it indicates that the free energy of non-specific pairing is optimized, and the sequence set S’ becomes the sequence set S. If E1<E0, it indicates that the free energy of non-specific pairing is not optimized. In this case, it is necessary to determine whether to set S’ to S according to the Metropolis criterion, and (z4) Decay the annealing temperature according to the set decay coefficient ΔT, and repeat steps (z1), (z2) and (z3) for S described in the previous step, that is, based on Monte Carlo simulated annealing, until the annealing temperature reaches the annealing end temperature, for the previous step described
Number
[0042] In another preferred example, when the optimization constraint parameter is a tetramer (n = 4), a symmetric array is used, and according to the above parameters, a set of arrays
Number
Advantages of the Invention
[0043] It should be understood that within the scope of the present invention, new or preferred technical solutions can be formed by combining each of the above technical features of the present invention with the technical features specifically described below (for example, in the embodiments). Due to space limitations, it will not be repeated here.
Brief Description of the Drawings
[0044]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
BEST MODE FOR CARRYING OUT THE INVENTION
[0045] As a result of extensive and detailed research, the inventors have developed for the first time multivalent protein drugs and their libraries, as well as their preparation methods and applications. Using the drug library and preparation method of the present invention, if necessary, short-acting protein drugs can be formed into multivalent complexes quickly, efficiently, at low cost, and in high yield to improve the drug half-life, or monomeric antigens can be formed into high-cost antigens to improve their immunogenicity. Based on this, the present invention has been completed.
[0046] Specifically, the present invention provides a multivalent protein drug comprising n protein drug units, wherein each drug unit comprises the same kind of drug element part and different nucleic acid element parts bound to the drug element part, n is a positive integer of ≧2, and the n different nucleic acid element parts form an n-mer by nucleic acid base complementarity, thereby constituting the multivalent protein drug. The multivalent protein drug of the present invention forms a stable complementary nucleic acid base pairing structure by rapid assembly (for example, within 1 minute) (not by complex peptide bonds or other chemical modifications, etc.). According to experiments, the drug of the present invention shows that the half-life in vivo of animals can be extended by increasing the molecular weight by increasing the price.
[0047] Also, based on the same embodiment, the drug element may also be an antigen used in the research and development of vaccines. The difference is that each antigen unit comprises the same or different antigen element parts and different nucleic acid element parts bound to the antigen element part, and the n different nucleic acid element parts form an n-mer by nucleic acid base complementarity, thereby constituting the multivalent antigen. Finally, the present invention provides a highly optimized nucleic acid sequence library containing a group of nucleic acid sequences that are efficient and accurately assembled into dimers to pentamers for rapidly and accurately self-assembling the above-described drug or antigen unit into a multivalent polymer complex.
[0048] The term Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. As used herein, when used in relation to a specifically recited value, the term "about" refers to the possibility that the value may vary by no more than 1% from the recited value. For example, as used herein, the expression "about 100" includes all values between 99 and 101 (e.g., 99.1, 99.2, 99.3, 99.4, etc.).
[0049] Drug D In the present invention, the drug element moiety is a protein drug or a polypeptide drug. Typically, the protein drug includes, but is not limited to, cytokines, hormones (e.g., insulin, growth hormone, etc.), antibody drugs, and polypeptides. In a preferred embodiment of the present invention, the protein drug is G-CSF for treating leukopenia.
[0050] Antigen library A The present invention provides an antigen library, which includes N antigen units, wherein the antigen unit includes an antigen element moiety and a nucleic acid element moiety bound to the antigen element moiety, and different nucleic acid element moieties form an n-mer by nucleic acid base complementarity, thereby constituting a multivalent antigen, wherein the antigen element moiety is a protein antigen or a polypeptide antigen, Typically, the protein antigen and polypeptide antigen include, but are not limited to, viral and bacterial proteins, their structural regions, and fragments.
[0051] In the present invention, the antigen element part is selected from M different antigen proteins of the library, M ≤ N, and the M different antigen proteins include different proteins of a specific virus or mutants of the same protein of different virus strains. In a preferred embodiment of the present invention, the protein antigen is derived from the novel coronavirus SARS-CoV-2, specifically, it is an expensive antigen formed by the receptor binding domain (RBD) of the viral spike protein.
[0052] Left-handed nucleic acid The left-handed nucleic acid refers to an existence like a mirror image with respect to the right-handed nucleic acid (D-nucleic acid) existing in nature, and is divided into left-handed DNA (L-DNA) and left-handed RNA (L-RNA). The left-handedness (chiral center) mainly exists in the deoxyribose or ribose part of the nucleic acid and is mirror-inverted. Therefore, the left-handed nucleic acid is not decomposed by nucleases (for example, exonuclease, endonuclease) ubiquitous in plasma.
[0053] Preparation method 1. Design and preparation of the L-nucleic acid strand framework According to the present invention, the L-nucleic acid strand framework is formed by base pairing of two or more L-nucleic acid single strands. The 5' or 3' end of each L-nucleic acid single strand is all activated to a group (such as NH2, etc.) for subsequent modification, and then one end of a linker (such as SMCC, SBAP, etc.) is coupled to the activated group on the L-nucleic acid single strand. The L-nucleic acid having a linker can be assembled into a desired L-nucleic acid strand framework. In another preferred example, the activating functional groups (such as aldehyde, maleimide, etc.) at the 5' or 3' end of the L-nucleic acid single strand are already included in nucleic acid synthesis. After it is confirmed that the L-nucleic acid having a linker can successfully self-assemble into a framework, the L-nucleic acid single strand having a linker can be respectively coupled to an antibody for subsequent assembly. The L-nucleic acid framework of the present invention can basically be prepared by the following steps.
[0054] 1.1. Design of L-nucleic acid single strands that can rapidly self-assemble Determine the required polyvalence n (such as trimer, tetramer), determine the number of L-nucleic acid single strands n required according to the polyvalence n, design the corresponding number of L-nucleic acid single strand sequences, and optimize the base pairing to adjust the stability of the target nucleic acid framework and reduce the possibility of non-specific pairs between nucleic acid strands. The details of nucleic acid sequence design will be specifically described in the summary of the invention and the examples.
[0055] 1.2. Activation of L-DNA or L-RNA The activation of the L-nucleic acid includes the modification of the active group at its 5' end (X1) or 3' end (X3), and subsequent linker coupling. The modification of the active group can be customized by nucleic acid synthesis companies. The linker generally has a bifunctional group, that is, one end can be coupled to the active group of the nucleic acid, and the other end can be bound to a specific site (such as NH3, SH) on the protein.
[0056] According to a preferred embodiment of the present invention, all L-nucleic acids constituting the framework can have an aldehyde modification added at the 5'-end to complete the activation of the L-nucleic acid and then be coupled to the N-terminal α-amine of the protein.
[0057] 2. Method for preparing protein-L-nucleic acid complex First, the 5'-end or 3'-end of the L-nucleic acid is modified with an aldehyde, and then, under low pH (5-6) conditions, the aldehyde group of the L-nucleic acid is specifically bound to the N-terminal NH3 of the protein through a reductive amination reaction.
[0058] Algorithm and nucleic acid sequences optimized by the algorithm The present invention further provides a method and apparatus for determining a single-stranded nucleic acid sequence for forming a multimer complex based on a complementary nucleic acid backbone. Preferably, the method includes the preferred algorithm of the present invention.
[0059] Typically, a nucleic acid sequence library optimized by the computer algorithm of the present invention can include (a) a nucleic acid sequence capable of forming a trimer by self-assembly of base complementary pairing, (b) a nucleic acid sequence capable of forming a tetramer by self-assembly of base complementary pairing, and (c) a nucleic acid sequence capable of forming a pentamer by self-assembly of base complementary pairing. Typical complexes formed by representative nucleic acid sequences are, for example, the trimer, tetramer, and pentamer molecules shown in FIG. 1.
[0060] Preferably, the nucleic acid sequence W described in the present invention has a structure shown in Formula 1, W = X1-R1-X2-R2-X3 (1) Here, R1 is a base complementary pairing region 1, R2 is a base complementary pairing region 2, X1, X2, and X3 are each independently none or redundant nucleic acids, The lengths of R1 and R2 are 14 to 16 bases, the lengths of X1 and X3 are 0 to 5 bases, the length of X2 is 0 to 3 bases, and the sequence is A, AA, AGA, or AAA, wherein R1 of the nucleic acid sequence forms a target pairing with R2 of a different nucleic acid sequence, and R2 of the nucleic acid sequence forms a target pairing with R1 of another nucleic acid sequence. In the present invention, the self-assembly of any region of the nucleic acid sequence belongs to non-target pairing, which needs to be avoided in the design.
[0061] Preferably, the nucleic acid sequence described in the present invention can be designed or optimized by a computer algorithm of simulated annealing (SA).
[0062] The computer algorithm minimizes the free energy (ΔG°) of the DNA double-stranded structure formed by target pairing, and at the same time maximizes the ΔG° of non-target pairing, Specifically, the stability of the DNA double-stranded structure depends on each nearest-neighbor base pair in the sequence, and there may be 10 different nearest-neighbor interactions in any Watson-Crick DNA double-stranded structure, and the interactions of these pairs are AA / TT; AT / TA; TA / AT; CA / GT; GT / CA; CT / GA; GA / CT; CG / GC; GC / CG; GG / CC.
[0063] More specifically, the ΔG° values of the 10 base pairs described above can be calculated by enthalpy (ΔH°) and entropy (ΔS°) at any temperature, and Table 1 statistically presents the enthalpy, entropy, and free energy data of 10 sequences.
Table 4
[0064] Using the thermodynamic values in Table 1, the enthalpy value ΔH° and free energy value ΔG° of the DNA oligomer can all be effectively predicted by the nearest neighbor method. Taking the complementary pairing of GGAATTCC / CCTTAAGG as an example, ΔG° = -14.63 kcal / mol is calculated by the nearest neighbor method.
[0065] In addition to optimizing the ΔG NS ° value, the simulated annealing implements the constraint of the dissociation temperature (T m ) of the nucleic acid sequence pair to ensure that T m > 50°C (when L = 14 bases) for the R1 and R2 regions.
[0066] The nearest neighbor model accurately predicts the stability of the DNA double strand based on thermodynamic calculations. The prediction of the model for a given base sequence is based on the nearest neighbor base pair. The calculation of its melting temperature is related to enthalpy (ΔH°) and entropy (ΔS°), and the calculation method is as follows:
Equation
[0067] Simulated annealing is a common probabilistic algorithm, which is a method for approximately solving optimization problems designed based on the Monte Carlo concept. Its purpose is to find an approximate optimal solution in a wide search space within a certain period. The idea of simulated annealing is derived from the annealing process of solid substances in physics: first, the solid is heated sufficiently and then cooled slowly. When heating, the internal energy of the particles inside the solid increases as they move freely. Then, as the temperature gradually decreases, the particles become orderly and tend to reach an equilibrium state at each temperature. When the temperature decrease rate near the condensation point is sufficiently slow, the ground state is reached and the internal energy is minimized. According to the Metropolis criterion, the probability that a particle tends to be in equilibrium at temperature T is exp(-ΔE / (KT)), where E is the internal energy at temperature T, ΔE is the change amount, and K is the Boltzmann constant. There is a similar process for combinatorial optimization problems. The solid microstate i is simulated as the solution X, the objective function corresponds to the internal energy in state i, and the solid temperature is simulated with the control parameter T. For each value of T, the iterative process of "generating a new solution → calculating the difference of the objective function → determining whether to accept it → accepting / rejecting" is repeated, and the T value is gradually attenuated. In the iterative process, according to the Metropolis criterion, more inappropriate solutions are marginally accepted, and the probability of more inappropriate solutions tends to gradually approach 0. Therefore, the algorithm can find the overall optimal solution as much as possible at the end.
[0068] In the present invention, the free energy of the non-target pairing region between single strands of the multimeric nucleic acid is simulated as the internal energy, and under the condition of ensuring the free energy and dissociation temperature of the target pairing region between nucleic acid single strands, the free energy of the non-target pairing region is optimized by iteration of simulated annealing. Finally, the optimized single strand is beneficial to the assembly of the multimer. The optimization objective function used is as follows.
Equation
[0069] ΔG°(S i, S j ) is the free energy of non-target pairing between array S i and array S j and is
Number
Number
[0070] The flowchart of a typical algorithm of the present invention is as shown in FIG. 2. In the present invention, a random element is introduced into simulated annealing, and in the process of each iterative update, a solution worse than the current one can be accepted with a certain probability, so it is possible to jump out of the local optimum solution and reach the global optimum solution.
[0071] Multivalent polymer complex The present invention further provides a multivalent polymer complex formed by mediating a protein drug using a nucleic acid multimer designed by the above algorithm, and the complex has an improved drug half-life and activity.
[0072] Preferably, the nucleic acid sequence is a group of nucleic acid sequences specifically assembled into an n-mer in a nucleic acid sequence library. Preferably, in the present invention, the protein drug is a protein drug that needs to be multivalent to improve the half-life or activity.
[0073] Typically, each nucleic acid strand of the nucleic acid sequence group binds to the protein drug respectively to form a protein drug-nucleic acid strand unit having the structure shown in Formula 2, D-[L-W i , i = 1 to n (2) Here, D is a protein drug element part, each Wi is independently a nucleic acid sequence, and the nucleic acid sequence is selected from the group consisting of left-handed nucleic acid, peptide nucleic acid, locked nucleic acid, sulfur-modified nucleic acid, 2'-fluoro-modified nucleic acid, 5-hydroxymethylcytosine nucleic acid, or a combination thereof, and the nucleic acid sequence has the structure shown in Formula 1 selected from the nucleic acid sequence group capable of forming the n-mer described above, L is a linker, and the linker part is already included in the synthesis or preparation of Wi and binds to X1 or X3 of Wi (refer to Formula 1), "-" is a covalent bond.
[0074] In another preferred example, the drug element part is selected from the group consisting of protein drugs and polypeptide drugs that need to increase the molecular weight to improve the half-life, and protein drugs and polypeptide drugs that need to form multivalency to improve the activity.
[0075] In another preferred example, L has a functional group such as aldehyde, NHS ester or the like near the D terminus, which is used to bind to the N-terminal α-amine or lysine e-amine on D.
[0076] In another preferred example, L has a maleimide functional group or a haloacetyl (e.g., bromoacetyl, iodoacetyl, etc.) functional group near the D terminus, which is used to bind to the free thiol (-SH) functional group on D.
[0077] In another preferred example, D is selected from the group consisting of natural proteins, recombinant proteins, chemically modified proteins, and synthetic polypeptides. In another preferred example, D can have site-specific modifications or site-specific additions of amino acids in order to bind to the L-W of Formula 1 i
[0078] The protein drug can bind to different L-Ws that can aggregate into an n-mer, i forming D-[L-W1], D-[L-W2], …, D-[L-W n which are self-assembly units of the protein drug, D-[L-W1], D-[L-W2], …, D-[L-W n are assembled into a multivalent molecular complex of the protein drug by equimolar mixing in solution.
[0079] The present invention uses a nucleic acid multimer designed by the above algorithm to form a multivalent polymer complex by mediating one or more types of antigens, thereby improving the effect of the vaccine when inducing neutralizing antibodies in vivo. Here, the nucleic acid sequence is a group of nucleic acid sequences that specifically aggregate into an n-mer in a nucleic acid sequence library. Here, the antigen is an antigen or an antigen library, the antigen library contains M different antigen proteins, and 1 ≤ M ≤ n.
[0080] Each nucleic acid strand of the nucleic acid sequence group binds to the antigen in the antigen library respectively, forming an antigen-nucleic acid strand unit having the structure shown in Formula 3. Ak-[L-W i , i = 1 - n, k = 1 - M (3) Here, Ak is the antigen k of the antigen library, and one Ak corresponds to one or more types of L-W i (for example: A1-[L-W1], A1-[L-W2], A2-[L-W3], A3-[L-W4]), Another embodiment of Formula 3 is the same as Formula 3 above, the antigen protein can aggregate into an n-mer and bind to different L-Ws i respectively to form antigen self-assembly units such as A1-[L-W1], A2-[L-W2], …, A3-[L-W n . The antigen self-assembly units are assembled into a multivalent antigen complex by equimolar mixing in solution.
[0081] The main advantages of the present invention are as follows. (1) The present invention can complete the multivalentization of a ready-made short-acting protein drug without the need for the reconstruction of a fusion protein or complex chemical modification and cross-linking, and can improve its half-life and activity. The aldehyde modification of L-nucleic acid can specifically bind to the N-terminal amine of the protein to form a protein drug unit that can self-assemble into an oligomer. (2) The protein drug unit (protein-nucleic acid linker) of the present invention can complete the multivalentization of the protein drug within 1 minute by utilizing the mediation of the left-handed nucleic acid strand. (3) For vaccine development, the present invention can form a monomeric protein antigen at high cost and improve its immunogenicity. (4) For vaccine development, the present invention can also aggregate antigen mutants and subtypes of different virus or bacterial strains into a multivalent antigen to induce a wider range of neutralizing antibodies.
[0082] Hereinafter, the present invention will be further described together with specific examples. It should be understood that these examples are only used to explain the present invention and do not limit the scope of the present invention. In the following examples, experimental methods without indicating specific conditions usually follow conventional conditions such as those described in, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or conditions proposed by the manufacturer. Unless otherwise specified, percentages and parts are percentages by weight and parts by weight.
[0083] Example 1: Design, synthesis and verification of trimeric nucleic acid backbone aggregates Design nucleic acids that can form three pairs according to the shape in Fig. 1(A). Specifically, the single-stranded R1 nucleic acid undergoes specific complementary pairing with the single-stranded R6 nucleic acid but does not pair with other single-stranded nucleic acids. Similarly, R3 and R5 each form specific complementary pairing with R2 and R4, respectively, but do not pair with other single-stranded nucleic acids. The free energy (ΔG S °) of specific complementary pairing is much smaller than the free energy (ΔG NS °) of non-specific pairing. The free energy (ΔG S °) of specific complementary pairing is less than -29 kcal / mol, while the free energy (ΔG NS °) of non-specific pairing is all greater than -7 kcal / mol. Thus, the trimeric form is the most stable in the reaction system. The specific implementation steps of trimer optimization are as described in the schematic diagram in Fig. 2, and its annealing parameters are: initial annealing temperature T o = 50 °C, final annealing temperature T f = 0.12 °C, annealing temperature decay coefficient ΔT = 0.98 (the initial annealing temperature decays to 0.98 of the current value each time), and the optimization constraint parameters are: length of the paired sequence L = 16 bases, dissociation temperature threshold T m > 54 °C, free energy threshold of the paired sequence: ΔG S ° < -29 kcal / mol, free energy threshold of non-specific pairing: ΔG NSIt is restricted to > -7 kcal / mol. The specific implementation steps of the optimization are as follows.
[0084] Array initialization: Initialize the parameter alignment array
Number
Number
[0085] Generation of a new solution: The new solution S’ is obtained by updating all the array set S. First,
Number
[0086] Optimization judgment: Calculate the objective function values E of sets S and S’ according to formula 2 i and E i+1 respectively, and for E i+1 -E i If ≧0, it means that the free energy of non-target pairing (ΔG NS °) is optimized by the update, then S←S’, and the new solution S’ becomes S. E i+1 -E i If <0, it means that a deteriorated solution is obtained by the update. According to the Metropolis criterion, calculate the probability p = exp(-ΔE / T o ), and at the same time generate r, rε[0,1) randomly. If p>r, accept this deteriorated solution, otherwise reject the deteriorated solution. Finally, the annealing initial temperature decays to the current 0.98, generates the next new solution, and obtains the optimized array set until it decays to the annealing end temperature.
[0087] The optimized arrays in Table 2_1 are obtained by the optimization of the above algorithm. The purpose of the 5'-terminal A of each array is the modification of the active group and subsequent linker coupling. In the free energy (ΔG NS °) row table and target pairing region parameter index table, some main parameter values are statistically calculated. For the trimer-optimized sequence-specific pairing pattern diagram, see Figure 3. From the corresponding nucleic acid backbone running diagram (Figure 4), it can be seen that Lane 9 is an artificially designed trimer, and its band is unclear and trailing. Lane 10 is the main band formed by the optimized sequences S1, S2 and S3, indicating that a trimer is formed and shows very high stability.
[0088]
Table 5
[0089]
Table 6
[0090]
Table 7
[0091] Example 2: Design, Synthesis, and Verification of a Tetrameric Nucleic Acid Skeleton Assembly Design nucleic acids that can pair according to the shape of Fig. 1(B). Here, the single-stranded nucleic acids R1, R3, R5, and R7 each perform specific complementary pairing with the single-stranded nucleic acids R8, R2, R4, and R6, but do not pair with other single-stranded nucleic acids. The free energy (ΔG S °) of specific complementary pairing is much smaller than the free energy (ΔG NS °) of non-specific pairing. The free energy (ΔG S °) of specific complementary pairing is less than -27.4 kcal / mol, while the free energy (ΔG NS °) of non-specific pairing is all greater than -7 kcal / mol. Thus, the tetrameric form is the most stable in the reaction system. The specific implementation steps of tetramer optimization are as described in the schematic diagram in Fig. 2. The annealing parameters are an initial annealing temperature T o = 50 °C, a final annealing temperature T f = 0.12 °C, and an annealing temperature decay coefficient ΔT = 0.98 (the initial annealing temperature decays to 0.98 of the current value each time). The optimization constraint parameters are a paired sequence length L = 14 bases, a dissociation temperature threshold T m > 52 °C, a free energy threshold for the paired sequence: ΔG S ° < -27.4 kcal / mol, and a free energy threshold for non-specific pairing ΔG NS ° > -7 kcal / mol.
[0092] The specific implementation steps of the optimization are as follows. Array initialization: Initialize the parameter paired array
Number
Number
[0093]
Table 8
[0094] New solution generation: The same as the generation of the new solution in Example 1. On the other hand, when the fixed core structure is used, for the update, it does not include the fixed core structure part. To strictly comply with the parameter constraints, the dissociation temperature (T m ) of the pairing region of the new nucleic acid sequence should be higher than 52 °C, and at the same time, the free energy (ΔG S °) of the pairing region should be less than -27.4 kcal / mol. Optimization judgment: The same as the optimization judgment in Example 1.
[0095] The sequences after optimization in Table 4_1 are obtained by the optimization of the above algorithm. The schematic diagram of its specific pair is as shown in FIG. 5. The optimized sequence of the tetramer, the broken line statistical chart of the free energy of the non-target during optimization is as shown in 6. When the present invention adds a connector, in order to pursue that it does not have a great impact on the free energy (ΔG NS °) of the optimized sequence without adding a connector, regardless of whether a connector is added, the free energy (ΔG NS(°) Since the sum of the rows and columns should be as large as possible, during optimization, the objective function values of the optimized arrays without adding connectors were statistically analyzed, and only the optimized arrays with added connectors were finally detected. From the corresponding nucleic acid backbone running diagram (Figure 7), lane 15 is a band of artificially designed tetramers with a trailing phenomenon, and lane 16 is a main band of approximately 100 bp formed by sequences S1, S2, S3, and S4 after algorithm optimization, indicating that a tetramer was formed and showing very high stability.
[0096] During the above-mentioned tetramer implementation stage, the main optimization is the free energy of non-specific pairing between sequences. However, the secondary structure formed by sequence self-folding also has a great impact on the tetramer assembly. If the dissociation temperature of the secondary structure formed by the sequence itself is too high, once such a stable secondary structure is formed, it becomes difficult to break such a state, and the tetramer assembly becomes difficult. Therefore, the dissociation temperature of the secondary structure corresponding to the four sequences that need to control the tetramer assembly should not be too high. In the case of tetramers, it is necessary to control the dissociation temperature of the secondary structures of the four nucleic acid sequences. When using a symmetric structure (R1, R3, R5, R7 in Figure 1 maintain symmetry with R4, R6, R8, R2 respectively), similar secondary structures appear between the two sequences, and there is no significant difference in the dissociation temperatures of the secondary structures of both. In this case, since the above-mentioned effect can be obtained only by controlling the secondary structures of the two nucleic acid sequences, symmetry is advantageous for controlling the secondary structure in tetramer optimization.
[0097]
Table 9
[0098]
Table 10
[0099]
Table 11
[0100] Example 3: Design, Synthesis, and Verification of a Pentameric Nucleic Acid Skeleton Aggregate The optimized sequence of the tetramer shows a very good aggregation effect mainly because the central part of the tetramer does not form pairing. That is, due to the fixed core structure, a complex structure is not formed in the central region, and sufficient freedom is given to the tetramer. Therefore, in order to integrate the tetramer sequence and the fixed core structure in the optimization of the pentamer sequence, two schemes using the tetramer sequence and the fixed core structure of Example 2 are considered and designed.
[0101] The first conversion scheme: Design nucleic acids that can form five pairs according to the shape of Fig. 8(B). This scheme retains the partial sequence of the tetramer and the complete fixed core structure of Example 2, and the core structure remains closed. At R1 and R8 of the original tetramer, except for the core structure, the tetramer is opened, and two nucleic acid sequences R9 and R with a length of 14 are added. Here, the single-stranded nucleic acids R1, R3, R5, R7, and R9 specifically pair with the single-stranded nucleic acids R, R2, R4, R6, and R8 respectively, but do not pair with other single-stranded nucleic acids. The free energy () of the specific complementary pairing is less than -27.4 kcal / mol, but the free energy (ΔG°) of the non-specific pairing is greater than -7 kcal / mol. Thus, the pentamer form is the most stable in the reaction system. The specific implementation steps for optimizing the first conversion scheme of the pentamer are as described in the schematic diagram 2, and its annealing parameters are: the initial annealing temperature T = 50 °C, the final annealing temperature T = 0.12 °C, and the annealing temperature decay coefficient ΔT = 0.9 (the initial annealing temperature decays to 0.9 of the current value each time, and the use of the tetramer partial sequence makes the updated region smaller and the decay of the annealing temperature faster). The optimization constraint parameters are: the length of the paired sequence L = 14 bases, the dissociation temperature threshold T > 52 °C, the free energy threshold of the paired sequence: ΔG° < -27.4 kcal / mol, and the free energy threshold of the non-specific pairing ΔG 10 are added. Here, the single-stranded nucleic acids R1, R3, R5, R7, and R9 specifically pair with the single-stranded nucleic acids R 10 , R2, R4, R6, and R8 respectively, but do not pair with other single-stranded nucleic acids. The free energy () of the specific complementary pairing is less than -27.4 kcal / mol, but the free energy (ΔG NS °) of the non-specific pairing is greater than -7 kcal / mol. Thus, the pentamer form is the most stable in the reaction system. The specific implementation steps for optimizing the first conversion scheme of the pentamer are as described in the schematic diagram 2, and its annealing parameters are: the initial annealing temperature T o = 50 °C, the final annealing temperature T f = 0.12 °C, and the annealing temperature decay coefficient ΔT = 0.9 (the initial annealing temperature decays to 0.9 of the current value each time, and the use of the tetramer partial sequence makes the updated region smaller and the decay of the annealing temperature faster). The optimization constraint parameters are: the length of the paired sequence L = 14 bases, the dissociation temperature threshold T m > 52 °C, the free energy threshold of the paired sequence: ΔG S ° < -27.4 kcal / mol, and the free energy threshold of the non-specific pairing ΔG NSRestrict to > -7 kcal / mol.
[0102] For such a scheme, a complete tetramer-fixed core structure is used, nucleic acid sequence W5 is developed based on nucleic acid sequence 4, nucleic acid sequence W6 is developed based on nucleic acid sequence 4, W5 has the structure of formula 7, W6 has the structure of formula 8, W3 = X1 - R1 - X2 - C1 - X2 - C2 - Q1 - X3(5) W4 = X1 - Q1 - C1 - X2 - C2 - X2 - R1 - X3(6) Such a scheme includes an array of four structures such as formula 1, formula 4, formula 5 and formula 6.
[0103]
Table 12
[0104] The specific implementation stages of optimization are as follows. Parameter initialization alignment array
Number
Number
[0105] Generation of a new solution: Randomly select and update one sequence from R1, R8, R9 and R 10 to obtain a new nucleic acid sequence, whether the dissociation temperature of this nucleic acid sequence is higher than 52°C, and at the same time the free energy of the pairing region (ΔG SCheck whether it is less than -27.4 kcal / mol, and if the dissociation temperature and the free energy of the pairing region (ΔG S °) do not meet the constraint requirements, repeat the update. When the dissociation temperature and the free energy of the pairing region (ΔG S °) meet the constraint requirements, update S according to the principle of base complementary pairing to finally obtain a new sequence set S’. Even after 15 updates, if the obtained new nucleic acid sequence still does not meet the constraint requirements of the dissociation temperature and the free energy of the pairing region (ΔG S °) all the time, set S to the new solution S’ to prevent infinite loop. Therefore, such a scheme needs to use a complete fixed core structure and part of the tetramer sequence, and ensure that the fixed core structure and the retained tetramer sequence part do not change during the generation of new solutions.
[0106] Optimization judgment: The same as the optimization judgment in Example 1. On the other hand, the initial annealing temperature decays to the current 0.9. The sequence after optimization of the first conversion scheme of the pentamer in Table 6_1 is obtained by the optimization of this algorithm. Figure 9 is a broken line statistical chart of the total free energy values of the non-target pairing regions between sequences during this optimization (the free energies of the non-target pairing regions between S3 and S3, S3 and S4, and S4 and S4 do not participate in the statistics). Lane 9 in Figure 10 is the main band formed by the sequence set after optimization, indicating that a pentamer is formed and showing very high stability.
[0107]
Table 13
[0108]
Table 14
[0109]
Table 15
[0110] Second conversion scheme: Design nucleic acids that can form five pairs according to the shape in Fig. 8(C). This scheme only retains the fixed core structure sequence of the tetramer. To adapt to the pentamer, the core structure is opened at R1 and R2, and the other parts are randomly generated and then the pentamer needs to be optimized. Here, the single-stranded nucleic acids R 10 , R2, R4, R6, and R8 specifically pair with the single-stranded nucleic acids R1, R3, R5, R7, and R9 respectively, but do not pair with other single-stranded nucleic acids. The free energy (ΔG S °) of specific complementary pairing is less than -27.4 kcal / mol, while the free energy (ΔG NS °) of non-specific pairing is greater than -7.2 kcal / mol. Thus, the pentamer form is the most stable in the reaction system. The specific implementation steps for optimizing the second conversion scheme of the pentamer are as described in schematic diagram 2, and its annealing parameters are: initial annealing temperature T o = 50 °C, final annealing temperature T f = 0.12 °C, annealing temperature decay coefficient ΔT = 0.98 (the initial annealing temperature decays to 0.98 of the current value each time), and the optimization constraint parameters are: length of the paired sequence L = 14 bases, dissociation temperature threshold T m > 52 °C, free energy threshold of the paired sequence: ΔG S ° < -27.4 kcal / mol, free energy threshold of non-specific pairing ΔG NS ° > -7.2 kcal / mol.
[0111]
Table 16
[0112] The specific implementation steps of the optimization are as follows. Array initialization: According to the optimization constraint parameters, a set of arrays with a length of 9 excluding the paired sequence R9 and the core structure
Number
Number
[0113] Generation of new solutions: The same as the generation of new solutions in Example 1. On the one hand, this scheme uses a fixed core structure of the tetramer, and it is necessary to ensure that the fixed core structure part does not change during the generation of new solutions. At the same time, strictly abide by the parameter constraints, and the dissociation temperature for updating the pairing region of the nucleic acid sequence exceeds 52 °C, and at the same time, the free energy (ΔG S °) of the pairing region is less than -27.4 kcal / mol. Optimization judgment: The same as the optimization judgment in Example 1.
[0114] The optimized sequence of the second conversion scheme of the pentamer in Table 8_1 is obtained by the optimization of this algorithm. The schematic diagram of its specific pairing is as shown in Figure 11. Figure 12 is a broken line statistical chart of the total free energy value of the non-target pairing region between sequences during this optimization. Lane 35 in Figure 13 is the main band formed by sequences S1, S2, S3, S4, and S5. The pentamer has a slight tail, but the aggregation effect is good.
[0115]
Table 17
[0116]
Table 18
[0117]
Table 19
[0118] Furthermore, by repeating Examples 1, 2, and 3, single-stranded nucleic acid sequences and sets thereof used to form complementary nucleic acid backbone-based trimer, tetramer, and pentamer complexes shown in Table 9-1, Table 9-2, and Table 9-3 (see above) are obtained.
[0119] Example 4: Coupling of G-CSF and L-DNA For the coupling of G-CSF and L-DNA, the reductive amination reaction method is adopted to selectively couple L-DNA with an aldehyde group modification at the 5'-end to the N-terminus of G-CSF. By methods such as dilution or gel filtration chromatography, the buffer solution of G-CSF is replaced with an acetate buffer solution (20 mM acetic acid, 150 mM NaCl, pH 5.0), and the sample is concentrated to 30 mg / mL. Take 100 OD (1 OD = 33 μg) of L-DNA dry powder with an aldehyde group modification at the 5'-end and dissolve it in 60 μL of acetate buffer solution. Dissolve 30 mg of sodium cyanoborohydride in acetate buffer solution and adjust the concentration to 800 mM. Take 50 μL, 60 μL, and 20 μL of G-CSF, L-DNA, and sodium cyanoborohydride at the above concentrations respectively, mix them uniformly, and incubate and react for 48 hours while rotating in the dark at room temperature. The samples before and after the reaction are verified for the effect of the coupling reaction by polyacrylamide gel electrophoresis. The (L-DNA)-(G-CSF) conjugate shows an obvious shift in the electrophoresis gel diagram compared with uncoupled G-CSF, and the coupling efficiency may reach about 70% - 80% (Figure 14).
[0120] Example 5: Purification of (L-DNA)-(G-CSF) Conjugate The purification of the (L-DNA)-(G-CSF) conjugate is divided into two steps. In the first step, Hitrap Q HP is used to remove unreacted G-CSF and the (L-DNA)2-(G-CSF) conjugate bound to two L-DNAs (Figure 15a). The reaction mixture obtained in Example 5 is diluted 10-fold with the loading buffer of the Q column and then loaded. 10 column volumes of the loading buffer are eluted to remove unreacted G-CSF. Then, using a linear gradient elution method from 0 to 100%, 50 column volumes are eluted to separate the (L-DNA)-(G-CSF) conjugate and the (L-DNA)2-(G-CSF) conjugate. Polyacrylamide gel electrophoresis is used to identify the types of components of each A280 absorption peak (Figure 15b), and the (L-DNA)-(G-CSF) conjugate (including unreacted nucleic acids) is collected. In the second step, Hiscreen Capto MMC is used to remove unreacted nucleic acids, and finally, a highly pure (L-DNA)-(G-CSF) conjugate (Figure 15c) is obtained. The sample collected in Step 1 is directly loaded onto the Hiscreen Capto MMC column, 10 column volumes of the loading buffer are eluted to remove unreacted nucleic acids, and then the (L-DNA)-(G-CSF) conjugate is eluted with 100% elution buffer.
[0121] The purification conditions are as shown in the following table.
Table 20
[0122] Using the (L-DNA)-(G-CSF) conjugate sample obtained by the two-step purification method, the purity of the sample was identified by 2% agarose gel electrophoresis (Figure 15d). The gel diagram shows that there is only one nucleic acid band in the sample, and the (L-DNA)-(G-CSF) conjugate shows an obvious shift in the electrophoresis gel diagram compared with the uncoupled L-DNA, which indicates that the unreacted nucleic acid and the (L-DNA)2-(G-CSF) conjugate have been removed.
[0123] Example 6: Assembly of monovalent, divalent and trivalent G-CSF conjugates The nucleic acid concentrations of S1-G-CSF, S3-G-CSF, S4-G-CSF, S2, S3, and S4 were measured by Nanodrop respectively. According to the structural design of the monovalent, divalent and trivalent protein conjugates, the above components were quantitatively taken and quantified, and assembled according to the molar ratio of 1:1:1:1. After mixing, each assembled unit was automatically assembled according to the principle of base complementary pairing. The samples before and after assembly were confirmed for the assembly effect and sample purity by polyacrylamide gel electrophoresis (Figure 16).
[0124] Example 7: In vitro activity evaluation of G-CSF M-NFS-60 cells (mouse leukemia lymphocytes / G-CSF-dependent cells) were inoculated into a resuscitation medium (RPMI1640 + 10% FBS + 15 ng / mL of G-CSF + 1X penicillin-streptomycin), and the cells were resuscitated under the conditions of 37 °C and 5% CO2. After the cell density reached 80% - 90%, the cells were passaged. After 2 - 3 passages, the cells were plated in 96 wells. The cell plating experiment used a corning 3599#96-well plate, and the cell plating density was 6000 cells / well. Various samples (GCSF, NAPPA4-GCSF, NAPPA4-GCSF2, NAPPA4-GCSF3) were serially diluted, and the working concentrations were (0.001, 0.01, 0.1, 1, 10, 100 ng / mL), and the final volume was 100 μL. PBS was used as a control. After culturing in a constant temperature incubator for 48 hours, 10 μL of CCK8 solution was added to each well, and the culture plate was incubated in the incubator for 1 - 4 hours. The absorbance at 450 nm was measured with a microplate reader, and the cell growth rate of various samples was calculated.
[0125] Cell growth rate (%) = [A (drug added) - A (0 drug added)] / [A (0 drug added) - A (blank)] × 100 A (drug added): Absorbance of the well with cells, CCK solution, and drug solution A (blank): Absorbance of the well with medium and CCK8 solution but without cells A (0 drug added): Absorbance of the well with cells and CCK8 solution but without drug solution
[0126] Using the above activity test method, the G-CSF-binding L-DNA tetramer framework was evaluated for activity, and it was found that the L-DNA tetramer framework had no effect on the activity of G-CSF (Figure 17). Using the same activity test method, it was found that the divalent and trivalent G-CSFs assembled with the L-DNA tetramer also had no negative effect on the activity of G-CSF (Figure 18).
[0127] Example 8: Coupling and purification of SM(PEG)2-PMO conjugate In this example, experiments are performed using phosphorodiamidate morpholino nucleic acids. Specifically, the following four PMO single-stranded sequences (from 5' to 3') are selected.
[0128] Strand 1 (PMO1): SEQ ID NO: 275 5'-AGCAGCCTCGTTGAATCGCCAAGACACC-3' Strand 2 (PMO2): SEQ ID NO: 276 5'-AGGTGTCTTGGCGAAAGTTGCTCCGACG-3' Strand 3 (PMO3): SEQ ID NO: 277 5'-ACGTCGGAGCAACTAAGCGGTTCTGTGG-3' Strand 4 (PMO4): SEQ ID NO: 278 5'-ACCACAGAACCGCTATCAACGAGGCTGC-3'
[0129] There is an NH2 group modification at the 5'-end for coupling to the NHS active group of SM(PEG)2. Dissolve the PMO single strand containing a 5'-terminal NH2 modification in a phosphate buffer (50 mM NaH2PO4, 150 mM NaCl, pH 7.4) to prepare a stock solution with a final concentration of 1 mM. Dissolve the SM(PEG)2 (linker molecule) powder in dimethyl sulfoxide (DMSO) to freshly prepare a 250 mM SM(PEG)2 stock solution. Add a 10 - to 50-fold molar amount of the SM(PEG)2 stock solution to the PMO single-strand stock solution, mix rapidly, and then react at room temperature for 30 minutes to 2 hours. After the reaction is completed, add 10% volume of 1 M Tris-HCl (pH 7.0) to the reaction solution, mix, and then incubate at room temperature for 20 minutes to stop the excess SM(PEG)2 and continue the reaction. After the incubation is completed, purify the SM(PEG)2-PMO using Hitrap Capto MMC. The unreacted SM(PEG)2 flows through the column without binding, and the SM(PEG)2-PMO bound to the column is eluted with an eluent (25 mM BICINE, 200 mM NH4Cl, 1 M Arginine monohydrochloride, pH 8.5). The elution results are as shown in Figure 19a.
[0130] Analyze the PMO samples before and after coupling by positive-ion mode liquid chromatography mass spectrometry. As shown in Figures 20a and 20b, the results indicate that the finally obtained molecular weight of SM(PEG)2-PMO is consistent with the theoretical value, showing a high coupling reaction efficiency.
[0131] Example 9: Preparation of Nanobody Mutants Introduce a cysteine mutation for nucleic acid coupling at the carboxy-terminal of the nanobody. Optimize the gene sequence of the anti-HSA nanobody to yeast-preferred codons and then subclone it into the pPICZ alpha A plasmid. The amino acid sequence of the anti-HSA nanobody is SEQ ID NO:279. Add a His tag to the N-terminal of the nanobody to facilitate purification.
[0132] SEQ ID NO:279, Amino Acid Sequence of Anti-HSA Nanobody Mutant: HHHHHHAVQLVESGGGLVQPGNSLRLSCAASGFTFRSFGMSWVRQAPGKEPEWVSSISGSGSDTLYADSVKGRFTISRDNAKTTLYLQMNSLKPEDTAVYYCTIGGSLSRSSQGTQVTVSSGSC
[0133] After linearizing the plasmid, it was electroporated into Pichia pastoris X33 and screened on a Zeocin concentration gradient YPD perspective plate to obtain strains with a high copy number of the target gene. Using GMGY medium, monoclonal strains were cultured at 30 °C and 250 rpm. After obtaining a sufficient number of bacteria, GMMY medium was used to induce the secretion and expression of the target nanobody at 20 °C and 250 rpm, and 1% methanol was added every 24 hours.
[0134] The yield of nanobody expression in a laboratory-grade glass flask of the strain screened by high copy can reach 40 - 80 mg / L. According to SDS-PAGE identification and analysis, the culture supernatant after 72 hours of induction contains a large amount of target nanobody monomers and nanobody dimers. The His-tag affinity column is used to purify the nanobody in the culture supernatant.
[0135] Example 10: Coupling and Purification of Nanobody-PMO Conjugate The nanobody sample eluted by His-tag affinity chromatography (Example 9) was dialyzed with a dialysis buffer (20 mM Tris, 15 mM NaCl, pH 7.4) containing a reducing agent. During dialysis, the C-terminal sulfhydryl group was reduced, and at the same time, small impurity molecules such as free -SH groups were removed. The reduced nanobody and SM(PEG)2-PMO single strand (prepared in Example 8) were mixed according to a molar ratio of 1:1 - 2, and after uniform mixing, the reaction was carried out at room temperature for 2 hours.
[0136] As shown in Fig. 19b by SDS-PAGE identification, the results indicate that the coupling efficiency may reach over 90%. Unreacted SM(PEG)2-PMO single strands are removed with a His-tag affinity column, and the nanobody and nanobody-PMO mixture are collected.
[0137] Superdex TM 75 Increase 10 / 300GL(Superdex TM Nanobody and nanobody-PMO were separated using 75 Increase 10 / 300 GL (Superdex 75 Increase 10 / 300 GL), and as shown in Fig. 21, the nanobody and nanobody-PMO were effectively separated. Nanobodies before and after nucleic acid coupling were analyzed with a cation mode liquid chromatography ultra analyzer, and as shown in Figs. 20c and d, the results indicate that the finally obtained molecular weight of the nanobody-PMO is consistent with the theoretical value.
[0138] Example 11: Self-assembly of NAPPA-PMO drugs Hereinafter, taking pmo-NAPPA4-HSA(1) as an example, the self-assembly process of NAPPA-PMO drugs will be introduced. The concentrations of anti-HSA Nb-PMO1, PMO2, PMO3, and PMO4 are measured respectively. After quantitatively taking the above components and preheating at 37 °C for 5 minutes, they are mixed at a molar ratio of 1:1 under 37 °C conditions and incubated for 1 minute. That is, the assembly of pmo-NAPPA4-HSA(1) is completed.
[0139] Similar to the assembly of pmo-NAPPA4-HSA(1,2,3), the necessary assembly modules are anti-HSA Nb-PMO1, anti-HSA Nb-PMO2, anti-HSA Nb-PMO3, and PMO4. Under low temperature conditions, SDS-PAGE was used to identify the assembly status of the samples, and as shown in Fig. 22, the results indicate that the bands of the assembled samples are uniform.
[0140] Example 12: Verification of the binding activity of the nanobody-PMO monomer and the NAPPA-PMO drug after assembly Hereinafter, taking pmo-NAPPA4-HSA(1) as an example, the ELISA method is used to detect the binding ability of the anti-HSA nanobody, the anti-HSA Nb-PMO1 monomer, and the assembled pmo-NAPPA4-HSA(1) with the HSA protein (ACROBiosystems, HSA-H5220).
[0141] Each well of a 96-well ELISA plate is coated with 100 ng of HSA protein at 4 °C overnight. After washing with a washing solution (PBS containing 0.05% Tween-20) and blocking with a blocking solution (PBS containing 3% BSA and 0.05% Tween-20), serially diluted anti-HSA nanobody, PMO-coupled nanobody anti-HSA Nb-PMO1, and assembled pmo-NAPPA4-HSA(1) are added and incubated at room temperature for 1 hour. After washing three times, a 1:5000 diluted horseradish peroxidase-coupled rabbit anti-camelid VHH antibody (GenScript, A02016) is added and incubated at room temperature for 1 hour. After washing three times, a tetramethylbenzidine substrate solution (Beyotime, P0209) is added to develop color, and the color development is stopped using a stop solution (Beyotime, P0215). The absorbance at 450 nm of each well is read using a microplate reader (Molecular Devices, SpectraMax i3x), and the corresponding EC50 is calculated.
[0142] As shown in Figure 23, the calculation results show that the EC50 values of the binding of the anti-HSA nanobody, the anti-HSA Nb-PMO1, and the assembled pmo-NAPPA4-HSA(1) with the HSA protein are 0.577 nM, 0.391 nM, and 0.529 nM, respectively, indicating that the PMO-coupled nanobody method does not affect the binding activity between the nanobody and the corresponding antigen, and the PMO assembly method also does not affect the binding activity between the nanobody and the corresponding antigen.
[0143] Example 13: Nuclease resistance experiment of NAPPA-PMO drugs To verify whether PMO as a nucleic acid derivative is resistant to the degradation by various nucleases and whether the assembled NAPPA-PMO drugs are also resistant to nucleases or depolymerization, the following experiment was designed. In the experiment, three common nucleases, DNase I (Thermo Scientific, EN0523), T7 endonuclease I (NEB, M0302S), and S1 Nuclease (Thermo Scientific, EN0321), were selected, and the PMO assembly sample pmo-NAPPA4-HSA(1) and the D-DNA assembly sample DDNA-NAPPA4 (control) were incubated at 37 °C for 1 hour. The incubated pmo-NAPPA4-HSA(1) and DDNA-NAPPA4 were analyzed by SDS-PAGE and 2% agarose electrophoresis, respectively.
[0144] As shown in Figure 24, pmo-NAPPA4-HSA(1) is not degraded by the three nucleases (left figure), while DDNA-NAPPA4 is completely degraded by DNase I and S1 Nuclease and cleaved into short fragments by T7 endonuclease I. Therefore, the experiment shows that the NAPPA-PMO drugs are resistant to the degradation by common nucleases.
[0145] All documents mentioned in the present invention are cited as references in this application as if each document was individually cited as a reference. Further, after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms are also included in the scope defined by the appended claims of this application.
Claims
**Claim 1** A multimer complex based on complementary nucleic acid backbones, wherein the complex is a multimer formed by complexing n monomers having complementary nucleic acid backbones, where each monomer is a polypeptide having a single-stranded nucleic acid, and n is a positive integer from 3 to 6. In the multimer, the single-stranded nucleic acid of each monomer and the single-stranded nucleic acids of two other monomers form complementary double-stranded structures through base complementarity to form a complementary nucleic acid backbone structure. Here, the single-stranded nucleic acid sequences in the multimer consist of a set of single-stranded nucleic acid sequences, and the set is selected from any one of the following Group 1 to Group 4 for forming a trimer complex, tetramer complex, or pentamer complex based on complementary nucleic acid backbones: (Group 1) Single-stranded nucleic acid sequences for forming a trimer complex based on complementary nucleic acid backbones: 【Table 1-1】 【Table 1-2】 (Group 2) Single-stranded nucleic acid sequences for forming a tetramer complex based on complementary nucleic acid backbones: 【Table 2-1】 【Table 2-2】 【Table 2-3】 (Group 3) Single-stranded nucleic acid sequences for forming a pentamer complex based on complementary nucleic acid backbones: 【Table 3-1】 【Table 3-2】 【Table 3-3】 【Table 3-4】 (Group 4) Single-stranded nucleic acid sequences for forming a tetramer complex based on complementary nucleic acid backbones: 【Table 4】 The multimer complex, characterized in that it is selected from any one of the sequence sets shown. **Claim 2** The monomer has the structure of Formula I, Z1-W (I) In the formula, Z1 is a polypeptide moiety, W is a single-stranded nucleic acid sequence, and "-" is a linker or a bond, The multimer complex according to Claim 1. **Claim 3** The nucleic acid sequence W has the structure shown in Formula 1, X1-R1-X2-R2-X3 (1) Here, R1 is a base complementary pairing region 1, R2 is a base complementary pairing region 2, X1, X2, and X3 are each independently either nothing or redundant nucleic acids, and "-" is a bond, The multimer complex according to Claim 2. **Claim 4** The sequence of X2 is selected from the group consisting of A, AA, AGA, or AAA, The multimer complex according to Claim 2. **Claim 5** The set of single-stranded nucleic acid sequences is selected from any one of sequence sets 3-1 to 3-20 in Group 1 for forming a trimer complex based on complementary nucleic acid backbones, The multimer complex according to any one of Claims 1 to 4. **Claim 6** The single-stranded nucleic acid sequence set is selected from any one of sequence sets 4-1 to 4-20 in group 2 for forming a complementary nucleic acid backbone-based tetramer complex, characterized in that The multimer complex according to any one of claims 1 to 4.
7. The single-stranded nucleic acid sequence set is selected from any one of sequence sets 5-1 to 5-20 in group 3 for forming a complementary nucleic acid backbone-based pentamer complex, characterized in that The multimer complex according to any one of claims 1 to 4.
8. The single-stranded nucleic acid sequence set is a sequence set in group 4 consisting of single-stranded nucleic acid sequences SEQ ID No: 275 to 278 for forming a complementary nucleic acid backbone-based tetramer complex, characterized in that The multimer complex according to any one of claims 1 to 4.
9. A pharmaceutical composition, The pharmaceutical composition, (a) The complementary nucleic acid backbone-based multimer complex according to claim 1, and (b) A pharmaceutically acceptable carrier, characterized in that the pharmaceutical composition.
10. A nucleic acid sequence library, The nucleic acid sequence library contains a nucleic acid sequence for forming the complementary nucleic acid backbone-based multimer complex according to claim 2, characterized in that the nucleic acid sequence library.
11. The nucleic acid sequence W has a structure represented by formula 1, X1-R1-X2-R2-X3 (1) Here, R1 is base complementary pairing region 1, R2 is base complementary pairing region 2, X1, X2 and X3 are each independently none or redundant nucleic acids, "-" is a bond, characterized in that The nucleic acid sequence library according to claim 10.
12. Use of the nucleic acid sequence library according to claim 10, It is used for the preparation of the multimer complex according to claim 1 or a pharmaceutical composition containing the multimer complex according to claim 1, characterized in that the use.
13. A method for determining a single-stranded nucleic acid sequence for forming a complementary nucleic acid backbone-based multimer complex, comprising the following steps: (a) Setting the parameters of simulated annealing: Setting the initial annealing temperature, the final annealing temperature and the annealing temperature decay coefficient ΔT, Setting the optimization constraint parameters, (1) The number n of single-stranded nucleic acids, where n is a positive integer from 3 to 6, (2) The length L of the inverted repeat sequence, where L is 12 to 16 bases, (3) Dissociation temperature threshold T of the pairing region m , (4) Free energy threshold value ΔG of the specific pairing region array S ° (5) Free energy threshold value ΔG for non-specific pairing NS ° (6) Connector X2, where connector X2 is selected from the group consisting of A, AA, and AAA, (7) Dissociation temperature threshold T of the secondary structure (hairpin) m-H , (8) CG ratio P within the paired array CG and P CG ranges from [0.4, 0.6), Initialize the sequence set according to the above parameters 【Number 1】 and (b) the objective function value E of the set S described in the previous step 0 Calculate the free energy of non-specific pairing between sequences and with themselves (ΔG NS ) and simultaneously calculate the free energy matrix C of non-specific pairings n×n and obtain S i and S j (1≦i≦n, 1≦j≦n) and S i and S j Non-specific pairing free energy ΔG NS ° (S i, S j ) according to S i Or S j randomly select a set of sequences S to perform an update operation to obtain new nucleic acid sequences and obtain an updated set of sequences S'; (c) Determine whether the sequences in the set S' described in the previous stage satisfy the optimization constraint parameter conditions set in stage (a), and the dissociation temperature T of the specific pairing region m , the free energy ΔG of the specific pairing region sequence S °, the dissociation temperature T of the secondary structure m-H and the CG ratio P CG Verify the parameters including. If the above parameters satisfy the constraint conditions, execute step (d). Otherwise, repeat step (c). If S' that satisfies the conditions cannot be obtained even after step (b) is executed 15 times continuously under a specific annealing temperature, in order to prevent infinite loops, set S as set S' and execute the next step. (d) The objective function value E of the set S' described in the previous stage 1 is calculated, and E 0 and E 1 are compared, and E 1 ≥ E 0 indicates that the free energy of non-specific pairing is optimized, and the array set S' becomes the array set S. When E 1 < E 0 indicates that the free energy of non-specific pairing is not optimized. In this case, it is determined whether to set S' to S according to the Metropolis criterion, and (e) Decay the annealing temperature according to the decay coefficient ΔT set in step (a), and repeat steps (b), (c), and (d) for S described in the previous step, that is, based on Monte Carlo simulated annealing, until the annealing temperature reaches the annealing end temperature, the 【Number 2】 described in the previous step becomes a single-stranded nucleic acid sequence for forming a complementary nucleic acid backbone-based multimer complex The method, characterized by comprising the same.
14. (9) When n = 4, use a symmetric sequence and further include the step of initializing the sequence set according to the above parameters [Number 3] The method according to claim 13, characterized by the above.
15. A set of nucleic acids consisting of a set of single-stranded nucleic acid sequences for forming a complementary nucleic acid backbone-based multimer complex, where the set of single-stranded nucleic acid sequences is one of the following Group 1 to Group 4 for forming a complementary nucleic acid backbone-based trimer complex, tetramer complex, or pentamer complex: (Group 1) A single-stranded nucleic acid sequence for forming a complementary nucleic acid backbone-based trimer complex: 【Table 5-1】 【Table 5-2】 (Group 2) A single-stranded nucleic acid sequence for forming a complementary nucleic acid backbone-based tetramer complex: 【Table 6-1】 【Table 6-2】 【Table 6-3】 (Group 3) A single-stranded nucleic acid sequence for forming a complementary nucleic acid backbone-based pentamer complex: 【Table 7-1】 【Table 7-2】 【Table 7-3】 【Table 7-4】 (Group 4) A single-stranded nucleic acid sequence for forming a complementary nucleic acid backbone-based tetramer complex: 【Table 8】 The set of nucleic acids, characterized by being selected from any one of the sequence sets shown above.
16. The single-stranded nucleic acid sequence is a sequence consisting of a nucleic acid selected from the group consisting of left-handed nucleic acid, peptide nucleic acid, locked nucleic acid, sulfur-modified nucleic acid, 2'-fluoro-modified nucleic acid, 5-hydroxymethylcytosine nucleic acid, phosphorodiamidate morpholino nucleic acid, and combinations thereof The multimer complex according to claim 1.
17. The single-stranded nucleic acid sequence is a sequence consisting of phosphorodiamidate morpholino nucleic acid The multimer complex according to claim 1.
Citation Information
Patent Citations
Multispecific protein drugs and libraries thereof, and methods of production and use
JP2020519696A