Protein sequencing library construction method and protein sequencing method
Through the method of endonuclease enzyme digestion and chemical coupling, a protein sequencing library suitable for nanopore sequencing was constructed, which solved the universality and stability of protein sequencing in unknown sequences and achieved efficient protein sequence sequencing.
Patent Information
- Application Number
- PCT/CN2023/143721
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-30
- Publication Date
- 2025-07-03
AI Technical Summary
The prior art is difficult to achieve efficient sequencing of proteins of unknown sequences, especially the lack of universality and stability based on nanopore sequencing technology, so it is impossible to build a polypeptide library suitable for proteins of different sequences and structures.
Through the endonuclease enzyme, polypeptide fragments containing free amino groups at the N-terminal and the same functional groups at the C-terminal end are obtained. These functional groups are coupled to nucleic acids to construct a protein sequencing library, including C-terminal restriction enzyme treatment of specific amino acids such as lysine, glutamate and tyrosine, and chemical bonds are formed by combining specific chemical reactions.
Sequence sequencing of the vast majority of proteins is achieved, the universality and stability of sequencing is improved, and a high-throughput protein sequencing library can be constructed, suitable for the detection of proteins in unknown sequences.
Smart Images

Figure CN2023143721_03072025_PF_FP_ABST
Abstract
Description
Method for constructing protein sequencing library and protein sequencing method Technical Field
[0001] The present invention relates to the field of protein sequencing, and in particular to a method for constructing a protein sequencing library and a protein sequencing method. Background Art
[0002] Unlike nucleic acid sequencing, protein sequencing currently faces a challenge due to the lack of effective sequence amplification methods, which has limited the sensitivity of protein detection. With the recent development of nanopores as single-molecule sensors, large arrays of nanopores can be formed on ultra-small devices, allowing for the simultaneous reading of large numbers of DNA strands. Their long read lengths have made them highly attractive for DNA and direct RNA sequencing. Nanopore-based protein sensing technology is still in its infancy and faces unique challenges unique to protein and proteomics sequencing. In particular, proteins span a wide range of sizes and possess stable three-dimensional folded structures. Unlike nucleic acids, the backbone of polypeptides is not naturally charged, complicating the possibility of single-molecule electric threading into nanopores. Similar to DNA nanopore sequencing, a bias is applied across the biological membrane, forcing a linearized amino acid sequence through the nanopore, generating an ionic current that allows the reading of unlabeled proteins or peptides. Currently, the engineered fragmented toxin C (FraC) nanopore has been used to demonstrate discrimination between peptides containing glutamic acid substituted for alanine, and the wild-type aerolysin protein has been shown to recognize single amino acid differences in short polyarginine peptides. Furthermore, Cardozo et al. established a library of approximately 20 proteins orthogonally barcoded with intrinsic peptide sequences and successfully read them using a nanopore sensor. Proteomics can precisely quantify different protein types at single-molecule resolution, enabling the identification of protein fingerprints. This facilitates biomarker analysis and plays a significant role in early disease detection and screening.
[0003] Due to the complex composition of proteins, they are not limited to the permutations and combinations of the 20 natural amino acids after translation. There are also hundreds of variants generated by post-translational modifications, cleavage, and splicing, such as D-isomers, enantiomers, and higher-order structures of post-translational modifications. We still need to describe the electrical signals from different sequences in large quantities in order to create a "coding table" corresponding to the electrical signals and protein sequences. Currently, nucleic acid-peptide complexes are commonly coupled through nanopores, which requires the modification of reactive groups on the side chains or both ends of the polypeptide. Due to the specificity of protein structure and sequence, there are currently no reports that can construct N-terminal to C-terminal or C-terminal to N-terminal double-end modification libraries for proteins of unknown sequence that are not affected by polypeptide side chain groups and specific sequences.
[0004] Among the above-mentioned existing technologies, protein detection based on nanopore sequencing technology mostly uses artificially synthesized or artificially modified known sequence polypeptides as test carriers, and fails to provide a biochemical full-process solution from proteins of unknown sequence to peptide libraries that can be used for sequencing. In addition, during the sequencing process, there are more or less requirements for the sequence of the polypeptide to be sequenced. The electrical properties of the polypeptides currently reported to be perforated by nanopores are generally electronegative and short peptides, which are conducive to the generation of characteristic signal peaks when entering the pore, and it is difficult to meet the universal demand for sequencing proteins with different sequences and structures. In addition, the stability and repeatability of existing sequencing methods are general, and they can only distinguish one type of test protein, which is difficult to meet the requirements of actual sequencing.
[0005] Summary of the Invention
[0006] The main purpose of the present invention is to provide a method for constructing a protein sequencing library and a protein sequencing method to solve the problem in the prior art that it is difficult to sequence natural proteins (i.e., proteins with unknown sequences).
[0007] To achieve the above objectives, according to a first aspect of the present invention, a method for constructing a protein sequencing library is provided, which comprises: enzymatically cleaving a target protein with an endonuclease to obtain two or more polypeptide fragments containing a free amino group at the N-terminus and the same functional group at the C-terminus, wherein the functional group comprises an amino group, an ε-carboxyl group, or a phenolic hydroxyl group; utilizing the functional group at the C-terminus and the free amino group at the N-terminus, the polypeptide fragments are coupled with nucleic acids to obtain a protein sequencing library; wherein the nucleic acids are coupled with the functional group at the C-terminus and the free amino group at the N-terminus of the polypeptide fragments, respectively.
[0008] Furthermore, the connection between the nucleic acid and the free amino group at the N-terminus of the polypeptide fragment includes: reacting a first N-terminal reactive group with the free amino group to form a first N-terminal chemical bond, and achieving coupling of the N-terminus of the polypeptide fragment with the nucleic acid through the first N-terminal chemical bond; or converting the free amino group into a first modifying group, reacting the first modifying group with the second N-terminal reactive group to form a second N-terminal chemical bond, and achieving coupling of the N-terminus of the polypeptide fragment with the nucleic acid through the second N-terminal chemical bond; preferably, the second N-terminal chemical bond includes: a triazole structure, an eight-membered nitrogen-containing heterocyclic structure, a β-hydroxyaldehyde structure, a β-hydroxyketone structure, an amino acid structure, a polyheterocyclic structure, an aromatic structure or an olefin structure; preferably, the polyheterocyclic structure includes a nitrogen-containing polyheterocyclic structure.
[0009] Furthermore, the amino acid at the C-terminus of the polypeptide fragment is lysine, the functional group is amino group, and the endonuclease is lysine C endonuclease.
[0010] Furthermore, the connection between the nucleic acid and the amino group at the C-terminus of the polypeptide fragment includes: reacting a first C-terminal reactive group with the amino group to form a first C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the first C-terminal chemical bond; or converting the amino group into a modifying group, reacting the modifying group with a second C-terminal reactive group to form a second C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the second C-terminal chemical bond; preferably, the second C-terminal chemical bond includes: a triazole structure, an eight-membered nitrogen-containing heterocyclic structure, a β-hydroxyaldehyde structure, a β-hydroxyketone structure, an amino acid structure, a polyheterocyclic structure, an aromatic structure or an olefin structure.
[0011] Furthermore, the amino acid at the C-terminus of the polypeptide fragment is glutamic acid, the functional group is ε-carboxyl, and the endonuclease is a glutamic acid C-terminal restriction endonuclease.
[0012] Furthermore, the connection between the nucleic acid and the ε-carboxyl group at the C-terminus of the polypeptide fragment includes: reacting the third C-terminal reactive group with the ε-carboxyl group to form a third C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the third C-terminal chemical bond; or converting the ε-carboxyl group into a modifying group, reacting the modifying group with the fourth C-terminal reactive group to form a fourth C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the fourth C-terminal chemical bond; preferably, the fourth C-terminal chemical bond includes a ketone structure.
[0013] Furthermore, the amino acid at the C-terminus of the polypeptide fragment is one or more of tyrosine, phenylalanine or tryptophan, the functional group is a phenolic hydroxyl group, and the endonuclease is chymotrypsin.
[0014] Furthermore, the connection between the nucleic acid and the phenolic hydroxyl group at the C-terminus of the polypeptide fragment includes: reacting the fifth C-terminal reactive group with the phenolic hydroxyl group to form a fifth C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the fifth C-terminal chemical bond; or converting the phenolic hydroxyl group into a modifying group, reacting the modifying group with the sixth C-terminal reactive group to form a sixth C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the sixth C-terminal chemical bond; preferably, the sixth C-terminal chemical bond includes a sulfonic acid bond or a triazole structure.
[0015] Furthermore, the nucleic acid includes a specific molecular tag.
[0016] Furthermore, after coupling, uncoupled polypeptide fragments are removed to obtain a protein sequencing library.
[0017] Furthermore, the target protein is multiple and a corresponding protein sequencing library is constructed for each target protein using a construction method using different nucleic acids. The protein sequencing libraries of the multiple target proteins are then mixed to obtain a high-throughput protein sequencing library.
[0018] To achieve the above-mentioned object, according to a second aspect of the present invention, a protein sequencing method is provided, which comprises: constructing a protein library using the above-mentioned protein sequencing library construction method to obtain a protein sequencing library; and sequencing the protein sequencing library using sequencing technology to obtain protein sequence information.
[0019] Furthermore, the technology includes nanopore protein sequencing technology or single-molecule fluorescence recognition and detection technology.
[0020] Furthermore, in the sequencing method, nanopore protein sequencing technology is used to pull the nucleic acid through the nanopore protein, so that the protein is directed to pass through the nanopore protein and complete the sequencing.
[0021] Applying the technical solution of the present invention and utilizing the aforementioned protein sequencing library construction method, the target protein is enzymatically digested with an endonuclease to obtain two or more polypeptide fragments containing free amino groups at the N-terminus and the same functional group at the C-terminus. By selectively modifying the N-terminus and C-terminus or simultaneously modifying both ends, the polypeptide fragments are coupled to nucleic acids to obtain a protein sequencing library suitable for subsequent protein sequencing. The protein sequencing library obtained using this construction method is capable of sequencing the sequences of the vast majority of proteins, not just specific proteins. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0023] FIG1 shows a schematic diagram of a protein sequencing process according to an embodiment of the present invention.
[0024] FIG2 shows a graph showing the HPLC results according to Example 1 of the present invention.
[0025] FIG3 shows a mass spectrum according to Example 1 of the present invention.
[0026] FIG4 shows the electrophoresis result according to Example 1 of the present invention.
[0027] FIG5 shows a graph of nanopore sequencing current signals according to Example 1 of the present invention.
[0028] FIG6 shows a graph showing the HPLC results according to Example 2 of the present invention.
[0029] FIG. 7 shows the electrophoresis results according to Example 2 of the present invention.
[0030] FIG8 shows a graph of nanopore sequencing current signals according to Example 2 of the present invention. DETAILED DESCRIPTION
[0031] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the embodiments.
[0032] Currently, mainstream protein sequencing methods still rely on qualitative and quantitative mass spectrometry analysis, with a few utilizing site-specific amino acid modification or identification to distinguish different protein sequence fingerprints. To improve the sensitivity of trace protein detection, nanopore detection has been used to validate the current signals of amino acids with specific site mutations, or to detect the signals of short charged polypeptides, enabling identification. However, conventional detection of proteins of unknown sequence has not yet been achieved.
[0033] Among the existing methods, methods for sequencing specific proteins include: 1) pre-in vitro chemical synthesis of a polypeptide sequence with modified groups on both ends, such as a serine amino acid at the N-terminus and an azide or alkyne side chain at the C-terminus, thereby coupling different nucleic acid sequences using orthogonal thiol-maleimide reactions and click reactions, and obtaining a pure OPO conjugate (nucleic acid-polypeptide-nucleic acid (Oligo-Peptide-Oligo) conjugate complex) after purification through the reaction, and finally performing signal detection in a confined nanopore; 2) designing a method called "Nanopore as a traceable protein tag for engineering reporter genes" through in vitro protein expression technology. This method designs four sequences expressed in bacteria or human hosts, including OsmY for recognizing the nanopore, Smt3 for restricting the passage of the pore, 20 different protein sequences, and a tail polyGSD for nanopore capture, thereby distinguishing different amino acid sites and different polypeptide sequences.
[0034] In the above-mentioned existing technologies, protein detection based on nanopore sequencing technology mostly uses artificially synthesized or artificially modified polypeptides of known sequences as test carriers, but fails to provide a biochemical full-process solution from proteins of unknown sequence to peptide libraries that can be used for sequencing. In addition, during the sequencing process, there are more or less requirements for the sequence of the polypeptide to be sequenced. The electrical properties of the polypeptides perforated by nanopores currently reported are generally electronegative and short peptides, which are conducive to the generation of characteristic signal peaks when entering the pore, and it is difficult to meet the universal demand for sequencing proteins with different sequences and structures. In addition, the stability and repeatability of existing sequencing methods are general, and they can only distinguish one type of test protein, which is difficult to meet the requirements of actual sequencing.
[0035] Although mass spectrometry is still the mainstream method for protein sequencing, there have been many early explorations of protein sequencing based on nanopore detection technology. In principle, this method relies on the change in the current signal generated when the polypeptide molecule passes through the nanopore to detect the type and sequence of the polypeptide. Technically, this method requires a helicase to provide power, guiding the polypeptide molecule through the pore at a steady speed under the traction of nucleic acids. This requires that the protein to be tested be converted into a collection of peptide segments coupled to nucleic acids for systematic protein sequencing. Currently, related detection methods are all based on polypeptides or specific proteins with known sequences, and their universality is poor.
[0036] Based on the above problems, the present application creatively proposes a full-process solution for converting proteins of ordinary unknown sequences into polypeptide libraries that can be used for nanopore sequencing, that is, the target protein is degraded into a linear polypeptide of primary structure by specific endonucleases such as lysine C-terminal restriction endonuclease, glutamate C-terminal restriction endonuclease, etc., to form a polypeptide library with consistent functional groups such as amino, ε-carboxyl, and phenolic hydroxyl groups on the terminal side chains of the carbon chain, and different chemical reaction sites are used to modify the functional groups on the N-terminal amino group and the C-terminal side chain, thereby coupling with modified nucleic acids to construct a sequencing library. By recognizing nanopore sequencing signals, different protein sequence information is determined. Different from the selectivity of existing library construction for protein sequences, the present application attempts to develop a method for constructing a library that can sequence most proteins, and proposes a series of protection schemes for the present application.
[0037] In a first typical embodiment of the present application, a method for constructing a protein sequencing library is provided, which comprises: enzymatically cleaving a target protein with an endonuclease to obtain two or more polypeptide fragments containing a free amino group at the N-terminus and the same functional group at the C-terminus, wherein the functional groups include an amino group, an ε-carboxyl group, or a phenolic hydroxyl group; utilizing the functional group at the C-terminus and the free amino group at the N-terminus, coupling the polypeptide fragments with nucleic acids to obtain a protein sequencing library; wherein the nucleic acids are coupled to the functional group at the C-terminus and the free amino group at the N-terminus of the polypeptide fragments, respectively.
[0038] The target protein contains a free amino group at the N-terminus and a free carboxyl group at the C-terminus. In the above construction method, the target protein is enzymatically cleaved by an endonuclease at specific amino acid sites to cut the target protein into multiple polypeptide fragments. This ensures that the amino acid at the C-terminus of the polypeptide fragment contains a free carboxyl group. Due to the specific structure of the amino acid, the C-terminus of the polypeptide fragment also has an amino, ε-carboxyl, or phenolic hydroxyl functional group. Furthermore, the free amino group at the N-terminus and the functional group at the C-terminus of the polypeptide fragment are used to couple the polypeptide fragment with a nucleic acid to form an OPO (nucleic acid-polypeptide-nucleic acid, Oligo-Peptide-Oligo) structure, thereby obtaining a protein sequencing library.
[0039] The "free amino group" in this application refers to the naturally occurring -NH2 group at the N-terminus of a polypeptide (i.e., the "H2N" in the protein formula H2NCHRCOOH). The "amino group" in this application refers to the functional group generated at the new C-terminus obtained by endonuclease cleavage. The actual source of this "amino group" is the -NH2 group carried by the "R" in the amino acid at the C-terminus of the polypeptide (the "R" in the above protein formula).
[0040] In a preferred embodiment, the connection between the nucleic acid and the free amino group at the N-terminus of the polypeptide fragment includes: reacting a first N-terminal reactive group with the free amino group to form a first N-terminal chemical bond, and achieving coupling of the N-terminus of the polypeptide fragment and the nucleic acid through the first N-terminal chemical bond; or converting the free amino group into a first modifying group, reacting the first modifying group with a second N-terminal reactive group to form a second N-terminal chemical bond, and achieving coupling of the N-terminus of the polypeptide fragment and the nucleic acid through the second N-terminal chemical bond; preferably, the second N-terminal chemical bond includes: a triazole structure, an eight-membered nitrogen-containing heterocyclic structure, a β-hydroxyaldehyde structure, a β-hydroxyketone structure, an amino acid structure, a polyheterocyclic structure, an aromatic structure or an olefin structure; preferably, the polyheterocyclic structure includes a nitrogen-containing polyheterocyclic structure.
[0041] For the free amino group at the N-terminus of the polypeptide fragment, the effect of linking the nucleic acid to the N-terminus of the polypeptide can be achieved by directly reacting the free amino group with the first N-terminal reactive group, or by first modifying the free amino group and then reacting it with the second N-terminal reactive group.
[0042] In a preferred embodiment, the amino acid at the C-terminus of the polypeptide fragment is lysine, the functional group is amino group, and the endonuclease is lysine C endonuclease.
[0043] By selecting a specific lysine C endonuclease to endonuclease lysine in a polypeptide, peptide fragments containing free amino groups at both the N-terminus and the C-terminus can be obtained. Since the lysine (K) content in target proteins of varying sources and sequences is moderate, enzymatic cleavage with a lysine-restricted C-terminal endonuclease can separate the protein into primary sequences of varying lengths, preventing the formation of secondary structures in peptides when the cleavage chain is too long. This facilitates the subsequent simultaneous modification of the N-terminus and C-terminus, or the selective modification of the C-terminal side chain amino group.
[0044] In a preferred embodiment, the construction method includes: reacting the reactive group with -NH2 (derived from an amino group and / or a free amino group) to form a first C-terminal chemical bond, and coupling the polypeptide fragment to the nucleic acid through the first C-terminal chemical bond; or converting -NH2 into a modifying group, reacting the modifying group with the reactive group to form a second C-terminal chemical bond, and coupling the polypeptide fragment to the nucleic acid through the second C-terminal chemical bond;
[0045] Preferably, the connection between the nucleic acid and the amino group at the C-terminus of the polypeptide fragment includes: reacting a first C-terminal reactive group with the amino group to form a first C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the first C-terminal chemical bond; or converting the amino group into a modifying group, reacting the modifying group with a second C-terminal reactive group to form a second C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the second C-terminal chemical bond; preferably, the second C-terminal chemical bond includes: a triazole structure, an eight-membered nitrogen-containing heterocyclic structure, a β-hydroxyaldehyde structure, a β-hydroxyketone structure, an amino acid structure, a polyheterocyclic structure, an aromatic structure or an olefin structure; preferably, the polyheterocyclic structure includes a nitrogen-containing polyheterocyclic structure.
[0046] For the amino functional group at the C-terminus of the polypeptide fragment, the effect of connecting the nucleic acid to the C-terminus of the polypeptide can be achieved by directly reacting the amino group with the first C-terminal reactive group, or first modifying the amino group and then reacting it with the second C-terminal reactive group.
[0047] Among them, the above-mentioned second C-terminal chemical bond includes but is not limited to: a triazole structure formed by a click reaction between an azide group and an alkynyl group, an eight-membered nitrogen-containing heterocyclic compound formed by an oxygen-bridged cyclization reaction between a nitrone and a cycloalkyne, a nucleophilic addition of an aldehyde or ketone to form a β-hydroxyaldehyde or β-hydroxyketone, a cycloaddition reaction between a tetrazine and a cyclic olefin or a cyclic alkyne to form a nitrogen-containing polycyclic heterocyclic compound, a click reaction between an amino group and a carbonyl / ketone of an isocyanide to form an amino acid analogue; a cycloaddition reaction between a tetracycloalkane and a dienophile group to form a polyheterocyclic compound; a substitution reaction between an aromatic or alkenyl boronic acid or boronate ester group and a halogen group such as chlorine, bromine, or iodine to form an aromatic or olefinic compound.
[0048] The first C-terminal chemical bond used in this application and the reaction of the reactive group and the modifying group to generate the second C-terminal chemical bond are both prior arts. Those skilled in the art can flexibly select similar reactions disclosed in the prior art to select the modifying group and the reactive group and obtain the corresponding first or second C-terminal chemical bond.
[0049] Preferably, the reactive group includes a group capable of binding to an amine via acylation or alkylation, more preferably an isothiocyanate, isocyanate, acyl azide, NHS ester, sulfonyl chloride, glyoxal, epoxide, oxirane, carbonate, aryl halide, imidoester, carbodiimide, anhydride, or fluorophenyl ester, as well as an aldehyde group capable of binding via reductive amination. The modifying groups correspond one to one with the reactive groups, and the above reaction is performed to produce a second C-terminal chemical bond.
[0050] In the above construction method, the nucleic acid and the polypeptide fragment are coupled. The reactive group on the nucleic acid reacts with the amino group, or a modified group obtained by further conversion of the amino group, to form a first or second C-terminal chemical bond. This chemical bond is used to couple the polypeptide fragment to the nucleic acid. This reaction can be achieved whether the reactive group is located at the 5' or 3' end of the nucleic acid.
[0051] In the above-mentioned coupling, the reactive group on the nucleic acid directly forms a first C-terminal chemical bond with the amino group, including but not limited to using DBCO to react with the amino group to form an azide group for coupling, or converting 1H-imidazole-1-sulfonyl azide hydrochloride into a bioorthogonal azide group, both of which do not affect the side chains of the protein sequence. Alternatively, the reactive group on the C-terminal side chain of the polypeptide can be first converted by enzymatically converting the -NH2 to a modifying group such as -CHO, which can then react with a reactive group such as -NH2 on the nucleic acid to form a second C-terminal chemical bond to achieve coupling.
[0052] In a preferred embodiment, the amino acid at the C-terminus of the polypeptide fragment is glutamic acid, the functional group is ε-carboxyl, and the endonuclease is a glutamic acid C-terminal restriction endonuclease.
[0053] Considering that glutamate plays an important physiological function in organisms and the terminal side chain effectively reduces the sequence steric effect, glutamate with a moderate protein content was selected as the C-chain terminal.
[0054] In a preferred embodiment, the connection between the nucleic acid and the ε-carboxyl group at the C-terminus of the polypeptide fragment includes: reacting the third C-terminal reactive group with the ε-carboxyl group to form a third C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the third C-terminal chemical bond; or converting the ε-carboxyl group into a modifying group, reacting the modifying group with the fourth C-terminal reactive group to form a fourth C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the fourth C-terminal chemical bond; preferably, the fourth C-terminal chemical bond includes a ketone structure.
[0055] The third C-terminal chemical bond corresponding to the above-mentioned ε-carboxyl functional group includes but is not limited to an amide structure formed by the reaction of primary amines with ε-carboxyl groups, or an α-amidoamide structure formed by the reaction of primary amines, primary aldehydes or ketones and isonitriles with ε-carboxyl groups, or an ester structure formed by the reaction of aromatic sulfonates with ε-carboxyl groups, or an amide-4-phenyl ketone structure formed by the reaction of p-phenylenediamine with ε-carboxyl groups, or an ester covalent metal structure formed by the reaction of transition metal palladium and platinum with ε-carboxyl groups.
[0056] Preferably, the fourth C-terminal chemical bond corresponding to the ε-carboxyl functional group comprises a ketone structure.
[0057] Preferably, the modification group includes but is not limited to acetonitrile. The ε-carboxyl group at the end of the carbon chain is reduced to acetonitrile by catalysis of vanadium chloride peroxidase in the coexistence of hydrogen peroxide and sodium bromide.
[0058] Preferably, the fourth C-terminal reactive group includes: primary amines of the combined condensation reagents of the amide reaction, combined primary amines, primary aldehydes and isonitriles of the Uggie reaction, aromatic sulfonates and paraphenylene cyclopropenes through nucleophilic reaction, covalent complexes formed by metal platinum palladium catalysis, etc.
[0059] Taking the above-mentioned modifying group acetonitrile as an example, acetonitrile as a modifying group forms a fourth C-terminal chemical bond with the fourth C-terminal reactive group, including but not limited to the cyano group and the halogen group undergoing a reduction reaction under metal catalysis to generate a ketone compound, and the cyano group and the alkyl sulfonate group undergoing a hydrolysis reaction under transition metal catalysis to generate a ketone compound.
[0060] In a preferred embodiment, the amino acid at the C-terminus of the polypeptide fragment is one or more of tyrosine, phenylalanine or tryptophan, the functional group is a phenolic hydroxyl group, and the endonuclease is chymotrypsin.
[0061] Since the content of phenylalanine and tryptophan is low and the content of tyrosine is high, and the selectivity and compatibility of tyrosine side chain modification are currently good, the enzymatically hydrolyzed peptides can be effectively labeled, which is beneficial to the classification and analysis of unknown protein sequences.
[0062] In a preferred embodiment, the connection between the nucleic acid and the phenolic hydroxyl group at the C-terminus of the polypeptide fragment includes: reacting the fifth C-terminal reactive group with the phenolic hydroxyl group to form a fifth C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the fifth C-terminal chemical bond; or converting the phenolic hydroxyl group into a modifying group, reacting the modifying group with the sixth C-terminal reactive group to form a sixth C-terminal chemical bond, and achieving coupling of the C-terminus of the polypeptide fragment and the nucleic acid through the sixth C-terminal chemical bond; preferably, the sixth C-terminal chemical bond includes a sulfonic acid bond or a triazole structure.
[0063] The fifth C-terminal chemical bond corresponding to the above-mentioned phenolic hydroxyl functional group includes but is not limited to the reaction of a primary amine or secondary amine or imine-containing group bound to formaldehyde with a phenolic hydroxyl group to form an ortho-aminomethylphenol structure, or the reaction of diazo with a phenolic hydroxyl group to form an oxime structure, or the reaction of thiazolidinedione with a phenolic hydroxyl group to form an ortho-phenolic hydroxythiazolidinedione structure, or the reaction of fluorinated sulfonamide with a phenolic hydroxyl group to form a phenyl sulfate structure, or the reaction of sulfonated triazole with a phenolic hydroxyl group to form a phenolsulfonated triazole structure, or α-D-fluoro-glycosides with a phenolic hydroxyl group to form a phenol-oxy-glycosyl structure, or the reaction of an electrophilic π-allylpalladium complex with a phenolic hydroxyl group to form an alkylphenyloxy ether structure, or the reaction of N,N-dimethylaniline with a phenolic hydroxyl group under the catalysis of cerium (IV) ammonium nitrate and tris(2,2'-bipyridine)ruthenium complex to form a dimethylaminediphenyl ether structure.
[0064] Preferably, the sixth C-terminal chemical bond corresponding to the phenolic hydroxyl functional group includes but is not limited to a sulfonic acid bond or a triazole structure.
[0065] The sixth C-terminal chemical bond includes but is not limited to a nucleophilic substitution reaction between an arylsulfonyl fluoride group and a hydroxyl group to form a sulfonic acid bond, or a click reaction between a phenolic azophenyl acetylene group and an azide group to form a triazole group.
[0066] Preferably, the modifying groups include primary or secondary amines or imine-containing compounds that bind to formaldehyde under acidic conditions through a Mannich-type reaction, diazo compounds under alkaline conditions, thiazolidinediones, fluorinated sulfonamides or sulfided triazole compounds, α-D-fluoro-glycosides, and aromatic nitrogen compounds based on the conversion of transition metals palladium, cesium, and rubidium.
[0067] Through the above-mentioned "enzyme cleavage-coupling" or "enzyme cleavage-modification-coupling" method, a complete polypeptide library can be constructed for most proteins, so as to carry out subsequent nanopore signal detection and other sequencing. Through the subsequent signal comparison and classification, different trace proteins can be identified.
[0068] The coupling of the above-mentioned nucleic acid and polypeptide fragment can occur at the N-terminus and C-terminus of the polypeptide fragment. By binding the nucleic acid fragment to the double-end of the polypeptide fragment, the following effects are achieved, including but not limited to: 1) uniform modification and coupling library construction of proteins from different sources in high-throughput protein sequencing; 2) determination of the start and end points of polypeptide pores by nucleic acid sequence, and identification of sequencing quality by nucleic acid sequence.
[0069] In a preferred embodiment, the nucleic acid comprises a unique molecular identifier (UMI).
[0070] In a preferred embodiment, after coupling, the uncoupled polypeptide fragments are removed to obtain a protein sequencing library.
[0071] After coupling, uncoupled polypeptide fragments are removed by molecular size differences. The purification process does not involve a solid-body purification step and does not require capture with an affinity probe. The construction method of this application can universally process unknown proteins, which is beneficial for the detection of unknown proteins.
[0072] In a preferred embodiment, there are multiple target proteins and a construction method is used to construct a corresponding protein sequencing library for each target protein using different nucleic acids. The protein sequencing libraries of the multiple target proteins are then mixed to obtain a high-throughput protein sequencing library.
[0073] In a preferred embodiment, each polypeptide fragment of the protein is coupled to a nucleic acid.
[0074] In the above-mentioned construction method, different target proteins can be respectively constructed into libraries to obtain protein sequencing libraries corresponding to the target proteins, and the sequence information in the protein sequencing library can reflect the sequence information of the target protein. In actual DNA sequencing, in order to improve the throughput of sequencing, DNA libraries with different sources are usually mixed and then high-throughput sequencing is performed together. In the above-mentioned construction method, similar operations can also be performed to obtain a high-throughput protein sequencing library mixed with multiple target protein sequence information. In the subsequent analysis of sequencing data, the sequence of the nucleic acid can be used to classify the sequencing information, thereby achieving the effect of high-throughput protein sequencing.
[0075] In a second exemplary embodiment of the present application, a protein sequencing method is provided, comprising: constructing a protein library using the above-described protein sequencing library construction method to obtain a protein sequencing library; and sequencing the protein sequencing library using a sequencing technology to obtain protein sequence information. The steps of the above-described construction method and sequencing method are shown in Figure 1.
[0076] In a preferred embodiment, the technology includes nanopore protein sequencing technology or single-molecule fluorescence recognition and detection technology; preferably, the single-molecule fluorescence recognition and detection technology includes total internal reflection fluorescence microscopy detection technology (TIRF), fluorescence resonance energy transfer detection technology (FRET), live localization microscopy detection technology (PALM) or stochastic optical reconstruction microscopy detection technology (STORM).
[0077] In a preferred embodiment, in the sequencing method, nanopore protein sequencing technology is used, by pulling the nucleic acid through the nanopore protein, so that the protein is directed to pass through the nanopore protein and the sequencing is completed.
[0078] In the above-mentioned protein sequencing library, the nucleic acid in OPO can pull the protein through the nanopore protein, improving the utilization rate of the sequencing pore and thus increasing the sequencing depth.
[0079] The beneficial effects of the present application will be further explained in detail below with reference to specific embodiments.
[0080] Example 1
[0081] This example uses the ATP-dependent Clp protease adaptor protein (ClpS) as an example to construct a sequencing library and detect its related sequences. The process includes:
[0082] 1. Protein hydrolysis into peptides
[0083] 100 μg of recombinant human CLPS protein (Abcam ab180330) was dissolved in 200 μL of distilled water to a concentration of 500 μM. 40 μL was added to a 1.5 mL centrifuge tube. 0.5 μL of 1 M dithiothreitol (Thermo R0861) and 10 μL of 8 M urea were added. The mixture was boiled in a 37°C metal bath for 10 minutes. 5 μL of 10 mM iodoacetamide (Aladdin I105563-5g) was added and the mixture was incubated at room temperature in the dark for 0.5 hour. Lys C endonuclease (NEB P8109S) was dissolved in 10 mM Tris-HCl, pH 8.0, to a concentration of 100 ng / μL. 1 μL was added to the centrifuge tube containing denatured CLPS. 0.5 μL of 1 M Tris-HCl, pH 8.0 was added and the mixture was incubated at 37°C for 16 hours.
[0084] 2. Purification of samples after enzymatic hydrolysis
[0085] The enzymatically digested sample was transferred to a 3K ultrafiltration tube (UFC500308), supplemented with 150 μL of distilled water, and centrifuged at 12,000 rpm for 15 minutes. The supernatant was collected and the centrifugation step was repeated three times. The supernatants were combined and lyophilized. The lyophilized sample was reconstituted in 300 μL of distilled water, desalted and purified using a peptide desalting column (Thermo 89851), and lyophilized.
[0086] 3. Sample site modification steps
[0087] Peptide double-end modification: Dissolve the peptide sample in 20 μL of anhydrous dimethyl sulfoxide (Thermo D12345), add 5 μL of 10 mM 1H-imidazole-1-sulfonyl azide hydrochloride (CAS No: 952234-36-5, Lot No. H305026-1g) and 2 μL of 100 mM triethylamine, and react at 25°C for 1 hour. Add 2 μL of 1 M Tris-HCl buffer (pH 8.0), a quencher, and react at room temperature for half an hour before lyophilization. The lyophilized sample was reconstituted in 300 μL of distilled water, desalted, purified using a peptide desalting column (Thermo 89851), and lyophilized.
[0088] Peptide Single-End Modification: Dissolve the peptide sample in 20 μL of anhydrous dimethyl sulfoxide (Thermo D12345), add 0.5 μL of 10 mM N-(2-((5-azidopentyl)amino)ethyl)-N-phenylethenesulfonamide (structure shown in Formula I), a C-terminal amino-selective modification reagent, and 2 μL of 100 mM triethylamine. Incubate at 25°C for 1 hour. Add 2 μL of 1 M Tris-HCl buffer (pH 8.0) as a quencher, incubate at room temperature for half an hour, and lyophilize. Redissolve the lyophilized sample in 300 μL of distilled water, desalt and purify using a peptide desalting column (Thermo 89851), and lyophilize.
[0089] IV. Sequencing Library Construction
[0090] Dissolve the desired nucleic acid sequences O1 and O2 in 1× PBS buffer (Thermo, 10010001) to a 200 μM concentration and remove 20 μL of each. Dissolve the lyophilized modified sample in 10 μL of dimethyl sulfoxide (Thermo D12345) and add it to the 1.5 mL centrifuge tube containing O1 and O2. Incubate overnight at room temperature. The reaction sample is divided into two tubes and desalted and purified using a nucleic acid purification column (NEB T1030L).
[0091] SEQ ID NO: 1: P-GCTTCTCGTG-DBCO;
[0092] SEQ ID NO: 2: DBCO-GCTGTCTTCTGTCGTCGTTTCCTTCTCTGC.
[0093] The sequence shown in SEQ ID NO: 1 is the nucleic acid sequence O1, wherein the 5' end of O1 has a phosphate modification and the 3' end has an alkynyl modification (DBCO, dibenzocyclooctyne);
[0094] The sequence shown in SEQ ID NO: 2 is the nucleic acid sequence O2, and the 5' end of O2 has an alkynyl modification (DBCO, dibenzocyclooctyne).
[0095] 5. Library Sequencing
[0096] Since nanopore sequencing is used in this example, the coupling product is added to the adapter and then nanopore sequencing is performed to obtain different current signals and perform data analysis.
[0097] 6. Experimental Results
[0098] 1. Results after peptide modification
[0099] In this example, a C18 column (059142) was used, with a mobile phase of water and acetonitrile. The gradient method was 5%-40% acetonitrile at 0.3 ml / min for 20 minutes. By comparing the difference in the elution time of the unmodified peptide and the modified peptide, the HPLC results are shown in Figure 2 and the mass spectrometry structure is shown in Figure 3. The peak at 16.1 min in the HPLC is the uncoupled modified peptide raw material, [M+H] + =717.47; the peak at 17.6 min in HPLC is the successfully modified peptide, [M+H] + =769.34.
[0100] 2. Library electrophoresis results
[0101] In this example, electrophoresis was performed using 12% polyacrylamide denaturing gel at 150 V for 90 minutes. The results are shown in FIG4 . In Figure 4, the first lane is a DNA marker, which in this example uses the ThermoFisher ultra-low range DNA molecular weight standard ladder (10597012); the second lane is the starting nucleic acid sequence O1; the third lane is the starting nucleic acid sequence O2; the fourth lane is a blank control in which O1 and O2 are added with reaction reagents, where O1 is a 10 bp oligo that runs off the bottom of the gel after ionization at 150V for 1 hour on a 12% Urea-PAGE, resulting in the absence of an O1 band in lanes 2 and 4; the fifth lane is a double complex band of O1-peptide-O1, the product of the reaction between O1 and a polypeptide; the sixth lane is a band resulting from the reaction between nucleic acid sequence O2 and a polypeptide, where the dotted-line box represents a single-sided coupling band of O2-peptide; and the seventh lane is a band resulting from the reaction between O1, O2, and a polypeptide, where the band in the solid-line box represents a double complex band of O1-peptide-O2, and the band in the dotted-line box represents a single-sided coupling band of O2-peptide.
[0102] 3. Sequencing results
[0103] The prepared library was subjected to nanopore sequencing, and different current signal diagrams as shown in Figure 5 were obtained through data analysis. The signals in the dotted boxes are signals of different polypeptide segments. Through data analysis, the corresponding polypeptide information can be obtained and the protein sequence information can be obtained by comparison.
[0104] Example 2
[0105] This example uses the ATP-dependent Clp protease adaptor protein (ClpS) as an example to construct a sequencing library and detect its related sequences. The process includes:
[0106] 1. Protein hydrolysis into peptides
[0107] 100 μg of recombinant human CLPS protein (Abcam ab180330) was dissolved in 200 μL of distilled water to a 500 μM concentration. 40 μL was added to a 1.5 mL centrifuge tube. 0.5 μL of 1M dithiothreitol (Thermo R0861) and 10 μL of 8M urea were added. The mixture was boiled in a 37°C metal bath for 10 minutes. 5 μL of 10 mM iodoacetamide (Aladdin I105563-5g) was added and the mixture was incubated at room temperature in the dark for 0.5 hour. Glu C endonuclease (NEB P8100S) was dissolved in 2× Glu C Reaction Buffer (NEB B8100S) and diluted to 1× with water to a concentration of 100 ng / μL of Glu C. 1 μL was then added to the centrifuge tube containing denatured ClpS and incubated at 37°C for 16 hours.
[0108] 2. Purification of samples after enzymatic hydrolysis
[0109] The enzymatically digested sample was transferred to a 3K ultrafiltration tube (UFC500308), supplemented with 150 μL of distilled water, and centrifuged at 12,000 rpm for 15 minutes. The supernatant was collected and the centrifugation step was repeated three times. The supernatants were combined and lyophilized. The lyophilized sample was reconstituted in 300 μL of distilled water, desalted and purified using a peptide desalting column (Thermo 89851), and lyophilized.
[0110] 3. Sample site modification steps
[0111] Dissolve the peptide sample in 20 μL of anhydrous dimethyl sulfoxide (Thermo D12345) and add 5 μL of 10 mM HHS-465 (MedKoo Cat#: 533540). Incubate at 25°C for 1 hour. Redissolve the lyophilized sample in 300 μL of distilled water, desalt and purify using a peptide desalting column (Thermo 89851), and lyophilize.
[0112] IV. Sequencing Library Construction
[0113] The desired nucleic acid sequences O3, O4, and O5 were dissolved in 1× PBS buffer (Thermo, 10010001) to a concentration of 200 μM. 20 μL of each was removed and incubated at 95°C for 10 minutes, followed by natural annealing and return to room temperature. The lyophilized modified sample was dissolved in 10 μL of dimethyl sulfoxide (Thermo D12345) and added to a 1.5 mL centrifuge tube containing the O3, O4, and O5 solutions. Then, 2 μL of 1 mM CuSO4, 2 μL of 10 mM sodium ascorbate, and 1 μL of 5 mM tris(3-hydroxypropyltriazolylmethyl)amine (THPTA, Sigma, Lot No. 762342) were added and allowed to react overnight at room temperature. The reaction samples were divided into two tubes and desalted and purified using nucleic acid purification columns (NEB T1030L).
[0114] SEQ ID NO: 3: P-GCTTCTCGTGNNTTTTTTCTCTC-NHS;
[0115] SEQ ID NO: 4: N3-CCCNNNTTTTTTTTTGCTGTCTCTGTCGTCGTTTC;
[0116] SEQ ID NO: 5: AAACGACGACAGAAGACAGCAAAAAAAAATGTGGGNGAGAGAAAAAAACAACACGAGAAGCA.
[0117] The sequence shown in SEQ ID NO: 3 is the nucleic acid sequence O3, wherein the 5' end of O3 has a phosphate modification and the 3' end has an N-hydroxysuccinimide ester group (NHS ester);
[0118] The sequence shown in SEQ ID NO: 4 is the nucleic acid sequence O4, and the 5' end of O4 has an azide modification;
[0119] The sequence shown in SEQ ID NO: 5 is the nucleic acid sequence O5.
[0120] O3 and O4 are used to connect to the polypeptide, and O5 forms a stable complementary structure with O3 and O4, bringing the reactive groups closer together and increasing reaction efficiency. The "N" in the O3, O4, and O5 sequences above represents an abasic site (AP site), a structure with no base but only deoxyribonucleic acid and phosphate, as shown in Formula II.
[0121] 5. Library Sequencing
[0122] Since nanopore sequencing is used in this example, the coupling product is added to the adapter and then nanopore sequencing is performed to obtain different current signals and perform data analysis.
[0123] 6. Experimental Results
[0124] 1. Results after peptide modification
[0125] In this example, a C18 column (059142) was used, with a mobile phase of water and acetonitrile. The gradient method was 5% to 40% acetonitrile by volume at 0.3 ml / min for 20 minutes. By comparing the difference in the peak elution time between the unmodified peptide and the double-end modified peptide samples, the HPLC results are shown in Figure 6, which shows that the modification conversion rate of the peptide is high.
[0126] 2. Library electrophoresis results
[0127] In this example, electrophoresis was performed on a 12% polyacrylamide denaturing gel at 150V for 90 minutes. The results are shown in Figure 7. In Figure 7, the first lane represents a DNA marker, in this example, a Thermofisher Ultra Low Range DNA Molecular Weight Standard Ladder (10597012); the second lane represents the starting nucleic acid sequence O3; the third lane represents the starting nucleic acid sequence O4; the fourth lane represents the starting nucleic acid sequence O5; the fifth lane represents a blank control in which O3, O4, and O5 are added to the reaction reagents; and the sixth lane represents the nucleic acid bands of the sample after the reaction of O3, O4, and O5 with the peptide. The bands in the solid-line frame represent the O3-peptide-O4 double complex band.
[0128] 3. Sequencing results
[0129] The prepared library is subjected to nanopore sequencing, and different current signal graphs as shown in FIG8 are obtained through data analysis. Through data analysis, the corresponding polypeptide information can be obtained and the sequence information of the protein can be obtained by comparison.
[0130] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects:
[0131] The present invention provides a full-process solution from proteins of unknown sequence to sequenceable polypeptide libraries. The library construction process is single. Since the content of lysine in proteins is moderate, it can be specifically enzymatically hydrolyzed into a polypeptide sequence library containing the same functional group (such as amino group, ε-carboxyl group or phenolic hydroxyl group) at the C-terminus, deconstructing the higher-order structure of trace proteins, which is beneficial to the interpretation of over-hole signals and avoiding negative signals caused by pore blockage or instability. This site-specific enzymatic hydrolysis step is applicable to all proteins and has no special site length requirements. The amino groups at one or both ends of the polypeptide of a single sample are modified simultaneously, such as using 1H-imidazole-1-sulfonyl azide hydrochloride to convert it into a double-terminal azide group, or using a reductive amination reaction to connect click groups such as cyclooctyne, thereby coupling specific nucleic acid sequences, which is beneficial to high-throughput library construction and saves time and effort. In addition, the above-mentioned library construction method and reaction system are mild and do not damage the protein sequence information.
[0132] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for constructing a protein sequencing library, characterized in that, The construction method includes: Using an endonuclease to digest a target protein to obtain two or more polypeptide fragments with a free amino group at the N-terminus and the same functional group at the C-terminus, where the functional group includes an amino group, an ε-carboxyl group, or a phenolic hydroxyl group; Using the functional group at the C-terminus and the free amino group at the N-terminus to couple the polypeptide fragments with nucleic acids to obtain the protein sequencing library; Wherein, the nucleic acids are respectively coupled with the functional group at the C-terminus and the free amino group at the N-terminus of the polypeptide fragments.
2. The construction method according to claim 1, characterized in that The connection between the nucleic acid and the free amino group at the N-terminus of the polypeptide fragment includes: Reacting a first N-terminal reaction group with the free amino group to form a first N-terminal chemical bond, and realizing the coupling of the N-terminus of the polypeptide fragment with the nucleic acid through the first N-terminal chemical bond; or Converting the free amino group into a first modification group, and the first modification group reacts with a second N-terminal reaction group to form a second N-terminal chemical bond, and realizing the coupling of the N-terminus of the polypeptide fragment with the nucleic acid through the second N-terminal chemical bond; Preferably, the second N-terminal chemical bond includes: a triazole structure, an eight-membered nitrogen-containing heterocyclic structure, a β-hydroxy aldehyde structure, a β-hydroxy ketone structure, an amino acid structure, a polyheterocyclic structure, an aryl structure, or an olefin structure.
3. The construction method according to claim 1, characterized in that, The amino acid at the C-terminus of the polypeptide fragment is lysine, the functional group is the amino group, and the endonuclease is lysine C endonuclease.
4. The construction method according to claim 3, characterized in that, The connection between the nucleic acid and the amino group at the C-terminus of the polypeptide fragment includes: Reacting a first C-terminal reaction group with the amino group to form a first C-terminal chemical bond, and realizing the coupling of the C-terminus of the polypeptide fragment with the nucleic acid through the first C-terminal chemical bond; or Converting the amino group into a modification group, and the modification group reacts with a second C-terminal reaction group to form a second C-terminal chemical bond, and realizing the coupling of the C-terminus of the polypeptide fragment with the nucleic acid through the second C-terminal chemical bond; Preferably, the second C-terminal chemical bond includes: a triazole structure, an eight-membered nitrogen-containing heterocyclic structure, a β-hydroxy aldehyde structure, a β-hydroxy ketone structure, an amino acid structure, a polyheterocyclic structure, an aryl structure, or an olefin structure.
5. The construction method according to claim 1, characterized in that, The amino acid at the C-terminus of the polypeptide fragment is glutamic acid, the functional group is the ε-carboxyl group, and the endonuclease is glutamic acid C-terminal restriction endonuclease.
6. The construction method according to claim 5, wherein The connection between the nucleic acid and the ε-carboxyl group at the C-terminus of the polypeptide fragment includes: Reacting a third C-terminal reaction group with the ε-carboxyl group to form a third C-terminal chemical bond, and realizing the coupling of the C-terminus of the polypeptide fragment with the nucleic acid through the third C-terminal chemical bond; or Converting the ε-carboxyl group into a modification group, and the modification group reacts with a fourth C-terminal reaction group to form a fourth C-terminal chemical bond, and realizing the coupling of the C-terminus of the polypeptide fragment with the nucleic acid through the fourth C-terminal chemical bond; Preferably, the fourth C-terminal chemical bond includes a ketone structure.
7. The construction method according to claim 1, wherein The amino acid at the C-terminus of the polypeptide fragment is one or more of tyrosine, phenylalanine, or tryptophan, the functional group is the phenolic hydroxyl group, and the endonuclease is chymotrypsin.
8. The construction method according to claim 7, wherein The connection between the nucleic acid and the phenolic hydroxyl group at the C-terminus of the polypeptide fragment includes: React the fifth C-terminal reactive group with the phenolic hydroxyl group to form a fifth C-terminal chemical bond, and achieve the coupling of the C-terminal of the polypeptide fragment with the nucleic acid through the fifth C-terminal chemical bond; or Convert the phenolic hydroxyl group into a modifying group, and the modifying group reacts with a sixth C-terminal reactive group to form a sixth C-terminal chemical bond, and achieve the coupling of the C-terminal of the polypeptide fragment with the nucleic acid through the sixth C-terminal chemical bond; Preferably, the sixth C-terminal chemical bond includes a sulfonic acid bond or a triazole structure.
9. The construction method according to claim 1, wherein The nucleic acid includes a specific molecular tag.
10. The construction method according to claim 1, characterized in that, After the coupling, remove the polypeptide fragments that have not undergone the coupling to obtain the protein sequencing library.
11. The construction method according to any one of claims 1 to 10, characterized in that For multiple target proteins, after constructing corresponding protein sequencing libraries for each target protein using different nucleic acids by using the construction method, mix the protein sequencing libraries of multiple target proteins to obtain a high-throughput protein sequencing library.
12. A method for sequencing a protein, characterized in that, The sequencing method includes: constructing a library for a protein using the protein sequencing library construction method according to any one of claims 1 to 11 to obtain the protein sequencing library; sequencing the protein sequencing library using a sequencing technology to obtain the sequence information of the protein.
13. The sequencing method according to claim 12, characterized in that, The technology includes nanopore protein sequencing technology or single molecule fluorescence recognition detection technology.
14. The sequencing method according to claim 13, wherein In the sequencing method, using the nanopore protein sequencing technology, by pulling the nucleic acid through the nanopore protein, the protein is directed to pass through the nanopore protein and complete the sequencing.
Citation Information
Patent Citations
Protein / polypeptide sequencing method adopting Aerolysin nanopores
CN112480204A
Nanopore channel monomolecular protein sequencer
CN112578106A
Nanopore proteomics
WO2022245209A2