High-throughput pan-sample protein interaction screening method
By constructing a generalized open reading frame and a multi-system screening system, combined with single-cell sequencing and droplet PCR amplification technology, high-throughput and full-coverage protein interaction screening was achieved, solving the problems of low throughput and limited scope in traditional methods, and making it suitable for PPI network identification in multiple species.
Patent Information
- Application Number
- CN202512030316.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-12-30
AI Technical Summary
Traditional protein interaction screening methods cannot achieve high-throughput, comprehensive library-to-library screening, and suffer from low throughput and limited scope. Furthermore, deep learning models do not perform well in predicting new protein interactions.
A cDNA library was constructed using a generalized open reading frame, and combined with a multi-system, high-accuracy library-to-library PPI screening system, high-throughput PPI identification was achieved through single-cell transcriptome sequencing (scPPI-seq), second-generation sequencing (dPPI-seq) for direct plasmid copy amplification by droplet PCR, and third-generation sequencing (dTPPI-seq).
It achieves high-throughput, comprehensive protein interaction screening, breaking through the bottleneck of traditional screening. It is applicable to the identification of PPI networks among different species and has the advantages of broad sample size, high throughput, high accuracy, and low cost.
Smart Images

Figure CN121428075A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of protein processing and screening, and in particular to a high-throughput general sample protein interaction screening method. Background Technology
[0002] Proteins, as the main executors of cellular functions, often participate in various intracellular life activities through interactions, such as signal transduction, metabolic regulation, transcription and translation, and cell cycle control. Therefore, protein-protein interactions (PPIs) within the host cell form the basis of the cellular functional network and play a crucial role in maintaining normal physiological functions and preventing the occurrence of pathological processes.
[0003] Meanwhile, an increasing number of studies are focusing on host-microbe protein-protein interactions (HM-PPIs). Human hosts coexist with a large number of symbiotic microorganisms (such as the gut microbiota) for extended periods. These microorganisms can directly or indirectly interact with host proteins through their secretions or surface proteins, thereby regulating immune, metabolic, neural, and barrier functions. Pathogens, on the other hand, can use their effector proteins to interfere with host signaling pathways via PPIs to evade immunity, promote colonization, or spread infection. For example, pathogens such as Helicobacter pylori, Salmonella, and Escherichia coli have all been shown to manipulate host cell functions through a series of sophisticated protein-protein interaction mechanisms.
[0004] Traditional protein interaction screening methods (such as Y2H and AP-MS) can only perform single-pair interaction screening, failing to achieve full coverage at the library-to-library level, and suffer from low throughput and limited scope. Although deep learning-based large language models (such as AlphaFold) have been widely used in PPI prediction in recent years, they still have problems such as limited binding affinity datasets, oversimplification of task difficulty, inability to handle "new" proteins, and inability to handle multi-protein collaborations. Therefore, establishing a high-throughput, general-sample library-to-library protein interaction screening method is of great significance for helping research institutions and enterprises improve R&D efficiency, accelerate the development of new drugs and materials, assist in the development of anti-aging related antibodies, molecular glues, cell therapies, vaccines, and diagnostics, and facilitate the development of new products in fields such as sustainable chemical synthesis and crop protection. Summary of the Invention
[0005] To address the problems existing in the background technology mentioned above, the present invention provides a high-throughput screening system for "library-to-library" protein-protein interactions and a method for using it. It mainly includes a general sample open reading frame acquisition system, a multi-system high-accuracy "library-to-library" PPI screening system, and single-cell transcriptome sequencing (scPPI-seq) after induction screening, as well as second-generation sequencing (dPPI-seq) based on droplet PCR for direct plasmid copy amplification and third-generation dTPPI-seq.
[0006] like Figure 1 As shown, the technical solution of the present invention is as follows: 1) Construction of generic sample libraries mRNA was extracted from the target sample in vitro, and then a pan-sample cDNA library was obtained using modified primers. 2) Homogenization of pansample cDNA libraries; 3) Plasmid modification: Modify the plasmid of the dual-hybrid system to obtain the modified plasmid; 4) Homologous recombination was performed on the sample library and the modified plasmid of the double hybrid system, and the plasmid was introduced into the double hybrid system organisms to screen for organisms with mutual interaction libraries (positive PPI clones) and single-cell labeled cDNA mutual interaction libraries. 5) After obtaining positive PPI clones, single-cell sequencing high-throughput PPI pair identification was performed on organisms with interaction libraries using three methods to construct a pan-sample protein interaction network.
[0007] Step 1) specifically refers to one of the following two processing procedures: The first method involves first extracting mRNA from the target sample, then using modified 3'RT Oligo-dT primers and TSO primers to obtain a full-length eukaryotic cDNA library as a pan-sample cDNA library using Smart-seq 3. This includes: Some dual-hybrid systems fuse the library to the N-terminus of a plasmid vector. When obtaining the library, Oligo(dT) magnetic beads are first used to capture eukaryotic mRNA specifically via the Poly-A tail. Then, a modified 3'RT random primer is used for reverse transcription to obtain a cDNA library with fragments of uneven length. The modified 3'RT random primer sequence is SEQ ID No.1, i.e., ACTCTGCGTTGATACCACTGCNNNNNNNN. This avoids "false negatives" caused by fusion sites and improves the full coverage and sensitivity of interaction detection.
[0008] The modified 3'RT Oligo-dT primer sequence is SEQ ID No. 7, namely ACTCTGCGTTGATACCACTGCTTTTTTTTTTTTTTTTTTTTT, and the modified TSO primer sequences include SEQ ID Nos. 8-10, namely TCGACTCTAGAGGATCCCrGrGrG, TCGACTCTAGAGGATCCCrG, and TCGACTCTAGAGGATCCCrGrG.
[0009] The above sequences were optimized based on the sequences of the plasmid vector. The 3'RT primer sequence was added to the vector through plasmid modification and also corresponds to the sequence on the hydrogel-encoded microspheres. The TSO primers were designed based on the 5' homologous arm sequence of the vector cloning site. The TSO primer sequence was changed according to the homologous arm for different vectors, while the 3'RT primer sequence remained unchanged.
[0010] The second method involves using a polycistronic cDNA library from prokaryotes. First, rRNA is digested using an rRNA digestion kit. Then, the 5' end of the mRNA is capped using the Faustovirus Capping Enzyme, Cap 2'-O-methyltransferase, and 3'biotin-GTP. Finally, a Poly-A tail is added to the 3' end of the mRNA using E. coli Poly-A polymerase and ATP. The purified mRNA is then purified using mRNA purification beads. A full-length cDNA library is constructed using the Smart-seq 3 protocol, specifically by using modified 3'RT Oligo-dT primers and TSO primers with Smart-seq 3 to obtain the full-length cDNA library from prokaryotes, which serves as a pan-sample cDNA library.
[0011] Step 2) specifically involves using Duplex-Specific Nuclease (DSN) to homogenize the cDNA in the pan-sample cDNA library, reducing the abundance of high-copy genes while retaining low-expression DNA molecules, thus making the gene concentration in the cDNA sample more uniform, improving screening efficiency and discovering interaction networks of extremely low-abundance genes.
[0012] Specifically, step 3) involves modifying the plasmids of bacterial or yeast double-hybrid systems for single-cell labeling by linking a linker sequence containing a homologous arm to form a modified plasmid. The linker sequence is SEQ ID No. 2, namely AAGCGTGGTATCAACGCAGAGT, which has a homologous sequence with the modified 3'RT Oligo-dT primer. The linker linker in the plasmid is located at the 3' end of the MSC multiple cloning site of the vector, and the linker site sequence is SEQ ID No. 3, namely ACTCTGCGTTGATACCACTGCTT.
[0013] The bacterial dual-hybrid systems mentioned include pMRBAD-Z-CGFP / pRT11a-Z-NGFP, pBT / pTRG, pKT25(pKNT25) / pUT18C(pUT18), pET11a-link-NGFP / pMRBAD-link-CGFP, pTET-GFP11 / pET-GFP1-10, pSLBLC-c-myc-hFADD / pSBLNL-hFasDD-His, pQE / pREP, and pACYC184 / pGSTDHFR. The yeast dual-hybrid system mentioned includes the GAL4 system. This scheme is compatible with all dual-plasmid bacterial dual-hybrid and yeast dual-hybrid screening schemes.
[0014] Step 4) specifically involves: firstly, homologous recombination of the target sample cDNA library with the modified plasmid of the dual-hybrid system to form their respective bait protein and prey protein-particle libraries; then, the plasmid library is transferred to the corresponding dual-hybrid system organism (bacteria or yeast) via electrotransduction (high efficiency) or chemical transduction, and positive PPI combination clones are screened in the selection medium to obtain organisms with protein-protein interaction pairs. Gene expression is induced by IPTG to obtain a single-cell labeled cDNA interaction library.
[0015] The screening medium is such as M63 medium (5xM63 medium formula: weigh 10 g of (NH4)2SO2, 68 g of KH2PO4, 2.5 mg of FeSO4·7H2O and 5 mg of vitamin B1, add deionized water to a final volume of 1 L, adjust the pH to 7.0 with KOH and then autoclave).
[0016] The modified plasmid contains the linker sequence AAGCGTGGTATCAACGCAGAGT, and the positive result is that the bacterial BACTH double hybrid system turns blue under the selection medium M63 medium. The bait protein library is composed of a modified plasmid pKT25 or pKNT25 from a dual-hybrid system and a cDNA library through homologous recombination. The prey protein is composed of a modified plasmid pUT18C or pUT18 from a dual-hybrid system and a cDNA library through homologous recombination.
[0017] Step 5) employs single-cell transcriptome sequencing based on random primers (scPPI-seq), second-generation sequencing based on droplet PCR for direct plasmid copy amplification (dPPI-seq), or third-generation sequencing (dTPPI-seq).
[0018] Step 5) specifically involves single-cell transcriptome sequencing (scPPI-seq) based on random primers: First, the screened bacterial or yeast organisms were fixed in 4% PFA solution for 8-16 hours to allow protein-nucleic acid cross-linking. Lysozyme and Zymolyase enzymes were then added to open the cell walls of the bacteria and yeast, respectively. Next, the cell membrane was permeabilized with 0.4% Triton solution. Then, in situ reverse transcription and TdT terminal transferase were performed to add Poly(A) to the 3' hydroxyl end of cDNA. Oligo(dT) hydrogel-encoded microspheres were then added to the bacteria or yeast for single-cell encapsulation and cDNA double-strand amplification, resulting in a cDNA library with cell barcode markers. The cDNA library was then used to repair ends and ligate P5 / P7 adapters to construct an Illumina sequencing library using the VAHTS Universal Pro DNA Library Prep Kit for Illumina. Finally, the obtained Illumina sequencing library was sequenced at paired ends (150 bp) using an Illumina X plus 25B sequencer, for a total of 300 bp. bp sequencing compares the sequenced data with the reference genome of the target sample species to identify interacting protein pairs, recognize cDNA combinations with the same cell barcode, and construct an interacting protein network.
[0019] The high-temperature resistant droplet generation oil is specifically QX200™ Droplet Generation Oil for EvaGreen (#1864005), the aqueous phase stabilizer is specifically a digital PCR-specific aqueous phase stabilizer purchased from Nanjing Stone Gene Technology Co., Ltd., and the fragmentation reagent is specifically UltraClean DNA Library Prep Kit V3 for Illumina (UTD521).
[0020] The second-generation sequencing based on droplet PCR direct plasmid copy amplification (dPPI-seq) in step 5) specifically involves: Since the transfected plasmid contains more than 5 copies, modified hydrogel-encoded microspheres containing linker sequences were directly added to the screened positive PPI bacteria or yeast for single-cell encapsulation along with heat-resistant droplet-generating oil and aqueous phase stabilizer, followed by PCR amplification. After amplification, the amplified products were fragmented using a transposase DNA library construction kit (Novi Station UTD521). Then, a cDNA sequencing library containing cell barcode sequences was constructed using P5 / P7 adapter primers containing cell barcode linkers. Finally, the obtained cDNA sequencing library was sequenced at both ends of 150 bp using an Illumina X plus 25B sequencer, for a total of 300 bp. The sequenced sequences were compared with the reference genome of the target sample species to identify interacting protein pairs and cDNA combinations with the same cell barcode, thus constructing an interacting protein network.
[0021] The P5 primer sequence is: SEQ ID No. 41, AATGATACGGCGACCACCGAGATCTACACCTCTCTATTCGTCGGCAGCGTCAGATGTGTATAAGAGACAGGTGAGTGATGGTTGAGGATGTGTGGAGATA; The P7 primer sequence is: SEQ ID No. 42, CAAGCAGAAGACGGCATACGAGATTAAGGCGAGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG.
[0022] The single-stranded DNA on the hydrogel-encoded microspheres containing the linker sequence consists of an upstream primer complement fragment, a barcode, a unique multiplex index (UMI), and a linker sequence. The upstream primer complement fragment binds to the upstream primer during PCR amplification. The barcode is used to label cDNA within the same cell septum. The UMI is a random sequence used to label each original cDNA. The linker sequence acts as a primer in PCR to complete single-cell labeling amplification.
[0023] In the above process, for yeast organisms, a mixture containing lysis buffer (mix) was added during PCR amplification. The mixture containing lysis buffer (mix) consisted of 2 mg / g cell Zymolase lysing enzyme reacting at 37°C for 30 min to lyse the cell wall.
[0024] For bacterial organisms, no mixture containing lysis buffer is added during PCR amplification.
[0025] In step 5), the third-generation sequencing based on droplet PCR for direct plasmid copy amplification (dTPPI-seq) employs one of the following methods: Since Illumina X plus 25B paired-end 150 bp sequencing can only identify PPI pairs, but cannot determine whether the identified protein interactions are false positives caused by frameshift mutations, performing third-generation full-length sequencing is the key to solving the problem. However, the full-length cDNA library is 300-3000 bp. Directly performing third-generation sequencing on PicBio or ONT will result in a large waste of chips and low data output due to the short sequence. To adapt to long-sequence sequencing on PicBio and ONT platforms,
[0026] A) The first type First, for the cDNA intergenic library obtained in step 4) with single-cell labeling, 15 pairs of adapter primers modified with deoxyuridine (dU) were used for secondary amplification. After amplification, the obtained libraries were mixed together and purified with magnetic beads. After purification, a single nucleotide gap was generated at the uridine position using USER enzyme. Finally, Endo VIII enzyme was used to cleave the double-stranded DNA to produce a single-stranded overhang that was separated from the double-stranded DNA and discarded, leaving the double-stranded DNA with sticky ends. The double-stranded DNA with sticky ends was then annealed and ligated with Taq high-fidelity DNA ligase (NEB M0647S) to form a long fragment of more than 10kb, which made the original 300 bp sequencing at least 15 times longer, reaching a DNA sequence of 10kb in length. Finally, third-generation sequencing was performed using Nanopore and PacBio platforms to improve the utilization rate of sequencing wells and sequencing efficiency of sequencing chips. The multiple pairs of IdeoxyU-modified adapter primer sequences are SEQ ID No. 4 and SEQ ID No. 5, respectively: F: GGAGTTGGAGTGAGTGGATGAGTGATG, R: GTGAGTGATGGTTGAGGATGTGTGGAGATA.
[0027] B) The second type First, the single-cell labeled cDNA intergenic library obtained in step 4) is processed using TdTase terminal transferase to add two sequences from either the dATP-dTTP pair or the dCTP-dGTP pair to the 3' hydroxyl end of the DNA molecules. The products with the two sequences added are then purified using magnetic beads. After purification, the two products are mixed and treated at 95°C for 1 min, followed by slow annealing at 0.1°C / s to 25°C. Finally, T4 DNA polymerase is added for DNA repair and ligation. T4 DNA ligase can directly tandemly generate long fragments of over 10 kb, facilitating third-generation sequencing. Finally, third-generation sequencing is performed using Nanopore and PacBio platforms to improve the utilization rate of sequencing wells and sequencing efficiency.
[0028] The high throughput mentioned in this invention refers to a high-throughput identification technology system that can identify tens of thousands to hundreds of thousands of positive PPI clones in a single operation.
[0029] The beneficial effects of this invention are: This invention constructs a general sample open reading frame acquisition system, a multi-system high-accuracy "library-to-library" PPI screening system, and a multi-mode droplet microfluidic high-throughput sequencing identification system. It is suitable for high-throughput screening after obtaining tens of thousands of positive PPI combination clones through "library-to-library" bacterial or yeast double hybridization. It can identify PPI network identification in a wide range of scenarios, including intra-species, inter-species (model species vs. non-model species, bacteria vs. host), and hybridization technology systems (bacterial two-hybrid, yeast one-hybrid). It breaks through the bottleneck of traditional screening that can only perform "one-to-one" or a few "one-to-many" screenings, and has the advantages of general sample, high throughput, high accuracy, low cost, and full coverage of protein interaction identification. Attached Figure Description
[0030] Figure 1 This is a schematic diagram of the present invention; in the diagram, mRNA: Messenger ribonucleic acid; cDNA: Complementary DNA; CB: Cell barcode; PPI: Protein-protein interaction networks; Figure 2 This is a schematic diagram of plasmid modification and homologous recombination. Figure 3 A graph of the full-length open reading frame for mice and humans is constructed; in the graph, k: Kilo; bp: Base pair; Figure 4 This is a screening diagram for positive clones; in the diagram, M: Marker; bp: Base pair; Figure 5 This is a screening plot of LA plates for the library; in the plot, Amp: Ampicillin; Kan: Kanamycin; IPTG: Isopropyl β-D-1-thiogalactopyranoside; X-gal: 5-bromo-4-chloro-3-indolyl β-D-galactopyranoside; Figure 6 This is a diagram showing library induction and screening in M63 medium; in the diagram, DsRed: Red fluorescent proteins; Figure 7 This is a single-cell sequencing diagram; Figure 8 Diagram of the microsphere and library structure encoded by dPPI hydrogel; Figure 9 Contamination rate assessment graph for single-cell high-throughput screening; in the graph, hs: human species; mm: Musmusculus; Figure 10 This is a diagram of the construction of a full-length cDNA library using third-generation sequencing; in the diagram, RFU: Relative fluorescence unit; bp: Base pair; Figure 11 This is a PCR clone diagram of the mouse candidate protein cDNA; in the diagram, M: Marker; bp: Base pair; Figure 12 The image shows the screening results of candidate proteins for yeast double hybridization. Detailed Implementation
[0031] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] The embodiments of the present invention are as follows: Example 1:
[0033] 1. Obtain the target library by direct TRIzol lysis. RNA from the target sample was obtained by direct lysis according to the TRIzol instructions. For eukaryotic samples, full-length open reading frames were obtained directly using the modified Smart-seq 3 method. For prokaryotic samples, to ensure the efficiency of 5' capping and 3' poly-A tailing, the rRNA was digested using the NEBNext® rRNA Depletion kit before reverse transcription.
[0034] Next, the 5' capping of mRNA was performed using the Faustovirus Capping Enzyme, Cap 2´-O-methyltransferase, and 3'biotin-GTP. A 3' Poly-A tail structure was added to the mRNA using E. coli Poly-A polymerase and ATP. A full-length cDNA library was constructed from the magnetically purified mRNA using the Smart-seq 3 protocol. The specific method is as follows: RT reaction is performed using an RNaseH-free MMLV RT enzyme (such as Maxima H-minus reversetranscriptase enzyme (Thermo Scientific)). The reaction system contains 25 mM Tris-HCl pH 8.0-8.4, 30 mM NaCl, 2.5 mM MgCl2, 1 mM GTP, 8 mM DTT, 0.25 U RNase inhibitor, 0.3 mM dNTPs, 0.1 uM TSO primer and 0.1 uM 3' RT primer, and RNA. RNA, dNTPs and 3' RT primer are denatured at 72°C for 5 min beforehand. Then, the system is immediately placed on ice and the remaining reagents are added. The RT reaction is performed at 42°C for 90 min, followed by 10 cycles of 50°C for 2 min and 42°C for 2 min. After the RT reaction, cDNA library was amplified for 25 cycles using 2X KAPA HiFi HotStart ReadyMix. After amplification, the cDNA library was purified using 0.6x VAHTS DNA Clean Beads (N411).
[0035] 2. To ensure the cDNA in the library is free of repetitive sequences and covers more transcript information, we used Duplex-Specific Nuclease (DSN) for cDNA homogenization. This was performed using the Puente Bio cDNA homogenization kit (FZ1031). First, 200 ng of the cDNA library was mixed with 4x Hybridization buffer and denatured at 98°C for 2 min, then incubated at 68°C for 5 h. Then, 1 μL of DSN and 1 μL of 10×DSN Reaction buffer were added, gently mixed, briefly centrifuged, and incubated at 68°C for 7-20 min to degrade the double-stranded cDNA. Finally, 10 μL of 2×DSN Stop Buffer was added, gently mixed, briefly centrifuged, and incubated at 65°C for 5 min to terminate the reaction. The remaining homogenized single-stranded cDNA was amplified by PCR for library cloning. Figure 3 ).
[0036] 3. High-throughput screening and sequencing solutions Currently, there are many low-throughput yeast and bacterial two-hybrid systems based on one-to-one protein interaction. This invention conducts library-to-library high-throughput single-cell screening based on traditional protein interaction bacterial and yeast two-hybrid systems. To evaluate the cross-contamination of single-cell sequencing technology, the BACTH system is used as an example.
[0037] Because the library-to-library bacterial double hybridization process makes it difficult to directly determine whether the recombinant plasmid is correctly expressed, the DsRed gene expressing red fluorescent protein, a human cDNA library, and a mouse cDNA library were homologously recombined into the bait expression vector pUT18C and electroporated into *E. coli* to extract plasmids. The mouse cDNA library was homologously recombined into the bait expression vector pKNT25 and electroporated into BTH101 cells to prepare competent cells. The DsRed-pUT18C, human cDNA library-pUT18C, and mouse cDNA library-pUT18C plasmids were then electroporated again into mouse cDNA library-pKNT25 BTH101 *E. coli* competent cells and plated on M63 LA plates containing X-gal+IPTG+Amp+Kan for positive clone screening. Figures 2-5 All positive PPI clones were collected and expanded in LB medium containing double antibiotics, followed by IPTG-induced expression. Red fluorescent protein was used as an indicator signal for gene expression. Figure 6 When the bacterial culture turns red, bacteria are collected simultaneously. The high throughput of the single-cell system and the low contamination rate are evaluated using human-mouse library double-hybrid Escherichia coli.
[0038] 3.1 Identification of protein interactions using random primer smRandom-seq (scPPI-seq, single-cell high-throughput protein interaction scPPI-seq) The collected bacteria were fixed overnight in 4% paraformaldehyde to cross-link the intracellular RNA, DNA, and proteins. Lysozyme was used to digest the cell wall, permeabilizing the fixed bacteria to facilitate the next step of in situ reverse transcription. The microorganisms were used as reaction vessels for the in situ reaction. Random primers were added to bind to the intracellular RNA, and total RNA was captured and synthesized into cDNA. A poly-A tail was added in situ to the 3' end of the cDNA using a terminal transferase (TdT). Each step of the process was followed by washing with buffer 3-8 times to prevent reagent residue from affecting subsequent reactions. Individual bacteria and labeled microbeads were encapsulated into droplets using a microfluidic device. Poly-T primers were released from the microbeads by enzymatic digestion, digesting the bacterial RNA to release cDNA from the bacteria. The poly-T primers bound to the poly-A tail at the end of the cDNA, subsequently extending the cDNA to add specific coding, and a molecular tag (UMI) was added to each cDNA. After demulsification, cDNA was collected and purified, expanded, and sequencing adapters were added to construct a sequencing library. The cDNA product of rRNA was digested with Cas9, and the cDNA product of mRNA was enriched for high-throughput sequencing. The obtained sequencing data were subjected to quality control and filtering before alignment to the Ecoli_bw25113 genome. Figure 7 Unmapping reads were extracted and aligned to human and mouse reference genes, respectively. Feature Counts software was used to count the number of mouse and human transcripts detected in the same barcode and their corresponding UMIs. A UMI threshold was designed to exclude false positives. The higher the UMI of mouse and human transcripts detected by a barcode, the fewer false positives generated by sequencing and the higher the reliability of the interaction pairs. Finally, an interaction network was constructed from the selected interaction pairs.
[0039] Comparison of valid Top 2000 barcodes revealed that 80.3% (1606) of the barcodes identified two or more genes, 84.4% (1688) of the barcodes contained the DsRed gene, and 75.9% (1281) of the barcodes containing two or more genes had the DsRed gene UMI number in the Top 2. A single-cell sequencing PPI identification method based on random primer droplet microfluidic “DsRed library” was successfully established (Table 1).
[0040] Table 1. Screening results of the Dsred-mouse library using dPPI-seq Barcode Dsred-UMI Mouse-gene-UMI Mouse-gene-UMI Barcode1 Dsred-2 Dync1h1-1478 Gm17111-2 Barcode2 Dsred-11 mt-Cytb-804 Barcode3 Dsred-48 mt-Cytb-43 Alb-7 Barcode4 Dsred-81 mt-Cytb-325 Alb-1 Barcode5 Dsred-1 Dync1h1-392 Barcode6 Dsred-47 mt-Cytb-47 mt-Co3-4 Barcode7 Dsred-1 Dync1h1-347 Barcode8 Dsred-8 Alb-298 mt-Cytb-11 Barcode9 Dsred-38 mt-Cytb-18 Alb-6 Barcode10 Dsred-1 mt-Cytb-39 Barcode11 Dsred-1 mt-Cytb-320 Barcode12 Dsred-144 Eef1a1-161 Slc22a27-3 Barcode13 Dsred-1 mt-Co1-300 mt-Cytb-15 Barcode14 Dsred-1 Eef1a1-294 Slc22a27-10 Barcode15 Dsred-2 mt-Cytb-303 mt-Co3-1 3.2 Protein interaction identification based on digital PCR system (dPPI-seq) Referring to the smRandom-seq article, 96*96*48 cell barcode sequences were synthesized. Based on vector characteristics, the last adapter sequence was modified to AAGCGTGGTATCAACGCAGAGT, SEQ ID No. 2. The first cell barcode segment was amino-modified at the 5' end and immobilized on polyacrylamide hydrogel microspheres. Different cell barcodes were ligated using T7 ligase according to the split pool method. After completion, NaOH treatment was used to form dsDNA single strands. Figure 8 ).
[0041] Using microfluidic devices, hydrogel-encoded microspheres and individual bacteria were directly encapsulated in water-in-oil droplets. To prevent droplet fusion at high temperatures, 10% aqueous phase stabilizer (purchased from Hangzhou Chuangyan Xinghe) was added to the aqueous phase, along with 25% OptiPrep density gradient medium, 2X KAPA HiFi HotStart ReadyMix, 400 E. coli cells / ul selected with double antibodies, 2.5ul PUT18C and 2.5ul PKNT25 vector primers, and 5ul USER enzyme (M5505L). The incubation process was 37℃ for 30 min, 95℃ for 3 min, followed by 30 cycles of 95℃ for 30 s, 60℃ for 20 s, 72℃ for 2 min 30 s, and a final incubation at 72℃ for 5 min, and storage at 4℃. After PCR amplification, the cDNA library was purified using 0.6x VAHTS DNA Clean Beads (N411). After amplification for three cycles using barcode adapter primers and PUT18C and PKNT25 vector primers, Illumina libraries were constructed using the VAHTSUniversal Pro DNA Library Prep Kit for Illumina.
[0042] Comparison of selected effective cells revealed that 81% of cells identified as having barcodes from one mouse and one human, 10% had empty vectors, and 9% contained cells with two or more genes (contamination). The data are reliable, and a protein interaction identification method based on a digital PCR system has been successfully established. Figure 9 ).
[0043] 3.3 Direct plasmid copy amplification using droplet PCR based on third-generation sequencing (dTPPI-seq) Since Illumina X plus 25B paired-end 150 bp sequencing can only identify PPI pairs but cannot determine whether there are frameshift mutations in the library during PCR amplification, performing third-generation full-length sequencing is the key to solving the problem. However, full-length cDNA libraries are 300-3000 bp. Directly performing third-generation sequencing on PicBio or ONT would result in a large amount of chip usage due to the short sequence length. Therefore, long-sequence sequencing compatible with PicBio and ONT platforms is necessary.
[0044] 3.3.1 After pre-amplification of cDNA, cDNA libraries were amplified again using adapter primers modified with deoxyuridine (dU). The adapter sequences are shown in Table 2, as shown in SEQ ID No. 20. All libraries were mixed together and purified. After purification, a single nucleotide gap was generated at the uridine position using USER enzyme and a single-stranded overhang was generated by EndoVIII cleavage. Then, sticky end annealing and ligase were used to ligate the overhang to form a long fragment (Sequence 1).
[0045] ONT sequencing yielded the full-length open reading frame of the human EEF1G gene, as shown in SEQ ID No. 6, specifically below, with the underlined and bolded sequence being the adapter sequence: GGTCGACTCTAGAGGATCCCGACTC TGCGTTGATACCACTGCTT ACTCTGCTTGATACCACTGCTACTCTGCGTTGATACCACTGCTACTCTGCGTTGATACCACTGCTTAGACGAGATTGTATCTTCTCTCCGTATTACCGACCAAGAGCCGTGT TATCTCCACACATCCTCAACC ATCACTCAC .
[0046] 3.3.2 After cDNA amplification, dATP-dTTP or dCTP-dGTP pairs were added to the 3' hydroxyl end of the DNA molecule using TdTase terminal transferase. After purification, the samples were mixed and annealed, and then subjected to DNA repair and ligation with T4 DNA polymerase to tandemly form long fragments exceeding 10 kb. Figure 10 ).
[0047] Table 2. 15 pairs of linker primers modified with deoxyuracil (dU) A-F AGCTTACTTGTGAAGATGGAGTTGGAGTGAGTGGATGAGTGATG B-F ACTTGTAAGCUGTCTAUGGAGTTGGAGTGAGTGGATGAGTGATG C-F ACTCTGUCAGGTCCGAUGGAGTTGGAGTGAGTGGATGAGTGATG D-F ACCTCCTCCUCCAGAAUGGAGTTGGAGTGAGTGGATGAGTGATG E-F AACCGGACACACUTAGUGGAGTTGGAGTGAGTGGATGAGTGATG F-F AGAGTCCAAUTCGCAGUGGAGTTGGAGTGAGTGGATGAGTGATG G-F AATCAAGGCUTAACGGUGGAGTTGGAGTGAGTGGATGAGTGATG H-F ATGTTGAAUCCTAGCGUGGAGTTGGAGTGATGGATGAGTGAT IF AGTGCGTUGCGAATTGUGGAGTTGGAGTGAGTGGATGAGTGATGATG JF AATTGCGUAGTTGGCCUGGAGTTGGAGTGATGAGTGATGAGTGATG KF ACACTTGGUCGCAATCUGGAGTTGGAGTGAGTGGATGAGTGATG LF AGTAAGCCUTCGTGTCUGGAGTTGGAGTGAGTGGATGAGTGATG MF ACCTAGAUCAGAGCCTUGGAGTTGGAGTGAGTGGATGAGTGATG NF AGGTAUGCCGGUTAAGUGGAGTTGGAGTGAGTGGATGAGTGATGATG OF AAGUCACCGGCACCUTUGGAGTTGGAGTGAGTGGATGAGTGATG BR ATAGACAGCUTACAAGUGTGAGTGATGGTTGAGGATGTGTGGAGATA CR ATCGGACCUGACAGAGUGTGAGTGATGGTTGAGGATGTGTGGAGATA DR ATTCUGGAGGAGGAGGUGTGAGTGATGGTTGAGGATGTGTGGAGATA IS ACTAAGTGUGTCCGGTUGTGAGTGATGGTTGAGGATGTGTGGAGATA FR ACTGCGAAUTGGACTCUGTGAGTGATGGTTGAGGATGTGTGGAGATA GR ACCGTUAAGCCTTGATUGTGAGTGATGGTGAGTGGTTGAGGATGTGTGGAGATA HR ACGCTAGGAUTCAACAUGTGAGTGATGGTTGAGGATGTGTGGAGATA IR ACAATCGCAACGCACUGTGAGTGATGGTTGAGGATGTGTGGAGATA JR AGGCCAACUACGCAATUTGAGTGATGGTTGAGGATGTGTGGAGATA KR AGATUGCGACCAAGTGUGTGAGTGATGGTTGAGGATGTGTGGAGATA LR AGACACGAAGGCUTACUGTGAGTGATGGTTGAGGATGTGTGGAGATA MRI AAGGCTCUGATCTAGGUGTGAGTGATGGTTGAGGATGTGTGGAGATA NR ACTUAACCGGCAUACCUGTGAGTGATGGTTGAGGATGTGTGGAGATA OR AAAGGUGCCGGUGACTUGTGAGTGATGGTTGAGGATGTGTGGAGATA PR ATCTCGAGCCACTTCATGTGAGTGATGGTTGAGGATGTGTGGAGATA Example 2: Construction of Mouse PPI Network This embodiment continues from Embodiment 1, constructing a real protein interaction network in mice and verifying its accuracy.
[0048] To verify the stability of the established high-throughput protein interaction screening method, a mouse library was homologously recombined into the bait expression vector pUT18C and electroporated into E. coli to extract plasmids. Simultaneously, the mouse library was homologously recombined again into the bait expression vector pKNT25 and competent cells were prepared. The mouse library-pUT18C plasmid was electroporated again into the mouse library-pKNT25BTH101 competent cells and plated on M63 LA plates containing X-gal+IPTG+Amp+Kan for positive clone screening. A mouse PPI network was constructed using dPPI-seq (Table 3).
[0049] Genes with the top two UMIs in the barcode pairs selected by dPPI-seq were used to predict the accuracy of the results using machine learning. Six genes were then randomly selected from the dPPI-seq results for cloning. Figure 11 Yeast two-hybridization was performed and found to have 3 pairs of strong interactions and 1 pair of weak interactions. Figure 12 The results corresponded to the dPPI-seq screening results, verifying the accuracy of this high-throughput screening.
[0050] Table 3 Mouse protein-protein interaction pairs Barcode Mouse-gene Mouse-gene Barcode1 Ywhag Hdgf Barcode2 mt-Cytb mt-Atp6 Barcode3 Kyat3 Etfb Barcode4 Hmgcs 2 Acaalb Barcode5 Ywhaq Ces1c Barcode6 Ftcd Aass Barcode7 Ywhaq Gclm Barcode8 Apoa4 Sugt1 Barcode9 Mapre2 Acol Barcode10 Edem2 mt-Cytb Barcode11 Thrsp Ptges3 Barcode12 Alb Ankrd17_ Barcode13 Rpl13a Rp16 Barcode14 Map2 S1c9a3r1 Barcode15 Serpinala Mtx2 Barcode16 Nubl Gda Barcode17 Fkbp4 Eef2 Barcode18 Alb Fabp1 Barcode19 Gm20425 Ndrg1 Barcode20 Apoa2 Tmem29 Barcode21 Alb Fam32a Barcode22 Gss Mup3 Barcode23 S1c9a3r1 Etfa Barcode24 Etfb Ptges3 Barcode25 Fbpl ubb Barcode26 Cyp2d9 Rp18 Barcode27 Acadvl Cyp2f2 Barcode28 Eif3e Sfpq Barcode29 Ldha Capn1 Barcode30 Apob Rp113a As shown in Table 4, the comparison results are as follows: Compared with existing technologies such as traditional yeast two-hybrid, AP-MS, and Alpha-Seq, the high-throughput, general-sample protein interaction screening method established in this example represents a significant performance breakthrough for the "library-to-library" protein interaction screening platform. It employs a comprehensive "library-to-library" screening mode, followed by single-cell sequencing based on droplet microfluidics, increasing throughput to an astonishing 10^65 times. 8 -10 9 Its level far exceeds the limited flux of other methods.
[0051] Meanwhile, the detection cost per interaction can be controlled to below $1, making it unparalleled in cost-effectiveness. More importantly, the platform can effectively achieve low false positives across multiple systems, solving the pain points of traditional methods such as high false positives and complex backgrounds. It enables broad-sample, high-throughput, high-accuracy, low-cost, and comprehensive protein interaction screening, providing an unprecedentedly powerful tool for large-scale protein interaction research.
[0052] Table 4. Performance Comparison between the "Library-to-Library" Protein Interaction Screening Platform and Existing Protein Interaction Screening Methods Methods / Platforms Filtering mode Flux Scale False positive / false negative cost Advantages Limitations Traditional yeast two-hybrid (Y2H) Single Bait to Prey (pair-by-pair filtering) Medium (10 4 -10 5 level)]]> There are many false positives, and membrane protein restriction. Low to medium ($50-$200) Classic method, simple to operate Limited throughput necessitates numerous repeated experiments. AP-MS / Co-IP-MS For known proteins or complexes Low to medium (10²–10³ level) The background is complex and requires strict comparison. High (approximately $1,000-$5,000) Strong physiological correlation, verifiable interaction Limited sensitivity, cumbersome operation, high price, and poor repeatability Alpha-Seq (A-Alpha Bio) Yeast pairing + NGS, pairwise batch processing High (10 6 -10 7 Tier / Month) Dependence on library matching quality Mid ($500-$1,000) Quantitative analysis is possible, suitable for molecular gels / drug sieves. It still belongs to the "pairwise" mode, not full interaction. "Book-to-Book" PPI Screening Library vs. Library (Comprehensive Coverage) extremely high (10 8 -10 9 level)]]> Low false positive rate across multiple systems Ultra-low (<$1) It can provide full coverage and quantitative analysis, saving time and costs. none The above specific embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.
[0053] The above description is only a preferred embodiment of the present invention. Therefore, all equivalent changes or modifications made to the structure, features and principles described in the claims of this patent application are included in the scope of this patent application.
[0054] The gene sequence involved in this invention is as follows: SEQ ID No.1; Name: Modified 3'RT random primer DNA sequence DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ctggatccaatggcatcttcaacacccgc SEQ ID No.2; Name: DNA sequence of the linker DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AAGCAGTGGTATCAACGCAGAGT SEQ ID No. 3; Name: Linker linker connection site sequence in plasmid DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACTCTGCGTTGATACCACTGCTT SEQ ID No.4; Name: DNA sequences of multiple pairs of IdeoxyU-modified adapter primers F DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct GGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 5; Name: DNA sequences of multiple pairs of IdeoxyU-modified adapter primers R DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct GTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 6; Name: DNA sequence of the full-length open reading frame of the human EEF1G gene DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct SEQ ID No.7; Name: DNA sequence of modified 3'RT Oligo-dT primers DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACTCTGCGTTGATACCACTGCTTTTTTTTTTTTTTTTTT SEQ ID No. 8; Name: DNA sequence of modified TSO primers 1 DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct TCGACTCTAGAGGATCCCrGrGrG SEQ ID No. 9; Name: DNA sequence of modified TSO primer 2 DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct TCGACTCTAGAGGATCCCrG SEQ ID No. 10; Name: DNA sequence 3 of the modified TSO primer DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct TCGACTCTAGAGGATCCCrGrG SEQ ID No. 11; Name: Deoxyuracil (dU) modified adapter primer AF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AGCTTACTTGTGAAGATGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 12; Name: Deoxyuracil (dU) modified adapter primer BF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACTTGTAAGCUGTCTAUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 13; Name: Deoxyuracil (dU) modified adapter primer CF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACTCTGUCAGGTCCGAUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 14; Name: Deoxyuridine (dU) modified adapter primer DF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACCTCCTCCUCCAGAAUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 15; Name: Deoxyuracil (dU) modified adapter primer EF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AACCGGACACACUTAGUGGAGTGTGGAGTGAGTGGATGAGTGATG SEQ ID No. 16; Name: Deoxyuracil (dU) modified adapter primer FF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AGAGTCCAAUTCGCAGUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 17; Name: Deoxyuracil (dU) modified linker primer GF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AATCAAGGCUTAACGGUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 18; Name: Deoxyuracil (dU) modified adapter primer HF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ATGTTGAAUCCTAGCGUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 19; Name: Deoxyuracil (dU) modified adapter primer IF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AGTGCGTUGCGAATTGUGGAGTGTGGAGTGAGTGGATGAGTGATG SEQ ID No. 20; Name: Deoxyuracil (dU) modified adapter primer JF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AATTGCGUAGTTGGCCUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 21; Name: Deoxyuracil (dU) modified adapter primer KF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACACTTGGUCGCAATCUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 22; Name: Deoxyuridine (dU) modified linker primer LF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AGTAAGCCUTCGTGTCUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 23; Name: Deoxyuracil (dU) modified adapter primer MF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACCTAGAUCAGAGCCTUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 24; Name: Deoxyuracil (dU) modified linker primer NF DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AGGTAUGCCGGUTAAGUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 25; Name: OF adapter primer modified with deoxyuracil (dU) DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AAGUCACCGGCACCUTUGGAGTTGGAGTGAGTGGATGAGTGATG SEQ ID No. 26; Name: Deoxyuracil (dU) modified adapter primer BR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ATAGACAGCUTACAAGUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 27; Name: Deoxyuridine (dU) modified adapter primer CR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ATCGGACCUGACAGAGUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 28; Name: Deoxyuracil (dU) modified adapter primer DR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ATTCUGGAGGAGGAGGUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 29; Name: Deoxyuridine (dU) modified adapter primer ER DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACTAAGTGUGTCCGGTUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 30; Name: Deoxyuracil (dU) modified adapter primer FR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACTGCGAAUTGGACTCUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 31; Name: Deoxyuracil (dU) modified adapter primer GR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACCGTUAAGCCTTGATUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 32; Name: Deoxyuridine (dU) modified adapter primer HR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACGCTAGGAUTCAACAUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 33; Name: Deoxyuracil (dU) Modified Linker Primer IR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACAATUCGCAACGCACUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 34; Name: Deoxyuracil (dU) Modified Proton Primer JR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AGGCCAACUACGCAATUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 35; Name: Deoxyuracil (dU) modified adapter primer KR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AGATUGCGACCAAGTGUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 36; Name: Deoxyuracil (dU) modified adapter primer LR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AGACACGAAGGCUTACUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 37; Name: Deoxyuracil (dU) modified adapter primer MR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AAGGCTCUGATCTAGGUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 38; Name: Deoxyuracil (dU) modified adapter primer NR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ACTUAACCGGCAUACCUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No. 39; Name: Deoxyuridine (dU) modified adapter primer OR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AAAGGUGCCGGUGACTUGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No.40; Name: Deoxyuracil (dU) Modified Proton Primer PR DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct ATCTCGAGCCACTTCATGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No.41; Name: P5 primer DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct AATGATACGGCGACCACCGAGATCTACACCTCTCTATTCGTCGGCAGCGTCAGATGTGTATAAGAGACAGGTGAGTGATGGTTGAGGATGTGTGGAGATA SEQ ID No.42; Name: P7 primer DNA type: other DNA Biological origin: Artificial Sequence / synthetic construct CAAGCAGAAGACGGCATACGAGATTAAGGCGAGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG.
Claims
1. A high-throughput pan-sample protein interaction screening method, characterized in that: 1) pan-sample library construction mRNA of a target sample is extracted, and a modified primer is used to obtain a pan-sample library; 2) pan-sample library homogenization treatment; 3) modification of a plasmid of a two-hybrid system to obtain a modified plasmid; 4) homologous recombination of the sample library and the modified plasmid of the two-hybrid system, and introduction into an organism to screen an organism with an interaction library and a single-cell labeled cDNA interaction library; 5) single-cell sequencing PPI pair identification of the organism with the interaction library, and construction of a pan-sample protein interaction network. The step 1) is specifically: mRNA of a target sample is extracted, and a modified 3'RT Oligo-dT primer and a TSO primer are used to obtain a full-length cDNA library of eukaryotes as a pan-sample library by using the Smart-seq 3 method, specifically including: first, the mRNA of eukaryotes is captured by Poly-A tail specificity using Oligo(dT) magnetic beads, and then a modified 3'RT random primer is used for reverse transcription to obtain a cDNA library with uneven length, and the sequence of the modified 3'RT random primer is SEQ ID No. 1, namely ACTCTGCGTTGATACCACTGCNNNNNNNN. The step 1) is specifically: for a prokaryotic polycistronic cDNA library, rRNA digestion is first performed using an rRNA digestion kit, a 5' end capping of mRNA is performed using a vaccinia virus capping system, and finally a 3' end Poly-A tail structure of mRNA is performed using an E. coli Poly-A polymerase and ATP, and the mRNA is obtained by using a modified 3'RT Oligo-dT primer and a TSO primer using the Smart-seq 3 method to obtain a full-length cDNA library of prokaryotes as a pan-sample library. The step 2) is specifically to use a Duplex-Specific Nuclease enzyme to homogenize the cDNA of the pan-sample library. The step 3) is specifically to modify the plasmid of a bacterial two-hybrid system or a yeast two-hybrid system, and link a linker containing a homologous arm to form a modified plasmid. The sequence of the linker is SEQ ID No. 2, namely AAGCAGTGGTATCAACGCAGAGT, and the connection site of the linker in the plasmid is the 3' end position of the MSC multiple cloning site of the vector, and the connection site sequence is SEQ ID No. 3, namely ACTCTGCGTTGATACCACTGCTT. 2. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: 3. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: 4. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: 5. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: 6. The high-throughput general sample protein interaction screening method according to claim 5, characterized in that: 7. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: The step 4) is specifically: first, the target sample cDNA library is respectively homologously recombined with the plasmid modified by the two-hybrid system to form a bait protein and prey protein plasmid library, then the plasmid library is transformed into the corresponding organism by electroporation or chemical transformation, and positive PPI combination clone screening is carried out in the screening medium, so as to obtain an organism with protein interaction pairs and induce gene expression by IPTG, and obtain a single-cell labeled cDNA interaction library.
8. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: The modified plasmid contains a linker sequence, and the positive bacteria BACTH two-hybrid system is blue under the screening of the screening medium. The bait protein library is composed of the modified plasmid pKT25 or pKNT25 of the two-hybrid system and the cDNA library homologous recombination, and the prey protein is composed of the modified plasmid pUT18C or pUT18 of the two-hybrid system and the cDNA library homologous recombination.
9. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: The step 5) adopts single-cell transcriptome sequencing based on random primers, second-generation sequencing or third-generation sequencing based on droplet PCR direct plasmid copy amplification.
10. The high-throughput general sample protein interaction screening method according to claim 9, characterized in that: The single-cell transcriptome sequencing based on random primers in the step 5) is specifically: First, the screened bacteria or yeast organisms are fixed by 4% PFA solution for 8-16 h to crosslink proteins and nucleic acids, lysozyme enzyme and Zymolyase enzyme are added to open the cell wall, then 0.4% triton solution is used for cell membrane permeabilization treatment, then in-situ reverse transcription reaction and TdT terminal transferase cDNA 3' hydroxyl end Poly(A) are carried out in sequence, then the bacteria or yeast are sequentially encapsulated and cDNA double-strand amplified by adding water gel coding microspheres of oligo(dT), to obtain a cDNA library with cell barcode markers, the cDNA library is again end-repaired and connected with P5 / P7 adapters to construct an Illumina sequencing library using the VAHTS Universal Pro DNA Library Prep Kit for Illumina library preparation kit, and finally the obtained Illumina sequencing library is sequenced, and the obtained sequences are compared with the reference genome of the target sample species to identify interaction protein pairs and cDNA combination pairs with the same cell barcode, and to construct an interaction protein network.
11. The high-throughput generic sample protein interaction screening method of claim 9, characterized in that: The second generation sequencing based on plasmid copy amplification by droplet PCR in step 5) is as follows: the microspheres of modified hydrogel containing linker sequences are directly added into the high-temperature-resistant droplet generation oil and water phase stabilizer, and then the positive PPI bacteria or yeast screened are encapsulated by single cell for PCR amplification, after amplification, the transposase DNA library construction kit is used for fragmenting and breaking the amplification product, then the P5 / P7 adapter primer containing cell barcode is used for cDNA sequencing library construction containing cell barcode sequence, and finally the obtained cDNA sequencing library is sequenced, and the sequences obtained by sequencing and the reference genome of the target sample species are compared to identify the interaction protein pair, identify the cDNA combination pair with the same cell barcode, and construct the interaction protein network.
12. The high-throughput generic sample protein interaction screening method of claim 11, wherein: For yeast organisms, a mixture containing a lysis solution is also added during PCR amplification, and the mixture containing the lysis solution is: 2 mg / g of cells Zymolase lysis enzyme is reacted at 37°C for 30 min to lyse the cell wall.
13. The high-throughput generic sample protein interaction screening method of claim 9, wherein: In step 5), the third generation sequencing based on plasmid copy amplification by droplet PCR is as follows: first, the cDNA interaction library obtained in step 4) is labeled by single cell, and the linker primer modified with deoxyuracil (dU) is used for secondary amplification, then the obtained library is mixed together and purified by magnetic beads, after magnetic bead purification, a single nucleotide gap is generated at the uracil position by USER enzyme, finally, Endo VIII enzyme is used to cleave a single strand overhang to separate from double-stranded, and the remaining double-stranded with sticky ends is reserved, then the double-stranded with sticky ends is annealed, and then Taq high-fidelity DNA ligase is used for ligation to form a long fragment, finally, Nanopore and PacBio platforms are used for third generation sequencing.
14. The high-throughput generic sample protein interaction screening method of claim 9, wherein: In the step 5), the third-generation sequencing based on droplet PCR directly amplifying plasmid copies is as follows: first, the single-cell labeled cDNA interaction library obtained in the step 4) is treated by using a TdTase terminal transferase to add two sequences in a dATP-dTTP pair or a dCTP-dGTP pair to the 3' hydroxyl end of the DNA molecule respectively, the product after adding two sequences is subjected to magnetic bead purification, after the magnetic bead purification, the two products are mixed, treated at 95°C for 1 min, and then slowly reduced to 25°C at a rate of 0.1°C / s for annealing, and finally, T4 DNA polymerase is added for DNA repair and ligation, and T4 DNA ligase is directly connected into a long fragment; finally, Nanopore and PacBio platforms are used for third-generation sequencing.
Citation Information
Patent Citations
Library versus library yeast two-hybrid massive interaction protein screening method
CN103774241A
Recombinant vector used for high-flux yeast two-hybrid technology and method for large-scale screening interacting protein thereof
CN106434734A
Method for screening mouse spermatogenesis-related interacting protein and application of mouse spermatogenesis-related interacting protein
CN116084026A
Methods for protein interaction determination
US20050164214A1