A high-throughput general sample protein interaction screening method

By employing a generalized open reading frame and multi-system screening approach, combined with single-cell sequencing and droplet PCR technology, the throughput and coverage issues of traditional protein interaction screening methods have been resolved. This approach enables high-throughput, full-coverage, and low-cost protein interaction screening, applicable to the identification of PPI networks in various species.

CN121428075BActive Publication Date: 2026-05-15LIANGZHU LAB
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LIANGZHU LAB
Filing Date
2025-12-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional protein interaction screening methods cannot achieve high-throughput, comprehensive library-to-library screening, and suffer from low throughput and limited scope. Furthermore, deep learning models do not perform well in predicting new protein interactions and multi-protein collaborations.

Method used

We adopted a generalized open reading frame construction method and combined it with a multi-system, high-accuracy library-to-library PPI screening system. We achieved high-throughput PPI screening through single-cell transcriptome sequencing (scPPI-seq), second-generation sequencing (dPPI-seq) and third-generation sequencing (dTPPI-seq) for direct plasmid copy amplification by droplet PCR.

Benefits of technology

It achieves high-throughput, full-coverage, and low-cost protein interaction screening, breaking through the bottleneck of traditional screening. It is applicable to the identification of PPI networks among different species and features broad sample coverage, high throughput, and high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121428075B_ABST
    Figure CN121428075B_ABST
Patent Text Reader

Abstract

The application discloses a high-throughput general sample protein interaction screening method. The method comprises the following steps: firstly, constructing a general sample open reading frame library, extracting mRNA of a target sample, and then obtaining a cDNA library by using a modified random primer or an oligo-dT primer; then, uniformly processing the general sample cDNA library; modifying plasmids of a double-hybrid system to obtain modified plasmids; homologously recombining the general sample library and the modified plasmids of the double-hybrid system, and introducing them into organisms to screen and obtain organisms with an interaction library; identifying single-cell PPI pairs of the screened organisms, and constructing a general sample protein interaction network. The application is suitable for screening and identifying the same Barcode cDNA combination pairs by using single-cell sequencing after obtaining tens of thousands of positive PPI combination clones by means of a "library-library" bacterial or yeast double-hybrid system, and realizing PPI network identification in different intra-species, inter-species and hybridization technology systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of protein processing and screening, and in particular to a high-throughput general sample protein interaction screening method. Background Technology

[0002] Proteins, as the main executors of cellular functions, often participate in various intracellular life activities through interactions, such as signal transduction, metabolic regulation, transcription and translation, and cell cycle control. Therefore, protein-protein interactions (PPIs) within the host cell form the basis of the cellular functional network and play a crucial role in maintaining normal physiological functions and preventing the occurrence of pathological processes.

[0003] Meanwhile, an increasing number of studies are focusing on host-microbe protein-protein interactions (HM-PPIs). Human hosts coexist with a large number of symbiotic microorganisms (such as the gut microbiota) for extended periods. These microorganisms can directly or indirectly interact with host proteins through their secretions or surface proteins, thereby regulating immune, metabolic, neural, and barrier functions. Pathogens, on the other hand, can use their effector proteins to interfere with host signaling pathways via PPIs to evade immunity, promote colonization, or spread infection. For example, pathogens such as Helicobacter pylori, Salmonella, and Escherichia coli have all been shown to manipulate host cell functions through a series of sophisticated protein-protein interaction mechanisms.

[0004] Traditional protein interaction screening methods (such as Y2H and AP-MS) can only perform single-pair interaction screening, failing to achieve full coverage at the library-to-library level, and suffer from low throughput and limited scope. Although deep learning-based large language models (such as AlphaFold) have been widely used in PPI prediction in recent years, they still have problems such as limited binding affinity datasets, oversimplification of task difficulty, inability to handle "new" proteins, and inability to handle multi-protein collaborations. Therefore, establishing a high-throughput, general-sample library-to-library protein interaction screening method is of great significance for helping research institutions and enterprises improve R&D efficiency, accelerate the development of new drugs and materials, assist in the development of anti-aging related antibodies, molecular glues, cell therapies, vaccines, and diagnostics, and facilitate the development of new products in fields such as sustainable chemical synthesis and crop protection. Summary of the Invention

[0005] To address the problems existing in the background technology mentioned above, the present invention provides a high-throughput screening system for "library-to-library" protein-protein interactions and a method for using it. It mainly includes a general sample open reading frame acquisition system, a multi-system high-accuracy "library-to-library" PPI screening system, and single-cell transcriptome sequencing (scPPI-seq) after induction screening, as well as second-generation sequencing (dPPI-seq) based on droplet PCR for direct plasmid copy amplification and third-generation dTPPI-seq.

[0006] like Figure 1 As shown, the technical solution of the present invention is as follows:

[0007] 1) Construction of generic sample libraries

[0008] mRNA was extracted from the target sample in vitro, and then a pan-sample cDNA library was obtained using modified primers.

[0009] 2) Homogenization of pansample cDNA libraries;

[0010] 3) Plasmid modification: Modify the plasmid of the dual-hybrid system to obtain the modified plasmid;

[0011] 4) Homologous recombination was performed on the sample library and the modified plasmid of the double hybrid system, and the plasmid was introduced into the double hybrid system organisms to screen for organisms with mutual interaction libraries (positive PPI clones) and single-cell labeled cDNA mutual interaction libraries.

[0012] 5) After obtaining positive PPI clones, single-cell sequencing high-throughput PPI pair identification was performed on organisms with interaction libraries using three methods to construct a pan-sample protein interaction network.

[0013] Step 1) specifically refers to one of the following two processing procedures:

[0014] The first method involves first extracting mRNA from the target sample, then using modified 3'RT Oligo-dT primers and TSO primers to obtain a full-length eukaryotic cDNA library as a pan-sample cDNA library using Smart-seq 3. This includes:

[0015] Some dual-hybrid systems fuse the library to the N-terminus of a plasmid vector. When obtaining the library, Oligo(dT) magnetic beads are first used to capture eukaryotic mRNA specifically via the Poly-A tail. Then, a modified 3'RT random primer is used for reverse transcription to obtain a cDNA library with fragments of uneven length. The modified 3'RT random primer sequence is SEQ ID No.1, i.e., ACTCTGCGTTGATACCACTGCNNNNNNNN. This avoids "false negatives" caused by fusion sites and improves the full coverage and sensitivity of interaction detection.

[0016] The modified 3'RT Oligo-dT primer sequence is SEQ ID No. 7, namely ACTCTGCGTTGATACCACTGCTTTTTTTTTTTTTTTTTTTTT, and the modified TSO primer sequences include SEQ ID Nos. 8-10, namely TCGACTCTAGAGGATCCCrGrGrG, TCGACTCTAGAGGATCCCrG, and TCGACTCTAGAGGATCCCrGrG.

[0017] The above sequences were optimized based on the sequences of the plasmid vector. The 3'RT primer sequence was added to the vector through plasmid modification and also corresponds to the sequence on the hydrogel-encoded microspheres. The TSO primers were designed based on the 5' homologous arm sequence of the vector cloning site. The TSO primer sequence was changed according to the homologous arm for different vectors, while the 3'RT primer sequence remained unchanged.

[0018] The second method involves using a polycistronic cDNA library from prokaryotes. First, rRNA is digested using an rRNA digestion kit. Then, the 5' end of the mRNA is capped using the Faustovirus Capping Enzyme, Cap 2'-O-methyltransferase, and 3'biotin-GTP. Finally, a Poly-A tail is added to the 3' end of the mRNA using E. coli Poly-A polymerase and ATP. The purified mRNA is then purified using mRNA purification beads. A full-length cDNA library is constructed using the Smart-seq 3 protocol, specifically by using modified 3'RT Oligo-dT primers and TSO primers with Smart-seq 3 to obtain the full-length cDNA library from prokaryotes, which serves as a pan-sample cDNA library.

[0019] Step 2) specifically involves using Duplex-Specific Nuclease (DSN) to homogenize the cDNA in the pan-sample cDNA library, reducing the abundance of high-copy genes while retaining low-expression DNA molecules, thus making the gene concentration in the cDNA sample more uniform, improving screening efficiency and discovering interaction networks of extremely low-abundance genes.

[0020] Specifically, step 3) involves modifying the plasmids of bacterial or yeast double-hybrid systems for single-cell labeling by linking a linker sequence containing a homologous arm to form a modified plasmid.

[0021] The linker sequence is SEQ ID No. 2, namely AAGCGTGGTATCAACGCAGAGT, which has a homologous sequence with the modified 3'RT Oligo-dT primer. The linker linker in the plasmid is located at the 3' end of the MSC multiple cloning site of the vector, and the linker site sequence is SEQ ID No. 3, namely ACTCTGCGTTGATACCACTGCTT.

[0022] The bacterial dual-hybrid systems mentioned include pMRBAD-Z-CGFP / pRT11a-Z-NGFP, pBT / pTRG, pKT25(pKNT25) / pUT18C(pUT18), pET11a-link-NGFP / pMRBAD-link-CGFP, pTET-GFP11 / pET-GFP1-10, pSLBLC-c-myc-hFADD / pSBLNL-hFasDD-His, pQE / pREP, and pACYC184 / pGSTDHFR. The yeast dual-hybrid system mentioned includes the GAL4 system. This scheme is compatible with all dual-plasmid bacterial dual-hybrid and yeast dual-hybrid screening schemes.

[0023] Step 4) specifically involves: firstly, homologous recombination of the target sample cDNA library with the modified plasmid of the dual-hybrid system to form their respective bait protein and prey protein-particle libraries; then, the plasmid library is transferred to the corresponding dual-hybrid system organism (bacteria or yeast) via electrotransduction (high efficiency) or chemical transduction, and positive PPI combination clones are screened in the selection medium to obtain organisms with protein-protein interaction pairs. Gene expression is induced by IPTG to obtain a single-cell labeled cDNA interaction library.

[0024] The screening medium is such as M63 medium (5xM63 medium formula: weigh 10 g of (NH4)2SO2, 68 g of KH2PO4, 2.5 mg of FeSO4·7H2O and 5 mg of vitamin B1, add deionized water to a final volume of 1 L, adjust the pH to 7.0 with KOH and then autoclave).

[0025] The modified plasmid contains the linker sequence AAGCGTGGTATCAACGCAGAGT, and the positive result is that the bacterial BACTH double hybrid system turns blue under the selection medium M63 medium.

[0026] The bait protein library is composed of a modified plasmid pKT25 or pKNT25 from a dual-hybrid system and a cDNA library through homologous recombination. The prey protein is composed of a modified plasmid pUT18C or pUT18 from a dual-hybrid system and a cDNA library through homologous recombination.

[0027] Step 5) employs single-cell transcriptome sequencing based on random primers (scPPI-seq), second-generation sequencing based on droplet PCR for direct plasmid copy amplification (dPPI-seq), or third-generation sequencing (dTPPI-seq).

[0028] Step 5) specifically involves single-cell transcriptome sequencing (scPPI-seq) based on random primers:

[0029] First, the screened bacterial or yeast organisms were fixed in 4% PFA solution for 8-16 hours to allow protein-nucleic acid cross-linking. Lysozyme and Zymolyase enzymes were then added to open the cell walls of the bacteria and yeast, respectively. Next, the cell membrane was permeabilized with 0.4% Triton solution. Then, in situ reverse transcription and TdT terminal transferase were performed to add Poly(A) to the 3' hydroxyl end of cDNA. Oligo(dT) hydrogel-encoded microspheres were then added to the bacteria or yeast for single-cell encapsulation and cDNA double-strand amplification, resulting in a cDNA library with cell barcode markers. The cDNA library was then used to repair ends and ligate P5 / P7 adapters to construct an Illumina sequencing library using the VAHTS Universal Pro DNA Library Prep Kit for Illumina. Finally, the obtained Illumina sequencing library was sequenced at paired ends (150 bp) using an Illumina X plus 25B sequencer, for a total of 300 bp. bp sequencing compares the sequenced data with the reference genome of the target sample species to identify interacting protein pairs, recognize cDNA combinations with the same cell barcode, and construct an interacting protein network.

[0030] The high-temperature resistant droplet generation oil is specifically QX200™ Droplet Generation Oil for EvaGreen (#1864005), the aqueous phase stabilizer is specifically a digital PCR-specific aqueous phase stabilizer purchased from Nanjing Stone Gene Technology Co., Ltd., and the fragmentation reagent is specifically UltraClean DNA Library Prep Kit V3 for Illumina (UTD521).

[0031] The second-generation sequencing based on droplet PCR direct plasmid copy amplification (dPPI-seq) in step 5) specifically involves:

[0032] Since the transfected plasmid contains more than 5 copies, modified hydrogel-encoded microspheres containing linker sequences were directly added to the screened positive PPI bacteria or yeast for single-cell encapsulation along with heat-resistant droplet-generating oil and aqueous phase stabilizer, followed by PCR amplification. After amplification, the amplified products were fragmented using a transposase DNA library construction kit (Novi Station UTD521). Then, a cDNA sequencing library containing cell barcode sequences was constructed using P5 / P7 adapter primers containing cell barcode linkers. Finally, the obtained cDNA sequencing library was sequenced at both ends of 150 bp using an Illumina X plus 25B sequencer, for a total of 300 bp. The sequenced sequences were compared with the reference genome of the target sample species to identify interacting protein pairs and cDNA combinations with the same cell barcode, thus constructing an interacting protein network.

[0033] The P5 primer sequence is: SEQ ID No. 41, AATGATACGGCGACCACCGAGATCTACACCTCTCTATTCGTCGGCAGCGTCAGATGTGTATAAGAGACAGGTGAGTGATGGTTGAGGATGTGTGGAGATA;

[0034] The P7 primer sequence is: SEQ ID No. 42, CAAGCAGAAGACGGCATACGAGATTAAGGCGAGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG.

[0035] The single-stranded DNA on the hydrogel-encoded microspheres containing the linker sequence consists of an upstream primer complement fragment, a barcode, a unique multiplex index (UMI), and a linker sequence. The upstream primer complement fragment binds to the upstream primer during PCR amplification. The barcode is used to label cDNA within the same cell septum. The UMI is a random sequence used to label each original cDNA. The linker sequence acts as a primer in PCR to complete single-cell labeling amplification.

[0036] In the above process, for yeast organisms, a mixture containing lysis buffer (mix) was added during PCR amplification. The mixture containing lysis buffer (mix) consisted of 2 mg / g cell Zymolase lysing enzyme reacting at 37°C for 30 min to lyse the cell wall.

[0037] For bacterial organisms, no mixture containing lysis buffer is added during PCR amplification.

[0038] In step 5), the third-generation sequencing based on droplet PCR for direct plasmid copy amplification (dTPPI-seq) employs one of the following methods:

[0039] Since Illumina X plus 25B paired-end 150 bp sequencing can only identify PPI pairs, but cannot determine whether the identified protein interactions are false positives caused by frameshift mutations, performing third-generation full-length sequencing is the key to solving the problem. However, the full-length cDNA library is 300-3000 bp. Directly performing third-generation sequencing on PicBio or ONT will result in a large waste of chips and low data output due to the short sequence. To adapt to long-sequence sequencing on PicBio and ONT platforms,

[0040] A) The first type

[0041] First, for the cDNA intergenic library obtained in step 4) with single-cell labeling, 15 pairs of adapter primers modified with deoxyuridine (dU) were used for secondary amplification. After amplification, the obtained libraries were mixed together and purified with magnetic beads. After purification, a single nucleotide gap was generated at the uridine position using USER enzyme. Finally, Endo VIII enzyme was used to cleave the double-stranded DNA to produce a single-stranded overhang that was separated from the double-stranded DNA and discarded, leaving the double-stranded DNA with sticky ends. The double-stranded DNA with sticky ends was then annealed and ligated with Taq high-fidelity DNA ligase (NEB M0647S) to form a long fragment of more than 10kb, which made the original 300 bp sequencing at least 15 times longer, reaching a DNA sequence length of 10k. Finally, third-generation sequencing was performed using Nanopore and PacBio platforms to improve the utilization rate of sequencing wells and sequencing efficiency of sequencing chips.

[0042] The multiple pairs of IdeoxyU-modified adapter primer sequences are SEQ ID No. 4 and SEQ ID No. 5, respectively: F: GGAGTTGGAGTGAGTGGATGAGTGATG, R: GTGAGTGATGGTTGAGGATGTGTGGAGATA.

[0043] B) The second type

[0044] First, the single-cell labeled cDNA intergenic library obtained in step 4) is processed using TdTase terminal transferase to add two sequences from either the dATP-dTTP pair or the dCTP-dGTP pair to the 3' hydroxyl end of the DNA molecules. The products with the two sequences added are then purified using magnetic beads. After purification, the two products are mixed and treated at 95°C for 1 min, followed by slow annealing at 0.1°C / s to 25°C. Finally, T4 DNA polymerase is added for DNA repair and ligation. T4 DNA ligase can directly tandemly generate long fragments of over 10 kb, facilitating third-generation sequencing. Finally, third-generation sequencing is performed using Nanopore and PacBio platforms to improve the utilization rate of sequencing wells and sequencing efficiency.

[0045] The high throughput mentioned in this invention refers to a high-throughput identification technology system that can identify tens of thousands to hundreds of thousands of positive PPI clones in a single operation.

[0046] The beneficial effects of this invention are:

[0047] This invention constructs a general sample open reading frame acquisition system, a multi-system high-accuracy "library-to-library" PPI screening system, and a multi-mode droplet microfluidic high-throughput sequencing identification system. It is suitable for high-throughput screening after obtaining tens of thousands of positive PPI combination clones through "library-to-library" bacterial or yeast double hybridization. It can identify PPI network identification in a wide range of scenarios, including intra-species, inter-species (model species vs. non-model species, bacteria vs. host), and hybridization technology systems (bacterial two-hybrid, yeast one-hybrid). It breaks through the bottleneck of traditional screening that can only perform "one-to-one" or a few "one-to-many" screenings, and has the advantages of general sample, high throughput, high accuracy, low cost, and full coverage of protein interaction identification. Attached Figure Description

[0048] Figure 1 This is a schematic diagram of the present invention; in the diagram, mRNA: Messenger ribonucleic acid; cDNA: Complementary DNA; CB: Cell barcode; PPI: Protein-protein interaction networks;

[0049] Figure 2 This is a schematic diagram of plasmid modification and homologous recombination;

[0050] Figure 3 A graph of the full-length open reading frame for mice and humans is constructed; in the graph, k: Kilo; bp: Base pair;

[0051] Figure 4This is a screening diagram for positive clones; in the diagram, M: Marker; bp: Base pair;

[0052] Figure 5 This is a screening plot of LA plates for the library; in the plot, Amp: Ampicillin; Kan: Kanamycin; IPTG: Isopropyl β-D-1-thiogalactopyranoside; X-gal: 5-bromo-4-chloro-3-indolyl β-D-galactopyranoside;

[0053] Figure 6 This is a diagram showing library induction and screening in M63 medium; in the diagram, DsRed: Red fluorescent proteins;

[0054] Figure 7 This is a single-cell sequencing diagram;

[0055] Figure 8 Diagram of the microsphere and library structure encoded by dPPI hydrogel;

[0056] Figure 9 Contamination rate assessment graph for single-cell high-throughput screening; in the graph, hs: human species; mm: Musmusculus;

[0057] Figure 10 This is a diagram of the construction of a full-length cDNA library using third-generation sequencing; in the diagram, RFU: Relative fluorescence unit; bp: Base pair;

[0058] Figure 11 This is a PCR clone diagram of the mouse candidate protein cDNA; in the diagram, M: Marker; bp: Base pair;

[0059] Figure 12 The image shows the screening results of candidate proteins for yeast double hybridization. Detailed Implementation

[0060] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] The embodiments of the present invention are as follows: Example

[0062] 1. Obtain the target library by direct TRIzol fragmentation.

[0063] RNA from the target sample was obtained by direct lysis according to the TRIzol instructions. For eukaryotic samples, full-length open reading frames were obtained directly using the modified Smart-seq 3 method. For prokaryotic samples, to ensure the efficiency of 5' capping and 3' poly-A tailing, the rRNA was digested using the NEBNext® rRNA Depletion kit before reverse transcription.

[0064] Next, the 5' capping of mRNA was performed using the Faustovirus Capping Enzyme, Cap 2´-O-methyltransferase, and 3'biotin-GTP. A 3' Poly-A tail structure was added to the mRNA using E. coli Poly-A polymerase and ATP. A full-length cDNA library was constructed from the magnetically purified mRNA using the Smart-seq 3 protocol. The specific method is as follows: RT reaction is performed using an RNaseH-free MMLV RT enzyme (such as Maxima H-minus reversetranscriptase enzyme (Thermo Scientific)). The reaction system contains 25 mM Tris-HCl pH 8.0-8.4, 30 mM NaCl, 2.5 mM MgCl2, 1 mM GTP, 8 mM DTT, 0.25 U RNase inhibitor, 0.3 mM dNTPs, 0.1 uM TSO primer and 0.1 uM 3' RT primer, and RNA. RNA, dNTPs and 3' RT primer are denatured at 72°C for 5 min beforehand. Then, the system is immediately placed on ice and the remaining reagents are added. The RT reaction is performed at 42°C for 90 min, followed by 10 cycles of 50°C for 2 min and 42°C for 2 min. After the RT reaction, cDNA library was amplified for 25 cycles using 2X KAPA HiFi HotStart ReadyMix. After amplification, the cDNA library was purified using 0.6x VAHTS DNA Clean Beads (N411).

[0065] 2. To ensure the cDNA in the library is free of repetitive sequences and covers more transcript information, we used Duplex-Specific Nuclease (DSN) for cDNA homogenization. This was performed using the Puente Bio cDNA homogenization kit (FZ1031). First, 200 ng of the cDNA library was mixed with 4x Hybridization buffer and denatured at 98°C for 2 min, then incubated at 68°C for 5 h. Then, 1 μL of DSN and 1 μL of 10×DSN Reaction buffer were added, gently mixed, briefly centrifuged, and incubated at 68°C for 7-20 min to degrade the double-stranded cDNA. Finally, 10 μL of 2×DSN Stop Buffer was added, gently mixed, briefly centrifuged, and incubated at 65°C for 5 min to terminate the reaction. The remaining homogenized single-stranded cDNA was amplified by PCR for library cloning. Figure 3 ).

[0066] 3. High-throughput screening and sequencing solutions

[0067] Currently, there are many low-throughput yeast and bacterial two-hybrid systems based on one-to-one protein interaction. This invention conducts library-to-library high-throughput single-cell screening based on traditional protein interaction bacterial and yeast two-hybrid systems. To evaluate the cross-contamination of single-cell sequencing technology, the BACTH system is used as an example.

[0068] Because the library-to-library bacterial double hybridization process makes it difficult to directly determine whether the recombinant plasmid is correctly expressed, the DsRed gene expressing red fluorescent protein, a human cDNA library, and a mouse cDNA library were homologously recombined into the bait expression vector pUT18C and electroporated into *E. coli* to extract plasmids. The mouse cDNA library was homologously recombined into the bait expression vector pKNT25 and electroporated into BTH101 cells to prepare competent cells. The DsRed-pUT18C, human cDNA library-pUT18C, and mouse cDNA library-pUT18C plasmids were then electroporated again into mouse cDNA library-pKNT25 BTH101 *E. coli* competent cells and plated on M63 LA plates containing X-gal+IPTG+Amp+Kan for positive clone screening. Figures 2 - 5 All positive PPI clones were collected and expanded in LB medium containing double antibiotics, followed by IPTG-induced expression. Red fluorescent protein was used as an indicator signal for gene expression. Figure 6 When the bacterial culture turns red, bacteria are collected simultaneously. The high throughput of the single-cell system and the low contamination rate are evaluated using human-mouse library double-hybrid Escherichia coli.

[0069] 3.1 Identification of protein interactions using random primer smRandom-seq (scPPI-seq, single-cell high-throughput protein interaction scPPI-seq)

[0070] The collected bacteria were fixed overnight in 4% paraformaldehyde to cross-link the intracellular RNA, DNA, and proteins. Lysozyme was used to digest the cell wall, permeabilizing the fixed bacteria to facilitate the next step of in situ reverse transcription. The microorganisms were used as reaction vessels for the in situ reaction. Random primers were added to bind to the intracellular RNA, and total RNA was captured and synthesized into cDNA. A poly-A tail was added in situ to the 3' end of the cDNA using a terminal transferase (TdT). Each step of the process was followed by washing with buffer 3-8 times to prevent reagent residue from affecting subsequent reactions. Individual bacteria and labeled microbeads were encapsulated into droplets using a microfluidic device. Poly-T primers were released from the microbeads by enzymatic digestion, digesting the bacterial RNA to release cDNA from the bacteria. The poly-T primers bound to the poly-A tail at the end of the cDNA, subsequently extending the cDNA to add specific coding, and a molecular tag (UMI) was added to each cDNA. After demulsification, cDNA was collected and purified, expanded, and sequencing adapters were added to construct a sequencing library. The cDNA product of rRNA was digested with Cas9, and the cDNA product of mRNA was enriched for high-throughput sequencing. The obtained sequencing data were subjected to quality control and filtering before alignment to the Ecoli_bw25113 genome. Figure 7 Unmapping reads were extracted and aligned to human and mouse reference genes, respectively. Feature Counts software was used to count the number of mouse and human transcripts detected in the same barcode and their corresponding UMIs. A UMI threshold was designed to exclude false positives. The higher the UMI of mouse and human transcripts detected by a barcode, the fewer false positives generated by sequencing and the higher the reliability of the interaction pairs. Finally, an interaction network was constructed from the selected interaction pairs.

[0071] Comparison of valid Top 2000 barcodes revealed that 80.3% (1606) of the barcodes identified two or more genes, 84.4% (1688) of the barcodes contained the DsRed gene, and 75.9% (1281) of the barcodes containing two or more genes had the DsRed gene UMI number in the Top 2. A single-cell sequencing PPI identification method based on random primer droplet microfluidic “DsRed library” was successfully established (Table 1).

[0072] Table 1. Screening results of the Dsred-mouse library using dPPI-seq

[0073] Barcode Dsred-UMI Mouse-gene-UMI Mouse-gene-UMI Barcode1 Dsred-2 Dync1h1-1478 Gm17111-2 Barcode2 Dsred-11 mt-Cytb-804 Barcode3 Dsred-48 mt-Cytb-43 Alb-7 Barcode4 Dsred-81 mt-Cytb-325 Alb-1 Barcode5 Dsred-1 Dync1h1-392 Barcode6 Dsred-47 mt-Cytb-47 mt-Co3-4 Barcode7 Dsred-1 Dync1h1-347 Barcode8 Dsred-8 Alb-298 mt-Cytb-11 Barcode9 Dsred-38 mt-Cytb-18 Alb-6 Barcode10 Dsred-1 mt-Cytb-39 Barcode11 Dsred-1 mt-Cytb-320 Barcode12 Dsred-144 Eef1a1-161 Slc22a27-3 Barcode13 Dsred-1 mt-Co1-300 mt-Cytb-15 Barcode14 Dsred-1 Eef1a1-294 Slc22a27-10 Barcode15 Dsred-2 mt-Cytb-303 mt-Co3-1

[0074] 3.2 Protein interaction identification based on digital PCR system (dPPI-seq)

[0075] Referring to the smRandom-seq article, 96*96*48 cell barcode sequences were synthesized. Based on vector characteristics, the last adapter sequence was modified to AAGCGTGGTATCAACGCAGAGT, SEQ ID No. 2. The first cell barcode segment was amino-modified at the 5' end and immobilized on polyacrylamide hydrogel microspheres. Different cell barcodes were ligated using T7 ligase according to the split pool method. After completion, NaOH treatment was used to form dsDNA single strands. Figure 8 ).

[0076] Using microfluidic devices, hydrogel-encoded microspheres and individual bacteria were directly encapsulated in water-in-oil droplets. To prevent droplet fusion at high temperatures, 10% aqueous phase stabilizer (purchased from Hangzhou Chuangyan Xinghe) was added to the aqueous phase, along with 25% OptiPrep density gradient medium, 2X KAPA HiFi HotStart ReadyMix, 400 E. coli cells / ul selected with double antibodies, 2.5ul PUT18C and 2.5ul PKNT25 vector primers, and 5ul USER enzyme (M5505L). The incubation process was 37℃ for 30 min, 95℃ for 3 min, followed by 30 cycles of 95℃ for 30 s, 60℃ for 20 s, 72℃ for 2 min 30 s, and a final incubation at 72℃ for 5 min, and storage at 4℃. After PCR amplification, the cDNA library was purified using 0.6x VAHTS DNA Clean Beads (N411). After amplification for three cycles using barcode adapter primers and PUT18C and PKNT25 vector primers, Illumina libraries were constructed using the VAHTSUniversal Pro DNA Library Prep Kit for Illumina.

[0077] Comparison of selected effective cells revealed that 81% of cells identified as having barcodes from one mouse and one human, 10% had empty vectors, and 9% contained cells with two or more genes (contamination). The data are reliable, and a protein interaction identification method based on a digital PCR system has been successfully established. Figure 9 ).

[0078] 3.3 Direct plasmid copy amplification using droplet PCR based on third-generation sequencing (dTPPI-seq)

[0079] Since Illumina X plus 25B paired-end 150 bp sequencing can only identify PPI pairs but cannot determine whether there are frameshift mutations in the library during PCR amplification, performing third-generation full-length sequencing is the key to solving the problem. However, full-length cDNA libraries are 300-3000 bp. Directly performing third-generation sequencing on PicBio or ONT would result in a large amount of chip usage due to the short sequence length. Therefore, long-sequence sequencing compatible with PicBio and ONT platforms is necessary.

[0080] 3.3.1 After pre-amplification of cDNA, cDNA libraries were amplified again using adapter primers modified with deoxyuridine (dU). The adapter sequences are shown in Table 2, as shown in SEQ ID No. 20. All libraries were mixed together and purified. After purification, a single nucleotide gap was generated at the uridine position using USER enzyme and a single-stranded overhang was generated by EndoVIII cleavage. Then, sticky end annealing and ligase were used to ligate the overhang to form a long fragment (Sequence 1).

[0081] ONT sequencing yielded the full-length open reading frame of the human EEF1G gene, as shown in SEQ ID No. 6, specifically below, with the underlined and bolded sequence being the adapter sequence:

[0082] GGTCGACTCTAGAGGATCCCGACTC TGCGTTGATACCACTGCTT ACTCTGCTTGATACCACTGCTACTCTGCGTTGATACCACTGCTACTCTGCGTTGATACCACTGCTTAGACGAGATTGTATCTTCTCTCCGTATTACCGACCAAGAGCCGTGT TATCTCCACACATCCTCAACC ATCACTCAC .

[0083] 3.3.2 After cDNA amplification, dATP-dTTP or dCTP-dGTP pairs were added to the 3' hydroxyl end of the DNA molecule using TdTase terminal transferase. After purification, the samples were mixed and annealed, and then subjected to DNA repair and ligation with T4 DNA polymerase to tandemly form long fragments exceeding 10 kb. ​ ).

[0084] Table 2 15 pairs of linker primers modified with deoxyuracil (dU)

[0085] ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ATGTTGAAUCCTAGCGUGGAGTTGGAGTGATGGATGAGTGAT IF AGTGCGTUGCGAATTGUGGAGTTGGAGTGAGTGGATGAGTGATGATG JF AATTGCGUAGTTGGCCUGGAGTTGGAGTGATGAGTGATGAGTGATG KF ACACTTGGUCGCAATCUGGAGTTGGAGTGAGTGGATGAGTGATG LF AGTAAGCCUTCGTGTCUGGAGTTGGAGTGAGTGGATGAGTGATG MF ACCTAGAUCAGAGCCTUGGAGTTGGAGTGAGTGGATGAGTGATG NF AGGTAUGCCGGUTAAGUGGAGTTGGAGTGAGTGGATGAGTGATGATG OF AAGUCACCGGCACCUTUGGAGTTGGAGTGAGTGGATGAGTGATG BR ATAGACAGCUTACAAGUGTGAGTGATGGTTGAGGATGTGTGGAGATA CR ATCGGACCUGACAGAGUGTGAGTGATGGTTGAGGATGTGTGGAGATA DR ATTCUGGAGGAGGAGGUGTGAGTGATGGTTGAGGATGTGTGGAGATA IS ACTAAGTGUGTCCGGTUGTGAGTGATGGTTGAGGATGTGTGGAGATA FR ACTGCGAAUTGGACTCUGTGAGTGATGGTTGAGGATGTGTGGAGATA GR ACCGTUAAGCCTTGATUGTGAGTGATGGTGAGTGGTTGAGGATGTGTGGAGATA HR ACGCTAGGAUTCAACAUGTGAGTGATGGTTGAGGATGTGTGGAGATA IR ACAATCGCAACGCACUGTGAGTGATGGTTGAGGATGTGTGGAGATA JR AGGCCAACUACGCAATUTGAGTGATGGTTGAGGATGTGTGGAGATA KR AGATUGCGACCAAGTGUGTGAGTGATGGTTGAGGATGTGTGGAGATA LR AGACACGAAGGCUTACUGTGAGTGATGGTTGAGGATGTGTGGAGATA MRI AAGGCTCUGATCTAGGUGTGAGTGATGGTTGAGGATGTGTGGAGATA NR ACTUAACCGGCAUACCUGTGAGTGATGGTTGAGGATGTGTGGAGATA OR AAAGGUGCCGGUGACTUGTGAGTGATGGTTGAGGATGTGTGGAGATA PR ATCTCGAGCCACTTCATGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0086] Example 2: Construction of Mouse PPI Network

[0087] This embodiment continues from Embodiment 1, constructing a real protein interaction network in mice and verifying its accuracy.

[0088] To verify the stability of the established high-throughput protein interaction screening method, a mouse library was homologously recombined into the bait expression vector pUT18C and electroporated into E. coli to extract plasmids. Simultaneously, the mouse library was homologously recombined again into the bait expression vector pKNT25 and competent cells were prepared. The mouse library-pUT18C plasmid was electroporated again into the mouse library-pKNT25BTH101 competent cells and plated on M63 LA plates containing X-gal+IPTG+Amp+Kan for positive clone screening. A mouse PPI network was constructed using dPPI-seq (Table 3).

[0089] Genes with the top two UMIs in the barcode pairs selected by dPPI-seq were then used to predict the accuracy of machine learning. Six genes were then randomly selected from the dPPI-seq results for cloning. Figure 11 Yeast two-hybridization was performed and found to have 3 pairs of strong interactions and 1 pair of weak interactions. Figure 12 The results corresponded to the dPPI-seq screening results, verifying the accuracy of this high-throughput screening.

[0090] Table 3 Mouse protein-protein interaction pairs

[0091] Barcode Mouse-gene Mouse-gene Barcode1 Ywhag Hdgf Barcode2 mt-Cytb mt-Atp6 Barcode3 Kyat3 Etfb Barcode4 Hmgcs 2 Acaalb Barcode5 Ywhaq Ces1c Barcode6 Ftcd Aass Barcode7 Ywhaq Gclm Barcode8 Apoa4 Sugt1 Barcode9 Mapre2 Acol Barcode10 Edem2 mt-Cytb Barcode11 Thrsp Ptges3 Barcode12 Alb Ankrd17_ Barcode13 Rpl13a Rp16 Barcode14 Map2 S1c9a3r1 Barcode15 Serpinala Mtx2 Barcode16 Nubl Gda Barcode17 Fkbp4 Eef2 Barcode18 Alb Fabp1 Barcode19 Gm20425 Ndrg1 Barcode20 Apoa2 Tmem29 Barcode21 Alb Fam32a Barcode22 Gss Mup3 Barcode23 S1c9a3r1 Etfa Barcode24 Etfb Ptges3 Barcode25 Fbpl ubb Barcode26 Cyp2d9 Rp18 Barcode27 Acadvl Cyp2f2 Barcode28 Eif3e Sfpq Barcode29 Ldha Capn1 Barcode30 Apob Rp113a

[0092] As shown in Table 4, the comparison results are as follows:

[0093] Compared with existing technologies such as traditional yeast two-hybrid, AP-MS, and Alpha-Seq, the high-throughput, general-sample protein interaction screening method established in this example represents a significant performance breakthrough for the "library-to-library" protein interaction screening platform. It employs a comprehensive "library-to-library" screening mode, followed by single-cell sequencing based on droplet microfluidics, increasing throughput to an astonishing 10^65 times. 8 -10 9 Its level far exceeds the limited flux of other methods.

[0094] Meanwhile, the detection cost per interaction can be controlled to below $1, making it unparalleled in cost-effectiveness. More importantly, the platform can effectively achieve low false positives across multiple systems, solving the pain points of traditional methods such as high false positives and complex backgrounds. It enables broad-sample, high-throughput, high-accuracy, low-cost, and comprehensive protein interaction screening, providing an unprecedentedly powerful tool for large-scale protein interaction research.

[0095] Table 4. Performance Comparison between the "Library-to-Library" Protein Interaction Screening Platform and Existing Protein Interaction Screening Methods

[0096] Methods / Platforms Filtering mode Flux Scale False positive / false negative cost Advantages Limitations Traditional yeast two-hybrid (Y2H) Single Bait to Prey (pair-by-pair filtering) <![CDATA[Medium (10 4 – 10 5 level)]]> There are many false positives, and membrane protein restriction. Low to medium ($50-$200) Classic method, simple to operate Limited throughput necessitates numerous repeated experiments. AP-MS / Co-IP-MS For known proteins or complexes Low to medium (10²–10³ level) The background is complex and requires strict comparison. High (approximately $1,000-$5,000) Strong physiological correlation, verifiable interaction Limited sensitivity, cumbersome operation, high price, and poor repeatability Alpha-Seq (A-AlphaBio) Yeast pairing + NGS, pairwise batch processing <![CDATA[High (10 6 – 10 7 Level / month)]]> Dependence on library matching quality Mid ($500-$1,000) Quantitative analysis is possible, suitable for molecular gels / drug sieves. It still belongs to the "pairwise" mode, not full interaction. "Book-to-Book" PPI Screening Library vs. Library (Full Coverage Combination) <![CDATA[Extremely high (10 8 –10 9 level)]]> Low false positive rate across multiple systems Super low ($1) It can provide full coverage and quantitative analysis, saving time and costs. none

[0097] The above specific embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.

[0098] The above description is only a preferred embodiment of the present invention. Therefore, all equivalent changes or modifications made to the structure, features and principles described in the claims of this patent application are included in the scope of this patent application.

[0099] The gene sequence involved in this invention is as follows:

[0100] SEQ ID No.1;

[0101] Name: Modified 3'RT random primer DNA sequence

[0102] DNA type: other DNA

[0103] Biological origin: Artificial Sequence / synthetic construct

[0104] ACTCTGCGTTGATACCACTGCNNNNNNNN

[0105] SEQ ID No.2;

[0106] Name: DNA sequence of the linker

[0107] DNA type: other DNA

[0108] Biological origin: Artificial Sequence / synthetic construct

[0109] AAGCAGTGGTATCAACGCAGAGT

[0110] SEQ ID No. 3;

[0111] Name: Linker linker connection site sequence in plasmid

[0112] DNA type: other DNA

[0113] Biological origin: Artificial Sequence / synthetic construct

[0114] ACTCTGCGTTGATACCACTGCTT

[0115] SEQ ID No.4;

[0116] Name: DNA sequences of multiple pairs of IdeoxyU-modified adapter primers F

[0117] DNA type: other DNA

[0118] Biological origin: Artificial Sequence / synthetic construct

[0119] GGAGTTGGAGTGAGTGGATGAGTGATG

[0120] SEQ ID No. 5;

[0121] Name: DNA sequences of multiple pairs of IdeoxyU-modified adapter primers R

[0122] DNA type: other DNA

[0123] Biological origin: Artificial Sequence / synthetic construct

[0124] GTGAGTGATGGTTGAGGATGTGTGGAGATA

[0125] SEQ ID No. 6;

[0126] Name: DNA sequence of the full-length open reading frame of the human EEF1G gene

[0127] DNA type: other DNA

[0128] Biological origin: Artificial Sequence / synthetic construct

[0129]

[0130] SEQ ID No.7;

[0131] Name: DNA sequence of modified 3'RT Oligo-dT primers

[0132] DNA type: other DNA

[0133] Biological origin: Artificial Sequence / synthetic construct

[0134] ACTCTGCGTTGATACCACTGCTTTTTTTTTTTTTTTTTT

[0135] SEQ ID No. 8;

[0136] Name: DNA sequence of modified TSO primers 1

[0137] DNA type: other DNA

[0138] Biological origin: Artificial Sequence / synthetic construct

[0139] TCGACTCTAGAGGATCCCrGrGrG

[0140] SEQ ID No. 9;

[0141] Name: DNA sequence of modified TSO primer 2

[0142] DNA type: other DNA

[0143] Biological origin: Artificial Sequence / synthetic construct

[0144] TCGACTCTAGAGGATCCCrG

[0145] SEQ ID No. 10;

[0146] Name: DNA sequence 3 of the modified TSO primer

[0147] DNA type: other DNA

[0148] Biological origin: Artificial Sequence / synthetic construct

[0149] TCGACTCTAGAGGATCCCrGrG

[0150] SEQ ID No. 11;

[0151] Name: Deoxyuracil (dU) modified adapter primer AF

[0152] DNA type: other DNA

[0153] Biological origin: Artificial Sequence / synthetic construct

[0154] AGCTTACTTGTGAAGATGGAGTTGGAGTGAGTGGATGAGTGATG

[0155] SEQ ID No. 12;

[0156] Name: Deoxyuracil (dU) modified adapter primer BF

[0157] DNA type: other DNA

[0158] Biological origin: Artificial Sequence / synthetic construct

[0159] ACTTGTAAGCUGTCTAUGGAGTTGGAGTGAGTGGATGAGTGATG

[0160] SEQ ID No. 13;

[0161] Name: Deoxyuracil (dU) modified adapter primer CF

[0162] DNA type: other DNA

[0163] Biological origin: Artificial Sequence / synthetic construct

[0164] ACTCTGUCAGGTCCGAUGGAGTTGGAGTGAGTGGATGAGTGATG

[0165] SEQ ID No. 14;

[0166] Name: Deoxyuridine (dU) modified adapter primer DF

[0167] DNA type: other DNA

[0168] Biological origin: Artificial Sequence / synthetic construct

[0169] ACCTCCTCCUCCAGAAUGGAGTTGGAGTGAGTGGATGAGTGATG

[0170] SEQ ID No. 15;

[0171] Name: Deoxyuracil (dU) modified adapter primer EF

[0172] DNA type: other DNA

[0173] Biological origin: Artificial Sequence / synthetic construct

[0174] AACCGGACACACUTAGUGGAGTGTGGAGTGAGTGGATGAGTGATG

[0175] SEQ ID No. 16;

[0176] Name: Deoxyuracil (dU) modified adapter primer FF

[0177] DNA type: other DNA

[0178] Biological origin: Artificial Sequence / synthetic construct

[0179] AGAGTCCAAUTCGCAGUGGAGTTGGAGTGAGTGGATGAGTGATG

[0180] SEQ ID No. 17;

[0181] Name: Deoxyuracil (dU) modified linker primer GF

[0182] DNA type: other DNA

[0183] Biological origin: Artificial Sequence / synthetic construct

[0184] AATCAAGGCUTAACGGUGGAGTTGGAGTGAGTGGATGAGTGATG

[0185] SEQ ID No. 18;

[0186] Name: Deoxyuracil (dU) modified adapter primer HF

[0187] DNA type: other DNA

[0188] Biological origin: Artificial Sequence / synthetic construct

[0189] ATGTTGAAUCCTAGCGUGGAGTTGGAGTGAGTGGATGAGTGATG

[0190] SEQ ID No. 19;

[0191] Name: Deoxyuracil (dU) modified adapter primer IF

[0192] DNA type: other DNA

[0193] Biological origin: Artificial Sequence / synthetic construct

[0194] AGTGCGTUGCGAATTGUGGAGTGTGGAGTGAGTGGATGAGTGATG

[0195] SEQ ID No. 20;

[0196] Name: Deoxyuracil (dU) modified adapter primer JF

[0197] DNA type: other DNA

[0198] Biological origin: Artificial Sequence / synthetic construct

[0199] AATTGCGUAGTTGGCCUGGAGTTGGAGTGAGTGGATGAGTGATG

[0200] SEQ ID No. 21;

[0201] Name: Deoxyuracil (dU) modified adapter primer KF

[0202] DNA type: other DNA

[0203] Biological origin: Artificial Sequence / synthetic construct

[0204] ACACTTGGUCGCAATCUGGAGTTGGAGTGAGTGGATGAGTGATG

[0205] SEQ ID No. 22;

[0206] Name: Deoxyuracil (dU) modified linker primer LF

[0207] DNA type: other DNA

[0208] Biological origin: Artificial Sequence / synthetic construct

[0209] AGTAAGCCUTCGTGTCUGGAGTTGGAGTGAGTGGATGAGTGATG

[0210] SEQ ID No. 23;

[0211] Name: Deoxyuracil (dU) modified adapter primer MF

[0212] DNA type: other DNA

[0213] Biological origin: Artificial Sequence / synthetic construct

[0214] ACCTAGAUCAGAGCCTUGGAGTTGGAGTGAGTGGATGAGTGATG

[0215] SEQ ID No. 24;

[0216] Name: Deoxyuracil (dU) modified linker primer NF

[0217] DNA type: other DNA

[0218] Biological origin: Artificial Sequence / synthetic construct

[0219] AGGTAUGCCGGUTAAGUGGAGTTGGAGTGAGTGGATGAGTGATG

[0220] SEQ ID No. 25;

[0221] Name: OF adapter primer modified with deoxyuracil (dU)

[0222] DNA type: other DNA

[0223] Biological origin: Artificial Sequence / synthetic construct

[0224] AAGUCACCGGCACCUTUGGAGTTGGAGTGAGTGGATGAGTGATG

[0225] SEQ ID No. 26;

[0226] Name: Deoxyuracil (dU) modified adapter primer BR

[0227] DNA type: other DNA

[0228] Biological origin: Artificial Sequence / synthetic construct

[0229] ATAGACAGCUTACAAGUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0230] SEQ ID No. 27;

[0231] Name: Deoxyuridine (dU) modified adapter primer CR

[0232] DNA type: other DNA

[0233] Biological origin: Artificial Sequence / synthetic construct

[0234] ATCGGACCUGACAGAGUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0235] SEQ ID No. 28;

[0236] Name: Deoxyuracil (dU) modified adapter primer DR

[0237] DNA type: other DNA

[0238] Biological origin: Artificial Sequence / synthetic construct

[0239] ATTCUGGAGGAGGAGGUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0240] SEQ ID No. 29;

[0241] Name: Deoxyuridine (dU) modified adapter primer ER

[0242] DNA type: other DNA

[0243] Biological origin: Artificial Sequence / synthetic construct

[0244] ACTAAGTGUGTCCGGTUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0245] SEQ ID No. 30;

[0246] Name: Deoxyuracil (dU) modified adapter primer FR

[0247] DNA type: other DNA

[0248] Biological origin: Artificial Sequence / synthetic construct

[0249] ACTGCGAAUTGGACTCUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0250] SEQ ID No. 31;

[0251] Name: Deoxyuracil (dU) modified adapter primer GR

[0252] DNA type: other DNA

[0253] Biological origin: Artificial Sequence / synthetic construct

[0254] ACCGTUAAGCCTTGATUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0255] SEQ ID No. 32;

[0256] Name: Deoxyuridine (dU) modified adapter primer HR

[0257] DNA type: other DNA

[0258] Biological origin: Artificial Sequence / synthetic construct

[0259] ACGCTAGGAUTCAACAUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0260] SEQ ID No. 33;

[0261] Name: Deoxyuracil (dU) Modified Linker Primer IR

[0262] DNA type: other DNA

[0263] Biological origin: Artificial Sequence / synthetic construct

[0264] ACAATUCGCAACGCACUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0265] SEQ ID No. 34;

[0266] Name: Deoxyuracil (dU) Modified Proton Primer JR

[0267] DNA type: other DNA

[0268] Biological origin: Artificial Sequence / synthetic construct

[0269] AGGCCAACUACGCAATUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0270] SEQ ID No. 35;

[0271] Name: Deoxyuracil (dU) modified adapter primer KR

[0272] DNA type: other DNA

[0273] Biological origin: Artificial Sequence / synthetic construct

[0274] AGATUGCGACCAAGTGUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0275] SEQ ID No. 36;

[0276] Name: Deoxyuridine (dU) modified adapter primer LR

[0277] DNA type: other DNA

[0278] Biological origin: Artificial Sequence / synthetic construct

[0279] AGACACGAAGGCUTACUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0280] SEQ ID No. 37;

[0281] Name: Deoxyuracil (dU) modified adapter primer MR

[0282] DNA type: other DNA

[0283] Biological origin: Artificial Sequence / synthetic construct

[0284] AAGGCTCUGATCTAGGUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0285] SEQ ID No. 38;

[0286] Name: Deoxyuracil (dU) modified adapter primer NR

[0287] DNA type: other DNA

[0288] Biological origin: Artificial Sequence / synthetic construct

[0289] ACTUAACCGGCAUACCUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0290] SEQ ID No. 39;

[0291] Name: Deoxyuridine (dU) modified adapter primer OR

[0292] DNA type: other DNA

[0293] Biological origin: Artificial Sequence / synthetic construct

[0294] AAAGGUGCCGGUGACTUGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0295] SEQ ID No.40;

[0296] Name: Proximity primers modified with deoxyuracil (dU)

[0297] DNA type: other DNA

[0298] Biological origin: Artificial Sequence / synthetic construct

[0299] ATCTCGAGCCACTTCATGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0300] SEQ ID No.41;

[0301] Name: P5 primer

[0302] DNA type: other DNA

[0303] Biological origin: Artificial Sequence / synthetic construct

[0304] AATGATACGGCGACCACCGAGATCTACACCTCTCTATTCGTCGGCAGCGTCAGATGTGTATAAGAGACAGGTGAGTGATGGTTGAGGATGTGTGGAGATA

[0305] SEQ ID No.42;

[0306] Name: P7 primer

[0307] DNA type: other DNA

[0308] Biological origin: Artificial Sequence / synthetic construct

[0309] CAAGCAGAAGACGGCATACGAGATTAAGGCGAGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG.

Claims

1. A high-throughput general sample protein interaction screening method, characterized in that: 1) Construction of generic sample libraries: mRNA was extracted from the target sample, and then a pan-sample library was obtained using modified primers. 2) Homogenization of generic sample libraries; 3) Modify the plasmid of the dual-hybrid system to obtain the modified plasmid; 4) Homologous recombination was performed on the sample library and the modified plasmid of the double hybrid system, and the plasmid was introduced into organisms to screen for organisms with mutual interaction libraries and cDNA mutual interaction libraries with single-cell markers. 5) Identify PPI pairs for single-cell sequencing of organisms with inter-text libraries and construct a pan-sample protein interaction network; Step 1) specifically involves: firstly, extracting the mRNA from the target sample, and then using modified 3'RT Oligo-dT primers and TSO primers to obtain a full-length eukaryotic cDNA library as a pan-sample library using Smart-seq 3. Specifically, this includes: firstly, using Oligo(dT) magnetic beads to specifically capture eukaryotic mRNA via the Poly-A tail, and then using modified 3'RT random primers for reverse transcription to obtain a cDNA library with uneven lengths. The sequence of the modified 3'RT random primer is SEQ ID No. 1, namely ACTCTGCGTTGATACCACTGCNNNNNNNN; Step 1) specifically involves: for a polycistronic cDNA library of prokaryotes, rRNA digestion is first performed using an rRNA digestion kit, the 5' end of the mRNA is capped using a vaccinia virus capping system, and finally, a Poly-A tail structure is added to the 3' end of the mRNA using E. coli Poly-A polymerase and ATP. The mRNA is then used with modified 3'RT Oligo-dT primers and TSO primers to obtain a full-length cDNA library of prokaryotes as a pan-sample library using Smart-seq 3. Step 2) specifically involves using Duplex-Specific Nuclease enzyme to homogenize the cDNA in the pansample library. Specifically, step 3) involves modifying the plasmid of the bacterial double-hybrid system or the yeast double-hybrid system by connecting it to a linker containing homologous arms to form a modified plasmid. The linker sequence is SEQ ID No. 2, namely AAGCGTGGTATCAACGCAGAGT. The linker linker in the plasmid is located at the 3' end of the MSC multiple cloning site of the vector, and the linker site sequence is SEQ ID No. 3, namely ACTCTGCGTTGATACCACTGCTT. Step 4) specifically involves: firstly, homologous recombination of the target sample cDNA library with the plasmid modified by the dual-hybrid system to form bait protein and prey protein-particle libraries, then transferring the plasmid library to the corresponding organisms via electrotransduction or chemical transduction, and screening for positive PPI combination clones in the screening medium to obtain organisms with protein-protein interaction pairs, and inducing gene expression using IPTG to obtain a single-cell labeled cDNA interaction library; The modified plasmid contains a linker sequence, and the positive result is that the bacterial BACTH double hybrid system turns blue under the screening medium. The bait protein particle library is composed of homologous recombination of the modified plasmid pKT25 or pKNT25 of the dual-hybrid system and a cDNA library, and the prey protein particle library is composed of homologous recombination of the modified plasmid pUT18C or pUT18 of the dual-hybrid system and a cDNA library. Step 5) employs single-cell transcriptome sequencing based on random primers, second-generation sequencing based on droplet PCR for direct plasmid copy amplification, or third-generation sequencing.

2. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: Step 5) specifically involves single-cell transcriptome sequencing based on random primers: First, the screened bacterial or yeast organisms were fixed in a 4% PFA solution for 8-16 h to allow proteins and nucleic acids to cross-link. Lysozyme and Zymolyase enzymes were then added to open the cell walls, followed by cell membrane permeation with a 0.4% Triton solution. In situ reverse transcription and TdT terminal transferase were then performed to add Poly(A) to the 3' hydroxyl end of cDNA. Oligo(dT) hydrogel-encoded microspheres were then added to the bacteria or yeast for single-cell encapsulation and cDNA double-strand amplification, resulting in a cDNA library with cell barcode markers. The cDNA library was then used to perform end repair and ligation of P5 / P7 adapters to construct an Illumina sequencing library. Finally, the obtained Illumina sequencing library was sequenced. The sequenced sequences were compared with the reference genome of the target sample species to identify interacting protein pairs and cDNA combinations with the same cell barcode, thus constructing an interacting protein network.

3. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: Step 5) of the second-generation sequencing based on droplet PCR for direct plasmid copy amplification specifically involves: using modified hydrogel-encoded microspheres containing linker sequences, along with heat-resistant droplet-generating oil and an aqueous phase stabilizer, to encapsulate selected positive PPI bacteria or yeast into single cells for PCR amplification. After amplification, the amplified products are fragmented using a transposase DNA library construction kit. Then, a cDNA sequencing library containing cell barcode sequences is constructed using P5 / P7 linker primers containing cell barcodes. Finally, the obtained cDNA sequencing library is sequenced, and the sequenced sequences are compared with the reference genome of the target sample species to identify interacting protein pairs and cDNA combinations with the same cell barcode, thus constructing an interacting protein network.

4. The high-throughput general sample protein interaction screening method according to claim 3, characterized in that: For yeast organisms, a mixture containing lysis buffer was added during PCR amplification, specifically: 2 mg / g cell Zymolase lysin was used to lyse the cell wall at 37°C for 30 min.

5. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: In step 5), the third-generation sequencing based on droplet PCR for direct plasmid copy amplification is as follows: First, for the cDNA intergenic library obtained in step 4) with single-cell labeling, secondary amplification is performed using adapter primers modified with deoxyuridine (dU). After amplification, the obtained libraries are mixed together and purified with magnetic beads. After purification, a single nucleotide gap is generated at the uridine position using USER enzyme. Finally, a single-stranded overhang is generated by cleavage with Endo VIII enzyme and discarded, leaving the double-stranded strand with sticky ends. Then, the double-stranded strand with sticky ends is annealed and ligated with Taq high-fidelity DNA ligase to form a long fragment. Finally, third-generation sequencing is performed using Nanopore and PacBio platforms.

6. The high-throughput general sample protein interaction screening method according to claim 1, characterized in that: In step 5), the third-generation sequencing based on droplet PCR for direct plasmid copy amplification specifically involves: first, using the single-cell labeled cDNA intergenic library obtained in step 4), TdTase terminal transferase is used to add two sequences from either the dATP-dTTP pair or the dCTP-dGTP pair to the 3' hydroxyl end of the DNA molecule. The products after adding the two sequences are purified by magnetic beads. After magnetic bead purification, the two products are mixed and treated at 95°C for 1 min, followed by slow annealing at 0.1°C / s to 25°C. Finally, T4 DNA polymerase is added for DNA repair and ligation, and the T4 DNA ligase directly tandemly grows fragments. Finally, third-generation sequencing is performed using Nanopore and PacBio platforms.