A method for bacterial single-cell transcriptome sequencing to map a mouse full-coverage protein interaction network

By constructing a mouse open reading frame library and modifying the BACTH bacterial two-hybrid system, combined with random primer single-cell transcriptome sequencing, high-throughput screening and identification of a full-coverage mouse protein interaction network was achieved, solving the problem of insufficient detection throughput in existing technologies and providing panoramic protein interaction information support.

CN121428064BActive Publication Date: 2026-05-19LIANGZHU LAB
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LIANGZHU LAB
Filing Date
2025-12-31
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies are insufficient for a systematic and global analysis of the full-coverage protein interaction network in mice. They have limited detection throughput, long experimental cycles, and high costs, and cannot cover all potential protein interaction relationships.

Method used

A high-throughput screening system using mouse open reading frame library construction, BACTH bacterial two-hybrid system modification, and random primer single-cell transcriptome sequencing was employed. Through library-to-library screening and high-throughput PPI identification, a mouse full-coverage protein interaction network was constructed.

Benefits of technology

It enables high-throughput, high-accuracy, and low-cost mapping of a full-coverage protein interaction network in mice, breaking through the bottleneck of traditional "one-to-one" detection, providing systematic and panoramic protein interaction information, and supporting mouse functional genomics and disease mechanism analysis as well as drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121428064B_ABST
    Figure CN121428064B_ABST
Patent Text Reader

Abstract

The application discloses a method for drawing a mouse full-coverage protein interaction network by bacterial single-cell transcriptome sequencing. The method comprises the following steps: extracting mRNA of a mouse general tissue, obtaining a cDNA library by using a modified random primer or an oligo-dT primer, uniformly processing a mouse protein library, modifying a plasmid of a BACTH bacterial double-hybrid system to obtain a modified plasmid, homologous recombining the sample library and the modified plasmid of the double-hybrid system, introducing into DHM1 E. coli to screen positive PPI clones with an interaction library, and performing single-cell transcriptome sequencing based on the random primer on the screened clones to construct a mouse protein interaction network. The application is used for high-throughput screening of tens of thousands of positive PPI combination clones obtained by a 'library versus library' bacterial or yeast double-hybrid method, and identification of the same Barcode cDNA combination to realize PPI network identification in different intra-species, inter-species and hybridization technology systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of protein interaction screening, and in particular to a method for mapping a mouse full-coverage protein interaction network using bacterial single-cell transcriptome sequencing. Background Technology

[0002] Proteins are the direct executors of life activities, and their interactions constitute the core network maintaining cellular homeostasis and physiological function. The protein-protein interaction network (PPI network) systematically describes the complex relationships formed by proteins within the cell and is a key foundation for revealing the molecular mechanisms of life processes. Mapping a high-precision, high-coverage PPI network not only helps in understanding fundamental life processes such as cell signaling, metabolic regulation, transcriptional regulation, and developmental differentiation, but also provides important references for research on the molecular mechanisms of diseases, drug target discovery, and the design of novel therapeutic strategies. Among mammals, mice, as one of the most important model organisms, are highly similar to humans in terms of genome structure, developmental physiology, and disease models, making them a core experimental animal for studying human diseases and validating candidate drug effects. Therefore, constructing a high-resolution, high-accuracy PPI network covering the entire mouse genome has irreplaceable scientific and applied value for advancing basic biological research, precision medicine, and drug development.

[0003] However, existing protein interaction research techniques still have significant limitations. Traditional methods such as yeast two-hybrid (Y2H), co-immunoprecipitation (Co-IP), and affinity purification-mass spectrometry (AP-MS) are all based on pairwise or limited combination interaction screening, representing a typical "one-to-one" detection model. These methods require prior knowledge of the coding sequences of interacting proteins and rely on heterologous expression systems for stepwise validation. They are time-consuming, labor-intensive, and have limited throughput, making it difficult to cover all potential protein interaction relationships. This methodological limitation results in previously obtained protein interaction maps often having significant coverage blind spots, reflecting only information on some key pathways or protein complexes, and failing to achieve a systematic and global analysis of the mouse protein interaction network.

[0004] Therefore, there is an urgent need for new methods that can achieve full-coverage protein interaction network mapping in mice, fundamentally breaking through the bottleneck of traditional "one-to-one" detection, solving the problems of limited throughput and insufficient coverage of existing technologies, and parsing tens of thousands of potential protein interaction events in parallel in a single experiment. This would achieve true full-coverage protein interaction network mapping, shorten the experimental cycle, reduce detection costs, provide systematic and panoramic protein interaction information, reveal the interaction patterns between different tissues and cell types in mice, as well as between microorganisms and the host, and provide comprehensive and accurate interaction information support for functional genomics, disease mechanism analysis, and drug development in mouse models. Summary of the Invention

[0005] To address the current lack of high-throughput library-to-library protein interaction screening methods, this invention provides a high-throughput library-to-library screening system and its usage method, and constructs a mouse full-coverage protein interaction network. The system mainly includes a mouse open reading frame acquisition system, a library-to-library PPI screening system, and a single-cell transcriptome sequencing identification system based on random primers.

[0006] like Figure 1 As shown, the technical solution of the present invention is as follows:

[0007] 1) Construction of mouse protein library

[0008] mRNA was extracted from whole mouse tissues in vitro, and then a mouse protein library was obtained using modified primers as the open reading frame of the mouse target sample;

[0009] 2) Homogenization of mouse protein libraries;

[0010] 3) Modify the plasmid of the BACTH bacterial two-hybrid system to obtain the modified plasmid;

[0011] 4) The mouse protein library and the modified T18 and T25 plasmids of the double hybrid system were homologously recombined and introduced into DHM1 Escherichia coli. The bacteria were screened in M63 medium containing Amp and Kan resistance and IPTG inducer to obtain mice with full coverage positive PPI clones.

[0012] 5) After obtaining bacteria with positive PPI clones, high-throughput PPI pair identification was performed using single-cell transcriptome sequencing based on random primers. A pair of open reading frames in a cell is a pair of protein interaction pairs, and a mouse full-coverage protein interaction network was constructed.

[0013] Step 1) Mouse protein library construction specifically includes: firstly, extracting mouse mRNA, and using modified 3'RT Oligo-dT primers and TSO primers to obtain a full-length mouse open reading frame library using Smart-seq 3; then, fusing the mouse open reading frame library to the N-terminus of a plasmid vector. When obtaining the library, Oligo(dT) magnetic beads are first used to specifically capture eukaryotic mRNA via the Poly-A tail, and then modified 3'RT random primers are used for reverse transcription to obtain a cDNA library as the mouse protein library. The modified 3'RT random primer sequence is SEQ ID No. 1, i.e., ACTCTGCGTTGATACCACTGCNNNNNNNN, to achieve high-sensitivity, low-false-negative protein interaction detection.

[0014] The modified 3'RT Oligo-dT primer sequence is SEQ ID No. 4, namely ACTCTGCGTTGATACCACTGCTTTTTTTTTTTTTTTTTTTTT, and the modified TSO primer sequences include SEQ ID Nos. 5-7, namely TCGACTCTAGAGGATCCCrGrGrG, TCGACTCTAGAGGATCCCrG, and TCGACTCTAGAGGATCCCrGrG.

[0015] The above sequences were optimized based on the sequence of the plasmid vector. The 3'RT sequence was added to the vector through modification and also corresponds to the sequence on the hydrogel-encoded microspheres. The TSO primers were designed based on the homologous arm sequence at the 5' end of the vector cloning site.

[0016] Step 2) specifically involves using a double-stranded specific nuclease to homogenize the mouse protein library, reducing the abundance of high-copy genes, making the gene concentration in the library more uniform, improving the efficiency of protein interaction network mapping, and discovering interaction networks of low-abundance genes.

[0017] Specifically, step 3) involves modifying the plasmid of the BACTH bacterial double hybrid system, linking it with a linker sequence containing a homologous arm, performing single-cell labeling, and forming the modified plasmid.

[0018] The linker sequence is SEQ ID No. 2, namely AAGCGTGGTATCAACGCAGAGT, which has a homologous sequence with the modified 3'RT Oligo-dT primer. The linker in the plasmid is located at the 3' position of the MSC multiple cloning site of the vector, and the linker sequence is SEQ ID No. 3, namely ACTCTGCGTTGATACCACTGCTT.

[0019] Step 4) specifically involves:

[0020] First, mouse protein libraries were homologously fused with the C-terminus of pUT18C and pKT25 plasmids modified using the BACTH bacterial double-hybrid system to form pUT18C and pKT25 protein-particle libraries. Alternatively, open reading frame libraries obtained using random primers were homologously fused with the N-terminus of pUT18 and pKNT25 plasmids modified using the BACTH bacterial double-hybrid system to form pUT18 and pKNT25 protein-particle libraries. Subsequently, the plasmid libraries were electrotransduced into DHM1 Escherichia coli, and positive PPI clones were screened in M63 medium containing Amp and Kan resistance and IPTG inducer to obtain bacteria with positive PPI clones.

[0021] The screening medium is such as M63 medium (5x M63 medium formula: weigh 10 g of (NH4)2SO4, 68 g of KH4PO4, 2.5 mg of FeSO4·7H2O and 5 mg of vitamin B1, add deionized water to make up to 1 L, adjust the pH to 7.0 with KOH and then autoclave).

[0022] The modified plasmid contains the linker sequence AAGCGTGGTATCAACGCAGAGT, and bacteria with positive PPI clones appear blue on the selection medium plate.

[0023] Step 5) specifically involves single-cell transcriptome sequencing based on random primers:

[0024] First, bacteria with positive PPI clones were transferred to liquid culture medium for amplification and fixation with 4% PFA solution for 8-16 h to allow protein-nucleic acid cross-linking. Then, bacterial cell membranes were permeabilized with PBS containing 0.04% PBS at -20°C, and lysozyme was used to perforate the E. coli cell walls. Next, in situ reverse transcription and tailing dA reactions were performed using reverse transcriptase and TdT terminal transferase. Then, oligo(dT) hydrogel-encoded microspheres were added to bacteria or yeast for single-cell encapsulation and cDNA double-strand amplification to obtain PCR libraries (referencing the Nature Communications article "Droplet-based high-throughput single microbe RNA sequencing by smRandom-seq"). The PCR libraries were then used to repair ends and ligate P5 / P7 adapters using the VAHTS Universal Pro DNA Library Preparation Kit (for Illumina) to construct Illumina sequencing libraries. Finally, the obtained Illumina sequencing libraries were sequenced using an Illumina X plus 25B sequencer for paired-end 150 bp sequencing, totaling 300 bp. bp sequencing involves comparing the sequenced protein with a reference genome of a mouse species to identify interacting protein pairs and construct an interacting protein network.

[0025] The high throughput mentioned in this invention refers to a high-throughput identification technology system that can identify tens of thousands to hundreds of thousands of positive PPI clones in a single operation.

[0026] The beneficial effects of this invention are:

[0027] This invention establishes a high-throughput library-to-library PPI screening method and constructs a mouse-wide PPI network. This method is applicable to high-throughput screening of tens of thousands of positive PPI combination clones obtained from library-to-library bacterial or yeast two-hybrids, and single-cell sequencing identification of identical barcode cDNA combination pairs. It can achieve PPI network identification within different species, between species (model species vs. non-model species, bacteria vs. hosts), and hybridization technology systems (bacterial two-hybrids, yeast one-hybrids), breaking through the bottleneck of traditional "one-to-one" and a few "one-to-many" screenings. It has the advantages of high throughput, high accuracy, low cost, and full coverage of protein interaction identification. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the present invention.

[0029] Figure 2 The image of the modified pKNT25 plasmid;

[0030] Figure 3 The image of the modified pKT25 plasmid;

[0031] Figure 4 The image of the modified pUT18 plasmid;

[0032] Figure 5 The image of the modified pUT18C plasmid;

[0033] Figure 6 The mouse PPI-positive clones obtained through screening;

[0034] Figure 7 PCR identification of the mouse PPI-interacting positive clones obtained through screening;

[0035] Figure 8 This represents a 20-pair PPI interaction network in mice. Detailed Implementation

[0036] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0037] like Figure 1 As shown, embodiments of the present invention are as follows:

[0038] 1. Mouse open reading frame library acquisition system

[0039] Mouse heart, liver, spleen, lung, kidney, muscle, bone, skin, brain, stomach, and intestinal tissues were collected. The tissues were ground into powder using liquid nitrogen, and mouse RNA was extracted by direct lysis according to the TRIzol instructions. RNA integrity and concentration were assessed by gel electrophoresis and UV spectrophotometry. Full-length open reading frames were then obtained directly using the optimized Smart-seq 3 method, as follows: RT reactions were performed using an RNase H-free MMLVRT enzyme (such as Maxima H-minus reverse transcriptase enzyme (Thermo Scientific)). The reaction system contained 25 mM Tris-HCl pH 8.0-8.4, 30 mM NaCl, 2.5 mM MgCl2, 1 mM GTP, 8 mM DTT, 0.25 U RNase inhibitor, 0.3 mM dNTPs, 0.1 μM TSO primer, and 0.1 μM 3'RT primer (3'RT Oligo-dT or 3'RT...). Random primers, RNA, dNTPs, and 3' RT primers were pre-denatured at 72°C for 5 min, then immediately placed on ice with the remaining reagents added. The reaction was performed at 42°C for 90 min, followed by 10 cycles of 50°C for 2 min and 42°C for 2 min. After the RT reaction, cDNA library amplification was performed for 25 cycles using a 2X KAPA HiFi HotStartReadyMix. After amplification, the cDNA library was purified using 0.6X VAHTS DNA Clean Beads (N411).

[0040] 2. To ensure the cDNA library is free of repetitive sequences and covers more transcript information, cDNA homogenization was performed using Duplex-Specific Nuclease (DSN). This was done using the PuYinTe Bio cDNA homogenization kit (FZ1031). First, 200 ng of the cDNA library was mixed with 4x Hybridization buffer and denatured at 98°C for 2 min, then incubated at 68°C for 5 h. Then, 1 μL of DSN and 1 μL of 10×DSN Reaction buffer were added, gently mixed, briefly centrifuged, and incubated at 68°C for 7-20 min to degrade the double-stranded cDNA. Finally, 10 μL of 2×DSN Stop Buffer was added, gently mixed, briefly centrifuged, and incubated at 65°C for 5 min to terminate the reaction. The remaining homogenized single-stranded cDNA was amplified by PCR for library cloning. After amplification, the cDNA library was purified using 0.6x VAHTS DNA Clean Beads (N411).

[0041] 3. Full-coverage PPI screening in mice

[0042] The purified mouse cDNA library was then combined with the modified T18 / T25 vector ( Figures 2-5 Homologous recombination was performed using the Basic Seamless Cloning and Assembly kit, and the cells were electroporated into DH5α competent cells. Positive clones were verified by PCR and Sanger sequencing. Bacteria were then collected for plasmid extraction from T18-mouse and T25-mouse cDNA libraries. The T18-mouse cDNA library was first electroporated into DH101 E. coli competent cells. DH101 E. coli containing the T18-mouse cDNA library were then used to prepare new competent cells using a supercompetent bacterial preparation kit (Beyotime, D0302). The T25-mouse cDNA library was again electroporated into DH101 competent cells containing the T18-mouse cDNA library. These cells were then plated on M63 medium containing Amp and Kan resistance and IPTG inducer for screening of positive PPI combination clones. Figure 6 ), select single clones and perform PCR again ( Figure 7 After verifying positive clones using Sanger sequencing, mouse PPI clones with full coverage were collected for high-throughput identification.

[0043] 4. Constructing a full-coverage protein-protein interaction map of mice using single-bacterial transcriptome sequencing based on random primers.

[0044] The collected bacteria were fixed overnight in 4% paraformaldehyde to cross-link the intracellular RNA, DNA, and proteins. Lysozyme was used to digest the cell wall, permeabilizing the fixed bacteria to facilitate the next step of in situ reverse transcription. The microorganisms were used as reaction vessels for the in situ reaction. Random primers were added to bind to the intracellular RNA, and total RNA was captured and synthesized into cDNA. A poly-A tail was added in situ to the 3' end of the cDNA using a terminal transferase (TdT). Each step of the process was followed by washing with buffer 3-8 times to prevent reagent residue from affecting subsequent reactions. Individual bacteria and labeled microbeads were encapsulated into droplets using a microfluidic device. Poly-T primers were released from the microbeads by enzymatic digestion, digesting the bacterial RNA to release cDNA from the bacteria. The poly-T primers bound to the poly-A tail at the end of the cDNA, subsequently extending the cDNA to add specific coding, and a molecular tag (UMI) was added to each cDNA. After demulsification, cDNA was collected and purified, expanded, and sequencing adapters were added to construct a sequencing library. The cDNA product of rRNA was digested using Cas9, and the cDNA product of mRNA was enriched for high-throughput sequencing. After quality control and filtering, the sequencing data was aligned to the Ecoli_bw25113 genome. Unmapping reads were extracted and aligned to human and mouse reference genes. Feature Counts software was used to count the number of mouse and human transcripts detected in the same barcode and their UMIs. A UMI threshold was designed to exclude false positives. Higher UMIs for mouse and human transcripts detected by a barcode resulted in fewer false positives and higher reliability of the interaction pairs. Finally, a mouse full-coverage protein interaction network was constructed from the selected interaction pairs. A total of 6932 interacting protein pairs were identified after one screening, including 1293 proteins exhibiting multi-protein interactions. Figure 8 ).

[0045] The above specific embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.

[0046] The above description is merely a preferred embodiment of the present invention. Therefore, all equivalent changes or modifications made to the structure, features, and principles described in the claims of this patent application are included within the scope of this patent application.

[0047] The gene sequence involved in this invention is as follows:

[0048] SEQ ID No. 1:

[0049] Name: Modified 3'RT random primer DNA sequence

[0050] DNA type: other DNA

[0051] Biological origin: Artificial Sequence / synthetic construct

[0052] ACTCTGCGTTGATACCACTGCNNNNNNNN

[0053] SEQ ID No. 2:

[0054] Name: DNA sequence of the linker

[0055] DNA type: other DNA

[0056] Biological origin: Artificial Sequence / synthetic construct

[0057] AAGCAGTGGTATCAACGCAGAGT

[0058] SEQ ID No. 3:

[0059] Name: Linker linker connection site sequence in plasmid

[0060] DNA type: other DNA

[0061] Biological origin: Artificial Sequence / synthetic construct

[0062] ACTCTGCGTTGATACCACTGCTT

[0063] SEQ ID No.4:

[0064] Name: DNA sequence of modified 3'RT Oligo-dT primers

[0065] DNA type: other DNA

[0066] Biological origin: Artificial Sequence / synthetic construct

[0067] ACTCTGCGTTGATACCACTGCTTTTTTTTTTTTTTTTTT

[0068] SEQ ID No. 5:

[0069] Name: DNA sequence of modified TSO primers 1

[0070] DNA type: other DNA

[0071] Biological origin: Artificial Sequence / synthetic construct

[0072] TCGACTCTAGAGGATCCCrGrGrG

[0073] SEQ ID No. 6:

[0074] Name: DNA sequence of modified TSO primer 2

[0075] DNA type: other DNA

[0076] Biological origin: Artificial Sequence / synthetic construct

[0077] TCGACTCTAGAGGATCCCrG

[0078] SEQ ID No. 7:

[0079] Name: DNA sequence 3 of the modified TSO primer

[0080] DNA type: other DNA

[0081] Biological origin: Artificial Sequence / synthetic construct

[0082] TCGACTCTAGAGGATCCCrGrG.

Claims

1. A method for mapping a mouse full-coverage protein interaction network using bacterial single-cell transcriptome sequencing, characterized in that: 1) Construction of mouse protein library mRNA was extracted from whole mouse tissues, and then a mouse protein library was obtained using modified primers. 2) Homogenization of mouse protein libraries; 3) Modify the plasmid of the BACTH bacterial two-hybrid system to obtain the modified plasmid; 4) The mouse protein library and the modified plasmid of the double hybrid system were homologously recombined and introduced into DHM1 Escherichia coli. The bacteria were screened in M63 medium containing Amp and Kan resistance and IPTG inducer to obtain mice with full coverage positive PPI clones. 5) After obtaining bacteria with positive PPI clones, high-throughput PPI pair identification was performed using single-cell transcriptome sequencing based on random primers. A pair of open reading frames in a cell is a pair of protein interaction pairs, and a mouse full-coverage protein interaction network was constructed. Step 1) Mouse protein library construction specifically includes: firstly, extracting mouse mRNA, and using modified 3'RT Oligo-dT primers and TSO primers to obtain a mouse open reading frame library using Smart-seq 3; then, fusing the mouse open reading frame library to the N-terminus of a plasmid vector, first using Oligo(dT) magnetic beads to specifically capture eukaryotic mRNA via Poly-A tail, and then using modified 3'RT random primers for reverse transcription to obtain a cDNA library as the mouse protein library. The modified 3'RT Oligo-dT primer sequence is SEQ ID No. 4, the modified TSO primer sequences include SEQ ID Nos. 5-7, and the modified 3'RT random primer sequence is SEQ ID No.

1. Step 2) specifically involves homogenizing the mouse protein library using a double-stranded specific nuclease. Specifically, step 3) involves modifying the plasmid of the BACTH bacterial double hybrid system by linking it with a linker containing a homologous arm to form a modified plasmid. The linker sequence is SEQ ID No. 2, namely AAGCGTGGTATCAACGCAGAGT. The linker linker in the plasmid is located at the 3' position of the MSC multiple cloning site of the vector, and the linker site sequence is SEQ ID No. 3, namely ACTCTGCGTTGATACCACTGCTT. Step 4) specifically involves: First, mouse protein libraries were homologously fused with the C-terminus of pUT18C and pKT25 plasmids modified by the BACTH bacterial double-hybrid system to form pUT18C and pKT25 protein-particle libraries. Alternatively, open reading frame libraries obtained using random primers were homologously fused with the N-terminus of pUT18 and pKNT25 plasmids modified by the BACTH bacterial double-hybrid system to form pUT18 and pKNT25 protein-particle libraries. Subsequently, the plasmid libraries were electrotransduced into DHM1 Escherichia coli, and positive PPI clones were screened in M63 medium containing Amp and Kan resistance and IPTG inducer to obtain bacteria with positive PPI clones. The modified plasmid contains the linker sequence AAGCGTGGTATCAACGCAGAGT, and bacteria with positive PPI clones appear blue on the selection medium plate.

2. The method for constructing a full-coverage mouse protein interaction network using bacterial single-cell transcriptome sequencing according to claim 1, characterized in that: Step 5) specifically involves single-cell transcriptome sequencing based on random primers: First, bacteria with positive PPI clones were transferred to liquid culture medium for amplification and fixation with 4% PFA solution for 8-16 hours to allow protein-nucleic acid cross-linking. Then, bacterial cell membranes were permeabilized with PBS containing 0.04% PBS at -20°C, and lysozyme was used to perforate the E. coli cell walls. Next, in situ reverse transcription and tailing dA reactions were performed using reverse transcriptase and TdT terminal transferase. Then, oligo(dT) hydrogel-encoded microspheres were added to bacteria or yeast for single-cell encapsulation and cDNA double-strand amplification to obtain PCR libraries. The PCR libraries were then used to perform end repair and ligation of P5 / P7 adapters using the VAHTS Universal ProDNA DNA Library Preparation Kit to construct Illumina sequencing libraries. Finally, the obtained Illumina sequencing libraries were sequenced, and the sequences were compared with the reference genome of a mouse species to identify interacting protein pairs and construct an interacting protein network.