A method for analyzing the proteome of a microorganism based on pba data
By using PBA detection technology to analyze the proteome of microscopic biological structures, the problem of identifying early disease biomarkers in existing technologies has been solved. This enables high-throughput, particle-size-free proteome detection, improving the sensitivity and accuracy of early disease diagnosis.
Patent Information
- Application Number
- CN202010624704.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-01
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2040-07-01
AI Technical Summary
Existing technologies struggle to identify a small number of specific protein markers from diseased tissues with high sensitivity and high resolution, making early disease detection more difficult, especially in microscopic biological structures such as single-cell extracellular vesicles, where detection methods are limited by the size and number of fluorescent particles.
Using PBA detection technology, we analyze the raw PBA data of microbial structures, extract nucleotide sequences, obtain protein expression data, and perform various data analyses, including quantification of protein combinations and subpopulations. We also utilize machine learning algorithms for data mining and visualization.
It enables high-throughput, particle-size-free analysis of the proteome of microbial structures, allowing for the screening of disease-related biomarkers and improving the sensitivity and accuracy of early disease diagnosis.
Smart Images

Figure CN113889182B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of bioinformatics analysis of proteome, and particularly relates to a method for analyzing micro-biological structure proteome based on PBA data. BACKGROUND
[0002] With the increasing maturity and cost reduction of liquid biopsy technology and high-throughput sequencing technology, the field of bioinformatics has more solutions in clinical medicine. The central dogma reveals the mode of genetic information transmission, i.e. DNA-RNA-protein, and protein is a direct participant in many life activities. Therefore, directly focusing on the expression of proteome and performing true protein quantification has high scientific research and clinical value.
[0003] Currently, micro-biological structures mainly include but are not limited to extracellular vesicles, intracellular vesicles, subcellular structures, multi-protein complexes, viruses, virus-vesicle complexes, bacteria, and single cells. The proteome of micro-biological structures has outstanding value for the prevention and treatment of major diseases. For example, extracellular vesicles (EVs) derived from different cells and tissues have high heterogeneity in composition. However, the current research method of protein biomarkers is still the extraction of total EV proteins, and the analysis of proteins still relies on total protein analysis methods such as Western Blot, ELISA, and mass spectrometry. In the early stage of disease, a small amount of exosomes from diseased cells or tissues will be mixed with a large amount of normal body fluid EVs, and the differential expression of disease biomarkers will be averaged into the total signal, increasing the difficulty of detection. High sensitivity and high resolution identification of a small amount of specific markers from diseased tissues is the key to early detection of diseases. Therefore, analyzing micro-biological structures such as single EVs is a feasible way to achieve early diagnosis of diseases, and has important scientific significance and clinical application prospects.
[0004] To solve this problem, scientists have explored different technical methods to analyze micro-biological structures such as single EVs. For example, the nano-plasmonic exosome quantification analysis method (nPLEX) captures and analyzes exosomes of specific antigens, but has the disadvantage of low detection throughput. Nanoparticle tracking analysis based on fluorescent labeling can detect the accurate particle size and concentration of nanoparticles, but can only detect single-color fluorescence. Modified flow cytometry can achieve single EV analysis similar to single cell detection to achieve high-throughput detection of multiple factors. However, the particle size of exosomes is 30-150 nm, which has exceeded the detection limit of the optical system, and biological structures with a small particle size of 30-80 nm cannot be detected; the types of proteins detected are also limited by the number of fluorescent signals that can be detected simultaneously.
[0005] Therefore, it is necessary to provide new technical solutions to overcome the technical problems existing in the existing technology. Summary of the Invention
[0006] To address the problems existing in the prior art, this invention provides an analysis method for the proteome of microbial structures based on PBA data. Based on the data generated by PBA detection technology, the method analyzes the total protein expression, protein combination forms, microbial structure subpopulations and their fingerprint characteristics and quantification at the proteome level of microbial structures, filling the gap in bioinformatics analysis in the field of microbial structure proteome detection.
[0007] The first aspect of this invention provides a method for analyzing the microbial structural proteome based on PBA data, comprising the following steps:
[0008] Step 1: Input the raw PBA data of the microbial structure, and extract the nucleotide sequence from the raw PBA data to obtain the intermediate transfer data;
[0009] Step 2: Extract protein expression data of the microbial structures from the intermediate transmission data;
[0010] Step 3: Obtain one or more of the following from the protein expression data of the aforementioned microbial structures: proteome data, total protein expression data, and multi-protein co-expression analysis data;
[0011] Step 4: Perform data analysis on one or more of the proteome data, total protein expression data, and multi-protein co-expression analysis data obtained in Step 3.
[0012] The technical effect of this step is as follows: This technical step can convert raw PBA data into protein expression data and prepare for the subsequent acquisition of other data.
[0013] Preferably, in the analytical method described above, the nucleotide sequence in step one includes one or more of the following: a nucleotide sequence encoding an antibody-like biomacromolecule, a nucleotide sequence encoding identification information of the microbial structure, a nucleotide sequence encoding sample identification, a nucleotide sequence encoding the number of times each sequence in the original PBA data has been sequenced, or a correction nucleotide sequence. This technical step achieves the technical effect of bioinformatics transformation of the original PBA data, that is, extracting the biological information represented by the DNA fragment encoding.
[0014] Preferably, in the analytical method described above, the antibody-like biomolecules can specifically bind to proteins derived from the microbial structures. Through the antibody-like biomolecules, the recognition and detection of proteins from the microbial structures can be achieved.
[0015] Preferably, in the analytical method described above, the coding nucleotide sequence conjugated to the antibody-like biomolecule includes a nucleotide sequence containing biological information, which represents information about the protein specifically recognized by the antibody-like biomolecule. Through this technical step, the coding nucleotide sequence of the antibody-like biomolecule can be read using sequencing methods, thereby further enabling the detection of protein expression information of the microbial structure.
[0016] Preferably, in the aforementioned analytical method, the corrected nucleotide sequence is used for quality control of the PBA data and to confirm the location of different coding nucleotide sequences in the nucleotide sequence described in step one. This technical step achieves the technical effects of filtering out erroneous sequences generated during sequencing and library amplification in the raw PBA data, and accurately locating other coding nucleotide sequences.
[0017] Preferably, in the aforementioned analytical method, the encoding nucleotide sequence used to trace the sequencing count of each sequence in the original PBA data is a nucleotide sequence used to identify the uniqueness of the binding reaction of the antibody-like biomolecule to the protein of the microbial structure in the PBA binding reaction, and is used to identify the sequencing count of each sequence in the original PBA data, and to perform deduplication processing and trace the uniqueness of each sequence. This technical step enables the deduplication of repetitive sequences generated by the amplification of the Chinese library from the original PBA data.
[0018] Preferably, for the aforementioned analytical method, the method for acquiring proteomic data in step three is as follows: in the protein expression data of the microbial structures, at least m proteins are expressed simultaneously, and the expression level of each of the m proteins must be greater than or equal to n to be included in the proteomic data; where m is any value greater than or equal to 1, and the upper limit of m is the number of protein types to be detected; n is any value greater than or equal to 1. This technical step enables the extraction, filtering, and quality control of proteomic data of microbial structures. In some embodiments, m and n can be selected empirically.
[0019] Preferably, for the aforementioned analytical method, the method for obtaining the total protein expression data in step three is as follows: in the protein expression data of the microbial structures, only proteins with an expression total greater than or equal to z can be included in the total protein expression data; where z is any value greater than or equal to 1. This technical step enables the calculation, filtering, and quality control of the total protein expression in microbial structures. In some embodiments, z can be selected based on experience.
[0020] Preferably, for the aforementioned analytical method, the method for obtaining the multi-protein co-expression analysis data in step three is as follows: In the protein expression data of the microbial structure, only when v proteins are simultaneously expressed, and the number of times the v proteins are simultaneously expressed is greater than or equal to u, can the multi-protein co-expression situation formed by the v proteins be included in the multi-protein co-expression analysis data. Here, v is any value greater than or equal to 1, and u is any value greater than or equal to 1. This technical step enables the calculation, filtering, and quality control of multi-protein co-expression data at the level of a single microbial structure. In some embodiments, u and v can be selected based on experience.
[0021] Preferably, the analytical method involves performing data analysis on the proteomic data to obtain subpopulation data of the microbial structures. The data analysis method includes machine learning algorithm analysis, and the subpopulation data of the microbial structures includes one or more of the following: protein fingerprint feature data of the subpopulations and quantitative data of the subpopulations. This technical step enables the determination of microbial structure subpopulations, as well as the quantification and determination of protein fingerprint features of the subpopulations. Preferably, the data analysis of the subpopulation data of the microbial structures includes, but is not limited to, one or more of the following: difference analysis, correlation analysis, dimensionality reduction analysis, and machine learning algorithm analysis.
[0022] Preferably, the analytical method involves performing data analysis on the total protein expression data to output the biological information of the microbial structure; wherein the data analysis method includes one or more of differential analysis, enrichment analysis, dimensionality reduction analysis, or machine learning algorithm analysis, and the biological information of the microbial structure includes one or more of the following: total protein data, protein clustering, sample clustering, and protein fingerprint feature data of the sample. More preferably, the machine learning algorithm includes, but is not limited to, one or more of the following: clustering analysis algorithms, distance definition algorithms, hierarchical clustering algorithms, or correlation analysis calculation methods. The clustering analysis algorithms include, but are not limited to, hierarchical clustering and k-means. The distance definition algorithms include, but are not limited to, Euclidean distance, Manhattan distance, Chebyshev distance, and Jaccard similarity coefficient. The hierarchical clustering algorithms include, but are not limited to, single linkage clustering, complete linkage clustering, average linkage clustering, and centroid linkage clustering. The enrichment analysis includes, but is not limited to, GO enrichment analysis and KEGG analysis. The correlation analysis calculation methods include, but are not limited to, Pearson correlation coefficient and Spearman correlation coefficient.
[0023] The above-mentioned data analysis techniques can be used to analyze and mine data on microbial structures.
[0024] Preferably, for the aforementioned analysis method, step four further includes: displaying the multi-protein co-expression analysis data as a heatmap and outputting the biological information of the microscopic biological structure; wherein the data analysis method includes one or more of differential analysis, correlation analysis, dimensionality reduction analysis, or machine learning algorithm analysis. The dimensionality reduction analysis includes, but is not limited to, one or more of Principal Component Analysis (PCA), Multidimensional Scaling (MDS), Linear Discriminant Analysis (LDA), and Uniform Manifold Approximation and Projection.
[0025] The technical effect of this step is described as follows: This technical step enables the visualization of microbial structural data through charts and graphs, and enables the analysis and mining of multi-protein co-expression data of microbial structural data.
[0026] Preferably, the analytical method further includes step five: screening for biomarkers based on one or more of the proteomic data, total protein expression data, multi-protein co-expression analysis data, and the output results of the data analysis; wherein the output results of the data analysis include one or more of the aforementioned microbial structure subpopulation data, the aforementioned microbial structure biological information, or the aforementioned microbial structure biological information. The technical effect of this step is described as follows: This technical step can achieve the technical effect of biomarker screening, improving the performance of microbial structures in disease prediction, screening, diagnosis, prognosis, and efficacy assessment.
[0027] Preferably, in the analytical method described above, the proteins derived from the microbial structures include both surface and internal proteins of the microbial structures. This technical step allows for the selection and identification of research subjects at all expressible protein levels, i.e., the detection of specific protein types.
[0028] Preferably, in the aforementioned analytical method, the different coding nucleotide sequences include one or more of the following: coding nucleotide sequences coupled to biological macromolecules, coding nucleotide sequences identifying the identity information of the microbial structures, coding nucleotide sequences with sample identification capabilities, coding nucleotide sequences tracing the number of times each sequence in the original PBA data was sequenced, and correction nucleotide sequences. Through this technical step, biological information can be encoded into sequenceable nucleotide sequences, enabling high-throughput reading and analysis of the biological information.
[0029] The beneficial effects of this invention are as follows:
[0030] The analytical method described in this invention, based on data generated by PBA detection technology, analyzes the total protein expression, proteomic data of individual microbial structures (such as extracellular vesicles), protein combination patterns of individual microbial structures (such as extracellular vesicles), subpopulations of individual microbial structures (such as extracellular vesicles), and protein fingerprint characteristics of samples at the proteomic level of individual microbial structures (such as extracellular vesicles). It also quantifies the subpopulations of individual microbial structures (such as extracellular vesicles), filling a gap in bioinformatics analysis in the field of microbial structure proteomic detection. This analytical workflow helps researchers explore the biological characteristics of microbial structures and provides a novel solution for seeking new biomarkers in clinical applications. The advantages of the analytical method described in this invention, using PBA detection technology, include high sample throughput, small sample volume, and detection not limited by the number of protein types or the nanoscale particle size limitations of microbial structures. Therefore, it can achieve large-scale screening of biomarkers for disease-related microbial structures. Attached Figure Description
[0031] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This provides protein expression data for extracellular vesicles.
[0033] Figure 2 This refers to the co-expression and extraction process of proteins from extracellular vesicles.
[0034] Figure 3 This describes the raw data mining and analysis process for PBA. Detailed Implementation
[0035] Unless otherwise specified, experimental methods in the following examples are generally performed according to national standards. If no corresponding national standard exists, then generally accepted international standards, standard conditions, or conditions recommended by the manufacturer shall apply.
[0036] Features mentioned in this invention or in the embodiments may be combined. All features disclosed in this specification may be used in any compositional form, and each feature disclosed in the specification may be replaced by any alternative feature that provides the same, equivalent, or similar purpose. Therefore, unless otherwise specified, the disclosed features are merely general examples of equivalent or similar features.
[0037] In this invention, unless otherwise specified, all the technical features and preferred features mentioned herein can be combined to form new technical solutions.
[0038] Unless otherwise specified, the microbial structures mentioned herein include, but are not limited to: extracellular vesicles such as exosomes, microvesicles, and apoptotic bodies; intracellular vesicles such as multivesicular bodies and lysosomes; subcellular structures such as mitochondria and inclusion bodies; multiprotein complexes; viruses; virus-vesicle complexes; bacteria; and single cells.
[0039] In this invention, unless otherwise specified, the ev-tag mentioned herein refers to a nucleotide sequence obtained in PBA (proximity barcoding assay) detection technology by extending the nucleotide sequence coupled to an antibody-like biomolecule using a PBA template. This ev-tag is a specific base sequence that identifies a microbial structure. Proteins of the same microbial structure are labeled with the same specific base sequence ev-tag because they are geographically close; this sequence serves as the identifier for each microbial structure.
[0040] In this invention, unless otherwise specified, the PBA template referred to herein refers to a multi-copy repetitive sequence, where each copy includes a fixed DNA sequence identical to that of every PBA template and a PBA template-specific DNA sequence, i.e., the inverse complementary sequence of the ev-tag. Multiple proteins on each microscopic biological structure are recognized by their corresponding antibody-like biomolecules. Due to their proximity, the coding nucleotide sequence coupled to each antibody-like biomolecule can be complementary to one copy of the multi-copy repetitive sequence of the PBA template. This copy can serve as a template, causing the coding nucleotide sequence on the antibody to undergo an extension reaction, thereby obtaining the ev-tag.
[0041] In this invention, unless otherwise specified, the mol-tag mentioned herein refers to a nucleotide sequence that can indicate the number of times a nucleotide sequence in a PBA sequencing library has been copied in the PBA (proximity barcoding assay) detection technology.
[0042] In this invention, unless otherwise specified, the protein-tag mentioned herein refers to a nucleotide sequence that indicates the specific binding of a protein to a biomolecule that represents the properties of an antibody.
[0043] In this invention, unless otherwise specified, the term "sample-tag" refers to a nucleotide sequence that indicates the specificity of a sample.
[0044] In this invention, unless otherwise specified, p1-tag as used herein refers to a fixed universal DNA sequence. The nucleotide sequence can be any known DNA sequence, including but not limited to: ACCTGAGACATCATAATAGCA.
[0045] In this invention, unless otherwise specified, each sequence in the PBA raw data mentioned herein includes, but is not limited to, reads, where reads refer to the sequence of a DNA fragment.
[0046] In this invention, unless otherwise specified, filtering as mentioned herein refers to the process of deleting low-quality data from the data.
[0047] In this invention, unless otherwise specified, the coding nucleotide sequence of the biomacromolecules mentioned herein refers to any biomacromolecule capable of specifically recognizing the target protein, such as antibodies, nucleic acid aptamers, and peptide aptamers.
[0048] In this invention, unless otherwise specified, the PBA binding reaction mentioned herein refers to a single binding of an antibody-like biomacromolecule to a protein in a corresponding microbial structure, and a single binding can be specifically labeled using a mol-tag.
[0049] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention is further described below in conjunction with specific embodiments, but this invention includes, but is not limited to, these embodiments.
[0050] This invention primarily addresses the analysis of total protein expression and protein composition in microscopic biological structures, as well as the classification and fingerprinting of these structures based on proteomics. Utilizing the methods described in this invention, the invention starts with microscopic structures in organisms, such as exosomes, microvesicles, and apoptotic bodies, to explore their protein fingerprint characteristics, thereby achieving the goal of classifying microscopic biological structures. Further clinical biological significance lies in the ability to classify and subtype diseases for which medical diagnostic indicators are difficult to determine or for which there are currently no effective early diagnostic indicators by analyzing the total protein expression, protein composition, and subpopulations of microsecretions in patients' bodies, such as exosomes and apoptotic bodies.
[0051] In this invention, unless otherwise specified, the method for obtaining PBA raw data mentioned in this invention includes, but is not limited to: obtaining a sample containing microbial structures, performing a PBA experiment (proximity barcoding assay) using the method described in CN105745334A, obtaining a sequencing library composed of DNA amplicon, and obtaining PBA raw data, including sequencer data, DNA sequence data, and quality-controlled DNA sequence data.
[0052] In this invention, unless otherwise specified, the samples of microbial structures include, but are not limited to, extracellular vesicle samples such as exosomes, microvesicles, and apoptotic bodies; intracellular vesicles such as multivesicular bodies and lysosomes extracted from biological samples; subcellular structures such as mitochondria and inclusion bodies extracted from biological samples; multi-protein complexes extracted from biological samples; cultured viruses; viruses extracted from biological samples; virus-vesicle complexes extracted from biological samples; cultured bacteria; bacteria extracted from biological samples; extracellular vesicles of cultured bacteria; and prepared single cells. Liquid samples containing extracellular vesicles (EVs) include, but are not limited to, cell lines, liquid samples derived from humans or other animals, and include, but are not limited to, plasma, serum, cerebrospinal fluid, saliva, urine, tears, sweat, body fluid, bronchoalveolar lavage fluid, nasal lavage fluid, joint lavage fluid, interstitial EV extract, and fecal EV extract.
[0053] Proximity barcoding assay (PBA) enables high-throughput, multi-factor, single-EV level proteomics detection. PBA uses DNA fragments to encode antibodies that label proteins, allowing for an unlimited number of factors to be detected simultaneously. Proteins on the same EV, being relatively close, receive the same DNA encoding for that EV during the PBA binding reaction. Therefore, we can read all combinations of single EV encoding and antibody encoding through sequencing, achieving proteomic analysis of all individual EVs in a sample. Compared with other protein detection techniques, PBA has the following advantages: (1) single exosome resolution; (2) true omics detection, with an unlimited number of proteins that can be detected simultaneously; (3) high sensitivity, theoretically single-molecule sensitivity for protein detection; (4) high throughput, analysis of all single exosomes in a sample; (5) parallel operation of samples using 96-well plates; (6) direct capture of exosomes in the sample without purification; (7) not limited by exosome particle size.
[0054] PBA technology enables the detection of proteomic structures at the microscopic level, yielding novel data that cannot be processed using existing bioinformatics software. Therefore, PBA raw data mining has been developed to achieve deeper interpretation of biological information. PBA raw data mining includes, but is not limited to, […]. Figure 3 The analysis process for development and optimization is shown.
[0055] In this invention, unless otherwise specified, the PBA raw data mentioned in this invention includes, but is not limited to, sequencing data, DNA sequence files, and quality-controlled DNA sequence files.
[0056] In this invention, unless otherwise specified, the PBA raw data mentioned herein also includes, but is not limited to, antibody-like biomolecules, which can specifically bind to proteins of the microbial structures. The antibody-like biomolecules refer to any biomolecule capable of specifically recognizing the target protein, such as antibodies, nucleic acid aptamers, and peptide aptamers. The target protein includes, but is not limited to, the proteins of the microbial structures, including, but not limited to, surface and internal proteins of the microbial structures.
[0057] In this invention, unless otherwise specified, the p1 correction file mentioned in this invention refers to a corrected nucleotide sequence, used to control the quality of the PBA raw data and to confirm the location of different coding nucleotide sequences in the nucleotide sequence contained in the PBA raw data.
[0058] In this invention, unless otherwise specified, the quality control of PBA raw data mentioned in this invention includes, but is not limited to, filtering and cleaning of PBA raw data.
[0059] In this invention, unless otherwise specified, the DNA sequence files in the PBA raw data mentioned herein include, but are not limited to, ev-tags, mol-tags, and protein-tags. ev-tags include, but are not limited to, variable nucleotide sequences. An ev-tag sequence represents the coding nucleotide sequence for the identity of the microbial structure (including but not limited to EVs), and its length is greater than or equal to 1 bp with no upper limit. A mol-tag is a variable nucleotide sequence used for tracing the origin of a coding molecule, and its length is greater than or equal to 4 bp with no upper limit. A protein-tag is a variable nucleotide sequence used to label the identified protein information, and its length is greater than or equal to 1 bp with no upper limit. The PBA raw data includes, but is not limited to, p1-tags. A p1-tag is a nucleotide sequence fixed in a PBA binding reaction; the sequence itself has no coding meaning, and its length is greater than or equal to 1 bp with no upper limit.
[0060] In the raw PBA data, the same mol-tag is randomly copied during library amplification to identify a unique sequencing read and thus determine the actual sample detection volume. The protein name represented by the protein-tag is obtained by comparing it with the antibody library information table (the protein-tag column and the protein info column, which contains information about the protein name recognized by the antibody).
[0061] The antibody library information table is a design information table corresponding to the detection target range of the PBA detection method. It is suitable for PBA detection technology and includes two columns of control information: a protein-tag specific base sequence and a corresponding protein information column. The method for obtaining the antibody library information table is as follows:
[0062] 1. To perform covalent coupling of antibody-like biomolecules and DNA sequences.
[0063] 2. During antibody-DNA sequence conjugation, it is necessary to record in detail the antibody name, antibody manufacturer and catalog number, the name of the protein recognized by the antibody, the name of the DNA sequence conjugated by the antibody, and the protein-tag information contained in the DNA sequence conjugated by the antibody, ensuring a one-to-one correspondence. Specifically, the antibody library information table includes a one-to-one correspondence between the protein name recognized by the antibody (i.e., the protein info column) and the protein-tag information in the DNA sequence conjugated by the antibody (i.e., the protein tag column).
[0064] Example 1. A method for analyzing single-sample extracellular vesicle proximity coding data.
[0065] This embodiment provides a method for analyzing proximity coding data of single-sample extracellular vesicles, including the following steps:
[0066] Step 1: Obtain the standard FASTQ file of single-sample extracellular vesicle PBA data. Using the standard FASTQ file of single-sample extracellular vesicles as input, extract the triplet tags, including...<ev-tag,mol-tag,protein-tag> The intermediate data stream is in the form of a sparse matrix triplet. The effective extraction length of the standard fastq file is the first 75 bp. The output is an intermediate file containing the base sequences of ev-tag, protein-tag, molecular-tag and p1-tag, which is the intermediate data.
[0067] The single-sample extracellular vesicle PBA data is second-generation sequencing data. The full length of the second-generation sequencing data is 75bp, including ev-tag, mol-tag and protein-tag. The specific positions of the tags in the 75bp length are shown in Table 1.
[0068] Table 1. Location information of different tags in PBA data
[0069] tag type position (position in 1-75 bp) ev-tag 1-16 p1-tag 17-37 mol-tag 38-45 protein-tag 46-53
[0070] The method for extracting the triplet tags is as follows: the reads are split into strings according to Table 1 (the position information table of different tags in PBA data), and the base strings at the corresponding positions obtained from the splitting are the triplet tags.
[0071] Step 2: Based on the antibody library information table, further protein information is compared in the intermediate transmission data. Tag information is converted into protein information through a step-by-step comparison. In this step, the protein comparison incorporates the antibody library information table. In the antibody library information table, each protein-tag corresponds to a specific protein-antibody information, which is an experimental record of DNA fragment conjugation modification of the antibody, i.e., protein name - the sequence of the DNA fragment conjugated to the antibody.
[0072] Before obtaining the standard FASTQ file for single-sample extracellular vesicle PBA data, the reads data need to be filtered. During the filtering process, adapter sequences containing the p1 sequence were not removed. The retention of the p1 sequence further validated the intermediate transport data. The entire sample library has a consistent adapter sequence; mismatched or shifted adapter sequences were removed.
[0073] After step two, the p1 calibration file (Sample_PBA_test1.pri_anti.tmp.fastq file) is obtained.
[0074] Step 3: Perform reads filtering, using the p1 correction file from Step 2 as the input file for a third round of file filtering. The specific steps are as follows:
[0075] In PBA library construction, the DNA fragment undergoes multiple rounds of PCR amplification, resulting in multiple copies. Sequencing achieves higher sequencing depth; a higher copy number indicates higher reliability. Reads with low copy numbers (e.g., reads with a copy number of 1) are removed from the library, and duplicate reads are eliminated. (Note that mol-tags provide a unique identifier for each PBA binding reaction, preventing identical proteins on the same evanescent cell from being merged into a single count; instead, a single PBA expression count is obtained.) The PBA binding reaction refers to a single antibody targeting a protein in a corresponding microscopic biological structure, where the protein specifically binds to the antibody to form an extracellular vesicle protein-antibody complex. A single binding can be specifically identified using mol-tags.
[0076] After step three, protein expression data for all individual extracellular vesicles in the sample were obtained, such as... Figure 1 As shown.
[0077] For extracellular vesicle protein expression data, the expression matrix is sparse, meaning that only a few proteins are expressed for each EV (i.e., the expression value is not 0). The data stream is stored and transmitted in the form of a triplet compression of the sparse matrix, and the data representation is (ev-tag, protein-info, value).
[0078] Step 4: Input the extracellular vesicle protein expression data obtained in Step 3 to obtain a multi-protein co-expression matrix. By pre-constructing one-dimensional, two-dimensional, and three-dimensional arrays, and mapping target proteins to array indices, multiple statistical cases can be achieved with a single data traversal. Taking the co-expression of two proteins as an example, to avoid duplicate counting, the protein indices are sorted, resulting in an upper triangular two-dimensional array, ultimately yielding the co-expression status of extracellular vesicle proteins, i.e., the protein combination form, such as... Figure 2 As shown.
[0079] Step 5: Input the extracellular vesicle protein expression data obtained in Step 3, and calculate the multi-protein co-expression matrix and the overall protein expression of the sample under constraints. The multi-protein co-expression is constrained only if the current extracellular vesicle expresses only this protein combination.
[0080] extracellular vesicle ev-tags (P1, P2, P3, ..., P) n The co-expression of proteins is as follows: In the context of protein assemblages, i represents the number of protein types on the surface, e.g., i = 2 indicates a protein assemblage formed by two surface proteins; k is a constant, defaulting to k = 3; n represents the number of proteins with non-zero signal values on single-cell extracellular vesicles, p n This refers to the nth protein. Under conditions of restricted co-expression of extracellular vesicle proteins, the current protein co-expression can only be one of the following: (P1, P2, P3, ..., P...). n ), where ev-tag is an oligonucleotide sequence that specifically marks extracellular vesicles, P1, P2…P n This refers to the detected protein expression levels. For example, to calculate the co-expression of proteins (P1, P2), we need to count the frequency of this protein combination in the sample. The calculation method follows the principle described above: first, construct an empty array, with the rows and columns of the matrix corresponding to the proteins. This allows us to quickly find the position of the protein combination (P1, P2) in the matrix. When it appears once in the sample, the count at the corresponding position in the matrix is incremented by one, thus achieving the purpose of frequency counting. The same principle applies to protein combinations with one protein (P1) and three proteins (P1, P2, P3), constructing one-dimensional and three-dimensional arrays respectively.
[0081] Regarding the limiting conditions, for an extracellular vesicle protein expression data ev-tag (P1, P2, P3) that detects three proteins, the normal protein combination should be (P1), (P2), (P3), (P1, P2), (P1, P3), (P2, P3), (P1, P2, P3) – these are the seven possibilities. The limiting condition only considers the co-expressed protein as (P1, P2, P3).
[0082] Step Six: Input the proteomic data of single-cell extracellular vesicles obtained in Step Three, and filter the proteomic data of single-cell extracellular vesicles to form filtered proteomic data of single-cell extracellular vesicles. The filtering process is as follows:
[0083] The weighted evaluation value W of single-cell extracellular vesicle feature information is calculated as follows. v is the minimum screening threshold for each component of the single-cell extracellular vesicle feature information, v = 1, 2, 3;
[0084] The proteome of a single extracellular vesicle is defined as a feature information vector, where x iLet be the i-th component of the feature information vector, k be the dimension of the feature vector, and m be the number of components that satisfy the minimum component threshold of the feature information, m = 1, 2, 3; remove information vectors with a weight evaluation value of 0, sort the remaining information vectors according to their weight evaluation values, and then filter the data according to the step size step.
[0085] Step 7: Apply machine learning algorithms to classify and genotype the proteomic data of single extracellular vesicles. Dimensionality reduction is performed on the high-dimensional extracellular vesicle data to explore the composition ratio of each sub-component and the expression of characteristic proteins in the samples.
[0086] Dimensionality reduction is performed using proteomic data of single extracellular vesicles as input. Dimensionality reduction algorithms include, but are not limited to, Principal Component Analysis (PCA), Factor Analysis, Linear Discriminant Analysis (LDA), Multidimensional Scaling (MDS), SNE algorithm, t-SNE (t-Distributed Stochastic Neighbor Embedding) algorithm, and UMAP (Uniform Manifold Approximation and Projection) algorithm. In this example, the t-SNE algorithm and the UMAP algorithm are used in conjunction to visualize the proteomic data of extracellular vesicles.
[0087] The specific function formulas for the t-SNE algorithm and the UMAP algorithm are as follows:
[0088] Data point x in high-dimensional space i and x j x i With conditional probability p j|i Select x j As its nearest point. j|i That is, at data point x i Given the existence of a data point x, there exists a given condition. j The conditional probability when x j The closer to x i conditional probability p j|i The larger it is. Consider x. i p is calculated using a Gaussian distribution centered at the point. j|i , in, This represents the variance of the Gaussian distribution, where k is the data point x. i Data points other than ||x i -xj ||For x i With x j The measure between, ||x i -x j || is the square root of the sum of squares of the components after subtracting the vector.
[0089] When data points in a high-dimensional space are mapped to a low-dimensional space, data point x i and x j The corresponding point in the low-dimensional space is y. i and y j Its conditional probability q j|i The calculation formula is as follows: In low-dimensional space, the variance of the Gaussian fraction is set as... It facilitates formula calculations.
[0090] To measure the similarity between the distribution of data points in high-dimensional space and the distribution of data points in low-dimensional space, the KL (Kullback-Leibler Divergence) distance is used, and the loss function C can be constructed.
[0091]
[0092] In practical applications, it has been found that the loss function cannot be optimized to the minimum value well, and always oscillates around the optimal solution. At this time, the momentum parameter is introduced, that is, the momentum gradient descent method.
[0093] The new momentum gradient change is:
[0094]
[0095]
[0096] in, The new gradient is the current one. Compared with the previous gradient The weighted summation, where the weight β ranges from 0 to 1; for the loss function C, with a random initial point x0, its gradient is... The gradient of the second iteration point x1 is: The gradient of the second iteration point x2 is: This process is iterated until a solution to the loss function C is found, yielding the final mapping of the data in a low-dimensional space. This includes the previously obtained momentum information. A β value of 0.9 is generally preferred. The problem can then be transformed into solving for the loss function using the gradient descent algorithm.
[0097] The t-SNE algorithm employs a more general joint probability distribution p. ji With q ijp ji =p ij q ji =q ij .
[0098] When the amount of data is large, the computational load is undoubtedly very large. For points that are far apart, the p... ji It is usually smaller, so just consider x i To determine the conditional probabilities of several points around the center, we introduce the parameter perplexity, which is crucial in the t-SNE algorithm. The probability distribution in the high-dimensional space can then be represented as follows:
[0099]
[0100] Where n is the total number of data points. N i For x i Let be the set of points under the perplexity of the center point. Through the above optimization, not only is the symmetry problem solved, but the computational cost is also reduced. In the low-dimensional space, to solve the crowding problem, the Gaussian distribution is replaced with a t-distribution with a higher tail; when the degrees of freedom are 1, it becomes the Cauchy distribution. At this point, the probability distribution in the low-dimensional space is:
[0101] Specific machine learning algorithms include, but are not limited to, K-Nearest Neighbor (KNN), Random Forest, Linear Regression, Logistic Regression, K-means, Gradient Descent, and Unsupervised Neural Networks. Loss functions include, but are not limited to, Mean Squared Error, Maximum Likelihood (MLA) for logarithmic loss, and Cross-Entropy. Distance definition algorithms include, but are not limited to, Euclidean Distance, Manhattan Distance, Chebyshev Distance, Jaccard Similarity Coefficient, Pearson Correlation Coefficient, and Cosine Similarity. (The above algorithms can be combined to form new technical solutions).
[0102] This example uses a combination of the Self-Organizing Map (SOM) algorithm and a consistency coefficient to classify and genotype single extracellular vesicles. A SOM network is constructed by inputting proteomic data of single extracellular vesicles. A predefined 10*10 two-dimensional matrix is used as the network output layer to train an unsupervised neural network. In the SOM network, the weight coefficient vector w of each output node... jThe numbers j = 1, 2, 3, ..., 100 form the probability structure of this input pattern space, and the weight coefficient vector w j The optimal reference vector for the input pattern. The weight coefficient vector w obtained from the self-group mapping network. j As input, the consistency similarity coefficient is calculated for different numbers of classes. By comparing the difference between the cumulative distribution function (CDF) of the consistency similarity coefficient and the area change enclosed by the x-axis for different numbers of classes, the optimal number of classes can be determined when the difference gradually flattens out. This allows for the determination of the classification of single-cell extracellular vesicles and the optimal number of classes. For the specific function formulas of the Self-Organizing Map (SOM) algorithm and the consistency coefficient, existing methods can be referenced, including but not limited to those described in Vikas Chaudhary, RS Bhatia, Anil K. Ahlawat. A novel Self-Organizing Map (SOM) learning algorithm with nearest and farthest neurons. Alexandria Engineering. Journal, Volume 53, Issue 4, December 2014, 827-831.
[0103] Step 8: Based on the clustering results of Step 7, explore the protein fingerprint characteristics of extracellular vesicle subpopulations to identify extracellular vesicle subpopulations with biomarker potential and their protein fingerprint characteristics.
[0104] Example 2. A method for analyzing multi-sample exosome data using proximity coding.
[0105] This embodiment describes a method for data analysis based on multiple PBA test samples. The general analysis process is as follows: Figure 3 As shown. The data for each sample comes from the figshare website, the specific link is [link to figshare website]. https: / / figshare.com / articles / FASTQ_files / 7956023 The sample data for this embodiment was obtained from figshare.
[0106] Step 1: Using the method described in Example 1, the exosome characteristics, characteristic spectrum, multi-protein co-expression, and protein expression of each sample are statistically analyzed.
[0107] Step 2: Merge the statistical results of each sample, specifically including the surface feature spectrum of a single exosome, the protein combination data of multi-protein co-expression of a single exosome, the sample protein expression matrix, and the sample quality control results, to obtain the output file data. The output file data includes the protein co-expression matrix (contain), the protein co-expression matrix under specified conditions (uniq), the sample protein expression matrix, the PBA quality control file, and the exosome proteome data.
[0108] Step 3: PBA sample control. Using the PBA quality control file as input text, calculate the total sequencing read throughput, the number of reads successfully aligned with the antibody library information table, the sample protein detection volume (unique reads: this will differ from the total sample protein calculated from the protein expression matrix. The total sample protein obtained by summing the protein expression matrix may be slightly smaller than the unique reads because the protein expression matrix performs threshold screening for proteins with low expression levels, which is 4 by default), and the sample ev detection throughput (ev count).
[0109] Step 4: Protein Expression Analysis. This analysis is based on the expression data matrix of each sample, and performs cluster analysis on the protein expression of the samples. This allows for the classification of samples and proteins. Through cluster analysis, the similarity of protein expression patterns among different samples can be revealed; proteins clustered in the same group may have similar biological functions.
[0110] The inputs are protein expression matrices. The complete algorithm and Euclidean distance in hierarchical clustering are used to obtain the similarity of protein expression patterns among different samples. Proteins clustered in the same cluster are considered to have similar biological functions.
[0111] Step 5: Protein Combination Analysis (MDS Multidimensional Scaling Analysis) MDS multidimensional scaling analysis is performed on the protein co-expression matrix under different conditions to explore the reproducibility of experimental procedures on the samples. Compared to t-SNE, which projects data based on minimizing the change in probability distribution, MDS uses distance as the standard. Its principle is to utilize the similarity between paired samples to construct a suitable low-dimensional space, ensuring that the distance between samples in this space is as consistent as possible with the similarity between samples in the high-dimensional space. The algorithm is as follows:
[0112] The protein co-expression matrix is set as follows: Where n is the number of samples, and d are the dimensions:
[0113] The distance between paired samples i and j is dist ij (The distances here are measured using the Euclidean distance system.)
[0114] Let Z be the matrix after dimensionality reduction. Generally, when the dimensionality is reduced to two dimensions, the size of the Z matrix is n*2; if it is three dimensions, it is n*3.
[0115] Now we obtain the inner product matrix B of matrix Z: B = Z. T Z,
[0116] The optimization objective is to maintain the same similarity between high-dimensional and low-dimensional samples as much as possible, and to make the similarity of paired samples in high-dimensional samples as equal as possible to that in low-dimensional samples. This yields the following:
[0117]
[0118] The solution approach is to obtain b from the distance matrix of the samples. ij That is, if we obtain the inner product moment B, we can achieve the dimensionality reduction goal by orthogonally diagonalizing the sample to obtain the matrix Z in the low-dimensional space.
[0119] In addition to using inner products to find low-dimensional mappings as described above, MDS algorithms can also be implemented by optimizing the objective function, such as t-SNE and UMAP, to solve for low-dimensional mappings.
[0120] In experiments with grouped designs, this method can verify the repeatability of samples within the same group.
[0121] In the command line, after creating a new output folder, it is used as an input parameter. Finally, the output path will output the multi-protein co-expression protein combination matrix under different conditions and the sample MDS clustering results.
[0122] Step Six: In the multi-sample analysis, machine learning methods are applied to classify and genotype the single exosome proteome data. Dimensionality reduction is performed on the high-dimensional exosome data to explore the composition ratio and protein fingerprint characteristics of each subpopulation within the samples.
[0123] Dimensionality reduction was performed using exosomal proteome data as input (the algorithm is the same as step seven in Example 1).
[0124] Step 7: Exploring protein fingerprint characteristics in exosome subpopulations. This step is to identify subpopulations with biomarker potential and their protein fingerprint characteristics (the algorithm is the same as step 8 in Example 1).
[0125] Example 3. PBA Methodology Analysis of Group Sample Exosome Data
[0126] The PBA methodology for analyzing exosome data from group samples adds a comparative analysis of sample groups to the multi-sample data analysis, using difference tests to examine the significance of differences between different sample groups. Its preliminary analytical steps are consistent with those of the multi-sample analysis.
[0127] This example uses blood samples from 9 patients with acute myeloid leukemia (AML) and 23 healthy individuals as comparative and control groups to demonstrate the application of the PBA methodology in group sample analysis. The specific operation method includes the following steps:
[0128] Step 1: Statistically analyze the extracellular vesicle protein expression data, single-cell extracellular vesicle proteome data, total protein expression data, and multi-protein co-expression data for each sample.
[0129] Step 2: Combine the statistical results of each sample, including single-cell extracellular vesicle proteome data, total protein expression data, multi-protein co-expression data, and sample quality control results.
[0130] Step 3: PBA sample quality control, which involves statistical analysis of PBA test data at each stage.
[0131] Step 4: Protein Expression Cluster Analysis. This analysis clusters protein expression patterns based on the total protein expression levels of each sample. It allows for the clustering of samples and proteins. Cluster analysis reveals the similarity of protein expression patterns among different samples; proteins clustered in the same group may have similar biological functions.
[0132] Step 5: Protein Co-expression Analysis (MDS Multidimensional Scaling Analysis) MDS multidimensional scaling analysis was performed on protein combinations under different conditions. Compared to t-SNE, which projects data based on minimizing changes in probability distribution, MDS uses distance as the standard. Its principle is to utilize the similarity between paired samples to construct a suitable low-dimensional space, ensuring that the distance between samples in this space is as consistent as possible with the similarity between samples in the high-dimensional space.
[0133] Step Six: Apply machine learning methods to genotype the proteomic data of single-cell extracellular vesicles. Dimensionality reduction is performed on the high-dimensional data to explore the composition ratio and protein fingerprint characteristics of each subpopulation in the sample.
[0134] Step 7: Investigation of protein fingerprint characteristics in extracellular vesicle subpopulations. This step is to identify subpopulations with biomarker potential and their protein fingerprint characteristics.
[0135] Step 8: Preparation for Differential Analysis. The three expression files—proteomic data of single extracellular vesicles, assemblage data of single extracellular vesicles, and total protein expression matrix—are integrated into one file to prepare for the next step of differential analysis.
[0136] The integrated file containing proteomic data of single extracellular vesicles, combined data of single extracellular vesicles, and total protein expression matrix of samples was used as the input file for differential analysis.
[0137] Step Nine: Difference Analysis. Difference analysis uses statistical tests to identify genes that produce differences between two groups of samples and to determine the significance of these differences. Based on quantitative gene information, differentially expressed genes can be screened, selecting functional genes under different experimental conditions. It is recommended to use 3–8 replicate samples; samples with fewer than 3 replicates should not be screened. Two differentially expressed gene screening methods were selected for different data, as detailed in Table 2 below.
[0138] Table 2. Screening methods for differentially expressed genes
[0139]
[0140] The exact test provided by edgeR is a difference test method based on the negative binomial distribution model. It is similar to the Fisher precision test, which is commonly used in enrichment analysis (i.e., whether the differentially expressed genes are significantly enriched in a certain pathway).
[0141] For the first test method, the Exact test, the analysis procedure is as follows:
[0142] 1. Obtain the data matrix and sample grouping
[0143] edgeR is a model-based standardization, so no other transformations are needed when inputting the data, such as rpkm, fpkm, or cpm values; the original matrix can be input directly.
[0144] 2. Filtering genes with low expression levels
[0145] To filter genes with low expression levels across all samples, you can define your own filtering criteria.
[0146] 3. Standardize the data using the TMM method.
[0147] The `calcNormFactors` function was used to standardize the data. TMM standardization is a method for standardizing the components. A standardization factor is calculated to standardize the library size. When the standardization factor is less than 1, it indicates that a small number of highly expressed genes in the library are monopolizing sequencing. Standardization is achieved by reducing the library size, as shown below:
[0148] The TMM standardization method is briefly described below: 1. After data filtering, the gene values of each sample are normalized by correcting for the total number of reads. 2. Then, the quartiles at the 75th percentile are found, the mean is calculated, and the sample whose quartile is closest to this mean is used as the reference sample (ref). 3. For each sample, a standardization factor is calculated: discrete genes are determined using log2(ref / sample) of each gene, and the geometric mean is used. After determining the transcription levels, select genes from the intermediate subset. Then calculate the weighted average of the remaining genes, which is obtained by multiplying the log2(ref / sample) value by the count value and dividing by the sum of the counts. The standardization factor is: scalefactor = 2. weightaverage 4. The final standardized factor can be obtained by centering the standardized factors of all samples.
[0149] 4. Calculate the dispersion
[0150] We fit our expression data using a negative binomial distribution. For sample i and gene g, the number of detection reads for sample i and gene g is y. gi Distribution model: Y gi ~NB(μ gi ,φ) where the expectation is: E(Y gi )=μ gi Variance: Var(Y) gi )=μ gi (1+μ gi φ)
[0151] μ gi =m i λ gj Where m i λ is the size of the library. gj This represents the true relative abundance of gene g in group j.
[0152] φ represents the dispersion, specifically the biological coefficient of variation (BCV), which reflects the differences inherent in the biological sample itself, contrasting with differences introduced by experimental techniques. Ideally, all genes share a common dispersion, but in reality, different genes may have different dispersions, necessitating the calculation of gene-specific tagwise dispersions. (See references: edgeR: a Bioconductor package for differential expression analysis of digital gene expression data; Moderated statistical tests for assessing differences in tag abundance).
[0153] 5. Conduct accuracy checks
[0154] Once the model and dispersion are ready, significance testing can be performed. edgeR provides an exact test similar to Fisher's exact test for significance testing (see: Small-sample estimation of negative binomial dispersion, with applications to SAGE data).
[0155] The second test, the T-test, is relatively simple and can be applied to protein combination data (protein combinations formed by two or three surface proteins) and microbial structural subpopulation data. It is also known as the t-test for grouped data.
[0156] Shapiro-Wilk test: also known as W test, mainly tests whether the research objects conform to a normal distribution;
[0157] F-test: Tests the homogeneity of variance between two groups of data. The prerequisite for the F-test is that both groups are normal populations.
[0158] Student's t-test: This is a grouped samples t-test used to examine whether the difference between the means of two samples and the populations they represent is significant. Its prerequisites are that the samples follow a normal distribution and the population variances are homogeneous.
[0159] Welch's t test: When population variances are not equal, the Welch t test, also known as the Aspin-Welch test, is used, and its degrees of freedom are adjusted.
[0160] Wilcoxon rank-sum test: The above two tests are parametric tests, while the Wilcoxon rank-sum test...
[0161] These are nonparametric tests. The difference between parametric and sample tests lies in the assumption that the distribution of the population is known, using population data and sample information to infer the overall situation. Nonparametric tests, on the other hand, cannot make necessary assumptions about the population distribution. The Wilcoxon rank-sum test is a method for testing whether the distributions of two populations are in the same location.
[0162] Inspection process:
[0163] 1. Perform the Shapiro-Wilk test on the two samples to determine whether they belong to a normal distribution. If they match, use the non-parametric test - the Wilcoxon rank-sum test.
[0164] 2. Under the premise that the normality assumption holds, perform homogeneity of variance analysis. If the variances are homogeneous, perform Student's t test. If the variances are not homogeneous, perform Welch's t test after degrees of freedom adjustment to obtain the significance p-value.
[0165] 3. Correct the p-values obtained from multiple hypothesis tests; the method used is the BH method.
[0166] Step 10: After identifying significant differences in expression, proceed with the subsequent chart expansion.
[0167] Example 4. A method for analyzing proximity coding technology data of microbial structures.
[0168] This embodiment provides a method for analyzing proximity coding technology data of microbial structures. The difference from Embodiment 1 is that the single-sample extracellular vesicle PBA data is replaced with microvesicle PBA data. The microvesicles are extracted according to the method described in the research literature of Lee, J. et al. (Lee, J. et al. Microvesicles from brain-extract-treated mesenchymal stem cells improve neurological functions in a rat model of ischemic stroke. Sci Rep 6, 33038 (2016).).
[0169] Example 5. A method for analyzing proximity coding technology data of microbial structures.
[0170] This embodiment provides a method for analyzing proximity coding technology data of microbial structures. The difference from Embodiment 1 is that the single-sample extracellular vesicle PBA data is replaced with PBA data of apoptotic bodies. The apoptotic bodies are extracted according to the method described in the research literature of Serrano-Heras, G. et al. (Serrano-Heras, G. et al., Isolation and quantification of blood apoptotic bodies in neurological patients. bioRxiv preprint, doi:https: / / doi.org / 10.1101 / 757872.).
[0171] Example 6. A method for analyzing proximity coding technology data of microbial structures.
[0172] This embodiment provides a method for analyzing proximity coding (PBA) data of microbial structures. The difference from Embodiment 1 is that the single-sample extracellular vesicle PBA data is replaced with mitochondrial PBA data, and the mitochondrial data is extracted using a kit (Thermo Scientific). TM Mitochondria Isolation Kit for Cultured Cells (Item No. 89874)
[0173] Example 7. A method for analyzing proximity coding technology data of microbial structures.
[0174] This embodiment provides an analysis method for proximity coding technology data of microbial structures. The difference from Embodiment 1 is that the single-sample exosome PBA data is replaced with inclusion body PBA data. The inclusion bodies are extracted according to the method described in the research literature of Hornick, CA et al. (Hornick, CA et al. Isolation and characterization of multivesicular bodies from rat hepatocytes: an organelle distinct from secretory vesicles of the Golgi apparatus. J Cell Biol. 1985 May 1; 100(5): 1558–1569.).
[0175] Example 8. A method for analyzing proximity coding technology data of microbial structures.
[0176] This embodiment provides a method for analyzing proximity coding technology data of microbial structures. The difference from Embodiment 1 is that the single-sample exosome PBA data is replaced with PBA data of a multi-protein complex, which is obtained from a multi-protein complex extracted by magnetic bead immunoprecipitation.
[0177] Example 9. A method for analyzing proximity coding technology data of microbial structures.
[0178] This embodiment provides a method for analyzing proximity coding technology data of microbial structures. The difference from Embodiment 1 is that the single-sample exosome PBA data is replaced with viral PBA data, and the virus is a bacteriophage culture medium.
[0179] Example 10. A method for analyzing proximity coding technology data of microbial structures.
[0180] This embodiment provides a method for analyzing proximity coding technology data of microbial structures. The difference from Embodiment 1 is that the single-sample exosome PBA data is replaced with PBA data of COVID19 virus-vesicle complex, which is obtained from the sputum of COVID19 infected individuals (20 infected individuals and 24 healthy controls).
[0181] Example 11. A method for analyzing proximity coding technology data of microbial structures.
[0182] This embodiment provides a method for analyzing proximity coding technology data of microbial structures. The difference from Embodiment 1 is that the single-sample exosome PBA data is replaced with PBA data of Pseudomonas aeruginosa bacteria, wherein the Pseudomonas aeruginosa bacteria is a culture medium of the strain.
[0183] Example 12. A method for analyzing proximity coding technology data of microbial structures.
[0184] This embodiment provides a method for analyzing proximity coding technology data of microbial structures. The difference from Embodiment 1 is that the single-sample exosome PBA data is replaced with ACHN, HEK293, caki-1, 769-P, 786-O single cells, and the cells are adherent cultured cells.
[0185] In this invention, the diseases for which the analytical methods described in any of the foregoing embodiments can be applied include, but are not limited to, cardiovascular and cerebrovascular diseases, ovarian cancer, lung cancer, and renal fibrosis.
[0186] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A method for analyzing the microbial structural proteome based on PBA data, characterized in that, Includes the following steps: Step 1: Input the raw PBA data for a single microbial structure, which is selected from cell-generated vesicles, protein complexes or aggregates, viruses or single cells; extract intermediate transfer data from the raw data containing at least the following nucleotide sequences: sequences conjugated to antibody-like biomolecules, ev-tag sequences for uniquely identifying a single microbial structure, sample identification sequences, mol-tag sequences for tracing sequencing counts and removing repetitive sequences, and correction sequences for correcting sequence positions; Step 2: Based on the antibody library information table, the nucleotide sequence information in the intermediate transmission data is converted into protein expression data of the individual micro-biological structure, forming associated data in ev-tag, protein-info, and value formats; Step 3: Obtain core feature data from the protein expression data, including: (1) Proteome data: satisfying that at least m proteins are simultaneously expressed on a single microbial structure, and the expression level of each protein is ≥ n, where m ≥ 1 and n ≥ 1; (2) Total protein expression data: The total expression of any protein in a single microorganism in this sample is ≥z, where z≥1; (3) Multi-protein co-expression data: satisfying that v proteins are simultaneously expressed on a single microbial structure, and the number of co-expressions is ≥u, where v≥1 and u≥1; Step 4: Analyze the core feature data: Use hierarchical clustering, k-means, or self-organizing map (SOM) algorithms to cluster the proteome data, divide it into microbial structural subgroups, and extract the specific protein fingerprint features and subgroup proportions of each subgroup; perform enrichment analysis on the total protein expression data; and perform correlation analysis and network visualization on the multi-protein co-expression data. Step 5: Based on the analysis results of Step 4, screen for biomarkers related to microbial structures.
2. The method for analyzing the microbial structural proteome based on PBA data according to claim 1, characterized in that, The antibody biomolecules in step one are selected from antibodies or nucleic acid aptamers, and the antibody library information table records the one-to-one correspondence between protein tags and corresponding proteins.
3. The method for analyzing the microbial structural proteome based on PBA data according to claim 1, characterized in that, In step three, the upper limit of m is the total number of protein types detected, and the upper limit of v is the total number of protein types detected.
4. The method according to claim 1, characterized in that, Step four also includes merging and analyzing the core feature data of multiple samples, and using t-tests or edgeR exact tests to test cross-sample differences.
5. The method for analyzing the microbial structural proteome based on PBA data according to claim 1, characterized in that, The biomarkers in step five are selected from specific protein combinations, microbial structural subgroups, or their protein fingerprint characteristics.
6. The method for analyzing the microbial structural proteome based on PBA data according to claim 1, characterized in that, Step four, the analysis of proteomic data, also includes dimensionality reduction analysis using t-SNE, UMAP, or PCA.
Citation Information
Patent Citations
Methods to profile molecular complexes via proximity dependant barcoding
CN105745334A
Analysis method for researching formation of variegation of nelumbo nucifera based on combination of transcriptomics with proteomics TMT
CN110331225A
Systems and Methods for the Analysis of Proximity Binding Assay Data
US20130304390A1