A method for diagnosing prostate cancer using exosome surface protein proteomics
By combining dual-marker orthogonal barcode capture and chemical glycoproteomics with high-resolution mass spectrometry, the problems of exosome specificity and surface proteome analysis in prostate cancer diagnosis have been solved, realizing a highly specific and sensitive diagnostic tool suitable for the precise diagnosis of prostate cancer.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-04-27
- Publication Date
- 2026-05-26
AI Technical Summary
Current technologies lack a unified and efficient method that can organically combine the high-specificity capture of prostate cancer-derived exosomes with high-throughput, panoramic analysis of their surface proteome, resulting in insufficient specificity and sensitivity in prostate cancer diagnosis.
We employed a dual-marker orthogonal barcode capture probe targeting CD63 protein and prostate-specific membrane antigen, combined with chemical glycoproteomics and high-resolution mass spectrometry, to achieve panoramic analysis of exosome surface proteins through in-situ magnetic bead processing, and used machine learning algorithms to construct a diagnostic model.
It achieves high-purity enrichment of prostate cancer exosomes and panoramic analysis of their surface proteome, significantly improving diagnostic specificity and sensitivity, reducing biological background noise, and enabling the discovery of new diagnostic biomarkers, making it suitable for large-scale clinical screening.
Smart Images

Figure CN122084901A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical detection technology, and in particular to a method for diagnosing prostate cancer using exosome surface protein spectrometry. Background Technology
[0002] Prostate cancer (PCa) is the second most common malignant tumor among men worldwide, and its early and accurate diagnosis is crucial for improving patient survival and quality of life. Currently, the clinical practice commonly uses serum prostate-specific antigen (PSA) testing combined with digital rectal examination for initial screening, but this method has significant limitations. PSA has low specificity as a biomarker, and its levels can be elevated in non-cancerous lesions such as benign prostatic hyperplasia and prostatitis. This leads to up to 65%-70% of unnecessary prostate biopsies, causing psychological and economic burdens on patients. Although multiparametric magnetic resonance imaging (mpMRI) and other techniques are used to assist in decision-making, it is still difficult to completely avoid overdiagnosis of indolent cancers and missed diagnosis of clinically significant cancers. Therefore, developing a highly specific, highly sensitive, and non-invasive diagnostic tool is a critical issue that urgently needs to be addressed in clinical practice.
[0003] Liquid biopsy, particularly analysis based on extracellular vesicles (EVs, including exosomes), offers a promising solution to the aforementioned problems. Exosomes are nanoscale (typically 30-200 nm) lipid bilayer vesicles actively secreted by cells and are widely distributed in bodily fluids such as blood and urine. They carry and protect various biomolecular information (such as proteins and nucleic acids) from their source cells, acting as "cellular messengers" and reflecting the physiological and pathological state of their parent cells. Compared to free circulating biomarkers, exosomes have the unique advantages of high stability and rich information content. Notably, proteins on the surface of exosomes, especially glycosylated transmembrane proteins or membrane-anchored proteins, not only directly participate in key biological processes such as exosome biogenesis and targeted delivery, but also become optimal targets for highly specific capture and recognition due to their direct exposure to the external environment. Existing studies have shown that over 90% of cell surface proteins are glycosylated, a characteristic that likely extends to the surface of exosomes.
[0004] However, translating exosome surface protein profiles into reliable clinical diagnostic tools faces multiple technical challenges, constituting the shortcomings of existing technologies: First, low separation specificity. Conventional exosome separation methods (such as ultracentrifugation and polymer precipitation) rely primarily on physical properties (size, density), failing to distinguish exosomes derived from prostate cancer from "background noise" exosomes from other cells or tissues, resulting in diluted target signals and limited detection sensitivity and specificity. Second, incomplete surface protein analysis. Traditional immunological methods (such as ELISA and Western blotting) typically only detect a few known protein markers (such as PSMA and EpCAM), employing a "hypothesis-driven" strategy that cannot comprehensively discover and identify unknown, diagnostically valuable surface proteins, especially low-abundance proteins. Third, cumbersome and low-throughput detection procedures. Analysis of surface proteins and internal nucleic acids (such as miRNAs) often requires separate and complex experimental procedures (e.g., capture followed by lysis, and separate proteomic and transcriptomic analyses), involving numerous steps, long processing times, and large sample requirements, making it difficult to meet the needs of large-scale clinical sample screening.
[0005] To address these shortcomings, some improvements have been attempted in existing technologies. For example, some studies have used magnetic beads modified with mixed antibodies (such as CD9, CD63, and CD81) for affinity capture, improving exosome enrichment efficiency; other studies have attempted to simultaneously detect exosome surface proteins and internal miRNAs, but have failed to fundamentally solve the problems of capture specificity and comprehensive coverage of the surface proteome. Other studies have proposed using two aptamers (targeting CD63 and EpCAM respectively) for orthogonal barcoding of tumor-derived exosomes, improving sorting specificity. A team at Southeast University developed a chemical glycoproteomics method called EVscope, which achieves efficient panoramic analysis of the urinary exosome surface proteome by oxidatively labeling the glycan chains of exosome surface glycoproteins, and discovered new prostate cancer biomarkers. However, this method still lacks strong specificity screening for prostate cancer-derived exosomes in the initial capture stage; its analysis still only covers the total exosome population in urine.
[0006] Therefore, existing technologies lack a unified and efficient method that organically combines the highly specific capture of exosomes derived from prostate cancer with high-throughput, panoramic analysis of their surface proteome (especially glycoproteins). This invention aims to solve this technical problem. Summary of the Invention
[0007] To achieve the above objectives, the present invention provides a method for diagnosing prostate cancer using exosome surface protein profiling, comprising the following steps: Step 1: Collect biological fluid samples from the subjects and preprocess them to remove impurities and obtain clear body fluid samples; Step 2: Prepare dual-marker orthogonal barcode capture probes. The first capture probe connects the first DNA aptamer targeting CD63 protein to the first magnetic bead via the first DNA barcode sequence. The second capture probe connects the second DNA aptamer targeting prostate-specific membrane antigen to the second magnetic bead via the second DNA barcode sequence. The first DNA barcode sequence and the second DNA barcode sequence do not hybridize. Step 3: Incubate the first and second capture probes together with a clear body fluid sample, separate them using a magnetic field and wash them to obtain a prostate cancer-derived exosome-magnetic bead complex that simultaneously binds the first and second capture probes. Step 4: Perform in-situ chemical glycoproteomics processing on the exosome-magnetic bead complex derived from prostate cancer, specifically including: oxidizing the glycan chains of the surface glycoproteins of prostate cancer-derived exosomes and converting them into aldehyde groups, labeling the aldehyde groups with biotinylate, cleaving the exosomes and enzymatically digesting them into a peptide mixture, and using streptavidin magnetic beads to specifically enrich the surface glycoprotein peptides carrying biotin labels from the peptide mixture; Step 5: Perform liquid chromatography-tandem mass spectrometry analysis on surface glycoprotein peptides to identify and quantify surface glycoproteins and generate surface glycoprotein spectrum data; Step 6: Input the surface glycoprotein spectrum data into the pre-trained diagnostic model, which is trained based on the surface glycoprotein spectrum data of prostate cancer patients and benign control populations, and output the diagnostic results of prostate cancer.
[0008] Preferably, in step 1, the biological fluid sample is midstream urine from the first urination in the morning or peripheral blood plasma; When the biological fluid sample is midstream urine from the first morning urination, the pretreatment includes two differential centrifugations: the first centrifugation is at a low speed, with a centrifugal force ranging from 200g to 500g and a centrifugation time ranging from 5 minutes to 15 minutes, to remove detached cells; the supernatant from the first centrifugation is then subjected to a second centrifugation at a medium speed, with a centrifugal force ranging from 1500g to 3000g and a centrifugation time ranging from 15 minutes to 30 minutes, to remove large particles; the supernatant from the second centrifugation is collected and diluted and stabilized with protease inhibitors and phosphate buffer. When the biological fluid sample is peripheral blood plasma, the pretreatment includes two centrifugations: the first centrifugation is a low-speed centrifugation with a centrifugal force ranging from 1500g to 2500g and a centrifugation time ranging from 10 minutes to 20 minutes, used to remove blood cells; the supernatant from the first centrifugation is taken for a second centrifugation, which is a high-speed centrifugation with a centrifugal force ranging from 10000g to 15000g and a centrifugation time ranging from 20 minutes to 40 minutes, used to remove platelets, cell debris, and large vesicles; the supernatant from the second centrifugation is collected to obtain platelet-free plasma.
[0009] Preferably, in step 2, both the first DNA barcode sequence and the second DNA barcode sequence contain a polyadenine sequence as a spacer arm; The first DNA aptamer is immobilized on the surface of the first magnetic bead coated with streptavidin through the interaction of biotin and streptavidin, and the second DNA aptamer is immobilized on the surface of the second magnetic bead coated with streptavidin through the interaction of biotin and streptavidin. The preparation process of the first capture probe and the second capture probe specifically includes: synthesizing the first DNA aptamer and the second DNA aptamer respectively; covalently linking the first DNA aptamer to the first DNA barcode sequence modified with biotin to form a first biotinylated aptamer-barcode complex; covalently linking the second DNA aptamer to the second DNA barcode sequence modified with biotin to form a second biotinylated aptamer-barcode complex; incubating the first biotinylated aptamer-barcode complex with the first magnetic beads coated with streptavidin, so that the first biotinylated aptamer-barcode complex is immobilized on the first magnetic beads by biotin-streptavidin binding; incubating the second biotinylated aptamer-barcode complex with the second magnetic beads coated with streptavidin, so that the second biotinylated aptamer-barcode complex is immobilized on the second magnetic beads by biotin-streptavidin binding.
[0010] Preferably, in step 3, the first capture probe and the second capture probe are mixed in equal volumes or equimolar amounts and then added to the clarified body fluid sample; The incubation process is carried out in a constant temperature rotary incubator, with an incubation temperature range of 20°C to 30°C and an incubation time range of 1 hour to 3 hours; After incubation, an external magnetic field is applied to aggregate the first magnetic bead, the second magnetic bead, and the bound prostate cancer-derived exosome-magnetic bead complex. The supernatant is then discarded. The aggregated complex is washed with a pre-cooled washing buffer, which is a phosphate buffer containing a surfactant, for at least three washes.
[0011] Preferably, in step 4, the oxidant used in the oxidation reaction is sodium periodate, the concentration of sodium periodate in the oxidation buffer is in the range of 1 mM to 5 mM, the oxidation reaction is carried out under light-proof and low-temperature conditions, the reaction temperature range is 0°C to 10°C, and the reaction time range is 20 minutes to 40 minutes. In the biotinylate labeling reaction, the concentration of biotinylate in the labeling buffer ranges from 2 mM to 10 mM, the labeling reaction is carried out at room temperature, and the reaction time ranges from 45 minutes to 90 minutes. The lysis buffer used in the lysis reaction contains a nonionic surfactant and a reducing agent. The protease used in the enzymatic hydrolysis reaction is trypsin. The enzymatic hydrolysis reaction temperature range is 35°C to 38°C, and the enzymatic hydrolysis reaction time ranges from 12 hours to 18 hours.
[0012] Preferably, in step 4, the specific enrichment process specifically includes: incubating the peptide mixture with streptavidin magnetic beads at room temperature for at least 1 hour; applying a magnetic field after incubation to aggregate the streptavidin magnetic beads bound to biotin-labeled peptides; removing the unbound non-glycosylated peptide solution; eluting the aggregated streptavidin magnetic beads with an acidic elution buffer to dissociate and collect the biotin-labeled surface glycoprotein peptides; The acidic eluent is an aqueous solution containing 1% to 5% acetonitrile by volume and 0.05% to 0.2% formic acid by mass and volume.
[0013] Preferably, in step 5, the liquid chromatography-tandem mass spectrometry analysis is performed using a nanoliter liquid chromatography system coupled with a high-resolution tandem mass spectrometer; The mass spectrometry data acquisition mode is either data-dependent acquisition mode or data-independent acquisition mode. In the data-dependent acquisition mode, the parent ion with the highest intensity is selected for fragmentation in real time, while in the data-independent acquisition mode, all ions are cyclically fragmented and scanned according to a predetermined mass-to-charge ratio window. The raw mass spectrometry data is used to identify peptides by searching a protein sequence database, and the surface glycoproteins corresponding to the peptides are relatively quantified by a label-free quantification method, generating the surface glycoprotein spectrum data containing a list of surface glycoprotein identifications and relative abundance information.
[0014] Preferably, in step 6, the process of constructing the pre-trained diagnostic model specifically includes: Step A: Collect and obtain the surface glycoprotein spectrum data of pathologically confirmed prostate cancer patient samples as a positive training set, and the surface glycoprotein spectrum data of benign prostate disease patient samples as a negative training set; Step B: Normalize and standardize the quantitative data of all surface glycoproteins in the positive training set and the negative training set; Step C: Use machine learning algorithms to perform feature selection on the preprocessed data, and screen out multiple surface glycoproteins whose expression differences between the positive training set and the negative training set are determined to be significant by the selected machine learning algorithm to form a diagnostic feature panel; Step D: Based on the quantitative data of multiple surface glycoproteins in the diagnostic feature panel, a classification algorithm is used to train and generate the diagnostic model that can distinguish between prostate cancer and benign conditions, and a classification threshold is determined for the diagnostic model.
[0015] Preferably, in step C, the machine learning algorithm is a random forest, support vector machine, or LASSO regression algorithm; The diagnostic feature panel contains 5 to 15 surface glycoproteins; The diagnostic feature panel contains surface glycoproteins selected from the group consisting of: extracellular 5'-nucleotidase, prostate stem cell antigen, integrin α6, protein tyrosine kinase 7, growth differentiation factor 15, CD276 molecule, and L-amino acid transporter 1.
[0016] Preferably, the method is also used to assess the invasive risk or grading of prostate cancer; The training set of the diagnostic model further includes surface glycoprotein profile data of prostate cancer patients with different clinical stages or Gleason scores; The diagnostic results output in step 6 include the risk level or invasiveness grade of prostate cancer.
[0017] The beneficial effects of this invention are: 1. This invention creatively employs an orthogonal recognition system using aptamer probes that target CD63 protein and prostate-specific membrane antigen (PSA) respectively. CD63 protein is a universal marker for exosomes, ensuring that the captured target is exosomes; PSA is a highly specific surface marker for prostate cells. Only when the same vesicle expresses both proteins simultaneously will it be co-captured and enriched. This design achieves "logical AND" gating at the molecular level, precisely screening prostate-derived exosome subpopulations that are highly likely to be cancerous from the complex body fluid matrix, while excluding exosomes from other tissues and organs that only express CD63, as well as free PSA proteins in non-vesicular forms. The result is high-purity enrichment of the target analyte from the analytical source, fundamentally reducing the biological background noise in subsequent detection, and providing a novel technical approach to address the core clinical pain points of insufficient specificity and high false-positive rates in traditional PSA detection.
[0018] 2. This invention employs a chemical glycoproteomics strategy, utilizing sodium periodate to oxidize the glycan chains of all glycoproteins on the surface of exosomes, followed by unbiased covalent labeling using biotinylate, achieving panoramic and unbiased coverage of the surface glycoproteome. Subsequently, combining the deep identification capabilities of high-resolution mass spectrometry with machine learning algorithms to mine massive amounts of data, hundreds of surface proteins can be discovered and quantified simultaneously. This method not only validates known biomarkers but, more importantly, systematically discovers novel, low-abundance, and diagnostically valuable combinations of surface glycoproteins. Finally, the diagnostic model constructed by integrating multivariate information through a machine learning model demonstrates discriminative power based on a complex network of biomarkers, rather than a single indicator. This results in significantly higher area under the curve, sensitivity, and specificity in independent tests compared to traditional prostate-specific antigen (PSA) detection, achieving a leap from "limited target detection" to "panoramic information mining and intelligent diagnosis."
[0019] 3. This invention designs a closed, integrated process using magnetic beads as a solid-phase carrier: starting with specific capture, subsequent oxidation labeling, lysis, enzymatic digestion, and even peptide affinity enrichment are all performed sequentially in situ on the magnetic bead surface, with step switching only achieved through magnetic separation and buffer washing. This "one-pot" design greatly reduces the loss of target analytes and the risk of contamination caused by sample transfer, significantly improving the repeatability and sensitivity of the analysis. Simultaneously, the standardization and relative simplification of the process require less sample volume and more controllable operation time, better meeting the practical requirements of clinical testing laboratories for high throughput, automation, and stability. This greatly enhances the feasibility and translational potential of this cutting-edge omics technology from laboratory to clinical application, providing a practical tool for non-invasive, large-scale screening and precise diagnosis of prostate cancer. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of the steps of the method of the present invention; Figure 2 This is a flowchart illustrating the steps involved in constructing the pre-trained diagnostic model in the method of this invention. Detailed Implementation
[0022] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should also be noted that, to make the embodiments more comprehensive, the following embodiments are the best and preferred embodiments, and those skilled in the art can use other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0023] Please see Figures 1-2 This invention provides a method for diagnosing prostate cancer using exosome surface protein profiling. The method begins with the collection and pretreatment of biological fluid samples, including urine or blood. Large particulate impurities such as cells and cell debris are removed using physical methods such as differential centrifugation to obtain a clear liquid matrix. A dual-marker orthogonal barcode capture system is then used to specifically separate the target exosomes.
[0024] The system consists of two functionalized magnetic bead probes: the first probe is linked to an aptamer targeting the CD63 protein via a first DNA barcode; the second probe is linked to an aptamer targeting the prostate-specific membrane antigen via a second DNA barcode. The two barcode sequences are carefully designed to ensure no hybridization, thus maintaining independence in solution. When both probes are incubated with a sample, only exosomes expressing both CD63 and prostate-specific membrane antigen on their surface are co-captured, forming a stable complex. After capture, in-situ chemical glycoproteomics processing is performed. This process involves oxidizing the glycan chains of exosome surface glycoproteins with sodium periodate to generate active aldehyde groups, followed by covalent labeling with biotinylate to achieve specific biotinylation of all surface glycoproteins.
[0025] Exosomes were then lysed and digested with trypsin to release a mixture of peptides. The peptides were then purified using streptavidin magnetic beads, with only biotin-tagged surface glycoprotein-derived peptides being selectively enriched. The enriched peptides were separated by nano-liquid chromatography and analyzed by high-resolution tandem mass spectrometry (HPLC) using a data-independent acquisition mode for full-spectrum scanning to ensure the detection of low-abundance peptides. The raw mass spectrometry data were processed using database searching and label-free quantification algorithms to generate surface glycoprotein profiles containing identification information and relative abundance values for each surface glycoprotein. Finally, this data was input into a pre-trained diagnostic model.
[0026] The diagnostic model is a classifier trained using machine learning algorithms based on surface glycoprotein profile data from pathologically confirmed prostate cancer patients and benign control groups. Upon receiving glycoprotein profile data from new samples, the model outputs a quantitative risk score or classification label as the diagnostic result for prostate cancer. This series of interconnected steps, from source-specific enrichment and panoramic surface protein labeling to big data model interpretation, constitutes a closed, efficient, and information-maximizing diagnostic chain.
[0027] First, the dual-marker capture strategy fundamentally solves the challenge of distinguishing prostate cancer-derived exosomes from complex bodily fluids, precisely focusing the analysis from "total exosomes" to "disease-related exosomes," significantly reducing background noise—a foundation for subsequent high-precision analysis. Second, in-situ chemical glycoproteomics processing using magnetic beads enables unbiased, panoramic labeling and enrichment of the surface proteome of target exosomes on a solid-phase support, avoiding losses caused by sample transfer in traditional methods and specifically highlighting glycoproteins as an information-rich functional molecule category. Finally, combining high-throughput mass spectrometry with machine learning models enables the automatic discovery of optimal diagnostic marker combinations from massive protein data, overcoming the limitations of traditional single or limited marker detection and significantly improving the robustness and accuracy of the diagnostic model. The entire approach provides diagnostic specificity and the ability to discover unknown markers far exceeding traditional PSA detection.
[0028] In one possible implementation, for midstream urine samples from the first morning urination, the pretreatment process includes two consecutive differential centrifugations. The first centrifugation uses a low-speed centrifugation, with the centrifugal force set sufficient to precipitate detached epithelial cells, leukocytes, and other formed cellular components, but not to cause excessive sedimentation of exosomes. After centrifugation, the supernatant is carefully transferred to a new container. The second centrifugation uses a medium-speed centrifugation, with the centrifugal force set to effectively precipitate larger cell debris, apoptotic bodies, and some large extracellular vesicles, while retaining smaller exosomes in the supernatant. The supernatant after the second centrifugation is collected, and a broad-spectrum protease inhibitor is immediately added to prevent protein degradation. Simultaneously, phosphate buffer is added for appropriate dilution and ionic strength adjustment to form a stable urine sample for testing.
[0029] For peripheral blood plasma samples, pretreatment also involves two centrifugation steps. The first, low-speed centrifugation aims to completely precipitate blood cell components such as red blood cells and white blood cells in whole blood under gentle conditions, obtaining platelet-rich supernatant plasma. The second, high-speed centrifugation aims to thoroughly remove platelets, cell membrane debris, and large protein polymers using higher centrifugal force, obtaining clear, platelet-free plasma. This step is crucial for removing impurities that may interfere with subsequent specific capture.
[0030] By employing appropriate centrifugation parameters and procedures for urine and blood, two different matrices, the removal of their respective unique interfering substances can be maximized. Differential centrifugation in urine pretreatment effectively removes a large number of non-vesicle particles, reducing sample viscosity and background. The platelet removal step in plasma pretreatment is particularly critical, as platelets themselves release a large number of vesicles, constituting the main background interference. Through this standardized pretreatment, samples from different sources are transformed into clear liquids with low impurity content and well-preserved exosomes, creating uniform and clean starting conditions for downstream high-specificity and high-sensitivity capture and detection steps, ensuring the reproducibility and reliability of experimental results.
[0031] In one possible implementation, the core portions of the first and second DNA barcode sequences are nucleotide sequences that are completely non-complementary to each other, ensuring no hybridization occurs in solution. At the end of each barcode sequence, a continuous polyadenine sequence is designed as a long-chain spacer arm. The introduction of the spacer arm increases the spatial distance between the DNA aptamer and the magnetic bead surface, reducing steric hindrance and allowing the aptamer to move and conformationally change more freely, thus binding to the target membrane protein more efficiently. Probe immobilization is achieved using a biotin-streptavidin system. First, the first and second DNA aptamers are prepared separately using solid-phase chemical synthesis, and then a complete barcode sequence modified with biotin is covalently linked to the end of the aptamer via an enzymatic ligation reaction, forming a biotinylated composite molecule.
[0032] Then, these two composite molecules were co-incubated with superparamagnetic magnetic beads coated with streptavidin. Streptavidin exhibits extremely high affinity and specificity for biotin; after incubation, the biotinylated aptamer-barcode complex was firmly and oriented uniformly onto the surface of the magnetic beads through this interaction. Excess unbound molecules were removed by magnetic separation and washing with buffer, ultimately yielding capture probe magnetic beads with a high surface density modified with specific aptamers.
[0033] The orthogonality of the barcode sequences ensures that the two probes do not interfere with each other when used in combination, and each functions independently. The introduction of the long-chain polyadenine spacer arm is a key design feature, significantly improving the binding efficiency and kinetics of the aptamer in recognizing its membrane protein target, overcoming recognition barriers that may arise from the solid-phase interface. Immobilization using the biotin-streptavidin system results in a binding strength far exceeding that of physical adsorption or many chemical couplings, ensuring probe stability during subsequent incubation and washing processes and preventing probe detachment that could lead to a decrease in capture efficiency. This stable, efficient, and specific probe construction method is the material basis for ensuring the high sensitivity and specificity of the entire dual-marker capture system.
[0034] In one possible implementation, the first and second capture probes are mixed in equal particle numbers or volumes to ensure balanced capture opportunities for both target sites on the target exosome. The mixed probes are incubated with a clear body fluid sample in a thermostatic rotary incubator. The incubation temperature is selected in the range of room temperature to slightly below body temperature to maintain protein activity and promote molecular diffusion and binding, while avoiding excessively high temperatures that could cause sample denaturation. The incubation time needs to be long enough to ensure that the aptamers and target proteins on the membrane fully bind to reach equilibrium. After incubation, an external magnetic field is applied, and all magnetic particles and their conjugates are rapidly enriched to the tube wall. After removing the supernatant containing unbound components, the magnetic bead complex is washed multiple times with a pre-cooled wash buffer. The wash buffer is a phosphate buffer containing a very low concentration of a nonionic surfactant, the concentration of which is precisely controlled below its critical concentration for micelle formation. This low concentration of surfactant effectively weakens and elutes nonspecific adsorption caused by hydrophobic interactions without disrupting the aptamer-antigen specific binding bond or dissolving the lipid bilayer.
[0035] This optimized incubation and washing protocol directly improved the purity and efficiency of the captured samples. Proportional mixing of probes ensured the symmetry of dual-target recognition, avoiding non-specific background caused by an overabundance of one probe. Isothermal rotation incubation provided a uniform reaction environment, promoting specific binding. Most importantly, the washing solution containing a surfactant with a critical micelle concentration effectively removed non-specific adsorbed impurities while perfectly protecting the specifically captured exosome complex and its fragile membrane structure. After several such rigorous washes, the resulting magnetic bead-exosome complex exhibited extremely high purity, minimizing background signals for downstream analysis and laying the foundation for obtaining high-quality omics data.
[0036] In one possible implementation, the oxidation reaction uses sodium periodate as the oxidant, its concentration set at a level sufficient to effectively oxidize the cis-diol structure at the glycan terminus without over-oxidizing and damaging the protein backbone or oxidizing other sensitive amino acid residues. The reaction is carried out at low temperature and in the dark to suppress side reactions of sodium periodate, ensuring the specificity and reproducibility of the oxidation. In the subsequent biotinylate labeling step, the concentration of biotinylate needs to be sufficiently high to drive its efficient covalent binding to the aldehyde group. The reaction is carried out at room temperature, and the time is optimized to ensure complete labeling. The buffer used in the lysis step contains a nonionic surfactant that effectively dissolves the lipid bilayer and denatures and dissolves membrane proteins while maintaining compatibility with subsequent enzymatic digestion steps. The buffer also contains a reducing agent to break disulfide bonds in the protein, linearizing the protein and increasing the efficiency of enzymatic digestion. Enzymatic digestion uses sequencing-grade modified trypsin, which is subjected to long-term digestion under isothermal conditions close to physiological temperature, ensuring that the complex protein mixture is thoroughly digested into peptides suitable for mass spectrometry analysis.
[0037] Mild and specific oxidation conditions ensure that only the glycans of surface glycoproteins are labeled, a prerequisite for surface protein-specific analysis. Efficient and thorough biotinylate labeling ensures high yields in subsequent enrichment steps. Effective lysis and reduction guarantee the complete dissolution and unfolding of all proteins, especially hydrophobic transmembrane proteins, crucial for comprehensive protein identification. Sufficient and specific enzymatic digestion produces a high-quality, information-complete mixture of peptides. The entire process is performed in situ on magnetic beads, with transitions between steps only involving magnetic separation and washing, greatly simplifying operations, reducing peptide loss and sample contamination, and improving reproducibility and throughput.
[0038] In one possible implementation, the enzymatically hydrolyzed peptide mixture is co-incubated with streptavidin-coated magnetic beads for a duration sufficient to ensure adequate binding between the biotinylated peptides and streptavidin. After incubation, a magnetic field is applied to separate the biotinylated peptides from the solution. The supernatant containing a large amount of non-biotinylated peptides is removed; this step removes approximately the majority of non-target peptides. Subsequently, the magnetic beads are eluted using an acidic eluent. The acidic eluent is typically a solution containing a low concentration of organic solvent and a weak acid. Its acidic environment protonates the key amino acid residues that bind streptavidin to biotin, thereby reversibly reducing their affinity and gently dissociating and collecting the biotinylated peptides. The organic solvent in the eluent helps maintain peptide solubility and improves elution efficiency.
[0039] This specific enrichment step enables the ultra-high purity separation of target peptides derived from surface glycoproteins from an extremely complex mixture of peptides. Streptavidin exhibits extremely high affinity for biotin and is highly stable after binding, allowing for a very rigorous washing process that almost completely removes non-specifically adsorbed peptides. Ultimately, the eluted peptide set consists almost entirely of peptides derived from surface glycoproteins, resulting in an exceptionally high signal-to-noise ratio. This significant improvement in purity directly enhances the sensitivity and accuracy of subsequent mass spectrometry analysis, enabling the reliable identification and quantification of even low-abundance surface biomarkers, which is crucial for the discovery of novel biomarkers.
[0040] In one possible implementation, peptides are first separated using a nano-liquid chromatography system. This system, employing an extremely fine-diameter column and very low flow rate, enables high-resolution separation of complex peptide mixtures based on hydrophobicity, significantly improving peak capacity and sensitivity for mass spectrometry detection. The separated peptides are then fed online into a high-resolution tandem mass spectrometer. Mass spectrometry acquisition can employ two main strategies: a data-dependent acquisition mode selects the most intense precursor ion for fragmentation in real time; while a data-independent acquisition mode performs indiscriminate, cyclic fragmentation scans of all ions according to a predetermined mass-to-charge ratio window. The latter provides more comprehensive fragment ion information, particularly beneficial for the quantitative reproducibility of low-abundance peptides. The obtained raw mass spectra are compared with protein sequence databases using a search algorithm to identify the peptide sequences. Simultaneously, label-free quantification techniques are used—that is, directly comparing the chromatographic peak intensity or mass spectrometric signal intensity of the same peptide in different samples—to achieve relative protein quantification, ultimately outputting surface glycoprotein spectrum data containing protein identity and abundance information.
[0041] This advanced analytical approach enables high-depth and high-precision identification and quantification of surface glycoprotein peptides. Nano-liquid chromatography provides superior separation capabilities and reduces ion suppression effects during mass spectrometry detection. High-resolution mass spectrometry ensures accurate mass determination and reliable spectral resolution. The application of data-independent acquisition modes enhances the comprehensiveness and quantitative reproducibility of complex sample analysis. Ultimately, this combination of technologies can obtain quantitative information on hundreds of surface glycoproteins from a single experiment, constructing extremely rich molecular characteristic spectra, providing a solid, high-quality data foundation for big data-based machine learning diagnostic models.
[0042] In one possible implementation, firstly, step A involves collecting sample data with clear clinical labels confirmed by pathological gold standard diagnosis, including surface glycoprotein profile data from prostate cancer patients as a positive training set, and corresponding data from patients with benign prostate disease as a negative training set. Step B involves data preprocessing, normalizing the protein quantification data of all samples in the training set to correct for systematic errors such as sample loading, and performing standardization transformation to ensure the data conforms to the distribution assumptions of subsequent algorithms. Step C is feature selection, using machine learning algorithms to process the preprocessed high-dimensional data. The algorithm evaluates the importance of each surface glycoprotein in distinguishing between positive and negative categories, automatically selecting proteins with significant differences in expression between the two groups and low correlation between them, forming a concise and effective diagnostic feature panel. Step D involves model training and threshold determination, using the selected feature protein data as input variables to construct a discriminative model using a classification algorithm. After model training, an optimal classification threshold is selected by calculating the relationship between the model's predicted score and the true label on the training or validation set; this threshold typically corresponds to the highest overall classification performance index.
[0043] This model building process successfully transforms high-throughput omics data into clinically applicable diagnostic tools. It overcomes the subjectivity and limitations of traditional methods that rely on expert experience to manually select biomarkers. Through algorithms, it automatically filters the optimal feature combinations from massive datasets. These combinations often include both known and unknown biomarkers, providing a more comprehensive characterization of disease states. Diagnostic models built based on multivariate combinations typically exhibit discriminative power far superior to any single biomarker, significantly improving diagnostic accuracy, stability, and generalization ability, representing a crucial step towards precision medicine.
[0044] In one possible implementation, the machine learning algorithm used for feature selection could be a random forest, which selects features by constructing multiple decision trees and evaluating the importance of each protein in the decision; it could also be a support vector machine, which finds the optimal hyperplane that maximizes the differentiation between two classes of samples, with the key elements in its weight vector corresponding to important features; or LASSO regression, which automatically compresses the coefficients of irrelevant features to zero by adding an L1 regularization term to the loss function, thus achieving feature selection. The final diagnostic feature panel typically contains between 5 and 15 proteins. This range ensures sufficient information while avoiding the curse of dimensionality and overfitting, which is beneficial for the model's stable performance on future new samples. Specific examples of proteins included in this feature panel include extracellular 5'-nucleotidase, prostate stem cell antigen, integrin α6, protein tyrosine kinase 7, growth differentiation factor 15, CD276 molecule, and L-amino acid transporter 1. These proteins are involved in multiple pathways related to cancer progression, such as cell adhesion, signal transduction, immune regulation, and metabolic reprogramming.
[0045] The clearly defined feature panel structure and methodology enhance the feasibility and reproducibility of the invention. The enumeration of specific algorithms provides implementers with a clear technical selection path. Limiting the range of feature counts offers best practice guidance on model complexity. The examples of specific proteins, particularly those molecules known to be closely related to the development and progression of prostate cancer, not only biologically validate the rationality of the model screening results but also reveal that the biomarker combination discovered in this invention has a solid biological basis and is not merely a mathematical coincidence. This provides biological interpretability to the diagnostic model and enhances its credibility in clinical translation.
[0046] In one possible implementation, to achieve this, the training set data used to build the diagnostic model needs to be further refined. In addition to samples categorized as cancerous or benign, the training set should also include samples representing different degrees of disease invasion, such as localized prostate cancer samples in low-risk, intermediate-risk, and high-risk groups, as well as samples that have metastasized. These samples are clinically labeled with corresponding Gleason scores, clinical stages, or metastatic status. Subsequently, a machine learning process similar to that used for classification models is employed, but the goal could be to build a regression model to predict a continuous risk score or a multi-classification model to distinguish different risk levels. When analyzing new samples, in addition to obtaining a "cancer / benign" diagnosis, a quantitative risk score or specific risk level can be further output.
[0047] This extended application greatly enhances the clinical value of the method of this invention. Prostate cancer is a highly heterogeneous disease, and distinguishing between indolent cancers and clinically significant cancers requiring active treatment is currently a challenge in clinical decision-making. The method of this invention, by capturing subtle molecular differences related to the protein profiles on the surface of exosomes, can provide prognostic information beyond simple diagnosis, assisting clinicians in making individualized treatment decisions, such as whether to choose active monitoring or immediate intervention. This is expected to improve treatment outcomes while reducing overtreatment of indolent diseases.
[0048] Example: Diagnosis of clinically significant prostate cancer based on urinary exosome surface glycoprotein profiles; This embodiment aims to develop a non-invasive liquid biopsy method that is superior to traditional PSA testing, to differentiate clinically significant prostate cancer from benign prostatic hyperplasia, thereby reducing unnecessary puncture biopsies.
[0049] 1. Sample collection and preprocessing; Sample Source: This study included 100 participants. Fifty patients with clinically significant prostate cancer (Gleason score ≥ 7) confirmed by transrectal ultrasound-guided prostate biopsy were included as the positive group. Another 50 patients diagnosed with benign prostatic hyperplasia (BPH) by comprehensive examination including ultrasound, PSA, and MRI were included as the benign control group. All participants provided informed consent.
[0050] Sample Type and Processing: Approximately 100 ml of midstream urine was collected from the first urination of all subjects upon waking. The urine sample processing strictly followed the differential centrifugation procedure for morning urine. Specifically, the urine sample was first centrifuged at 400 x g for 12 minutes at 4°C. This step aims to use relatively low centrifugal force to allow detached epithelial cells, leukocytes, and other formed elements in the urine to settle to the bottom of the tube, while exosomes, due to their small size and low sedimentation coefficient, remain in the supernatant. The supernatant was collected, taking care to avoid touching the sediment at the bottom of the tube. Subsequently, the supernatant was transferred to a new centrifuge tube and centrifuged at 2500 x g for 25 minutes at 4°C. This moderate centrifugation force aims to remove any "noise" particles in the urine, such as cell debris, apoptotic bodies, and larger extracellular vesicles, resulting in a highly clear urine supernatant. Finally, a broad-spectrum protease inhibitor mixture (added according to the instructions) and an equal volume of pre-chilled phosphate buffer were added to the supernatant. After thorough mixing, the mixture was aliquoted and stored at -80°C until subsequent analysis. This pretreatment process maximizes the preservation of exosome integrity and bioactivity while removing most impurities that would interfere with subsequent specific capture.
[0051] 2. Construct a dual-marker orthogonal barcode capture probe; Probe design principle: The first capture probe targets the CD63 protein, and the second capture probe targets the prostate-specific membrane antigen. The ingenuity of this design lies in constructing a "logical AND" gating system: only vesicles that simultaneously express CD63 (ensuring they are exosomes) and prostate-specific membrane antigen (ensuring they originate from the prostate and are highly likely to be cancerous tissue) can be co-captured, thereby achieving specific recognition of the target exosome subset.
[0052] DNA barcode sequences and spacer arms: The first DNA barcode sequence is designed as "Bio-T10-AAAAAACGTACG", and the second DNA barcode sequence is designed as "Bio-T10-CCCCCGACTAGT". Here, "Bio" represents the biotinylation at the 5' end; "T10" represents a flexible linker arm composed of 10 thymine deoxynucleotides, which provides spatial freedom for the subsequent aptamer, preventing steric hindrance from affecting its binding efficiency to the target protein; "AAAAAACGTACG" and "CCCCCGACTAGT" are core orthogonal sequences. Bioinformatics software verification confirmed that there are no consecutive complementary base pairs of four or more between them, thus preventing hybridization and ensuring that the two probes do not interfere with each other in solution.
[0053] Aptamer Immobilization: First, high-affinity anti-CD63 DNA aptamer sequences and anti-prostate-specific membrane antigen DNA aptamer sequences were synthesized using a solid-phase synthesis method. Then, via enzymatic ligation, the first DNA aptamer was covalently linked to a biotinylated first DNA barcode sequence to form a first biotinylated aptamer-barcode complex; similarly, a second biotinylated aptamer-barcode complex was prepared. Next, both complexes were incubated with commercially available streptavidin-coated superparamagnetic beads (bead diameter, for example, 2.8 μm). Streptavidin exhibits extremely high affinity for biotin, resulting in rapid and irreversible binding. After incubation, magnetic separation and thorough washing yielded a first capture probe (CD63 probe beads) and a second capture probe (prostate-specific membrane antigen probe beads) with a high surface density immobilized with the corresponding aptamer probes. This biotin-streptavidin system provides stable and uniform probe immobilization.
[0054] 3. Specific capture and enrichment of exosomes; Capture Incubation: Take 1 mL of pretreated urine sample and allow it to return to room temperature. Add 50 μL each of the prepared first and second capture probe magnetic bead suspensions (ensuring approximately equal molar numbers of the two types of magnetic beads) to the urine sample simultaneously. Place the mixture on a horizontal rotary mixer and gently incubate at 25°C at 15 rpm for 2.5 hours. This gentle rotational incubation condition promotes sufficient contact and binding between the probe and the target exosomes, while avoiding damage to the exosome membrane structure or increased nonspecific adsorption due to vigorous shaking.
[0055] Separation and Washing: After incubation, place the reaction tube on a strong magnetic rack and let it stand for 5 minutes. During this time, all magnetic beads and their aptamer-bound exosomes (especially target exosomes captured by both probes) are attracted to the tube wall. Carefully discard all supernatant. Then, while the magnetic beads remain fixed by the magnetic field, add 1 mL of pre-cooled wash buffer. This buffer is a phosphate buffer containing 0.01% polysorbate 20. Polysorbate 20 is a nonionic surfactant with a concentration below the critical micelle concentration, effectively weakening and eluting impurity proteins or irrelevant particles non-specifically adsorbed through hydrophobic interactions, without disrupting the specific binding between aptamers and antigens or the integrity of the exosome membrane. Repeat this washing process four times until a pure magnetic bead-exosome complex is obtained.
[0056] 4. In-situ chemical glycoproteomics processing of magnetic beads; This step is the key technology for achieving panoramic analysis of surface glycoproteins. All operations are performed on magnetic beads, eliminating the need for exosome elution and greatly reducing sample loss.
[0057] In-situ oxidation of surface glycoproteins: 200 μL of freshly prepared oxidation buffer containing 3 mmol / L sodium periodate was added to the washed magnetic bead complex. The reaction tube was placed on ice and wrapped in aluminum foil to protect it from light, and incubated for 30 minutes. Sodium periodate specifically and mildly oxidizes the sialic acid or other cis-diol structures at the glycan terminals of exosome surface glycoproteins, converting them into highly reactive aldehyde groups. Low temperature and light protection are used to prevent excessive oxidation or self-decomposition of sodium periodate, ensuring the specificity and controllability of the reaction.
[0058] Biotinylhydrazide labeling: After the oxidation reaction, magnetic separation and discard the oxidation buffer, followed by a rapid wash with pre-cooled phosphate buffer. Immediately add 200 μL of labeling buffer containing 5 mmol / L biotinylhydrazide. Incubate at room temperature in the dark for 75 minutes. The hydrazide group in the biotinylhydrazide molecule undergoes a highly efficient Schiff base reaction with the aldehyde group generated in the previous step, forming a stable covalent bond. This allows the biotin tag to be permanently and in situ labeled on all oxidized surface glycoproteins, acting like a "molecular fishing hook."
[0059] In situ lysis and enzymatic digestion: After labeling, the magnetic beads were washed. 100 μL of lysis buffer containing 1% RapiGestSF surfactant and 10 mmol / L dithiothreitol was added directly to the beads. The mixture was heated at 95°C for 10 minutes. This step completely disrupted the lipid bilayer of the exosomes, releasing all membrane and luminal proteins, while dithiothreitol reduced disulfide bonds in the proteins. Subsequently, trypsin was added, and the mixture was incubated overnight (16 hours) in a 37°C shaker. Trypsin cleaved the proteins into a mixture of peptides suitable for mass spectrometry analysis.
[0060] Specific enrichment of labeled peptides: After enzymatic digestion, no complex processing such as peptide desalting is required. The entire enzymatic digestion system is directly mixed with 50 μL of streptavidin magnetic bead suspension and incubated at room temperature for 90 minutes. Due to the extremely high affinity between biotin and streptavidin, all peptides labeled with biotinylate in the first step (i.e., peptides derived from exosome surface glycoproteins) are specifically captured onto the streptavidin magnetic beads. Magnetic separation easily removes all unlabeled non-glycosylated protein peptides, internal protein peptides, and components of the digestion buffer. Finally, elution is performed using an aqueous solution containing 2% acetonitrile and 0.1% formic acid to detach the pure surface glycoprotein-derived peptides from the streptavidin magnetic beads for subsequent mass spectrometry analysis. This enrichment step is crucial for obtaining high signal-to-noise ratio surface glycoprotein profiles.
[0061] 5. Liquid chromatography-tandem mass spectrometry analysis and data acquisition; Instrumentation and parameters: The enriched peptide solution was separated using a nano-flow rate liquid chromatography system. A reversed-phase C18 capillary column was used, with elution employing a 150-minute acetonitrile gradient. The separated peptides were detected by online electrospray injection into a high-resolution orbital trap tandem mass spectrometer.
[0062] Data Acquisition Mode: This embodiment employs a data-independent acquisition mode. In this mode, the mass spectrometer cyclically performs indiscriminate secondary fragmentation scans on all incoming ions according to a predetermined mass-to-charge ratio window, thereby obtaining fragmentation information for all peptides (including low-abundance peptides). Compared to traditional data-dependent acquisition modes, data-independent acquisition modes offer higher quantitative repeatability and a more complete spectral acquisition rate, making them particularly suitable for in-depth proteomics analysis of complex biological samples.
[0063] Data Analysis: The obtained raw mass spectrometry data were analyzed using specialized software (such as Spectronaut, DIA-NN, etc.). The software compared the spectra with the human protein sequence database to identify peptides and their corresponding proteins. Simultaneously, label-free quantification was performed by integrating the chromatographic peaks of the peptide precursor ions, i.e., calculating the relative abundance based on the intensity of the peptide signal. Finally, a "Surface Glycoprotein Spectrum Data" file containing hundreds of identified surface glycoproteins and their relative quantification values was generated for each sample.
[0064] 6. Diagnostic model construction and result interpretation; Model building process: Data preparation (Step A): 100 samples (50 cancer, 50 benign) were randomly divided into a training set (70 cases, 35 cancer / 35 benign) and an independent test set (30 cases, 15 cancer / 15 benign).
[0065] Data preprocessing (Step B): The surface glycoprotein quantification data of all samples in the training set are preprocessed. First, median normalization is performed to correct for differences in total protein content among different samples. Then, logarithmic transformation is performed to make the data distribution closer to a normal distribution.
[0066] Feature Selection (Step C): This is one of the most crucial steps in model building. This embodiment uses the LASSO regression algorithm for feature selection. LASSO regression automatically filters variables by adding an L1 regularization term to the loss function, compressing the coefficients of unimportant variables to zero. We input the protein expression data matrix (rows are samples, columns are hundreds of proteins) and sample labels (cancer / benign) from the training set into the LASSO regression model. Through 10-fold cross-validation, we selected the regularization strength λ that minimizes the model's average error. Ultimately, the model selected eight surface glycoproteins with non-zero coefficients, forming the diagnostic feature panel. These eight proteins include: extracellular 5'-nucleotidase, prostate stem cell antigen, integrin α6, protein tyrosine kinase 7, growth differentiation factor 15, CD276 molecule, L-amino acid transporter 1, and another transmembrane protein annotated in the database as having an unknown function.
[0067] Model Training and Threshold Determination (Step D): Using the quantitative data of the above eight proteins as input features, the final diagnostic model is trained on the training set using a logistic regression classification algorithm. The model output is a probability value between 0 and 1, representing the probability that the sample is cancerous. By analyzing the predicted probability distribution of the training set samples, we select the point with the largest Youden index as the classification threshold. In this example, the threshold is determined to be 0.52. That is, when the model's predicted probability for a new sample is greater than or equal to 0.52, it is interpreted as "high risk of prostate cancer"; when it is less than 0.52, it is interpreted as "high probability of benign".
[0068] Model Validation and Result Interpretation: Surface glycoprotein profiles (data for the aforementioned eight characteristic proteins) from 30 independent test samples that were not involved in model training were input into a trained logistic regression model to calculate predicted probabilities. Interpretations were then made based on a threshold of 0.52, and the results were compared with the gold standard pathological diagnosis.
[0069] Performance verification: To demonstrate the superiority of the method of the present invention, we set up the following three comparative examples and compared their performance with the method of the present invention (example) on the same test set (30 cases).
[0070] Comparative Example 1: Traditional serum PSA detection; Methods: Peripheral blood was collected from all subjects in the test set, and serum total PSA concentration was measured using a clinically accepted chemiluminescent immunoassay. The clinically common diagnostic threshold of 4.0 ng / mL was used for interpretation.
[0071] Comparative Example 2: Detection of exosomes using a single biomarker; Methods: A traditional immunomagnetic bead capture method was employed. Exosomes were captured from the same batch of urine samples using only magnetic beads coated with anti-prostate-specific membrane antigen antibodies. After capture, the exosomes were lysed, and the levels of five known prostate cancer-related proteins (including PSA and prostate acid phosphatase) in the lysate were detected using multiplex liquid chromatography-array technology. A simple linear discriminant function was then established for diagnosis.
[0072] Comparative Example 3: Non-specific enrichment of exosomal glycoproteomics; Methods: Total exosomes were nonspecifically precipitated from urine using a common polymer precipitation method (e.g., PEG precipitation). Subsequently, the precipitated total exosomes underwent the same chemical glycoproteomics processing and mass spectrometry analysis as in steps 4 and 5 of this embodiment. Using the same training set and LASSO regression algorithm, features were screened from the surface glycoprotein profiles of the total exosomes to construct a diagnostic model, which was then validated on a test set.
[0073] Table 1: Comparison of diagnostic performance of the embodiments of the present invention and comparative examples on independent test sets. Effect Analysis: Compared to Comparative Example 1 (PSA): the diagnostic performance of this invention (AUC 0.93) is significantly better than that of the traditional PSA test (AUC 0.71), mainly reflected in a substantial improvement in specificity (93.3% vs. 60.0%). This demonstrates that this method can effectively distinguish between cancer and benign hyperplasia, and is expected to significantly reduce false positives.
[0074] Compared to Comparative Example 2 (single biomarker): the embodiments of the present invention show improvements in both sensitivity (86.7% vs. 73.3%) and AUC. This indicates that by orthogonal capture with dual biomarkers, a purer population of target exosomes is obtained, and combined with panoramic glycoprotein profiling analysis (8 biomarker combinations), the information content and stability are superior to targeted detection of a few known proteins.
[0075] Compared to Comparative Example 3 (non-specific enrichment), the embodiments of this invention demonstrate superior performance in AUC, sensitivity, and specificity. This strongly proves the necessity and inventive value of the "dual-marker orthogonal barcode capture" step. Non-specific enrichment introduces a large amount of background noise exosomes, and even when using the same glycoproteomics technology, the screened markers are contaminated, leading to a decrease in the discriminative ability of the diagnostic model. This invention purifies the target analyte from the source, which is key to obtaining a highly specific diagnostic spectrum.
[0076] Furthermore, the method of this invention also has the potential to assess disease invasiveness. In an extended analysis, we found that the expression levels of growth differentiation factor 15 and CD276 molecules in the diagnostic feature panel were significantly higher in patients with a Gleason score of 8 or higher or those with existing micrometastases than in patients with localized carcinoma. This suggests that future research could expand the sample size and train a specialized risk stratification model to achieve an integrated "diagnosis-stratification" application.
[0077] The above embodiments fully demonstrate the completeness, feasibility, and significant progress and inventiveness of the technical solution of the present invention compared with the prior art.
[0078] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0079] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for diagnosing prostate cancer using exosome surface protein profiling, characterized in that, Includes the following steps: Step 1: Collect biological fluid samples from the subjects and preprocess them to remove impurities and obtain clear body fluid samples; Step 2: Prepare dual-marker orthogonal barcode capture probes. The first capture probe connects the first DNA aptamer targeting CD63 protein to the first magnetic bead via the first DNA barcode sequence. The second capture probe connects the second DNA aptamer targeting prostate-specific membrane antigen to the second magnetic bead via the second DNA barcode sequence. The first DNA barcode sequence and the second DNA barcode sequence do not hybridize. Step 3: Incubate the first and second capture probes together with a clear body fluid sample, separate them using a magnetic field and wash them to obtain a prostate cancer-derived exosome-magnetic bead complex that simultaneously binds the first and second capture probes. Step 4: Perform in-situ chemical glycoproteomics processing on the exosome-magnetic bead complex derived from prostate cancer, specifically including: oxidizing the glycan chains of the surface glycoproteins of prostate cancer-derived exosomes and converting them into aldehyde groups, labeling the aldehyde groups with biotinylate, cleaving the exosomes and enzymatically digesting them into a peptide mixture, and using streptavidin magnetic beads to specifically enrich the surface glycoprotein peptides carrying biotin labels from the peptide mixture; Step 5: Perform liquid chromatography-tandem mass spectrometry analysis on surface glycoprotein peptides to identify and quantify surface glycoproteins and generate surface glycoprotein spectrum data; Step 6: Input the surface glycoprotein spectrum data into the pre-trained diagnostic model, which is trained based on the surface glycoprotein spectrum data of prostate cancer patients and benign control populations, and output the diagnostic results of prostate cancer.
2. The method for diagnosing prostate cancer using exosome surface protein profiling according to claim 1, characterized in that, In step 1, the biological fluid sample is either midstream urine from the first urination in the morning or peripheral blood plasma; When the biological fluid sample is midstream urine from the first urination in the morning, the pretreatment includes two differential centrifugations: the first centrifugation is a low-speed centrifugation with a centrifugal force ranging from 200g to 500g and a centrifugation time ranging from 5 minutes to 15 minutes, used to remove exfoliated cells; the supernatant from the first centrifugation is then subjected to a second centrifugation, which is a medium-speed centrifugation with a centrifugal force ranging from 1500g to 3000g and a centrifugation time ranging from 15 minutes to 30 minutes, used to remove large particles of debris; Collect the supernatant from the second centrifugation, and dilute and stabilize it with protease inhibitor and phosphate buffer. When the biological fluid sample is peripheral blood plasma, the pretreatment includes two centrifugations: the first centrifugation is a low-speed centrifugation with a centrifugal force ranging from 1500g to 2500g and a centrifugation time ranging from 10 minutes to 20 minutes, used to remove blood cells; the supernatant from the first centrifugation is taken for a second centrifugation, which is a high-speed centrifugation with a centrifugal force ranging from 10000g to 15000g and a centrifugation time ranging from 20 minutes to 40 minutes, used to remove platelets, cell debris, and large vesicles; the supernatant from the second centrifugation is collected to obtain platelet-free plasma.
3. The method for diagnosing prostate cancer using exosome surface protein profiling according to claim 1, characterized in that, In step 2, both the first DNA barcode sequence and the second DNA barcode sequence contain a polyadenine sequence as a spacer arm; The first DNA aptamer is immobilized on the surface of the first magnetic bead coated with streptavidin through the interaction of biotin and streptavidin, and the second DNA aptamer is immobilized on the surface of the second magnetic bead coated with streptavidin through the interaction of biotin and streptavidin. The preparation process of the first capture probe and the second capture probe specifically includes: synthesizing the first DNA aptamer and the second DNA aptamer respectively; covalently linking the first DNA aptamer to the first DNA barcode sequence modified with biotin to form a first biotinylated aptamer-barcode complex; covalently linking the second DNA aptamer to the second DNA barcode sequence modified with biotin to form a second biotinylated aptamer-barcode complex; incubating the first biotinylated aptamer-barcode complex with the first magnetic beads coated with streptavidin, so that the first biotinylated aptamer-barcode complex is immobilized on the first magnetic beads by biotin-streptavidin binding; incubating the second biotinylated aptamer-barcode complex with the second magnetic beads coated with streptavidin, so that the second biotinylated aptamer-barcode complex is immobilized on the second magnetic beads by biotin-streptavidin binding.
4. The method for diagnosing prostate cancer using exosome surface protein profiling according to claim 3, characterized in that, In step 3, the first capture probe and the second capture probe are mixed in equal volumes or equimolar amounts and then added to the clarified body fluid sample; The incubation process is carried out in a constant temperature rotary incubator, with an incubation temperature range of 20°C to 30°C and an incubation time range of 1 hour to 3 hours; After incubation, an external magnetic field is applied to aggregate the first magnetic bead, the second magnetic bead, and the bound prostate cancer-derived exosome-magnetic bead complex. The supernatant is then discarded. The aggregated complex is washed with a pre-cooled washing buffer, which is a phosphate buffer containing a surfactant, for at least three washes.
5. The method for diagnosing prostate cancer using exosome surface protein profiling according to claim 1, characterized in that, In step 4, the oxidant used in the oxidation reaction is sodium periodate. The concentration of sodium periodate in the oxidation buffer ranges from 1 mM to 5 mM. The oxidation reaction is carried out under light-proof and low-temperature conditions, with a reaction temperature range from 0°C to 10°C and a reaction time range from 20 minutes to 40 minutes. In the biotinylate labeling reaction, the concentration of biotinylate in the labeling buffer ranges from 2 mM to 10 mM, the labeling reaction is carried out at room temperature, and the reaction time ranges from 45 minutes to 90 minutes. The lysis buffer used in the lysis reaction contains a nonionic surfactant and a reducing agent. The protease used in the enzymatic hydrolysis reaction is trypsin. The enzymatic hydrolysis reaction temperature range is 35°C to 38°C, and the enzymatic hydrolysis reaction time ranges from 12 hours to 18 hours.
6. The method for diagnosing prostate cancer using exosome surface protein profiling according to claim 5, characterized in that, In step 4, the specific enrichment process specifically includes: incubating the peptide mixture with streptavidin magnetic beads at room temperature for at least 1 hour; after incubation, applying a magnetic field to aggregate the streptavidin magnetic beads bound to biotin-labeled peptides; removing the unbound non-glycosylated peptide solution; eluting the aggregated streptavidin magnetic beads with an acidic elution buffer to dissociate and collect the biotin-labeled surface glycoprotein peptides; The acidic eluent is an aqueous solution containing 1% to 5% acetonitrile by volume and 0.05% to 0.2% formic acid by mass and volume.
7. The method for diagnosing prostate cancer using exosome surface protein profiling according to claim 1, characterized in that, In step 5, the liquid chromatography-tandem mass spectrometry analysis is performed using a nanoliter liquid chromatography system coupled with a high-resolution tandem mass spectrometer; The mass spectrometry data acquisition mode is either data-dependent or data-independent. In the data-dependent acquisition mode, the parent ion with the highest intensity is selected for fragmentation in real time, while in the data-independent acquisition mode, all ions are cyclically fragmented and scanned according to a predetermined mass-to-charge ratio window. The raw mass spectrometry data is used to identify peptides by searching a protein sequence database, and the surface glycoproteins corresponding to the peptides are relatively quantified by a label-free quantification method, generating the surface glycoprotein spectrum data containing a list of surface glycoprotein identifications and relative abundance information.
8. The method for diagnosing prostate cancer using exosome surface protein profiling according to claim 1, characterized in that, Step 6, the process of constructing the pre-trained diagnostic model specifically includes: Step A: Collect and obtain the surface glycoprotein spectrum data of pathologically confirmed prostate cancer patient samples as a positive training set, and the surface glycoprotein spectrum data of benign prostate disease patient samples as a negative training set; Step B: Normalize and standardize the quantitative data of all surface glycoproteins in the positive training set and the negative training set; Step C: Use machine learning algorithms to perform feature selection on the preprocessed data, and screen out multiple surface glycoproteins whose expression differences between the positive training set and the negative training set are determined to be significant by the selected machine learning algorithm to form a diagnostic feature panel; Step D: Based on the quantitative data of multiple surface glycoproteins in the diagnostic feature panel, a classification algorithm is used to train and generate the diagnostic model that can distinguish between prostate cancer and benign conditions, and a classification threshold is determined for the diagnostic model.
9. A method for diagnosing prostate cancer using exosome surface protein profiling according to claim 8, characterized in that, In step C, the machine learning algorithm is random forest, support vector machine, or LASSO regression algorithm; The diagnostic feature panel contains 5 to 15 surface glycoproteins; The diagnostic feature panel contains surface glycoproteins selected from the group consisting of: extracellular 5'-nucleotidase, prostate stem cell antigen, integrin α6, protein tyrosine kinase 7, growth differentiation factor 15, CD276 molecule, and L-amino acid transporter 1.
10. A method for diagnosing prostate cancer using exosome surface protein profiling according to any one of claims 1 to 9, characterized in that, The method is also used to assess the invasive risk or grade of prostate cancer; The training set of the diagnostic model further includes surface glycoprotein profile data of prostate cancer patients at different clinical stages or Gleason scores; The diagnostic results output in step 6 include the risk level or invasiveness grade of prostate cancer.