Analysis method and system

By combining third-generation sequencing and mass spectrometry analysis with a specific protein sequence database, the problems of incomplete transcript information and low protein detection throughput in single-cell omics have been solved. This enables direct association between unbiased, full-length transcripts and high-throughput peptide sequences, making it suitable for omics research in rare cells.

CN121905286APending Publication Date: 2026-04-21AIXINBO (BINHAI) BIOMEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AIXINBO (BINHAI) BIOMEDICAL TECH CO LTD
Filing Date
2025-12-29
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing single-cell omics analysis methods suffer from incomplete transcript information, low and biased protein detection throughput, and difficulty in correlating omics data. This is especially true for rare cell types such as CTC, where it is difficult to obtain full-length transcript information while simultaneously performing unbiased, high-throughput protein expression verification.

Method used

Third-generation sequencing technology was used to obtain full-length transcript information, and mass spectrometry analysis was combined to obtain high-throughput peptide sequence information. By comparing specific protein sequence databases with peptide sequence data, direct correlation between transcript and peptide levels was achieved. A customized database of the same cell type was used for verification and classification.

Benefits of technology

It achieves simultaneous acquisition of unbiased, full-length transcript information and high-throughput peptide sequence information, enabling direct correlation at the transcript-peptide level, verification of novel transcript isoforms and gene fusions, and adapting to the research needs of rare cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The embodiment of the invention relates to an analysis method and system, and the method comprises the following steps: a comparison step: carrying out comparison analysis based on specific protein sequence data and peptide fragment sequence data, the specific protein sequence data and the peptide fragment sequence data being derived from cells of the same type; and a result output step: outputting an analysis result. According to the method, unbiased and full-length transcript information and unbiased and high-throughput peptide fragment sequence information can be obtained from cells of the same type at the same time, and the transcript information and the peptide fragment sequence information can be directly associated on the transcript-peptide fragment level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics and relates to an analysis method and system. Background Technology

[0002] The main problems with existing single-cell genome analysis methods include: Incomplete transcript information: Traditional single-cell RNA sequencing (scRNA-seq) is based on the inherent defects of second-generation sequencing, resulting in short read lengths, making it difficult to accurately assemble full-length transcripts. At the same time, it cannot effectively identify complex information such as alternative splicing, gene fusion, and new isoforms.

[0003] Protein detection has low throughput and is biased: Existing single-cell protein detection technologies (such as CITE-seq and REAP-seq) rely on known antibodies and can only detect a few pre-defined proteins, making it impossible to perform unbiased, discovery-based proteomics analysis.

[0004] The challenge of linking omics data lies in the fact that transcriptomic and proteomic data come from different technology platforms. There is a lack of reliable methods for effectively integrating and accurately linking these two types of heterogeneous data from the same tissue and cell type. This is especially true for rare cell types like CTCs, where sample sizes are extremely low and heterogeneity is high, making it impossible to verify protein expression while obtaining full-length transcriptional information using traditional methods. Summary of the Invention

[0005] To address one of the aforementioned technical problems in the prior art, this invention provides an analysis method that can simultaneously obtain unbiased, full-length transcript information and unbiased, high-throughput peptide sequence information from the same type of cells, and can directly correlate the two at the transcript-peptide level.

[0006] Firstly, in some embodiments, an analysis method is provided, including: The comparison steps include comparison analysis based on a specific protein sequence database and peptide sequence data, wherein the specific protein sequence database and the peptide sequence data are derived from the same type of cell. Results output steps: including outputting the analysis results.

[0007] In some embodiments, the comparison step includes at least one of the following analyses: transcript statistics, gene type classification, marker transcript statistics, marker transcript classification, and protein database verification.

[0008] In some embodiments, the transcript statistics include counting transcripts and distinguishing between new transcripts and known transcripts.

[0009] In some embodiments, the gene type classification includes a classification analysis of transcript gene types.

[0010] In some embodiments, the marker transcript statistics include identifying and counting marker transcripts.

[0011] In some embodiments, the marker transcript classification includes classifying marker transcripts by cellular subgroup.

[0012] In some embodiments, the protein database verification includes comparing the self-tested transcripts with transcripts in a protein database to obtain verification results.

[0013] In some embodiments, the transcript statistics include counting all transcripts in each sample and distinguishing whether the transcripts are new transcripts to obtain new transcripts and known transcripts.

[0014] In some embodiments, the gene type classification includes classifying the gene types corresponding to the transcripts of each sample.

[0015] In some embodiments, the marker transcript statistics include counting the number of transcripts that can be used as markers for cell subpopulations.

[0016] In some embodiments, the tagged transcript classification includes classifying and statistically analyzing the tagged transcripts according to cell subpopulations.

[0017] In some embodiments, the protein database verification includes taking the intersection of the self-tested transcripts with the transcript sequences in the protein database and performing statistical analysis.

[0018] In some embodiments, the transcript statistics include: obtaining transcript information from the Seurat object, obtaining sample information, counting transcripts for each sample, loading a list of known transcripts from an external database, obtaining the cell index of the current sample, extracting the expression matrix of the current sample, calculating expressed transcripts, distinguishing between new transcripts and known transcripts, calculating statistics, and obtaining transcript statistics results.

[0019] In some embodiments, the gene type classification includes: obtaining a transcript-gene type mapping, separating new transcripts and known transcripts, initializing a result data frame, obtaining a list of transcripts for the sample, classifying them as new transcripts and known transcripts, statistically analyzing the gene type distribution of each type of transcript, obtaining the gene type of each transcript, calculating the percentage, and obtaining the gene type classification result.

[0020] In some embodiments, the marker transcript statistics include: using Seurat's FindAllMarkers function to identify marker genes, assuming that transcript-level data has been converted to gene-level data, or using transcript-level differential expression analysis, identifying marker transcripts of cell subpopulations, setting cell identities, searching for marker genes, screening for significant markers, counting the number of markers in each subpopulation, and obtaining marker transcript statistics results; In some embodiments, the labeled transcript classification includes: organizing labeled transcripts by cell subpopulation, creating a transcript-subpopulation association table, counting the number of cell subpopulations used as labels in each transcript, and obtaining the labeled transcript classification results.

[0021] In some embodiments, the protein database verification includes: loading transcript statistics into a protein database, obtaining all detected transcripts, separating new transcripts and known transcripts, comparing them with the protein database, verifying known transcripts and new transcripts, obtaining detailed matching information, checking for matches for each transcript, and obtaining verification results.

[0022] In some embodiments, the protein database includes, but is not limited to, the Uniprot database.

[0023] In some embodiments, the method for obtaining the specific protein sequence data includes, but is not limited to, first-generation, second-generation, or third-generation sequencing, preferably third-generation sequencing.

[0024] In some embodiments, the third-generation sequencing includes, but is not limited to, Nanopore sequencing.

[0025] In some embodiments, the method for obtaining the specific protein sequence data includes lysing cells, amplifying full-length cDNA, sequencing, and bioinformatics analysis to obtain full-length transcript information.

[0026] In some embodiments, the sequencing includes, but is not limited to, first-generation, second-generation, or third-generation sequencing, preferably third-generation sequencing.

[0027] In some embodiments, the analysis results include at least one of the following: transcript-peptide correspondence data association results, novel isoform verification results, fusion gene translation verification results, and transcript-peptide expression correlation verification results.

[0028] In some embodiments, the cells of the same type originate from the same tissue.

[0029] In some embodiments, the same type of cells includes rare cells.

[0030] In some embodiments, the rare cells include at least one of blood-derived rare cells and non-blood-derived rare cells.

[0031] In some embodiments, the non-hematogenous rare cells include circulating tumor cells, neural stem cells, mesenchymal stem cells, epithelial stem cells, or rare immune cell subtypes.

[0032] In some embodiments, the method for screening to obtain the same type of cells includes fluorescence in situ hybridization imaging analysis.

[0033] In some embodiments, the method for obtaining the peptide sequence data includes, but is not limited to, mass spectrometry analysis.

[0034] Secondly, in some embodiments, an analysis system is provided, comprising: The alignment module is used to perform alignment analysis based on specific protein data and peptide sequence data, wherein the specific protein data and the peptide sequence data originate from the same type of cell. Results output module: Used to output analysis results.

[0035] Thirdly, in some embodiments, an electronic device is provided, comprising: memory, and A processor coupled to the memory, the processor being configured to execute the method described in the first aspect based on instructions stored in the memory.

[0036] Fourthly, in some embodiments, a non-volatile computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method described in the first aspect.

[0037] Fifthly, in some embodiments, a drug target analysis method is provided, comprising: analyzing and obtaining targets that can be used to prepare drugs for treating diseases based on the results obtained by the method described in the first aspect.

[0038] In some embodiments, the disease includes at least one of cancer and inflammation.

[0039] In some embodiments, the cancer includes at least one of solid tumors and non-solid tumors.

[0040] In some embodiments, compared with the prior art, the beneficial effects of the present invention are: it can simultaneously obtain unbiased, full-length transcript information and unbiased, high-throughput peptide sequence information from the same type of cells, and can directly correlate the two at the transcript-peptide level. Attached Figure Description

[0041] Figure 1 A schematic diagram of the correlation analysis process is shown.

[0042] Figure 2The MYLK gene is shown.

[0043] Figure 3 The TTR gene is shown.

[0044] Figure 4 The UQCRQ gene is shown.

[0045] Figure 5 This diagram illustrates the combined application of full-length transcripts and proteomic profiles of the same cell type. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention in any way. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of this disclosure. Such structures and techniques have also been described in many publications.

[0047] definition "Non-hematopoietic rare cells" refer to rare cell types originating from non-blood tissues, typically found in specific organs or tissues rather than in circulating blood. For example, in hematopoietic system research, non-hematopoietic cells include osteoblasts, osteoclasts, adipocytes, endothelial cells, reticular cells, and tissue basophils (i.e., mast cells). These cells constitute a very low proportion of the blood or do not directly participate in blood cell circulation, but may enter the bloodstream under pathological conditions. In some embodiments, non-hematopoietic rare cells may also include, but are not limited to, circulating tumor cells, tumor stem cells, circulating endothelial cells, hematopoietic stem and progenitor cells, and specific immune cells such as antigen-specific T cells and iNKT cells. "Blood-derived rare cells" refer to rare cell types originating from the hematopoietic system. These cells are present in small numbers in the blood but perform important physiological or pathological functions. In some embodiments, blood-derived rare cells refer to cell types present in extremely low concentrations in the blood, typically less than 0.01% (i.e., no more than 10 per 100,000 blood cells). "Non-blooded" emphasizes non-blood-derived origin (such as tissue residency). Non-blood-derived rare cells can encompass tissue-specific cells, such as certain neuronal subtypes or epithelial stem cells.

[0048] Unless otherwise specified, "single-cell" in this article refers to cells derived from the same type, such as cells from the same tissue and of the same type. In some embodiments, an integration method and application system are provided that can simultaneously obtain unbiased, full-length transcript information and unbiased, high-throughput peptide sequence information from the same tissue and the same type of cells, and can directly correlate the two at the transcript-peptide level.

[0049] In some embodiments, a complete solution is provided, the process of which is as follows: Figure 1 As shown, by first obtaining a single-cell (i.e., cell of the same type) specific sequence library and identifying the obtained peptide sequences, a drug delivery target that can be used for CTC specificity can be obtained.

[0050] In some embodiments, a first-of-its-kind multi-omics integration platform for the same tissue and the same cell type is provided: for the first time, the advantages of long reads and full lengths of third-generation sequencing are combined with the advantages of unbiased protein discovery by mass spectrometry at the same tissue and the same cell type level.

[0051] In some embodiments, a revolutionary application of customized databases is realized: abandoning traditional general protein databases, a customized database is constructed using full-length transcript sequences generated by cells of the same type in the same tissue, for mass spectrometry data retrieval, which greatly improves the accuracy and sensitivity of peptide identification and can discover novel, cell-specific peptides.

[0052] In some embodiments, the present invention has a strong mutual verification capability, specifically including: From Transcripts to Proteins: Confirming whether novel transcript isoforms and gene fusions discovered by third-generation sequencing are truly translated into proteins (i.e., important evidence of “drug-worthy” targets).

[0053] From protein to transcript: Tracing back to the exact transcript variants of detected rare proteins.

[0054] Extremely low initial sample size requirement: Perfectly suited for research on rare cells such as CTCs, cells of the same type from the same tissue can simultaneously obtain two layers of omics information and establish correlations. Example

[0055] Unless otherwise stated, the present invention will be carried out using conventional techniques of cell biology, cell culture, molecular biology, and laboratory animal science, techniques described in the literature or performed in accordance with product instructions. For example, see J. Sambrook's *Molecular Cloning: A Laboratory Manual* (4th edition, Science Press). Unless otherwise specified, the materials, reagents, or instruments used in the examples are all commercially available conventional products, and other equivalents may be used.

[0056] 1. Split or parallel process individual cells.

[0057] Differential phase enrichment (SE) of non-hematogenous rare cells in body fluid specimens from tumor patients or tumor animal models (mice) (specifically, CTCs, circulating tumor cells in this example).

[0058] Using domestically produced or Zeiss fully automated 3D 6-channel image scanning systems, we can identify various aneuploid CTCs (circulating tumor cells) and CTECs (circulating tumor vascular endothelial cells) stained with i·FISH (immunofluorescence staining-chromosome fluorescence in situ hybridization).

[0059] Aneuploid CTC and CTEC subclasses with different clinical significance and expressing different tumor markers were identified and picked from the same glass slide.

[0060] Single cells of the same type isolated from the same tissue were placed in a 96-well plate and centrifuged at 800×g, 37℃, for 1 min to reach the bottom of the 96-well plate. 0.5 μl of Aixinbo enzyme digest was added to the well plate beforehand.

[0061] Place the 96-well plate in a 37°C constant temperature metal bath and let it stand for 2 hours to allow the same type of single cells to be fully lysed into peptides. After lysis, add 3.5 μl of Aixinbo stop acid solution to each well plate and then place the 96-well plate in a centrifuge at 800×g, 37°C, for 1 min to centrifuge the liquid to the bottom of the 96-well plate.

[0062] After centrifugation, the 96-well plate was placed in an ultra-low temperature freezer at -80°C.

[0063] 2. Specific kits, cycle numbers, and detailed protocols for third-generation sequencing library preparation for cDNA amplification of the same type of cells in the same tissue.

[0064] 2.1 Preparations before the experiment 2.1.1. Reagents & Consumables: RNA evaluation: Agilent RNA 6000 Pico Kit (Agilent, Code. 5067-1513) or other equivalent products; Magnetic bead purification: NucleoMag NGS Clean-up and Size Select (Macherey-Nagel, Code. 744970.50), magnetic rack; DNA evaluation: High Sensitivity DNA Kit (Agilent, Code. 5067-4626) and QubitdsDNA HS Assay Kit (Thermo Fisher Scientific, Cat. No. Q32854); DNA library preparation kits, such as Nextera XT DNA Library Preparation Kit (Illumina, Code. FC-131-1024).

[0065] Other materials: low-adsorption tubes, 80% ethanol, 0.2 ml RNase-free PCR tubes, 1.5 ml centrifuge tubes.

[0066] 2.1.2 Instruments: PCR instrument, Qubit 4 Fluorometer, Agilent Technologies 2100 Bioanalyzer, pipettes, vortex mixer, mini desktop centrifuge.

[0067] 2.2 Operating Procedures 2.2.1 First-strand cDNA Synthesis 2.2.1.1 All reagents required for thawing on ice After thawing the required reagents on ice, briefly centrifuge, mix well, and then place on ice for later use. For the 5' Template Swinging Oligo, AccuNext Reverse Transcriptase, and RNase Inhibitor, gently tap the tube wall or use a pipette to mix; do not vortex. The remaining reagents can be gently vortexed to mix.

[0068] 5X First-Strand Synthesis Buffer may form a precipitate. Shake well to dissolve the precipitate before use.

[0069] 2.2.1.2 Cell lysis and RNA denaturation and annealing 1) Prepare the Reaction Buffer on ice according to the table below. After preparation, gently mix with a pipette and collect by short-term centrifugation. Avoid generating air bubbles during mixing.

[0070] Table 1

[0071] Prepare the mixture according to Table 1 and name it 10X Reaction Buffer.

[0072] 2) Cell samples are selected for this step (using cells as a template, with 1 to 1000 cells in a single reaction): a) Prepare the reaction solution on ice according to the table below, and mix thoroughly; Table 2

[0073] *c: If resuspending cells in PBS: the added cell solution should not exceed 5 μL; because the culture medium and other components inhibit reverse transcription and PCR reactions, the cells should be resuspended in PBS without Mg2+. 2+ Ca 2+ Wash twice with 1×PBS and resuspend; add as little as possible during the reaction. The cell count should not exceed 1000 cells, as too many cells will inhibit the reaction.

[0074] If using flow cytometry to sort cells: cells can be directly sorted into the above 10.5 μL reaction solution and gently vortex to mix. Proceed to subsequent experiments immediately to avoid RNA degradation; if not proceeding immediately, the solution prepared in this step can be stored at -80°C. When using, follow step b).

[0075] b) Place the sample on ice, add 3' Oligo (dT) Primer*d, mix gently and centrifuge, then immediately proceed to step 4).

[0076] *d: If PCR amplification is less than 17 cycles, add 2 μL of 3' Oligo (dT) Primer. If the initial template volume is a single cell, or the number of PCR amplification cycles is ≥17, add 1 μL of 3' Oligo (dT) Primer and bring the volume to 12.5 μL using nuclease-free water.

[0077] 3) Following the procedure below, place the sample into a preheated PCR instrument for denaturation and annealing.

[0078] Table 3

[0079] 4) Incubate at 72℃ for 3 min to finish. After the PCR instrument temperature drops to 4℃, immediately place the sample on ice.

[0080] 5) During incubation, follow step 2.2.1.3. Reverse transcription reaction, and first prepare the RT Master Mix (AccuNextReverse Transcriptase is added before the template is added).

[0081] 6) After step 5), immediately proceed with step 2.2.1.3 below: reverse transcription reaction.

[0082] 2.2.1.3. Reverse transcription reaction 1) Prepare RT Master Mix: Table 4

[0083] *e: When performing steps 2) to 5) in 2.2.1.2, these four groups can be prepared into a mixture (AccuNext Reverse Transcriptase is added before the template is added). After gently mixing, take 7.5 μL of the mixture and add it to the denatured and annealed reaction solution in step 3) of 2.2.1.2 above. Gently mix with a pipette or shaker and centrifuge. Proceed to the subsequent steps immediately.

[0084] 2) Run the following program in the PCR instrument (set the reaction program in advance): Table 5

[0085] 3) Before the next reaction is ready, place the reaction product on ice or store it overnight at -20°C (to avoid cDNA degradation, proceed with the next steps as soon as possible).

[0086] Safety pause point: Samples can be stored overnight at -20℃, and for long-term storage, they can be stored at -80℃.

[0087] 2.3 PCR amplification of cDNA 2.3.1. Thaw 2X AccuNext PCR Buffer and PCR Primers on ice, gently vortex to mix, centrifuge, and place on ice. Gently pipette AccuNext DNA Polymerase to mix and place on ice for later use.

[0088] 2.3.2. Prepare the PCR Master Mix on ice according to the table below: Table 6

[0089] *f: These four components can be prepared into a mixture first, gently vortexed to mix, and then 30 μl of the mixture can be added to the above 20 μL cDNA synthesis product.

[0090] 2.3.3. Gently tap the tube wall to mix or gently vortex to mix, then centrifuge.

[0091] 2.3.4. Immediately run the following reaction program in the PCR instrument (set the reaction program in advance before preparing the PCR Master Mix): Table 7

[0092] 2.3.5. After the reaction is complete, place the product on ice in preparation for the DNA purification step (to avoid DNA degradation, proceed with the subsequent steps as soon as possible).

[0093] Safety pause point: Samples can be stored overnight at -20℃, and for long-term storage, they can be stored at -80℃.

[0094] 2.4 Purification of DNA Amplification Products PCR products were purified using magnetic beads. Taking the NucleoMag kit for cleanup and size selection of NGS library prep reactions (Macherey-Nagel, Code. 744970.50) as an example, the steps are as follows: Preparation before the experiment: For first-time use, the magnetic beads can be dispensed into 1.5 ml centrifuge tubes and stored at 4℃, depending on the experimental requirements. Before each experiment, fresh 80% ethanol can be prepared according to the experimental volume, with 400 μL required for each sample.

[0095] Operating steps: 2.4.1. After thoroughly vortexing and mixing the magnetic beads, place them at room temperature for equilibration for 30 minutes. After equilibration, vortex and mix them again. 2.4.2. Add 50 μL of equilibrated magnetic beads to the above 50 μL PCR amplification product (according to a magnetic bead:sample volume ratio of 1:1), vortex to mix thoroughly, and centrifuge briefly. 2.4.3. Incubate the magnetic bead / DNA mixture at room temperature for 10 minutes to allow the DNA to bind to the magnetic beads; 2.4.4. Place the sample on the magnetic rack for at least 5 minutes, until the liquid is completely clear and there are no magnetic beads in the supernatant. Carefully remove the supernatant without breaking the magnetic beads; 2.4.5. Keep the PCR tube on the magnetic rack at all times, add 200 μL of freshly prepared 80% ethanol (adding ethanol will not interfere with the magnetic beads), incubate at room temperature for 30 seconds, and carefully remove the supernatant; 2.4.6. Repeat step 5) once; 2.4.7. Keep the PCR tubes on the magnetic rack at all times, and open the caps and dry the magnetic beads for 5 to 10 minutes until there is no ethanol residue*h; *h: When there is no ethanol residue, the surface of the magnetic beads is dull; if the ethanol is not completely dried, it may affect the elution efficiency of DNA and downstream experiments. However, the magnetic beads should not be over-dried to avoid surface cracking and reduced DNA elution efficiency.

[0096] 2.4.8. After the magnetic beads have dried, remove the PCR tube from the magnetic rack, add 17 μL of Nuclease-free water to cover the magnetic beads, mix the magnetic beads by pipetting, and incubate at room temperature for 2 min (if the magnetic beads dry and crack, extend the incubation time appropriately). 2.4.9. Briefly centrifuge the PCR tube, place it on a magnetic rack, and separate the magnetic beads and liquid until the solution is clear (about 5 min); 2.4.10. Carefully aspirate 15 μL of the supernatant into a new low-adsorption tube (avoiding magnetic beads) and store at -20°C. To prevent DNA degradation, construct the library as soon as possible.

[0097] Safety pause point: Samples can be stored overnight at -20℃, and for long-term storage, they can be stored at -80℃.

[0098] 2.5 Quality Detection of DNA Amplification Products 2.5.1. Take 2 μl of the purified DNA amplification product and use a Qubit 4 Fluorometer and Qubit1X dsDNA HS Assay Kit (Thermo Fisher, Code. Q33231) to detect the concentration of the PCR product. Please refer to the product instructions for specific procedures.

[0099] 2.5.2. Take 1 μL of the purified DNA amplification product and verify it using the Agilent Technologies 2100 Bioanalyzer and Agilent High Sensitivity DNA Kit (Agilent, Code. 5067-4626). Please refer to the product instructions for specific procedures.

[0100] 2.5.3. Depending on the amount of template added, a successful experimental reaction will produce 2 to 20 ng of amplification product, with fragment sizes ranging from 400 to 10,000 bp and a peak size around 1,000 to 2,000 bp. The negative control produces no amplification product.

[0101] 5. Single-cell proteomics pretreatment steps (such as low-volume digestion, carrier-free protocols).

[0102] 5.1 Using a low-adsorption pipette tip, aspirate 2.5 μL of Ixinbo enzyme digest into cells of the same type from the same tissue; 5.2 Place the isolated single cells in a centrifuge at 800×g, 37℃, for 1 min to centrifuge the single cells to the bottom of the 96-well plate.

[0103] 5.3 Place the 96-well plate in a 37°C constant temperature metal bath for 2 hours to allow individual cells to fully lyse into peptides. If a constant temperature metal bath is not available, place it at room temperature for 4 hours for lysis.

[0104] 5.4 After lysis, add 7 μL of Ixinbo stop acid solution to each well plate, and then place the 96-well plate in a centrifuge at 800×g, 37℃, for 1 min to centrifuge the liquid to the bottom of the 96-well plate.

[0105] 5.5 Place the 96-well plate after centrifugation in an ultra-low temperature freezer at -80℃.

[0106] 6. Mass spectrometry data retrieval parameters, false positive rate control, and scripts for building customized databases (e.g., how to translate transcript sequences into 6-frame or 3-frame protein sequence databases).

[0107] LC-MS / mass spectrometry analysis was performed using an Orbitrap Astral mass spectrometer coupled with an EvoSep One system (EvoSep Biosystems). Samples were analyzed at 40 SPD (31 min gradient) using commercial analytical columns (AuroraElite TS, IonOpticks) connected online via an EASY-Spray™ source. The Orbitrap Astral mass spectrometer had a full MS resolution of 240,000 m / s and a full scan range of 380–980 m / s. Full MS AGC was set to 500%. The isolation window for MS / MS scan recording was 2Th, and the maximum ion implantation time (IIT) for CTC samples was 3 ms. The MS / MS scan range was 380–980 m / s. Separated ions were fragmented using an HCD containing 27% NCE.

[0108] Methods for visualizing associated results.

[0109] The specific methods for association analysis are as follows: # Step 1: Transcription statistics, corresponding to the functions in Table 8 # Step 2: Gene type classification, achieving functional correspondence (Table 9) # Step 3: Marker transcript statistics, corresponding to the functions in Table 10 # Step 4: Marker transcript classification, achieving the functional correspondence in Table 11 # Step 5: Uniprot database verification, implement functions corresponding to Table 12-15 Associative analysis scripts can be displayed: Implementation of a Single-Cell Transcript Analysis System in R Code # Filename: sc_transcript_analysis_system.R # Load necessary packages library(dplyr) library(tidyr) library(ggplot2) library(data.table) library(Biostrings) library(AnnotationDbi) library(org.Hs.eg.db) # Function 1: Main function for transcript statistics #' Single-cell transcript analysis system main function #' @param seurat_obj Seurat single-cell object #' @param uniprot_db Uniprot database file path #' @param output_dir Output directory #' @return A list containing all statistical results sc_transcript_analysis_system<- function(seurat_obj, uniprot_db =NULL, output_dir = "results") { cat("=== Single-cell transcript analysis system v1.0 ===\n") cat("Start Time:", format(Sys.time(), "%Y-%m-%d %H:%M:%S"), "\n") # Create output directory if (!dir.exists(output_dir)) { dir.create(output_dir, recursive = TRUE) } # Initialize the list of results results<- list() # Step 1: Transcription Statistics cat("\n[Step 1] Count transcripts and distinguish between new and known transcripts...\n") results$table1<- step1_transcript_statistics(seurat_obj) # Step 2: Gene Type Classification cat("\n[Step 2] Transcript Gene Type Classification Analysis...\n") results$table2<- step2_gene_type_classification(seurat_obj, results$table1) # Step 3: Marker Transcript Statistics cat("\n[Step 3] Identify and count marker transcripts...\n") results$table3<- step3_marker_transcript_identification(seurat_obj) # Step 4: Marker Transcript Classification cat("\n[Step 4] Marker transcripts are categorized by cell subsets...\n") results$table4<- step4_marker_by_subpopulation(results$table3) # Step 5: Uniprot database verification if (!is.null(uniprot_db)) { cat("\n[Step 5] Verify against the Uniprot database...\n") results$table5<- step5_uniprot_validation(results$table1, uniprot_db) } # Generate a comprehensive report cat("\n[Step 6] Generate analysis report...\n") results$report<- generate_comprehensive_report(results, output_dir) cat("Analysis complete! Results saved in directory:", output_dir, "\n") cat("Completion Time:", format(Sys.time(), "%Y-%m-%d %H:%M:%S"), "\n") return(results) } # Step 1: Transcription Statistics #' Statistical analysis of transcripts and differentiation between new and known transcripts #' @description Count all transcripts in each sample and distinguish whether they are new transcripts based on the transcripts. #' @param seurat_obj Seurat object #' @return Transcription Statistics Table 1 step1_transcript_statistics<- function(seurat_obj) { # Retrieve transcript information from a Seurat object transcript_counts<- as.data.frame(seurat_obj@assays$RNA@counts) transcript_metadata<- seurat_obj@assays$RNA@meta.features # Get sample information sample_info<- as.data.frame(seurat_obj@meta.data$orig.ident) colnames(sample_info)<- "sample" # Count the transcripts of each sample transcript_stats<- data.frame( sample = character(), total_transcripts = integer(), novel_transcripts = integer(), known_transcripts = integer(), novel_percentage = numeric(), stringsAsFactors = FALSE ) # In practical applications, this involves loading a list of known transcripts from an external database. known_transcripts<- load_known_transcript_database() for (sample in unique(sample_info$sample)) { # Get the cell index of the current sample sample_cells<- which(sample_info$sample == sample) # Extract the expression matrix of the current sample sample_expr<- transcript_counts[, sample_cells, drop = FALSE] # Calculate expressed transcripts (expression level > 0) expressed_transcripts<- rownames(sample_expr)[rowSums(sample_expr)>0] # Differentiate between new transcripts and known transcripts novel_transcripts<- setdiff(expressed_transcripts, known_transcripts) known_in_sample<- intersect(expressed_transcripts, known_transcripts) # Calculate statistics total_count<- length(expressed_transcripts) novel_count<- length(novel_transcripts) known_count<- length(known_in_sample) novel_pct<- ifelse(total_count>0, novel_count / total_count * 100,0) # Add to results table transcript_stats<- rbind(transcript_stats, data.frame( sample = sample, total_transcripts = total_count, novel_transcripts = novel_count, known_transcripts = known_count, novel_percentage = round(novel_pct, 2) )) } # Save results write.csv(transcript_stats, file.path(output_dir, "table1_transcript_statistics.csv"), row.names = FALSE) # Visualization plot_transcript_statistics(transcript_stats, output_dir) return(transcript_stats) } # Step 2: Gene Type Classification # Transcript Gene Type Classification #' @description Classify the gene types corresponding to the transcripts of each sample. #' @param seurat_obj Seurat object #' @param transcript_stats Results of Step 1 #' @return Gene type statistics table step2_gene_type_classification<- function(seurat_obj, transcript_stats) { # Obtain Transcript-Gene Type Mapping transcript_info<- get_transcript_annotation(seurat_obj) # Separating new transcripts from known transcripts known_transcripts<- load_known_transcript_database() # Initialize the result data frame gene_type_stats <- data.frame( sample = character(), gene_type = character(), transcript_class = character(), count = integer(), percentage = numeric(), stringsAsFactors = FALSE ) for (sample in unique(transcript_stats$sample)) { # Get the transcript list for this sample sample_transcripts<- get_sample_transcripts(seurat_obj, sample) # Classified as new / known transcripts novel_in_sample<- setdiff(sample_transcripts, known_transcripts) known_in_sample<- intersect(sample_transcripts, known_transcripts) # Statistical analysis of gene type distribution of various transcripts for (transcript_class in c("novel", "known")) { transcripts<- if (transcript_class == "novel") novel_in_sampleelse known_in_sample if (length(transcripts)>0) { # Obtain the gene type of these transcripts transcript_subset<- transcript_info[transcript_info$transcript_id %in% transcripts, ] if (nrow(transcript_subset)>0) { type_counts<- as.data.frame(table(transcript_subset$gene_type)) colnames(type_counts)<- c("gene_type", "count") # Calculate percentage type_counts$percentage<- round(type_counts$count / sum(type_counts$count) * 100, 2) type_counts$sample<- sample type_counts$transcript_class<- transcript_class gene_type_stats<- rbind(gene_type_stats, type_counts) } } } } # Save results write.csv(gene_type_stats, file.path(output_dir, "table2_gene_type_classification.csv"), row.names = FALSE) # Visualization plot_gene_type_distribution(gene_type_stats, output_dir) return(gene_type_stats) } # Step 3: Marker Transcript Identification #' Count the number of marker transcripts #' @description Count the number of transcripts that can serve as cell subset markers #' @param seurat_obj Seurat object #' @return marker transcript statistics table step3_marker_transcript_identification<- function(seurat_obj) { # Identify marker genes using Seurat's FindAllMarkers function # This assumes that transcript-level data has been converted to gene-level data. # Alternatively, use differential expression analysis at the transcript level. cat("marker transcripts for recognizing cell subsets...\n") # Set cell identity Idents(seurat_obj)<- seurat_obj$seurat_clusters # Searching for marker genes markers <- FindAllMarkers( seurat_obj, only.pos = TRUE, min.pct = 0.25, logfc.threshold = 0.25 ) # Filter for prominent markers significant_markers<- markers %>% filter(p_val_adj<0.05) %>% group_by(cluster) %>% slice_max(n = 50, order_by = avg_log2FC) # Count the number of markers in each cluster marker_counts<- significant_markers %>% group_by(cluster) %>% summarise( marker_count = n(), avg_log2FC_mean = mean(avg_log2FC), pct_exp_mean = mean(pct.1) ) # Save results write.csv(significant_markers, file.path(output_dir, "table3_marker_transcripts.csv"), row.names = FALSE) write.csv(marker_counts, file.path(output_dir, "table3_marker_counts_by_cluster.csv"), row.names = FALSE) # Visualization plot_marker_statistics(significant_markers, marker_counts, output_dir) return(list( marker_details = significant_markers, cluster_summary = marker_counts )) } # Step 4: Marker Transcript Classification #'Marker transcripts categorized by cell subpopulation' #' @description Transcripts used as markers will be categorized and statistically analyzed according to cell subpopulations. #' @param marker_results Results of step 3 #' @return Marker transcript table categorized by cell subpopulation step4_marker_by_subpopulation<- function(marker_results) { marker_details<- marker_results$marker_details # Organize marker transcripts by cell cluster marker_by_cluster<- marker_details %>% group_by(cluster) %>% summarise( marker_genes = paste(gene, collapse = ", "), top_5_markers = paste(head(gene, 5), collapse = ", "), count = n(), avg_log2FC = mean(avg_log2FC), max_log2FC = max(avg_log2FC) ) %>% arrange(desc(count)) # Create a detailed transcript-subgroup association table transcript_cluster_map<- marker_details %>% select(gene, cluster, avg_log2FC, pct.1, pct.2, p_val_adj) %>% arrange(cluster, desc(avg_log2FC)) # Count how many clusters each transcript serves as a marker transcript_frequency<- marker_details %>% group_by(gene) %>% summarise( cluster_count = n(), clusters = paste(sort(unique(cluster)), collapse = ","), max_avg_log2FC = max(avg_log2FC), min_p_val_adj = min(p_val_adj) ) %>% arrange(desc(cluster_count), desc(max_avg_log2FC)) # Save the results write.csv(marker_by_cluster, file.path(output_dir, "table4_markers_by_cluster.csv"), row.names = FALSE) write.csv(transcript_cluster_map, file.path(output_dir, "table4_transcript_cluster_mapping.csv"), row.names = FALSE) write.csv(transcript_frequency, file.path(output_dir, "table4_transcript_frequency.csv"), row.names = FALSE) # Visualization plot_marker_cluster_distribution(marker_by_cluster, output_dir) return(list( cluster_summary = marker_by_cluster,[[ID=4来1]] transcript_map = transcript_cluster_map, frequency = transcript_frequency )) } # Step 5: Uniprot database verification #' @description This involves statistically analyzing the intersection of self-tested transcripts and transcript sequences from the Uniprot database. #' @param transcript_stats Transcript statistics #' @param uniprot_db Uniprot database file path #' @return Verification Result Table step5_uniprot_validation<- function(transcript_stats, uniprot_db) { cat("Loading Unirot database...\n") # Load the Uniprot database uniprot_data<- load_uniprot_database(uniprot_db) # Get all detected transcripts all_transcripts<- get_all_detected_transcripts(seurat_obj) # Isolate new / known transcripts known_db<- load_known_transcript_database() novel_transcripts<- setdiff(all_transcripts, known_db) known_transcripts<- intersect(all_transcripts, known_db) # Comparison with Uniprot validation_results <- data.frame(transcript_class = character(), total_count = integer(), uniprot_match = integer(), uniprot_percentage = numeric(), stringsAsFactors = FALSE) # Validate known transcripts if (length(known_transcripts) > 0) { known_match <- intersect(known_transcripts, uniprot_data$transcript_ids) known_pct <- length(known_match) / length(known_transcripts) * 100 validation_results <- rbind(validation_results, data.frame( transcript_class = "known", total_count = length(known_transcripts), uniprot_match = length(known_match), uniprot_percentage = round(known_pct, 2) )) } # Validate novel transcripts if (length(novel_transcripts) > 0) { novel_match <- intersect(novel_transcripts, uniprot_data$transcript_ids) novel_pct <- length(novel_match) / length(novel_transcripts) * 100 validation_results <- rbind(validation_results, data.frame( transcript_class = "novel", total_count = length(novel_transcripts), uniprot_match = length(novel_match), uniprot_percentage = round(novel_pct, 2) )) } # Detailed matching information detailed_matches<- data.frame( transcript_id = character(), transcript_class = character(), in_uniprot = logical(), uniprot_accession = character(), protein_name = character(), stringsAsFactors = FALSE ) # Check matches for each transcript for (transcript in all_transcripts) { match_info<- check_uniprot_match(transcript, uniprot_data) detailed_matches<- rbind(detailed_matches, data.frame( transcript_id = transcript, transcript_class = ifelse(transcript %in% novel_transcripts, "novel", "known"), in_uniprot = match_info$matched, uniprot_accession = ifelse(match_info$matched, match_info$accession, NA), protein_name = ifelse(match_info$matched, match_info$protein_name, NA) )) } # Save results write.csv(validation_results, file.path(output_dir, "table5_uniprot_validation_summary.csv"), row.names = FALSE) write.csv(detailed_matches, file.path(output_dir, "table5_detailed_uniprot_matches.csv"), row.names = FALSE) # Visualization plot_uniprot_validation(validation_results, detailed_matches,output_dir) return(list( summary = validation_results, details = detailed_matches )) } } main_example<- function() { # Load necessary packages library(Seurat) library(dplyr) # Load test data seu<- readRDS("path / to / your / seurat_object.rds") # Set output directory output_dir<- "transcript_analysis_results" # Running Analysis System results<- sc_transcript_analysis_system( seurat_obj = seu, uniprot_db = "path / to / uniprot_database.fasta", output_dir = output_dir ) # View results print("=== Summary of Analysis Results===") print(paste("Number of samples:", nrow(results$table1))) print(paste("Total Transcripts:", sum(results$table1$total_transcripts))) print(paste("New Transcription Percentage:", mean(results$table1$novel_percentage),"%")) print(paste("Number of marker transcripts found:", nrow(results$table3$marker_details))) print(paste("Uniprot verification match rate:", results$table5$summary$uniprot_percentage[1], "%")) return(results) } 7.2 Relevant statistics on transcripts: Table 8

[0110] Table 9

[0111] Table 10

[0112] Table 11

[0113] 7.3 Transcriptional products of mass spectrometry data: Table 12

[0114] Table 13

[0115] Table 14

[0116] Table 15

[0117] 7.4 Transcriptional-associated mass spectrometry: MYLK gene, such as Figure 2 As shown.

[0118] 7.5 Transcriptional-associated mass spectrometry: TTR genes such as Figure 3 As shown.

[0119] 7.6 Transcriptional association mass spectrometry: UQCRQ gene, such as Figure 4 As shown.

[0120] 7.7 Combined application of full-length transcripts and proteomic profiles of the same cell type, such as... Figure 5 As shown.

[0121] 7.8 Results Explanation: 1. Transcript data for the full length of the same cell type: First, Table 8 shows the information of all transcripts in the full-length transcriptome of the same type of cells in this sample. Taking patients S1 and S2 as examples, S1-Nor and S2-Nor are their adjacent normal tissues; S1-Tumor and S2-Tumor are their tumor tissues. Table 8 shows the total number of known transcripts and the total number of unknown transcripts. Secondly, as shown in Table 9, the genes whose transcript information we detected generally come from the following categories: secretory proteins, secretory membrane proteins, secretory protein transcription factors, membrane proteins, membrane protein transcription factors, and transcription factors. Table 10 shows transcripts that can serve as single-cell subset markers: Marker_No cannot be used as a transcript for identifying specific subsets, while Marker_Yes can be used to identify specific subsets. Table 11 is a detailed breakdown of which subgroups have specific transcripts based on the Marker_Yes in Table 10; 2. Mass spectrometry data for single-cell proteins: Table 12: Verification of whether all transcript information detected in Table 8 above has the corresponding protein-coding product production: Among them, Pro_No revealed no protein production of the corresponding products, including 104965 new transcripts, 11974 known transcripts, and 116939 unknown transcripts; Pro_Yes revealed corresponding transcripts with corresponding protein production, including 18856 new transcripts, 3457 known transcripts, and 22313 unknown transcripts. Table 13 shows the corresponding protein-coding products that transcripts that can act as maker can produce; Table 14: The maker transcripts that produce proteins mainly come from the following: secretory proteins, membrane proteins, and transcription factors; Table 15: The maker transcripts that produce proteins mainly come from the cell subpopulations listed in this table.

[0122] In summary, based on the full-length single-cell transcriptome data obtained in step 1 and the single-cell proteome data obtained in step 2, the data was integrated and analyzed using the relevant scripts for correlation analysis mentioned in step 7.1 to obtain the transcript information of three representative CTCs in steps 7.4-7.6. It was also verified that a certain transcript of this gene showed upregulation and high expression in our data. This demonstrates that the experimental and analytical workflow provided in this embodiment can be used for rare cells, such as circulating tumor cells (CTCs), to perform in-depth molecular characterization and identify their unique targets, providing new ideas for subsequent drug development.

[0123] 3. Explanation of the relevance to CTC research In some embodiments, this invention is transformative for CTC research, and its relevance is reflected in: Unraveling the extreme heterogeneity of CTCs: CTCs are not homogeneous in patients, with significant differences in transcriptional and protein expression profiles among different CTCs. This invention allows for the most in-depth joint analysis of individual CTCs (i.e., CTCs of the same type), revealing their unique molecular characteristics and thus understanding the microscopic mechanisms of tumor metastasis.

[0124] Directly verifying the function of oncogenic driver variants: In CTCs, third-generation sequencing may reveal a large number of gene fusions, point mutations, and novel alternative splicing isoforms. However, whether these variants are "functional" can be directly demonstrated by detecting the specific peptides they encode, proving that these oncogenic transcripts have been successfully translated into proteins, thus confirming that they are true driving factors rather than non-functional transcriptional "noise." This is the gold standard evidence for identifying druggable targets.

[0125] Discovering novel biomarkers: Through an unbiased discovery model, novel peptide / protein variants encoded by new transcripts can be found in CTCs that have never been reported before. These molecules have great potential to become novel liquid biopsy biomarkers with high specificity and sensitivity for early cancer diagnosis, treatment monitoring, and prognosis.

[0126] Guiding Precision Treatment: By analyzing multiple CTCs from the same patient, a target atlas of the CTC population can be created. For example, it can identify which CTCs express specific drug targets (such as specific isoforms or fusion proteins of HER2), thus matching the patient with the most effective targeted drugs. Simultaneously, it can also reveal protein variants that lead to drug resistance, allowing for timely adjustments to treatment plans.

[0127] Understanding the biological processes of metastasis: Association analysis allows us to investigate which transcripts in CTCs are efficiently translated (e.g., proteins associated with migration, invasion, and drug resistance) and which are post-transcriptionally regulated. This provides an unprecedented perspective on understanding the mechanisms by which tumor cells survive in the bloodstream and colonize distant sites.

[0128] In some embodiments, this invention ushers in a new era for CTC and even the entire rare cell research. It connects the previously isolated two layers of omics information (transcripts and proteins) at the single-cell level, producing a "1+1>>2" effect, and providing a powerful tool for basic cancer research and clinical translational applications.

[0129] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.

Claims

1. An analytical method, characterized in that, include: The comparison steps include comparison analysis based on specific protein sequence data and peptide sequence data, wherein the specific protein sequence data and the peptide sequence data originate from the same type of cell. Results output steps: including outputting the analysis results.

2. The method as described in claim 1, characterized in that, In the comparison step, the comparison analysis includes at least one of the following analyses: transcript statistics, gene type classification, marker transcript statistics, marker transcript classification, and protein database verification.

3. The method as described in claim 2, characterized in that, The transcript statistics include statistical analysis of transcripts and differentiation between new transcripts and known transcripts; Optionally, the gene type classification includes a classification analysis of transcript gene types; Optionally, the marker transcript statistics include identifying and counting marker transcripts; Optionally, the marker transcript classification includes classifying marker transcripts by cellular subpopulation; Optionally, the protein database verification includes comparing the transcripts obtained from the self-test with transcripts in the protein database to obtain verification results.

4. The method as described in claim 2 or 3, characterized in that, The transcript statistics include counting all transcripts in each sample, and distinguishing whether the transcripts are new transcripts to obtain new transcripts and known transcripts; Optionally, the gene type classification includes classifying the gene types corresponding to the transcripts of each sample; Optionally, the marker transcript statistics include counting the number of transcripts that can be used as markers for cell subpopulations; Optionally, the marker transcript classification includes classifying and statistically analyzing the marker transcripts according to cell subpopulations; Optionally, the protein database verification includes taking the intersection statistics of the transcripts obtained from the self-test and the transcript sequences in the protein database.

5. The method as described in claim 2 or 3, characterized in that, The transcript statistics include: obtaining transcript information from the Seurat object, obtaining sample information, counting transcripts for each sample, loading a list of known transcripts from an external database, obtaining the cell index of the current sample, extracting the expression matrix of the current sample, calculating expressed transcripts, distinguishing between new transcripts and known transcripts, calculating statistics, and obtaining transcript statistics results. Optionally, the gene type classification includes: obtaining transcript-gene type mapping, separating new transcripts and known transcripts, initializing the result data frame, obtaining the transcript list of the sample, classifying it into new transcripts and known transcripts, statistically analyzing the gene type distribution of each type of transcript, obtaining the gene type of each transcript, calculating the percentage, and obtaining the gene type classification result; Optionally, the marker transcript statistics include: using Seurat's FindAllMarkers function to identify marker genes, assuming that transcript-level data has been converted to gene-level data, or using transcript-level differential expression analysis to identify marker transcripts in cell subpopulations, setting cell identities, searching for marker genes, screening for significant markers, counting the number of markers in each subpopulation, and obtaining marker transcript statistics results; Optionally, the labeled transcript classification includes: organizing labeled transcripts according to cell subpopulations, creating a transcript-subpopulation association table, counting the number of cell subpopulations used as labels in each transcript, and obtaining the labeled transcript classification results; Optionally, the protein database verification includes: loading transcript statistics into a protein database, obtaining all detected transcripts, separating new transcripts and known transcripts, comparing them with the protein database, verifying known transcripts and new transcripts, obtaining detailed matching information, checking the match for each transcript, and obtaining the verification result; Optionally, the protein database includes the Uniprot database; Optionally, the method for obtaining the specific protein sequence data includes third-generation sequencing; Optionally, the third-generation sequencing includes Nanopore sequencing; Optionally, the method for obtaining the specific protein sequence data includes lysing cells, amplifying full-length cDNA, sequencing, and bioinformatics analysis to obtain full-length transcript information; Optionally, the sequencing includes third-generation sequencing.

6. The method as described in claim 1, characterized in that, The analysis results include at least one of the following: transcript-peptide correspondence data correlation results, novel isoform verification results, fusion gene translation verification results, and transcript-peptide expression correlation verification results. Optionally, the cells of the same type are derived from the same tissue; Optionally, the same type of cells includes rare cells; Optionally, the rare cells include at least one of blood-derived rare cells and non-blood-derived rare cells; Optionally, the non-hematogenous rare cells include circulating tumor cells, neural stem cells, mesenchymal stem cells, epithelial stem cells, or rare immune cell subsets. Optionally, the method for screening to obtain cells of the same type includes fluorescence in situ hybridization imaging analysis; Optionally, the method for obtaining the peptide sequence data includes mass spectrometry analysis.

7. An analysis system, characterized in that, include: The alignment module is used to perform alignment analysis based on specific protein data and peptide sequence data, wherein the specific protein data and the peptide sequence data originate from the same type of cell. Results output module: Used to output analysis results.

8. An electronic device, characterized in that, include: Memory, and A processor coupled to the memory, the processor being configured to perform the method of any one of claims 1 to 6 based on instructions stored in the memory.

9. A non-volatile computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of any one of claims 1 to 6.

10. A method for analyzing drug targets, characterized in that, include: The results obtained by the method according to any one of claims 1 to 6 are analyzed to obtain targets that can be used to prepare drugs for treating diseases; Optionally, the disease includes at least one of cancer and inflammation.