PathSeq microbial identification and abundance detection method and system based on single cell transcriptome sequencing and application of PathSeq microbial identification and abundance detection method and system

By matching and visually adjusting the results of PathSeq software with single-cell transcriptome data, a one-click analysis workflow is generated, which solves the problems of accuracy and efficiency in microbial detection in single-cell transcriptome sequencing, realizes efficient and accurate microbial identification and abundance detection, and expands the application of advanced single-cell transcriptome analysis.

CN120954501APending Publication Date: 2025-11-14SHANGHAI OE BIOTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410590721.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-13
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing single-cell transcriptome sequencing technologies have low accuracy and uncertain results in microbial detection. The lack of a PathSeq-based microbial detection and analysis workflow leads to cumbersome operation and uncertain results.

Method used

The results from PathSeq software were matched with the RDS analysis files of single-cell transcriptomics and visualized to generate a one-click analysis workflow, including data preparation, quality filtering, microbial classification and result processing. The GATK pipeline was used for microbial identification and abundance assessment, and the data was integrated and visualized using the standard single-cell transcriptomics analysis software Seurat.

Benefits of technology

It improves the accuracy and efficiency of microbial detection, reduces the error rate, broadens the application scope of advanced single-cell transcriptome analysis, increases workflow speed by more than 50%, simplifies operation steps, and reduces the requirements for programming skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004837188280000081
    Figure BDA0004837188280000081
  • Figure BDA0004837188280000092
    Figure BDA0004837188280000092
  • Figure BDA0004837188280000101
    Figure BDA0004837188280000101
Patent Text Reader

Abstract

The invention discloses a PathSeq microbial identification and abundance detection method based on single cell transcriptome sequencing. The method comprises the following steps: preparing a bam file which stores sequencing data which is compared to a reference genome; carrying out PathSeq analysis by utilizing a GATK pipeline, and identifying the existence and abundance of microorganisms in the evaluation sample; a result file obtained through Pathseq analysis is sorted, and information about microbial infection and expression on the single cell level is obtained; preparing and reading an RDS file, wherein the RDS file comprises a cell name and sample information; importing and summarizing UMI and sequencing data information of all samples; counting the microbial content data of the whole sample level; counting content data of different microbial genus at a sample level; counting the conditions of the first ten microorganisms and drawing a visual graph; a Bacteriareads distribution diagram and a BacteriaUMIs distribution diagram and a violin diagram in the sample are drawn; and storing an RDS file containing information of different microbe genus. The invention further discloses a system for implementing the method and application of the method or the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of single-cell transcriptome sequencing technology, and relates to a method, system and application of PathSeq microbial identification and abundance detection based on single-cell transcriptome sequencing. Background Technology

[0002] Microorganisms within a host organism (such as a human) include bacteria, fungi, and viruses, forming a complex symbiotic relationship with the host. Taking humans as an example, human microorganisms aid digestion, maintain the immune system, prevent infection, produce physiologically active substances, and influence physiological health. The gut contains the largest number of microorganisms in the human body, reaching trillions, including beneficial microorganisms such as Bifidobacterium, Bacillus subtilis, lactic acid bacteria, and yeast, as well as harmful microorganisms such as Escherichia coli, Salmonella, and putrefactive bacteria. In short, human microorganisms are an integral part of the human body, performing various physiological functions through a symbiotic relationship, and playing a vital role in human health.

[0003] PathSeq is a bioinformatics analysis tool based on GATK software for detecting microbial infections. It is used to detect microorganisms in high-throughput gene sequencing data samples collected from host organisms (such as humans). PathSeq can effectively identify microorganisms against a background of millions of DNA or RNA sequence data, determine microbial abundance, and discover new microbial sequences, thereby rapidly and accurately identifying microorganisms that may be present in various sample types, such as the gut, blood, and respiratory tract.

[0004] Single-cell transcriptome sequencing is a high-throughput technology that can quantitatively sequence RNA molecules in a single cell, allowing for the study of gene expression within a single cell and intercellular heterogeneity. Current single-cell workflows typically employ methods such as direct alignment, assembly, and assembly diagrams for microbial detection. Compared to traditional sequencing technologies, single-cell microbial detection methods have relatively low accuracy and uncertain results. Pathseq has already been applied to spatial transcriptomics, demonstrating its ability to rapidly and accurately detect microorganisms in spatial transcriptomes. Microbial detection and analysis are also necessary in single-cell sequencing research. Currently, there is no Pathseq-based microbial detection and analysis workflow in this field; therefore, it is necessary to establish a method and system for PathSeq-based microbial identification and abundance detection using single-cell transcriptome sequencing. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a method for microbial identification and abundance detection based on single-cell transcriptome sequencing using PathSeq. This method can rapidly and accurately detect microorganisms and their abundance in different host organisms, overcoming the limitations of microbial detection in single-cell transcriptome analysis. Thus, it enables rapid and efficient detection of microorganisms using single-cell sequencing data.

[0006] This invention is the first to propose the application of Pathseq in single-cell transcriptomics. Due to the differences in cell barcoding and visualization between single-cell and spatial transcriptomics, compared to its application in spatial transcriptomics, this method, for the first time, matches the results obtained from PathSeq software with the cells in the RDS analysis files of single-cell transcriptomics and performs visualization adjustments, displaying the results on UMAP or tSNE plots (dot plots obtained from dimensionality reduction clustering), making the results clearer and more intuitive. Furthermore, this invention generates a one-click Pathseq analysis workflow, facilitating the rapid generation of detection results directly from single-cell files and the generation of result reports with a single click. This not only improves researchers' work efficiency and reduces time spent but also reduces the error rate.

[0007] To achieve the above objectives, this invention proposes a PathSeq microbial identification and abundance detection method based on single-cell transcriptome sequencing, comprising the following steps:

[0008] Step 1: Prepare a BAM file containing sequencing data aligned to a reference genome as the original BAM file:

[0009] The main task of this step is to prepare the raw BAM file containing sequencing data aligned to the reference genome from the host single-cell sequencing results, and to provide necessary reference genome information, k-mer libraries, and microbial databases for subsequent analysis. This step yields a raw BAM file already aligned to the reference genome and provides the necessary reference databases and environmental configuration information for subsequent analysis. This step lays the foundation for subsequent data processing and computation, while also providing crucial input data and information, making subsequent analysis more efficient and accurate.

[0010] Prepare the BAM file (which stores sequencing data, i.e., Reads information, and is the result file aligned to the reference genome) from the host single-cell sequencing results. This BAM file is called the raw BAM file. BAM files can also be generated by providing the raw FASTQ file and the reference genome information of the human or host species (generated using the 10x Genomics official quality control software CellRanger). Other components include a BWA index mirror of the host reference genome (used for rapid and accurate alignment of sequencing data to the reference genome), a k-mer library constructed based on the host reference genome (used for sequence alignment and variant detection), a reference microbial genome sequence file and its index mirror file (used to identify or discover different microorganisms by comparing sequencing data with genome sequences in the reference database), a microbial taxonomy database file (used for microbial alignment and classification), and environmental configuration information, all provided by pre-obtained data and information.

[0011] Step 2: Use the GATK pipeline to perform Pathseq analysis to identify and assess the presence and abundance of microorganisms in the sample:

[0012] This step uses the GATK pipeline to perform quality filtering and microbial taxonomy abundance assessment on sequencing data from non-host sources. Specifically, by comparing the sequencing data in the input BAM file with the host reference sequence, host-source sequencing data is removed, and the remaining non-host sequencing data is compared with existing microbial reference sequences to obtain a table of microbial taxonomy and a BAM file containing high-quality non-host sequencing data. This step can detect and assess the presence and abundance levels of microorganisms in the sample.

[0013] include:

[0014] Step 2.1: Perform quality filtering by comparing the sequencing data in the input BAM file with the host reference sequence, and subtract the sequencing data from the host from the input sequencing data to obtain sequencing data from non-host sources.

[0015] Step 2.2: Align the remaining non-host sequencing data with the microbial reference sequence, record the presence of microbial-related sequences in the non-host sequencing data, generate a table of detected microorganisms, and perform classification abundance scoring. The resulting file consists of a BAM file (pathseq.complete.bam) containing all high-quality non-host sequencing data aligned with the microbial reference sequence and a microbial classification table (pathseq.complete.csv). The microbial classification table in the result file sequentially displays the microbial classification number, taxonomic information (presented in a hierarchical structure), microbial type (kingdom, phylum, class, order, family, genus, species, etc.), microbial name, biological kingdom, score in the sample (indicating the amount of evidence for the presence of this taxonomic unit), score after normalization, number of reads explicitly matching the reference sequence, and length of the discovered microbial reference sequence.

[0016] Step 3: Organize the results files obtained from the Pathseq analysis in Step 2 to obtain information about microbial infection and expression at the single-cell level:

[0017] This step organizes and processes the results of the Pathseq analysis, obtaining files on microbial infection and expression at the single-cell level. This step yields a collection of detailed files on microbial infection and expression at the single-cell level. This information is of significant guiding importance for further analysis targeting microbial infection and expression at the single-cell level.

[0018] In this invention, by organizing the Pathseq results, we can obtain visium.readname (a file containing sequencing data, pathogens, and alignment score information), visium.unmap_cbub.bam (a BAM file containing pathogens), visium.unmap_cbub.fasta (a FATA file containing pathogens), visium.list (the corresponding cell file), visium.raw.readnamepath (cell name, UMI sequence, Pathseq classification, alignment score, and genus list), visium.genus.cell (infected cells and related information), visium.genus.csv (pathogen UMI count file), and visium.validate.csv (barcodes, UMIs, and corresponding pathogen files of infected cells).

[0019] Step three specifically includes the following steps:

[0020] Step 3.1: Prepare the cell tag file: Prepare the cell tag file (barcodes.tsv.gz) generated upstream of Cell Ranger, which is used to extract the corresponding cell names.

[0021] Step 3.2: Create a dictionary of bacterial genus names: Read the information at the "genus" level from the Pathseq classification results and create a dictionary containing the names of all bacterial genera.

[0022] Step 3.3: Process the BAM file generated by Pathseq: Read the pathseq.complete.bam file generated by Pathseq, and extract the sequencing data name, the original sample information (YP tag) of the sequencing data (i.e., the rows containing "YP" selected from the pathseq.complete.bam generated in step two), and the quality score of the sequencing data alignment with the reference sequence (a higher value indicates a more reliable alignment result). Output the total number of alignment results and the number of alignment results containing the "YP" tag. The "YP" tag is a general format used to identify the original sample information of sequencing data, and is usually used as key information for sample management and traceability. The "YP" tag stores the original information of each sample, including sample source, collection time, and sample type. By extracting the "YP" tag, data from different samples can be distinguished, thereby understanding the differences in pathogen composition under different conditions (such as different samples, different collection times, etc.).

[0023] Step 3.4: Generate a file containing sequencing data and alignment scores: Extract and generate a file containing sequencing data information, pathogen information, and alignment score information from the pathseq.complete.bam and pathseq.complete.csv files, namely the visium.readname file.

[0024] Step 3.5: Screening and Counting Unaligned Sequences: From the original BAM file, screen the alignment results of unique molecular tag sequences (UB) and cell barcode information (CB) tags based on sequencing data. If no genome alignment is found, count the number of unaligned sequences and write them into the BAM (unmap_cbub.bam) and FASTA (unmap_cbub.fasta) files for unaligned sequences. If the cell name of the alignment result appears in the PathSeq BAM file, add the cell name to the cell set of infected cells through the cell tag file (barcodes.tsv.gz). Finally, obtain the aligned cell name, UMI sequence, PathSeq classification, and alignment score (view the alignment results of each cell with each microorganism based on the score. The higher the score, the higher the probability that the cell contains that microorganism). This step yields visium.unmap_cbub.bam (the BAM file containing the pathogen), visium.unmap_cbub.fasta (the FASTA file containing the pathogen), and visium.list (the corresponding cell file).

[0025] Step 3.6: Integrate data to generate a new file: Integrate sequencing data, cell name, UMI sequence, Pathseq classification, alignment score, and genus list into one file to obtain the visium.raw.readnamepath file.

[0026] Step 3.7: Extract information on infected cells and pathogens: Extract the genus of the infected cells and pathogens, UMI count, and the corresponding pathogen genus count, and integrate them into a file to obtain the visium.genus.cell file.

[0027] Step 3.8: Organize the UMI count files: Organize the UMI counts of each infected cell in each pathogen (such as Mycoplasma, Hematologic Bacillus, Staphylococcus, etc.) category to obtain visium.genus.csv (pathogen UMI count file) and visium.validate.csv (which includes the barcodes, UMIs and corresponding pathogen files of the infected cells).

[0028] The UMI count refers to the counting of unique molecular identifiers (UMIs), which are sequences of 4-10 random nucleotides. During single-cell sequencing, the same mRNA molecule may yield multiple similar but not identical sequencing data. These sequencing data can be misclassified as different mRNAs during data processing. To prevent misclassification, a UMI is introduced for each mRNA molecule. Therefore, expression level data can be accurately calculated using UMI counting, thereby improving the accuracy and reliability of the data.

[0029] Step 4: Prepare and import the RDS file:

[0030] Prepare and read in the RDS file generated by Seurat, a standard single-cell transcriptome analysis software. The RDS file contains cell names and sample information, and facilitates the storage of UMI and sequencing data for subsequent image creation. This is crucial for understanding cell expression patterns, cell type clustering and classification, and the relationship between microbial infection and expression and the host cell transcriptome at the single-cell level.

[0031] Step 5: Import and summarize the UMI and sequencing data of all samples:

[0032] This step summarizes the UMI and sequencing data of all samples and adds them to the meta.data file of the RDS. This step produces a summary file containing the UMI and sequencing data of all samples, which helps to more comprehensively assess the expression patterns and abundance levels of microorganisms in different cell types.

[0033] Read the corresponding sample name, read the visium.validate.csv file (the first column is broken down into two columns, barcode and FASTA) and the visium.raw.readnamepath file, filter and integrate the information in the barcode and FASTA columns of the two files, keep only the rows that are common to both files, and generate the visium.filtered.readnamepath file.

[0034] Read the microbial UMI count file visium.genus.csv for the corresponding cell, format the cell names to match the cell names in the RDS, remove special characters from pathogen names, and generate a matrix of cell × microbial UMI counts. If the matrix contains NA (i.e., no match for the microorganism, a null value), convert it to 0 for easier summation. Sum the microbial UMI counts for each cell, generating a new column Bacteria_UMIs, and record all results with a count of 0 as NA. Add the cell × microbial UMI count matrix and the Bacteria_UMIs information to the meta.data file of the RDS.

[0035] For the previously generated visium.filtered.readnamepath file, calculate the number of sequencing data for all microorganisms in each cell, generate a new column Bacteria_reads, organize the cell names into the cell name format in RDS, and add the sequencing data information to the meta.data of RDS.

[0036] Step Six: Compile statistical data on microbial content across the entire sample level, calculating the number of cells detected in each sample, the number of cells containing at least one microbial genus, the sum of sequencing data for all microorganisms, and the sum of UMI counts for all microorganisms. This step provides statistical data on microbial content at the sample level.

[0037] And / or,

[0038] By statistically analyzing the abundance data of different microbial genera at the sample level, and integrating the UMI counts and sequencing data of different microbial genera in the sample, the relative abundance information of each microbial genera was obtained:

[0039] The relative abundance information of different microbial genera obtained in this step can comprehensively assess the expression patterns and abundance levels of microorganisms in the sample, as well as determine the proportion of microorganisms in the sample.

[0040] For the cell × microbial UMI count matrix, cells consistent with those in the RDS are selected. Additionally, a combination of the current sample and microbial genus is generated, and the number of each microbial genus captured in the current sample and the number of UMIs for that genus are calculated. Only samples with a UMI count greater than 0 are retained.

[0041] The visium.filtered.readnamepath file is read sequentially, the data files of each sample are processed, the sample name corresponding to each cell is added, and the total number of sequencing data of that microbial genus in the current sample is extracted.

[0042] The number of each microbial genus captured in the current sample, the number of UMIs, and the total number of sequencing data are integrated into a table. The number of UMIs is divided by the sum of the UMI counts of all microorganisms, and the result is multiplied by 100 to convert it into a percentage. The relative abundance of each microbial genus in the current sample is then calculated.

[0043] And / or,

[0044] Statistics on the top 10 microorganisms and visualization chart:

[0045] This step analyzed the top 10 microorganisms and created visualizations, including a pie chart showing the relative abundance and distribution of the top 10 genera of microorganisms by UMI number in the current sample, and a bar chart showing the number of UMIs and the total number of sequencing data. This step used visual graphics to present the results, making the conclusions clearer.

[0046] The top 10 genera of microorganisms with the highest number of UMIs in the current sample are selected, and duplicate elements are removed. The number and relative abundance of UMIs captured under each of these genera are generated, and the total number of sequencing data for each genera in the current sample is calculated. The UMI numbers of all other genera are summed and recorded as "others," and their relative abundance is calculated accordingly. If the relative abundance is less than 1%, its labels are counted as NA.

[0047] Plot the relative abundance pie chart of the top 10 microbial genera in terms of UMI quantity in the current sample and save it as a PNG file; plot the distribution of the top 10 microbial UMIs in each cell of the current sample's RDS, and display it as a UMAP plot or tSNE plot; plot bar charts of the number of UMIs captured and the total number of sequencing data under different microbial genera in the current sample.

[0048] And / or, plot the distribution of Bacteria_reads and Bacteria_UMIs in the sample and create a violin plot:

[0049] This step generates a UMAP or tSNE distribution map and a violin plot for the sum of microbial UMI counts (Bacteria_UMIs) for each cell in the current sample (which can be displayed separately by group, sample, clusters, cell type, etc.); and a UMAP or tSNE distribution map and a violin plot for the number of sequencing data (Bacteria_reads) for all microorganisms in each cell of the current sample. This step provides a more intuitive and clear view of the distribution at the sample level.

[0050] Step six, including the statistical analysis of microbial content data at the entire sample level, the statistical analysis of microbial genera at different sample levels and the statistical analysis of the top ten microorganisms and plotting, and the plotting of different distribution maps in the sample, can be performed in parallel without any order.

[0051] Step 7: Save the RDS file containing information on different microbial genera:

[0052] The saveRDS function can be used to save RDS files containing information on different microbial genera, which facilitates subsequent plotting, other analyses, and later retrieval.

[0053] Steps four through seven involve organizing the code based on the tabular data generated by the Pathseq software, and visualizing the top 10 data points with statistical significance.

[0054] The present invention also provides a system for realizing the above-mentioned microbial identification and abundance detection, the system comprising:

[0055] The genome alignment module uses Pathseq software to perform quality filtering on the sequencing data, removing sequencing data from the host organism, and aligning the non-host sequencing data with the microbial reference genome.

[0056] The data cleaning and result organization module cleans and organizes the results generated by Pathseq.

[0057] The results visualization module performs visualization processing and statistical analysis.

[0058] The present invention also provides the above-mentioned methods for identifying and detecting the abundance of microorganisms, or the above-mentioned methods for identifying and detecting the abundance of microorganisms, in the identification and detection of human gastrointestinal microorganisms.

[0059] The beneficial effects of this invention include:

[0060] This invention fills the gap in the field of single-cell transcriptome analysis for a streamlined process for microbial identification and abundance detection. Current methods for microbial identification based on single-cell transcriptome sequencing are cumbersome, lack tutorials, and exhibit significant uncertainty in results. Furthermore, they require a certain level of programming skill from the operators. This invention optimizes the microbial identification process, requiring only the provision of an RDS file and a BAM file (the remaining files are provided by the Pathseq software). The process automatically performs subsequent analyses based on the provided files, and the analysis speed is significantly faster: compared to the existing SAHMI algorithm, the speed is improved by more than 50%.

[0061] This invention broadens the application scope of advanced single-cell transcriptomics analysis: Typical advanced single-cell transcriptomics analyses include pseudo-time series analysis, GSVA analysis, SCENIC transcription factor regulation analysis, CellChat cell communication analysis, and Scran cell cycle analysis. Because single-cell transcriptomics is an emerging technology, it has a shorter history compared to other transcriptomics methods. Furthermore, microbial detection at the single-cell level is a combined analysis requiring simultaneous attention to both single-cell and microbial research; therefore, microbial detection analysis is currently relatively rare. This invention effectively expands the application scope of advanced single-cell transcriptomics analysis. Attached Figure Description

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0063] Figure 1 This is a diagram showing the distribution of the top ten microorganisms in an embodiment of the present invention.

[0064] Figure 2 This is a graph showing the total number of sequencing data for the top ten microorganisms in this embodiment of the invention.

[0065] Figure 3 This is a UMI count chart of the top ten microorganisms in an embodiment of the present invention.

[0066] Figure 4 This is a UMAP distribution diagram of the number of Fusobacterium UMIs at the cellular level in an embodiment of the present invention.

[0067] Figure 5 This is a schematic diagram of the process of the present invention. Detailed Implementation

[0068] The present invention will be further described in detail below with reference to the specific embodiments and accompanying drawings. Except for the contents specifically mentioned below, the processes, conditions, and experimental methods for implementing the present invention are all common knowledge and general knowledge in the art, and the present invention does not have any particular limitations.

[0069] Example

[0070] In this embodiment, microorganisms such as Fusobacterium, Bacteroides, and Prevotella were identified in human colon tissue.

[0071] 1) Read in the human BAM file generated by single-cell transcriptome analysis using CellRanger. The remaining files are provided in this invention.

[0072] 2) Pathseq analysis was performed using GATK, including BAM files of all high-quality non-host reads aligned with the microbial reference and tables of microbial classification. Due to the large amount of content in the microbial tables, only three rows of results are shown, as shown in Table 1 below. The table sequentially displays the microbial classification number (75, 147538, and 157, respectively), taxonomic information (presented in a hierarchical structure), microbial type (genus and subphylum, respectively), microbial name (*Stemobacter*, *Dermocytotrichum*, and *Treponema*), biological kingdom (bacteria and fungi, respectively), score in the sample (reflecting the relative abundance of the microorganism, 0.462, 26.041, and 5722.407, respectively), standardized score (0.002, 46.565, and 27.207, respectively), number of reads in the sample (1, 30, and 5724, respectively), number of explicit reads in the sample (0, 20, and 5718, respectively), and reference sequence length (all 0).

[0073] Table 1

[0074]

[0075]

[0076] 3) Prepare cell tag files and organize the Pathseq results. Extract the information at the "genus" level from the Pathseq classification results and generate the following files: visium.readname (including sequencing data, pathogen, and alignment score information), visium.unmap_cbub.bam (containing the pathogen's BAM file), visium.unmap_cbub.fasta (containing the pathogen's FATA file), visium.list (corresponding cell file), visium.raw.readnamepath (cell name, UMI sequence, Pathseq classification, alignment score, and genus list), visium.genus.cell (infected cells and related information), visium.genus.csv (pathogen UMI count file), and visium.validate.csv (barcodes, UMI, and corresponding pathogen file for infected cells). Due to the large amount of data in the table, only three rows of results are shown. They are as follows:

[0077] The visium.readname file contains sequencing data information, pathogen species IDs, and alignment scores. In the sequencing data information, A00988 is the instrument ID, 49 is the process unit ID, HH3LNDSX2 is the chip serial number, 3 is the experiment ID, 1339 is the cluster ID, and 21468 / 6762 are the X / Y coordinates, as shown in Table 2.

[0078] Table 2

[0079] Sequencing data information Pathogen species numbering Compare scoring information A00988:49:HH3LNDSX2:3:1339:21468:6762 412133 0 A00988:49:HH3LNDSX2:3:2418:7166:6637 412133 0 A00988:49:HH3LNDSX2:3:2170:26775:26443 1244083 29

[0080] visium.list is the file for the corresponding cells, representing the cells where microorganisms were detected, as shown in Table 3.

[0081] Table 3

[0082] Cell sequence identification TTACCATTCGATCCAA-1 GACCGTGGTGCCTAAT-1 CGGAGAACAATTGTGC-1

[0083] SEQ ID NO.1:TTACCATTCGATCCAA

[0084] SEQ ID NO.2: GACCGTGGTGCCTAAT

[0085] SEQ ID NO.3: CGGAGAACAATTGTGC

[0086] The visium.raw.readnamepath file includes sequencing data information, cell sequence identifiers, UMI sequences, Pathseq classification, alignment scores, and a genus list, as shown in Table 4. In Table 4, column 2 is the cell sequence identifier, and column 3 is the UMI sequence of the transcript molecule, which can be used to determine whether the detected microorganism originated from that cell.

[0087] Table 4

[0088]

[0089]

[0090] SEQ ID NO.4: CACCCCTCCAGT

[0091] SEQ ID NO.5: TTAACCGTTCAA

[0092] SEQ ID NO.6: AAAACTTGTGGT

[0093] The visium.genus.cell file contains information about infected cells, including cell sequence identifier, pathogen genus, UMI count, and the corresponding pathogen genus count, as shown in Table 5.

[0094] Table 5

[0095]

[0096] SEQ ID NO.7:ACGTTCCAGAGTGTTA

[0097] SEQ ID NO.8: CGGAACCAGAATCGAT

[0098] SEQ ID NO.9: GAGAAATCATTCATCT

[0099] The visium.genus.csv file is a count matrix of infected cells and their corresponding pathogens. A value of 0 indicates that the corresponding cell has not been infected by that pathogen, while a value other than 0 indicates the UMI count of the corresponding pathogen for that cell, as shown in Table 6.

[0100] Table 6

[0101] Cell sequence identifier (barcode) Aeromicrobium Akkermansia Alistipes Alloprevotella Atopobium GAGGCAAGTCGCAACC-1 1 0 0 0 0 TGGGAGAGTCGCATTA-1 0 0 0 1 0 ATTCATCAGGGACTGT-1 0 0 0 0 3

[0102] SEQ ID NO.10: GAGGCAAGTCGCAACC

[0103] SEQ ID NO.11: TGGGAGAGTCGCATTA

[0104] SEQ ID NO.12:ATTCATCAGGGACTGT

[0105] The visium.validate.csv file contains the barcodes of infected cells, the UMI of the pathogen, and the genus information to which the pathogen belongs, as shown in Table 7. In Table 7, the first column indicates that the UMI sequence after the plus sign comes from the cell represented by the preceding cell sequence identifier.

[0106] Table 7

[0107] Barcodes+UMI genus ATAGACCGTGAATGTA-1+CCAGAGGTACAT Treponema ATAGACCGTGAATGTA-1+GCAGGCCTGCTG Prevotella CGGAGAACAATTGTGC-1+AGTAGTATTGCT Nocardioides

[0108] SEQ ID NO.13:ATAGACCGTGAATGTA

[0109] SEQ ID NO.14: CCAGAGGTACAT

[0110] SEQ ID NO.15: GCAGGCCTGCTG

[0111] SEQ ID NO.16: AGTAGTATTGCT

[0112] 4) Prepare the RDS file generated by the standard single-cell transcriptome analysis software Seurat, and read it into the RDS file. Integrate the information in the barcode and FASTA columns of the visium.validate.csv and visium.raw.readnamepath files, keeping only the rows common to both files, to generate the visium.filtered.readnamepath file. This file is similar to the visium.raw.readnamepath file, except that it retains the common information. Save the corresponding Bacteria_UMIs and Bacteria_reads information to the meta.data file of the RDS.

[0113] 5) Calculate the number of cells detected in each sample, the number of cells with at least one microbial genus captured, the sum of sequencing data for all microorganisms, and the sum of UMI counts for all microorganisms. The generated table is shown below, where the sample name is CRC, the number of cells with at least one microbial genus captured is 2309, the sum of sequencing data for all microorganisms is 610, and the sum of UMI counts for all microorganisms is 284, as shown in Table 8.

[0114] Table 8

[0115]

[0116] 6) Statistically analyze the microbial genera detected in the current sample, the number of cells captured for each microorganism, the number of UMIs, the total number of sequencing data, and the relative abundance of the microbial genera. For example, in the Fusobacterium genus, 12 cell numbers were captured, the number of UMIs detected was 161, the total number of sequencing data was 1326, and the relative abundance was 56.69%, as shown in Table 9.

[0117] Table 9

[0118]

[0119]

[0120] 7) Analyze the top ten microorganisms. Select the top 10 genera in terms of UMI count in the current sample, remove duplicates, and generate the number and relative abundance of UMIs captured under each genera. Calculate the total number of sequencing data for each genera in the current sample. (Since 9 genera have a UMI count of 1 in this example, it actually represents the top 14 microorganisms.) The table below shows the UMI count and relative abundance of the genera.

[0121] Table 10

[0122]

[0123]

[0124] The table below shows the number of sequencing data for each microbial genera:

[0125] Table 11

[0126]

[0127] 8) Draw a graph showing the top ten microorganisms and create a corresponding pie chart. Figure 1 The pie chart shows the relative abundance of the corresponding microbial genera. The chart indicates that the genus *Fusobacterium* had the highest relative abundance in this sample, accounting for 56.69%, followed by the genus *Treponema*, accounting for 25.35%.

[0128] Plot the corresponding bar chart, with the vertical axis representing the corresponding microbial genus and the horizontal axis representing the total number of sequencing data. Figure 2 ) and UMI count ( Figure 3 The figure shows that the total number of sequencing data of Fusobacterium microorganisms in this sample is approximately between 1300 and 1500 (1326 as shown in Table 11), and the UMI count is above 150 (161 as shown in Table 10).

[0129] The corresponding distribution map is drawn, showing the distribution of each microbial genus at the cellular level (UMAP map). Figure 4 The figure shows the UMI count of Fusobacterium microorganisms in each cell of the sample. The bluer the color, the higher the UMI count of Fusobacterium microorganisms in that cell, and the more likely the cell is to be infected with this microorganism.

[0130] 9) Save RDS files containing information on different microbial genera.

[0131] Comparative Example

[0132] SAHMI is a single-cell analysis algorithm for host-microbiota interactions. The following table shows the simultaneous Pathseq and SAHMI analyses performed on six samples, calculating the cell count, runtime, and CPU time for each sample (Table 12). Runtime is calculated as the program's end time minus its start time, and CPU time is calculated as the number of CPU cores multiplied by the computation time. Although the degree of microbial infection varies among the samples, affecting runtime and CPU time, it is evident that Pathseq analysis significantly outperforms SAHMI analysis in terms of both runtime and CPU time.

[0133] Table 12

[0134]

[0135] The scope of protection of this invention is not limited to the above embodiments. Any variations and advantages that can be conceived by those skilled in the art without departing from the spirit and scope of this invention are included in this invention and are protected by the appended claims.

Claims

1. A PathSeq method for microbial identification and abundance detection based on single-cell transcriptome sequencing, characterized in that, The method includes the following steps: Step 1: Prepare a BAM file containing sequencing data aligned to the reference genome as the original BAM file; Step 2: Use the GATK pipeline to perform PathSeq analysis to identify and assess the presence and abundance of microorganisms in the sample; Step 3: Organize the results files obtained from the Pathseq analysis in Step 2 to obtain information about microbial infection and expression at the single-cell level; Step 4: Prepare and read in the RDS file, which contains cell names and sample information; Step 5: Import and summarize the UMI and sequencing data of all samples; Step Six: Compile data on the microbial content of the entire sample level; and / or, Statistical analysis of microbial genera content data at different sample levels, statistical analysis of the top ten microorganisms and visualization plot; and / or, Plot the distribution and violin plot of Bacteria_reads and Bacteria_UMIs in the sample; Step 7: Save the RDS file containing information on different microbial genera.

2. The method as described in claim 1, characterized in that, In step one, the BAM file contains sequencing data aligned to a reference genome; the BAM file is obtained through host single-cell sequencing and / or generated by providing the original FASTQ file and reference genome information of a human or host species. The files prepared in step one also include the BWA index image of the host reference genome, the k-mer library constructed based on the host reference genome, the reference microbial genome sequence file and index image file, the microbial classification database file, and environmental configuration information.

3. The method as described in claim 1, characterized in that, Step two involves comparing the sequencing data in the input BAM file with the host reference sequence, removing the host-derived sequencing data, and comparing the remaining non-host sequencing data with the microbial reference sequence to obtain a table of microbial classification and a BAM file containing high-quality non-host sequencing data; this includes the following steps: Step 2.1: Perform quality filtering by comparing the sequencing data in the input bam file with the host reference sequence, and subtract the sequencing data from the host from the input sequencing data; Step 2.2: Compare the remaining non-host sequencing data with the microbial reference sequence, record the presence of microbial-related sequences in the non-host sequencing data, generate a table of detected microorganisms, and perform classification and abundance scoring.

4. The method as described in claim 1, characterized in that, Step 3 includes preparing cell tag files, creating a bacterial genus name dictionary, processing the Pathseq-generated BAM file, generating a file containing sequencing data and alignment scores, filtering and counting unaligned sequences, integrating data to generate a new file, extracting infected cell and pathogen information, and organizing UMI count files; obtaining the following steps: a visium.readname file containing sequencing data, pathogen, and alignment score information; a visium.unmap_cbub.bam file containing pathogens; a visium.unmap_cbub.fasta file containing pathogens; a visium.list cell file; a visium.raw.readnamepath file including cell name, UMI sequence, Pathseq classification, alignment score, and genus list; a visium.genus.cell file including infected cells and information; a visium.genus.csv file including pathogen UMI counts; and a visium.validate.csv file including infected cell barcodes, UMIs, and corresponding pathogens; and / or, In step four, the RDS file generated by the standard single-cell transcriptome analysis software Seurat is prepared and read in. The RDS file contains cell names and sample information; and / or, In step five, the visium.validate.csv and visium.raw.readnamepath files are read and processed. After filtering and sorting, the common rows of the two files are retained to generate the visium.filtered.readnamepath file. Based on the microbial UMI count file visium.genus.csv, a matrix of cell × microbial UMI counts is generated. The microbial UMI count results for each cell are summed to generate a new column Bacteria_UMIs. All results with a count of 0 are recorded as NA. The cell × microbial UMI count matrix and the Bacteria_UMIs results are added to the meta.data of the RDS. Calculate the number of sequencing data for all microorganisms in each cell, generate a new column Bacteria_reads, organize the cell names into the cell name format in RDS, and add the sequencing data information to the meta.data of RDS.

5. The method as described in claim 1, characterized in that, In step six, the number of cells detected in each sample, the number of cells in which at least one microbial genus was captured, the sum of the sequencing data of all microorganisms, and the sum of the UMI counts of all microorganisms are calculated.

6. The method as described in claim 1, characterized in that, In step six, UMI counts and sequencing data of different microbial genera in the sample were calculated and integrated to obtain the relative abundance information of each microbial genera; the number of each microbial genera captured in the current sample, the number of UMIs, and the total number of sequencing data were integrated into a table, and the number of UMIs was divided by the sum of the UMI counts of all microorganisms, the result was multiplied by 100, and converted into a percentage to calculate the relative abundance of each microbial genera in the current sample; and / or, Take the top 10 genera of microorganisms with the most UMIs in the current sample, remove duplicate elements, generate the number and relative abundance of UMIs captured under the current microbial genera, and calculate the total number of sequencing data of the top 10 genera of microorganisms in the current sample; draw a pie chart of the relative abundance of the top 10 genera of microorganisms in the current sample and save it as a PNG file; draw the distribution of the top 10 microbial UMIs in each cell of the RDS of the current sample and display it as a UMAP plot or tSNE plot; draw bar charts of the number of UMIs captured and the total number of sequencing data under different microbial genera in the current sample respectively.

7. The method as described in claim 1, characterized in that, In step six, plot the UMAP or tSNE distribution map and violin plot of the sum of microbial UMI counts for each cell in the current sample (Bacteria_UMIs); plot the UMAP or tSNE distribution map and violin plot of the sequencing data of all microorganisms for each cell in the current sample (Bacteria_reads).

8. The method as described in claim 1, characterized in that, In step seven, the saveRDS function is used to save the RDS file containing information about different microbial genera.

9. A system for implementing the method according to any one of claims 1-8, the system comprising a genome alignment module, a data cleaning and result processing module, and a result visualization module; The genome alignment module uses Pathseq software to perform quality filtering on sequencing data, subtracting sequencing data from the host organism, and aligning the non-host sequencing data with the microbial reference genome. The data cleaning and result processing module is used to clean and process the results generated by Pathseq. The results visualization module is used to perform data visualization processing and statistical analysis.

10. The method according to any one of claims 1-8, or the system according to claim 9, in the identification and detection of human gastrointestinal microorganisms and upper respiratory tract microorganisms.