Data processing method and device for quantifying macro-genomic species components and abundance, and storage medium
By performing sequence dimensionality reduction processing on metagenomic sample sequencing data and comparing it with species-specific molecular tag databases, the time-consuming problems of database installation and sequence alignment in existing technologies have been solved, enabling rapid and accurate quantification of species composition and abundance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AGRICULTURAL GENOMICS INSTITUTE AT SHENZHEN CHINESE ACADEMY OF AGRICULTURAL SCIENCES (SHENZHEN BRANCH GUANGDONG LABORATORY FOR LINGNAN MODERN AGRICULTURE)
- Filing Date
- 2022-12-30
- Publication Date
- 2026-05-29
AI Technical Summary
Existing methods for quantifying species composition and abundance in metagenomics require downloading large databases and time-consuming sequence alignment, resulting in difficulties in use, slow operation speed, and inaccuracy.
Sample sketches are obtained by sequence dimensionality reduction and compared with species-specific molecular tag databases, which reduces computational resource consumption and improves data processing efficiency.
It greatly accelerated the computing speed, reduced the consumption of computing resources, improved the data processing efficiency, and enabled rapid and accurate analysis of species composition and abundance.
Smart Images

Figure CN115966255B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, and in particular relates to a data processing method, apparatus and storage medium for quantifying species composition and abundance in metagenomics. Background Technology
[0002] Metagenomic sample sequencing and metagenomic species composition and abundance quantification have wide applications in biomedical research. There are many methods for quantifying the microbial species composition and abundance in samples using metagenomic sequencing data, such as MetaPhlAn (metagenomic phylogenetic analysis) and methods for analyzing microbial abundance, activity, and community genomes based on mOTUs (operational taxonomic units).
[0003] However, on the one hand, the above methods all require the prior download and installation of a very large microbial reference genome or taxonomic molecular marker database (approximately several gigabytes in size), which is difficult for users with poor internet connections to complete. On the other hand, these methods also require the use of additional sequence alignment tools (such as Bowtie or BWA, which are usually computationally intensive and very time-consuming) to compare the sample metagenomic data with the taxonomic molecular marker database before calculating the species composition and abundance of each part of the sample. Therefore, current methods generally suffer from being difficult to use, slow in operation, and relatively inaccurate. Summary of the Invention
[0004] This application aims to provide a data processing method, apparatus, and storage medium for quantifying the species composition and abundance of metagenomics, which can greatly accelerate the computing speed, reduce the consumption of computing resources, and improve data processing efficiency.
[0005] This application provides a data processing method for quantifying species composition and abundance in metagenomics, including:
[0006] Obtain metagenomic sample sequencing data and perform sequence dimensionality reduction processing on the metagenomic sample sequencing data to obtain a sample sketch of the metagenomic sample sequencing data;
[0007] By using the sample sketch to query the species-specific molecular tag database, the target species contained in the metagenomic sample sequencing data and the abundance of each target species are determined and output, wherein the species-specific molecular tag database is constructed based on reference sketches of reference genome data of the same and / or different species.
[0008] This application also provides an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program includes:
[0009] The sample analysis module is used to acquire metagenomic sample sequencing data, perform sequence dimensionality reduction processing on the metagenomic sample sequencing data to obtain a sample sketch of the metagenomic sample sequencing data, and use the sample sketch to query a species-specific molecular tag database to determine the target species contained in the metagenomic sample sequencing data and the abundance of each target species, wherein the species-specific molecular tag database is constructed based on reference sketches of reference genome data of the same and / or different species.
[0010] The output module is used to output information about each target species and the abundance of each target species.
[0011] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the data processing method for quantifying metagenomic species composition and abundance as shown in the above embodiments.
[0012] In the embodiments described above, sequence dimensionality reduction processing is performed on the obtained metagenomic sample sequencing data to obtain a sample sketch of the metagenomic sample sequencing data. Then, the target species and abundance of each target species in the metagenomic sample sequencing data are determined and output by querying a species-specific molecular tag database using the sample sketch. Since the comparison is performed using a sketch with a relatively small amount of data obtained through sequence dimensionality reduction processing, the calculation speed can be greatly accelerated, the consumption of computing resources can be reduced, and the data processing efficiency can be improved. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention.
[0014] Figure 1 This is a schematic flowchart of a data processing method for quantifying species composition and abundance in metagenomics according to an embodiment of this application;
[0015] Figure 2 This is a schematic flowchart of a data processing method for quantifying species composition and abundance in metagenomics, provided in another embodiment of this application.
[0016] Figure 3 yes Figure 2 A schematic diagram of the implementation process of step S202 in the method shown;
[0017] Figure 4 yes Figure 2 A schematic diagram of database construction in the method shown;
[0018] Figure 5 yes Figure 2 A schematic diagram of sample analysis in the method shown;
[0019] Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0020] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0021] See Figure 1 , Figure 1 This is a flowchart illustrating a data processing method for quantifying metagenomic species composition and abundance according to an embodiment of this application. This method can be applied to devices such as desktop computers, laptops, tablets, personal computers, servers, and other computer equipment capable of data processing in mobile or non-mobile environments. Figure 1 As shown, the method mainly includes the following steps:
[0022] Step S101: Obtain metagenomic sample sequencing data and perform sequence dimensionality reduction processing on the metagenomic sample sequencing data to obtain a sample sketch of the metagenomic sample sequencing data;
[0023] Step S102: By using the sample sketch to query the species-specific molecular tag database, the target species included in the metagenomic sample sequencing data and the abundance of each target species are determined and output.
[0024] Specifically, a preset sequence dimensionality reduction algorithm, such as KSSD (K-mer Substring Space Decomposition), can be used to perform sequence dimensionality reduction on the acquired metagenomic sample sequencing data to obtain a sample sketch of the metagenomic sample sequencing data.
[0025] This species-specific molecular tag database is constructed based on reference sketches of reference genome data from the same and / or different species. These reference sketches are similar to sample sketches, except that they are obtained by performing sequence dimensionality reduction on the reference genome data.
[0026] Understandably, the species-specific molecular tag database also stores the correspondence between each sketch and its corresponding species. By comparing the sample sketch with the species-specific molecular tag database, one or more target species contained in the sample sketch can be identified, and the abundance of each target species can be obtained. Then, the relevant information of each identified target species and its abundance can be output according to a preset output method, such as generating, but not limited to, a report file containing descriptive information of each identified target species and its quantitative abundance, and outputting it to a local or cloud server for storage, etc.
[0027] In this embodiment, the sequence dimensionality reduction of the acquired metagenomic sample sequencing data is performed to obtain a sample sketch of the metagenomic sample sequencing data. Then, the target species and abundance of each target species in the metagenomic sample sequencing data are determined and output by querying the species-specific molecular tag database using the sample sketch. Since the comparison is performed using a sketch with a small amount of data obtained through sequence dimensionality reduction, the calculation speed can be greatly accelerated, the consumption of computing resources can be reduced, and the data processing efficiency can be improved.
[0028] See Figure 2 , Figure 2 This is a flowchart illustrating a data processing method for quantifying metagenomic species composition and abundance, provided in another embodiment of this application. This method can be applied to devices such as desktop computers, laptops, tablets, personal computers, servers, and other computer equipment capable of data processing in mobile or non-mobile environments. Figure 2 As shown, the method mainly includes the following steps:
[0029] Step S201: Obtain reference genome data and perform sequence dimensionality reduction processing on each reference genome data to obtain a reference sketch of each reference genome data.
[0030] Specifically, multiple reference genome data can be obtained from local or cloud databases, and the sequence dimensionality reduction of each reference genome data can be performed using a preset sequence sketching algorithm to obtain the reference sketch corresponding to each reference genome data.
[0031] Optionally, reference genome data can be obtained from a public database, which can be configured on a cloud server. For example, multiple reference genome datasets can be downloaded from a public database provided by NCBI (National Center for Biotechnology Information). In practical applications, reference genome data can come from the genomes of the same and / or different species; this application does not impose specific limitations.
[0032] The sequence sketching algorithm described above can be, but is not limited to, other sequence sketching algorithms such as KSSD or Minhash. In this embodiment, the reference sketch can also be called a reference genome sketch, which is a subset of k-mers sampled from all k-mers of the reference genome using a sequence sketching algorithm. This subset can be used to replace the entire genome sequence in subsequent calculations. Specifically, algorithms such as KSSD or Minhash can be used. Because this k-mer subset can be thousands of times smaller than the original genome sequence, it can greatly speed up the calculation and reduce the consumption of computing resources.
[0033] Step S202: Construct a species-specific molecular tag database based on the obtained reference sketches.
[0034] Understandably, a species-specific molecular tag database only needs to be built once, and subsequent analyses of multiple samples can be performed using this species-specific molecular tag database.
[0035] Specifically, such as Figure 3 As shown, step S202 may include the following steps:
[0036] Step S2021: Take the union of multiple reference sketches from the same species to obtain the pangenome sketch of the species;
[0037] Step S2022: Subtract the pangenome sketches of other species from the pangenome sketches of each species to obtain the species-specific pangenome sketches.
[0038] Step S2023: Construct a species-specific molecular tag database by indexing all specific pangenome sketches obtained in step S2022.
[0039] First, the pangenome sketch of a species is obtained by taking the union of multiple genome sketches (i.e., pangenome sketches) from the same species. This operation is applied to all species to obtain pangenome sketches for all species. Then, the pangenome sketches of other species are subtracted from the pangenome sketch of each species to obtain the species-specific pangenome sketch (i.e., species-specific molecular tag). Finally, a database, namely the species-specific molecular tag database, is constructed using the indexed species-specific pangenome sketches of all species.
[0040] Step S203: Obtain metagenomic sample sequencing data and perform sequence dimensionality reduction processing on the metagenomic sample sequencing data to obtain the sample sketch corresponding to the metagenomic sample sequencing data.
[0041] Specifically, metagenomic sample sequencing data can be generated through in-house sequencing or downloaded from publicly available data on cloud servers. In-house sequencing involves using sequencing tools to sequence the metagenomic sample to obtain sequencing data. These sequencing tools can include, but are not limited to, second-generation (e.g., Illumina) or third-generation (e.g., Nanopore) sequencers, as well as single-cell sequencing platforms (e.g., 10X Genomics). The sequencing data obtained through a sequencer is generally a collection of reads, in formats such as FASTQ.
[0042] In this embodiment, a sequence sketching algorithm is used to perform sequence dimensionality reduction on the acquired metagenomic sample sequencing data to obtain a sketch of the metagenomic sample sequencing data. Specifically, while obtaining the sample sketch (i.e., k-mers sampled according to a preset method) of each metagenomic sample sequencing data using the sequence sketching algorithm, it is also necessary to record the frequency of each k-mer in the sample sketch within the metagenomic sample sequencing data for subsequent calculation of the target species abundance.
[0043] Specifically, when using KSSD to obtain sample sketches (i.e., sampled k-mers) of metagenomic sample sequencing data, the number of times each sampled k-mer appears in the metagenomic sample sequencing data can be recorded.
[0044] Step S204: Use the sample sketch to query the species-specific molecular tag database, determine the target species included in the metagenomic sample sequencing data and the abundance of each target species, and output the results.
[0045] Specifically, for each sample sketch being analyzed, the intersection of the sample sketch with the species-specific pangenome sketches in the species-specific molecular tag database is used to determine which species in the database share k-mers with the sample sketch, thereby determining which target species are included in the sample sketch or the metagenomic sample sequencing data.
[0046] Furthermore, the occurrence frequency of common k-mers of each target species in the metagenomic sample sequencing data is counted and a summary statistic is calculated. The summary statistic can be, for example, but not limited to, the mean, median, etc., as the abundance of each target species.
[0047] Specifically, the abundance of a target species can be determined by counting the frequency of its median or other percentile k-mers in the metagenomic sequencing data. The median is the 50th percentile, and similarly, the Nth percentile can be used. Alternatively, the average of the Nth to Mth percentiles can be taken, where N and M are numbers in the range of 0-100.
[0048] Furthermore, in other embodiments, after obtaining the abundance of each target species, normalization or normalization can be performed to obtain the total abundance of each target species.
[0049] For example, assuming the abundances of three target species in the metagenomic sample sequencing data are calculated to be a, b, and c, the total abundances a', b', and c' of each of these three target species can be obtained by normalization. The sum of the three total abundances is 1. Specifically, a is normalized to a' = a / (a+b+c), b is normalized to b' = b / (a+b+c), and c is normalized to c' = c / (a+b+c).
[0050] Normalization yields a total abundance of 100. Specifically, normalize a to a' = a * 100 / (a + b + c), and similarly, b' and c' can be obtained.
[0051] Because it uses sketches for comparison, it avoids the need for computationally intensive sequence alignment tools, thus enabling rapid and accurate analysis of sample species composition and abundance.
[0052] Furthermore, in other embodiments of this application, reference genome data updated in real time in a public database can be obtained, and sequence dimensionality reduction processing can be performed on the updated reference genome data to obtain a sketch of the updated reference genome data; using the obtained sketch, the species-specific molecular tag database can be updated in real time. In this way, by utilizing the characteristic that reference genomes in public databases are constantly being updated, the species-specific molecular tag database can be updated in real time, thereby making species abundance quantification more accurate.
[0053] Specifically, updated reference genome data can be downloaded periodically or based on subscription information sent by the public database. A sketch of the updated reference genome data is then obtained using the same method as in step S201 above. Then, the above... Figure 3 The method shown yields a species-specific pangenome sketch of the updated reference genome data and indexes it in a species-specific molecular tag database.
[0054] It should be noted that steps S201, S202 and S203, S204 are independent of each other and can be completed by two independent modules. To aid understanding, the following will combine... Figure 4 and Figure 5 The data processing method provided in this embodiment will be further explained.
[0055] Regarding the database construction section, as shown in Figure 4, firstly, the genomic data used as a reference is downloaded from a public database through the species-specific molecular tag database construction module (i.e., Figure 4(Known microbial genomes in the genome).
[0056] Then, the KSSD algorithm was used to perform sequence dimensionality reduction on the acquired genome data. After KSSD dimensionality reduction, each genome was transformed into a much smaller genome sketch, which is a small sample of the k-mer set of that genome, such as... Figure 4 g in e.coli (GAT...), g k.pneumoniea (ATG...), g s.tphimurium (GAT ATG...). Figure 4 The example used is k=3. In practical applications, k can take other integer values, such as k=20. It is important to note that once the value of k is determined, it will not change throughout the entire process.
[0057] Then, the union of multiple genome sketches from the same species is taken to obtain a pan-genome sketch of the species, such as... Figure 4 As shown, by taking the union of all genome sketches from species 1, the pan-genome sketch P1 = g of species 1 is obtained. 11 ∪g 12 ∪……∪g 1x Taking the union of all genome sketches from species 2, we obtain the pan-genome sketch P2 = g of species 2. 21 ∪g 22 ∪……∪g 2y This process continues until the union of all genome sketches from species m is taken, resulting in the pan-genome sketch P of species m. m =g m1 ∪g m2 ∪……∪g mz Then, the pan-genome sketches of each species were subtracted from the pan-genome sketches of other species to obtain the species-specific pan-genome sketches, i.e. Figure 4 The specific molecular tag M in i =P i -∪ j≠i P j (where P) i Represents a pan-genome sketch of species i, ∪ j≠i P j The species-specific molecular tag database mentioned above can be constructed by indexing all species-specific molecular tags (representing the union of pan-genome sketches of other species besides this species).
[0058] Regarding the sample analysis section, such as Figure 5 As shown, the KSSD algorithm or other similar sequence dimensionality reduction algorithms are used to process the obtained sample sequencing data (i.e., Figure 5The metagenomic sample data in the sample is subjected to sequence dimensionality reduction to transform it into a much smaller sample sketch. This sample sketch is a subset of k-mers extracted from all k-mers in the sample sequencing data, and records the occurrence frequency of each k-mer in the sample sequencing data, such as: Figure 5 In the sample sketches, the k-mers TTC appear 2 times; GAT appears 41 times; TCG appears 15 times, and so on. Next, the sample sketches to be analyzed are compared with the species-specific pan-genome sketches of each record in the species-specific molecular tag database, and their intersections are calculated to obtain the common k-mers. If the number of common k-mers with a species-specific pan-genome sketch of a record in the database is greater than or equal to a certain threshold, then the sample sketch can be identified as containing the species to which the species-specific pan-genome sketch of that record belongs (i.e., the target species). The threshold is a fixed positive integer, such as 9. Finally, the frequency of occurrences of the common k-mers of each target species with the sample sketch in the sample sequencing data is summarized, and the resulting summary statistics are used as the abundance of each target species contained in the sample sequencing data.
[0059] It should be noted that each k-mer stored in the species-specific molecular tag database is unique and its species origin is already labeled. Therefore, given any k-mer, searching the database will determine whether the k-mer exists in the database, and if so, which species it belongs to.
[0060] As is understandable, metagenomics refers to a genome composed of multiple species, and a sample sketch typically contains multiple k-mers from different species. The method provided in this application aims to analyze which species are included in a sample sketch.
[0061] Typically, a sample sketch to be analyzed shares common k-mers with multiple species-specific tags in the aforementioned database. For each species, the common k-mers between the species-specific tag (which is also some k-mers) and the sample sketch are found, and the statistical abundance of these common k-mers is calculated. Possible abundance statistics (i.e., abstract statistics) may include, but are not limited to, the mean, median, and the mean of the 98th and 99th percentiles.
[0062] by Figure 5 For example, Figure 5 k-mer matching in this context involves searching the aforementioned database to determine the M values for each species stored in that database. i Species that share the same k-mer as the sample sketch and whose number of such identical k-mers exceeds a preset threshold are considered target species, such as... Figure 5M1, M2 and M3 correspond to species 1, species 2 and species 3.
[0063] Figure 5 The M1, M2, and M3 in the circular dashed box refer to species-specific molecular tags of species 1, species 2, and species 3, respectively, namely, species-specific k-mers, species-specific k-mers, and species-specific k-mers.
[0064] The count on the left side of the circular dashed box represents the number of times the species-specific molecular tags of each target species, which share k-mers with the sample sketch, appear in the metagenomic sample sequencing data (hereinafter referred to as sample data) to which the sample sketch belongs. For example, the k-mers shared with the sample sketch in M1 of target species 1, such as the number of times ATG and CGG appear in the sample data, and the k-mers shared with the sample sketch in M2 of target species 2, such as the number of times TTC and GAT appear in the sample data, etc.
[0065] The count statistics on the right side of the circular dashed box represent the summary statistics (i.e., abundance) of each target species obtained from the counts on the left. Figure 5 The abundance of each target species is the mean of the counts of k-mers at the 98th and 99th percentiles. However, in practical applications, this is not the only factor to consider.
[0066] The data processing method for quantifying metagenomic species composition and abundance provided in this embodiment has the following advantages:
[0067] 1) By using sequence dimensionality reduction to greatly reduce the dimensionality of genome or sequencing data to obtain a sketch, and then using the sketch with much smaller data volume to replace the original genome (or sequencing data) for calculation, the computation time and space overhead can be greatly reduced, thus having the advantages of saving computation time, memory and storage space consumption.
[0068] 2) Since sketches with much smaller data volume are used instead of the original genome (or sequencing data), it is easier to share sketches over the network, thus having the advantage of facilitating data transmission over the network.
[0069] 3) By taking advantage of the fact that reference genomes on public databases are constantly being updated, this application can add new reference genome sketches in real time to obtain a real-time updated species-specific molecular tag database, thereby making species abundance quantification more accurate. Therefore, this application also has the advantages of species-specific molecular tag database being updated in real time and quantification being more accurate.
[0070] See Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device 30 includes: a memory 50, a processor 60, and a computer program 40 stored in the memory and executable on the processor. When the processor executes the computer program, it implements the data processing method for quantifying metagenomic species composition and abundance as described in the foregoing embodiments.
[0071] The computer program 40 includes:
[0072] The sample analysis module 401 is used to acquire metagenomic sample sequencing data, perform sequence dimensionality reduction processing on the metagenomic sample sequencing data to obtain a sample sketch of the metagenomic sample sequencing data, and use the sample sketch to query a species-specific molecular tag database to determine the target species included in the metagenomic sample sequencing data and the abundance of each target species. The species-specific molecular tag database is constructed based on reference sketches of reference genome data of the same and / or different species.
[0073] Output module 402 is used to output information on each target species and the abundance of each target species.
[0074] Furthermore, computer program 40 also includes:
[0075] The database construction module 403 is used to obtain multiple reference genome data from public databases, perform sequence dimensionality reduction processing on each reference genome data to obtain reference sketches of each reference genome data, and construct a species-specific molecular tag database based on the obtained reference sketches.
[0076] The specific processes by which the sample analysis module 401, output module 402, and database construction module 403 implement their respective functions can be found in the above description. Figures 1 to 5 The relevant descriptions in the illustrated embodiments will not be repeated here.
[0077] Furthermore, the electronic device also includes at least one input device and at least one output device.
[0078] The aforementioned memory, processor, input device, and output device are connected via a bus.
[0079] Input devices can specifically include cameras, touch panels, physical buttons, etc. Output devices can specifically include monitors, printers, radio frequency modules, etc., where monitors can include, but are not limited to, touch or non-touch CRT monitors, liquid crystal displays (LCDs), LED displays, etc.
[0080] The memory can be high-speed random access memory (RAM) or non-volatile memory, such as disk storage. Memory is used to store a set of executable program code, and the processor is coupled to the memory.
[0081] In this embodiment, the sequence dimensionality reduction of the acquired metagenomic sample sequencing data is performed to obtain a sample sketch of the metagenomic sample sequencing data. Then, the target species and abundance of each target species in the metagenomic sample sequencing data are determined and output by querying the species-specific molecular tag database using the sample sketch. Since the comparison is performed using a sketch with a small amount of data obtained through sequence dimensionality reduction, the calculation speed can be greatly accelerated, the consumption of computing resources can be reduced, and the data processing efficiency can be improved.
[0082] Furthermore, this application embodiment also provides a computer-readable storage medium, which may be disposed in the aforementioned electronic device, and may be the memory of the aforementioned electronic device. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the data processing method for quantifying metagenomic species composition and abundance described in the foregoing embodiments. Furthermore, the computer-readable storage medium may also be a USB flash drive, a portable hard drive, a read-only memory (ROM), RAM, a magnetic disk, or an optical disk, or any other medium capable of storing program code.
[0083] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0084] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0085] The above is a description of the data processing method, electronic device, and computer-readable storage medium for quantifying metagenomic species composition and abundance provided by the present invention. For those skilled in the art, based on the ideas of the embodiments of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A data processing method for quantifying species composition and abundance in metagenomics, characterized in that, include: Multiple reference genome data are obtained from public databases, and the sequence sketching algorithm is used to perform sequence dimensionality reduction processing on each of the reference genome data to obtain a reference sketch of each of the reference genome data. The union of multiple reference sketches from the same species is taken to obtain a pan-genome sketch of the species; Subtract the pangenome sketches of other species from the pangenome sketches of each species to obtain the species-specific pangenome sketches. Construct a species-specific molecular tag database by indexing all the obtained species-specific pangenome sketches. Obtain metagenomic sample sequencing data, and use a sequence sketching algorithm to perform sequence dimensionality reduction on the metagenomic sample sequencing data to obtain a sample sketch of the metagenomic sample sequencing data. The sample sketch is a k-mer sampled according to a preset method, and the number of times each k-mer appears in the metagenomic sample sequencing data is recorded. By querying the species-specific molecular tag database using the sample sketch, the target species contained in the metagenomic sample sequencing data and the abundance of each target species are determined and output. Specifically, for each sample sketch being analyzed, the intersection of the sample sketch with each species-specific pangenome sketch in the species-specific molecular tag database is used to determine which species in the species-specific molecular tag database share k-mers with the sample sketch, thereby determining the target species contained in the sample sketch or the metagenomic sample sequencing data. The occurrence frequency of common k-mers of each target species in the metagenomic sample sequencing data is counted, and the summary statistics are calculated as the abundance of each target species.
2. The method as described in claim 1, characterized in that, The step of performing sequence dimensionality reduction processing on each of the reference genome data to obtain a reference sketch of each of the reference genome data includes: The reference genome data are subjected to sequence dimensionality reduction processing using a subsequence spatial decomposition algorithm to obtain the reference sketch.
3. The method as described in claim 2, characterized in that, The sequence dimensionality reduction processing of the metagenomic sample sequencing data includes: The subsequence spatial decomposition algorithm is used to perform sequence dimensionality reduction on the metagenomic sample sequencing data.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain the reference genome data updated in real time from the public database, and perform sequence dimensionality reduction processing on the updated reference genome data to obtain a sketch of the updated reference genome data; The species-specific molecular tag database is updated in real time using the sketches.
5. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The abundance of each target species is normalized or normalized to 100% to obtain the total abundance of each target species and output it.
6. An electronic device, the electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the computer program comprises: The sample analysis module is used to acquire multiple reference genome data from public databases, perform sequence dimensionality reduction on each reference genome data using a sequence sketching algorithm to obtain reference sketches for each reference genome data, take the union of multiple reference sketches from the same species to obtain a pan-genome sketch for that species, subtract pan-genome sketches of other species from each species' pan-genome sketch to obtain species-specific pan-genome sketches, construct a species-specific molecular tag database by indexing all the obtained species-specific pan-genome sketches, acquire metagenomic sample sequencing data, perform sequence dimensionality reduction on the metagenomic sample sequencing data using a sequence sketching algorithm to obtain sample sketches of the metagenomic sample sequencing data, wherein the sample sketches are k-mers sampled according to a preset method, and simultaneously record the... The number of times each k-mer in the sample sketch appears in the metagenomic sample sequencing data is determined, and the target species and abundance of each target species are determined by querying a species-specific molecular tag database using the sample sketch. Specifically, for each analyzed sample sketch, the intersection of the sample sketch with each species-specific pangenome sketch in the species-specific molecular tag database is used to determine which species in the species-specific molecular tag database share k-mers, thus identifying the target species contained in the sample sketch or the metagenomic sample sequencing data. The number of times the shared k-mers of each target species in the metagenomic sample sequencing data appear in the metagenomic sample sequencing data is counted, and a summary statistic is calculated as the abundance of each target species. The output module is used to output information about each target species and the abundance of each target species.
7. The apparatus as claimed in claim 6, characterized in that, The computer program also includes: The database construction module is used to obtain multiple reference genome data from public databases, perform sequence dimensionality reduction processing on each reference genome data to obtain reference sketches of each reference genome data, and construct the species-specific molecular tag database based on the obtained reference sketches.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data processing method for quantifying the species composition and abundance of metagenomics as described in any one of claims 1 to 5.