Reference sequence construction method, metagenome data compression method and electronic equipment

CN119998885APending Publication Date: 2025-05-13MGI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280100940.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing technologies are difficult to effectively compress metagenomic data, especially because the species composition in metagenomic data is complex and the compression efficiency cannot be significantly improved through stable reference sequences, resulting in huge challenges in data storage and transmission.

Method used

By constructing a project-specific reference sequence database and index, combined with conditional quality value lossy compression, efficient compression of metagenomic data is achieved. Specific methods include constructing a basic reference sequence database based on the source of metagenomic data samples, constructing an index, determining sequence abundance distribution through comparison results, selecting representative reference genomes to construct compressed reference sequences, and performing degeneracy processing on quality values.

Benefits of technology

The compression efficiency of metagenomic data has been greatly improved, achieving an average compression ratio of nearly 4 times that of traditional compression, easing the storage and transmission pressure of large sample data while ensuring high fidelity and accuracy of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998885A_ABST
    Figure CN119998885A_ABST
Patent Text Reader

Abstract

The invention provides a construction method of a reference sequence for metagenome data compression, which comprises the following steps: constructing a basic reference sequence database according to a sample source of metagenome data; based on the basic reference sequence database, constructing an index of the basic reference sequence database; according to the index of the basic reference sequence database, a first read length sequence is compared with the basic reference sequence database, a comparison result is obtained, and the first read length sequence is a read length sequence of a part of samples randomly selected in metagenome data to be compressed; and according to the comparison result, determining the sequence abundance distribution of the first read length sequence, and constructing the reference sequence for metagenome data compression.
Need to check novelty before this filing date? Find Prior Art

Description

Reference sequence construction method, metagenomic data compression method and electronic device Technical Field

[0001] The present disclosure relates to the technical field of biological data compression, and in particular to a reference sequence construction method, a metagenomic data compression method, and an electronic device. Background Art

[0002] The metagenome is the sum of all microbial genomes in an environment. Metagenomics is a new microbial research method that studies the genomes of microbial populations in environmental samples, using functional gene screening and / or sequencing analysis to investigate microbial diversity, population structure, evolutionary relationships, functional activity, interactions, and relationships with the environment. The study of metagenomic data allows researchers to transcend species boundaries, more effectively exploit multispecies genetic resources, and uncover the laws governing life at a higher and more complex level.

[0003] The rapid decline in the cost of high-throughput sequencing has led to a massive increase in the output of genomic data, posing a huge challenge to data storage and transmission. Genetic data is primarily stored in the Fastq format, and the distribution of its sequence information and quality values ​​is highly random, making it impossible to efficiently compress it using general compression software such as gzip. The index-based Fastq file compression tool in related technologies improves compression efficiency by aligning short read sequences (Reads) with a reference genome and converting sequence information into positional information. This strategy is highly dependent on the integrity of the reference gene sequence, but the species composition in metagenomic data is relatively complex, and a significant improvement in compression efficiency cannot be achieved through a stable reference sequence.

[0004] Therefore, it is urgent to develop a method for constructing an effective metagenomic reference sequence and a compression method for metagenomic data based on the sequence to improve data compression efficiency.

[0005] Summary of the Invention

[0006] To this end, embodiments of the present disclosure provide a method for constructing a reference sequence for metagenomic data compression, a metagenomic data compression method, a metagenomic data compression device, an electronic device, a non-transitory computer-readable storage medium, a computer program product, and a computer program.

[0007] An embodiment of the first aspect of the present disclosure proposes a method for constructing a reference sequence for metagenomic data compression, comprising: constructing a basic reference sequence database based on the sample source of the metagenomic data; constructing an index of the basic reference sequence database based on the basic reference sequence database; comparing a first read sequence with the basic reference sequence database according to the index of the basic reference sequence database to obtain a comparison result, wherein the first read sequence is a read sequence of a randomly selected portion of samples in the metagenomic data to be compressed; and determining the sequence abundance distribution of the first read sequence according to the comparison result to construct the reference sequence for metagenomic data compression.

[0008] In some embodiments, constructing a basic reference sequence database based on the sample source of the metagenomic data includes: obtaining corresponding reference genomes from public databases and summarizing them according to the sample source of the metagenomic data to obtain the basic reference sequence database.

[0009] In some embodiments, based on the basic reference sequence database, an index of the basic reference sequence database is constructed, including: a single reference genome in the basic reference sequence database includes a first subsequence and a second subsequence, the first subsequence and the second subsequence are merged, and the number of the reference genome is retained to obtain a subsequence-merged reference genome; based on the subsequence-merged reference genome, an index of the basic reference sequence database is constructed.

[0010] In some embodiments, according to the index of the basic reference sequence database, the first read sequence is compared with the basic reference sequence database, including: based on the index of the basic reference sequence database, the first read sequence is compared to each of the subsequence merged reference genomes; based on the alignment of the first read sequence to the subsequence merged reference genome, the number of the reference genome to which the read sequence is aligned is recorded.

[0011] In some embodiments, determining the sequence abundance distribution of the first read sequence based on the alignment results and constructing the reference sequence for metagenomic data compression includes: counting the number of times the first read sequence aligns to each of the reference genomes in the alignment results to obtain the sequence abundance distribution of the first read sequence; and sorting the reference genomes based on the sequence abundance, selecting the top X reference genomes to construct the reference sequence for metagenomic data compression. In some embodiments, X can be 1000.

[0012] In some embodiments, constructing the reference sequence for metagenome data compression further includes: selecting, based on the sorting, reference genomes whose sum of sequence abundance percentages is greater than Y% to construct the reference sequence for metagenome data compression. In some embodiments, Y can be 80.

[0013] In some embodiments, the method for constructing a reference sequence for metagenomic data compression also includes: splitting the basic reference sequence database into sub-basic reference sequence databases; constructing indexes of the sub-reference sequence databases based on the split sub-basic reference sequence databases; based on the indexes of the sub-reference sequence databases, aligning the first read sequence with each of the sub-basic reference sequence databases to obtain a second alignment result, wherein the second alignment result includes sub-result files based on each of the sub-basic reference sequence databases.

[0014] In some embodiments, the method for constructing a reference sequence for metagenomic data compression further includes: respectively counting the number of the first read sequence in each of the sub-result files that is aligned to each of the sub-basic reference sequence databases to obtain the sequence abundance distribution of the first read sequence in each of the sub-result files; performing a first sorting of the reference genome according to the sequence abundance in each of the sub-result files, and selecting the reference genomes with the top X positions in sequence abundance in each of the sub-result files to construct a sub-reference sequence database; performing a second sorting of the reference genomes in the sub-reference sequence database according to the sequence abundance; and selecting the reference genomes with the top X positions in the sequence abundance distribution in the sub-reference sequence database to construct the reference sequence for metagenomic data compression.

[0015] In some embodiments, constructing the reference sequence for metagenome data compression further comprises: selecting, according to the first sorting, reference genomes whose sum of sequence abundance ratios in each sub-result file is greater than Y% to construct the sub-reference sequence database, and

[0016] According to the second ranking, a reference genome having a sum of sequence abundance percentages greater than Y% in the sub-reference sequence database is selected to construct the reference sequence for metagenome data compression. In some embodiments, Y can be 80.

[0017] In some embodiments, the method for constructing a reference sequence for metagenomic data compression further comprises: performing a first and / or second screening on the alignment results, wherein the first screening comprises: selecting the read sequences without insertions and / or deletions in the alignment results; and the second screening comprises: selecting the read sequences below a mismatch threshold. In some embodiments, the mismatch threshold may be 3.

[0018] An embodiment of the second aspect of the present disclosure proposes a metagenomic data compression method, the method comprising: constructing a reference sequence for metagenomic data compression according to the method for constructing a reference sequence for metagenomic data compression proposed in any embodiment of the first aspect of the present disclosure above; aligning a second read sequence with the reference sequence and recording the alignment result to obtain compressed data of the metagenomic data, wherein the second read sequence is a read sequence of a sample to be compressed in the metagenomic data.

[0019] In some embodiments, the second read sequence is aligned with the reference sequence and the alignment results are recorded, including: when the number of mismatched bases between the second read sequence and the reference sequence is less than R1, recording the position of the second read sequence on the reference sequence; when the number of mismatched bases between the second read sequence and the reference sequence is greater than R1 and less than R2, recording the position of the paired base in the second read sequence on the reference sequence and recording the base information of the mismatched base; when the number of mismatched bases between the second read sequence and the reference sequence is greater than R2, recording the second read sequence. In some embodiments, R1, R2, and R3 are all integers greater than or equal to 0. In some embodiments, R1 is 0 to 5, and R2 is 3 to 10. In some embodiments, R1 is 0 to 2, and R2 is 3 to 8. In some embodiments, R1 is 0 and R2 is 3.

[0020] In some embodiments, the metagenomic data compression method further includes degenerating the quality value of the metagenomic data.

[0021] In some embodiments, degenerating the quality values ​​of the metagenomic data comprises: performing statistics on the quality values ​​of the bases in the metagenomic data to obtain a distribution of the quality values ​​within M quality value ranges; and mapping the quality values ​​within the M ranges to M mapping values ​​to degenerate the quality values ​​of the metagenomic data. In some embodiments, M is an integer greater than 0.

[0022] In some embodiments, the metagenomic data compression method further includes: when the proportion of bases with quality values ​​lower than Q accounts for less than a set proportion N of all bases in the metagenomic data, mapping the quality values ​​of all bases in the metagenomic data to degenerate the quality values ​​of the metagenomic data.

[0023] In some embodiments, the metagenomic data compression method further includes: when the proportion of bases with quality values ​​lower than Q accounts for a proportion of all bases in the metagenomic data that is higher than or equal to a set proportion N, mapping the quality values ​​of the bases with quality values ​​higher than Q in the metagenomic data to degenerate the quality values ​​of the metagenomic data.

[0024] In some embodiments, the metagenomic data compression method further includes: retaining the original quality values ​​of the bases with quality values ​​lower than Q in the metagenomic data when the proportion of bases with quality values ​​lower than Q accounts for a proportion higher than or equal to a set proportion N of all bases in the metagenomic data.

[0025] In some embodiments, Q is a quality value corresponding to a base error probability of 0.01% to 1%. In some embodiments, N is greater than or equal to 10%. In some embodiments, N is greater than or equal to 20%.

[0026] The third aspect of the present disclosure provides a metagenomic data compression device, comprising: a reference sequence construction module for constructing a reference sequence for metagenomic data compression according to the method for constructing a reference sequence for metagenomic data compression according to any embodiment of the first aspect of the present disclosure; and

[0027] A data compression module is used to compare the read sequence in the metagenome data with the reference sequence and record the comparison results to obtain compressed data of the metagenome data.

[0028] In some embodiments, the apparatus further comprises: a quality value degeneration module for degenerating the quality values ​​of the metagenomic data.

[0029] An embodiment of the fourth aspect of the present disclosure proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for constructing a reference sequence for metagenomic data compression as described in any embodiment of the first aspect of the present disclosure is implemented, the method comprising: constructing a basic reference sequence database according to the sample source of the metagenomic data; constructing an index of the basic reference sequence database based on the basic reference sequence database; comparing a first read sequence with the basic reference sequence database according to the index of the basic reference sequence database to obtain a comparison result, wherein the first read sequence is a read sequence of a randomly selected portion of samples in the metagenomic data to be compressed; and determining the sequence abundance distribution of the first read sequence according to the comparison result to construct the reference sequence for metagenomic data compression.

[0030] An embodiment of the fifth aspect of the present disclosure proposes a non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method for constructing a reference sequence for metagenomic data compression as described in any embodiment of the first aspect of the present disclosure is implemented.

[0031] The sixth aspect of the present disclosure provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the method for constructing a reference sequence for metagenomic data compression as described in any embodiment of the first aspect of the present disclosure.

[0032] The embodiments of the present disclosure achieve the following beneficial effects:

[0033] The method for constructing an effective metagenomic reference sequence and compressing metagenomic data based on the sequence proposed in the present disclosure can construct an effective metagenomic data compression reference sequence. By using index-dependent compression tools, the compression efficiency of metagenomic data can be greatly improved (the average compression ratio achieved is nearly 4 times that of traditional compression ratio), thereby alleviating the storage and transmission pressure of metagenomic data with large sample sizes. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0035] FIG1 is a diagram of a method for constructing a reference sequence for metagenome data compression according to an embodiment of the present disclosure;

[0036] FIG2 is a technical solution diagram for constructing a reference sequence for metagenome data compression according to an embodiment of the present disclosure;

[0037] FIG3 is a diagram of a method for constructing a reference sequence based on a reference genome with high sequence abundance according to an embodiment of the present disclosure;

[0038] FIG4 is a flow chart of a metagenomic data compression method according to an embodiment of the present disclosure;

[0039] FIG5 is a flow chart of metagenomic data compression based on reference sequences according to an embodiment of the present disclosure;

[0040] FIG6 is an example diagram of a quality value mapping table according to an embodiment of the present disclosure;

[0041] FIG7 is a flowchart of quality value degeneration according to an embodiment of the present disclosure;

[0042] FIG8 is a flowchart of conditional quality value degeneration according to one embodiment of the present disclosure;

[0043] FIG9 is a flowchart of conditional quality value degeneration according to another embodiment of the present disclosure;

[0044] FIG10 is a diagram of a metagenomic data compression method according to another embodiment of the present disclosure;

[0045] FIG11 is a structural diagram of a metagenomic data compression device according to an embodiment of the present disclosure;

[0046] FIG12 illustrates a block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure;

[0047] FIG13 is a diagram showing a specific data compression ratio distribution according to an embodiment of the present disclosure;

[0048] Figure 14 is a statistical graph of the Pearson correlation coefficient of the species composition of 233 samples before and after mass value degeneration. DETAILED DESCRIPTION

[0049] The present disclosure is further described in detail below in conjunction with specific embodiments. The examples provided are only for the purpose of illustrating the present disclosure and are not intended to limit the scope of the present disclosure. The examples provided below can serve as a guide for further improvements by those of ordinary skill in the art and do not in any way limit the present disclosure.

[0050] The present disclosure is made based on the following knowledge of the inventors:

[0051] In the related art, index-based (also referred to as index-dependent in the disclosed embodiments) compression tools are used to compress metagenomic data. In index-dependent compression tools for metagenomic data, the construction of reference sequences and data compression are usually achieved through the following two methods.

[0052] Method 1: Construct a universal reference sequence based on a public database. For example, for data with a clear source, such as intestinal microorganisms, a reference sequence can be constructed by compiling the genomes of all possible species in the database.

[0053] Method 2: Constructing sample-specific reference sequences based on species composition and sequence assembly. MetaCRAM (Kim, M. et al., 2016) and MCUIUC (Ligo, J. et al., 2013) first use metagenomic species identification tools to rapidly identify the species composition of the data. Based on the species identification results, the user selects species with abundance above a specific threshold as reference genome sources for constructing appropriate reference genomes. Reads that failed alignment are then assembled de novo to construct new reference sequences. Finally, metagenome data are compressed based on the reference sequences selected from the database and the de novo constructed reference sequences, respectively.

[0054] However, although the strategy of constructing a universal reference sequence based on a public database in Method 1 can cover as many species as possible by expanding the number of reference genomes, due to the large variety of microbial species, the final constructed reference sequence file is extremely large, which places very high demands on computer configuration (especially memory), making it difficult for users using small-scale computing clusters or personal computers to operate.

[0055] While Method 2's strategy of constructing sample-specific reference sequences based on species composition and sequence assembly achieves optimal compression efficiency while keeping memory requirements within acceptable limits, in practice, species identification and de novo sequence assembly are time-consuming, ultimately resulting in slow data compression. For example, MetaCRAM takes 73 minutes to compress an 8,230MB Fastq file.

[0056] The method for constructing a reference sequence for metagenomic data compression proposed in the embodiments of the present disclosure achieves efficient, index-dependent data compression of metagenomic data by constructing a project-specific reference sequence and combining it with conditional quality value lossy compression. The method for constructing a reference sequence for metagenomic data compression proposed in the embodiments of the present disclosure and the metagenomic data compression method based on the constructed reference sequence significantly improve the compression efficiency of metagenomic data and effectively alleviate the storage and transmission pressure of large sample sizes of metagenomic data.

[0057] The first embodiment of the present disclosure proposes a method for constructing a reference sequence for metagenomic data compression.

[0058] Figure 1 is a schematic diagram of a method for constructing a reference sequence for metagenome data compression according to an embodiment of the present disclosure. As shown in Figure 1 , the method may include steps 101-104.

[0059] Step 101: Construct a basic reference sequence database based on the sample source of the metagenomic data.

[0060] In the disclosed embodiments, the "sample source" refers to the environment from which the metagenome data sample to be compressed was extracted. In the disclosed embodiments, the sample may be intestinal microorganisms, water source microorganisms, soil microorganisms, etc., and the sample source may be the intestinal tract, water source, soil, etc.

[0061] In the embodiments of the present disclosure, based on the project background information (such as intestinal microorganisms, water source microorganisms, soil microorganisms, etc.) or the sample source, a corresponding public database can be selected and commonly used sequences can be downloaded and aggregated as a basic reference sequence library for the construction of the comparison index. In the embodiments of the present disclosure, the intestinal microorganism database can be GMrepo (Dai, D. et al., 2022), gutMEGA- (Zhang, Q. et al., 2021), and uhgg (Almeida, A. et al., 2021).

[0062] Step 102: Based on the basic reference sequence database, construct an index of the basic reference sequence database.

[0063] In the disclosed embodiments, an index-dependent alignment software or script is used to index the basic reference sequence database. In some embodiments, the index-dependent alignment software can be bwa (Burrows-Wheeler Aligner, Li H. and Durbin R. (2009) Fast and accurate short read alignment with Burrows-Wheeler Transform. Bioinformatics, 25: 1754-60. [PMID: 19451168]), Bowtie (Langmead B, Trapnell C, Pop M, Salzberg SL. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol 10: R25), or Bowtie2 (Langmead B, Salzberg S. Fast gapped-read alignment with Bowtie 2. Nature Methods. 2012, 9: 357-359).

[0064] Step 103: According to the index of the basic reference sequence database, the first read sequence is compared with the basic reference sequence database to obtain an alignment result, wherein the first read sequence is the read sequence of a portion of samples randomly selected from the metagenome data to be compressed.

[0065] In the disclosed embodiment, data from a portion of the samples to be compressed (i.e., reads) can be randomly selected for comparison with a base reference sequence database. It is understood that, compared to using a large sample size of the entire sample, randomly selecting a specific number of samples for preliminary comparison can effectively improve comparison efficiency and save computing resources.

[0066] Step 104: Determine the sequence abundance distribution of the first read sequence based on the alignment results, and construct a reference sequence for metagenomic data compression.

[0067] In the disclosed embodiment, sequence abundance refers to the number of sample reads aligned to each reference genome in the alignment. In the disclosed embodiment, the sequence abundance distribution of the reads of the selected sample is obtained by counting the number of reads (i.e., the first read sequence) of the selected sample in the alignment result outputted in step 103 aligned to each reference genome; the reference genomes in the basic reference sequence database are sorted according to sequence abundance, and the top-ranked reference genomes are selected according to the user's own computing configuration, compression ratio requirements or other personalized needs to construct a reference sequence for compressing the metagenome data of all samples.

[0068] Figure 2 is a technical solution diagram for constructing a reference sequence for metagenomic data compression according to an embodiment of the present disclosure. As shown in Figure 2, the method for constructing a reference sequence for metagenomic data compression proposed in the embodiment of the present disclosure may include determining a microbial database of a specific major category from a public microbial database based on project information, and obtaining a basic reference sequence database from the microbial database of a specific major category; using part of the samples to be compressed (i.e., test samples) to compare with the basic reference sequence database, obtaining the sequence abundance of each reference genome in the basic reference sequence database of the reads of the part of the samples and sorting them to obtain the sequence abundance distribution of the part of the samples; selecting the reference genomes of high-abundance species (i.e., the reference genomes in the basic reference sequence database with the highest sequence abundance ranking) for merging, thereby obtaining a project-specific (i.e., for the project) reference sequence for subsequent index-dependent data compression of metagenomic data. The method proposed in the embodiment of the present disclosure can effectively improve the comparison efficiency and save computing resources by randomly selecting a specific number of samples for preliminary comparison and constructing a reference sequence based on sequence abundance. At the same time, by selecting the top-ranked representative reference genomes based on the sequence abundance distribution of some samples to construct reference sequences for compression, the data volume of the reference sequences is greatly reduced, which is conducive to efficient alignment and compression in the later stage.

[0069] In an embodiment of the present disclosure, step S102 may also include: a single reference genome in the basic reference sequence database includes a first subsequence and a second subsequence, merging the first subsequence and the second subsequence, and retaining the number of the reference genome to obtain a subsequence-merged reference genome; based on the subsequence-merged reference genome, constructing an index of the basic reference sequence database.

[0070] In the disclosed embodiments, the first subsequence or the second subsequence can be a fragment sequence in a Fastq file of a single reference genome, such as the sequence of each chromosome in the reference genome. In the disclosed embodiments, by merging the subsequences in the Fastq file of a single reference genome, only the reference genome number is retained as the only sequence description line, which effectively reduces the size of the basic reference sequence database and facilitates the statistics of subsequent alignment results.

[0071] In an embodiment of the present disclosure, the basic reference sequence database can also be split into several sub-basic reference sequence databases and indexes of the sub-reference sequence databases can be constructed based on the split sub-basic reference sequence databases respectively; based on the indexes of the sub-reference sequence databases, the read sequences of some randomly selected samples (i.e., the first read sequences) are respectively compared with each sub-basic reference sequence database to obtain a second comparison result, wherein the second comparison result includes sub-result files based on each of the sub-basic reference sequence databases. It can be understood that in actual applications, the computing configuration of some users is not sufficient to perform operations based on a larger basic reference sequence database. Therefore, by splitting the basic reference sequence database into sub-basic reference sequence databases and performing operations based on the sub-basic reference sequence databases respectively, the requirements for the user's computing configuration are effectively reduced, so that the application threshold of the index construction method proposed in the embodiment of the present disclosure is lowered, and its application range is wider.

[0072] Figure 3 is a diagram of a method for constructing a reference sequence based on a high-sequence-abundance reference genome according to an embodiment of the present disclosure. As shown in Figure 3, the number of randomly selected samples to be compressed (i.e., the number of test samples) is A, and the number of base reference sequence databases (i.e., the number of base sequence index files) is B, where B = 1 corresponds to not splitting the base reference sequence database; B ≥ 2 corresponds to splitting the base reference sequence database into several sub-base reference sequence databases. X is the sequence abundance selection threshold selected by the user for the reference genome.

[0073] In an embodiment of the present disclosure, it is possible to perform alignment and sequence abundance screening for a single overall basic reference sequence database to determine a reference sequence database for metagenome data compression without splitting the basic reference sequence database. Specifically, when the subsequences of a single reference genome are merged and only the number of a single genome is retained, the read sequence (i.e., the first read sequence) of a randomly selected portion of the sample can be aligned to each subsequence merged reference genome based on the index of the basic reference sequence database; when the read sequence of the portion of the sample is aligned to the subsequence merged reference genome, the number of the reference genome to which the read sequence is aligned is recorded. The alignment software can be a script or software that is index-dependent and written locally, such as Bwa, Bowtie, Bowtie2, or a locally written script or software. In an embodiment of the present disclosure, after obtaining the alignment result, the number of the read sequences of the portion of the sample aligned to the number of each reference genome in the alignment result is counted to obtain the sequence abundance distribution of the read sequences of the portion of the sample; the reference genome is sorted according to the sequence abundance, and the reference genomes in the top X positions are selected to construct a reference sequence for metagenome data compression.

[0074] As shown in Figure 3, when the number B of the basic reference sequence database is 1, each sample in the test sample (i.e., part of the samples, number A) is compared with the basic reference sequence database to obtain the sequence abundance of each reference genome in the basic reference sequence database for the A samples; the sequence abundance of the A samples is merged and sorted to obtain the overall sequence abundance distribution of the A test samples; the first X reference genomes are selected to construct the reference sequence for metagenomic data compression.

[0075] In an embodiment of the present disclosure, when a basic reference sequence database can be split, the sub-basic reference sequence database after the split is compared and the sequence abundance screening is performed to determine the reference sequence database for metagenome data compression. Specifically, as shown in Figure 3, when B≥2, each sample in the test sample (i.e., a portion of the sample, the number is A) is compared with each sub-basic reference sequence database to obtain a sub-result file; the number of read sequences of the test sample in each sub-result file is respectively counted to each sub-basic reference sequence database to obtain the sequence abundance distribution of the read sequence of the test sample in each sub-result file, wherein there are B sub-result files, including A*B sequence abundance distributions; the reference genome is first sorted according to the sequence abundance in the B sub-result files, and the reference genomes with the top X positions in the sequence abundance in the B sub-result files are respectively selected to construct a sub-reference sequence database, that is, the sub-reference sequence database includes B*X reference genomes; according to the sequence abundance, the B*X reference genomes in the sub-reference sequence database are second sorted, and the reference genomes with the top X positions in the sequence abundance distribution in the sub-reference sequence database are selected to construct a reference sequence for metagenome data compression.

[0076] It can be understood that in the embodiments of the present disclosure, when the basic reference sequence database is split into several sub-basic reference sequence databases, the subsequences within a single reference genome in the basic reference sequence database can be merged before the splitting, and the number of the reference genome can be retained for subsequent comparison; or after the splitting, the subsequences within a single reference genome in the sub-basic reference sequence databases obtained by the splitting can be merged, and the number of the reference genome can be retained for subsequent comparison.

[0077] In the disclosed embodiments, the sequence abundance selection threshold X can be selected by the user based on data conditions, personal computing resources, or compression requirements. In some embodiments, X can be between 200 and 5000. In some embodiments, X can be between 500 and 3000. In some embodiments, X can be 1000.

[0078] In the embodiment of the present disclosure, based on the statistics and sorting results of sequence abundance, reference genomes with a sum of sequence abundance ratios greater than Y% can be selected to construct reference sequences for metagenomic data compression, where the ratio is the ratio of the sequence abundance corresponding to a certain reference genome to the total sequence abundance. Selecting a reference genome with a sum of sequence abundance ratios greater than Y% means selecting the top several reference genomes according to the statistics and sorting of sequence abundance, so that the sum of the sequence abundance ratios of the top several selected reference genomes is greater than Y%.

[0079] It will be appreciated that, in the disclosed embodiments, Y can be determined based on the sample size, the desired compression ratio, and the user's computing resources. In some embodiments, Y can be between 20 and 80. In some embodiments, Y can be between 40 and 80. In some embodiments, Y can be 80. In the disclosed embodiments, compared to using all reference genomes in the basic reference sequence database, the use of representative reference genomes does not affect the accuracy of subsequent data compression, that is, the compression performed based on the index constructed based on all reference genomes in the basic reference sequence database with a huge amount of data, the compressed data composition is highly correlated with the data composition after compression using the representative reference genome in the disclosed embodiments. Therefore, by selecting a representative reference genome ranked high in sequence abundance to construct a compressed index, the volume of the compressed index is effectively reduced, the amount of subsequent compression operations is greatly reduced, and the high fidelity of the compressed data is ensured.

[0080] In an embodiment of the present disclosure, after comparing a randomly selected portion of samples with a basic reference sequence database or a sub-basic reference sequence database, the comparison results may be subjected to a first and / or second screening, wherein the first screening includes: selecting read sequences without insertions and / or deletions in the comparison results; and the second screening includes: selecting read sequences below a mismatch threshold.

[0081] In the disclosed embodiments, after obtaining the alignment result after Bwa, Bowtie2 or the script of the same function, the result file (such as Bam or Sam format) produced by comparison can be compared to perform the first and / or second screening to carry out quality control of the comparison result. In certain embodiments, in the first screening, Reads without insertion and / or deletion can be selected according to the Cigar value of the result file (Concise Idiosyncratic Gapped Alignment Report), wherein no insertion and / or deletion is represented by 100M or 150M (100 and 150 represent Reads length of 100bp and 150bp, M represents Match, and 100M or 150M represents that the full-length sequence of 100bp or 150bp of Reads is completely matched with the reference sequence). In certain embodiments, in the second screening, Reads whose number of mismatches is lower than the mismatch threshold can be selected according to the N:M value of the result file. In certain embodiments, the mismatch threshold can be 1 to 10. In certain embodiments, the mismatch threshold can be 1 to 5. In certain embodiments, the mismatch threshold can be 3. It is understandable that the screening of reads in the comparison results removes reads with high mismatches, thereby improving the overall credibility of the reads and making the selection of the reference genome based on the sequence abundance distribution of the filtered high-confidence reads more accurate.

[0082] The method for constructing a reference sequence for metagenomic data compression proposed in the embodiments of the present disclosure effectively solves the problem that the basic reference sequence database has a large amount of data and users of small computing clusters or personal computers cannot construct the index required for alignment for a single Fastq file containing tens of thousands of reference genomes at one time by merging the subsequences of a single reference genome in a basic reference sequence database and retaining only their numbers, and / or splitting the basic reference sequence database into multiple sub-basic reference sequence databases. At the same time, the method randomly selects some samples for preliminary alignment and reference sequence construction, which greatly reduces the input and output of data in the alignment while ensuring that the constructed reference sequence has the greatest possible coverage of the data to be compressed, thereby improving the efficiency of reference sequence construction and saving computing and storage resources.

[0083] The second embodiment of the present disclosure provides a method for compressing metagenome data. FIG4 is a flow chart of the method for compressing metagenome data according to an embodiment of the present disclosure. As shown in FIG3 , the method includes:

[0084] Step 201: constructing a reference sequence for metagenomic data compression according to the method for constructing a reference sequence for metagenomic data compression described in any one of the embodiments of the first aspect above;

[0085] Step 202: align the second read sequence with the reference sequence and record the alignment result to obtain compressed data of the metagenomic data, wherein the second read sequence is the read sequence of the sample to be compressed in the metagenomic data.

[0086] In an embodiment of the present disclosure, after constructing a reference sequence for metagenome data compression based on a first read sequence, that is, a read sequence of a portion of samples randomly selected from the metagenome data to be compressed, the read sequences of some or all of the samples in the metagenome data can be compressed based on the reference sequence, that is, the second read sequence is compressed. It is understandable that the second read sequence can be the same as or different from the first read sequence. In an embodiment of the present disclosure, based on the reference sequence constructed for compression, all or part of the samples in the metagenome data can be compressed according to user needs, thereby achieving efficient compression while improving the flexibility of compression.

[0087] FIG5 is a flowchart of a reference sequence-based metagenomic data compression process according to an embodiment of the present disclosure. As shown in FIG5 , after constructing a reference sequence according to any of the embodiments of the first aspect of the present disclosure, the reads (Fastq files) in the metagenomic data to be compressed are input and aligned with the constructed reference sequence.

[0088] In some embodiments, when the number of mismatched bases between the read sequence (i.e., the second read sequence) in the metagenomic data and the reference sequence is less than R1, the position of the read sequence on the reference sequence is recorded; when the number of mismatched bases between the read sequence in the metagenomic data and the reference sequence is greater than R1 and less than R2, the position of the paired base in the read sequence on the reference sequence is recorded, and the base information of the mismatched bases in the read is recorded; when the number of mismatched bases between the read sequence in the metagenomic data and the reference sequence is greater than R2, the read sequence is recorded. In some embodiments, R1, R2, and R3 are all integers greater than or equal to 0. In some embodiments, R1 is 0 to 5, and R2 is 3 to 10. In some embodiments, R1 is 0 to 2, and R2 is 3 to 8. In some embodiments, R1 is 0 and R2 is 3.

[0089] In some embodiments, i. when the read sequence in the metagenomic data completely matches the reference sequence (i.e., R1=0), the position of the read sequence on the reference sequence is recorded; ii. when the number of mismatched bases between the read sequence in the metagenomic data and the reference sequence is greater than or equal to 3 (i.e., R2=3), the position of the paired base in the read sequence on the reference sequence is recorded, and the base information of the mismatched base is recorded; iii. when the read sequence in the metagenomic data cannot match the reference sequence (i.e., the number of mismatched bases is greater than 3), the read sequence is recorded.

[0090] In some embodiments, in step ii, if a read has a mismatch with the reference sequence and the number of mismatched bases is less than 5 (i.e., R1 = 1-4), the position of the read's paired base on the reference sequence is recorded, i.e., the matched base is converted into position information for storage, and the actual base information of the mismatched base is recorded. In some embodiments, the number of mismatched bases in step ii can be 1 to 3 (i.e., R1 = 1, 2, or 3).

[0091] In some embodiments, in step iii, if a read has a mismatch with the reference sequence and the number of mismatched bases is greater than 5 (i.e., R2 ≥ 5), the sequence information of the read is recorded, i.e., the actual base information of the read is retained. In some embodiments, the number of mismatched bases in step iii can be a positive integer greater than 3 (i.e., R2 > 3).

[0092] In an embodiment of the present disclosure, the metagenomic data compression method further includes: degenerating the quality value of the metagenomic data.

[0093] It is understandable that metagenomic data is often stored in the form of Fastq files. The Fastq format is divided into four lines, where the characters in the fourth line correspond to the probability of each base in the sequence being misidentified, i.e., the base quality score (Q-score). In other words, the base quality score is an integer mapping of the probability of base call error, which can be Q = -10*lgP, where P is the probability of base call error.

[0094] Base quality values ​​are represented using different systems depending on the sequencing platform, such as the Phred33 system and the Phred64 system. These systems use different characters to represent base quality values, but all can be converted to the base error probability using the formula Q = -10 * lgP. In the disclosed embodiment, the base quality value is divided into 0 to 40 based on the probability of base error, where 0 represents a 100% error probability and 40 represents a 0.01% error probability.

[0095] In an embodiment of the present disclosure, the quality values ​​of the metagenomic data are degenerated, including: performing statistics on the base quality values ​​in the metagenomic data to obtain the distribution of the quality values ​​within M quality value ranges; and mapping the quality values ​​within the M ranges to M mapping values ​​respectively to degenerate the quality values ​​of the metagenomic data.

[0096] In an embodiment of the present disclosure, M quality value ranges are set according to different base error probabilities, and corresponding M specific mapping values ​​are set to map the base quality values ​​to complete degeneration, where M can be an integer greater than 0, such as any integer from 1 to 100. In one embodiment of the present disclosure, all quality values ​​are divided into four grades (i.e., M=4) according to the error probability they represent, namely 0 to 3 (error probability >50%), 4 to 19 (error probability 1% to 40%), 20 to 30 (error probability 0.1% to 1%), and 30 to 40 (error probability 0.01% to 0.1%). In another embodiment, all quality values ​​are divided into three grades (i.e., M=3) according to the error probability they represent. It will be understood that the specific value of M and the M specific ranges can be determined and adjusted according to actual needs.

[0097] In the disclosed embodiment, the M specific mapping values ​​can be adjusted by the user based on the actual data situation, and this disclosure does not impose any restrictions on this. Figure 6 is an example of a quality value mapping table according to an embodiment of the disclosed embodiment. As shown in Figure 6, the quality values ​​0 to 40 can be divided into M quality value ranges, with Q1, Q2, ..., QM as the corresponding specific mapping values.

[0098] FIG7 is a flowchart of the quality value degeneration according to an embodiment of the present disclosure. As shown in FIG7 , the quality values ​​of the Reads to be compressed are counted and different threshold ranges are defined, such as [a, b], [c, d], [e, f], ..., etc., a total of M, where af represents different quality values. For example, when the quality values ​​of the bases are divided into 0 to 40 and M=3, [a, b] can be 0 to 10; [c, d] can be 11 to 20; [e, f] can be 21 to 40. After the base quality values ​​of the Reads to be compressed are divided into M threshold ranges, the bases falling into the same threshold range are mapped to the same specific mapping value, thereby degenerating the Reads to be compressed, thereby reducing the volume of the data to be compressed and reducing the amount of redundant calculations.

[0099] During specific calculations, the inventors of the present disclosure discovered that in metagenomic data with low overall quality values, fluctuations in medium and low-level quality values ​​will affect the alignment quality values ​​of some alignment software (such as when using Bowtie2, the alignment quality value is expressed in MAPQ), thereby affecting downstream analysis. Therefore, in the degeneration of metagenomic data quality values, the embodiments of the present disclosure also propose a technical solution for conditional degeneration of metagenomic data quality values ​​to reduce the impact of lossy compression of quality values ​​on downstream analysis.

[0100] Specifically, in the embodiment of the present disclosure, after the quality values ​​of the reads to be compressed are counted and before the reads are degenerated, it also includes: when the proportion of bases with quality values ​​lower than Q to all bases in the metagenomic data is lower than the set proportion N, the quality values ​​of all bases in the metagenomic data are mapped to degenerate the quality values ​​of the metagenomic data.

[0101] In an embodiment of the present disclosure, when the proportion of bases with quality values ​​lower than Q accounts for a proportion of all bases in the metagenomic data that is higher than or equal to a set proportion N, the quality values ​​of the bases with quality values ​​higher than Q in the metagenomic data are mapped to degenerate the quality values ​​of the metagenomic data.

[0102] In an embodiment of the present disclosure, when the proportion of bases with quality values ​​lower than Q accounts for a proportion of all bases in the metagenomic data that is higher than or equal to a set proportion N, the original quality values ​​of the bases with quality values ​​lower than Q in the metagenomic data are retained.

[0103] It is understood that in the embodiments of the present disclosure, Q can be determined based on the actual quality value distribution of the metagenomic data and the desired degeneracy. For example, when the quality value of the base is divided into 0 to 40, Q can be any integer between 0 and 40, that is, corresponding to a range of base error probability of 100% to 0.01%. In the embodiments of the present disclosure, Q can be the quality value corresponding to a base error probability of 0.01% to 1%. In some embodiments, Q can be the quality value corresponding to a base error probability of 0.1% to 1%.

[0104] In the embodiment of the present disclosure, the ratio N is set to be greater than or equal to 20%. In other embodiments, N is greater than or equal to 10%.

[0105] FIG8 is a flow chart of conditional quality value degeneration according to one embodiment of the present disclosure. As shown in FIG8 , the quality values ​​of bases range from 0 to 40, and the quality values ​​0 to 40 are divided into four quality value ranges (i.e., M = 4), namely 0 to 3 (error probability > 50%), 4 to 19 (error probability 1% to 40%), 20 to 30 (error probability 0.1% to 1%), and 30 to 40 (error probability 0.01% to 0.1%), and the mapping values ​​corresponding to the four quality value ranges are Q1, Q2, Q3, and Q4, respectively. According to Figure 7, the quality value statistics of the reads to be compressed are performed to obtain the distribution of reads in four quality value ranges, namely R1, R2, R3 and R4; it is determined whether the sum of the proportions of bases with quality values ​​lower than Q=29 in the metagenomic data is greater than or equal to the set proportion N, that is, whether R1%+R2%+R3% is greater than or equal to N%; if not, all bases in the compressed data are degenerated according to the quality value mapping table, that is, the quality values ​​of bases with quality values ​​of 0 to 3 will be mapped and degenerated to Q1, and the quality values ​​of bases with quality values ​​of 4 to 19 will be mapped and degenerated to Q2. The quality value of the base with a quality value of Q=29 will be mapped and degenerated to Q2, the quality value of the base with a quality value of 20 to 29 will be mapped and degenerated to Q3, and the quality value of the base with a quality value of 30 to 40 will be mapped and degenerated to Q4; if R1%+R2%+R3% is greater than or equal to N%, the bases with a quality value less than or equal to Q=29 will not be degenerated, that is, the original quality value of the base with a quality value less than or equal to Q=29 in the metagenomic data will be retained, and the bases with a quality value greater than Q=29 (i.e. greater than or equal to 30) will be degenerated according to the quality value mapping table.

[0106] FIG9 is a flow chart of conditional quality value degeneration according to another embodiment of the present disclosure. As shown in FIG9 , this flow differs from the flow shown in FIG8 only in that if R1% + R2% + R3% is greater than or equal to N%, the original quality values ​​of all bases in the compressed data are retained without degeneration.

[0107] Figure 10 is a diagram of a metagenomic data compression method according to an embodiment of the present disclosure. As shown in Figure 10, the method may include constructing an index for compression, conditional quality value degeneration of the data to be compressed (Fastq file), and data compression based on the constructed reference index.

[0108] The metagenomic data compression method proposed in the second embodiment of the present disclosure is a reference sequence constructed by the construction method of the reference sequence for metagenomic data compression described in any embodiment of the first embodiment above, and the read to be compressed is quickly compared with the constructed reference sequence. If it can be accurately compared to the corresponding position, it is only necessary to record the position information of the corresponding read on the reference sequence; if there is a small amount of mismatch, while recording the position information of the remaining paired bases, the information of the mismatched bases is retained; for the reads that cannot be accurately compared to the reference sequence, all sequence information is recorded, thereby greatly improving the compression efficiency of the metagenomic data and alleviating the storage pressure of metagenomic data with large sample sizes. In addition, the metagenomic data compression method proposed in the embodiment of the present disclosure conditionally degenerates the base quality values ​​before compression, that is, by setting a threshold, degenerates the bases with high quality values, and retains the original quality values ​​of the bases with medium and low quality values, thereby achieving simplification and reduction of the data to be compressed without affecting subsequent comparisons; at the same time, the data to be compressed based on the degenerate quality value further improves the compression efficiency.

[0109] The third embodiment of the present disclosure proposes a metagenomic data compression device. Figure 11 is a structural diagram of the metagenomic data compression device according to the embodiment of the present disclosure. As shown in Figure 11, the metagenomic data compression device 90 may include: a reference sequence construction module 901, which is used to construct a reference sequence for metagenomic data compression according to the construction method of the reference sequence for metagenomic data compression described in any embodiment of the first embodiment; and a data compression module 902, which is used to compare the read sequence in the metagenomic data with the reference sequence and record the comparison results to obtain compressed data of the metagenomic data.

[0110] In the embodiment of the present disclosure, the apparatus 90 may further include: a quality value degeneration module 903 for degenerating the quality values ​​of the metagenomic data.

[0111] The metagenomic data compression device proposed in the third embodiment of the present disclosure is a reference sequence constructed by the reference sequence construction method for metagenomic data compression described in any embodiment of the first embodiment, and the read to be compressed is quickly compared with the constructed reference sequence. If it can be accurately compared to the corresponding position, it is only necessary to record the position information of the corresponding read on the reference sequence; if there is a small amount of mismatch, while recording the position information of the remaining paired bases, the information of the mismatched bases is retained; for the reads that cannot be accurately compared to the reference sequence, all sequence information is recorded, thereby greatly improving the compression efficiency of the metagenomic data and alleviating the storage pressure of metagenomic data with large sample sizes. In addition, the metagenomic data compression device proposed in the embodiment of the present disclosure conditionally degenerates the base quality values ​​before compression, that is, by setting a threshold, degenerates the bases with high quality values, and retains the original quality values ​​of the bases with medium and low quality values, thereby achieving simplification and reduction of the data to be compressed without affecting subsequent comparisons; at the same time, the data to be compressed based on the degenerate quality value further improves the compression efficiency.

[0112] In order to implement the above embodiments, the embodiments of the present disclosure also propose an electronic device, including: a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the method for constructing a reference sequence for metagenomic data compression proposed in the first embodiment of the present disclosure or the method for compressing metagenomic data proposed in the second embodiment of the present disclosure.

[0113] In order to implement the above embodiments, the embodiments of the present disclosure also propose a non-transitory computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, it implements the method for constructing a reference sequence for metagenomic data compression as proposed in the first embodiment of the present disclosure or the metagenomic data compression method as proposed in the second embodiment of the present disclosure.

[0114] In order to implement the above embodiments, the embodiments of the present disclosure also propose a computer program product. When the instruction processor in the computer program product is executed, it executes the method for constructing a reference sequence for metagenomic data compression proposed in the first embodiment of the present disclosure or the metagenomic data compression method proposed in the second embodiment of the present disclosure.

[0115] In order to implement the above embodiments, the embodiments of the present disclosure also propose a computer program, which includes computer program code. When the computer program code is run on a computer, it enables the computer to execute the method for constructing a reference sequence for metagenomic data compression proposed in the first embodiment of the present disclosure or the method for compressing metagenomic data proposed in the second embodiment of the present disclosure.

[0116] Figure 12 shows a block diagram of an exemplary computer device suitable for implementing the embodiments of the present disclosure. The electronic device 12 shown in Figure 12 is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0117] 12 , electronic device 12 is implemented as a general-purpose computing device. Components of electronic device 12 may include, but are not limited to, one or more processors or processing units 16 , system memory 28 , and a bus 18 connecting various system components (including system memory 28 and processing unit 16 ).

[0118] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnection (PCI) bus.

[0119] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0120] Memory 28 may include computer-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer-readable storage media. By way of example only, storage system 34 may be configured to read and write to non-removable, non-volatile magnetic media (not shown in FIG. 10 , and commonly referred to as a "hard drive").

[0121] Although not shown in FIG12 , a disk drive for reading and writing to a removable non-volatile disk (e.g., a floppy disk) and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a Compact Disc Read Only Memory (CD-ROM), a Digital Video Disc Read Only Memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data media interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present disclosure.

[0122] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 42 generally implement the functions and / or methods of the embodiments described herein.

[0123] The electronic device 12 can also communicate with one or more external devices 14 (e.g., a keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with the electronic device 12, and / or any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). This communication can occur via an input / output (I / O) interface 22. Furthermore, the electronic device 12 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the electronic device 12 via the bus 18. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with the electronic device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0124] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the training method of the prediction model mentioned in the above embodiment.

[0125] It should be noted that the above explanations of the method for constructing a reference sequence for metagenomic data compression and the embodiments of the metagenomic data compression method are also applicable to the devices, electronic devices, non-transitory computer-readable storage media, computer program products and computer programs in the above embodiments and will not be repeated here.

[0126] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0127] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

[0128] It should be noted that, in the description of this disclosure, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of this disclosure, unless otherwise specified, "plurality" means two or more.

[0129] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.

[0130] It should be understood that various parts of the present disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0131] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0132] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0133] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0134] Unless otherwise specified, the experimental methods in the following examples are conventional methods and are performed according to the techniques or conditions described in the literature in the field or according to the product instructions.

[0135] Unless otherwise specified, the quantitative tests in the following examples were performed three times, and the results were averaged.

[0136] Example

[0137] This example describes the implementation of a specific solution using data from a gut microbiome project published on the China National GeneBank Big Data Platform (db.cngb.org / search / project / CNP0000497 / ). The project contains 233 samples and 466 files, totaling 6.32TB of raw data, which, after gzip compression, is 2.25TB.

[0138] (1) Construction of basic reference sequence database and its index

[0139] This example uses the reference data set provided by Metaphlan3 as the source of the basic reference sequence database (github.com / biobakery / MetaPhlAn / wiki / MetaPhlAn-3.0).

[0140] By back-tracing the microbial marker genes in mpa_v30_CHOCOPhlAn_201901_marker_info.txt.bz2, the reference genome number of the marker gene source in NCBI was obtained. According to the NCBI genome number, the corresponding ftp link was obtained from the website ftp.ncbi.nih.gov / genomes / genbank / bacteria / assembly_summary.txt, thereby batch downloading the reference genome sequence. In this example, a total of 25,435 reference genomes were downloaded. The subsequences within the reference genome were merged using a Python script, and the merging rules were as follows:

[0141] a. First, determine the number of subsequences (usually contigs or scaffolds) in the genome file based on the number of ">". If there is only one, change the content after ">" to the genome number (usually starting with GCA);

[0142] b. If the number of ">" is greater than one, first add 10 "N" characters as separators at the end of each subsequence, then delete all lines containing ">" except the first one, merge the subsequences, and replace the content after the first ">" with the genome number.

[0143] After completing the internal subsequence merging of a single genome, use the cat command in the shell to merge all reference genomes into a total Fastq file to obtain the final basic reference sequence file for alignment.

[0144] (2) Sequence alignment

[0145] Based on the basic reference sequence constructed in (1), this example uses the alignment software Bwa (Heng, L. et al 2009) to randomly select 50 Fastq files of test samples for alignment. The number of reads aligned to each genome sequence is counted, and the reference genome is sorted according to the number of reads aligned.

[0146] (3) Construction of project-specific compressed reference sequences

[0147] Based on the statistical results in (2), this example selects the top 1000 reference genomes by sequence abundance to construct a project-specific compressed reference sequence. The specific selection criteria are shown in Figures 2 and 3. The final constructed Fastq file size is 1.7GB, which is only 1.6% of the basic reference sequence.

[0148] (4) Data compression test

[0149] The quality value degenerate parameters set in this embodiment are as follows:

[0150] a. The quality value mapping scheme is: 0-3 are merged into 0, 4-19 are simplified to 11, 20-29 are simplified to 23, and 30-40 are simplified to 37;

[0151] b. The judgment criteria for low-quality reads are: when the proportion of bases in a read with quality values ​​in the range of 4 to 29 is greater than or equal to 20%, the bases in the range of 4 to 29 in the read are not degenerated, and the remaining bases are degenerated according to the original rules.

[0152] After completing the construction of the compressed reference sequence, this embodiment uses the index-dependent open source compression tool genozip (genozip.Readthedocs.io / ) to perform compression tests on all samples of the project (i.e., the above-mentioned total data volume of 6.32TB of raw data). Other similar tools include GTZ (github.com / Genetalks / gtz) and LW_FQZIP (github.com / Zhuzxlab / LW-FQZip2). Figure 13 shows a specific data compression ratio distribution diagram, where GZIP compression refers to directly compressing all sample data; Genozip index-free compression refers to using the Genozip tool to compress all sample data without using the project-specific compressed reference sequence constructed in step (3) above; Genozip indexed compression refers to using the Genozip tool to compress all sample data using the project-specific compressed reference sequence constructed in step (3) above. As shown in Figure 13, using the compression scheme designed in this disclosure, the average compression ratio for 233 samples (466 files in total) is 10.46, which is 3.72 times that of gzip (2.81). Compared to the case without using a reference sequence (6.73), the average compression ratio is increased by approximately 35%. This shows that the reference index and the compression scheme based on this index proposed in the embodiments of this disclosure can achieve efficient data compression.

[0153] (5) Evaluation of the impact of mass value degeneration on species composition analysis

[0154] This example uses the Metaphlan-based species identification process (github.com / MGI-EU / MMHP_SOP_rmhost) to obtain the species composition of each sample using the Fastq files before and after mass degeneration as input. Correlation statistics are then performed on the analysis results of the mass value data before and after degeneration for each sample. The statistical method is as follows:

[0155] a. First, log-transform the species abundance of each sample to make the data conform to a normal distribution.

[0156] b. Use the pearsonr function in the Python module scipy to calculate the Pearson correlation coefficient.

[0157] Figure 14 shows a statistical plot of the Pearson correlation coefficient for the species composition of 233 samples before and after mass value degeneration. As shown in Figure 14, the correlation coefficients for species composition for all samples before and after mass value degeneration were >0.999, indicating that the lossy compression scheme employed in this embodiment has little impact on downstream species composition analysis. Thus, the reference index and the compression scheme based on this index in the disclosed embodiments achieve efficient compression without affecting the composition of the data, thus ensuring high integrity, accuracy, and fidelity of the compressed data.

[0158] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and features of different embodiments or examples, unless they are mutually inconsistent.

[0159] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0160] Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are illustrative and are not to be construed as limitations on the present disclosure. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present disclosure.

Claims

1. A method for constructing a reference sequence for metagenome data compression, comprising: Constructing a basic reference sequence database based on the sample source of the metagenomic data; Based on the basic reference sequence database, constructing an index of the basic reference sequence database; Comparing a first read sequence with the basic reference sequence database according to the index of the basic reference sequence database to obtain an alignment result, wherein the first read sequence is a read sequence of a portion of samples randomly selected from the metagenome data to be compressed; and Based on the comparison results, the sequence abundance distribution of the first read length sequence is determined, and the reference sequence for metagenomic data compression is constructed.

2. The method according to claim 1, wherein constructing a basic reference sequence database based on the sample source of the metagenomic data comprises: According to the sample source of the metagenomic data, the corresponding reference genome is obtained from the public database and summarized to obtain the basic reference sequence database.

3. The method according to claim 2, wherein constructing an index of the basic reference sequence database based on the basic reference sequence database comprises: The single reference genome in the basic reference sequence database includes a first subsequence and a second subsequence, and the first subsequence and the second subsequence are merged, and the number of the reference genome is retained to obtain a subsequence merged reference genome; Based on the subsequences and reference genome, an index of the basic reference sequence database is constructed.

4. The method of claim 3, wherein aligning the first read sequence with the base reference sequence database according to the index of the base reference sequence database comprises: Based on the index of the basic reference sequence database, aligning the first read sequence to each of the subsequence merged reference genomes; Based on the alignment of the first read sequence to the subsequence merged reference genome, the number of the reference genome to which the read sequence is aligned is recorded.

5. The method according to claim 4, wherein determining the sequence abundance distribution of the first read sequence based on the alignment result and constructing the reference sequence for metagenomic data compression comprises: Counting the number of the numbers of the first read sequence aligned to each of the reference genomes in the alignment result to obtain the sequence abundance distribution of the first read sequence; The reference genomes are sorted according to the sequence abundance, and the top X reference genomes are selected to construct the reference sequences for metagenomic data compression.

6. The method according to claim 5, wherein constructing the reference sequence for metagenomic data compression further comprises: According to the ranking, a reference genome whose sum of sequence abundance ratios is greater than Y% is selected to construct the reference sequence for metagenomic data compression.

7. The method according to any one of claims 1 to 6, further comprising: Splitting the basic reference sequence database into sub-basic reference sequence databases; Constructing indexes of the sub-reference sequence databases based on the split sub-basic reference sequence databases respectively; Based on the index of the sub-reference sequence database, the first read sequence is respectively compared with each of the sub-base reference sequence databases to obtain a second comparison result, wherein the second comparison result includes a sub-result file based on each of the sub-base reference sequence databases.

8. The method according to claim 7, further comprising: Counting the number of alignments of the first read sequence in each of the sub-result files to each of the sub-base reference sequence databases, respectively, to obtain the sequence abundance distribution of the first read sequence in each of the sub-result files; Performing a first sort on the reference genomes according to the sequence abundances in each of the sub-result files, and selecting the reference genomes with the top X sequence abundances in each of the sub-result files to construct a sub-reference sequence database; performing a second sorting of the reference genomes in the sub-reference sequence database according to the sequence abundance; The reference genomes in the top X positions in the sequence abundance distribution in the sub-reference sequence database are selected to construct the reference sequence for metagenomic data compression.

9. The method according to claim 8, wherein constructing the reference sequence for metagenomic data compression further comprises: According to the first ranking, the reference genomes whose sum of sequence abundance ratios in each sub-result file is greater than Y% are selected to construct the sub-reference sequence database, and According to the second sorting, a reference genome whose sum of sequence abundance proportions in the sub-reference sequence database is greater than Y% is selected to construct the reference sequence for metagenomic data compression.

10. The method according to any one of claims 1 to 9, further comprising: The comparison results are subjected to a first and / or second screening, wherein The first screening comprises: selecting the read sequence without insertions and / or deletions in the alignment result; The second screening includes: selecting the read length sequence below the mismatch threshold.

11. A method for compressing metagenomic data, the method comprising: The method for constructing a reference sequence for metagenomic data compression according to claim 1, constructing a reference sequence for metagenomic data compression; The second read sequence is compared with the reference sequence and the comparison result is recorded to obtain compressed data of the metagenomic data, wherein the second read sequence is the read sequence of the sample to be compressed in the metagenomic data.

12. The method according to claim 11, wherein aligning the second read sequence with the reference sequence and recording the alignment result comprises: When the number of mismatched bases between the second read sequence and the reference sequence is less than R1, recording the position of the second read sequence on the reference sequence; When the number of mismatched bases between the second read sequence and the reference sequence is greater than R1 and less than R2, recording the position of the paired base in the second read sequence on the reference sequence and recording the base information of the mismatched base; When the number of mismatched bases between the second read sequence and the reference sequence is greater than R2, the second read sequence is recorded.

13. The method according to claim 11, further comprising degenerating the quality value of the metagenomic data, wherein the degenerating comprises: Performing statistics on the base quality values ​​in the metagenomic data to obtain a distribution of the quality values ​​within M quality value ranges; The quality values ​​within the M ranges are mapped to M mapping values ​​respectively to degenerate the quality values ​​of the metagenomic data.

14. The method according to claim 13, further comprising: When the proportion of bases with quality values ​​lower than Q to all bases in the metagenomic data is lower than a set proportion N, the quality values ​​of all bases in the metagenomic data are mapped to degenerate the quality values ​​of the metagenomic data.

15. The method according to claim 14, further comprising: When the proportion of bases with quality values ​​lower than Q accounts for a proportion of all bases in the metagenomic data that is higher than or equal to a set proportion N, the quality values ​​of the bases with quality values ​​higher than Q in the metagenomic data are mapped to degenerate the quality values ​​of the metagenomic data.

16. The method according to claim 15, further comprising: When the proportion of bases with quality values ​​lower than Q to all bases in the metagenomic data is higher than or equal to a set proportion N, the original quality values ​​of the bases with quality values ​​lower than Q in the metagenomic data are retained.

17. The method according to any one of claim 16, wherein Q is a quality value corresponding to a base error probability of 0.01% to 1%.

18. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for constructing a reference sequence for metagenomic data compression according to claim 1 is implemented, the method comprising: Constructing a basic reference sequence database based on the sample source of the metagenomic data; Based on the basic reference sequence database, constructing an index of the basic reference sequence database; Comparing a first read sequence with the basic reference sequence database according to the index of the basic reference sequence database to obtain an alignment result, wherein the first read sequence is a read sequence of a randomly selected portion of the samples in the metagenome data to be compressed; and Based on the comparison results, the sequence abundance distribution of the first read length sequence is determined, and the reference sequence for metagenomic data compression is constructed. 。 19. A non-transitory computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method for constructing a reference sequence for metagenomic data compression as claimed in claim 1 is implemented.

20. A computer program product, comprising a computer program, which, when executed by a processor, implements the method for constructing a reference sequence for metagenomic data compression as claimed in claim 1.