Metagenomic sequencing data screening method, device, equipment and readable storage medium

By obtaining and utilizing the alignment relationship between similar sequence comparison information and reference human genome sequences, non-human genome sequences in metagenomic sequencing data can be accurately screened out, solving the problem of non-human sequences being mistakenly screened as human sequences in existing technologies, and improving the accuracy of data screening and the reliability of analysis.

CN119541640BActive Publication Date: 2025-09-23SANSURE BIOTECH INC +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411598250.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-09-23
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

In metagenomic sequencing data, existing technologies make it difficult to accurately screen and remove human sequences, resulting in non-human genome sequences being mistakenly screened as human sequences, affecting the accuracy and reliability of subsequent analysis.

Method used

By obtaining similar sequence alignment information of multiple non-human genome similar sequences and combining them with reference human genome sequences, potential non-human genome sequences are detected and screened in the original metagenomic sequencing data, and data screening is performed using the alignment relationship between similar sequence alignment information and reference non-human genome sequences.

Benefits of technology

The accuracy of metagenomic sequencing data screening has been improved, ensuring that non-human genome sequences are not mistakenly screened as human sequences, thereby improving the accuracy and reliability of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541640B_ABST
    Figure CN119541640B_ABST
Patent Text Reader

Abstract

The present application relates to a method, device, equipment and readable storage medium for screening metagenomic sequencing data. The method includes: obtaining similar sequence alignment information corresponding to multiple non-human genome similar sequences, wherein the non-human genome similar sequence is a non-human genome sequence in all non-human genome fragments that has a similar relationship with the reference human genome sequence; based on the similar sequence alignment information and the reference human genome sequence, detecting the original non-human genome similar sequence and the original non-human genome sequence in the original metagenomic sequencing data, and taking the detected original non-human genome similar sequence and the original non-human genome sequence as potential non-human genome sequences; based on the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence, screening the original metagenomic sequencing data to obtain a data screening result. The use of this method improves the accuracy of metagenomic sequencing data screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of metagenomic sequencing technology, and in particular to a metagenomic sequencing data screening method, apparatus, computer equipment, and computer-readable storage medium. Background Art

[0002] Metagenomic sequencing technology is a technology that directly extracts all genomic information from environmental samples for analysis. Due to its comprehensiveness, high throughput and high sensitivity, it is widely used in the study of the composition and function of complex microbial communities. In the process of processing metagenomic sequencing data of human host samples, the screening and removal of human sequences is a very important part of metagenomic sequencing analysis. The accuracy of this step directly affects the accuracy and reliability of subsequent analysis of microbial community structure and function.

[0003] At present, in the process of screening metagenomic sequencing data, it is usually based on the comparison results between the metagenomic sequencing data and the reference human genome sequence, that is, the data sequences in the metagenomic sequencing data that match the reference human genome sequence are screened out and removed. However, due to the similarity between human genome sequences and non-human genome sequences, it is difficult to determine the type of data sequence that matches the reference human genome sequence, which makes it easy to mistakenly screen non-human genome sequences as human genome sequences that need to be removed. Therefore, the current accuracy of metagenomic sequencing data screening can be further improved. Summary of the Invention

[0004] Based on this, it is necessary to provide a metagenomic sequencing data screening method, device, computer equipment and computer-readable storage medium to improve the accuracy of metagenomic sequencing data screening in response to the above technical problems.

[0005] In a first aspect, the present application provides a method for screening metagenomic sequencing data, comprising:

[0006] Obtaining similar sequence alignment information corresponding to multiple non-human genome similar sequences, wherein the non-human genome similar sequences are non-human genome sequences in all non-human genome fragments that have a similar relationship with the reference human genome sequence;

[0007] Detecting original non-human genome similar sequences and original non-human genome sequences in the original metagenomic sequencing data based on the similar sequence alignment information and the reference human genome sequence, and using the detected original non-human genome similar sequences and original non-human genome sequences together as potential non-human genome sequences;

[0008] Based on the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence, the original metagenomic sequencing data is screened to obtain a data screening result.

[0009] In one embodiment, obtaining similar sequence alignment information corresponding to multiple non-human genome similar sequences includes:

[0010] Determine a plurality of candidate non-human genome similar sequences in all the non-human genome fragments, and obtain initial sequence alignment information for each of the candidate non-human genome similar sequences, wherein the candidate non-human genome similar sequence is a non-human genome sequence aligned with the reference human genome sequence;

[0011] According to the first sequence screening condition between the candidate non-human genome similar sequence and the reference human genome sequence, the similar sequence alignment information is screened in all the initial sequence alignment information.

[0012] In one embodiment, the first sequence screening condition includes a reference sequence screening condition and an additional sequence screening condition; and screening similar sequence alignment information in all initial sequence alignment information based on the sequence screening condition between the candidate non-human genome similar sequence and the reference human genome sequence includes:

[0013] In the case where an additional human genome sequence exists among the multiple candidate non-human genome similar sequences, the reference sequence screening condition and the additional sequence screening condition are obtained, wherein the additional human genome sequence is a genome sequence that is different from the reference human genome sequence among the multiple candidate non-human genome similar sequences, the reference sequence screening condition refers to the first sequence screening condition between any of the candidate non-human genome similar sequences and the reference human genome sequence, and the additional sequence screening condition refers to the first sequence screening condition between any of the candidate non-human genome similar sequences and the additional human genome sequence;

[0014] The similar sequence alignment information is screened in all initial sequence alignment information according to the reference sequence screening condition and the additional sequence screening condition.

[0015] In one embodiment, the initial sequence alignment information includes a sequence region alignment index, a first sequence base alignment index, and a second sequence base alignment index, and the reference sequence screening condition includes one of the following:

[0016] The sequence region alignment index is greater than a preset sequence region alignment index threshold, and the first sequence base alignment index is less than or equal to a first preset sequence base alignment threshold;

[0017] The sequence region alignment index is equal to the preset sequence region alignment index threshold, the first sequence base alignment index is less than or equal to the second preset sequence base alignment threshold, and the second sequence base alignment index is greater than the second preset sequence base alignment threshold;

[0018] The second preset sequence base alignment threshold is smaller than the first preset sequence base alignment threshold.

[0019] In one embodiment, before determining a plurality of candidate non-human genome similar sequences in all non-human genome fragments, the method further comprises:

[0020] Dividing all non-human genome fragments into a plurality of non-human genome test sequences with a preset sequencing length, and generating sequence information to be aligned for each of the plurality of non-human genome test sequences;

[0021] According to the parallel comparison results between each of the sequence information to be compared and the reference sequence information corresponding to the reference human genome sequence, the candidate non-human genome similar sequence is detected in the multiple non-human genome test sequences.

[0022] In one embodiment, detecting the original non-human genome similar sequence and the original non-human genome sequence in the original metagenomic sequencing data based on the similar sequence alignment information and the reference human genome sequence includes:

[0023] According to the reference human genome sequence, a plurality of potential human genome sequences and original non-human genome sequences are detected in the original metagenomic sequencing data, wherein the potential human genome sequence is an original genome sequence aligned with the reference human genome sequence, and the original non-human genome sequence is an original genome sequence not aligned with the reference human genome sequence;

[0024] Selecting a plurality of target potential human genome sequences that meet a sequence truncation condition from the plurality of potential human genome sequences;

[0025] generating a second sequence screening condition based on the multiple target potential human genome sequences and the similar sequence comparison information;

[0026] According to the second sequence screening condition, the original non-human genome similar sequence is screened among the multiple target potential human genome sequences.

[0027] In one embodiment, the similar sequence alignment information includes a sequence region alignment index, a first sequence base alignment index, and a third sequence base alignment index, and the second sequence screening condition includes one of the following:

[0028] The sequence region alignment index is equal to a preset sequence region alignment index threshold, and a first index difference between the first sequence base alignment index and the third sequence base alignment index is less than a first preset index difference threshold;

[0029] The sequence region alignment index is greater than the preset sequence region alignment index threshold, and a second index difference between the first sequence base alignment index and the third sequence base alignment index is greater than a second preset index difference threshold;

[0030] The second preset indicator difference threshold is smaller than the first preset indicator difference threshold.

[0031] In one embodiment, the similar sequence alignment information includes a second sequence base alignment index; the raw metagenomic sequencing data is screened based on the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence to obtain a data screening result, including:

[0032] If the potential non-human genome sequence is aligned with the reference non-human genome sequence, then, if it is detected that the potential non-human genome sequence carries a sequence tag, obtaining an index and value between the second sequence base alignment index and the third sequence base alignment index of the potential non-human genome sequence;

[0033] Detecting the genome sequence type of the potential non-human genome sequence according to a magnitude relationship between the index and value and a preset index and value threshold;

[0034] The original metagenomic sequencing data is screened according to the genome sequence type to obtain a data screening result.

[0035] In one embodiment, the method further comprises:

[0036] The method further comprises:

[0037] Using the genome sequence aligned with the reference human genome sequence as the genome sequence to be tested;

[0038] Splitting the genome sequence to be detected according to a preset sequence splitting length to obtain a preset number of genome subsequences to be detected;

[0039] Extracting a target genomic subsequence to be detected from each of the genomic subsequences to be detected;

[0040] A selection step: selecting a plurality of genome sliding sequences located within a preset sequence sliding window from the target genome subsequence to be detected, wherein the genome sliding sequences are continuous;

[0041] When it is detected that the non-redundant sequence amount of the multiple genome sliding sequences is less than a preset non-redundant sequence amount threshold, determining that the region sequence to be detected composed of the multiple genome sliding sequences is a repeated region sequence;

[0042] Return to execute the selection step until all the target genome subsequences to be detected are selected.

[0043] In a second aspect, the present application also provides a metagenomic sequencing data screening device, comprising:

[0044] An acquisition module is used to obtain similar sequence alignment information corresponding to multiple non-human genome similar sequences, wherein the non-human genome similar sequences are non-human genome sequences in all non-human genome fragments that have a similar relationship with the reference human genome sequence;

[0045] a detection module, configured to detect original non-human genome similar sequences and original non-human genome sequences in the original metagenomic sequencing data based on the similar sequence alignment information and the reference human genome sequence, and to use the detected original non-human genome similar sequences and original non-human genome sequences together as potential non-human genome sequences;

[0046] The screening module is used to screen the original metagenomic sequencing data based on the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence to obtain a data screening result.

[0047] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0048] Obtain similar sequence alignment information corresponding to multiple non-human genome similar sequences, wherein the non-human genome similar sequence is a non-human genome sequence in all non-human genome fragments that has a similar relationship with a reference human genome sequence; based on the similar sequence alignment information and the reference human genome sequence, detect the original non-human genome similar sequence and the original non-human genome sequence in the original metagenomic sequencing data, and use the detected original non-human genome similar sequence and the original non-human genome sequence together as potential non-human genome sequences; based on the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence, screen the original metagenomic sequencing data to obtain a data screening result.

[0049] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0050] Obtain similar sequence alignment information corresponding to multiple non-human genome similar sequences, wherein the non-human genome similar sequence is a non-human genome sequence in all non-human genome fragments that has a similar relationship with a reference human genome sequence; based on the similar sequence alignment information and the reference human genome sequence, detect the original non-human genome similar sequence and the original non-human genome sequence in the original metagenomic sequencing data, and use the detected original non-human genome similar sequence and the original non-human genome sequence together as potential non-human genome sequences; based on the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence, screen the original metagenomic sequencing data to obtain a data screening result.

[0051] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:

[0052] Obtain similar sequence alignment information corresponding to multiple non-human genome similar sequences, wherein the non-human genome similar sequence is a non-human genome sequence in all non-human genome fragments that has a similar relationship with a reference human genome sequence; based on the similar sequence alignment information and the reference human genome sequence, detect the original non-human genome similar sequence and the original non-human genome sequence in the original metagenomic sequencing data, and use the detected original non-human genome similar sequence and the original non-human genome sequence together as potential non-human genome sequences; based on the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence, screen the original metagenomic sequencing data to obtain a data screening result.

[0053] The above-mentioned metagenomic sequencing data screening method, device, computer equipment and computer-readable storage medium first obtain the similar sequence alignment information corresponding to multiple non-human genome similar sequences aligned with the reference human genome sequence in all non-human genome fragments. Since the similar sequence alignment information is used to characterize the sequence alignment between the reference human genome sequence and each non-human genome similar sequence, and the non-human genome similar sequence refers to the genome sequence aligned with the reference human genome sequence in all non-human genome fragments, the similar sequence alignment information can be used to determine the reference human genome sequence. The non-human genome sequences in the sequence are then detected in the process of detecting the original metagenomic sequencing data based on the similar sequence alignment information and the reference human genome sequence, and the non-human genome similar sequences and the non-human genome sequences are collectively used as the non-human genome sequences, that is, the data sequences in the original metagenomic sequencing data that are not aligned with the reference human genome sequence, and the data sequences whose actual alignment conditions of the data sequences aligned with the reference human genome sequence and the similar sequence alignment information are consistent are collectively used as potential In the non-human genome sequence, that is, through similar sequence comparison information and the reference human genome sequence, the purpose of matching potential non-human genome sequences that may be non-human genome sequences in the original metagenomic sequencing data is achieved. Finally, through the comparison relationship between the potential non-human genome sequence and the reference non-human genome sequence, the original metagenomic sequencing data is screened to obtain the data screening result, that is, through the reference non-human genome sequence, the sequence type of the potential non-human genome sequence can be accurately screened, and then through the combined comparison process of the reference human genome sequence and the non-reference human genome sequence and the similar sequence comparison information, the purpose of accurately screening out the non-human gene sequence that is aligned with the reference human genome sequence in the original metagenomic sequencing data can be achieved, rather than being unable to accurately define the type of the data sequence aligned with the reference human genome sequence. Therefore, the technical defect that it is difficult to determine the type of the data sequence aligned with the reference human genome sequence due to the similarity between the human genome sequence and the non-human genome sequence is overcome, which makes it easy to mistakenly screen the non-human genome sequence as a human genome sequence that needs to be removed. Therefore, the accuracy of metagenomic sequencing data screening is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0055] Figure 1 Schematic diagram of a process for screening metagenomic sequencing data in one embodiment;

[0056] Figure 2 A schematic diagram of a process for screening metagenomic sequencing data in another embodiment;

[0057] Figure 3 This is a structural block diagram of a metagenomic sequencing data screening device in one embodiment;

[0058] Figure 4 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0060] First of all, it should be understood that metagenomic sequencing technology is of great significance in the study of microbial ecosystems in the intestine, oral cavity, skin, etc. It can comprehensively analyze all microbial DNA (Deoxyribonucleotide) in human samples. Acid, deoxyribonucleic acid), without the need to isolate or culture individual microorganisms. This technology helps to reveal the relationship between human microorganisms and health and disease, such as the association between intestinal microorganisms and obesity, diabetes, and inflammatory bowel disease. At the same time, metagenomic sequencing can help medical researchers identify pathogens and study antimicrobial resistance, and develop personalized treatment strategies, or improve disease conditions by adjusting the microbiome. In the process of processing metagenomic sequencing data of human host samples, screening and removal of human sequences is a very important part of metagenomic sequencing analysis. Its main function is to avoid contamination of human genome sequences, thereby ensuring that subsequent metagenomic analysis focuses on the genetic information of the microbiome. Since sequencing samples usually come from human tissues such as the intestine, skin, or mouth, a large number of human DNA fragments are usually mixed in the sequencing data. Removing these sequences can avoid errors, thereby improving the accuracy and scientificity of the analysis. However, the similarity between different sequences may make it difficult to accurately distinguish between human and microbial sequences, especially when there are homologous regions between human genome sequences and certain bacterial genome sequences, it is impossible to distinguish whether the data sequence to be analyzed is a human genome sequence or a bacterial genome sequence. Errors in screening here will lead to removal of human sequences. During the removal process, non-human genome sequences may be mistakenly identified as human genome sequences and mistakenly removed, thereby affecting the accuracy of the data. Even if only a small part of the data sequences to be analyzed are similar to the human genome sequences, the above situation may occur due to the following reasons: 1) Due to the polymorphic characteristics of human genome sequences, the human genome, which serves as the index of the data sequences to be analyzed, is difficult to cover all human genome sequences. Therefore, the situation where only a small part of the data sequences to be analyzed match the human genome sequences may be due to incomplete indexing; 2) The contamination of human genome sequences may also greatly affect the results of subsequent analysis. Once the missed human sequences are aligned with the non-human genome, they may be incorrectly detected; 3) Due to the huge amount of second-generation sequencing data, it may be impossible to manually verify the specific situation accurately. Therefore, when the alignment of the data sequences to be analyzed and the non-human genome sequences is unclear, only the data sequences that are aligned with the human genome sequences can be removed. Therefore, in the process of metagenomic sequencing data screening, it is very important to accurately screen the types of data sequences obtained from metagenomic sequencing data. Therefore, there is an urgent need for a method to improve the accuracy of metagenomic sequencing data screening.

[0061] In one embodiment, Figure 1As shown, a method for screening metagenomic sequencing data is provided. This embodiment takes the method applied to a terminal as an example, and the terminal includes but is not limited to a personal computer, a laptop computer, a smart phone and a tablet computer, etc. The terminal includes an acquisition module, a detection module and a screening module, wherein the acquisition module is used to obtain similar sequence comparison information corresponding to multiple non-human genome similar sequences, wherein the non-human genome similar sequence is a non-human genome sequence in all non-human genome fragments that has a similar relationship with the reference human genome sequence, the detection module is used to detect the original non-human genome similar sequence and the original non-human genome sequence in the original metagenomic sequencing data according to the similar sequence comparison information and the reference human genome sequence, and use the detected original non-human genome similar sequence and the original non-human genome sequence together as potential non-human genome sequences, and the screening module is used to screen the original metagenomic sequencing data according to the comparison relationship between the potential non-human genome sequence and the reference non-human genome sequence. Screening, obtaining data screening results, and then through the information interaction between the acquisition module, the detection module and the screening module, it is possible to detect potential non-human genome sequences in the original metagenome sequencing data through similar sequence comparison information and the reference human genome sequence, and to obtain non-human genome similar sequences that are aligned with the reference human genome sequence in the potential non-human genome sequence by reference to the non-human genome sequence. Finally, after clearly defining the non-human genome similar sequences that are aligned with the reference human genome sequence, the original metagenome sequencing data is screened, and the obtained data screening results can accurately screen the types of each data sequence in the original metagenome sequencing data, thereby solving the problem of incorrectly screening non-human genome sequences as human genome sequences that need to be removed. Therefore, the accuracy of metagenome sequencing data screening can be improved. It is understandable that this method can also be applied to servers, and can also be applied to systems including terminals and servers, and is implemented through the interaction between terminals and servers. In this embodiment, the method includes the following steps:

[0062] Step 202: Obtain similar sequence alignment information corresponding to multiple non-human genome similar sequences, wherein the non-human genome similar sequence is a non-human genome sequence in all non-human genome fragments that has a similar relationship with the reference human genome sequence.

[0063] It should be noted that the non-human genome sequence in the non-human genome fragment may be mistakenly identified as a human genome sequence due to contamination by the reference human genome sequence or the presence of homologous regions with the reference human genome sequence. In order to understand the sequence alignment between different non-human genome sequences aligned with the reference human genome sequence and the reference human genome sequence, similar sequence alignment information can be recorded in advance, thereby determining all non-human genome sequences that have a similar relationship with the reference human genome sequence before screening the metagenomic sequencing data. The non-human genome similar sequence can specifically be the non-human genome sequence in the non-human genome fragment aligned with the reference human genome sequence. It can be understood that a non-human genome fragment may include one or more non-human genome similar sequences.

[0064] It should be noted that the similar sequence comparison information is used to characterize the sequence comparison between the reference human genome sequence and multiple non-human genome similar sequences. For example, in one feasible method, the similar sequence comparison information may include the human chromosome ID of the comparison, the human chromosome position of the comparison, the alignment of the non-human similar sequence when compared to the human genome, and the mismatch of the non-human similar sequence when compared to the human genome, etc. The similar sequence comparison information can be used to understand the sequence comparison between different non-human genome similar sequences and the reference human genome sequence in all non-human genome fragments, wherein all non-human genome fragments may specifically include bacterial genome fragments, viral genome fragments, fungal genome fragments, and protist genome fragments, etc. The similar sequence comparison information can be stored in the form of a text file on the terminal where the metagenomic sequencing data screening method is deployed, or it can be stored in the form of a database on the terminal where the metagenomic sequencing data screening method is deployed. For example, in one feasible method, the similar sequence comparison information is stored as a fastq file.

[0065] As an example, step 202 includes: obtaining similar sequence alignment information corresponding to multiple non-human genome similar sequences aligned with the reference human genome sequence in all non-human genome fragments by reading a preset similar sequence alignment file, wherein the preset similar sequence alignment file can specifically be a fastq file that stores similar sequence alignment information.

[0066] In step 204, based on the similar sequence alignment information and the reference human genome sequence, the original non-human genome similar sequence and the original non-human genome sequence are detected in the original metagenomic sequencing data, and the detected original non-human genome similar sequence and the original non-human genome sequence are collectively used as potential non-human genome sequences.

[0067] It should be noted that the reference human genome sequence refers to a standardized sequence representing the structure of the human genome. By comparing the data sequence of the original metagenomic sequencing data with the reference human genome sequence, it is possible to obtain a comparison result of whether the data sequence is aligned with the reference human genome sequence. It can be understood that if the data sequence is not aligned with the reference human genome sequence, the data sequence can be directly used as a potential non-human genome sequence. If the data sequence is aligned with the reference human genome sequence, due to the similarity between the human genome sequence and the non-human genome sequence, the data sequence cannot be directly used as the original human genome. The group sequence needs to further rely on the similar sequence comparison information to determine whether the data sequence is an original non-human genome similar sequence. That is, if the similar sequence comparison information records the sequence comparison of the data sequence to the reference human genome sequence, it means that the data sequence has a similar relationship with the reference human genome sequence, and the data sequence can also be regarded as a potential non-human genome sequence. It can be understood that the original non-human genome similar sequence refers to the non-human genome similar sequence in the original metagenome sequencing data, and the original non-human genome sequence refers to the non-human genome sequence in the original metagenome sequencing data.

[0068] It can be understood that potential non-human genome sequences refer to data sequences in the original metagenomic sequencing data that may be non-human genome sequences. The original metagenomic sequencing data can be metagenomic sequencing data of human host samples. The process of aligning similar sequences of non-human genomes in the original metagenomic sequencing data based on similar sequence alignment information can be based on multiple alignment information. For example, in one feasible method, if the data sequence in the original metagenomic sequencing data is aligned with the specified position of the human chromosome recorded in the similar sequence alignment information, it is further determined whether the alignment of the data sequence and the reference human genome sequence is consistent with the alignment and mismatch recorded in the similar sequence alignment information. If the situations are consistent, it indicates that the data sequence can be used as a potential non-human genome sequence. For example, the similar sequence comparison information can be set as "59M16S:9A9C20T2T2T12", where "59M" represents 59 aligned sequence regions, "16S" represents 16 non-aligned regions, and "9A9C20T2T2T12" represents the base mutation in the aligned region, that is, the 10th base in the aligned region mutates to A, the 20th base mutates to C, and the 41st, 44th and 47th bases mutate to T. If the actual comparison between the data sequence and the reference human genome sequence matches the similar sequence comparison information, the data sequence can be used as a potential non-human genome sequence.

[0069] As an example, step 204 includes: detecting multiple data sequences in the original metagenomic sequencing data and the reference human genome sequence to obtain multiple first potential non-human genome sequences that are not aligned with the reference human genome sequence, and screening multiple second potential non-human genome sequences that match similar sequence alignment information from multiple candidate human genome sequences aligned with the reference human genome sequence, and taking the multiple first potential non-human genome sequences and the multiple second potential non-human genome sequences together as potential non-human genome sequences.

[0070] Step 206 , based on the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence, the original metagenomic sequencing data is screened to obtain a data screening result.

[0071] It should be noted that after obtaining the potential non-human genome sequence, the potential non-human genome sequence needs to be tested to ensure that the data sequence recovered by relying on similar sequence comparison information and the reference human genome sequence is consistent with the reference non-human genome sequence, so as to meet the screening expectations of screening the original metagenomic sequencing data. It can be understood that the data screening result can specifically be that all data sequences of the original metagenomic sequencing data are marked with sequence tags to indicate that different data sequences are non-human genome sequences or human genome sequences.

[0072] As an example, step 206 includes: detecting whether the genome sequence identifiers of the potential non-human genome sequence and the reference non-human genome sequence are consistent; if they are consistent, the potential non-human genome sequence is identified as a non-human genome sequence; if they are inconsistent, the potential non-human genome sequence is identified as a human genome sequence; based on the sequence identification result, the original metagenomic sequencing data is screened to obtain a data screening result.

[0073] In the above-mentioned metagenomic sequencing data screening method, first, by reading the preset similar sequence comparison file, similar sequence comparison information corresponding to multiple non-human genome similar sequences that are compared with the reference human genome sequence in all non-human genome fragments is obtained, and then multiple data sequences and the reference human genome sequence in the original metagenomic sequencing data are detected to obtain multiple first potential non-human genome sequences that are not compared with the reference human genome sequence, and multiple second potential non-human genome sequences that match the similar sequence comparison information are screened from multiple candidate human genome sequences that are compared with the reference human genome sequence, and the multiple first potential non-human genome sequences and the multiple second potential non-human genome sequences are taken together as potential non-human genome sequences, and finally, based on the comparison relationship between the potential non-human genome sequences and the reference non-human genome sequences, the original metagenomic sequencing data is screened. The group sequencing data is screened to obtain data screening results, that is, by referring to the non-human genome sequence, the sequence type of the potential non-human genome sequence can be accurately detected, and then through the combined comparison process of the reference human genome sequence and the non-reference human genome sequence and the similar sequence comparison information, the purpose of accurately screening out the non-human gene sequences that are aligned with the reference human genome sequence in the original metagenomic sequencing data can be achieved, rather than being unable to accurately define the type of the data sequence aligned with the reference human genome sequence. Therefore, the technical defect of being difficult to determine the type of the data sequence aligned with the reference human genome sequence due to the similarity between the human genome sequence and the non-human genome sequence is overcome, which makes it easy to mistakenly screen the non-human genome sequence as the human genome sequence that needs to be removed. Therefore, the accuracy of metagenomic sequencing data screening is improved.

[0074] In one embodiment, Figure 2 The following example shows how to obtain similar sequence alignment information corresponding to multiple non-human genome similar sequences, including:

[0075] Step 302: determine multiple candidate non-human genome similar sequences in all non-human genome fragments, and obtain initial sequence alignment information for each candidate non-human genome similar sequence, wherein the candidate non-human genome similar sequence is a non-human genome sequence aligned with the reference human genome sequence.

[0076] It should be noted that after all non-human genome fragments are aligned with the reference human genome sequence, the non-human genome sequence aligned with the reference human genome sequence will be used as a candidate human genome similar sequence, and the initial sequence alignment information of the candidate human genome similar sequence will be recorded, where the initial sequence text information can be stored in the form of a file.

[0077] As an example, step 302 includes: intercepting all non-human genome fragments to obtain multiple non-human genome sequences, using multiple genome sequences in the multiple non-human genome sequences that are aligned with the reference human genome sequence as multiple candidate non-human genome similar sequences, and obtaining sequence alignment information of each candidate non-human genome similar sequence.

[0078] Step 304 : Based on the first sequence screening condition between the candidate non-human genome similar sequence and the reference human genome sequence, similar sequence alignment information is screened from all initial sequence alignment information.

[0079] It should be noted that in order to store the similar sequence alignment information used for screening of metagenomic sequencing data in advance, multiple conditions can be used to screen each initial sequence alignment information, thereby obtaining similar sequence alignment information corresponding to multiple non-human genome similar sequences, and obtaining similar sequence alignment information in the simplest way that is conducive to the next step of reading. Among them, when screening each initial sequence alignment information, it is based on the first sequence screening condition. It can be understood that one or more first sequence screening conditions can be set. When multiple first sequence screening conditions are set, if any candidate non-human genome similar sequence meets all the first sequence screening conditions, the candidate non-human genome similar sequence will be used as a non-human genome similar sequence, and the initial sequence alignment information of the candidate non-human genome similar sequence will be used as the similar sequence alignment information.

[0080] As an example, step 304 includes: for multiple candidate non-human genome similar sequences, selecting a genome sequence whose sequence alignment with the reference human genome sequence satisfies the first sequence screening condition as a non-human genome similar sequence, and screening the initial sequence alignment information of the multiple non-human genome similar sequences into commonly corresponding similar sequence alignment information.

[0081] In this embodiment, similar sequence alignment information is stored on the terminal before the metagenomic sequencing data is screened. Specifically, multiple initial sequence alignment information recorded by the alignment of all non-human genome fragments and reference human genome fragments can be first obtained, and then sequence screening conditions between sequences are set in advance to screen all the initial sequence alignment information, thereby obtaining the similar sequence comparison file required for screening potential non-human genome sequences. That is, by setting the sequence screening conditions, similar sequence alignment information actually used for screening potential non-human genome sequences is screened out from multiple initial sequence alignment information, thereby reducing the storage amount of similar sequence alignment information.

[0082] In one embodiment, the first sequence screening condition includes a reference sequence screening condition and an additional sequence screening condition; based on the sequence screening condition between the candidate non-human genome similar sequence and the reference human genome sequence, similar sequence alignment information is screened in all initial sequence alignment information, including: when there is an additional human genome sequence in multiple candidate non-human genome similar sequences, obtaining the reference sequence screening condition and the additional sequence screening condition, wherein the additional human genome sequence is a genome sequence that is different from the reference human genome sequence in multiple candidate non-human genome similar sequences, the reference sequence screening condition refers to the first sequence screening condition between any candidate non-human genome similar sequence and the reference human genome sequence, and the additional sequence screening condition refers to the first sequence screening condition between any candidate non-human genome similar sequence and the additional human genome sequence; based on the reference sequence screening condition and the additional sequence screening condition, similar sequence alignment information is screened in all initial sequence alignment information.

[0083] It should be noted that for candidate non-human genome-similar sequences, there may be genome sequences that are aligned with the reference human genome sequence, and there may also be genome sequences that are aligned with additional human genome sequences. The process of screening each initial sequence alignment information varies depending on whether the candidate non-human genome-similar sequence is aligned with the additional human genome sequence. If there is an additional human genome sequence that can be aligned, this candidate non-human genome-similar sequence is required to be able to simultaneously align with the reference human genome sequence and the additional human genome sequence, and to pass the multi-condition screening of the reference human genome sequence and the additional human genome sequence, that is, the similar sequence alignment information is screened with the help of the reference sequence screening condition and the additional sequence screening condition, wherein the reference sequence screening condition refers to the sequence screening condition between any candidate non-human genome-similar sequence and the reference human genome sequence, and the additional sequence screening condition refers to the sequence screening condition between any candidate non-human genome-similar sequence and the additional human genome sequence.

[0084] As an example, when it is detected that there are additional human genome sequences different from the reference human genome sequences in multiple candidate non-human genome similar sequences, the reference sequence screening conditions corresponding to each of the multiple candidate non-human genome similar sequences are obtained, and the additional sequence screening conditions corresponding to each of the multiple candidate non-human genome similar sequences are obtained; for any candidate non-human genome similar sequence, when it is determined that the candidate non-human genome similar sequence meets the corresponding reference sequence screening conditions and the corresponding additional sequence screening conditions, the initial sequence alignment information of the candidate non-human genome is screened as similar sequence alignment information, and by integrating the screening results of multiple candidate non-human genome similar sequences, the similar sequence alignment information corresponding to more than one non-human genome similar sequences is obtained.

[0085] In one feasible method, assuming that both the initial similar sequence alignment information and the similar sequence alignment information are stored in the form of files, the genome sequences that are aligned to both and do not pass the additional sequence screening conditions can be selected from the key information of the additional human genome sequence, and recorded separately in a directional file. Then, when the initial sequence alignment information is screened based on the reference sequence screening conditions, the sites existing in the directional file can be excluded first, thereby improving the screening efficiency of the initial sequence alignment information. The specific steps are as follows: 1) prepare records that do not meet the additional sequence screening conditions and read the non-human genome sequence of the non-human genome fragment; 2) read the additional human genome sequence 1) The key information of the human genome sequence is read, and the sites that do not meet the additional sequence screening conditions and are aligned with the reference human genome sequence are recorded; 2) The records are written for subsequent key information screening of the human genome sequence; 3) The records are written for subsequent key information screening of the human genome sequence; 4) The key information of the human genome sequence is selected, and the records that do not meet the additional sequence screening conditions are read; 5) The non-human genome sequence is read for subsequent complexity testing; 6) The key information of the human genome sequence is read, and the information that does not meet the additional sequence screening conditions and meets the reference sequence screening conditions is selected; 7) The preset index file is written to obtain a similar sequence alignment file that stores similar sequence alignment information, where the similar sequence alignment information can specifically be chromosome number, chromosome site, alignment status, and mismatch status.

[0086] In this embodiment, in the presence of additional human genome sequences, the reference sequence screening conditions and additional sequence screening conditions corresponding to multiple candidate non-human genome similar sequences are used together as sequence screening conditions for screening each initial sequence comparison information, thereby ensuring that the screened similar sequence comparison information is simultaneously aligned with the reference human genome sequence and the additional human genome sequence, and simultaneously meets the reference sequence screening conditions and the additional sequence screening conditions. Therefore, the genome sequence aligned with the additional human genome sequence can also be obtained through the similar sequence comparison information index, thereby further improving the accuracy of metagenomic sequencing data screening.

[0087] In one embodiment, the initial sequence alignment information includes a sequence region alignment index, a first sequence base alignment index, and a second sequence base alignment index, and the reference sequence screening condition includes one of the following: the sequence region alignment index is greater than a preset sequence region alignment index threshold, and the first sequence base alignment index is less than or equal to the first preset sequence base alignment threshold; the sequence region alignment index is equal to the preset sequence region alignment index threshold, the first sequence base alignment index is less than or equal to the second preset sequence base alignment threshold, and the second sequence base alignment index is greater than the second preset sequence base alignment threshold; wherein the second preset sequence base alignment threshold is less than the first preset sequence base alignment threshold.

[0088] It should be noted that, in the initial sequence alignment information, the information used to construct the reference sequence screening conditions can exist in the form of indicators, wherein the initial sequence alignment information of any candidate non-human genome similar sequence includes a sequence region alignment index, a first sequence base alignment index, and a second sequence base alignment index. For example, in one practicable manner, the sequence region alignment index can specifically be the number of non-aligned regions, the first sequence base alignment index can specifically be the number of aligned bases, and the second sequence base alignment index can specifically be the number of mismatched bases in the aligned region. The reference sequence screening conditions can be set by the staff. It is understandable that, as the number of non-aligned regions decreases, the second sequence base alignment index also decreases. Therefore, when different reference sequence screening conditions are set, the second preset sequence base alignment threshold is less than the first preset sequence base alignment threshold. At the same time, when the data length of the original metagenomic sequencing data changes, the reference sequence screening conditions, sequence region alignment index, first sequence base alignment index, and second sequence base alignment index also change accordingly. When the data length is proportionally amplified, the settings of the above indicators and thresholds can also be proportionally amplified.

[0089] In one feasible manner, taking 75bp as an example, the reference sequence screening condition can be specifically that the number of non-aligned regions is equal to 1, the number of aligned bases is less than or equal to 60, and the number of mismatched bases in the aligned regions is greater than 0; the reference sequence screening condition can also be specifically that the number of non-aligned regions is greater than 1, and the number of aligned bases is less than or equal to 70.

[0090] In one embodiment, before determining multiple candidate non-human genome similar sequences in all non-human genome fragments, the method also includes: dividing all non-human genome fragments into multiple non-human genome test sequences with a preset sequencing length, and generating sequence information to be compared for each of the multiple non-human genome test sequences; based on the parallel comparison results between each sequence information to be compared and the reference sequence information corresponding to the reference human genome sequence, detecting candidate non-human genome similar sequences in the multiple non-human genome test sequences.

[0091] It should be noted that, considering the huge workload of comparing non-human genome fragments and reference human genome sequences, the process of comparing non-human genome sequences and reference human genome sequences can be separated, that is, before obtaining the initial sequence comparison information of multiple candidate non-human genome similar sequences, the non-human genome sequence and the reference human genome are compared, and the initial sequence comparison information is recorded, wherein the non-human genome test sequence refers to the non-human genome sequence for comparison test, the sequence information to be compared refers to the key information waiting for comparison, and the reference sequence information refers to the key information of the reference human genome sequence. Then, when the non-human genome test sequence is compared with the reference human genome sequence, the non-human genome test data is used as a candidate non-human genome similar sequence, and its corresponding initial sequence comparison information is recorded.

[0092] As an example, all non-human genome fragments are divided into non-human genome test sequences with a preset sequencing length, and sequence information to be compared is generated for each non-human genome test sequence; each sequence information to be compared and the reference sequence information corresponding to the reference human genome sequence are compared in parallel, and the non-human genome test sequences that are compared to the reference human genome sequence among the multiple non-human genome test sequences are used as multiple candidate non-human genome similar sequences.

[0093] In one feasible method, the initial sequence alignment information of multiple candidate non-human genome similar sequences can be stored and recorded. Since all non-human genome fragments are large, after being split and converted into fastq files, in order to optimize the storage and use of the terminal, the sequence information to be aligned of different non-human genome test sequences can be aligned in parallel, that is, while aligning a part of the genome sequence, the fastq file of the next part of the genome sequence is prepared. In order to improve the flexibility of the alignment between non-human genome sequences and human genome sequences, a function of restoring to the pre-interruption stage after interruption can be added. The specific steps can be: 1) Start multi-threading and read the non-human genome sequence , and generate a fastq file of approximately 300GB, sending information to activate the main thread to perform the alignment; 2) After receiving the information, the main thread changes the name of the fastq file and sends information to activate the multi-threaded generation of the next fastq file; 3) The fastq file is aligned with the human genome file, the results are read synchronously, and key information is recorded; 4) After the alignment is completed, the current program status is recorded for recovery after interruption; 5) Receive multi-threaded information. If the fastq file generation is complete, the next loop is executed. If all fastq files have been successfully generated, the loop process ends; 6) The recorded information is written to the index file for storage.

[0094] In one embodiment, based on similar sequence alignment information and a reference human genome sequence, original non-human genome similar sequences and original non-human genome sequences are detected in original metagenomic sequencing data, including: based on the reference human genome sequence, detecting multiple potential human genome sequences and original non-human genome sequences in the original metagenomic sequencing data, wherein the potential human genome sequence is an original genome sequence aligned with the reference human genome sequence, and the original non-human genome sequence is an original genome sequence not aligned with the reference human genome sequence; selecting multiple target potential human genome sequences that meet a sequence truncation condition from multiple potential human genome sequences; generating a second sequence screening condition based on the alignment information of the multiple target potential human genome sequences and similar sequences; and screening the multiple target potential human genome sequences for original non-human genome similar sequences based on the second sequence screening condition.

[0095] It should be noted that the purpose of relying on similar sequence alignment information and the reference human genome sequence for comparison is essentially to perform sequence recovery, that is, to recover potential non-human genome sequences in the original metagenome data. In this step, the sequencing data in the sam format that has not been aligned with the human genome sequence is converted into the fastq format, and the potential non-human genome sequences indexed by the similar sequence alignment information can be recovered. In the process of screening potential non-human genome sequences in the original metagenome sequencing data, the original metagenome sequencing data and the reference human genome sequence are first compared, wherein the original metagenome sequencing data has a preset sequence length, such as 75bp. When performing the recovery comparison, the data type and data source of the original metagenome sequencing data can be pre-determined, for example, to determine whether the original metagenome sequencing data is truncated. Sequence, and determine whether the original metagenome sequencing data ends with soft shearing or hard shearing. When the original metagenome sequencing data is a truncated sequence and ends with a shearing region, the subsequent detection process can be executed, and then by comparing with the reference human genome sequence, the potential human genome sequence and the original non-human genome sequence can be simply distinguished, wherein the potential human genome sequence refers to the genome sequence compared with the reference human genome sequence, and then the set sequence truncation conditions can be used to determine whether the potential human genome sequence meets the expected truncation requirements, and according to the set sequence comparison conditions, whether any target human genome sequence can be used as the original non-human genome similar sequence. It can be understood that the sequence comparison conditions can be one or more, and the target potential human genome sequence refers to the potential human genome sequence that meets the expected truncation requirements.

[0096] As an example, the original metagenomic sequencing data are respectively detected with the reference human genome sequence to obtain the original non-human genome sequence that is not aligned with the reference human genome sequence and the potential human genome sequence aligned with the reference human genome sequence; at least one target potential human genome sequence that meets the sequence truncation condition is selected from multiple potential human genome sequences, wherein the sequence truncation condition can be specifically: judging whether the potential human genome sequence is a truncated sequence and ends with soft clipping or hard clipping, and if it is truncated and ends with a clipping region, determining that the sequence truncation condition is met; through the matching relationship between the sequence alignment conditions constructed by each target potential human genome sequence and similar sequence alignment information, multiple target potential human genome sequences are aligned to the original non-human genome similar sequence; the original non-human genome sequence and the original non-human genome similar sequence are jointly identified as a potential non-human genome sequence.

[0097] In an practicable manner, the specific steps for aligning potential non-human genome sequences in the original metagenome can be: 1) aligning the original genome sequence and the reference human genome sequence, and using the recovery module deployed on the terminal to read the output alignment result; 2) determining whether the original genome sequence and the reference human genome sequence are aligned. If the original genome sequence is not aligned with the reference human genome sequence, converting the original genome sequence into fastq format and directly outputting the alignment result, that is, directly outputting the original genome sequence as a potential non-human genome sequence; 3) if the original genome sequence is aligned with the reference human genome sequence, determining whether the chromosome and site of the original genome sequence are located in the similar sequence alignment information. If it is determined that the chromosome and site are located in the similar sequence alignment information, then Execute the next step; 4) Determine whether the original genome sequence is a truncated sequence and meets the sequence extension conditions. If the original genome sequence is a truncated sequence and meets the sequence extension conditions, extend the original genome sequence; 5) Determine the alignment and mismatch of the original genome sequence. If the original genome sequence is consistent with the records of the specified chromosome and site in the similar sequence comparison information, execute the next step; 6) Determine whether the original genome sequence is a forward comparison. If the original genome sequence is a forward comparison, output the original genome sequence as a potential non-human genome sequence. If the original genome sequence is not a forward comparison, it means that the original genome sequence was reversely compared during the comparison to the index sequence, and then the reverse complementary genome sequence of the original genome sequence needs to be generated for output.

[0098] In this embodiment, by setting sequence truncation conditions and sequence alignment conditions, relying on the reference human genome sequence and similar sequence alignment information, original non-human genome sequences that are not aligned with the reference human genome sequence and original non-human genome similar sequences that are aligned with the reference human genome sequence are screened from multiple original genome sequences intercepted from the original metagenome sequencing data, and the two together constitute potential non-human genome sequences, thereby achieving the purpose of recovering non-human genome similar sequences in the original metagenome sequencing data, thereby laying the foundation for improving the accuracy of metagenome sequencing data screening.

[0099] In one embodiment, the similar sequence alignment information includes a sequence region alignment index, a first sequence base alignment index, and a third sequence base alignment index, and the second sequence screening condition includes one of the following: the sequence region alignment index is equal to a preset sequence region alignment index threshold, and a first index difference between the first sequence base alignment index and the third sequence base alignment index is less than a first preset index difference threshold; the sequence region alignment index is greater than the preset sequence region alignment index threshold, and a second index difference between the first sequence base alignment index and the third sequence base alignment index is a second preset index difference threshold; wherein, the second preset index difference threshold is less than the first preset index difference threshold.

[0100] It should be noted that in the process of comparing the original metagenome sequences, similar sequence comparison information can exist in the form of indicators. For example, in one feasible method, the sequence region comparison index can be specifically the number of non-aligned regions, the first sequence base comparison index can be specifically the number of aligned bases, and the third sequence base comparison index can be specifically the sequence length.

[0101] Among them, the second sequence screening condition can be specifically: the number of non-aligned regions is equal to 1, and the sequence length minus the number of aligned bases is less than 15. The second sequence screening condition can also be specifically: the number of non-aligned regions is greater than 1, and the sequence length minus the number of aligned bases is less than 5.

[0102] In one embodiment, the similar sequence alignment information includes a second sequence base alignment index; based on the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence, the original metagenomic sequencing data is screened to obtain a data screening result, including: if the potential non-human genome sequence is aligned with the reference non-human genome sequence, then when it is detected that the potential non-human genome sequence carries a sequence tag, the index and value between the second sequence base alignment index and the third sequence base alignment index of the potential non-human genome sequence are obtained; based on the size relationship between the index and value and a preset index and value threshold, the genome sequence type of the potential non-human genome sequence is detected; based on the genome sequence type, the original metagenomic sequencing data is screened to obtain a data screening result.

[0103] It should be noted that the potential non-human genome sequence meets the recovery expectations only when the potential non-human genome sequence and the reference non-human genome sequence are aligned. For example, in one feasible method, the specific steps for aligning the potential non-human genome sequence and the reference non-human genome sequence can be: 1) reading the result of aligning to the non-human genome sequence; 2) if it is aligned to the non-human genome and is not a recovered sequence, it is directly output; if it is aligned to the non-human genome and is a recovered sequence, the next step is executed; 3) calculating the sum of the number of mismatched bases and the number of unaligned bases, and outputting it if the sum is less than a specified threshold. It can be understood that the sequence tag is used to identify the sequence type of the non-human genome sequence, specifically to identify the non-human genome sequence as a recovered sequence or a non-recovered sequence.

[0104] As an example: detect whether the potential non-human genome sequence is aligned with the reference non-human genome sequence. If the potential non-human genome sequence is detected to be aligned with the reference non-human genome sequence, then when it is detected that the potential non-human genome sequence carries a sequence tag, obtain the index and value between the second sequence base alignment index and the third sequence base alignment index of the potential non-human genome sequence; take multiple candidate non-human similar sequences whose respective indexes and values ​​are less than a preset index and value threshold as non-human genome similar sequences; screen the original metagenomic sequencing data using multiple non-human genome sequences and multiple non-human genome similar sequences to obtain a data screening result.

[0105] In one embodiment, the method further includes: using a genome sequence aligned with a reference human genome sequence as a genome sequence to be detected; splitting the genome sequence to be detected according to a preset sequence splitting length to obtain a preset number of genome subsequences to be detected; extracting a target genome subsequence to be detected from each genome subsequence to be detected; a selection step: selecting multiple genome sliding sequences located within a preset sequence sliding window from the target genome subsequence to be detected, wherein the genome sliding sequences are continuous; when it is detected that the amount of non-redundant sequences of the multiple genome sliding sequences is less than a preset non-redundant sequence amount threshold, determining that the region sequence to be detected composed of the multiple genome sliding sequences is a repeated region sequence; returning to execute the selection step until all target genome subsequences to be detected are selected.

[0106] It should be noted that in the process of recording the initial sequence alignment information by aligning the non-human genome sequence and the human genome sequence and screening the potential non-human genome sequences in the original metagenomic sequencing data, in order to eliminate the interference of sequence duplication regions on the screening of metagenomic sequencing data, the genome sequence of the reference human genome sequence can be compared with the set simple duplication detection method to perform deduplication detection of simple duplication sequences, wherein the genome sequence compared with the reference human genome sequence can specifically be the data sequence of the original metagenomic sequencing data or the data sequence in all non-human genome fragments, etc. The genome subsequence to be detected is any subsequence obtained by splitting the genome sequence to be detected, and the target genome subsequence to be detected refers to the genome subsequence set to be detected after removing part of the genome subsequence to be detected.

[0107] As an example, a genome sequence compared to a reference human genome sequence is used as a genome sequence to be detected; the genome sequence to be detected is split into a preset number of genome subsequences to be detected with a preset sequence splitting length; each genome subsequence to be detected is sorted, and multiple genome subsequences to be detected except the head and tail sequences are selected as multiple target genome subsequences to be detected; a selection step: selecting multiple genome sliding sequences located within a preset sequence sliding window from the multiple target genome subsequences to be detected, wherein the multiple genome sliding sequences are continuous; the non-redundant sequence amounts of the multiple genome sliding sequences located within the preset sequence sliding window are detected, and when it is detected that the non-redundant sequence amount is less than a preset non-redundant sequence amount threshold, it is determined that the region sequence to be detected composed of the multiple genome sliding sequences is a repeated region sequence; and the selection step is returned to execute until all target genome subsequences to be detected are extracted.

[0108] In one feasible manner, the method for detecting and comparing the genome sequence with the reference human genome sequence can be specifically as follows: after splitting the genome sequence according to a fixed length, multiple 3-base sequences are obtained, and each genome subsequence to be detected is obtained after subtracting two 3-base sequences, and then the 5-base sequence is incremented in sequence. Each time, the non-redundant number of 23 3-base sequences is detected to see if it is less than 12. If it is less than 12, it is determined that a simple repeat region exists, that is, the sequence of the region to be detected composed of multiple genome sliding sequences is a repeat region sequence, wherein the 23 3-base sequences represent the genome sliding sequence.

[0109] In one feasible approach, the above-mentioned metagenomic sequencing data screening method can be used to increase the number of detected species and the number of newly detected key species. For example, assuming that there are an average of 10 newly detected species in each batch of samples and an average of 1 newly detected key species in each batch of samples, the missed detection of species can be reduced; assuming that among the 32 newly detected key species, only 3 recovered sequences are matched to the human genome through blastn, and the matching situation is poor, it shows that the threshold of similarity with human origin is operating well.

[0110] By using similar sequence alignment information and reference human genome sequences, the purpose of aligning potential non-human genome sequences that may be non-human genome sequences in the original metagenomic sequencing data is achieved. Finally, the original metagenomic sequencing data is screened by comparing the potential non-human genome sequences with the reference non-human genome sequences to obtain data screening results. That is, by using the reference non-human genome sequence, the sequence type of the potential non-human genome sequence can be accurately screened. Then, by using the combined alignment process of the reference human genome sequence and the non-reference human genome sequence and similar sequence alignment information, the purpose of accurately screening out the non-human gene sequences aligned with the reference human genome sequence in the original metagenomic sequencing data can be achieved, rather than being unable to accurately define the type of the data sequence aligned with the reference human genome sequence. Therefore, the technical defect that it is difficult to determine the type of the data sequence aligned with the reference human genome sequence due to the similarity between the human genome sequence and the non-human genome sequence is overcome, which makes it easy to mistakenly screen the non-human genome sequence as a human genome sequence that needs to be removed. Therefore, the accuracy of metagenomic sequencing data screening is improved.

[0111] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0112] Based on the same inventive concept, embodiments of the present application also provide a metagenomic sequencing data screening device for implementing the aforementioned metagenomic sequencing data screening method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more metagenomic sequencing data screening device embodiments provided below can be found in the above-described limitations of the metagenomic sequencing data screening method and will not be further elaborated here.

[0113] In an exemplary embodiment, Figure 3 As shown, a metagenomic sequencing data screening device is provided, comprising: an acquisition module 401, a detection module 402 and a screening module 403, wherein:

[0114] An acquisition module 401 is used to obtain similar sequence alignment information corresponding to multiple non-human genome similar sequences;

[0115] A detection module 402 is configured to detect original non-human genome similar sequences and original non-human genome sequences in the original metagenomic sequencing data based on the similar sequence alignment information and the reference human genome sequence, and use the detected original non-human genome similar sequences and original non-human genome sequences as potential non-human genome sequences;

[0116] The screening module 403 is used to screen the original metagenomic sequencing data according to the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence to obtain a data screening result.

[0117] In one embodiment, the acquisition module 401 is further configured to:

[0118] A plurality of candidate non-human genome similar sequences are determined in all the non-human genome fragments, and initial sequence alignment information of each candidate non-human genome similar sequence is obtained, wherein the candidate non-human genome similar sequence is a non-human genome sequence aligned with the reference human genome sequence; and the similar sequence alignment information is screened in all the initial sequence alignment information according to a first sequence screening condition between the candidate non-human genome similar sequence and the reference human genome sequence.

[0119] In one embodiment, the first sequence screening condition includes a reference sequence screening condition and an additional sequence screening condition; the acquisition module 401 is further configured to:

[0120] In the case where an additional human genome sequence exists among the multiple candidate non-human genome similar sequences, the reference sequence screening condition and the additional sequence screening condition are obtained, wherein the additional human genome sequence is a genome sequence among the multiple candidate non-human genome similar sequences that is different from the reference human genome sequence, the reference sequence screening condition refers to the first sequence screening condition between any of the candidate non-human genome similar sequences and the reference human genome sequence, and the additional sequence screening condition refers to the first sequence screening condition between any of the candidate non-human genome similar sequences and the additional human genome sequence; based on the reference sequence screening condition and the additional sequence screening condition, the similar sequence alignment information is screened in all initial sequence alignment information.

[0121] In one embodiment, the initial sequence alignment information includes a sequence region alignment index, a first sequence base alignment index, and a second sequence base alignment index, and the reference sequence screening condition includes one of the following: the sequence region alignment index is greater than a preset sequence region alignment index threshold, and the first sequence base alignment index is less than or equal to the first preset sequence base alignment threshold; the sequence region alignment index is equal to the preset sequence region alignment index threshold, the first sequence base alignment index is less than or equal to the second preset sequence base alignment threshold, and the second sequence base alignment index is greater than the second preset sequence base alignment threshold; wherein the second preset sequence base alignment threshold is less than the first preset sequence base alignment threshold.

[0122] In one embodiment, the metagenomic sequencing data screening device is further used to:

[0123] All non-human genome fragments are divided into multiple non-human genome test sequences with a preset sequencing length, and sequence information to be compared is generated for each of the multiple non-human genome test sequences; based on the parallel comparison results between each of the sequence information to be compared and the reference sequence information corresponding to the reference human genome sequence, the candidate non-human genome similar sequence is detected in the multiple non-human genome test sequences.

[0124] In one embodiment, the detection module 402 is further configured to:

[0125] According to the reference human genome sequence, multiple potential human genome sequences and original non-human genome sequences are detected in the original metagenomic sequencing data, wherein the potential human genome sequence is the original genome sequence aligned with the reference human genome sequence, and the original non-human genome sequence is the original genome sequence not aligned with the reference human genome sequence; multiple target potential human genome sequences that meet the sequence truncation condition are selected from the multiple potential human genome sequences; based on the alignment information of the multiple target potential human genome sequences and the similar sequences, a second sequence screening condition is generated; and based on the second sequence screening condition, the multiple target potential human genome sequences are screened for the original non-human genome similar sequence.

[0126] In one embodiment, the similar sequence alignment information includes a sequence region alignment index, a first sequence base alignment index, and a third sequence base alignment index, and the second sequence screening condition includes one of the following:

[0127] The sequence region comparison index is equal to a preset sequence region comparison index threshold, and a first index difference between the first sequence base comparison index and the third sequence base comparison index is less than a first preset index difference threshold; the sequence region comparison index is greater than the preset sequence region comparison index threshold, and a second index difference between the first sequence base comparison index and the third sequence base comparison index is a second preset index difference threshold; wherein, the second preset index difference threshold is less than the first preset index difference threshold.

[0128] In one embodiment, the screening module 403 is further configured to:

[0129] The similar sequence alignment information includes a second sequence base alignment index; the original metagenomic sequencing data is screened according to the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence to obtain a data screening result, including: if the potential non-human genome sequence is aligned with the reference non-human genome sequence, then when it is detected that the potential non-human genome sequence carries a sequence tag, the index and value between the second sequence base alignment index and the third sequence base alignment index of the potential non-human genome sequence are obtained; based on the size relationship between the index and value and a preset index and value threshold, the genome sequence type of the potential non-human genome sequence is detected; based on the genome sequence type, the original metagenomic sequencing data is screened to obtain a data screening result.

[0130] In one embodiment, the metagenomic sequencing data screening device is further used to:

[0131] A genome sequence that is compared with the reference human genome sequence is used as a genome sequence to be detected; the genome sequence to be detected is split according to a preset sequence splitting length to obtain a preset number of genome subsequences to be detected; a target genome subsequence to be detected is extracted from each genome subsequence to be detected; a selection step: a plurality of genome sliding sequences located within a preset sequence sliding window are selected from the target genome subsequence to be detected, wherein the genome sliding sequences are continuous; when it is detected that the non-redundant sequence amount of the plurality of genome sliding sequences is less than a preset non-redundant sequence amount threshold, it is determined that the region sequence to be detected composed of the plurality of genome sliding sequences is a repeated region sequence; and the selection step is returned to execute until all the target genome subsequences to be detected are selected.

[0132] Each module in the aforementioned metagenomic sequencing data screening device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0133] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 4 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, it realizes a method for screening metagenomic sequencing data. Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0134] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0135] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0136] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0137] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0138] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0139] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for screening metagenomic sequencing data, characterized in that: The method comprises: Obtaining similar sequence alignment information corresponding to multiple non-human genome similar sequences, wherein the non-human genome similar sequences are non-human genome sequences in all non-human genome fragments that have a similar relationship with the reference human genome sequence; Detecting original non-human genome similar sequences and original non-human genome sequences in the original metagenomic sequencing data based on the similar sequence alignment information and the reference human genome sequence, and using the detected original non-human genome similar sequences and original non-human genome sequences together as potential non-human genome sequences; Screening the original metagenomic sequencing data based on an alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence to obtain a data screening result; The method of detecting original non-human genome similar sequences and original non-human genome sequences in original metagenomic sequencing data according to the similar sequence alignment information and the reference human genome sequence includes: detecting multiple potential human genome sequences and original non-human genome sequences in the original metagenomic sequencing data according to the reference human genome sequence, wherein the potential human genome sequence is the original genome sequence aligned with the reference human genome sequence, and the original non-human genome sequence is the original genome sequence not aligned with the reference human genome sequence; selecting multiple target potential human genome sequences that meet the sequence truncation condition from the multiple potential human genome sequences; generating a second sequence screening condition according to the multiple target potential human genome sequences and the similar sequence alignment information; and screening the multiple target potential human genome sequences for the original non-human genome similar sequence according to the second sequence screening condition; The similar sequence alignment information includes a second sequence base alignment index; the original metagenomic sequencing data is screened according to the alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence to obtain a data screening result, including: if the potential non-human genome sequence is aligned with the reference non-human genome sequence, then when it is detected that the potential non-human genome sequence carries a sequence tag, the index and value between the second sequence base alignment index and the third sequence base alignment index of the potential non-human genome sequence are obtained; based on the size relationship between the index and value and a preset index and value threshold, the genome sequence type of the potential non-human genome sequence is detected; based on the genome sequence type, the original metagenomic sequencing data is screened to obtain a data screening result.

2. The method according to claim 1, characterized in that The obtaining of similar sequence alignment information corresponding to a plurality of non-human genome similar sequences includes: Determine a plurality of candidate non-human genome similar sequences in all the non-human genome fragments, and obtain initial sequence alignment information for each of the candidate non-human genome similar sequences, wherein the candidate non-human genome similar sequence is a non-human genome sequence aligned with the reference human genome sequence; According to the first sequence screening condition between the candidate non-human genome similar sequence and the reference human genome sequence, the similar sequence alignment information is screened in all the initial sequence alignment information.

3. The method according to claim 2, characterized in that The first sequence screening condition includes a reference sequence screening condition and an additional sequence screening condition; the screening of similar sequence alignment information in all initial sequence alignment information based on the sequence screening condition between the candidate non-human genome similar sequence and the reference human genome sequence includes: In the case where an additional human genome sequence exists among the multiple candidate non-human genome similar sequences, the reference sequence screening condition and the additional sequence screening condition are obtained, wherein the additional human genome sequence is a genome sequence that is different from the reference human genome sequence among the multiple candidate non-human genome similar sequences, the reference sequence screening condition refers to the first sequence screening condition between any of the candidate non-human genome similar sequences and the reference human genome sequence, and the additional sequence screening condition refers to the first sequence screening condition between any of the candidate non-human genome similar sequences and the additional human genome sequence; The similar sequence alignment information is screened in all initial sequence alignment information according to the reference sequence screening condition and the additional sequence screening condition.

4. The method according to claim 3, characterized in that The initial sequence alignment information includes a sequence region alignment index, a first sequence base alignment index, and a second sequence base alignment index, and the reference sequence screening condition includes one of the following: The sequence region alignment index is greater than a preset sequence region alignment index threshold, and the first sequence base alignment index is less than or equal to a first preset sequence base alignment threshold; The sequence region alignment index is equal to the preset sequence region alignment index threshold, the first sequence base alignment index is less than or equal to the second preset sequence base alignment threshold, and the second sequence base alignment index is greater than the second preset sequence base alignment threshold; The second preset sequence base alignment threshold is smaller than the first preset sequence base alignment threshold.

5. The method according to claim 2, characterized in that Before determining a plurality of candidate non-human genome similar sequences in all the non-human genome fragments, the method further comprises: Dividing all non-human genome fragments into a plurality of non-human genome test sequences with a preset sequencing length, and generating sequence information to be aligned for each of the plurality of non-human genome test sequences; According to the parallel comparison results between each of the sequence information to be compared and the reference sequence information corresponding to the reference human genome sequence, the candidate non-human genome similar sequence is detected in the multiple non-human genome test sequences.

6. The method according to claim 1, wherein The similar sequence comparison information includes a sequence region comparison index, a first sequence base comparison index, and a third sequence base comparison index, and the second sequence screening condition includes one of the following: The sequence region alignment index is equal to a preset sequence region alignment index threshold, and a first index difference between the first sequence base alignment index and the third sequence base alignment index is less than a first preset index difference threshold; The sequence region alignment index is greater than the preset sequence region alignment index threshold, and a second index difference between the first sequence base alignment index and the third sequence base alignment index is greater than a second preset index difference threshold; The second preset indicator difference threshold is smaller than the first preset indicator difference threshold.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Using the genome sequence aligned with the reference human genome sequence as the genome sequence to be tested; Splitting the genome sequence to be detected according to a preset sequence splitting length to obtain a preset number of genome subsequences to be detected; Extracting a target genomic subsequence to be detected from each of the genomic subsequences to be detected; A selection step: selecting a plurality of genome sliding sequences located within a preset sequence sliding window from the target genome subsequence to be detected, wherein the genome sliding sequences are continuous; When it is detected that the non-redundant sequence amount of the multiple genome sliding sequences is less than a preset non-redundant sequence amount threshold, determining that the region sequence to be detected composed of the multiple genome sliding sequences is a repeated region sequence; Return to the selection step until all target genome subsequences to be detected are selected.

8. A metagenomic sequencing data screening device, characterized in that: The device comprises: An acquisition module is used to obtain similar sequence alignment information corresponding to multiple non-human genome similar sequences, wherein the non-human genome similar sequences are non-human genome sequences in all non-human genome fragments that have a similar relationship with the reference human genome sequence; a detection module, configured to detect original non-human genome similar sequences and original non-human genome sequences in the original metagenomic sequencing data based on the similar sequence alignment information and the reference human genome sequence, and to use the detected original non-human genome similar sequences and original non-human genome sequences together as potential non-human genome sequences; A screening module, configured to screen the raw metagenomic sequencing data based on an alignment relationship between the potential non-human genome sequence and the reference non-human genome sequence to obtain a data screening result; The detection module is further used to detect multiple potential human genome sequences and original non-human genome sequences in the original metagenome sequencing data based on the reference human genome sequence, wherein the potential human genome sequence is the original genome sequence aligned with the reference human genome sequence, and the original non-human genome sequence is the original genome sequence not aligned with the reference human genome sequence; select multiple target potential human genome sequences that meet the sequence truncation condition from the multiple potential human genome sequences; generate a second sequence screening condition based on the alignment information of the multiple target potential human genome sequences and the similar sequences; and screen the multiple target potential human genome sequences for the original non-human genome similar sequence based on the second sequence screening condition; The similar sequence alignment information includes a second sequence base alignment index; the screening module is further used to obtain the index and value between the second sequence base alignment index and the third sequence base alignment index of the potential non-human genome sequence if the potential non-human genome sequence is aligned with the reference non-human genome sequence and, when it is detected that the potential non-human genome sequence carries a sequence tag; detect the genome sequence type of the potential non-human genome sequence based on the size relationship between the index and value and a preset index and value threshold; and screen the original metagenomic sequencing data based on the genome sequence type to obtain a data screening result.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method for removing human gene sequence in macro genome sequencing data

    CN108197434A

  • Nucleic acid sequence alignment method

    CN110875084A