Analysis method, device and equipment for pathogen targeted high-throughput sequencing data and medium
Through the integrated reaction of chimeric primers and pathogen samples, combined with the analysis conditions set by quality control indicators, accurate analysis of high-throughput sequencing data of pathogens is achieved, which solves the problems of sample exposure and contamination risks in traditional methods and improves the accuracy of data analysis.
Patent Information
- Application Number
- CN202510895171.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-23
Smart Images

Figure CN120690291A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of pathogen infection detection, and in particular to an analysis method, apparatus, computer equipment, and computer-readable storage medium for pathogen-targeted high-throughput sequencing data. Background Art
[0002] Targeted Next-Generation Sequencing (tNGS) is a pathogen detection technology based on a high-throughput sequencing platform. It uses customized probes or primers to capture nucleic acid fragments of pathogens of clinical concern, and combines high-throughput sequencing with bioinformatics analysis to achieve rapid identification and quantitative analysis of pathogens.
[0003] Traditional targeted metagenomic sequencing is usually based on a two-step experiment. Specifically, the sample nucleic acid is first non-specifically amplified using random primers or degenerate primers to obtain a sufficient concentration of DNA (deoxyribonucleic acid) for subsequent targeted operations. Customized probes or primers are then used to specifically capture the target pathogen sequence in the pre-amplification product. However, due to the numerous sample processing and sequencing steps in the two-step method, the chance of sample exposure increases, which in turn leads to an increased risk of exogenous contamination. Therefore, the current data analysis of pathogen-targeted high-throughput sequencing data has low analytical accuracy. Summary of the Invention
[0004] Based on this, it is necessary to provide an analysis method, device, computer equipment and computer-readable storage medium for pathogen-targeted high-throughput sequencing data to improve the analysis accuracy of data analysis on pathogen-targeted high-throughput sequencing data in response to the above technical problems.
[0005] In a first aspect, the present application provides a method for analyzing pathogen-targeted high-throughput sequencing data, comprising:
[0006] Generate targeted and enriched pathogen high-throughput sequencing data through an integrated reaction of a chimeric primer and a pathogen sample to be analyzed, wherein the chimeric primer is composed of a target sequence and an adapter sequence;
[0007] Matching corresponding analysis conditions to the pathogen species detected in the pathogen high-throughput sequencing data, wherein the analysis conditions are set based on the quality control indicators of the pathogen sample to be analyzed, and the quality control indicators are obtained by comparing all pathogen test samples with standard pathogen samples;
[0008] According to the analysis conditions, the pathogen high-throughput sequencing data is analyzed to obtain data analysis results.
[0009] In a second aspect, the present application also provides an analysis device for pathogen-targeted high-throughput sequencing data, comprising:
[0010] A generation module, configured to generate targeted enriched pathogen high-throughput sequencing data through an integrated reaction of a chimeric primer with a pathogen sample to be analyzed, wherein the chimeric primer is composed of a target sequence and an adapter sequence;
[0011] a matching module, configured to match corresponding analysis conditions to pathogen species detected by the pathogen high-throughput sequencing data, wherein the analysis conditions are set based on the quality control indicators of the pathogen sample to be analyzed, and the quality control indicators are obtained by comparing and analyzing all pathogen test samples with standard pathogen samples;
[0012] The analysis module is used to analyze the pathogen high-throughput sequencing data according to the analysis conditions to obtain data analysis results.
[0013] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0014] Through the integrated reaction of chimeric primers and pathogen samples to be analyzed, targeted enriched pathogen high-throughput sequencing data is generated, wherein the chimeric primers are composed of a targeting sequence and a linker sequence; corresponding analysis conditions are matched to the pathogen species detected in the pathogen high-throughput sequencing data, wherein the analysis conditions are set based on the quality control indicators of the pathogen samples to be analyzed, and the quality control indicators are obtained by comparing and analyzing all pathogen test samples with standard pathogen samples; according to the analysis conditions, the pathogen high-throughput sequencing data is analyzed to obtain data analysis results.
[0015] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0016] Through the integrated reaction of chimeric primers and pathogen samples to be analyzed, targeted enriched pathogen high-throughput sequencing data is generated, wherein the chimeric primers are composed of a targeting sequence and a linker sequence; corresponding analysis conditions are matched to the pathogen species detected in the pathogen high-throughput sequencing data, wherein the analysis conditions are set based on the quality control indicators of the pathogen samples to be analyzed, and the quality control indicators are obtained by comparing and analyzing all pathogen test samples with standard pathogen samples; according to the analysis conditions, the pathogen high-throughput sequencing data is analyzed to obtain data analysis results.
[0017] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0018] Through the integrated reaction of chimeric primers and pathogen samples to be analyzed, targeted enriched pathogen high-throughput sequencing data is generated, wherein the chimeric primers are composed of a targeting sequence and a linker sequence; corresponding analysis conditions are matched to the pathogen species detected in the pathogen high-throughput sequencing data, wherein the analysis conditions are set based on the quality control indicators of the pathogen samples to be analyzed, and the quality control indicators are obtained by comparing and analyzing all pathogen test samples with standard pathogen samples; according to the analysis conditions, the pathogen high-throughput sequencing data is analyzed to obtain data analysis results.
[0019] The above-mentioned analysis method, device, computer equipment and computer-readable storage medium for pathogen targeted high-throughput sequencing data first generate targeted enriched pathogen high-throughput sequencing data through an integrated reaction between chimeric primers and pathogen samples to be analyzed, wherein the chimeric primers are composed of target sequences and linker sequences, and a one-step reaction system can be constructed through the integrated reaction between the chimeric primers and the pathogen samples to be analyzed, thereby obtaining pathogen high-throughput sequencing data; and then matching corresponding analysis conditions for pathogen species detected in the pathogen high-throughput sequencing data, wherein the analysis conditions are set based on the quality control indicators of the pathogen samples to be analyzed, and the quality control indicators are obtained by comparing and analyzing all pathogen test samples with standard pathogen samples, so that the pathogen species detected in the pathogen high-throughput sequencing data can be accurately matched with the analysis conditions set based on the quality control indicators of the samples to be analyzed; finally, based on the analysis analysis conditions, analyze the pathogen high-throughput sequencing data, and obtain data analysis results; since the chimeric primer is composed of a target sequence and a linker sequence, it can simultaneously complete the nonspecific amplification and specific capture processes in the same reaction tube, and directly generate pathogen high-throughput sequencing data, that is, the sample exposure time is greatly shortened through the one-step tNGS pathogen sequencing method. At the same time, in the data analysis stage, the system can match the corresponding analysis conditions for the pathogen high-throughput sequencing data based on the multiple quality control indicators of the sample to be analyzed, thereby ensuring the accurate analysis of the pathogen high-throughput sequencing data. Therefore, it overcomes the technical defect that the two-step method has many steps in sample processing and sequencing, which increases the chance of sample exposure and thus increases the risk of exogenous contamination. Therefore, the analysis accuracy of data analysis of pathogen targeted high-throughput sequencing data is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 Schematic diagram of a process for analyzing pathogen-targeted high-throughput sequencing data in one embodiment;
[0022] Figure 2 A schematic diagram of analysis dimensions of problem areas in an analysis method for pathogen-targeted high-throughput sequencing data in another embodiment;
[0023] Figure 3 Schematic diagram of a process for analyzing pathogen-targeted high-throughput sequencing data in another embodiment;
[0024] Figure 4 A schematic diagram of a process for analyzing pathogen-targeted high-throughput sequencing data according to another embodiment of a method for analyzing pathogen-targeted high-throughput sequencing data;
[0025] Figure 5 1 is a structural block diagram of an analysis device for pathogen-targeted high-throughput sequencing data in one embodiment;
[0026] Figure 6 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0028] First, in the process of data analysis for pathogen-targeted high-throughput sequencing data, the traditional tNGS two-step method plays an important role in early pathogen detection through the staged operation of pre-amplification → targeted enrichment. The specific operation process is as follows: 1) In the pre-amplification stage, random primers or degenerate primers are used to non-specifically amplify trace clinical sample nucleic acids to solve the problem of insufficient sample nucleic acid; 2) In the targeted enrichment stage, biotin-labeled probes or specific primers are used to specifically capture the target pathogen conserved sequences in the pre-amplification products, thereby reducing the host DNA background interference; However, the traditional tNGS two-step method still has certain technical bottlenecks, such as For example, the contamination control problem is caused by the cumbersome stage-by-stage operation steps of the traditional tNGS two-step method, which leads to a high risk of sample exposure. For example, the sample nucleic acid needs to be non-specifically amplified in the pre-amplification stage. After completion, the tube needs to be opened and transferred to the targeted enrichment system, and then multiple pipetting operations such as probe hybridization and magnetic bead washing are performed. Each time the lid is opened, trace pathogens in the laboratory environment may be introduced. After the sample is contaminated, the targeted metagenomic sequencing analysis of the sample will be abnormal. Therefore, there is an urgent need for an analysis method for pathogen targeted high-throughput sequencing data that can improve the analysis accuracy of data analysis of pathogen targeted high-throughput sequencing data.
[0029] In one embodiment, Figure 1As shown, a method for analyzing pathogen-targeted high-throughput sequencing data is provided. This embodiment takes the application of this method to a terminal as an example. The terminal includes but is not limited to a personal computer and a laptop computer. The terminal includes a generation module, a matching module and an analysis module. The generation module is used to generate targeted and enriched pathogen high-throughput sequencing data through an integrated reaction of a chimeric primer and a pathogen sample to be analyzed. The chimeric primer is composed of a target sequence and a linker sequence. The matching module is used to match corresponding analysis conditions for pathogen species detected in the pathogen high-throughput sequencing data. The analysis conditions are obtained based on the quality control index setting of the pathogen sample to be analyzed. The quality control index is obtained by comparing all pathogen test samples with standard pathogen samples. The comparison analysis is obtained, and the analysis module is used to analyze the pathogen high-throughput sequencing data according to the analysis conditions to obtain data analysis results; through the information interaction between the generation module, the matching module and the analysis module, the non-specific amplification and specific capture processes can be completed synchronously in the same reaction tube, and the pathogen high-throughput sequencing data can be directly generated. In the data analysis stage, the system can match the corresponding analysis conditions for the pathogen high-throughput sequencing data based on the multiple quality control indicators of the sample to be analyzed, thereby ensuring the accurate analysis of the pathogen high-throughput sequencing data, that is, realizing one-step tNGS pathogen sequencing, so that the analysis accuracy of the data analysis of pathogen-targeted high-throughput sequencing data can be improved. It can be understood that this method can also be applied to servers, and can also be applied to systems including terminals and servers, and is implemented through the interaction between terminals and servers. In this embodiment, the method includes the following steps:
[0030] Step 102 : Generate targeted and enriched pathogen high-throughput sequencing data through an integrated reaction of a chimeric primer and a pathogen sample to be analyzed, wherein the chimeric primer is composed of a target sequence and a linker sequence.
[0031] It should be noted that the pathogen sample to be analyzed represents a sample collected from an organism or environment containing a pathogen, and which requires analysis of pathogen high-throughput sequencing data, specifically blood, sputum or tissue extracts, etc.; a chimeric primer refers to a special primer consisting of two sequences, a target sequence and a linker sequence, wherein the target sequence is used to specifically bind to a specific target nucleic acid sequence in the pathogen sample to be analyzed, thereby completing the targeted identification of the target pathogen nucleic acid, and the linker sequence may specifically include the universal sequence required by the high-throughput sequencing platform, which can make the nucleic acid fragments that have been targeted and enriched compatible with the sequencing platform; the integrated reaction of the chimeric primer and the pathogen sample to be analyzed refers to the process of completing a series of operations from nucleic acid extraction, primer binding to target sequence amplification after mixing the pathogen sample to be analyzed and the chimeric primer under certain reaction conditions. For example, through RNA reverse transcription and target sequence amplification are performed simultaneously using a one-step RT-PCR kit; pathogen high-throughput sequencing data refers to sequence data measured by a high-throughput sequencing platform; it can be understood that the one-step tNGS pathogen sequencing can be specifically performed as follows: 1) The extracted sample nucleic acid is mixed with the chimeric primer and added to a reaction system containing components such as DNA polymerase, dNTPs and buffer, and a PCR amplification reaction is performed. During the reaction, the target sequence of the chimeric primer specifically binds to the target pathogen nucleic acid, and the sample nucleic acid is used as a template and amplified under the action of DNA polymerase, so that the nucleic acid sequence of the target pathogen is enriched and amplified, and the adapter sequence is also introduced into the amplified product; 2) The amplified product is purified to remove impurities in the reaction system. According to the requirements of the high-throughput sequencing platform, the purified product is end-repaired, A-tailed and connected to the adapter to construct a library suitable for sequencing; 3) The constructed sequencing library is loaded onto a high-throughput sequencer and sequenced according to the corresponding sequencing process, thereby generating a large amount of pathogen high-throughput sequencing data.
[0032] As an example, step 102 includes: placing a chimeric primer consisting of a target sequence and a linker sequence and a pathogen sample to be analyzed in the same reaction system, performing an integrated reaction of sample nucleic acid extraction, primer binding, and nucleic acid amplification, constructing a sequencing library, and sequencing the sequencing library on a machine to obtain high-throughput sequencing data of targeted enriched pathogens.
[0033] Step 104 is to match the pathogen species detected by the pathogen high-throughput sequencing data with corresponding analysis conditions, wherein the analysis conditions are set based on the quality control indicators of the pathogen sample to be analyzed, and the quality control indicators are obtained by comparing and analyzing all pathogen test samples with standard pathogen samples.
[0034] It should be noted that, due to the innovation of the sequencing process, the one-step tNGS pathogen sequencing integrates the sample extraction, amplification and sequencing processes into a closed system. This method can greatly shorten the sample exposure time, streamline the operation links, and reduce the possibility of contamination from the root. However, due to changes in the reaction mode between the primers and the pathogen samples to be analyzed, it is necessary to adaptively adjust the analysis conditions to obtain analysis conditions that match the one-step detection characteristics in order to accurately complete the analysis of pathogen high-throughput sequencing data; analysis conditions refer to the data analysis parameters and rules set for different pathogen species; quality control indicators Refers to a quantitative indicator used to evaluate the reliability of pathogen high-throughput sequencing data, which can be specifically set based on the analysis requirements for analyzing pathogen high-throughput sequencing data; pathogen test samples refer to samples that are actually collected and subjected to high-throughput sequencing, which contain information on unknown pathogens and are the objects of analysis; standard pathogen samples refer to reference samples with known pathogen types, gene sequences, and various quality control parameters, which serve as a benchmark for evaluating test samples and are used to determine quality control indicators and analysis conditions; for example, in one feasible method, all pathogen test samples can be manually compared and divided with the standard and gold standard detection results one by one, from Relying on quality control indicators to write universal analysis conditions can increase the positive rate of pathogen high-throughput sequencing data without the occurrence of false positives. Among them, quality control indicators may include setting thresholds, primer detection conditions, strong positive interference, negative control background detection conditions and normalized sequence numbers. Among them, primer detection conditions reflect the effectiveness of primer design and reaction system, strong positive interference is used to define the false positive risk caused by high concentrations of positive nucleic acids in samples, negative control background detection conditions are used to define the situation where non-target sequences are detected in negative controls, and normalized sequence numbers represent the number of valid sequences after normalization processing, reflecting data uniformity and comparability. Data quality indicators are obtained by sequencing and analyzing standard samples such as bacteria, fungi, viruses, and mycobacteria before analyzing high-throughput sequencing data of pathogens, thereby setting multiple analysis conditions to serve the data analysis stage. The analysis conditions are set based on the quality control indicators of the pathogen samples to be analyzed, and the quality control indicators are obtained by comparing and analyzing all pathogen test samples with standard pathogen samples. It can be understood that the analysis conditions can be described as the results output by the Xinbaoyang model, that is, for high-throughput sequencing data of different pathogens, the Xinbaoyang model will output corresponding analysis conditions to provide a basis for subsequent data analysis processes.
[0035] As an example, step 104 includes: searching a preset analysis condition table for analysis conditions that match pathogen species detected by pathogen high-throughput sequencing data.
[0036] Step 106: Analyze the pathogen high-throughput sequencing data according to the analysis conditions to obtain data analysis results.
[0037] It should be noted that the data analysis results refer to various types of information about pathogens obtained by processing and interpreting pathogen high-throughput sequencing data according to specific analysis conditions; in some feasible embodiments, the data analysis results include determining whether the pathogen high-throughput sequencing data contains target pathogens, suspected pathogens, or is unreliable data, wherein the presence of target pathogens in pathogen high-throughput sequencing data can also be described as a "main report", the presence of suspected pathogens in pathogen high-throughput sequencing data can also be described as a "suspected pathogen", and the presence of pathogen high-throughput sequencing data as unreliable data can also be described as "not displayed". For example, in one feasible method, in pathogen detection, nucleic acid is extracted from the patient's throat swab sample and high-throughput sequencing is performed to obtain pathogen high-throughput sequencing data, which is matched to the set analysis conditions: the sequencing data is aligned to the reference genome using the BWA algorithm, and the minimum allele frequency is required to be ≥0.05 and the coverage is ≥20× when identifying variant sites. The pathogen high-throughput sequencing data is analyzed based on the above conditions to ultimately obtain the data analysis results.
[0038] As an example, step 106 includes: analyzing pathogen high-throughput sequencing data based on preset analysis conditions to obtain data analysis results.
[0039] In one feasible method, before analyzing the high-throughput sequencing data of pathogens, sequencing and analysis work is first carried out on standard samples such as bacteria, fungi, viruses, and mycobacteria. Based on the characteristics of the detection data, a sequence number is selected as the initial positive threshold for standard pathogen samples. The goal is to ensure that most standard pathogens can accurately report positive when the concentration reaches 1000 copies / ml, effectively avoiding the occurrence of false negative results. After experimental verification, the one-step tNGS first determined that the threshold is 300, that is, when the number of microbial sequences detected is greater than 300, it is considered that there is a target in the high-throughput sequencing data of pathogens; further, when the threshold set above can stably and easily make the test results consistent with the expected results, in order to further improve the detection sensitivity, the threshold is lowered; the goal after adjustment is to enable most pathogens to be detected at a concentration of 500 copies / ml. After experimental testing, In the experiment, the one-step tNGS method further lowered the threshold to 100, that is, when the number of detected microbial sequences was greater than 100, it was considered as the main report; however, in the actual detection process, it was found that when the number of sequences was greater than 100, the main report was reported, and some pathogens were missed or falsely reported. Therefore, in order to further improve the analysis accuracy of pathogen high-throughput sequencing data, it is necessary to systematically set the analysis condition items. Specifically, based on the problem area, the analysis condition items in the analysis conditions can be set in detail, so as to supplement and improve the positive reporting standards from multiple dimensions to ensure the accuracy and reliability of the test results. Figure 2 , Figure 2This is a schematic diagram of the analysis dimensions of the problem area. Taking into account the sample type, the thresholds set for different samples may be different. Specifically, they can be divided into direct infection site standards and non-infection sites. Among them, the threshold for non-infection sites can be appropriately lowered. In addition, the genomes of species of the same genus are highly similar, and the proportion of microorganisms in the same category, high abundance is more likely to be pathogenic bacteria.
[0040] In one embodiment, Figure 3 As shown, according to the analysis conditions, the pathogen high-throughput sequencing data is analyzed to obtain data analysis results, including:
[0041] Step 202 : Screen out sequencing data sequences that match the pathogen species from the pathogen high-throughput sequencing data, and determine the number of sequencing sequences that the sequencing data sequences map to the target pathogen gene.
[0042] It should be noted that the sequencing data sequence represents each specific nucleic acid sequence fragment in the high-throughput sequencing data of the pathogen; matching refers to comparing the sequencing data sequence with the genome sequence of the known pathogen species, judging the degree of similarity between the two through an algorithm, and determining that the sequence belongs to the pathogen species when the similarity reaches a certain standard; mapping refers to locating the screened sequencing data sequence to the specific position of the targeted pathogen gene, determining its start and end sites on the gene, and clarifying the specific region of the gene corresponding to the sequence; the targeted pathogen gene represents a base with characteristic or research value in the pre-selected pathogen; the number of sequencing sequences is the number of sequencing data sequences successfully mapped to the targeted pathogen gene. This value can reflect the sequencing coverage of the target gene region and is used to evaluate data quality and analyze information such as gene variation. It can be expressed as Reads. The number of sequencing sequences can specifically be the normalized sequence number. Specifically, the normalized sequence number refers to the number of sequences of the pathogen species contained in every 500K (0.5M) of original sequencing data sequence. It can be understood that the higher the value of the normalized sequence number, the stronger the detection signal of the pathogen.
[0043] As an example, step 202 includes: screening out sequencing data sequences that match the pathogen species from the pathogen high-throughput sequencing data, and determining the number of sequencing sequences that the sequencing data sequences are mapped to the target pathogen gene.
[0044] Step 204: extract at least one analysis condition item of the pathogen high-throughput sequencing data from the analysis conditions.
[0045] It should be noted that the analysis condition item refers to the smallest component unit that constitutes the analysis condition, which can be set based on the required analysis dimension before data analysis. For example, "the minimum allele frequency threshold of the variant site is 5%" can be used as an analysis condition item in the analysis condition. After the analysis condition is matched, at least one analysis condition item that constitutes the analysis condition can be divided based on the field identifier in the analysis condition. It can be understood that the more analysis condition items are set, the greater the analysis value of high-throughput sequencing data of pathogens.
[0046] As an example, step 204 includes: extracting a condition item identification field from the analysis condition, and dividing the analysis condition into at least one analysis condition item of the pathogen high-throughput sequencing data based on the condition item identification field.
[0047] Step 206 : Based on the correspondence between the number of sequence numbers and each analysis condition item, condition matching analysis is performed on the pathogen high-throughput sequencing data to obtain data analysis results.
[0048] It should be noted that in the process of analyzing pathogen high-throughput sequencing data, the sequencing sequence numbers and analysis condition items are compared and judged according to the corresponding relationship to complete the data analysis of the pathogen high-throughput sequencing data.
[0049] As an example, step 206 includes: performing condition matching analysis on the pathogen high-throughput sequencing data according to the correspondence between the number of sequencing sequences and each analysis condition item to obtain a data analysis result.
[0050] In one embodiment, each analysis condition item includes at least one of a historical detection frequency condition item, a sequencing sequence number condition item, a background contamination condition item, a negative quality control condition item, and a primer coverage condition item. The data analysis result includes determining whether the pathogen high-throughput sequencing data contains a target pathogen, a suspected pathogen, or is unreliable data:
[0051] Based on the correspondence between the number of sequence numbers and each analysis condition, condition matching analysis is performed on the pathogen high-throughput sequencing data to obtain data analysis results, including one of the following:
[0052] If a frequently detected species is detected in the pathogen high-throughput sequencing data, and the detection frequency of the frequently detected species in the historical detection frequency condition item is greater than the preset detection frequency threshold, then the pathogen high-throughput sequencing data is determined to be unreliable data, where the preset detection frequency threshold represents the critical value of the occurrence frequency of the frequently detected species in the historical detection data;
[0053] It should be noted that high-frequency detection species represent pathogen species that appear frequently in historical detection data, which may be common environmental bacteria or laboratory contamination bacteria (such as Escherichia coli or Staphylococcus aureus, etc.); setting the historical detection frequency condition item in the analysis conditions can compare the current pathogen high-throughput sequencing data, and the preset detection frequency threshold can be set based on data analysis requirements, such as 75%, 76% or 77%.
[0054] Specifically, in the process of analyzing high-throughput sequencing data of pathogens, for the same batch of frequently detected species, the historical detection batches are checked, and the detection frequency of the frequently detected species is compared with the preset detection frequency threshold set in advance. If the detection frequency is greater than 75%, the high-throughput sequencing data of the pathogen is determined to be unreliable data, that is, "not displayed".
[0055] If, in the sequencing sequence number condition item, the number of sequenced sequences detected is greater than or equal to the primary abundance threshold, and in the background contamination condition item, the sequence number ratio between the number of sequenced sequences detected and the number of negative control sequences is greater than or equal to the primary contamination correction threshold, or if, in the negative quality control condition item, the number of negative control sequences detected is equal to the preset negative quality control exemption threshold, then it is determined that the target pathogen exists in the pathogen high-throughput sequencing data, where the primary abundance threshold represents the minimum effective abundance standard of the target pathogen sequence number, and the primary contamination correction threshold represents the minimum standard of the sequence number ratio for excluding contamination in high-abundance data;
[0056] It should be noted that the background contamination condition item refers to the condition item for judging the degree of contamination by the ratio of the number of sequencing sequences to the number of negative control sequences. The first-level abundance threshold and the first-level contamination correction threshold can be set in advance based on the analysis requirements, among which the first-level abundance threshold can be 150, 160 or 170, etc., and the first-level contamination correction threshold can be 5, 6 or 7, etc.; the number of negative control sequences refers to the number of sequencing sequences of the negative control sample, among which the negative control sample can be water, and the number of negative control sequences can be expressed as Reads water; the preset negative quality control exemption threshold is usually set to 0 or an extremely low value (such as 1). When the number of negative control sequences is equal to this value, the quality control is considered to have passed.
[0057] Specifically, when Reads ≥ 150 and Reads / Reads water ≥ 5, or when Reads water = 0, the pathogen high-throughput sequencing data was determined to contain the target pathogen, that is, it was considered as the main report.
[0058] If, in the sequencing sequence number condition item, the number of sequencing sequences detected is greater than or equal to the secondary abundance threshold, and in the background contamination condition item, the sequence number ratio between the number of sequencing sequences detected and the number of negative control sequences is greater than or equal to the secondary contamination correction threshold, or if, in the negative quality control condition item, the number of negative control sequences detected is equal to the preset negative quality control exemption threshold, and in the primer coverage condition item, the primer coverage corresponding to the number of sequencing sequences detected is greater than or equal to the preset primer coverage threshold, then it is determined that the target pathogen exists in the pathogen high-throughput sequencing data, wherein the secondary abundance threshold represents the critical standard for distinguishing target pathogens from suspected pathogens, the secondary contamination correction threshold represents the ratio standard for confirming true infection with medium and high abundance, the preset negative quality control exemption threshold represents the critical standard for the number of sequences of negative control without contamination, and the preset primer coverage threshold represents the minimum proportion standard for effective coverage of the primer binding region;
[0059] It should be noted that when the number of sequenced sequences is greater than or equal to the secondary abundance threshold, it indicates the possible presence of pathogens of real infection; the secondary contamination correction threshold represents the minimum sequence number ratio standard for confirming real infection with medium and high abundance data; the primer coverage condition item is used to evaluate the coverage degree of primers on the target gene binding region, and the primer coverage rate is used as the analysis condition for the quality control indicator, among which the primary abundance threshold is greater than the secondary abundance threshold, the primary contamination correction threshold is greater than the secondary contamination correction threshold, the primary abundance threshold can be 50, 60 or 70, etc., the primary contamination correction threshold can be 10, 11 or 12, etc., and the preset primer coverage rate threshold can be 60%, 62% or 65%, etc.
[0060] Specifically, if Reads ≥ 50 and Reads / Reads ≥ 10, or Reads = 0 and {multi-primer detection ≥ 60% detection}, it is determined that the target pathogen exists in the pathogen high-throughput sequencing data, that is, it is considered as the main report.
[0061] If, in the sequencing sequence number condition item, the number of sequencing sequences detected is greater than or equal to the secondary abundance threshold, and in the background contamination condition item, the sequence number ratio between the number of sequencing sequences detected and the number of negative control sequences is greater than or equal to the third-level contamination correction threshold and less than the second-level contamination correction threshold, or if, in the negative quality control condition item, the number of negative control sequences detected is equal to the preset negative quality control exemption threshold, then it is determined that the pathogen high-throughput sequencing data contains suspected pathogens, wherein the third-level contamination correction threshold represents the lower limit standard of the contamination ratio for suspected infections with medium and high abundance;
[0062] It should be noted that a suspected pathogen refers to a pathogen whose sequencing data meets some infection conditions but has not yet reached the target pathogen confirmation standard, and may be a pathogen caused by real infection or contamination; the third-level contamination correction threshold is less than the first-level contamination correction threshold, and the third-level contamination correction threshold can specifically be 1, 2 or 3, etc. The threshold setting is repeated in the above embodiment, and this embodiment will not repeat it here.
[0063] Specifically, when Reads≥50 and 10>Reads / Readswater≥1, or Readswater=0, it is determined that there is a suspected pathogen in the pathogen high-throughput sequencing data, that is, it is considered to be a suspected pathogen.
[0064] If, in the sequencing sequence number condition item, the number of sequencing sequences detected is greater than or equal to the third-level abundance threshold and less than the second-level abundance threshold, and in the background contamination condition item, the sequence number ratio between the sequencing sequence number and the negative control sequence number is detected to be greater than or equal to the fourth-level contamination correction threshold, or if, in the negative quality control condition item, the number of negative control sequences detected is equal to the preset negative quality control exemption threshold, and in the primer coverage condition item, the primer coverage corresponding to the sequencing sequence number is detected to be greater than or equal to the preset primer coverage threshold, then it is determined that there are suspected pathogens in the pathogen high-throughput sequencing data, where the third-level abundance threshold represents the minimum effective abundance standard for the number of suspected pathogen sequences, and the fourth-level contamination correction threshold represents the contamination ratio filtering standard for low-abundance suspected infections.
[0065] It should be noted that the third-level abundance threshold is smaller than the second-level abundance threshold, and the fourth-level pollution correction threshold is larger than the second-level pollution correction threshold. Among them, the third-level abundance threshold can be specifically 10, 11 or 12, and the fourth-level pollution correction threshold can be specifically 25, 26 or 27.
[0066] Specifically, when 50>Reads≥10 and Reads / Reads water ≥25, or when Reads water = 0 and {multi-primer detection ≥60% detection}, it is determined that there is a suspected pathogen in the pathogen high-throughput sequencing data, that is, it is considered to be a suspected pathogen.
[0067] In this way, by dynamically matching the number of sequencing sequences with the analysis condition items, a complete logical chain is formed from data credibility filtering to pathogen classification determination. Specifically, through the historical frequency of high-frequency species and the graded correction of contamination ratios, the false positive rate can be reduced to below 3%, avoiding interference from reagents / environmental bacteria. At the same time, the sensitivity of data analysis is improved by utilizing pathogen classification determination. In addition, the dual verification of negative quality control and primer coverage can further improve the analytical reliability of high-throughput sequencing data of pathogens.
[0068] It is understandable that in order to avoid misjudgments or missed judgments caused by single-dimensional data analysis, it is necessary to build a more comprehensive data analysis system. For example, based on dimensions such as the control of the number of different species of the same genus, the determination of associations between unreliable species, contamination correction and data filtering, and the dynamic release of high-frequency detected species, specific analysis processes can be set up to further improve the data analysis results.
[0069] In one embodiment, the data analysis result further includes one of the following:
[0070] The number of species of the same genus and different species detected in the pathogen high-throughput sequencing data is less than a first preset number threshold, and the number of the first species in the target pathogen is less than or equal to a second preset number threshold, and the number of the second species in the suspected pathogen is less than or equal to a third preset number threshold, wherein the first preset number threshold represents a maximum allowable detection number critical value of the same genus and different species, the second preset number threshold represents a maximum allowable number critical value of the first species, and the third preset number threshold represents a maximum allowable number critical value of the second species;
[0071] It should be noted that heterologous species of the same genus refer to pathogens that belong to the same genus but different species in taxonomy. For example, Escherichia coli and Salmonella belong to the Enterobacteriaceae family. Since the gene sequences of heterologous species of the same genus have high homology, misjudgment may be caused by sequencing errors or primer cross-binding. Therefore, the analysis conditions can be set in a targeted manner; the first preset quantity threshold can be specifically 2 or 3 species, etc., the second preset quantity threshold and the third preset quantity threshold can be the same or different, and can be specifically 6 or 7 species, etc.
[0072] Specifically, in the data analysis results, no more than two species of the same genus are reported, the number of first species among target pathogens is less than or equal to 6, and the number of second species among suspected pathogens is less than or equal to 6, that is, the number of main reported / suspected pathogens does not exceed 6.
[0073] A first species and a second species of the same genus but different species are detected in pathogen high-throughput sequencing data, and the second species is determined to be an untrustworthy species, wherein the target sequences of the first species and the second species are similar, and a sequence number ratio between the first species and the second species is greater than or equal to a preset sequence ratio threshold, and the preset sequence ratio threshold represents a criterion for determining the sequence number ratio of the same genus but different species;
[0074] It should be noted that the target sequence similarity characterizes the degree of homology between the target gene sequences of the first species and the second species, and is usually expressed as a percentage of sequence identity; untrustworthy species refer to species that are judged to be non-real infections due to high sequence similarity with the target species, and the preset sequence ratio threshold can be 100, 110 or 120, etc.
[0075] Specifically, the number of sequencing sequences of the first species A is represented as ReadsA, and the number of sequencing sequences of the second species B is represented as ReadsB. For different species of the same genus and similar targets; when ReadsA / ReadsB ≥ 100, B is judged as an unreliable species.
[0076] Determine whether the species detected in the pathogen high-throughput sequencing data participates in the conditional matching analysis based on the relationship between the sequence number ratio between the sequence number and the negative control sequence number and the secondary contamination correction threshold;
[0077] It should be noted that the number of sequencing sequences refers to the number of valid sequences that match the target pathogen genes after high-throughput sequencing, reflecting the nucleic acid abundance of the pathogen in the sample; it is understandable that all species detected in the same batch should be reported with caution, and the secondary contamination correction threshold can be used to determine whether the species detected in the pathogen high-throughput sequencing data participate in the conditional matching analysis.
[0078] Specifically, if the Reads / Reads ratio is ≥5, the species detected in the pathogen high-throughput sequencing data is determined to participate in the conditional matching analysis, that is, it is normally reported as positive, otherwise it is filtered.
[0079] In the pathogen high-throughput sequencing data, the number of sequencing sequences of the frequently detected species detected in the current batch is greater than the number of sequencing sequences of the frequently detected species in all historical batches, and the frequently detected species is released.
[0080] It should be noted that the analysis process of the above data analysis can be specifically described as follows: 1) Data acquisition and preprocessing: obtain the current batch of pathogen high-throughput sequencing data, filter low-quality sequences, and sort the sequencing sequence numbers of each species; 2) Query the sequencing sequence numbers of the target species in all historical batches, and count its historical maximum sequence number or average sequence number; 3) Compare the sequence numbers of high-frequency detected species in the current batch with the corresponding values in the historical data; 4) If the sequence number of the current batch is greater than the sequence number in all historical batches, the release mechanism is triggered; 5) The species that meet the conditions are included in the main report and confirmed as valid detections.
[0081] Specifically, when the number of reads is greater than the number of reads corresponding to all batches, the frequently detected species are released.
[0082] Understandably, because universal pathogens are ubiquitous in the environment and the human body, they are susceptible to interference from colonization or contamination during testing. However, for specific pathogens, if a positive test is positive, contamination must be ruled out first. Consequently, analytical conditions are more stringent, and even multiple rounds of testing and verification are required to ensure the accuracy and biosafety of the results. Consequently, analytical conditions for universal and specific pathogens differ to accommodate the data analysis requirements of each species.
[0083] In one embodiment, the pathogen species include common pathogen species and special pathogen species, and each analysis condition item further includes at least one of a strong positive interference condition item and a weak positive interference condition item;
[0084] Based on the correspondence between the number of sequence numbers and each analysis condition item, condition matching analysis is performed on the pathogen high-throughput sequencing data to obtain data analysis results, which also include one of the following items;
[0085] If a universal pathogen species is detected as a strong positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the strong positive species is greater than or equal to the first strong positive interference threshold of the strong positive interference condition, the data analysis results are adjusted, where the first strong positive interference threshold represents the critical value for distinguishing the strong positive species interference judgment among universal pathogen species;
[0086] It should be noted that strongly positive species indicate pathogens with extremely high numbers of sequenced sequences, specifically, the number of sequenced sequences may account for more than 50% of the total number of sequences. Since they may be the dominant bacteria of real infection, or they may be false positives caused by contamination or experimental errors, strong positive interference condition items are set for evaluation. Among them, the strong positive interference condition item is used to evaluate whether the strong positive species interferes with the analysis condition set for the detection of other pathogens. The first strong positive interference threshold can be 300, 350 or 400, etc.
[0087] Specifically, when the number of reads is 300 times, the universal pathogen species is defined as a strong positive species. At this time, the analysis conditions need to be downgraded, for example, the "main report" is downgraded to "suspected pathogen".
[0088] If a universal pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequenced sequences of the weak positive species is less than or equal to the first weak positive interference threshold of the weak positive interference condition item, then the data analysis result is retained, wherein the first weak positive interference threshold represents the primary critical value for distinguishing weak positive interference determinations in universal pathogen species;
[0089] It should be noted that weakly positive species refer to pathogens with a low number of sequenced sequences, for example, if the number of sequenced sequences accounts for less than 10% of the total number of sequences, they may be non-dominant bacteria caused by low-abundance infection, environmental pollution or experimental errors, and then targeted weak positive interference condition items are set for data analysis. Among them, the weak positive interference condition item is a set of analysis conditions used to evaluate whether the weakly positive species interferes with the detection of other pathogens. The first weak positive interference threshold can be 500, 600 or 700, etc.
[0090] Specifically, the number of sequencing sequences of weakly positive species can be expressed as weak reads. When weak reads ≥ 500, no degradation is required.
[0091] If a universal pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the weak positive species is less than or equal to the second weak positive interference threshold of the weak positive interference condition item, then the pathogen high-throughput sequencing data is determined to be unreliable data, wherein the second weak positive interference threshold represents the secondary critical value for distinguishing weak positive interference judgments in universal pathogen species;
[0092] It should be noted that the second weak positive interference threshold is smaller than the first weak positive interference threshold, and the second weak positive interference threshold may specifically be 5, 6, or 7.
[0093] Specifically, the number of sequencing sequences of weakly positive species can be expressed as weak Reads. When weak Reads ≤ 10, it is downgraded by two levels, for example, from the main report to not displayed.
[0094] If a special pathogen species is detected as a strong positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the strong positive species is greater than or equal to the second strong positive interference threshold of the strong positive interference condition item, the data analysis results are adjusted, where the second strong positive interference threshold represents the critical value for distinguishing the interference judgment of the strong positive species among the special pathogen species;
[0095] It should be noted that the second strong positive interference threshold is different from the first strong positive interference threshold. The second strong positive interference threshold may be 500, 600 or 700, etc.
[0096] Specifically, when the number of reads is 700 times, the special pathogen species is defined as a strong positive species. At this time, the analysis conditions need to be downgraded, for example, the "main report" is downgraded to "suspected pathogen".
[0097] If a special pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the weak positive species is less than or equal to the third weak positive interference threshold of the weak positive interference condition item, then the data analysis result is retained, wherein the third weak positive interference threshold represents the first-level critical value for distinguishing the weak positive interference judgment in the special pathogen species;
[0098] It should be noted that the third weak positive interference threshold and the first weak positive interference threshold may be the same or different.
[0099] Specifically, the number of sequencing sequences of weakly positive species can be expressed as weak reads. When weak reads ≥ 500, no degradation is required.
[0100] If a special pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the weak positive species is less than or equal to the fourth weak positive interference threshold of the weak positive interference condition item, the pathogen high-throughput sequencing data is determined to be unreliable data, where the fourth weak positive interference threshold represents the secondary critical value for distinguishing the weak positive interference judgment in the special pathogen species.
[0101] Specifically, if there are strong positive results ≥2000 times in the same batch, they will not be displayed (except for 100% target detection and ≥2 primer pairs, and weak reads ≥500, except for all targets detected).
[0102] In this way, pathogens whose detection results are inconsistent with the expected results are regarded as special pathogens and require additional conditions to be written. Similarly, the analysis conditions are set according to quality control indicators such as thresholds, primer detection, strong positive interference, and negative control background detection. This can further lay the foundation for improving the accuracy of data analysis of pathogen-targeted high-throughput sequencing data.
[0103] It is understandable that for common pathogen species, targeted analysis conditions can be set based on actual conditions.
[0104] In one embodiment, the pathogen species comprises a universal pathogen species; and the analysis conditions further comprise one of the following:
[0105] If the universal pathogen species belongs to the first type of pathogen, the analysis conditions are set as follows: the number of chimeric primers used is greater than or equal to a first preset primer quantity threshold, and the sequence substitution rate between each set of primers and the human metapneumovirus genome is greater than or equal to a preset primer matching threshold, wherein the first preset primer quantity threshold represents the standard number of chimeric primers required for the detection of the first type of pathogen;
[0106] If the universal pathogen species belongs to the second type of pathogen, the analysis conditions are set as follows: the number of chimeric primers used is greater than or equal to a second preset primer quantity threshold, and the sequence substitution rate between each set of primers and the human metapneumovirus genome is greater than or equal to a preset primer matching threshold, wherein the second preset primer quantity threshold represents the standard number of chimeric primers required for the detection of the second type of pathogen;
[0107] If the universal pathogen species belongs to the third type of pathogen, the analysis conditions are set as follows: the number of chimeric primers used is equal to the second preset primer number threshold, and the sequence substitution rate between each set of primers and the target pathogen genome is greater than or equal to the preset primer matching threshold, wherein the third preset primer number threshold represents the standard number of chimeric primers required for the detection of the third type of pathogen;
[0108] If the universal pathogen species is a DNA virus, the analysis condition is set as follows: the sequence signal intensity of the DNA virus in the same batch reaches a preset strong positive signal threshold, where the preset strong positive signal threshold represents the critical value of the sequence signal intensity for DNA virus detection;
[0109] If the universal pathogen species is Mycobacterium tuberculosis, the analysis conditions are set as follows: the sequence signal intensity is lower than the first signal intensity upper limit value and higher than the second signal intensity lower limit value, and no sequence signal of the pathogen is detected in the environmental control sample, wherein the first signal intensity upper limit value represents the upper limit standard for the suspected determination of Mycobacterium tuberculosis, and the second signal intensity lower limit value represents the lower limit standard for the suspected determination of Mycobacterium tuberculosis;
[0110] If the universal pathogen species is Mycobacterium fortuitum, the analysis conditions are set as follows: among the chimeric primers used, the ratio of the number of primers that actually detect the target sequence to the total number of primers is greater than or equal to a first preset primer detection threshold, and the sequence substitution rate between each set of detection primers and the Mycobacterium fortuitum genome is greater than or equal to a preset primer matching threshold, wherein the preset primer matching threshold represents the minimum standard for the sequence substitution rate between the chimeric primers and the target genome;
[0111] If the universal pathogen species is Mycobacterium fortuitum, the analysis conditions are set as follows: among the chimeric primers used, the number of primers that actually detect the target sequence is greater than or equal to a fourth preset primer quantity threshold, and the sequence replacement rate between each detection primer and the Mycobacterium fortuitum genome is greater than or equal to a preset primer matching threshold, wherein the fourth preset primer quantity threshold represents the minimum primer quantity standard for actually detecting the target sequence in the chimeric primers when detecting Mycobacterium fortuitum.
[0112] It should be noted that the first type of pathogens, the second type of pathogens and the third type of pathogens are different categories of pathogens divided based on the genetic characteristics of the pathogens and analysis requirements. Specifically, the first type of pathogens may be human metapneumovirus, the second type of pathogens may include rhinovirus A, Mycobacterium tuberculosis and Mycobacterium avium, etc., and the third type of pathogens may include Escherichia coli and Staphylococcus aureus, etc.; the first preset primer quantity threshold, the second preset primer quantity threshold, the third threshold primer quantity threshold and the fourth preset primer quantity threshold refer to the minimum number standards of chimeric primers required for the detection of different types of pathogens, which may be the same or different.
[0113] Specifically, 1) Human metapneumovirus multi-primer situation ≥ 2 and sequence replacement rate ≥ 60%; 2) Rhinovirus A, Mycobacterium tuberculosis or Mycobacterium avium multi-primer situation ≥ 1 and sequence replacement rate ≥ 60%; 3) Escherichia coli or Staphylococcus aureus multi-primer situation = 100% replacement ≥ 60%; 4) DNA virus in the same batch has a strong positive detection (200 times Reads is defined as a strong positive); 5) Mycobacterium tuberculosis 10>Reads ≥ 3 and Re Ads water = 0 is considered a suspected pathogen; 6) Mycobacterium fortuitum (14. Multi-primer detection rate ≥50%, and replacement multi-primer detection rate ≥60%; 7) Mycobacterium fortuitum (15. Multi-primer detection rate ≥1, and replacement multi-primer detection rate ≥60%; 8) Aspergillus fumigatus (8. When there is a strong positive detection in the same batch (800 times reads is defined as a strong positive), the above conditions need to be downgraded. When the number of weak reads is ≥500, no downgrade is required).
[0114] In one embodiment, the pathogen species comprises a drug-resistant pathogen species; and the analytical conditions further comprise one of the following:
[0115] Detection of an associated pathogen of a drug-resistant pathogen species, and the associated pathogen is the target pathogen;
[0116] The number of sequencing sequences of the drug-resistant pathogen species detected is greater than a sequencing sequence number threshold, wherein the sequencing sequence number threshold represents the minimum effective standard for the number of sequencing sequences of the drug-resistant pathogen;
[0117] The sequence variation frequency of the drug-resistant site of the drug-resistant pathogen species detected is greater than or equal to a preset mutation frequency threshold, wherein the preset mutation frequency threshold represents the critical standard of the sequence variation frequency of the drug-resistant site.
[0118] It should be noted that due to the significant differences in the drug resistance mechanisms of different pathogens, for example, bacteria acquire drug resistance through gene mutations, while viruses acquire resistance through envelope protein mutations. Therefore, analysis conditions need to be set for specific resistance sites and genes; drug-resistant pathogen species refer to pathogen types that are resistant to at least one antimicrobial drug, and can acquire drug resistance through gene mutations or acquisition of drug-resistant genes; associated pathogens characterize pathogens that are directly associated with drug-resistant pathogen species in the process of evolution, symbiosis or transmission; the sequencing sequence number threshold is used to exclude false negatives or random errors caused by insufficient sequencing depth; resistance sites refer to specific nucleotide positions in the pathogen genome that are directly related to drug resistance.
[0119] Specifically, 1) the associated pathogen corresponding to the detected drug-pathogen species is located in the main report; 2) the number of drug-resistant gene reads is ≥50, and the rest are not displayed; 3) the mutation frequency is ≥10% and reported normally.
[0120] In this way, by confirming the associated pathogens, verifying the sequencing sequence numbers, and analyzing the frequency of drug-resistant site mutations, we can eliminate interference from non-target pathogens, reduce the false positive / false negative rate, and ensure the reliability of drug-resistant test results. This further lays the foundation for improving the accuracy of data analysis of pathogen-targeted high-throughput sequencing data.
[0121] In one practicable manner, referring to Figure 4 , Figure 4 This is a schematic diagram of the analysis process of pathogen-targeted high-throughput sequencing data, which goes through data preparation, data quality control, pathogen database comparison, close relatives, off-target sequence analysis, pathogen positive reporting, supplementary positive reporting, pathogen downgrading and drug-resistant gene positive reporting, among which pathogen positive reporting, supplementary positive reporting, pathogen downgrading and drug-resistant gene positive reporting are sub-processes involved in the data analysis process. It can be understood that the filtered pathogenic microorganisms are released from the filter list and re-participate in the positive reporting. According to different types of pathogen species and different data analysis requirements, the positive reporting threshold is adaptively set to construct the corresponding quality control indicators, and then different analysis condition items can be set based on the quality control indicators. Finally, multiple analysis conditions are integrated and stored in the preset analysis condition table before data analysis, so that the pathogen species detected in the pathogen high-throughput sequencing data can be matched with the corresponding analysis conditions in the subsequent data analysis stage, and the pathogen high-throughput sequencing data can be analyzed according to the analysis conditions to obtain data analysis results.
[0122] Since the chimeric primer is composed of a target sequence and a linker sequence, it can simultaneously complete nonspecific amplification and specific capture processes in the same reaction tube, and directly generate pathogen high-throughput sequencing data. That is, the sample exposure time is greatly shortened through the one-step tNGS pathogen sequencing method. At the same time, in the data analysis stage, the system can match the corresponding analysis conditions for the pathogen high-throughput sequencing data based on the multiple quality control indicators of the sample to be analyzed, thereby ensuring the accurate analysis of the pathogen high-throughput sequencing data. Therefore, it overcomes the technical defect that the two-step method has many steps in sample processing and sequencing, which increases the chance of sample exposure and thus increases the risk of exogenous contamination. Therefore, the analysis accuracy of pathogen targeted high-throughput sequencing data is improved.
[0123] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0124] Based on the same inventive concept, the present application also provides an analysis device for pathogen-targeted high-throughput sequencing data for implementing the aforementioned analysis method for pathogen-targeted high-throughput sequencing data. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of the embodiments of one or more analysis devices for testing pathogen-targeted high-throughput sequencing data provided below can be found in the above-mentioned limitations of the analysis method for pathogen-targeted high-throughput sequencing data, and will not be repeated here.
[0125] In an exemplary embodiment, Figure 5 As shown, a device for analyzing pathogen-targeted high-throughput sequencing data is provided, comprising: a generation module 301, a matching module 302, and an analysis module 303, wherein:
[0126] A generation module 301 is used to generate targeted enriched pathogen high-throughput sequencing data through an integrated reaction between a chimeric primer and a pathogen sample to be analyzed, wherein the chimeric primer is composed of a target sequence and an adapter sequence;
[0127] Matching module 302, for matching corresponding analysis conditions to pathogen species detected in pathogen high-throughput sequencing data, wherein the analysis conditions are set based on the quality control indicators of the pathogen samples to be analyzed, and the quality control indicators are obtained by comparing and analyzing all pathogen test samples with standard pathogen samples;
[0128] The analysis module 303 is used to analyze the pathogen high-throughput sequencing data according to the analysis conditions to obtain data analysis results.
[0129] In one embodiment, the analysis module 303 is further configured to:
[0130] Screening data sequences that match pathogen species from pathogen high-throughput sequencing data, and determine the number of sequencing sequences that map the sequencing data sequences to the targeted pathogen genes;
[0131] At least one analysis condition item of pathogen high-throughput sequencing data is extracted from the analysis conditions; and condition matching analysis is performed on the pathogen high-throughput sequencing data according to the correspondence between the sequencing sequence number and each analysis condition item to obtain a data analysis result.
[0132] In one embodiment, each analysis condition item includes at least one of a historical detection frequency condition item, a sequencing sequence number condition item, a background contamination condition item, a negative quality control condition item, and a primer coverage condition item; the data analysis result includes determining whether the pathogen high-throughput sequencing data contains a target pathogen, a suspected pathogen, or is unreliable data; the analysis module 303 is further configured to:
[0133] If a frequently detected species is detected in the pathogen high-throughput sequencing data, and the detection frequency of the frequently detected species in the historical detection frequency condition item is greater than the preset detection frequency threshold, then the pathogen high-throughput sequencing data is determined to be unreliable data, where the preset detection frequency threshold represents the critical value of the occurrence frequency of the frequently detected species in the historical detection data;
[0134] If, in the sequencing sequence number condition item, the number of sequenced sequences detected is greater than or equal to the primary abundance threshold, and in the background contamination condition item, the sequence number ratio between the number of sequenced sequences detected and the number of negative control sequences is greater than or equal to the primary contamination correction threshold, or if, in the negative quality control condition item, the number of negative control sequences detected is equal to the preset negative quality control exemption threshold, then it is determined that the target pathogen exists in the pathogen high-throughput sequencing data, where the primary abundance threshold represents the minimum effective abundance standard of the target pathogen sequence number, and the primary contamination correction threshold represents the minimum standard of the sequence number ratio for excluding contamination in high-abundance data;
[0135] If, in the sequencing sequence number condition item, the number of sequencing sequences detected is greater than or equal to the secondary abundance threshold, and in the background contamination condition item, the sequence number ratio between the number of sequencing sequences detected and the number of negative control sequences is greater than or equal to the secondary contamination correction threshold, or if, in the negative quality control condition item, the number of negative control sequences detected is equal to the preset negative quality control exemption threshold, and in the primer coverage condition item, the primer coverage corresponding to the number of sequencing sequences detected is greater than or equal to the preset primer coverage threshold, then it is determined that the target pathogen exists in the pathogen high-throughput sequencing data, wherein the secondary abundance threshold represents the critical standard for distinguishing target pathogens from suspected pathogens, the secondary contamination correction threshold represents the ratio standard for confirming true infection with medium and high abundance, the preset negative quality control exemption threshold represents the critical standard for the number of sequences of negative control without contamination, and the preset primer coverage threshold represents the minimum proportion standard for effective coverage of the primer binding region;
[0136] If, in the sequencing sequence number condition item, the number of sequencing sequences detected is greater than or equal to the secondary abundance threshold, and in the background contamination condition item, the sequence number ratio between the number of sequencing sequences detected and the number of negative control sequences is greater than or equal to the third-level contamination correction threshold and less than the second-level contamination correction threshold, or if, in the negative quality control condition item, the number of negative control sequences detected is equal to the preset negative quality control exemption threshold, then it is determined that the pathogen high-throughput sequencing data contains suspected pathogens, wherein the third-level contamination correction threshold represents the lower limit standard of the contamination ratio for suspected infections with medium and high abundance;
[0137] If, in the sequencing sequence number condition item, the number of sequencing sequences detected is greater than or equal to the third-level abundance threshold and less than the second-level abundance threshold, and in the background contamination condition item, the sequence number ratio between the sequencing sequence number and the negative control sequence number is detected to be greater than or equal to the fourth-level contamination correction threshold, or if, in the negative quality control condition item, the number of negative control sequences detected is equal to the preset negative quality control exemption threshold, and in the primer coverage condition item, the primer coverage corresponding to the sequencing sequence number is detected to be greater than or equal to the preset primer coverage threshold, then it is determined that there are suspected pathogens in the pathogen high-throughput sequencing data, where the third-level abundance threshold represents the minimum effective abundance standard for the number of suspected pathogen sequences, and the fourth-level contamination correction threshold represents the contamination ratio filtering standard for low-abundance suspected infections.
[0138] In one embodiment, the data analysis result further includes one of the following:
[0139] The number of species of the same genus and different species detected in the pathogen high-throughput sequencing data is less than a first preset number threshold, and the number of the first species in the target pathogen is less than or equal to a second preset number threshold, and the number of the second species in the suspected pathogen is less than or equal to a third preset number threshold, wherein the first preset number threshold represents a maximum allowable detection number critical value of the same genus and different species, the second preset number threshold represents a maximum allowable number critical value of the first species, and the third preset number threshold represents a maximum allowable number critical value of the second species;
[0140] A first species and a second species of the same genus but different species are detected in pathogen high-throughput sequencing data, and the second species is determined to be an untrustworthy species, wherein the target sequences of the first species and the second species are similar, and a sequence number ratio between the first species and the second species is greater than or equal to a preset sequence ratio threshold, and the preset sequence ratio threshold represents a criterion for determining the sequence number ratio of the same genus but different species;
[0141] Determine whether the species detected in the pathogen high-throughput sequencing data participates in the conditional matching analysis based on the relationship between the sequence number ratio between the sequence number and the negative control sequence number and the secondary contamination correction threshold;
[0142] In the pathogen high-throughput sequencing data, the number of sequencing sequences of the frequently detected species detected in the current batch is greater than the number of sequencing sequences of the frequently detected species in all historical batches, and the frequently detected species is released.
[0143] In one embodiment, the pathogen species include common pathogen species and specific pathogen species, and each analysis condition item further includes at least one of a strong positive interference condition item and a weak positive interference condition item; the analysis module 303 is further configured to:
[0144] If a universal pathogen species is detected as a strong positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the strong positive species is greater than or equal to the first strong positive interference threshold of the strong positive interference condition, the data analysis results are adjusted, where the first strong positive interference threshold represents the critical value for distinguishing the strong positive species interference judgment among universal pathogen species;
[0145] If a universal pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequenced sequences of the weak positive species is less than or equal to the first weak positive interference threshold of the weak positive interference condition item, then the data analysis result is retained, wherein the first weak positive interference threshold represents the primary critical value for distinguishing weak positive interference determinations in universal pathogen species;
[0146] If a universal pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the weak positive species is less than or equal to the second weak positive interference threshold of the weak positive interference condition item, then the pathogen high-throughput sequencing data is determined to be unreliable data, wherein the second weak positive interference threshold represents the secondary critical value for distinguishing weak positive interference judgments in universal pathogen species;
[0147] If a special pathogen species is detected as a strong positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the strong positive species is greater than or equal to the second strong positive interference threshold of the strong positive interference condition item, the data analysis results are adjusted, where the second strong positive interference threshold represents the critical value for distinguishing the interference judgment of the strong positive species among the special pathogen species;
[0148] If a special pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the weak positive species is less than or equal to the third weak positive interference threshold of the weak positive interference condition item, then the data analysis result is retained, wherein the third weak positive interference threshold represents the first-level critical value for distinguishing the weak positive interference judgment in the special pathogen species;
[0149] If a special pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the weak positive species is less than or equal to the fourth weak positive interference threshold of the weak positive interference condition item, the pathogen high-throughput sequencing data is determined to be unreliable data, where the fourth weak positive interference threshold represents the secondary critical value for distinguishing the weak positive interference judgment in the special pathogen species.
[0150] In one embodiment, the pathogen species includes a universal pathogen species; and the analysis conditions further include one of the following:
[0151] If the universal pathogen species belongs to the first type of pathogen, the analysis conditions are set as follows: the number of chimeric primers used is greater than or equal to a first preset primer quantity threshold, and the sequence substitution rate between each set of primers and the human metapneumovirus genome is greater than or equal to a preset primer matching threshold, wherein the first preset primer quantity threshold represents the standard number of chimeric primers required for the detection of the first type of pathogen;
[0152] If the universal pathogen species belongs to the second type of pathogen, the analysis conditions are set as follows: the number of chimeric primers used is greater than or equal to a second preset primer quantity threshold, and the sequence substitution rate between each set of primers and the human metapneumovirus genome is greater than or equal to a preset primer matching threshold, wherein the second preset primer quantity threshold represents the standard number of chimeric primers required for the detection of the second type of pathogen;
[0153] If the universal pathogen species belongs to the third type of pathogen, the analysis conditions are set as follows: the number of chimeric primers used is equal to the second preset primer number threshold, and the sequence substitution rate between each set of primers and the target pathogen genome is greater than or equal to the preset primer matching threshold, wherein the third preset primer number threshold represents the standard number of chimeric primers required for the detection of the third type of pathogen;
[0154] If the universal pathogen species is a DNA virus, the analysis condition is set as follows: the sequence signal intensity of the DNA virus in the same batch reaches a preset strong positive signal threshold, where the preset strong positive signal threshold represents the critical value of the sequence signal intensity for DNA virus detection;
[0155] If the universal pathogen species is Mycobacterium tuberculosis, the analysis conditions are set as follows: the sequence signal intensity is lower than the first signal intensity upper limit value and higher than the second signal intensity lower limit value, and no sequence signal of the pathogen is detected in the environmental control sample, wherein the first signal intensity upper limit value represents the upper limit standard for the suspected determination of Mycobacterium tuberculosis, and the second signal intensity lower limit value represents the lower limit standard for the suspected determination of Mycobacterium tuberculosis;
[0156] If the universal pathogen species is Mycobacterium fortuitum, the analysis conditions are set as follows: among the chimeric primers used, the ratio of the number of primers that actually detect the target sequence to the total number of primers is greater than or equal to a first preset primer detection threshold, and the sequence substitution rate between each set of detection primers and the Mycobacterium fortuitum genome is greater than or equal to a preset primer matching threshold, wherein the preset primer matching threshold represents the minimum standard for the sequence substitution rate between the chimeric primers and the target genome;
[0157] If the universal pathogen species is Mycobacterium fortuitum, the analysis conditions are set as follows: among the chimeric primers used, the number of primers that actually detect the target sequence is greater than or equal to a fourth preset primer quantity threshold, and the sequence replacement rate between each detection primer and the Mycobacterium fortuitum genome is greater than or equal to a preset primer matching threshold, wherein the fourth preset primer quantity threshold represents the minimum primer quantity standard for actually detecting the target sequence in the chimeric primers when detecting Mycobacterium fortuitum.
[0158] In one embodiment, the pathogen species includes a drug-resistant pathogen species; and the analysis conditions further include one of the following:
[0159] Detection of an associated pathogen of a drug-resistant pathogen species, and the associated pathogen is the target pathogen;
[0160] The number of sequencing sequences of the drug-resistant pathogen species detected is greater than a sequencing sequence number threshold, wherein the sequencing sequence number threshold represents the minimum effective standard for the number of sequencing sequences of the drug-resistant pathogen;
[0161] The sequence variation frequency of the drug-resistant site of the drug-resistant pathogen species detected is greater than or equal to a preset mutation frequency threshold, wherein the preset mutation frequency threshold represents the critical standard of the sequence variation frequency of the drug-resistant site.
[0162] Each module in the aforementioned pathogen-targeted high-throughput sequencing data analysis device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0163] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 6As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, it realizes a method for analyzing pathogen-targeted high-throughput sequencing data. Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0164] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0165] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0166] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0167] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0168] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0169] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for analyzing pathogen-targeted high-throughput sequencing data, characterized in that: The method comprises: Generate targeted and enriched pathogen high-throughput sequencing data through an integrated reaction of a chimeric primer and a pathogen sample to be analyzed, wherein the chimeric primer is composed of a target sequence and an adapter sequence; Matching corresponding analysis conditions to the pathogen species detected in the pathogen high-throughput sequencing data, wherein the analysis conditions are set based on the quality control indicators of the pathogen sample to be analyzed, and the quality control indicators are obtained by comparing all pathogen test samples with standard pathogen samples; According to the analysis conditions, the pathogen high-throughput sequencing data is analyzed to obtain data analysis results.
2. The method according to claim 1, characterized in that Analyzing the pathogen high-throughput sequencing data according to the analysis conditions to obtain data analysis results includes: Screening data sequences that match the pathogen species from the pathogen high-throughput sequencing data, and determining the number of sequencing sequences that the sequencing data sequences map to the target pathogen gene; extracting at least one analysis condition item of the pathogen high-throughput sequencing data from the analysis conditions; According to the correspondence between the sequencing sequence number and each of the analysis condition items, condition matching analysis is performed on the pathogen high-throughput sequencing data to obtain data analysis results.
3. The method according to claim 2, characterized in that Each of the analysis condition items includes at least one of the following five conditions: a historical detection frequency condition item, a sequencing sequence number condition item, a background contamination condition item, a negative quality control condition item, and a primer coverage condition item. The data analysis result includes determining that the pathogen high-throughput sequencing data contains a target pathogen, a suspected pathogen, or is unreliable data; The condition matching analysis is performed on the pathogen high-throughput sequencing data according to the correspondence between the sequencing sequence number and each of the analysis condition items to obtain a data analysis result, including one of the following: If a frequently detected species is detected in the pathogen high-throughput sequencing data, and in the historical detection frequency condition item, the detection frequency of the frequently detected species is greater than a preset detection frequency threshold, then the pathogen high-throughput sequencing data is determined to be untrustworthy data, wherein the preset detection frequency threshold represents the critical value of the occurrence frequency of the frequently detected species in the historical detection data; If, in the sequencing sequence number condition item, it is detected that the sequencing sequence number is greater than or equal to the primary abundance threshold, and in the background contamination condition item, it is detected that the sequence number ratio between the sequencing sequence number and the negative control sequence number is greater than or equal to the primary contamination correction threshold, or if, in the negative quality control condition item, it is detected that the negative control sequence number is equal to the preset negative quality control exemption threshold, then it is determined that the target pathogen exists in the pathogen high-throughput sequencing data, wherein the primary abundance threshold represents the minimum effective abundance standard of the target pathogen sequence number, and the primary contamination correction threshold represents the minimum standard of the sequence number ratio for excluding contamination in high-abundance data; If, in the sequencing sequence number condition item, it is detected that the sequencing sequence number is greater than or equal to the secondary abundance threshold, and in the background contamination condition item, it is detected that the sequence number ratio between the sequencing sequence number and the negative control sequence number is greater than or equal to the secondary contamination correction threshold, or if, in the negative quality control condition item, it is detected that the negative control sequence number is equal to the preset negative quality control exemption threshold, and in the primer coverage condition item, it is detected that the primer coverage corresponding to the sequencing sequence number is greater than or equal to the preset primer coverage threshold, then it is determined that the target pathogen exists in the pathogen high-throughput sequencing data, wherein the secondary abundance threshold represents the critical standard for distinguishing the target pathogen from the suspected pathogen with medium and high abundance sequences, the secondary contamination correction threshold represents the ratio standard for confirming real infection with medium and high abundance, the preset negative quality control exemption threshold represents the critical standard for the number of sequences of negative control without contamination, and the preset primer coverage threshold represents the minimum proportion standard for effective coverage of the primer binding region; If, in the sequencing sequence number condition item, it is detected that the sequencing sequence number is greater than or equal to the secondary abundance threshold, and in the background contamination condition item, it is detected that the sequence number ratio between the sequencing sequence number and the negative control sequence number is greater than or equal to the third-level contamination correction threshold and less than the second-level contamination correction threshold, or if, in the negative quality control condition item, it is detected that the negative control sequence number is equal to the preset negative quality control exemption threshold, then it is determined that the pathogen high-throughput sequencing data contains suspected pathogens, wherein the third-level contamination correction threshold represents the lower limit standard of the contamination ratio for suspected infections with medium and high abundance; If, in the sequencing sequence number condition item, it is detected that the sequencing sequence number is greater than or equal to the third-level abundance threshold and less than the second-level abundance threshold, and in the background contamination condition item, it is detected that the sequence number ratio between the sequencing sequence number and the negative control sequence number is greater than or equal to the fourth-level contamination correction threshold, or if, in the negative quality control condition item, it is detected that the negative control sequence number is equal to the preset negative quality control exemption threshold, and in the primer coverage condition item, it is detected that the primer coverage corresponding to the sequencing sequence number is greater than or equal to the preset primer coverage threshold, then it is determined that there is a suspected pathogen in the pathogen high-throughput sequencing data, wherein the third-level abundance threshold represents the minimum effective abundance standard for the number of suspected pathogen sequences, and the fourth-level contamination correction threshold represents the contamination ratio filtering standard for low-abundance suspected infections.
4. The method according to claim 3, characterized in that The data analysis results also include one of the following: The number of species of the same genus and different species detected in the pathogen high-throughput sequencing data is less than a first preset number threshold, and the number of the first species in the target pathogen is less than or equal to a second preset number threshold, and the number of the second species in the suspected pathogen is less than or equal to a third preset number threshold, wherein the first preset number threshold represents a maximum allowable detection number critical value of the same genus and different species, the second preset number threshold represents a maximum allowable number critical value of the first species, and the third preset number threshold represents a maximum allowable number critical value of the second species; A first species and a second species of the same genus but different species are detected in the pathogen high-throughput sequencing data, and the second species is determined to be an untrustworthy species, wherein the target sequences of the first species and the second species are similar, and a sequence number ratio between the first species and the second species is greater than or equal to a preset sequence ratio threshold, wherein the preset sequence ratio threshold represents a criterion for determining the sequence number ratio of the same genus but different species; Determining whether the species detected in the pathogen high-throughput sequencing data participates in the conditional matching analysis based on the relationship between the sequence number ratio between the sequence number and the negative control sequence number and the secondary contamination correction threshold; If the number of sequencing sequences of the frequently detected species detected in the current batch in the high-throughput sequencing data of the pathogen is greater than the number of sequencing sequences of the frequently detected species in all historical batches, the frequently detected species is released.
5. The method according to claim 3, characterized in that The pathogen species include common pathogen species and special pathogen species, and each of the analysis condition items further includes at least one of a strong positive interference condition item and a weak positive interference condition item; The step of performing condition matching analysis on the pathogen high-throughput sequencing data according to the correspondence between the sequencing sequence number and each of the analysis condition items to obtain a data analysis result further includes one of the following items: If the universal pathogen species is detected as a strong positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the strong positive species is greater than or equal to the first strong positive interference threshold of the strong positive interference condition item, then adjusting the data analysis result, wherein the first strong positive interference threshold represents the critical value for distinguishing the interference determination of the strong positive species in the universal pathogen species; If the universal pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the weak positive species is less than or equal to the first weak positive interference threshold of the weak positive interference condition item, then the data analysis result is retained, wherein the first weak positive interference threshold represents a primary critical value for distinguishing weak positive interference determinations in the universal pathogen species; If the universal pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the weak positive species is less than or equal to the second weak positive interference threshold of the weak positive interference condition item, then the pathogen high-throughput sequencing data is determined to be untrustworthy data, wherein the second weak positive interference threshold represents a secondary critical value for distinguishing weak positive interference determinations in the universal pathogen species; If the special pathogen species is detected as a strong positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the strong positive species is greater than or equal to the second strong positive interference threshold of the strong positive interference condition item, then adjust the data analysis result, wherein the second strong positive interference threshold represents the critical value for distinguishing the interference determination of the strong positive species in the special pathogen species; If the special pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the weak positive species is less than or equal to the third weak positive interference threshold of the weak positive interference condition item, then the data analysis result is retained, wherein the third weak positive interference threshold represents a primary critical value for distinguishing the weak positive interference judgment in the special pathogen species; If the special pathogen species is detected as a weak positive species in the pathogen high-throughput sequencing data, and the number of sequencing sequences of the weak positive species is less than or equal to the fourth weak positive interference threshold of the weak positive interference condition item, then the pathogen high-throughput sequencing data is determined to be unreliable data, wherein the fourth weak positive interference threshold represents the secondary critical value for distinguishing the weak positive interference judgment in the special pathogen species.
6. The method according to claim 1, wherein The pathogen species includes common pathogen species; and the analysis conditions further include one of the following: If the universal pathogen species belongs to the first type of pathogen, the analysis conditions are set as follows: the number of chimeric primers used is greater than or equal to a first preset primer quantity threshold, and the sequence substitution rate between each set of primers and the human metapneumovirus genome is greater than or equal to a preset primer matching threshold, wherein the first preset primer quantity threshold represents the chimeric primer quantity standard required for detecting the first type of pathogen; If the universal pathogen species belongs to the second type of pathogen, the analysis conditions are set as follows: the number of chimeric primers used is greater than or equal to a second preset primer quantity threshold, and the sequence substitution rate between each set of primers and the human metapneumovirus genome is greater than or equal to a preset primer matching threshold, wherein the second preset primer quantity threshold represents the chimeric primer quantity standard required for the detection of the second type of pathogen; If the universal pathogen species belongs to the third type of pathogen, the analysis conditions are set as follows: the number of chimeric primers used is equal to a second preset primer number threshold, and the sequence substitution rate between each set of primers and the target pathogen genome is greater than or equal to a preset primer matching threshold, wherein the third preset primer number threshold represents the standard number of chimeric primers required for detecting the third type of pathogen; If the universal pathogen species is a DNA virus, the analysis condition is set as follows: the sequence signal intensity of the DNA virus in the same batch reaches a preset strong positive signal threshold, wherein the preset strong positive signal threshold represents the critical value of the sequence signal intensity when detecting the DNA virus; If the universal pathogen species is Mycobacterium tuberculosis, the analysis conditions are set as follows: the sequence signal intensity is lower than a first signal intensity upper limit value and higher than a second signal intensity lower limit value, and no sequence signal of the pathogen is detected in the environmental control sample, wherein the first signal intensity upper limit value represents the signal intensity upper limit standard for the suspected determination of Mycobacterium tuberculosis, and the second signal intensity lower limit value represents the signal intensity lower limit standard for the suspected determination of Mycobacterium tuberculosis; If the universal pathogen species is Mycobacterium fortuitum, the analysis conditions are set as follows: among the chimeric primers used, the ratio of the number of primers that actually detect the target sequence to the total number of primers is greater than or equal to a first preset primer detection threshold, and the sequence replacement rate between each set of detection primers and the Mycobacterium fortuitum genome is greater than or equal to a preset primer matching threshold, wherein the preset primer matching threshold represents the minimum standard for the sequence replacement rate between the chimeric primers and the target genome; If the universal pathogen species is Mycobacterium fortuitum, the analysis conditions are set as follows: among the chimeric primers used, the number of primers that actually detect the target sequence is greater than or equal to a fourth preset primer quantity threshold, and the sequence replacement rate between each detection primer and the Mycobacterium fortuitum genome is greater than or equal to a preset primer matching threshold, wherein the fourth preset primer quantity threshold represents the minimum primer quantity standard for actually detecting the target sequence in the chimeric primers when detecting Mycobacterium fortuitum.
7. The method according to claim 1, characterized in that The pathogen species includes a drug-resistant pathogen species; and the analysis conditions further include one of the following: Detecting an associated pathogen of the drug-resistant pathogen species, wherein the associated pathogen is the target pathogen; The number of sequencing sequences of the drug-resistant pathogen species detected is greater than a sequencing sequence number threshold, wherein the sequencing sequence number threshold represents the minimum effective standard for the number of sequencing sequences of the drug-resistant pathogen; The sequence variation frequency of the drug-resistant site of the drug-resistant pathogen species detected is greater than or equal to a preset mutation frequency threshold, wherein the preset mutation frequency threshold represents a critical standard for the sequence variation frequency of the drug-resistant site.
8. An analysis device for pathogen-targeted high-throughput sequencing data, characterized in that: The device comprises: A generation module, configured to generate targeted enriched pathogen high-throughput sequencing data through an integrated reaction of a chimeric primer with a pathogen sample to be analyzed, wherein the chimeric primer is composed of a target sequence and an adapter sequence; a matching module, configured to match corresponding analysis conditions to pathogen species detected by the pathogen high-throughput sequencing data, wherein the analysis conditions are set based on the quality control indicators of the pathogen sample to be analyzed, and the quality control indicators are obtained by comparing and analyzing all pathogen test samples with standard pathogen samples; The analysis module is used to analyze the pathogen high-throughput sequencing data according to the analysis conditions to obtain data analysis results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Composition for detecting respiratory pathogens based on high-throughput sequencing
CN117625848A
Analysis method for sequencing data of targeted detection microorganisms and application
CN120126562A
System and method for combating mycobacterium tuberculosis infections
US20220380786A1
Bacterial 16s rrna gene sequence-based bacterial "species" level detection and analysis method
WO2022262491A1
Cited By
Positive reporting processing method for pathogen targeted high-throughput sequencing data, computer equipment and readable storage medium
CN121687189A
Pathogen-targeted high-throughput sequencing data positive report processing method, computer device and readable storage medium
CN121687189B