A pathogenic microorganism analysis and identification system and its application

By designing a pathogenic microbial analysis and identification system, the automated analysis of mNGS technology's down-of-machine data is realized, the dependence on professional and technical personnel is solved, the analysis efficiency and result reliability are improved, and the wide application of mNGS technology in clinical etiological diagnosis is promoted.

CN115862739BActive Publication Date: 2025-06-03SHENZHEN GENEPLUS CLINICAL LAB +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211377592.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-06-03
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

In pathogenic diagnosis, mNGS technology relies on professional and technical personnel for the analysis of original down-machine data, which limits its wide application in clinical pathogenic diagnosis.

Method used

A pathogenic microbial analysis and identification system was designed, including data statistics module, sample management module, experimental management module, information analysis module, report management module and system management module. The sample analysis submodule in the information analysis module is automated, including data quality control, sequence classification and credibility screening and other steps.

Benefits of technology

The automated analysis of mNGS original off-machine data is realized, which reduces the dependence on professional and technical personnel, improves the analysis efficiency and reliability of results, and solves the application limitations of mNGS technology in clinical etiological diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862739B_ABST
    Figure CN115862739B_ABST
Patent Text Reader

Abstract

The present application discloses a pathogenic microorganism analysis and identification system and its application. The system of the present application includes a data statistics module, a sample management module, an experiment management module, an information analysis module, a report management module, and a system management module; the information analysis module includes a sample analysis sub-module, which is used to implement data quality control steps, host removal steps, sequence classification steps, calculation of classification credibility index steps, and filtering of non-pathogenic bacteria steps; the credibility is screened according to the number of reads at the species level and genus level, the number of unique Kmers at the species level, the genome size of the corresponding species, the read coverage size, the number of windows containing reads, and the microbial genome reference genome by the number of reads. The present application realizes the automatic analysis of mNGS raw off-machine data through each module, obtains reliable classification results according to the credibility screening scheme, automatically reports the microorganism, and solves the dependence on professional technicians for the analysis of mNGS raw off-machine data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of pathogenic microorganism analysis and identification, and in particular to a pathogenic microorganism analysis and identification system and its application. Background Art

[0002] Infectious diseases are a major global disease burden, and currently a variety of molecular detection schemes have been applied to clinical etiology diagnosis, such as blood culture, immunology, PCR, etc. However, these methods generally have shortcomings such as high false negatives, long culture cycles, and low detection performance.

[0003] Metagenomic next-generation sequencing (mNGS) is an environmental microbial sequencing technology based on high-throughput sequencing. In etiological diagnosis, mNGS technology can directly extract the nucleic acid of all microorganisms in the infection site, sequence them on a high-throughput sequencing platform, and obtain the species information of suspected pathogenic microorganisms through microbial database comparison and algorithm analysis. mNGS technology plays an important role in etiological diagnosis with its wide detection range, high accuracy, good specificity, and strong detection ability.

[0004] At present, mNGS technology is becoming more and more mature, and it has been clinically recognized and gradually used in many fields such as bloodstream infection, respiratory tract infection, central nervous system infection, fever of unknown cause and pneumonia; mNGS technology has been used for pathogen detection of infectious diseases in many consensuses and guidelines. However, the analysis and detection of mNGS usually requires professional bioinformatics personnel and interpreters, which greatly limits the widespread application of mNGS technology in clinical etiology diagnosis.

[0005] Therefore, how to solve the dependence of raw off-machine data analysis in mNGS technology on professional technicians is a technical problem that needs to be solved urgently in this field. Summary of the invention

[0006] The purpose of this application is to provide a new pathogenic microorganism analysis and identification system and its application.

[0007] In order to achieve the above objectives, this application adopts the following technical solutions:

[0008] One aspect of the present application discloses a pathogenic microorganism analysis and identification system, including a data statistics module, a sample management module, an experiment management module, an information analysis module, a report management module, and a system management module; among them, the information analysis module includes a sample analysis sub-module, and the sample analysis sub-module includes steps for implementing the following: a data quality control step, including 1) removing low-quality reads according to the data sequencing quality; 2) removing residual adapter sequences; 3) marking low-complexity sequences, including breaking the genomic library into a sequence set with a length of 31bp - 50bp and a step size of 1bp for shifting, comparing each sequence in the sequence set with the human reference genome, if 100% match, replacing the base at the corresponding position of the sequence on the fusion genome with N to achieve the purpose of shielding the human homologous region sequence and reducing the interference of the human homologous region; shielding plasmid sequences, shielding the bacterial plasmid homologous region in the high-quality genomic library through the PLSDB database; a de-host step, including directly aligning with the human genome sequence to remove the human genomic sequence in the sequencing data; a sequence classification step, including 1) classifying species based on the kmer alignment method; 2) filtering the publicly available genomic sequences, including filtering non-specific sequences, filtering low-complexity sequences, marking homologous sequences, marking highly pathogenic bacterial sequences, marking background bacterial sequences; a step of calculating classification credibility indicators, including screening the credibility of the alignment results using the following indicators:

[0009] SG = number of reads at the species level / number of reads at the genus level, > 0.6

[0010] UN = number of reads at the species level / number of unique Kmers at the species level, < 0.8

[0011] Depth = number of reads at the species level × 50 / size of the corresponding species genome

[0012] Coverage = size of reads coverage / size of the corresponding species genome, > 3 regions

[0013] SpeciesReadsDiscreteness = number of windows containing reads / number of windows obtained by dividing the microbial reference genome according to the number of reads, > 0.4

[0014] Comprehensively screening according to the above indicators, a reliable classification result is finally obtained;

[0015] A step of filtering non-pathogenic bacteria, including filtering non-pathogenic bacteria according to the reagent background database, the symbiotic bacteria database, and the conditional / important pathogenic bacteria library, combined with the number of output reads, to determine the finally reported microorganisms.

[0016] It should be noted that after screening according to SG, UN, Depth, Coverage, and SpeciesReadsDiscreteness, the highly credible and truly existing microbial information in the sample can be obtained, and at the same time, false-positive microorganisms caused by factors such as bioinformatics comparison methods and databases can be excluded.

[0017] It should also be noted that through the automated operation of each module, the pathogenic microorganism analysis and identification system of the present application can achieve the automated analysis of the original mNGS off-machine data, and finally obtain a reliable classification result according to the credibility screening scheme of the present application, and report the microorganism; it solves the dependence of the analysis of the original mNGS off-machine data on professional technicians.

[0018] In one implementation manner of the present application, the data quality control step for masking plasmid sequences includes: 1) removing the sequences in the genomic fasta file sequence name description information that contain the keywords "Plasmid" or "plasmid"; 2) breaking the genomic sequence excluding plasmids into a sequence set with a length of 31bp - 50bp and a shifting step of 1bp, comparing each sequence in the sequence set with the plasmid database, and if there is a 100% match, replacing the bases at the corresponding positions of the sequence on the fused genome with N to achieve the purpose of masking the plasmid homologous region sequences.

[0019] In one implementation manner of the present application, the sequence classification step classifies species based on the kmer alignment method, including breaking 50bp reads into 14 segments with a length of 37bp, and the kmer sequences moving by 1bp; when aligning, comparing the kmer sequences with the genomic sequence library, and statistically analyzing the alignment results of each kmer according to the reads; the classification of each read is determined according to the species with the largest number of matches of the kmer of the read.

[0020] It should be noted that since the present application breaks the reads into shorter kmers for alignment, faster alignment speed and higher accuracy can be obtained. For example, in one implementation manner of the present application, for 20 samples, with an average of 34M reads and SE50 data, the test results on a node with 188G internal reference and 96 CPUs show that the serial analysis takes 1 hour and 10 minutes, and the parallel analysis takes 30 minutes; the alignment speed is significantly improved and the accuracy is very high. It can be understood that breaking 50bp reads into 14 segments with a length of 37bp is only a specific scheme and parameter in one implementation manner of the present application. On this basis, reads of other different lengths can also be broken into more or fewer segments of kmer sequences that are shorter than the original reads. As for the shifting step, it is generally 1bp, but it can also be longer.

[0021] In an implementation of the present application, the sequence classification step filters the publicly available genomic sequences, where the publicly available genomic sequences are sequences from public databases. For example, the public databases include NCBI and FDA-ARGOS.

[0022] It should be noted that the original genomic sequence data comes from public databases such as NCBI and FDA-ARGOS. Due to reasons such as assembly, the original genomic data may introduce incorrect sequences, resulting in classification errors during alignment. Therefore, it is necessary to filter the publicly available genomic sequences. The present application preferably filters the publicly available genomic sequences through the following steps: filtering non-specific sequences, filtering low-complexity sequences, marking homologous sequences, marking highly pathogenic bacteria sequences, and marking background bacteria sequences. In an implementation of the present application, the genomes obtained through the above steps are integrated into an alignment database, and the final obtained microorganisms include: 8,525 species of bacteria, 6,913 species of viruses, 440 species of fungi, and 160 species of parasites.

[0023] In an implementation of the present application, the rules for reporting and not reporting in the step of filtering non-pathogenic bacteria are as follows:

[0024] 1) Microorganisms with ≤ 3 reads are not reported;

[0025] 2) Reagent bacteria are not reported;

[0026] 3) If it is both a reagent bacterium and a conditional / important pathogenic bacterium, it is reported;

[0027] 4) Non-reagent bacteria and conditional / important pathogenic bacteria are not reported;

[0028] 5) The following microorganisms do not consider the minimum number of reads and must be reported once detected: Mycobacterium tuberculosis, Mycobacterium tuberculosis complex, Mycobacterium avium, Mycobacterium avium-intracellulare, Mycobacterium intracellulare, Mycobacterium africanum, Mycobacterium orygis, Mycobacterium bovis, Mycobacterium microti, Mycobacterium canetti, Mycobacterium caprae, Mycobacterium pinnipedii, Mycobacterium suricattae, Mycobacterium mungi.

[0029] In an implementation manner of the present application, the data statistics module includes those for sample statistics, disk space statistics, sample submission statistics, and sample type statistics.

[0030] In an implementation manner of the present application, the sample statistics includes statistics of the following information: (1) Statistics of new samples within a period: including the number of new samples within a period; (2) Statistics of the total number of samples: including the number of all samples in the system; (3) Statistics of the total number of issued reports: including the number of sample reports that have passed the review; (4) Statistics of the number of batches put on the machine within a period: including the number of batches put on the machine within a period; (5) Statistics of samples under analysis: including the number of analysis tasks with the analysis status of "under analysis"; (6) Statistics of analyzed samples: including the number of analysis tasks with the analysis status of "analysis completed"; The disk space statistics includes statistics of the disk usage and / or the remaining disk space in the system; The sample submission statistics includes, according to a set time length, statistics of the number of new samples in each period within the set time length; The sample type statistics includes statistics of the occupancy ratio of samples corresponding to each sample type.

[0031] For example, the sample statistics mainly includes statistics of new samples this week: the number of new samples this week; the total number of samples: the number of all samples in the system; the total number of issued reports: the number of sample reports that have passed the review; the number of batches put on the machine this week: the number of batches put on the machine this week; samples under analysis: the number of analysis tasks with the analysis status of "under analysis"; analyzed samples: the number of analysis tasks with the analysis status of "analysis completed". Disk space: Statistics of the disk usage in Gene+Box. Sample submission statistics: Statistics of the number of new samples in each of the past 12 months starting from the current month on a monthly basis. Sample type statistics: Statistics of the occupancy ratio of samples corresponding to each sample type. The specific time period can be adjusted according to requirements and is not specifically limited here.

[0032] In an implementation manner of the present application, the sample management module includes those for adding single sample information, batch importing sample information, modifying sample information, deleting sample information, viewing the details of sample information, and querying sample information according to single or multiple conditions.

[0033] In an implementation manner of the present application, the experiment management module includes those for creating a task of putting on the machine, importing a task of putting on the machine, modifying a task of putting on the machine, deleting a task of putting on the machine, viewing the details of the task of putting on the machine, and performing data splitting.

[0034] In an implementation manner of the present application, the information analysis module further includes an analysis task issuing sub-module, and the analysis task issuing sub-module includes those for analysis issuing, query function, and deletion; Analysis issuing includes: sample marking function, sample pairing function, and sample adding function.

[0035] In one implementation of the present application, the sample analysis sub-module further includes functions for viewing analysis tasks, re-analyzing, deleting, querying, performing information analysis, and generating reports for execution interpretation.

[0036] In one implementation of the present application, the report management module includes functions for result verification, report review, report download, importing into Excel, and viewing report details; among them, result verification includes: reporting mark function, report generation function, and report query function.

[0037] In one implementation of the present application, the system management module includes functions for account management, unit management, role management, project management, data dictionary, and report management.

[0038] Preferably, account management includes creating new accounts, associating roles, querying, modifying accounts, resetting passwords, deleting accounts, and viewing account details.

[0039] Preferably, unit management includes adding new cooperative units, querying, viewing unit details, modifying units, and deleting units.

[0040] Preferably, role management includes adding new roles, querying, modifying roles, deleting roles, and associating permissions.

[0041] Preferably, project management includes adding new projects, importing background bacteria, modifying projects, and deleting projects.

[0042] Preferably, the data dictionary includes adding, modifying, deleting, and refreshing items in the classification.

[0043] Preferably, report management includes adding templates, modifying templates, downloading templates, and deleting templates.

[0044] On the other hand, the present application discloses the application of the pathogenic microorganism analysis and identification system of the present application in the detection or identification of at least one of bacteria, viruses, fungi, and parasites.

[0045] Due to the above technical solutions, the beneficial effects of the present application are as follows:

[0046] The pathogenic microorganism analysis and identification system of the present application realizes the automated analysis of mNGS raw off-machine data through each module, then obtains reliable classification results according to the credibility screening scheme, and automatically reports the microorganisms, effectively solving the dependence on professional technical personnel for the analysis of mNGS raw off-machine data. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic structural diagram of the pathogenic microorganism analysis and identification system in the embodiments of the present application;

[0048] Figure 2It is a screenshot of the interface for adding new samples in the sample management of the pathogenic microorganism analysis and identification system in the embodiments of the present application;

[0049] Figure 3 It is a screenshot of the interface for entering sample information in the pathogenic microorganism analysis and identification system in the embodiments of the present application;

[0050] Figure 4 It is a screenshot of the interface for batch importing sample information in the pathogenic microorganism analysis and identification system in the embodiments of the present application;

[0051] Figure 5 It is a screenshot of the interface for setting the information of the sample to be analyzed in the pathogenic microorganism analysis and identification system in the embodiments of the present application;

[0052] Figure 6 It is a screenshot of the interface for issuing analysis tasks in the pathogenic microorganism analysis and identification system in the embodiments of the present application;

[0053] Figure 7 It is a screenshot of the interface for querying the analysis progress through system query in the pathogenic microorganism analysis and identification system in the embodiments of the present application;

[0054] Figure 8 It is a screenshot of the interface for directly interpreting the results by importing the results in the pathogenic microorganism analysis and identification system in the embodiments of the present application;

[0055] Figure 9 It is a screenshot of the interface for reporting the situation in the pathogenic microorganism analysis and identification system in the embodiments of the present application. Detailed implementation manners

[0056] The present application will be further described in detail below in conjunction with the accompanying drawings through specific implementation manners. In the following implementation manners, many detailed descriptions are provided to enable a better understanding of the present application. However, those skilled in the art can easily recognize that some of these features can be omitted in different situations, or can be replaced by other devices, materials, or methods. In some cases, some operations related to the present application are not shown or described in the specification to avoid overwhelming the core part of the present application with excessive descriptions. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.

[0057] For the original off-machine data of mNGS, currently, there is a lack of a simple, easy-to-use, and automated analysis and interpretation platform on the market; therefore, the analysis and detection of mNGS greatly rely on professional technical personnel, thus restricting the wide application of mNGS technology in clinical pathogen diagnosis.

[0058] Based on the above research and understanding, the present application has developed and provided a new pathogenic microorganism analysis and identification system, namely the Gene+box Pathogen Analysis and Interpretation All-in-One Machine, which is embedded with 6 mature information management system modules, including a mature biological information process. Relying on a powerful database, it can effectively analyze bacteria, viruses, fungi, parasites, etc. in samples, thereby accurately generating reports and achieving fast, stable, and reliable clinical delivery.

[0059] Embodiment

[0060] The pathogenic microorganism analysis and identification system in this example, as Figure 1 shown, includes a data statistics module 11, a sample management module 12, an experiment management module 13, an information analysis module 14, a report management module 15, and a system management module 16. The functions and roles of each module are detailed as follows:

[0061] The data statistics module 11 includes functions for: sample statistics, disk space, sample submission statistics, and sample type statistics.

[0062] Among them, sample statistics, for example, mainly includes: new samples this week: the number of new samples this week; total sample volume: the total number of samples in the system; total number of issued reports: the number of sample reports that have passed the review; number of batches on the machine this week: the number of batches on the machine this week; samples in analysis: the number of analysis tasks with the analysis status of "in analysis"; analyzed samples: the number of analysis tasks with the analysis status of "analysis completed".

[0063] Disk space statistics, for example, the disk usage statistics in Gene+Box. Sample submission statistics, for example, monthly statistics of the number of new samples in each of the past 12 months starting from the current month. Sample type statistics, for example, statistics of the occupancy ratio of samples corresponding to each sample type.

[0064] The sample management module 12 includes functions for: adding new samples, batch importing, modifying samples, deleting samples, viewing details, and querying / advanced querying.

[0065] The experiment management module 13 includes functions for: creating a new task on the machine, importing a task on the machine, modifying a task on the machine, deleting a task on the machine, viewing details of the task on the machine, and performing data splitting.

[0066] The information analysis module 14 includes an analysis task assignment sub-module and a sample analysis sub-module. Among them, the analysis task assignment sub-module mainly includes functions for: analysis assignment, querying, and deleting; among which, analysis assignment includes: marking samples, sample pairing, and adding samples.

[0067] The sample analysis sub-module includes functions for: viewing analysis tasks, re-analysis, deleting, querying, performing information analysis, and performing interpretation to generate reports.

[0068] Moreover, the sample analysis sub-module mainly includes steps for implementing the following:

[0069] Data quality control steps, including: 1) removing low-quality reads according to the data sequencing quality; 2) removing residual adapter sequences. Some reads may have residual adapter sequences, which affect subsequent alignment, so they need to be removed; 3) marking low-complexity sequences. Low-complexity sequences have a greater impact on the result of host removal in the next step. Higher low-complexity regions may lead to incomplete host removal and thus classification errors. Therefore, these regions are marked to obtain a more accurate classification result compared with existing similar solutions. Specifically, it includes breaking the genomic DNA in the library into a sequence set with a length of 31bp - 50bp and a sliding step of 1bp. Each sequence in the sequence set is compared with the human reference genome. If there is a 100% match, the base at the corresponding position of the sequence on the fusion genome is replaced with N to achieve the purpose of shielding the sequence in the human homologous region and reducing the interference of the human homologous region. In addition, the plasmid sequence is also shielded. The homologous region of bacterial plasmids in the high-quality genomic library is shielded through the PLSDB database. Species classified as fungi, viruses, and parasites do not contain plasmid sequences, so there is no need to remove plasmid homologous sequences.

[0070] Among them, shielding the plasmid sequence includes: 1) removing sequences containing the keywords "Plasmid" or "plasmid" in the sequence name description information of the genomic fasta file; 2) breaking the genomic sequence excluding plasmids into a sequence set with a length of 31bp - 50bp, specifically 31bp in this example, and a sliding step of 1bp. Each sequence in the sequence set is compared with the plasmid database. If there is a 100% match, the base at the corresponding position of the sequence on the fusion genome is replaced with N to achieve the purpose of shielding the plasmid homologous region sequence.

[0071] Host removal step, including directly aligning with the human genomic sequence to remove the human genomic sequence in the sequencing data, reducing the time and classification accuracy for the next classification.

[0072] Sequence classification steps, including: 1) classifying species based on the kmer alignment method; 2) filtering the publicly available genomic sequences, including filtering non-specific sequences, filtering low-complexity sequences, marking homologous sequences, marking highly pathogenic bacteria sequences, and marking background bacteria sequences. Among them, the publicly available genomic sequences are sequences from public databases, such as the NCBI and FDA-ARGOS public databases.

[0073] Among them, classifying species based on the kmer alignment method includes breaking 50bp reads into 14 segments with a length of 37bp, and the kmer sequences moving 1bp at a time; when aligning, the kmer sequences are aligned with the genomic sequence library, and the alignment results of each kmer are statistically analyzed according to the reads; the classification of each read is determined according to the species with the largest number of matches of the kmer of the read. Since the reads are broken into shorter kmers for alignment, faster alignment speed and higher accuracy are obtained. The test results of this example based on 20 samples, with an average of 34M reads and SE50 data, on a node with 188G internal reference and 96 CPUs show that the serial analysis takes 1 hour and 10 minutes, and the parallel analysis takes 30 minutes.

[0074] The original genomic sequence data is sourced from public databases such as NCBI, FDA-ARGOS, etc. Due to reasons such as assembly, the original genomic data will introduce error sequences, resulting in incorrect classification during alignment. Therefore, it is necessary to filter the publicly available genomic sequences. This example uses the following steps to filter the publicly available genomic sequences: filtering non-specific sequences, filtering low-complexity sequences, marking homologous sequences, marking highly pathogenic bacteria sequences, and marking background bacteria sequences. The genomes obtained through the above steps are then integrated into an alignment database, and finally the following numbers of microorganisms are obtained: 8525 species of bacteria, 6913 species of viruses, 440 species of fungi, and 160 species of parasites.

[0075] Inevitably, incorrect classification results will occur due to incorrect alignment in the alignment-based method. This example uses the following indicators to screen the credibility of the alignment results, so as to provide more accurate alignment results. That is, the steps for calculating the classification credibility index include using the following indicators to screen the credibility of the alignment results:

[0076] SG = number of reads at the species level / number of reads at the genus level, > 0.6

[0077] UN = number of reads at the species level / number of unique Kmers at the species level, < 0.8

[0078] Depth = number of reads at the species level × 50 / size of the corresponding species genome

[0079] Coverage = size of reads coverage / size of the corresponding species genome, > 3 regions

[0080] SpeciesReadsDiscreteness = number of windows containing reads / number of windows obtained by dividing the microbial reference genome by the number of reads, > 0.4

[0081] According to the comprehensive screening of the above indicators, reliable classification results are finally obtained.

[0082] Step of filtering non-pathogenic bacteria, including filtering non-pathogenic bacteria according to the reagent background database, symbiotic bacteria database, conditional / important pathogenic bacteria database, and combining the number of output reads to determine the finally reported microorganisms.

[0083] Among them, in the step of filtering non-pathogenic bacteria, the rules for reporting and not reporting are as follows:

[0084] 1) Microorganisms with the number of reads ≤ 3 are not reported;

[0085] 2) Reagent bacteria are not reported;

[0086] 3) If it is both a reagent bacterium and a conditional / important pathogenic bacterium, it is reported;

[0087] 4) Non-reagent bacteria and conditional / important pathogenic bacteria are not reported;

[0088] 5) The following microorganisms do not consider the minimum number of reads and must be reported once detected: Mycobacterium tuberculosis, Mycobacterium tuberculosis complex, Mycobacterium avium, Mycobacterium avium-intracellulare, Mycobacterium intracellulare, Mycobacterium africanum, Mycobacterium orygis, Mycobacterium bovis, Mycobacterium microti, Mycobacterium canetti, Mycobacterium caprae, Mycobacterium pinnipedii, Mycobacterium suricattae, Mycobacterium mungi.

[0089] Report management module 15, including for: result verification, report review, report download, import to Excel, view report details. Among them, result verification includes: reporting mark, report generation, query function.

[0090] The system management module 16 includes functions for: account management, unit management, role management, project management, data dictionary, and report management. Among them, account management includes creating new accounts, associating roles, querying, modifying accounts, resetting passwords, deleting accounts, and viewing account details; unit management includes adding cooperative units, querying, viewing unit details, modifying units, and deleting units; role management includes adding roles, querying, modifying roles, deleting roles, and associating permissions; project management includes adding projects, importing background bacteria, modifying projects, and deleting projects; data dictionary includes adding, modifying, deleting, and refreshing items in the classification; report management includes adding templates, modifying templates, downloading templates, and deleting templates.

[0091] The pathogenic microorganism analysis and identification system in this example, namely the Gene+box pathogen analysis and interpretation all-in-one machine, has the following system architecture design:

[0092] The system adopts a front-end and back-end separated development architecture. The back-end uses the Spring Cloud microservice framework, introducing the idea of componentization to achieve high cohesion and low coupling, and integrating the front-end and back-end through the Nginx reverse proxy software.

[0093] 1.1 Front-end

[0094] Front-end framework: VUE

[0095] Front-end templates: ElementUI, VantUI

[0096] 1.2 Back-end

[0097] Project framework: Spring Boot + Spring Cloud + MyBatis-Plus

[0098] Database: MySQL

[0099] Distributed registry: Spring Cloud Eureka

[0100] Project deployment: Docker containers

[0101] The entire project uses the microservice framework. When users access the system through a browser or a mobile device, they are first routed by Nginx to the front-end static resource page or the back-end service.

[0102] The back-end service adopts the SpringCloud microservice architecture. Among them, the background module corresponds to user authorization and authentication, and then business processing is carried out. According to the different business selections of users, a series of automatic operations will be performed to call different modules for analysis, and finally the results will be returned.

[0103] The registry is used for each module to register on it, so that each module can discover each other.

[0104] The pathogenic microorganism analysis and identification system in this example has completed the entire process from analysis to interpretation and finally to report generation, significantly improving work efficiency, simplifying the operation process, reducing labor costs, and enhancing the user experience. The system has achieved the entire process from analysis to interpretation and finally to report generation, enabling fast, stable, and reliable clinical delivery.

[0105] Usage Example

[0106] The usage method of this system is simple. The following is a specific usage example:

[0107] 1. Sample Management

[0108] Add new samples. Specifically, as Figure 2 shown, it mainly includes entering sample information or batch importing sample information. Moreover, sample information can also be modified or deleted. For sample information entry, as Figure 3 shown, fill in the sample number, test items, blood pathogen detection, respiratory pathogen detection, etc. Batch importing of sample information can also be carried out as Figure 4 shown.

[0109] 2. Laboratory Management

[0110] Set the information for the instrument run. As Figure 5 shown, similarly, the information for the instrument run can be entered or imported, and the information for the instrument run can also be modified or deleted, etc.

[0111] 3. Analysis and Interpretation

[0112] Mainly issue analysis tasks. As Figure 6 shown, it includes issuing analysis tasks, selecting the samples to be analyzed, and the samples can also be deleted. Or, as Figure 7 shown, query the analysis progress through the system query. After the analysis is completed, the results can be interpreted, or the results can be directly interpreted by importing. As Figure 8 shown, interpret the samples, download the reports, and Excel tables that have been interpreted can also be imported.

[0113] For the entire process in this example, it will automatically give whether to report the microorganism, and the interpreter only needs to confirm, or the reporting situation can also be manually changed, as Figure 9 shown.

[0114] 4. Clinical Performance

[0115] Collect 24 bronchoalveolar lavage fluids from infected patients in the respiratory department of the same hospital. Use the analysis method to extract and analyze the pathogenic microorganisms in the bronchoalveolar lavage fluid, and finally observe whether the pathogenic microorganisms found by this method are consistent with the clinical results. The patient information is shown in Table 1.

[0116] Table 1 Clinical Sample Information Table

[0117]

[0118]

[0119] Based on the final reporting situation of 24 clinical samples, as shown in Table 2, the sensitivity is 100%, the specificity is 100%, and the positive agreement rate is 91.6%.

[0120] Table 2 Reported Results of Clinical Samples

[0121] TP: 22 FP: 2 FN: 0 TN: 0

[0122] The above results show that the pathogenic microorganism analysis and identification system in this example can automate the whole process from analysis to interpretation and finally report generation, achieving fast, stable and reliable clinical delivery, with both sensitivity and specificity as high as 100%.

[0123] The above content is a further detailed description of the present application in combination with specific implementation manners, and it cannot be determined that the specific implementation of the present application is only limited to these descriptions. For those of ordinary skill in the technical field to which the present application belongs, without departing from the concept of the present application, several simple deductions or substitutions can also be made.

Claims

1. A pathogenic microorganism analysis and identification system, characterized in that: it includes a data statistics module, a sample management module, an experiment management module, an information analysis module, a report management module and a system management module; the information analysis module includes a sample analysis sub-module, and the sample analysis sub-module includes steps for implementing the following; The data quality control step includes: 1) removing low-quality reads according to the data sequencing quality; 2) removing residual adapter sequences; 3) marking low-complexity sequences, including breaking the genomic DNA in the library into a sequence set with a length of 31bp - 50bp and a shifting step size of 1bp, comparing each sequence in the sequence set with the human reference genome, if 100% match, then replacing the base at the corresponding position of the sequence on the fusion genome with N, so as to shield the sequences in the human homologous region and reduce the interference of the human homologous region; shield the plasmid sequences by shielding the homologous regions of bacterial plasmids in the high-quality genomic library through the PLSDB database; The host removal step includes directly comparing with the human genome sequence to remove the human genomic sequences in the sequencing data; The sequence classification step includes: 1) classifying species based on the kmer alignment method; 2) filtering the publicly available genomic sequences, including filtering non-specific sequences, filtering low-complexity sequences, marking homologous sequences, marking highly pathogenic bacterial sequences, and marking background bacterial sequences; The step of calculating the classification confidence index includes screening the confidence of the alignment results using the following indicators, SG = number of reads at the species level / number of reads at the genus level, > 0.6 UN = number of reads at the species level / number of unique Kmers at the species level, < 0.8 Coverage = reads coverage size / corresponding species genome size, > 3 regions SpeciesReadsDiscreteness = number of windows containing reads / number of windows obtained by dividing the microbial reference genome by the number of reads, > 0.4 Comprehensively screen according to the above indicators to finally obtain a reliable classification result; The step of filtering non-pathogenic bacteria includes filtering non-pathogenic bacteria according to the reagent background database, symbiotic bacteria database, conditional / important pathogenic bacteria library, combined with the number of output reads, to determine the finally reported microorganisms; In the data quality control step, shielding the plasmid sequences includes: 1) removing the sequences whose sequence name description information in the genomic fasta file contains the keywords "Plasmid" or "plasmid"; 2) breaking the genomic sequence excluding plasmids into a sequence set with a length of 31bp - 50bp and a shifting step size of 1bp, comparing each sequence in the sequence set with the plasmid database, if 100% match, then replacing the base at the corresponding position of the sequence on the fusion genome with N, so as to shield the plasmid homologous region sequences; 2. The pathogenic microorganism analysis and identification system according to claim 1, characterized in that: In the sequence classification step, species are classified based on the kmer alignment method, including breaking 50bp reads into 14 segments with a length of 37bp, and generating kmer sequences that move 1bp at a time; when aligning, the kmer sequences are aligned with the genomic sequence library, and the alignment results of each kmer are statistically analyzed according to the reads; the classification of each read is determined according to the species with the largest number of alignable kmers of the read.

3. The pathogenic microorganism analysis and identification system according to claim 1, characterized in that: The publicly available genomic sequences are sequences from public databases.

4. The pathogenic microorganism analysis and identification system according to claim 3, characterized in that: The public databases include NCBI and FDA-ARGOS.

5. The pathogenic microorganism analysis and identification system according to claim 1, characterized in that: In the step of filtering non-pathogenic bacteria, the rules for reporting and not reporting are as follows: 1) Microorganisms with ≤3 reads are not reported; 2) Reagent bacteria are not reported; 3) If it is both a reagent bacterium and a conditional / important pathogenic bacterium, it is reported; 4) Non-reagent bacteria and conditional / important pathogenic bacteria are not reported; 5) The following microorganisms do not consider the minimum number of reads and must be reported if detected: Mycobacterium tuberculosis, Mycobacterium tuberculosis complex, Mycobacterium avium, Mycobacterium avium-intracellulare, Mycobacterium intracellulare, Mycobacterium africanum, Mycobacterium orygis, Mycobacterium bovis, Mycobacterium microti, Mycobacterium canetti, Mycobacterium caprae, Mycobacterium pinnipedii, Mycobacterium suricattae, Mycobacterium mungi.

6. The pathogenic microorganism analysis and identification system according to any one of claims 1-5, characterized in that: The data statistics module includes those for sample statistics, disk space statistics, sample submission statistics, and sample type statistics.

7. The pathogenic microorganism analysis and identification system according to claim 6, characterized in that: The sample statistics include the following information: (1) New sample statistics over a period of time: including the number of new samples in a period of time; (2) Total sample statistics: including the total number of all samples in the system; (3) Total issued report statistics: including the number of sample reports that have passed the review; (4) Number of batches put into the machine over a period of time: including the number of batches put into the machine in a period of time; (5) Samples in analysis: including the number of analysis tasks with the analysis status of "analyzing"; (6) Analyzed samples: including the number of analysis tasks with the analysis status of "analysis completed"; The disk space statistics include the statistics of disk usage and / or disk remaining in the system; The sample submission statistics include the statistics of the number of new samples in each period within a set time length according to the set time length; The sample type statistics include the statistics of the occupancy ratio of samples corresponding to each sample type.

8. The pathogenic microorganism analysis and identification system according to any one of claims 1-5, characterized in that: The sample management module includes functions for adding single sample information, batch importing sample information, modifying sample information, deleting sample information, viewing detailed sample information, and querying sample information according to single or multiple conditions.

9. The pathogenic microorganism analysis and identification system according to any one of claims 1-5, characterized in that: The experiment management module includes functions for creating a task of putting into the machine, importing a task of putting into the machine, modifying a task of putting into the machine, deleting a task of putting into the machine, viewing the details of putting into the machine, and performing data splitting.

10. The pathogenic microorganism analysis and identification system according to any one of claims 1-5, characterized in that: The information analysis module further includes an analysis task issuing sub-module, and the analysis task issuing sub-module includes functions for analysis issuing, querying, and deleting; The analysis issuing includes: sample marking function, sample pairing function, and sample adding function; The sample analysis sub-module further includes functions for viewing analysis tasks, re-analyzing, deleting, querying, performing information analysis, and performing interpretation to generate a report.

11. The pathogenic microorganism analysis and identification system according to any one of claims 1-5, characterized in that: The report management module includes functions for result verification, report review, report download, Excel import, and viewing report details; Among them, the result verification includes: reported marking function, report generation function, and report query function.

12. The pathogenic microorganism analysis and identification system according to claim 11, characterized in that: The system management module includes functions for account management, unit management, role management, project management, data dictionary, and report management.

13. The pathogenic microorganism analysis and identification system according to claim 12, characterized in that: The account management includes creating a new account, associating roles, querying, modifying the account, resetting the password, deleting the account, and viewing account details.

14. The pathogenic microorganism analysis and identification system according to claim 12, characterized in that: The unit management includes adding a cooperative unit, querying, viewing unit details, modifying the unit, and deleting the unit.

15. The pathogenic microorganism analysis and identification system according to claim 12, characterized in that: the role management includes adding a role, querying, modifying a role, deleting a role, and associating permissions.

16. The pathogenic microorganism analysis and identification system according to claim 12, characterized in that: the project management includes adding a project, importing background bacteria, modifying a project, and deleting a project.

17. The pathogenic microorganism analysis and identification system according to claim 12, characterized in that: the data dictionary includes adding, modifying, deleting, and refreshing items in the classification.

18. The pathogenic microorganism analysis and identification system according to claim 12, characterized in that: the report management includes adding a template, modifying a template, downloading a template, and deleting a template.

Citation Information

Patent Citations

  • Method and device for microbiological analysis of host sample

    CN111009286A

  • Pathogenic microorganism analysis and identification system and application thereof

    CN111462821A