Conjoint analysis method for pathogen recognition and patient genotype characteristics

By integrating host genotype analysis into metagenomic detection processes, the problem of the independence of pathogen and host detection has been solved, enabling efficient and accurate simultaneous detection of pathogen and host information, providing a basis for personalized medicine, and improving the diagnosis and treatment of infectious diseases.

CN120913649APending Publication Date: 2025-11-07中国人民解放军总医院第八医学中心
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510974956.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies have several drawbacks in pathogen detection and host interaction research, including the need for independent pathogen detection and host analysis, the requirement for additional experimental verification in host immune gene association studies, the high proportion of host DNA in clinical samples that is not utilized, unsatisfactory detection accuracy in low signal-to-noise ratio scenarios, and limitations in drug resistance/virulence detection. These issues result in long diagnosis times, high costs, and frequent adjustments to treatment plans.

Method used

By integrating host genotype variation feature analysis into the metagenomic pathogen detection process, and using a self-developed bioinformatics noise reduction analysis method and system, combined with deep learning algorithms, pathogenic pathogens and their drug resistance genes and virulence factors are identified, and host genotype analysis is performed to achieve simultaneous detection of pathogen and host information.

Benefits of technology

It improves the efficiency and accuracy of pathogen and host gene detection, reduces the false positive rate, rapidly identifies drug resistance genes and virulence factors, clarifies host genotype variation characteristics, provides a basis for the diagnosis of infectious diseases and personalized medicine, and improves diagnostic efficiency and treatment effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913649A_ABST
    Figure CN120913649A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of bioinformatics, and discloses a conjoint analysis method for pathogen recognition and patient genotype characteristics. The method comprises the following steps: firstly, carrying out metagenome sequencing on a positive clinical sample, obtaining original sequencing data, respectively carrying out quality control treatment and data classification on the original sequencing data, identifying a pathogenic pathogen DNA sequence and a host DNA sequence, then carrying out noise reduction treatment on the pathogenic pathogen DNA sequence, identifying the pathogenic pathogen and drug-resistant genes and virulence factors carried by the pathogenic pathogen, and identifying the drug-resistant genes and virulence factors carried by the pathogenic pathogen. Meanwhile, host genotype analysis is carried out on the basis of host DNA sequence characteristics, and finally composition information of pathogenic pathogens, a list of drug-resistant genes and virulence factors and host genotype variation characteristics are obtained. According to the method, a host gene variation analysis process and a conventional metagenome pathogenic pathogen detection analysis process are integrated, so that comprehensive analysis of host and pathogenic pathogen characteristics is realized, and the accuracy and efficiency of clinical diagnosis and treatment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of bioinformatics, and particularly relates to a pathogen identification and patient genotype feature combined analysis method. BACKGROUND

[0002] Pathogenic pathogen detection and host interaction research has become the core direction of infectious disease diagnosis and treatment. The current mainstream technical system mainly includes the following modules: metagenomic sequencing analysis: relying on bioinformatics species annotation software Kraken2, Bracken2 and genomic cloud platform Blast tools, detecting bacteria, viruses, fungi and other pathogenic pathogens; drug resistance / virulence factor detection: using Abricate tools based on sequence alignment method in comprehensive drug resistance database CARD and virulence database VFDB for screening and filtering, and combining deep learning tools DeepARG and VirulenHunter to determine the host information associated with drug resistance genes and virulence factors; host genome analysis: using GATK variant detection process, DRAGEN hardware and software jointly accelerated variant detection tools and other standard processes for variant detection; data noise reduction processing: using quality control tools such as DeBlur, DADA2 algorithm to filter low-quality signals.

[0003] However, there are still many technical bottlenecks in pathogenic pathogen detection and host interaction research: first, pathogen detection and host analysis need to be carried out independently; host immune genes (such as HLA typing) and pathogen susceptibility association research need additional experimental verification; the proportion of host DNA in clinical samples is often more than 90%, but the existing process directly discards this part of data. In addition, the existing method performs poorly in low signal-to-noise ratio scenarios, such as traditional filtering threshold (such as 0.001% abundance) leading to high false positive rate and missed detection rate, and sequencing errors and microbial similar sequences causing cross comparison, resulting in unsatisfactory detection result accuracy. Finally, there are obvious limitations in drug resistance / virulence detection: for example, the host genome analysis process (such as GATK) is not optimized for infection scenarios, and the selection mutation under pathogen pressure and insufficient coverage of immune-related genes are ignored.

[0004] Based on the retrospective study, complex infection cases need to obtain pathogen and host information simultaneously, and the existing technical system has defects such as longer average diagnosis time, higher detection cost and more treatment scheme adjustment times, so it is necessary to analyze the host-pathogen interaction mechanism in infectious diseases and improve the diagnosis efficiency of infectious diseases. SUMMARY

[0005] To solve the above problems, the present application provides a pathogen identification and patient (i.e. host) genotype feature combined analysis method, which integrates host genotype variation feature analysis into the traditional macro-genome pathogenic pathogen detection analysis process, comprehensively depicts the pathogenic pathogen composition characteristics and patient genotype information, and provides a basis for predicting the development trend, severity and personalized medical treatment of the disease. Specifically, the present application provides the following specific technical solutions:

[0006] In the first aspect, the present application provides a pathogen identification and host genotype feature combined analysis method, which comprises the following steps:

[0007] Sequencing: macro-genome sequencing of positive clinical samples to obtain raw sequencing data;

[0008] Quality control: quality control processing of raw sequencing data, including: removing low-quality sequences, filtering low-complexity sequences and trimming adapters;

[0009] Data classification: classifying the data after quality control processing to identify pathogenic pathogen DNA sequences and host DNA sequences;

[0010] Information processing: after noise reduction processing of the pathogenic pathogen DNA sequences, identifying the pathogenic pathogen and its carried drug resistance genes and virulence factors, and performing host genotype analysis based on the host DNA sequence characteristics;

[0011] Result output: outputting the composition information of the pathogenic pathogen, the list of drug resistance genes and virulence factors, and the host genotype variation characteristics.

[0012] Further, the combined analysis method is suitable for infectious diseases.

[0013] Further, in the quality control step:

[0014] Remove low-quality sequences: after quality assessment using FastQC, filter using Trimmomatic v0.39;

[0015] Filter low-complexity sequences: identify and filter simple repeat sequences by DustMasker algorithm;

[0016] Adapter trimming: use Cutadapt v3.7 to match and remove universal adapter sequences.

[0017] Further, in the data classification step, algorithm and database matching technology are used to identify the types of pathogenic pathogens and their relative abundance.

[0018] Further, in the noise reduction processing of the information processing step, the undirected maximal subgraph noise reduction algorithm is used to realize data noise reduction.

[0019] Further, in the drug resistance gene and virulence factor identification of the information processing step, the identified pathogenic pathogen DNA sequence is compared to the drug resistance gene and virulence factor database, and the drug resistance gene and virulence factor are screened out by global alignment and deep learning algorithm.

[0020] Further, the screening of drug resistance genes includes: comparing the identified pathogenic pathogen DNA sequence with the public comprehensive drug resistance database to screen out a list of potential drug resistance genes;

[0021] The screening of virulence factors includes: comparing the identified pathogenic pathogen DNA sequence with the public comprehensive virulence factor database to screen out a list of potential virulence factors;

[0022] Then, BLASTn is used for global sequence alignment to screen out non-redundant sequences with similarity and consistency of >80%; the candidate target genes meeting the conditions are manually verified; DeepARG is used to predict potential drug resistance genes and their associated pathogenic hosts, and VirulenHunter is used to identify potential virulence factors and their associated pathogenic hosts; the screening results and deep learning prediction results are combined to determine the final list of drug resistance genes and virulence factors.

[0023] Further, in the host genotype analysis step of information processing, host genotype analysis is realized by combining read alignment mapping and hidden Markov algorithm.

[0024] Further, the host genotype analysis includes genomic difference feature analysis, and the genomic difference features include: single base mutation, deletion, insertion, fragment duplication, insertion duplication, transversion, and copy number variation.

[0025] In a second aspect, the present application provides a pathogen identification and patient genotype feature combined analysis system, which comprises:

[0026] A sequencing module is used to perform metagenomic sequencing on positive clinical samples to obtain raw sequencing data.

[0027] A quality control module is used to perform quality control processing on the raw sequencing data, including: removing low-quality sequences, filtering low-complexity sequences, and trimming adapters.

[0028] A data classification module is used to classify the data after quality control processing to identify pathogenic pathogen DNA sequences and host DNA sequences.

[0029] An information processing module is used to identify pathogenic pathogens and their carried drug resistance genes and virulence factors after noise reduction processing of the pathogenic pathogen DNA sequences, and to perform host genotype analysis based on host DNA sequence characteristics.

[0030] A result output module is configured to output a list of the composition information of the pathogenic pathogens, the drug resistance genes, and the virulence factors; and the host genotype variation characteristics.

[0031] In a third aspect, the present application provides a computer storage medium storing a computer program, the computer program comprising program instructions, which when executed by a processor, perform the joint analysis method described above.

[0032] Compared with the prior art, the joint analysis method of pathogen identification and patient genotype characteristics has the following beneficial effects:

[0033] The present application involves the following key innovations:

[0034] 1) By improving the metagenomic sequencing analysis process, the pathogen and host gene information are detected simultaneously;

[0035] 2) The bioinformatics noise reduction analysis method and system based on hidden subgroups are self-developed, which can retain true positives while reducing more than 95% of false positive detections;

[0036] 3) The self-developed drug resistance gene and virulence factor detection algorithm can quickly identify the drug resistance genes and virulence factors carried in the sample, with an accuracy of more than 90%;

[0037] 4) The host genotype analysis process is integrated, which can determine the host genotype variation characteristics;

[0038] On the basis of the basic pathogenic pathogen metagenomic sequencing analysis, the present application adds data noise reduction, drug resistance gene and virulence factor gene identification, and host genotype characteristic analysis, which can more comprehensively capture the composition information of pathogenic pathogens, the list of drug resistance genes and virulence factors, and the host genotype variation characteristics of severe clinical samples. The present application provides a basis for predicting the development trend, severity, and personalized medicine of infectious diseases. The method provided by the present application not only helps to improve the diagnosis efficiency of infectious diseases, but also helps to understand the host-pathogen interaction mechanism in infectious diseases, and has important clinical and scientific research value. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 The figure is a flowchart of the joint analysis method of the present application. DETAILED DESCRIPTION

[0040] The technical solutions of the present application will be described clearly and completely below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0041] The pathogen identification and patient genotype feature analysis method provided in the embodiments of the present application, a flowchart of the method is shown in Figure 1 , and the method basically comprises the following steps:

[0042] Sequencing: performing metagenomic sequencing on the positive clinical sample to obtain raw sequencing data;

[0043] Quality control: performing quality control processing on the raw sequencing data, including removing low-quality sequences, filtering low-complexity sequences, and trimming adapters;

[0044] Data classification: classifying the data after quality control processing to identify pathogenic pathogen DNA sequences and host DNA sequences;

[0045] Information processing: after noise reduction processing of the pathogenic pathogen DNA sequences, identifying the pathogenic pathogen and the drug resistance genes and virulence factors carried by the pathogenic pathogen, and performing host genotype analysis based on the host DNA sequence characteristics;

[0046] Result output: outputting the composition information of the pathogenic pathogen, the list of drug resistance genes and virulence factors, and the host genotype variation characteristics.

[0047] In a possible implementation, the joint analysis method is applicable to infectious diseases.

[0048] In a possible implementation, in the quality control step:

[0049] Removing low-quality sequences: after quality assessment using FastQC, filtering is performed using Trimmomatic v0.39;

[0050] Filtering low-complexity sequences: identifying and filtering simple repeat sequences through the DustMasker algorithm;

[0051] Adapter trimming: using Cutadapt v3.7 to match and cut off universal adapter sequences.

[0052] In a possible implementation, in the data classification step, algorithm and database matching techniques are used to identify the types of pathogenic pathogens and their relative abundances.

[0053] In a possible implementation, in the noise reduction processing of the information processing step, a noise reduction algorithm of an undirected maximal subgraph is used to realize data noise reduction.

[0054] In a possible implementation, in the identification of drug resistance genes and virulence factors in the information processing step, the identified pathogenic pathogen DNA sequences are compared to a drug resistance gene and virulence factor database, and global alignment and deep learning algorithms are used to screen out drug resistance genes and virulence factors.

[0055] In a possible implementation, the screening of the drug resistance gene comprises: comparing the identified pathogenic pathogen DNA sequence with a public comprehensive drug resistance database to screen out a list of potential drug resistance genes;

[0056] The screening of the virulence factor comprises: comparing the identified pathogenic pathogen DNA sequence with a public comprehensive virulence factor database to screen out a list of potential virulence factors;

[0057] Then, BLASTn is used for global sequence alignment to screen out non-redundant sequences with similarity and consistency of more than 80%; the candidate target genes meeting the conditions are manually verified; DeepARG is used to predict potential drug resistance genes and associated pathogenic hosts, and VirulenHunter is used to identify potential virulence factors and associated pathogenic hosts; and the screening results and the deep learning prediction results are comprehensively compared to determine the final list of drug resistance genes and virulence factors.

[0058] In a possible implementation, in the host genotype analysis step of the information processing, the host genotype analysis is realized by combining read alignment mapping and a hidden Markov algorithm.

[0059] In a possible implementation, the host genotype analysis comprises genomic difference feature analysis, and the genomic difference features comprise: single base mutation, deletion, insertion, fragment duplication, insertion duplication, transversion, and copy number variation.

[0060] The pathogen identification and joint analysis system of patient genotype features provided in the embodiments of the present application basically comprises the following functional modules:

[0061] A sequencing module is configured to perform metagenomic sequencing on a positive clinical sample to obtain raw sequencing data.

[0062] A quality control module is configured to perform quality control processing on the raw sequencing data, including: removing low-quality sequences, filtering low-complexity sequences, and trimming adapters.

[0063] A data classification module is configured to classify the data after quality control processing to identify pathogenic pathogen DNA sequences and host DNA sequences.

[0064] An information processing module is configured to identify pathogenic pathogens and the drug resistance genes and virulence factors carried by the pathogenic pathogens after noise reduction processing of the pathogenic pathogen DNA sequences, and perform host genotype analysis based on host DNA sequence features.

[0065] A result output module is configured to output the composition information of the pathogenic pathogens, the list of drug resistance genes and virulence factors, and host genotype variation features.

[0066] The present application will be described in detail below with reference to specific embodiments.

[0067] Example 1

[0068] This example describes the specific application of the combined analysis method of pathogen identification and host genotype characteristics in respiratory tract infection.

[0069] 1) Sequencing: Perform metagenomic sequencing on positive clinical samples to obtain raw sequencing data.

[0070] Use positive samples of respiratory tract infection from a hospital (N=112 cases), sample types: sputum, bronchoalveolar lavage fluid (BALF). Use QIAamp DNA Mini Kit to extract total DNA, quality control requirements: DNA concentration ≥0.5 ng / μL, fragment length >500 bp. Then use Illumina NovaSeq 6000 platform for PE150 sequencing to obtain raw sequencing data. Raw data format: FASTQ file, average Q30>85%, data volume fluctuation range ±15%, data capacity 20G / sample.

[0071] 2) Quality control: Perform quality control processing on raw sequencing data, including: removing low-quality sequences, filtering low-complexity sequences, and trimming adapters.

[0072] (1) Remove low-quality sequences

[0073] Use Trimmomatic v0.39 to remove bases with quality value <Q20 at the beginning and end, sliding window width 4bp, truncate when average quality <Q20, and retain reads with length ≥50bp.

[0074] (2) Filter low-complexity sequences

[0075] Use DustMasker to filter low-complexity sequences with a complexity threshold of 0.7.

[0076] (3) Adapter trimming

[0077] Use Cutadapt v3.7 for adapter trimming, allowing 10% adapter matching error, and retain reads with length ≥30bp after trimming.

[0078] 3) Data classification: Classify the data after quality control processing to identify pathogenic pathogen DNA sequences and host DNA sequences.

[0079] Parallel process the FASTQ file after quality control (average 18±3 Gb / sample, Q30>85%), including:

[0080] (1) Pathogenic pathogen identification channel

[0081] Kraken2 (v2.1.2) combined with self-built clinical pathogen database (containing 5821 microbial genomes); parameter setting: --confidence 0.5 --memory-mapping; species abundance identification was performed using the matching tool Bracken (v2.7). The classification of pathogenic pathogens took 45±5 min / sample.

[0082] (2) Host sequence identification channel

[0083] BWA-MEM (v0.7.17) was aligned with the hg38 reference genome, and samtools (v1.15) was used to extract reads that were not aligned to the host. The separation of host sequences took 30±3 min / sample.

[0084] After parallel processing, the disputed sequences were reviewed using Blastn (v2.13.0) to ensure the accuracy of the classification.

[0085] Classification results:

[0086] (1) Identification results of pathogenic pathogens

[0087] Sensitivity: 92.3% (112 clinical verifications), typical detection: Klebsiella pneumoniae, methicillin-resistant Staphylococcus aureus, Candida albicans.

[0088] (2) Host sequence analysis

[0089] Host DNA proportion range: 15-60% (difference between sputum and BALF); SNP detection amount: average 12000 sites / sample (covering key genes such as HLA and TLR).

[0090] 4) Information processing: After noise reduction processing of pathogenic pathogen DNA sequences, perform list screening of drug resistance genes and virulence factors, and host genotype analysis.

[0091] (1) Noise reduction processing

[0092] Use CD-HIT clustering (parameters: -c 0.9-n 5) to group the original reads by 90% similarity, generate sequence clusters; take clusters as nodes, establish undirected edges between clusters with edit distance ≤3bp, form a similarity network graph; apply Bron-Kerbosch algorithm to identify all maximal complete subgraphs (cliques) in the graph, requiring the minimum size of the clique to be ≥5 nodes, and remove isolated nodes and small-scale connected components.

[0093] (2) Drug resistance gene detection

[0094] Consensus sequences after denoising were Blastn aligned to CARD database, and drug-resistant genes with matching degree ≥85% and coverage ≥80% were screened. Known drug-resistant mutations (such as Mycobacterium tuberculosis rpoB S450L) were identified by SNP calling.

[0095] (3) Virulence factor identification

[0096] The VFDB core dataset was aligned using ABRicate, and virulence genes with E-value <1e-10 were retained, and the virulence factor type was labeled.

[0097] (4) Sequence alignment and variation detection

[0098] Host sequences were aligned to the GRCh38 reference genome by BWA-MEM, and GATK aplotypeCaller was used to detect SNPs / InDels, with filtering criteria: QUAL>30, DP>10, population frequency <1%. Then ANNOVAR was used to annotate drug metabolism related genes, and the PharmGKB database was used to predict the phenotype effect.

[0099] 5) Result output: Output the composition information of pathogenic pathogens, drug-resistant genes and virulence factors; host genotype variation characteristics.

[0100] (1) Pathogenic pathogen composition information (Top3 high frequency detection)

[0101] Klebsiella pneumoniae: detection rate 28.6% (32 / 112); subtype distribution: ST11-KL64 (72%), ST23-KL1 (19%);

[0102] Methicillin-resistant Staphylococcus aureus: detection rate 21.4% (24 / 112); SCCmec typing: type IV (58%), type II (33%);

[0103] Candida albicans: detection rate 15.2% (17 / 112); proportion of drug-resistant strains: fluconazole resistance rate 41%.

[0104] (2) Drug-resistant genes and virulence factors

[0105] Klebsiella pneumoniae blaKPC-2 (89%), blaNDM-1 (11%), rmpA2 (67%), aerobactin (92%);

[0106] Pseudomonas aeruginosa blaVIM (76%), mexR mutation (54%), exoU (38%), lasR mutation (29%).

[0107] (3) Host genotype variation characteristics

[0108] Immune-related variants: TLR2 rs5743708 AG type: 34.8%, significantly associated with Gram-positive bacterial infection;

[0109] IFNG rs2430561 TT type: 22.3%, associated with increased risk of delayed fungal clearance;

[0110] Drug metabolism genes: CYP2C19*2 carriers: 18.7%, associated with risk of voriconazole toxicity;

[0111] FUT2 rs601338 AA: 29.5%, associated with increased susceptibility to norovirus.

[0112] The joint analysis method and results provided by the embodiment not only improve the detection efficiency and accuracy of pathogenic pathogens and host genes, but also provide precise medication guidance, personalized treatment plans, disease development trend prediction, and medical resource allocation optimization for clinical treatment, and have important clinical and scientific research value. For example, for IFNG rs2430561 TT genotype, when Candida albicans infection occurs, consider changing the clinical sample of fluconazole-resistant strain to echinocandin class, and prolong the treatment course. For IFNG rs2430561 TT genotype, consider KPC-producing Klebsiella pneumoniae prevention and control, such as contact isolation + environmental disinfection, and adjust the treatment plan, such as prolonging the infusion of meropenem.

[0113] Embodiment 2

[0114] The embodiment describes a joint analysis system of pathogen identification and patient genotype characteristics.

[0115] The joint analysis system of pathogen identification and patient genotype characteristics comprises the following functional modules:

[0116] A sequencing module for performing metagenomic sequencing on positive clinical samples to obtain raw sequencing data;

[0117] A quality control module for quality control processing of the raw sequencing data, including: removing low-quality sequences, filtering low-complexity sequences, and trimming adapters;

[0118] A data classification module for classifying the data after quality control processing to identify pathogenic pathogen DNA sequences and host DNA sequences;

[0119] An information processing module for identifying pathogenic pathogens and their carried drug resistance genes and virulence factors after noise reduction processing of the pathogenic pathogen DNA sequences, and performing host genotype analysis based on the characteristics of the host DNA sequences;

[0120] ​​​​​​​​​​​​The result output module is configured to output the composition information of the pathogenic pathogen, the list of drug resistance genes and virulence factors, and the host genotype variation characteristics.

[0121] The above-described embodiments are only some of the embodiments of the present application, but not all the embodiments. The detailed description of the embodiments of the present application is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the application. All other embodiments obtained by those of ordinary skill in the art based on the relevant deductions and substitutions made under the concept of the present application, without making creative efforts, fall within the scope of protection of the present application.

Claims

1. A method for combined analysis of pathogen recognition and host genotype characteristics, characterized in that, The method comprises the following steps: sequencing: performing metagenomic sequencing on positive clinical samples to obtain raw sequencing data; quality control: performing quality control processing on the raw sequencing data, including: removing low-quality sequences, filtering low-complexity sequences, and trimming adapters; data classification: classifying the data after quality control processing to identify pathogenic pathogen DNA sequences and host DNA sequences; information processing: after noise reduction processing of the pathogenic pathogen DNA sequences, identifying pathogenic pathogens and the drug resistance genes and virulence factors carried by the pathogens, and performing host genotype analysis based on the characteristics of the host DNA sequences; result output: outputting the composition information of the pathogenic pathogens, the list of drug resistance genes, and the virulence factors; and the host genotype variation characteristics.

2. The joint analysis method according to claim 1, characterized in that, The combined analysis method is suitable for infectious diseases.

3. The joint analysis method of claim 1, wherein, In the quality control step: removing low-quality sequences: after quality assessment using FastQC, filtering is performed using Trimmomatic v0.39; filtering low-complexity sequences: identifying and filtering simple repeat sequences through the DustMasker algorithm; adapter trimming: using Cutadapt v3.7 to match and remove universal adapter sequences.

4. The joint analysis method of claim 1, wherein, In the data classification step, algorithm and database matching techniques are used to identify the types of pathogenic pathogens and their relative abundances.

5. The joint analysis method of claim 1, wherein, In the noise reduction processing of the information processing step, the undirected maximal subgraph noise reduction algorithm is used to achieve data noise reduction.

6. The joint analysis method of claim 1, wherein, In the identification of drug resistance genes and virulence factors in the information processing step, the identified pathogenic pathogen DNA sequences are compared to the drug resistance gene and virulence factor database, and global alignment and deep learning algorithms are used to screen drug resistance genes and virulence factors.

7. The combined analysis method according to claim 6, characterized in that: the screening of drug resistance genes includes: comparing the identified pathogenic pathogen DNA sequences with a public comprehensive drug resistance database to screen a list of potential drug resistance genes; the screening of virulence factors includes: comparing the identified pathogenic pathogen DNA sequences with a public comprehensive virulence factor database to screen a list of potential virulence factors; then performing global sequence alignment using BLASTn to screen non-redundant sequences with similarity and consistency greater than 80%; manually verifying the candidate target genes that meet the conditions; using DeepARG to predict potential drug resistance genes and their associated pathogenic hosts, and using VirulenHunter to identify potential virulence factors and their associated pathogenic hosts; and combining the screening results and the deep learning prediction results to determine the final list of drug resistance genes and virulence factors.

8. The joint analysis method of claim 1, wherein, In the host genotype analysis step of the information processing, read alignment mapping and hidden Markov algorithm are combined to achieve host genotype analysis.

9. The joint analysis method of claim 8, wherein, Host genotype analysis includes genomic difference characteristic analysis, and the genomic difference characteristics include: single base mutation, deletion, insertion, fragment duplication, insertion duplication, transversion, and copy number variation.

10. A system for combined analysis of pathogen identification and patient genotype characteristics, comprising: The system comprises: a sequencing module for performing metagenomic sequencing on positive clinical samples to obtain raw sequencing data; a quality control module for performing quality control processing on the raw sequencing data, including: removing low-quality sequences, filtering low-complexity sequences, and trimming adapters; a data classification module, configured to classify the data after the quality control processing, and identify pathogenic pathogen DNA sequences and host DNA sequences; an information processing module, configured to identify pathogenic pathogens and the drug resistance genes and virulence factors carried by the pathogenic pathogens after noise reduction processing of the pathogenic pathogen DNA sequences, and perform host genotype analysis based on characteristics of the host DNA sequences; a result output module, configured to output a list of composition information, drug resistance genes, and virulence factors of the pathogenic pathogens; host genotype variation characteristics.

Citation Information

Patent Citations

  • Application of rs1143634 polymorphism in screening of patients suffering from severe fever with thrombocytopenia syndrome

    CN106222289A

  • HBV infection-related gene haplotype and application thereof

    CN108070647A

  • Method and device for detecting pathogenic microorganisms based on metagenomics

    CN113744807A

  • Sequencing method for simultaneously detecting expression quantities of pathogenic bacteria and host genes and application of sequencing method in diagnosis and prognosis of bacterial meningitis

    CN115537462A

  • Hidden subgroup-based raw signal noise reduction analysis method and system

    CN115719614A