Pathogenic fungus generic genome analysis method, device and equipment and readable storage medium
By cleaning and clustering the sequencing genomic data of pathogenic fungi, the problem of low analysis accuracy in existing technologies has been solved, achieving higher analysis accuracy and visualization of genetic structure.
Patent Information
- Application Number
- CN202511635071.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-03-03
AI Technical Summary
Existing pan-genome analysis methods for pathogenic fungi suffer from low accuracy.
By acquiring the sequencing genome data to be analyzed, removing contig sequences from humans and bacteria, performing gene prediction and pan-genome clustering analysis, using gff protein sequence files for clustering, and constructing a phylogenetic tree to display the genetic developmental structure.
It improves the accuracy and completeness of pan-genome analysis of pathogenic fungi, ensures the reliability and traceability of analysis results, and supports intuitive genetic tracing analysis.
Smart Images

Figure CN121601026A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of genome sequencing and analysis technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for pan-genome analysis of pathogenic fungi. Background Technology
[0002] Pathogenic fungi, as an important class of pathogenic microorganisms, pose a serious threat to human health. With the increasing number of immunocompromised individuals, the overuse of antifungal drugs, and the impact of global climate change, the incidence of fungal infections is rising annually, and their treatment is becoming increasingly difficult. Therefore, in-depth research into pathogenic fungi that cause human diseases, revealing their genetic diversity, pathogenic mechanisms, and drug resistance characteristics, is of great significance for disease prevention, diagnosis, and treatment. By constructing and analyzing pan-genomes of pathogenic fungi that cause human diseases, we can comprehensively understand the genetic diversity of these species, discover new pathogenic genes and drug resistance genes, and provide a theoretical basis for new drug development and disease treatment.
[0003] However, current pan-genome analysis methods for pathogenic fungi suffer from low accuracy. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for pan-genome analysis of pathogenic fungi that can improve the accuracy of analysis, in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a method for pan-genome analysis of pathogenic fungi, including:
[0006] Acquire the sequencing genomic data to be analyzed; the sequencing genomic data includes multiple contig sequences;
[0007] The contig sequences of humans and bacteria were removed from the sequenced genome data to obtain the cleaned sequenced genome data;
[0008] Gene prediction is performed on the cleaned sequencing genomic data to obtain the corresponding gff protein sequence file;
[0009] A pan-genome clustering analysis was performed on the gff protein sequence file to obtain the clustering analysis results of the sequenced genome data.
[0010] In one embodiment, gene prediction is performed on the cleaned sequencing genomic data, including:
[0011] To determine the completeness of the cleaned sequencing genomic data;
[0012] Gene prediction is performed on the cleaned sequencing genomic data when the integrity level is greater than or equal to a preset integrity level threshold.
[0013] In one embodiment, gene prediction is performed on the cleaned sequencing genomic data, including:
[0014] When there is a reference sequence corresponding to the sequencing genome data, the prediction model to be trained is trained using the reference sequence, and the cleaned sequencing genome data is input into the trained prediction model for gene prediction.
[0015] In cases where no corresponding reference sequence is available for the sequenced genome data, statistical analysis is performed on the cleaned sequenced genome data to predict gene patterns based on the statistically obtained patterns.
[0016] In one exemplary embodiment, pan-genome clustering analysis is performed on the gff protein sequence file to obtain the clustering analysis results of the sequenced genome data, including:
[0017] Obtain the similarity between protein sequences in a gff protein sequence file;
[0018] Based on the protein sequences and the degree of similarity between them, a protein sequence similarity map is constructed.
[0019] Clustering of protein sequence similarity maps yields multiple homologous gene groups; these groups include orthologous genes and paralogous genes.
[0020] Homologous gene groups with a frequency equal to the first value are classified as unique genes; homologous gene groups with a frequency greater than the first value and less than or equal to the second value are classified as accessory genes; and homologous gene groups with a frequency greater than the second value are classified as core genes; the first value is less than the second value.
[0021] In one embodiment, for any homologous gene group, the homologous gene group is any one of specific genes, accessory genes, and core genes, and the homologous gene group includes multiple protein sequences; the method further includes:
[0022] Protein annotation was performed on each protein sequence to obtain the protein annotation results of homologous gene groups;
[0023] Obtain the coverage of each protein annotation result;
[0024] Based on the coverage level, homologous gene groups were labeled using the protein annotation results.
[0025] In one embodiment, the method further includes:
[0026] Obtain metadata for each sequenced genome data; metadata includes time, location, isolation source, host, and molecular typing;
[0027] Based on the metadata of core genes and each sequenced genome data, a phylogenetic tree is drawn to show the genetic and developmental structure of each sequenced genome data.
[0028] Secondly, this application also provides a pan-genome analysis device for pathogenic fungi, comprising:
[0029] Acquire the sequencing genomic data to be analyzed; the sequencing genomic data includes multiple contig sequences;
[0030] The contig sequences of humans and bacteria were removed from the sequenced genome data to obtain the cleaned sequenced genome data;
[0031] Gene prediction is performed on the cleaned sequencing genomic data to obtain the corresponding gff protein sequence file;
[0032] A pan-genome clustering analysis was performed on the gff protein sequence file to obtain the clustering analysis results of the sequenced genome data.
[0033] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0034] Acquire the sequencing genomic data to be analyzed; the sequencing genomic data includes multiple contig sequences;
[0035] The contig sequences of humans and bacteria were removed from the sequenced genome data to obtain the cleaned sequenced genome data;
[0036] Gene prediction is performed on the cleaned sequencing genomic data to obtain the corresponding gff protein sequence file;
[0037] A pan-genome clustering analysis was performed on the gff protein sequence file to obtain the clustering analysis results of the sequenced genome data. Fourthly, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, performs the following steps:
[0038] Acquire the sequencing genomic data to be analyzed; the sequencing genomic data includes multiple contig sequences;
[0039] The contig sequences of humans and bacteria were removed from the sequenced genome data to obtain the cleaned sequenced genome data;
[0040] Gene prediction is performed on the cleaned sequencing genomic data to obtain the corresponding gff protein sequence file;
[0041] A pan-genome clustering analysis was performed on the gff protein sequence file to obtain the clustering analysis results of the sequenced genome data.
[0042] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0043] Acquire the sequencing genomic data to be analyzed; the sequencing genomic data includes multiple contig sequences;
[0044] The contig sequences of humans and bacteria were removed from the sequenced genome data to obtain the cleaned sequenced genome data;
[0045] Gene prediction is performed on the cleaned sequencing genomic data to obtain the corresponding gff protein sequence file;
[0046] A pan-genome clustering analysis was performed on the gff protein sequence file to obtain the clustering analysis results of the sequenced genome data.
[0047] The aforementioned method, apparatus, computer equipment, computer-readable storage medium, and computer program product for pan-genome analysis of pathogenic fungi acquire sequencing genomic data containing multiple contig sequences, remove human and bacterial contig sequences from the sequencing genomic data to obtain cleaned sequencing genomic data, perform gene prediction on the cleaned sequencing genomic data to obtain the corresponding GFF protein sequence files, and finally perform pan-genome cluster analysis on the GFF protein sequence files to obtain the cluster analysis results of the sequencing genomic data. By cleaning the sequencing genomic data to be analyzed from human and bacterial sequences, retaining only fungal gene sequences, and further predicting the cleaned data to obtain the corresponding GFF protein sequence files, cluster analysis is performed based on the GFF protein sequence files, thereby achieving pan-genome analysis of fungi and improving the accuracy and completeness of the analysis. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1This is a diagram illustrating the application environment of a pan-genome analysis method for pathogenic fungi in one embodiment;
[0050] Figure 2 This is a flowchart illustrating a pan-genome analysis method for pathogenic fungi in one embodiment;
[0051] Figure 3 This is a flowchart illustrating the pan-genome analysis method for pathogenic fungi in another embodiment;
[0052] Figure 4 This is a structural block diagram of a pan-genome analysis device for pathogenic fungi in one embodiment;
[0053] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0055] The pan-genome analysis method for pathogenic fungi provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Server 104 obtains the sequencing genome data to be analyzed from terminal 102. The sequencing genome data includes multiple contig sequences. Human and bacterial contig sequences are removed from the sequencing genome data to obtain cleaned sequencing genome data. Gene prediction is performed on the cleaned sequencing genome data to obtain the corresponding GFF protein sequence file. Finally, pan-genome clustering analysis is performed on the GFF protein sequence file to obtain the clustering analysis results of the sequencing genome data. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Headset devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0056] In one exemplary embodiment, such as Figure 2 As shown, a pan-genome analysis method for pathogenic fungi is provided, which can be applied to... Figure 1 The following steps, S201 to S204, are used to illustrate 104 examples of servers.
[0057] Step S201: Obtain the sequencing genomic data to be analyzed; the sequencing genomic data includes multiple contig sequences.
[0058] Among them, sequencing genomic data can be understood as an approximate representation of the genome, usually in the form of contigs, scaffolds, and finally pseudo-chromosome level assemblies. A contig sequence can be understood as a long sequence formed by overlapping and splicing several read sequences.
[0059] For example, server 104 acquires the original sample, performs sequencing reads on the original sample to form multiple read sequences, and then splices the multiple read sequences to form sequencing genomic data containing a longer contig sequence.
[0060] Step S202: Remove human and bacterial contig sequences from the sequencing genome data to obtain cleaned sequencing genome data.
[0061] Optionally, server 104 compares the contig sequences in the sequencing genome data with human and self-built bacterial contig sequence reference libraries, removes the human and bacterial contig sequences from the sequencing genome data, and thus obtains cleaned sequencing genome data.
[0062] Step S203: Perform gene prediction on the cleaned sequencing genomic data to obtain the gff protein sequence file corresponding to the cleaned sequencing genomic data.
[0063] Among them, ggf (General Feature Format) protein sequence files can be understood as location information used to describe various biological features (such as genes, exons, CDS, etc.) on a DNA sequence.
[0064] For example, server 104 obtains the integrity of the cleaned sequencing genome data. If the integrity is greater than or equal to a preset integrity threshold, server 104 performs gene prediction on the cleaned sequencing genome data to obtain the gff protein sequence file corresponding to the cleaned sequencing genome data; if the integrity is less than the integrity threshold, the process is terminated.
[0065] Step S204: Perform pan-genome cluster analysis on the gff protein sequence file to obtain the cluster analysis results of the sequenced genome data.
[0066] Pan-genome clustering analysis can be understood as classifying genes based on their frequency of occurrence.
[0067] Optionally, server 104 calculates the similarity between each protein sequence in the gff protein sequence file, constructs a protein sequence similarity map based on the similarity between each protein sequence and each protein sequence, then clusters the protein sequence similarity map to obtain multiple homologous gene groups, and then divides each homologous gene group into three categories: core genes, accessory genes, and unique genes based on the frequency of occurrence of homologous gene groups.
[0068] In the aforementioned method for pan-genome analysis of pathogenic fungi, the sequencing genome data containing multiple contig sequences is acquired. Human and bacterial contig sequences are removed from the sequenced genome data to obtain cleaned sequencing genome data. Gene prediction is performed on the cleaned sequencing genome data to obtain the corresponding GFF protein sequence files. Finally, pan-genome cluster analysis is performed on the GFF protein sequence files to obtain the cluster analysis results of the sequencing genome data. By cleaning the sequenced genome to be analyzed from human and bacterial sequences, retaining only fungal gene sequences, and further predicting the cleaned data to obtain the corresponding GFF protein sequence files, cluster analysis is performed based on the GFF protein sequence files, thereby achieving pan-genome analysis of fungi and improving the accuracy and completeness of the analysis.
[0069] In one embodiment, gene prediction on cleaned sequencing genomic data includes: obtaining the integrity level of the cleaned sequencing genomic data; and performing gene prediction on the cleaned sequencing genomic data if the integrity level is greater than or equal to a preset integrity level threshold.
[0070] In this context, completeness can be understood as a quantitative representation of the number and types of genes contained in the cleaned sequencing genomic data.
[0071] For example, server 104 obtains the integrity level of the cleaned sequencing genome data and compares it with a preset integrity level threshold. If the integrity level is greater than or equal to the threshold, gene prediction is performed on the cleaned sequencing genome data; if the integrity level is less than the threshold, the process terminates. By performing quality control on the cleaned sequencing genome data, i.e., setting controls on gene integrity, the subsequent gene analysis is ensured to be meaningful, unnecessary resource consumption is avoided, thereby reducing analysis costs and improving analysis accuracy.
[0072] In one embodiment, gene prediction is performed on the cleaned sequencing genome data, including: when the sequencing genome data corresponds to a reference sequence, the prediction model to be trained is trained using the reference sequence, and the cleaned sequencing genome data is input into the trained prediction model for gene prediction; when the sequencing genome data does not correspond to a reference sequence, the cleaned sequencing genome data is statistically analyzed, and gene prediction is performed based on the statistically obtained change patterns.
[0073] The reference sequence can be understood as the sequence information that a gene would present under normal, ideal conditions.
[0074] Optionally, when the sequenced genome data corresponds to a reference sequence, the server 104 uses the reference sequence to train the prediction model to be trained, thereby obtaining a trained prediction model, and then inputs the cleaned sequenced genome data into the trained prediction model for genome prediction. When the sequenced genome data does not correspond to a reference sequence, the server 104 performs statistical analysis on the cleaned sequenced genome data and performs gene prediction based on the statistically obtained change patterns. By selecting the appropriate method to perform gene prediction under different circumstances, the timeliness of gene prediction is ensured and the accuracy of gene prediction is improved.
[0075] In an exemplary embodiment, pan-genome clustering analysis is performed on the gff protein sequence file to obtain the clustering analysis results of the sequenced genome data, including: obtaining the similarity between each protein sequence in the gff protein sequence file; constructing a protein sequence similarity map based on each protein sequence and the similarity between protein sequences; clustering the protein sequence similarity map to obtain multiple homologous gene groups; the homologous gene groups include orthologous genes and paralogous genes; classifying homologous gene groups with a frequency equal to a first value as unique genes, classifying homologous gene groups with a frequency greater than the first value and less than or equal to a second value as accessory genes, and classifying homologous gene groups with a frequency greater than the second value as core genes; the first value is less than the second value.
[0076] Among them, the protein sequence similarity map can be understood as an image composed of the similarity between various protein sequences as edges and the various protein sequences as nodes. Orthologous genes can be understood as gene pairs generated by the same ancestral gene in different species through speciation events. Paralogous genes can be understood as gene pairs generated by gene duplication events within the same species.
[0077] For example, server 104 obtains the similarity between protein sequences in the protein sequence file. Using each protein sequence as a node, when the similarity between protein sequences is greater than a preset similarity threshold, two protein sequences are connected to form a protein sequence similarity map. Then, the protein sequence similarity map is clustered to obtain multiple homologous gene groups containing orthologous and paralogous genes. The frequency of each homologous gene group in the sequencing gene group is then obtained. Homologous gene groups with a frequency equal to a first value are classified as unique genes; homologous gene groups with a frequency greater than the first value and less than or equal to a second value are classified as unique genes; and homologous gene groups with a frequency greater than the second value are classified as core genes, where the first value is less than the second value. By using frequency as the standard for gene classification, the accuracy of gene clustering analysis is ensured, and the importance of genes in heredity is clearly defined.
[0078] In one embodiment, for any homologous gene group, which is any one of a unique gene, an accessory gene, and a core gene, and which includes multiple protein sequences, the method further includes: annotating each protein sequence to obtain the protein annotation results of the homologous gene group; obtaining the coverage of each protein annotation result; and labeling the homologous gene group based on each coverage and the protein annotation results.
[0079] Protein annotation can be understood as the process of identifying and providing protein-related information at the genomic or transcriptomic level.
[0080] Optionally, for any homologous gene group among the specific gene, accessory gene, and core gene, protein annotation is performed on multiple protein sequences contained in the homologous gene group to obtain the protein annotation results of the homologous gene group. Then, the coverage of each protein annotation result is obtained. If the coverage is less than a preset third value, the homologous gene group is divided into new gene clusters; if the coverage is greater than or equal to the third value and less than or equal to the fourth value, the corresponding protein sequence is labeled with the corresponding protein annotation result; if the coverage is greater than the fourth value, the entire homologous gene group is labeled with the corresponding protein annotation result, where the third value is less than the fourth value. By setting numerical values to perform protein and functional annotation on homologous gene groups, it is ensured that relevant personnel can clearly understand the detailed information of relevant genes through homologous gene groups, facilitating the implementation of corresponding interventions to prevent the host from being infected by pathogens.
[0081] In one embodiment, the method further includes: acquiring metadata for each sequenced genome data; the metadata includes time, location, isolator, host, and molecular typing; and drawing a phylogenetic tree based on the core genes and the metadata of each sequenced genome data to display the genetic developmental structure of each sequenced genome data.
[0082] Here, "source of isolation" can be understood as the initial origin of the sample, typically describing the category of environmental or biological material source, which can include human patients, animal hosts, environmental samples, medical environment samples, etc.; "host" can be understood as the species or individual of the organism from which it originated. While sometimes overlapping with "source of isolation," the emphasis differs: "host" can be at the species level (e.g., human, pig, bird) or a more specific systemic level (e.g., a subgroup of humans, a host in a specific disease state); "molecular typing" can be understood as the molecular typing information of the sample, i.e., typing the sample's genome / specific genomic regions to distinguish different variants or lineages, which can include mitochondrial / nuclear gene sequence typing (e.g., mtDNA haplogroup, MLST sequence type, etc.), pathogen molecular typing profiles (e.g., allele combinations at loci, SNP genotypes, whole-genome variation patterns), and classifications defined in reference databases (e.g., GISAID / Pangolin typing for SARS-CoV-2, MLSTtype for bacteria, etc.).
[0083] For example, server 104 obtains metadata for each sequenced genome data, including time, location, isolator, host, and molecular typing. A phylogenetic tree is constructed based on the core genes, and the metadata for each sequenced genome data is added to the phylogenetic tree. By constructing the phylogenetic tree using the core genes and the metadata of each sequenced genome data, the genetic developmental structure of each sequenced genome data is visualized, supporting intuitive source tracing analysis and accelerating genetic tracking.
[0084] In one exemplary embodiment, such as Figure 3 As shown, a specific implementation of a pan-genome analysis method for pathogenic fungi is provided, wherein the specific analysis steps are as follows (the following data are all specific examples and are not limited to this case to implement the technical solution of this application):
[0085] I. Supports the analysis of Assembly genomic data in FASTA format. The number of input files is ≥3, with no upper limit.
[0086] 2. Removal of human and bacterial host sequences: Blast is compared with human and self-built bacterial reference libraries to remove human and bacterial contig sequences from the genome data.
[0087] 3. Use BUSCO software to perform quality control on the genomic data after removing human and bacterial contigs, with a completeness of ≥90%.
[0088] IV. Gene prediction, which is divided into two cases:
[0089] 1) For genomes with reference sequences, use Augustus software to perform gene prediction;
[0090] 2) For genomes without a reference sequence, use Genemark-ES software for gene prediction.
[0091] 3) After gene prediction is completed, output a gff protein sequence file in FAA format.
[0092] 5. Use the OrthFinder software's BLASTP all vs all function to perform pan-genome cluster analysis on the FAA protein sequence files from step 4. The marker gene type conditions are as follows:
[0093] 1) Core genes are orthologous genes with a frequency of >95% in the genome;
[0094] 2) Dispensable genes are paralogous genes, and their frequency of occurrence in the genome is between 95% and 1%.
[0095] 3) Unique genes are paralogous genes that appear only in one genome dataset.
[0096] VI. Use Emapper to annotate the protein and function of gene clusters after pan-genome analysis.
[0097] 1) Emapper's gene cluster screening parameters: E-value = 1 × 10⁵; Identity > 70%; min query coverage > 50%; min subject coverage > 50%.
[0098] 2) Functional protein result annotation parameters:
[0099] a) If more than 70% of the sequences under a gene cluster can be annotated as the same protein, then all sequences in this gene cluster are considered to be that one protein.
[0100] b) If 50%-70% of the sequences under a gene cluster are annotated as the same protein, then only those 50%-70% of the sequences are considered to be that one protein, and the other sequences are not labeled.
[0101] c) If less than 50% of the sequences under a gene cluster can be annotated to the same protein, or no protein is matched, then it is considered a new gene cluster and needs to be marked as a new gene cluster.
[0102] 7. Draw a phylogenetic tree based on core genes.
[0103] Based on core genes, phylogenetic trees were constructed using the FastTree software. Each phylogenetic tree carries metadata for each genome, including time, location, phenotypic origin, host, and molecular typing. Each column in the metadata table contains one piece of metadata. The genetic and developmental structure of all genomic data can be directly visualized, supporting intuitive etymological analysis.
[0104] Compared with the prior art, this application has the following technical advantages:
[0105] 1. By performing data cleaning on the sequencing genome data to be analyzed, bacterial and human gene data are removed while fungal gene data is retained, enabling targeted analysis of the fungal pan-genome and improving the accuracy of the analysis.
[0106] 2. By performing quality control on the cleaned genomic data, and only proceeding to subsequent gene prediction and cluster analysis if the data passes quality control, the integrity of the data used for prediction and analysis is ensured, thereby guaranteeing the accuracy and reliability of the analysis.
[0107] 3. Using the predicted gff protein sequence file as the dominant representation in cluster analysis improves the traceability of genome cluster analysis and thus avoids analytical errors.
[0108] 4. By drawing a phylogenetic tree based on core genes, the genetic developmental structure of all sequenced genome data can be visualized, supporting intuitive source tracing analysis.
[0109] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0110] Based on the same inventive concept, this application also provides a pathogenic fungal pan-genome analysis device for implementing the aforementioned pathogenic fungal pan-genome analysis method. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of one or more pathogenic fungal pan-genome analysis device embodiments provided below can be found in the limitations of the pathogenic fungal pan-genome analysis method described above, and will not be repeated here.
[0111] In one exemplary embodiment, such as Figure 4 As shown, a pan-genome analysis device for pathogenic fungi is provided, including: an acquisition module 401, a cleaning module 402, a prediction module 403, and an analysis module 404, wherein:
[0112] The acquisition module 401 is used to acquire the sequencing genomic data to be analyzed; the sequencing genomic data includes multiple contig sequences;
[0113] The cleaning module 402 is used to remove human and bacterial contig sequences from the sequencing genome data to obtain cleaned sequencing genome data.
[0114] Prediction module 403 is used to predict genes in the cleaned sequencing genomic data to obtain the gff protein sequence file corresponding to the cleaned sequencing genomic data;
[0115] Analysis module 404 is used to perform pan-genome clustering analysis on the gff protein sequence file to obtain the clustering analysis results of the sequenced genome data.
[0116] In one embodiment, the prediction module 403 is further configured to obtain the integrity level of the cleaned sequencing genomic data; and to perform gene prediction on the cleaned sequencing genomic data if the integrity level is greater than or equal to a preset integrity level threshold.
[0117] In one embodiment, the prediction module 403 is further configured to, when the sequencing genomic data corresponds to a reference sequence, train the prediction model to be trained using the reference sequence, and input the cleaned sequencing genomic data into the trained prediction model for gene prediction; when the sequencing genomic data does not correspond to a reference sequence, perform regularity statistics on the cleaned sequencing genomic data, and perform gene prediction based on the statistically obtained change patterns.
[0118] In an exemplary embodiment, the analysis module 404 is further configured to obtain the similarity between various protein sequences in the gff protein sequence file; construct a protein sequence similarity map based on each protein sequence and the similarity between them; cluster the protein sequence similarity map to obtain multiple homologous gene groups; the homologous gene groups include orthologous genes and paralogous genes; classify homologous gene groups whose occurrence frequency is equal to a first value as unique genes, classify homologous gene groups whose occurrence frequency is greater than the first value and less than or equal to a second value as accessory genes, and classify homologous gene groups whose occurrence frequency is greater than the second value as core genes; the first value is less than the second value.
[0119] In one embodiment, for any homologous gene group, which is any one of the unique gene, accessory gene, and core gene, the homologous gene group includes multiple protein sequences. The pathogenic fungal pan-genome analysis device is also used to annotate each protein sequence to obtain the protein annotation results of the homologous gene group; obtain the coverage of each protein annotation result; and label the homologous gene group according to each coverage.
[0120] In one embodiment, the pathogenic fungal pangenome analysis device is also used to acquire metadata for each sequenced genome data; the metadata includes time, location, isolation source, host, and molecular typing; and based on the core gene and the metadata of each sequenced genome data, a phylogenetic tree is drawn to show the genetic developmental structure of each sequenced genome data.
[0121] Each module in the aforementioned pan-genome analysis device for pathogenic fungi can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0122] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores sequencing genome data to be analyzed, human contig sequences, bacterial contig sequences, washed sequencing genome data, GFF protein sequence files, and clustering analysis results of the sequencing genome data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a pan-genome analysis method for pathogenic fungi.
[0123] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0124] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the pathogenic fungal pangenome analysis method of the above embodiments.
[0125] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the pathogenic fungal pangenome analysis method of the above embodiments.
[0126] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the pan-genome analysis method for pathogenic fungi described above.
[0127] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0128] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0129] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0130] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for pan-genome analysis of pathogenic fungi, characterized in that, The method includes: Acquire sequencing genomic data to be analyzed; the sequencing genomic data includes multiple contig sequences; The human and bacterial contig sequences were removed from the sequenced genome data to obtain the cleaned sequenced genome data; Gene prediction is performed on the cleaned sequencing genomic data to obtain the gff protein sequence file corresponding to the cleaned sequencing genomic data; A pan-genome clustering analysis was performed on the gff protein sequence file to obtain the clustering analysis results of the sequenced genome data.
2. The method according to claim 1, characterized in that, The gene prediction of the cleaned sequencing genomic data includes: To determine the completeness of the cleaned sequencing genomic data; If the integrity level is greater than or equal to a preset integrity level threshold, gene prediction is performed on the cleaned sequencing genomic data.
3. The method according to claim 2, characterized in that, The gene prediction of the cleaned sequencing genomic data includes: When the sequencing genome data corresponds to a reference sequence, the reference sequence is used to train the prediction model to be trained, and the cleaned sequencing genome data is input into the trained prediction model for gene prediction. In cases where no reference sequence is available for the sequenced genome data, statistical analysis is performed on the cleaned sequenced genome data to predict gene patterns based on the statistically obtained patterns.
4. The method according to claim 1, characterized in that, The pan-genome clustering analysis of the gff protein sequence file yields the clustering analysis results of the sequenced genome data, including: Obtain the similarity between the protein sequences in the gff protein sequence file; A protein sequence similarity map is constructed based on each protein sequence and the degree of similarity between them. Clustering of the protein sequence similarity map yields multiple homologous gene groups; the homologous gene groups include orthologous genes and paralogous genes. Homologous gene groups with a frequency equal to a first value are classified as unique genes, homologous gene groups with a frequency greater than the first value and less than or equal to a second value are classified as accessory genes, and homologous gene groups with a frequency greater than the second value are classified as core genes; the first value is less than the second value.
5. The method according to claim 4, characterized in that, For any homologous gene group, wherein the homologous gene group is any one of the specific gene, the accessory gene, and the core gene, and the homologous gene group includes multiple protein sequences, the method further includes: Protein annotation was performed on each of the protein sequences to obtain the protein annotation results of the homologous gene group; Obtain the coverage of each protein annotation result; Based on the coverage level described, the homologous gene groups are labeled using the protein annotation results described.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Obtain metadata for each of the sequenced genome data; the metadata includes time, location, isolation source, host, and molecular typing. Based on the core genes and metadata of each sequenced genome data, a phylogenetic tree is drawn to show the genetic and developmental structure of each sequenced genome data.
7. A pan-genome analysis device for pathogenic fungi, characterized in that, The device includes: An acquisition module is used to acquire the sequencing genomic data to be analyzed; the sequencing genomic data includes multiple contig sequences; The cleaning module is used to remove human and bacterial contig sequences from the sequenced genome data to obtain cleaned sequenced genome data. The prediction module is used to perform gene prediction on the cleaned sequencing genomic data to obtain the gff protein sequence file corresponding to the cleaned sequencing genomic data; The analysis module is used to perform pan-genome clustering analysis on the gff protein sequence file to obtain the clustering analysis results of the sequenced genome data.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.