A method, system, device, and medium for fungal multi-gene joint analysis
By constructing a multi-gene joint analysis method for fungi, acquiring and comparing target sequence data of fungal samples, and constructing a multi-gene phylogenetic tree, the problem of low efficiency and insufficient accuracy in identifying complex fungal species in existing technologies is solved, and efficient and accurate multi-sample species identification is achieved.
Patent Information
- Application Number
- CN202510194871.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Existing single-gene DNA barcoding is ineffective in identifying closely related species and is difficult to accurately distinguish complex fungal species. Existing multi-gene analysis systems are inefficient, inaccurate, and have limited applicability, cannot support the simultaneous identification of multiple samples, and lack reference sequences for type strains.
A multi-gene joint analysis method for fungi was used to obtain target sequence data of fungal samples to be identified, construct a sequence alignment reference library for each target marker gene, perform multi-sample multi-gene pairwise sequence alignment and obtain closely related sequences, construct a multi-gene phylogenetic tree, and achieve species identification.
It significantly improves the accuracy and efficiency of species identification, supports simultaneous analysis of multiple samples, and can quickly and accurately identify complex fungal species. It is suitable for import and export pathogen quarantine, food pathogen monitoring, biosafety, and biological product production.
Smart Images

Figure CN120126559B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bioinformatics, and in particular to a method, system, device and medium for multi-gene joint analysis of fungi. BACKGROUND
[0002] DNA barcoding technology provides reliable technical support in the fields of import and export pathogen quarantine, food pathogen monitoring, biological safety and biological product production due to its advantages of simplicity, rapidness and low cost. However, the traditional single-gene DNA barcoding often has poor effect in identifying closely related species, resulting in inaccurate identification results. Many difficult-to-distinguish species are important pathogenic fungi or ecologically significant fungal resources, such as Aspergillus, Cladosporium and Fusarium. The ITS sequences of fungi in these genera have low variability and cannot be used for accurate species identification. In addition, some fungi are difficult to accurately identify by single-gene sequence analysis due to their complex genetic background.
[0003] In order to improve the accuracy and reliability of species identification, combined analysis of multiple gene sequences becomes the key. However, multi-gene sequence joint analysis involves complex operation steps, which are difficult for non-professionals to master, limiting its wide application. The existing one-stop multi-gene analysis system also has many limitations: first, the system does not support simultaneous identification of multiple samples, and can only analyze one sample at a time, which is low in efficiency; second, when the gene is missing, the system will use a loose mode for species identification, which reduces the accuracy of identification; in addition, the approximate genus range of the sequence to be identified needs to be known before identification, because the sequence reference library is only constructed for specific groups, and cannot support the identification of a wide range of fungal species; finally, the system lacks reference sequences of model strains, which is an important basis for species identification. These limitations make it difficult for existing tools to meet the demand for accurate identification, and it is urgent to develop more efficient and comprehensive multi-gene analysis methods to cope with the identification challenges of complex fungal species. SUMMARY
[0004] The present application provides a multi-gene joint analysis method, system, device and medium for fungi to solve the defects of low efficiency, insufficient accuracy and limited scope of existing multi-gene analysis methods.
[0005] The multi-gene joint analysis method for fungi provided by the present application comprises:
[0006] obtaining target sequence data of a fungal sample to be identified, wherein the target sequence data of the fungal sample to be identified at least includes sequence data of two target marker genes and gene ID information;
[0007] performing pairwise sequence alignment between the target sequence data of the fungal sample to be identified and the original reference sequence to obtain a sequence alignment reference library for each target marker gene;
[0008] According to the target sequence data of the to-be-identified fungal sample, a sequence alignment reference library for each target marker gene is obtained by performing pair sequence alignment on the original reference sequence and the sequence alignment reference library for each target marker gene.
[0009] According to the target sequence data of the to-be-identified fungal sample and the proximal sequence set of all target marker genes of the to-be-identified fungal sample, multi-sequence alignment and sequence trimming are performed respectively, and different gene sequences of the same voucher sample are concatenated to construct a multi-gene phylogenetic tree, thereby achieving species identification of the to-be-identified fungal sample.
[0010] According to the fungal multi-gene joint analysis method provided by the application,
[0011] The pair sequence alignment of the target sequence data of the to-be-identified fungal sample and the original reference sequence comprises the following steps:
[0012] The original reference sequence, the voucher sample thereof and the reference sequence gene ID information are obtained, and the original reference sequence is subjected to quality control processing to obtain a quality-controlled reference sequence, and a total reference sequence library is formed;
[0013] According to the quality-controlled reference sequence, the voucher sample thereof and the reference sequence gene ID information, a voucher sample gene matrix table is obtained;
[0014] According to the voucher sample gene matrix table, the total reference sequence library and the gene ID information of the target marker gene, the quality-controlled reference sequence having all the target marker genes is extracted to form a sequence alignment reference library for each target marker gene.
[0015] According to the fungal multi-gene joint analysis method provided by the application, the quality control processing of the original reference sequence comprises any one or any combination of the following:
[0016] The original reference sequence is subjected to classification correction;
[0017] The sequence format of the original reference sequence is unified.
[0018] According to the fungal multi-gene joint analysis method provided by the application, the total reference sequence library comprises any one or any combination of the following: a pathogenic fungus core library, a model pathogenic fungus core library, a fungus library and a model fungus library.
[0019] According to the fungal multi-gene joint analysis method provided by the application,
[0020] The gene ID information of the target marker gene is extracted according to the voucher sample gene matrix table, the total reference sequence library and the target marker gene, the quality-controlled reference sequence with all target marker genes is formed, and a sequence alignment reference library for each target marker gene is formed.
[0021] The quality-controlled reference sequence with all target marker genes is extracted and subjected to BLAST reference library format conversion to form a sequence alignment reference library for each marker gene.
[0022] According to the fungus multi-gene joint analysis method provided by the application,
[0023] The target sequence data of the to-be-identified fungus sample is used to perform multi-sample multi-gene paired sequence alignment and close sequence acquisition by using the original reference sequence and the sequence alignment reference library for each target marker gene, and the close sequence union of all target marker genes in the to-be-identified fungus sample is obtained.
[0024] The target sequence data of the to-be-identified fungus sample is used to perform sequence alignment for each target marker gene by using the sequence alignment reference library for each target marker gene, and the close reference sequence data of each target marker gene is obtained.
[0025] The close reference sequence data of each target marker gene is used to obtain the close reference sequence of all target marker genes.
[0026] The close reference sequence of all target marker genes is used to obtain the reference sequence of the close reference sequence of the voucher sample of all target marker genes by using the total reference sequence library, and a total close sequence set is formed as the close sequence union of all marker genes in the to-be-identified fungus sample.
[0027] According to the fungus multi-gene joint analysis method provided by the application, the target sequence data of the to-be-identified fungus sample and the close sequence union of all target marker genes thereof are used to perform multi-sequence alignment and sequence pruning, and different gene sequences of the same voucher sample are concatenated to construct a multi-gene phylogenetic tree, so that the species identification of the to-be-identified fungus sample is realized.
[0028] The close sequence union of all marker genes in the to-be-identified fungus sample is subjected to multi-sequence alignment processing to obtain sequence data after multi-sequence alignment processing.
[0029] According to the sequence data after multi-sequence alignment processing, the gene sequences are pruned and concatenated to construct a multi-gene phylogenetic tree, so that the species identification of the to-be-identified fungus sample is realized.
[0030] According to the fungus multi-gene joint analysis method provided by the application,
[0031] The multi-sequence alignment processing of the proximal sequence set of all marker genes in the fungus sample to be identified includes any one or any combination thereof:
[0032] The sequence number format processing of the multi-sequence alignment of the proximal sequence set of all marker genes in the fungus sample to be identified;
[0033] The target sequence and reference sequence merging of the same gene of the proximal sequence set of all marker genes in the fungus sample to be identified;
[0034] The sequence alignment of the proximal sequence set of all marker genes in the fungus sample to be identified.
[0035] According to the fungus multi-gene joint analysis method provided by the application,
[0036] According to the sequence data after the multi-sequence alignment processing, the pruning and the concatenated gene sequence are constructed to obtain a multi-gene phylogenetic tree, and the species identification of the fungus sample to be identified is realized, including:
[0037] The tree sequence format processing is performed on the sequence data after the multi-sequence alignment processing;
[0038] The pruning processing is performed on the sequence data after the multi-sequence alignment processing;
[0039] According to the sequence data after the pruning processing, the joint gene matrix of the concatenated gene is constructed to obtain the concatenated joint gene sequence data;
[0040] According to the concatenated joint gene sequence data, a multi-gene phylogenetic tree is constructed;
[0041] According to the relative position between the known voucher sample sequence on the multi-gene phylogenetic tree, the species identification of the fungus sample to be identified is realized.
[0042] The application also provides a fungus multi-gene joint analysis system, comprising:
[0043] The data acquisition module is used for acquiring target sequence data of the fungus sample to be identified, wherein the target sequence data of the fungus sample to be identified at least includes sequence data and gene ID information of two target marker genes;
[0044] The real-time library construction module is used for performing pairwise sequence alignment on the target sequence data of the fungus sample to be identified and the original reference sequence to obtain a sequence alignment reference library for each target marker gene;
[0045] The multi-gene joint analysis module is used for: according to target sequence data of a to-be-identified fungus sample, using original reference sequences and a sequence alignment reference library for each target marker gene, performing multi-sample multi-gene pair sequence alignment and obtaining close sequences, obtaining a close sequence set of all target marker genes in the to-be-identified fungus sample;
[0046] The tree building module is used for: according to target sequence data of a to-be-identified fungus sample and the close sequence set of all target marker genes thereof, respectively performing multi-sequence alignment and pruning sequences, concatenating different gene sequences of the same voucher sample, and constructing a multi-gene phylogenetic tree to achieve species identification of the to-be-identified fungus sample.
[0047] The present application also provides an electronic device comprising a processor and a memory storing a computer program, wherein the processor implements any of the above-mentioned fungus multi-gene joint analysis methods when executing the computer program.
[0048] The present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement any of the above-mentioned fungus multi-gene joint analysis methods.
[0049] The present application also provides a computer program product comprising a computer program, wherein the computer program can be stored on a non-transitory computer-readable storage medium, and the computer program is executed by a processor to enable a computer to execute any of the above-mentioned fungus multi-gene joint analysis methods.
[0050] The present application provides a fungus multi-gene joint analysis method, system, device and medium, which acquires target sequence data of a to-be-identified fungus sample, uses original reference sequences to construct a precise sequence alignment reference library for each target marker gene in real time, performs accurate multi-sample multi-gene pair sequence alignment and close sequence acquisition, finally constructs a multi-gene phylogenetic tree, and achieves comprehensive species identification, significantly improving the accuracy and efficiency of species identification.
[0051] Specifically, when constructing the sequence alignment reference library of the target marker gene, the reference sequences meeting the multi-gene condition are screened by introducing an auxiliary text sample gene matrix table, which can ensure the accuracy of the identification result; through multi-sample multi-gene pair sequence alignment, the close species sequences of the target sequences are automatically screened, and the sequence alignment results of multi-sample and multi-gene are integrated by combining the real-time constructed sequence alignment reference library, which significantly improves the identification speed and efficiency; the present application supports multi-sample multi-gene joint analysis at the same time, and all sequence alignment results are visualized on the same phylogenetic tree, and each sequence of each gene of each sample has detailed alignment results, which can provide comprehensive analysis data for users.
[0052] The present application supports accurate identification of multiple samples, does not need to know the sample genus range in advance, and has the expandability of other biological groups, can realize rapid and accurate identification, and can be more widely applied in scenes with accurate and rapid identification requirements, such as import and export pathogen quarantine, food pathogen monitoring, biological safety, and biological product production process. BRIEF DESCRIPTION OF DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0054] Figure 1 A flowchart of a fungal multi-gene joint analysis method provided by the present application.
[0055] Figure 2 A flowchart of a fungal multi-gene joint analysis method provided by the present application.
[0056] Figure 3 An analysis interface of a fungal multi-gene joint analysis system provided by the present application.
[0057] Figure 4 One of the fungal multi-gene joint analysis results obtained by the fungal multi-gene joint analysis method provided by the present application is a multi-gene phylogenetic tree.
[0058] Figure 5 One of the details of the voucher sample matched on the multi-gene phylogenetic tree.
[0059] Figure 6 The second of the details of the voucher sample matched on the multi-gene phylogenetic tree.
[0060] Figure 7 The second of the fungal multi-gene joint analysis results obtained by the fungal multi-gene joint analysis method provided by the present application is sequence alignment result one.
[0061] Figure 8 The second of the fungal multi-gene joint analysis results obtained by the fungal multi-gene joint analysis method provided by the present application is sequence alignment result two.
[0062] Figure 9 A structural diagram of a fungal multi-gene joint analysis system provided by the present application.
[0063] Figure 10 A structural diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0064] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application, and should not be understood as a limitation on the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application. In the description of the present application, it should be understood that the terms used are only for the purpose of description and should not be understood as indicating or implying the relative importance.
[0065] Figure 1 and 2 A flowchart of a fungal multi-gene joint analysis method provided by the present application. The execution subject of the fungal multi-gene joint analysis method provided by the present application can be any applicable terminal-side device or network-side device, such as a fungal multi-gene joint analysis device, etc.
[0066] Referring to Figure 1 and 2 The fungal multi-gene joint analysis method provided by the present application can include:
[0067] S110, obtaining target sequence data of a to-be-identified fungal sample, wherein the target sequence data of the to-be-identified fungal sample at least includes sequence data of two target marker genes and gene ID information.
[0068] The number of fungal samples to be identified can be one or more. In this embodiment, the number of fungal samples to be identified is two, named Q1 and Q2, and the marker genes are ITS and TUB2, and each fungal sample to be identified has its own ITS and TUB2 marker genes. Specifically, this embodiment organizes the fasta files of the two target genes of the two fungal samples to be identified, one file for each target gene, and names them according to the standardized gene names of this embodiment (5_8S, ACT, AOX_fmt, ApMat, CAL, CaM, CHS_1, CMDA, CO1, GAPDH, GPD, GS, HIS3, ITS, ITS_LSU, ITS1, ITS2, LSU, mat1_2_1, MS208, MS277, Nad6, RPB1, RPB2, SOD, SSU, TEF1a, TS, TUB). This embodiment takes Q1 (Q1_MFLUCC170318) and Q2 (Q2_MFLUCC170205) and their two marker genes ITS and TUB2 as examples, and writes the sequences of the ITS marker genes of Q1 and Q2 in fasta format into the file ITS.fasta, and the sequences of the TUB2 marker genes of Q1 and Q2 in fasta format into the file TUB2.fasta.
[0069] S120, performing pairwise sequence alignment between the target sequence data of the fungal sample to be identified and the original reference sequence, to obtain a sequence alignment reference library for each target marker gene.
[0070] In one embodiment, S120 can include:
[0071] Obtaining the original reference sequence, its voucher sample, reference sequence gene ID information, and performing quality control processing on the original reference sequence to obtain a quality controlled reference sequence, and forming a total reference sequence library. In this embodiment, the quality control processing on the original reference sequence includes: classification correction of the original reference sequence, and unification of the sequence format of the original reference sequence. The total reference sequence library includes any one or any combination thereof: pathogenic fungus core library, model pathogenic fungus core library, fungus library, model fungus library;
[0072] According to the quality controlled reference sequence, its voucher sample, reference sequence gene ID information, a voucher sample gene matrix table is obtained;
[0073] According to the voucher sample gene matrix table, the total reference sequence library and the gene ID information of the target marker gene, the quality-controlled reference sequences with all target marker genes are extracted to form a sequence alignment reference library for each target marker gene. Specifically, the quality-controlled reference sequences with all target marker genes extracted are converted into BLAST reference library format to form a sequence alignment reference library for each marker gene.
[0074] The acquisition method of the original reference sequence can be: 1. Collecting DNA barcode sequences of related species based on key pathogenic fungi species in the global quarantine species list. For example, obtaining DNA barcode sequence data of related species from biological databases such as GenBank and BOLD. These data cover a large number of DNA barcode sequences and related information from around the world, providing rich resources for constructing sequence libraries. 2. Obtaining self-test DNA barcode data through sampling, DNA extraction, PCR amplification, sequencing and other processes of collected voucher specimens or type specimens. Self-test data will greatly improve the comprehensiveness, authority and accuracy of the data set.
[0075] The quality control processing of the original reference sequence can be quality control and standardization processing of the collected data. For example, the collected data is sorted and low-quality sequences are removed; the species classification of the sequence is corrected by constructing a phylogenetic tree; all voucher sample numbers are standardized to remove special characters, spaces, etc.; the sequence species and its classification information are standardized to current name by matching with the fungi name of the international fungi name registration website Fungal Names, ensuring the accuracy and reliability of the data; the gene name is standardized, and the same gene from different data sources may have multiple writing methods, which needs to be unified to the same name for subsequent analysis.
[0076] From the self-test data, the present embodiment selects 5.8S (The intervening 5.8S nrRNA gene), ACT (Actin gene), AOX-fmt (Alternative oxidase gene), ApMat (Apn2-Mat1-2-1 mating type gene), CAL (Calmodulin gene), CaM (Calmodulin gene), CHS-1 (Chitin synthase gene), CMDA (Calmodulin gene), CO1 (Cytochrome oxidase subunit I gene), GAPDH (Glycerol-3-phosphate dehydrogenase gene), GPD (Glyceraldehyde-3-phosphate dehydrogenase gene), GS (Glutamine synthetase gene), HIS3 (Histone H3 gene), ITS (Internal transcribed spacer regions and intervening 5.8S nrRNA gene), ITS-LSU (Internal transcribed spacers\rRNA large subunit), ITS1 (Internal transcribed spacer regions 1), ITS2 (Internal transcribed spacer regions 2), LSU (Large subunit ribosomal RNA), mat1-2-1 (Mating type gene MAT1-2-1), MS208 (DNA replication-licensing factor required for DNA replication initiation and cell proliferation), MS277 (Single-copy protein coding genes for rRNA accumulation during biogenesis of the ribosome and localized on scaffold 7 in the M.The marker genes, including *Larici-populina* genome, Nad6 (NADH dehydrogenase subunit 6), RPB1 (DNA-directed RNA polymerase I second largest subunit gene), RPB2 (DNA-directed RNA polymerase II second largest subunit gene), SOD (Superoxide dismutase gene), SSU (Small subunit ribosomal RNA), TEF1a (Transcription elongation factor 1a), TS (Tryptophan synthase gene), and TUB (Beta tubulin gene), were amplified by PCR and sequenced; these genes are widely used as major DNA barcodes for plant pathogenic fungi in species identification.
[0077] Regarding the core libraries of pathogenic bacteria, the core libraries of model pathogenic bacteria, fungal libraries, and model fungal libraries, among which...
[0078] 1) Pathogenic fungi core library, including 2500 pathogenic fungi species of 646 genera in 8 phyla (Ascomycota, Basidiomycota, Mucoromycota, Glomeromycota, Monoblepharomycota, Mortierellomycota, Chytridiomycota, Blastocladiomycota), involving 279681 sequences of 29 marker gene fragments, which can realize accurate identification of related species through multi-gene joint analysis; the 29 marker gene fragments are 5.8S (The intervening 5.8S nrRNA gene), ACT (Actin gene), AOX-fmt (Alternative oxidase gene), ApMat (Apn2-Mat1-2-1 mating type gene), CAL (Calmodulin gene), CaM (Calmodulin gene), CHS-1 (Chitin synthase gene), CMDA (Calmodulin gene), CO1 (Cytochrome oxidase subunit I gene), GAPDH (Glycerol-3-phosphate dehydrogenase gene), GPD (Glyceraldehyde-3-phosphate dehydrogenase gene), GS (Glutamine synthetase gene), HIS3 (Histone H3 gene), ITS (Internal transcribed spacer regions and intervening 5.8 S rRNA gene), ITS-LSU (Internal transcribed spacers \ rRNA large subunit), ITS1 (Internal transcribed spacer regions 1), ITS2 (Internal transcribed spacer regions 2), LSU (Large subunit ribosomal RNA), matl-2-1 (Mating type gene MAT1-2-1), MS208 (DNA replication-licensing factor required for DNA replication initiation and cell proliferation), MS277 (Single-copy protein coding genes for rRNA accumulation during biogenesis of the ribosome and localized on scaffold 7 in the M. larici-populina genome), Nad6 (NADH dehydrogenase subunit 6), RPB1 (DNA-directed RNA polymerase I second largest subunit gene), RPB2 (DNA-directed RNA polymerase II second largest subunit gene), SOD (Superoxide dismutase gene), SSU (Small subunit ribosomal RNA), TEF1a (Transcription elongation factor 1a), TS (Tryptophan synthase gene), and TUB (Beta tubulin gene).
[0079] 2) Mode pathogen core library, providing a reference sequence library of model strains, the most important reference standard for species identification, including 22 marker genes of 623 species.
[0080] 3) Fungi library, providing ITS sequence reference library of UNITE database, which can achieve identification of all existing fungal species with ITS sequencing when users choose rDNA ITS sequence for species identification.
[0081] 4) Type Fungi library, providing reference sequences of marker genes of pathogenic fungi, non-pathogenic fungi.
[0082] The diverse high-quality reference library makes up for the limitation of existing reference libraries to a certain genus and cannot achieve accurate species identification of a wider range of species: the present embodiment includes four types of reference libraries, each of which consists of two files: XXX_SampGeneTab.txt and XXX_db.fasta. XXX represents four different data sets: Core_pathogenQT_all_db.fasta, Core_pathogenQT_Type_all_db.fasta, FungiQT_all_db.fasta, and FungiQT_Type_all_SampGeneTab.fasta. The sequence and gene numbers are in the specific format of the present embodiment. Sequence number format: >sample_id|genetaxon_name; gene number format: 5_8S, ACT, AOX_fmt, ApMat, CAL, CaM, CHS_1, CMDA, CO1, GAPDH, GPD, GS, HIS3, ITS, ITS_LSU, ITS1, ITS2, LSU, mat1_2_1, MS208, MS277, Nad6, RPB1, RPB2, SOD, SSU, TEF1a, TS, TUB.
[0083] The total sequence reference library in the present embodiment includes 2500 pathogenic fungal species of 646 genera of 8 phyla of fungal fungi, involves 29 gene fragments, and includes a total of 279681 sequences, providing a rich reference sequence library for accurate species identification. It is helpful to efficiently and accurately analyze multiple samples, multiple gene DNA fragments for accurate identification and tracing at the species level and strain level.
[0084] S130, according to the target sequence data of the fungal sample to be identified, using the original reference sequence and the sequence alignment reference library for each target marker gene, performing multiple sample and multiple gene pairwise sequence alignment and obtaining close sequences, obtaining the union set of close sequences of all target marker genes in the fungal sample to be identified.
[0085] In one embodiment, S130 can include:
[0086] According to the target sequence data of the to-be-identified fungal sample, the sequence of each target marker gene is aligned respectively by using a sequence alignment reference library for each target marker gene through a sequence alignment tool (for example, BLAST software), to obtain the proximal reference sequence data of each target marker gene;
[0087] According to the proximal reference sequence data of each target marker gene, a proximal reference sequence voucher sample ID set of all target marker genes is obtained;
[0088] According to the proximal reference sequence voucher sample ID set of all target marker genes, a reference sequence of the proximal reference sequence voucher sample of all target marker genes is obtained by using a total reference sequence library, to form a total proximal sequence set as a proximal sequence set of all marker genes in the to-be-identified fungal sample.
[0089] S140, according to the target sequence data of the to-be-identified fungal sample and the proximal sequence set of all target marker genes thereof, respectively, a multi-sequence alignment and sequence pruning are performed, and different gene sequences of the same voucher sample are concatenated to construct a multi-gene phylogenetic tree, so as to realize the species identification of the to-be-identified fungal sample.
[0090] In an embodiment, S140 can include:
[0091] The proximal sequence set of all marker genes in the to-be-identified fungal sample is subjected to multi-sequence alignment processing to obtain sequence data after multi-sequence alignment processing. Specifically, the multi-sequence alignment processing of the proximal sequence set of all target marker genes in the to-be-identified fungal sample includes: multi-sequence alignment sequence number format processing of the proximal sequence set of all target marker genes in the to-be-identified fungal sample, target sequence and reference sequence merging of the same gene of the proximal sequence set of all target marker genes in the to-be-identified fungal sample, and sequence alignment of the proximal sequence set of all target marker genes in the to-be-identified fungal sample through a sequence alignment tool (for example, MAFFT software).
[0092] According to the sequence data after multi-sequence alignment processing, the gene sequence is pruned and concatenated to construct a multi-gene phylogenetic tree, so as to realize the species identification of the to-be-identified fungal sample. Specifically, the sequence data after multi-sequence alignment processing is subjected to tree sequence format processing and GAP pruning processing by using a sequence alignment software (trimAI software). According to the sequence data after pruning processing, a joint gene matrix of concatenated genes is constructed to obtain concatenated joint gene sequence data, and then a multi-gene phylogenetic tree is constructed. Then, according to the phase position between each marker gene in the multi-gene phylogenetic tree, the species identification of the to-be-identified fungal sample is realized.
[0093] Referring to Figures 3-8This embodiment uses two fungal samples Q1 and Q2 (with ITS and TUB2 as marker genes) to perform multi-gene joint analysis of fungi (the maximum number of target sequences for multi-gene pairwise sequence alignment in multi-sample multi-gene alignment is set to 20, that is, a maximum of 20 alignment results are output for each gene sequence), and obtains two parts of results: 1) Phylogenetic tree ( Figure 4 (Multigene phylogenetic tree): Construct a multigene phylogenetic tree for the ITS and TUB2 genes by aligning two identification samples (red) with corresponding voucher samples (blue). Clicking on a voucher sample displays detailed information about the voucher. Figure 5 , 6 1) Classification information, specimen status, quarantine country, preservation information, host information, collection information, identification information, etc.; 2) Alignment results: The alignment results of different sequences of the two target genes of each fungal sample to be identified are given. The gene selection is shown in the upper left corner of the visualization results, and the sample selection is shown in the "RESULTS FOR" option at the top center of the visualization results. Figure 7 Clicking on each aligned sequence will take you to the middle section of the visualization results, where you will see the sequence alignment score for the corresponding sample. Figure 7 Clicking on the score result will take you to the bottom of the visualization results to see the base pairing comparison results between that alignment sequence and the sample to be identified. Figure 8 ).
[0094] This invention provides a method, system, device, and medium for combined multi-gene analysis of fungi. By acquiring target sequence data from at least two fungal samples to be identified, a precise sequence alignment reference library for each marker gene is constructed in real time using the original reference sequence. Accurate multi-sample multi-gene paired sequence alignment and acquisition of closely related sequences are performed, and finally, a multi-gene phylogenetic tree is constructed to achieve comprehensive species identification, significantly improving the accuracy and efficiency of species identification.
[0095] Specifically, when constructing the sequence alignment reference library for marker genes, an auxiliary text sample gene matrix table is introduced to screen reference sequences that meet the multi-gene criteria, ensuring the accuracy of the identification results. Through multi-sample, multi-gene paired sequence alignment, closely related sequences of the target sequence are automatically screened, and combined with the real-time constructed sequence alignment reference library, the sequence alignment results of multiple samples and multiple genes are integrated, significantly improving the identification speed and efficiency. This invention supports simultaneous multi-gene joint analysis of multiple samples and visualizes all sequence alignment results on the same phylogenetic tree. At the same time, each sequence of each gene in each sample has detailed alignment results, providing users with comprehensive analytical data.
[0096] The present application provides a multi-gene joint analysis method which can realize multi-sample and accurate identification without prior knowledge of the genus range of the fungal sample, has a high-quality sequence reference library, and an auxiliary text file structure format; the sequence of the reference library is limited according to the content of the auxiliary text sample-gene matrix table to realize real-time construction of a reference sequence set for accurate identification, so as to realize automatic screening of the target gene sequence of the target sample and entering subsequent analysis; the present application is not only suitable for multi-sample and multi-gene joint analysis for accurate identification of fungi, but also suitable for accurate identification of generalized fungi and other groups based on multi-gene joint analysis, which can be realized only by updating the data according to the data file format of the present application; the method of the present application is suitable for but not limited to the following scenarios: a. pathogenic bacteria identification in medical health; b. food pathogenic bacteria monitoring; c. import and export pathogenic bacteria quarantine; d. biological safety assessment; e. production process monitoring of biological products.
[0097] The present application supports multi-sample accurate identification, does not require prior knowledge of the genus range of the sample, and has the expandability of other biological groups, can realize rapid and accurate identification, and can be more widely applied in scenarios with accurate and rapid identification requirements, such as import and export pathogenic bacteria quarantine, food pathogenic bacteria monitoring, biological safety, and production process of biological products.
[0098] The fungal multi-gene joint analysis system provided by the present application is described below, and the fungal multi-gene joint analysis system described below can be referred to each other.
[0099] Referring to Figure 9 The fungal multi-gene joint analysis system provided by the present application can include:
[0100] The data acquisition module is configured to: acquire target sequence data of a fungal sample to be identified, wherein the target sequence data of the fungal sample to be identified at least includes sequence data of two target marker genes and gene ID information;
[0101] The real-time library construction module is configured to: perform pairwise sequence alignment on the target sequence data of the fungal sample to be identified and the original reference sequence to obtain a sequence alignment reference library for each target marker gene;
[0102] The multi-gene joint analysis module is configured to: according to the target sequence data of the fungal sample to be identified, using the original reference sequence and the sequence alignment reference library for each target marker gene, performing multi-sample and multi-gene pairwise sequence alignment and obtaining close sequence, obtaining the union set of close sequences of all target marker genes in the fungal sample to be identified;
[0103] The building module is used for: according to the target sequence data of the to-be-identified fungal sample and the proximal sequence set of all target marker genes thereof, respectively performing multi-sequence alignment and trimming sequence, and concatenating different gene sequences of the same voucher sample, to construct a multi-gene phylogenetic tree, and to realize species identification of the to-be-identified fungal sample.
[0104] Figure 10 An example of an entity structure diagram of an electronic device is shown as Figure 10 The electronic device can include a processor 810, a communications interface 820, a memory 830, and a communications bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communications bus 840. The processor 810 can invoke the logic instructions in the memory 830 to perform the following steps:
[0105] Obtaining target sequence data of a to-be-identified fungal sample, wherein the target sequence data of the to-be-identified fungal sample at least includes sequence data and gene ID information of two target marker genes;
[0106] Performing pairwise sequence alignment between the target sequence data of the to-be-identified fungal sample and the original reference sequence to obtain a sequence alignment reference library for each target marker gene;
[0107] According to the target sequence data of the to-be-identified fungal sample, using the original reference sequence and the sequence alignment reference library for each target marker gene, performing multi-sample multi-gene pairwise sequence alignment and proximal sequence acquisition to obtain a proximal sequence set of all target marker genes in the to-be-identified fungal sample;
[0108] According to the target sequence data of the to-be-identified fungal sample and the proximal sequence set of all target marker genes thereof, respectively performing multi-sequence alignment and trimming sequence, and concatenating different gene sequences of the same voucher sample, to construct a multi-gene phylogenetic tree, and to realize species identification of the to-be-identified fungal sample.
[0109] Further, the logic instructions in the memory 830 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0110] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to perform the following steps:
[0111] obtaining target sequence data of a to-be-identified fungal sample, wherein the target sequence data of the to-be-identified fungal sample at least includes sequence data of two target marker genes and gene ID information;
[0112] performing pairwise sequence alignment on the target sequence data of the to-be-identified fungal sample and the original reference sequence to obtain a sequence alignment reference library for each target marker gene;
[0113] performing multi-sample multi-gene pairwise sequence alignment and obtaining close sequences according to the target sequence data of the to-be-identified fungal sample, the original reference sequence and the sequence alignment reference library for each target marker gene to obtain a union set of close sequences of all target marker genes in the to-be-identified fungal sample;
[0114] performing multi-sequence alignment and pruning sequences according to the target sequence data of the to-be-identified fungal sample and the union set of close sequences of all target marker genes, concatenating different gene sequences of the same voucher sample, and constructing a multi-gene phylogenetic tree to achieve species identification of the to-be-identified fungal sample.
[0115] In another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to perform the following steps:
[0116] obtaining target sequence data of a to-be-identified fungal sample, wherein the target sequence data of the to-be-identified fungal sample at least includes sequence data of two target marker genes and gene ID information;
[0117] Pairing sequence alignment is performed between the target sequence data of the to-be-identified fungal sample and the original reference sequence, to obtain a sequence alignment reference library for each target marker gene;
[0118] According to the target sequence data of the to-be-identified fungal sample, the original reference sequence and the sequence alignment reference library for each target marker gene are used to perform multi-sample multi-gene pairing sequence alignment and acquisition of close sequences, to obtain a union set of close sequences of all target marker genes in the to-be-identified fungal sample;
[0119] According to the target sequence data of the to-be-identified fungal sample and the union set of close sequences of all target marker genes, multi-sequence alignment and sequence pruning are performed respectively, and different gene sequences of the same voucher sample are concatenated to construct a multi-gene phylogenetic tree, so as to realize species identification of the to-be-identified fungal sample.
[0120] The device embodiments described above are only schematic, wherein the units illustrated as separate components can or can not be physically separate, and the components illustrated as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software product, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, server, or network device, etc.) execute the method described in each embodiment or some part of the embodiment.
[0122] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for fungal multi-gene association analysis, characterized by, The method comprises the following steps: obtaining target sequence data of a to-be-identified fungal sample, wherein the target sequence data of the to-be-identified fungal sample at least comprises sequence data of two target marker genes and gene ID information; performing pairwise sequence alignment on the target sequence data of the to-be-identified fungal sample and original reference sequences to obtain a sequence alignment reference library for each target marker gene; performing multi-sample multi-gene pairwise sequence alignment and obtaining proximal sequences based on the target sequence data of the to-be-identified fungal sample, the original reference sequences and the sequence alignment reference library for each target marker gene to obtain a proximal sequence union of all target marker genes in the to-be-identified fungal sample; performing multi-sequence alignment and trimming sequences based on the target sequence data of the to-be-identified fungal sample and the proximal sequence union of all target marker genes, respectively, and concatenating different gene sequences of the same voucher sample to construct a multi-gene phylogenetic tree, thereby achieving species identification of the to-be-identified fungal sample; wherein the step of performing pairwise sequence alignment on the target sequence data of the to-be-identified fungal sample and original reference sequences to obtain a sequence alignment reference library for each target marker gene comprises the following steps: obtaining original reference sequences, voucher samples and reference sequence gene ID information, and performing quality control processing on the original reference sequences to obtain quality-controlled reference sequences and form a total reference sequence library; obtaining a voucher sample gene matrix table based on the quality-controlled reference sequences, voucher samples and reference sequence gene ID information; extracting quality-controlled reference sequences that simultaneously have all target marker genes based on the voucher sample gene matrix table, the total reference sequence library and the gene ID information of the target marker genes to form a sequence alignment reference library for each target marker gene; and the step of performing multi-sample multi-gene pairwise sequence alignment and obtaining proximal sequences based on the target sequence data of the to-be-identified fungal sample, the original reference sequences and the sequence alignment reference library for each target marker gene to obtain a proximal sequence union of all target marker genes in the to-be-identified fungal sample comprises the following steps: performing sequence alignment on each target marker gene based on the target sequence data of the to-be-identified fungal sample and the sequence alignment reference library for each target marker gene to obtain proximal reference sequence data of each target marker gene; obtaining a voucher sample ID union of proximal reference sequences of all target marker genes based on the proximal reference sequence data of each target marker gene; obtaining reference sequences of voucher samples of proximal reference sequences of all target marker genes based on the voucher sample ID union of proximal reference sequences of all target marker genes and the total reference sequence library to form a total proximal sequence set as a proximal sequence union of all marker genes in the to-be-identified fungal sample.
2. The fungal multi-gene association analysis method of claim 1, wherein, the step of extracting quality-controlled reference sequences that simultaneously have all target marker genes based on the voucher sample gene matrix table, the total reference sequence library and the gene ID information of the target marker genes to form a sequence alignment reference library for each target marker gene comprises the following steps: The quality-controlled reference sequence extracted and having all target marker genes is converted into BLAST reference library format to form a sequence alignment reference library for each marker gene.
3. The fungal multi-gene association analysis method of claim 1, wherein, The target sequence data of the to-be-identified fungus sample and the proximal sequences of all target marker genes thereof are collected, multi-sequence alignment is performed on each, and the different gene sequences of the same voucher sample are concatenated to construct a multi-gene phylogenetic tree, thereby achieving species identification of the to-be-identified fungus sample, including: The proximal sequences of all marker genes in the to-be-identified fungus sample are collected and subjected to multi-sequence alignment processing to obtain sequence data after multi-sequence alignment processing; According to the sequence data after multi-sequence alignment processing, the gene sequences are trimmed and concatenated to construct a multi-gene phylogenetic tree, and species identification of the to-be-identified fungus sample is achieved according to the relative positional relationship between the sequences on the phylogenetic tree and the known voucher sample sequences.
4. The fungal multi-gene association method of claim 3, wherein, The proximal sequences of all marker genes in the to-be-identified fungus sample are collected and subjected to multi-sequence alignment processing, including any one or any combination thereof: The proximal sequences of all marker genes in the to-be-identified fungus sample are subjected to multi-sequence alignment processing on the sequence number format; The proximal sequences of all marker genes in the to-be-identified fungus sample are subjected to multi-sequence alignment processing on the sequence number format; The proximal sequences of all marker genes in the to-be-identified fungus sample are subjected to multi-sequence alignment processing on the sequence number format.
5. The fungal multi-gene association analysis method of claim 3, wherein, The proximal sequences of all marker genes in the to-be-identified fungus sample are subjected to multi-sequence alignment processing on the sequence number format. The sequence data after multi-sequence alignment processing is subjected to tree sequence format processing; The sequence data after multi-sequence alignment processing is subjected to tree sequence format processing; According to the sequence data after multi-sequence alignment processing, the gene sequences are trimmed and concatenated to construct a multi-gene phylogenetic tree, and species identification of the to-be-identified fungus sample is achieved according to the relative positional relationship between the sequences on the phylogenetic tree and the known voucher sample sequences. Including: A data acquisition module is configured to acquire target sequence data of a to-be-identified fungus sample, wherein the target sequence data of the to-be-identified fungus sample at least includes sequence data and gene ID information of two target marker genes; 6. A fungal multi-gene association analysis system, characterized by, A real-time library building module is configured to perform pairwise sequence alignment on the target sequence data of the to-be-identified fungus sample and the original reference sequence to obtain a sequence alignment reference library for each target marker gene; A multi-gene joint analysis module is configured to perform multi-sample multi-gene pairwise sequence alignment and proximal sequence acquisition on the target sequence data of the to-be-identified fungus sample using the original reference sequence and the sequence alignment reference library for each target marker gene to obtain a proximal sequence set of all target marker genes in the to-be-identified fungus sample; The building module is used for: according to the target sequence data of the to-be-identified fungal sample and the union of the close sequences of all target marker genes thereof, respectively performing multi-sequence alignment and trimming the sequences, and concatenating different gene sequences of the same voucher sample to construct a multi-gene phylogenetic tree, thereby realizing species identification of the to-be-identified fungal sample. The pair-wise sequence alignment of the target sequence data of the to-be-identified fungal sample and the original reference sequence comprises: obtaining the original reference sequence, the voucher sample thereof and the reference sequence gene ID information, performing quality control processing on the original reference sequence to obtain the quality-controlled reference sequence, and forming a total reference sequence library; obtaining the voucher sample gene matrix table according to the quality-controlled reference sequence, the voucher sample thereof and the reference sequence gene ID information; extracting the quality-controlled reference sequence with all target marker genes from the total reference sequence library according to the voucher sample gene matrix table, the total reference sequence library and the gene ID information of the target marker genes, and forming the sequence alignment reference library for each target marker gene; and the pair-wise sequence alignment of the target sequence data of the to-be-identified fungal sample, the original reference sequence and the sequence alignment reference library for each target marker gene, comprises: performing sequence alignment on each target marker gene according to the target sequence data of the to-be-identified fungal sample and the sequence alignment reference library for each target marker gene, to obtain the close reference sequence data of each target marker gene; obtaining the voucher sample ID set of the close reference sequences of all target marker genes according to the close reference sequence data of each target marker gene; obtaining the reference sequence of the voucher sample of the close reference sequences of all target marker genes from the total reference sequence library according to the voucher sample ID set of the close reference sequences of all target marker genes, to form a total close sequence set as the union of the close sequences of all target marker genes of the to-be-identified fungal sample.
7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to realize the fungal multi-gene joint analysis method according to any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the fungal multi-gene joint analysis method according to any one of claims 1 to 5.
Citation Information
Patent Citations
16S sequencing and ITS sequencing data correlation analysis method and system
CN112466391A
Method for screening core group of stress-resistant endophytic fungus strain of salt-tolerant plant
CN115747367A