A method for identifying Aspergillus flavus at the species level and its application
Through genome-wide analysis and specific target sequence screening, the problem of Aspergillus species-level identification was solved, and the accurate identification of Aspergillus at the species-level was achieved, which was suitable for pathogenic bacteria diagnosis and food security detection.
Patent Information
- Application Number
- CN202410791988.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-06-19
AI Technical Summary
The prior art is difficult to accurately identify Aspergillus aflatoxin at the species level, especially because the intraspecies variation and genomic similarity with the relative species Aspergillus oryzae leads to the ineffective distinction between morphological and molecular identification methods.
The genome-wide analysis method was used to screen out specific target sequences, design primer pairs for PCR amplification, and verify through CRISPR/Cas12a detection or Sanger sequencing to construct a specific candidate target sequence library for Aspergillus aflatoxin for species-level identification of Aspergillus aflatoxin.
It has achieved accurate identification of Aspergillus aflatoxin at the species level, can be applied within a wide range of species, improves the accuracy and specificity of identification results, and is suitable for pathogenic bacteria diagnosis, food security and detection of Chinese medicinal materials pollutants.
Smart Images

Figure CN118600086B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fungal species identification, and more particularly to a method for identifying Aspergillus flavus at the species level and its application. Background Art
[0002] Intraspecific variation refers to phenotypic or genetic differences between individuals within a species. It is fundamental to biological evolution and a crucial component of maintaining biodiversity. However, widespread intraspecific variation poses significant challenges to species classification and identification. This internal variability not only complicates morphological classification and identification but also impacts the accuracy of genetic marker-based identification. Intraspecific variation is a widely recognized phenomenon in fungi.
[0003] Aspergillus flavus is a common source of contamination in crops such as corn, peanuts, and nuts. It has attracted widespread attention due to the potential threat to food safety posed by the strong carcinogen aflatoxin it produces. However, accurate identification of Aspergillus flavus at the species level is very difficult for the following reasons: (1) Neither the classic morphological identification method nor the currently widely used molecular identification method can accurately identify Aspergillus flavus at the species level. (2) There is a wide range of variation within Aspergillus flavus species, which further increases the difficulty of accurately identifying the species at the species level. (3) The key to accurate species identification lies in whether it can be accurately distinguished from closely related species. For Aspergillus flavus, Aspergillus oryzae, which belongs to the subgenus Aspergillus, poses a greater challenge to species identification than distinguishing from other species of the same genus. Aspergillus oryzae and Aspergillus flavus have highly similar genomes, and morphology and existing molecular markers cannot accurately distinguish the two.
[0004] In summary, how to provide an accurate method for identifying Aspergillus flavus at the species level is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a method and application for identifying Aspergillus flavus at the species level.
[0006] Currently, there are as many as 238 genomes of Aspergillus flavus that have been published, and 77 species of the same genus have published multiple genomes. The rich genomic resources lay the foundation for us to carry out research on whole genome analysis methods based on different individuals within the species using Aspergillus flavus as a model species.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A method for identifying Aspergillus flavus at the species level comprises the following steps:
[0009] (1) Genome acquisition
[0010] The genomes of Aspergillus flavus, Aspergillus fumigatus, Aspergillus oryzae, and Aspergillus parasiticus were downloaded from the NCBI database;
[0011] (2) Screening of the whole genome of Aspergillus flavus
[0012] The genomes of the A. flavus strains were constructed based on the maximum likelihood method, and Aspergillus fumigatus, Aspergillus oryzae, and Aspergillus parasiticus were selected as external groups to filter out potentially erroneous or unreliable A. flavus genomes.
[0013] (3) Whole genome analysis of Aspergillus flavus
[0014] The final screened Aspergillus flavus genome was decomposed into 25 bp fragments, and all fragments containing PAM sequences were extracted, i.e., candidate target sequences;
[0015] (4) Screening of specific candidate target sequences within the genus Aspergillus
[0016] The candidate target sequence was compared with the whole genome of 76 other species of Aspergillus, and sequences with at least three nucleotide differences from other sequences were selected as the genus-specific candidate target sequences;
[0017] (5) Construction of aflatoxin-specific candidate target sequence library
[0018] Comparing and analyzing the specific candidate target sequences within the genus with the genomes of 10 standard strains of Aspergillus oryzae, and selecting sequences that differ by at least three nucleotides from homologous sequences in Aspergillus oryzae as specific candidate target sequences;
[0019] (6) Selecting sequences with high occurrence frequency and good specificity in NCBI BLAST results from the specific candidate target sequences as specific target sequences, and the selected specific target sequences can be used to accurately distinguish Aspergillus flavus from other species in the biological world;
[0020] (7) Designing a primer pair based on the specific target sequence, extracting genomic DNA of the sample to be tested and amplifying it using the primer pair, and detecting whether the amplified product contains the specific target sequence. If so, the sample to be tested is Aspergillus flavus.
[0021] Furthermore, in step (1), the genomes of all different strains of Aspergillus flavus from NCBI are downloaded.
[0022] Furthermore, in step (3), the PAM sequence is a sequence with TTTV at the 5' end, wherein V represents G, C or A.
[0023] Furthermore, the specific target sequence is as follows:
[0024] TTTCTCTAGGTTAAACGGTGACTCT, SEQ ID NO.1;
[0025] TTTAGTGGTAGAGGTCCATAGACTA, SEQ ID NO.2;
[0026] TTTCGAAGCTTAGTGCTGATTTGGT, SEQ ID NO.3;
[0027] TTTCCGAGTACTTTCCAAAGCCGCT, SEQ ID NO.4;
[0028] TTTGGAAAGTACTCGGAAAATCTGG, SEQ ID NO.5;
[0029] TTTAGCGAACTAAGACGAGCTTTTC, SEQ ID NO.6;
[0030] TTTACTCGTACTTACACCTCGCTGT, SEQ ID NO.7;
[0031] TTTGCAGCGGCTTTGGAAAGTACTC, SEQ ID NO.8;
[0032] TTTGGGACAGACCTAAAGATAGCGT, SEQ ID NO.9;
[0033] TTTGATGTACTGTAATTCGGTGAGT, SEQ ID NO. 10.
[0034] It should be noted that the specific target sequences obtained according to the method of the present invention have the same identification capabilities. That is, other specific target sequences obtained using the same method also fall within the scope of protection of the present invention and are not limited to the above 10 sequences.
[0035] Furthermore, the nucleotide sequences of the primer pairs are shown in SEQ ID NO.11 to SEQ ID NO.26; wherein,
[0036] SEQ ID NO.11 to SEQ ID NO.12 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.1;
[0037] SEQ ID NO.13 to SEQ ID NO.14 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.2;
[0038] SEQ ID NO.15 to SEQ ID NO.16 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.3;
[0039] SEQ ID NO.17 to SEQ ID NO.18 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.4;
[0040] SEQ ID NO.17 to SEQ ID NO.18 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.5;
[0041] SEQ ID NO.19 to SEQ ID NO.20 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.6;
[0042] SEQ ID NO.21 to SEQ ID NO.22 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.7;
[0043] SEQ ID NO.17 to SEQ ID NO.18 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.8;
[0044] SEQ ID NO.23 to SEQ ID NO.24 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.9;
[0045] SEQ ID NO. 25 to SEQ ID NO. 26 are used to amplify a specific target sequence having a nucleotide sequence of SEQ ID NO. 10.
[0046] Furthermore, the detection method in step (7) includes a CRISPR / Cas12a detection method and a Sanger sequencing detection method.
[0047] It should be noted that any method that can identify whether a specific target sequence exists in a sample to be tested can be used for the final detection and falls within the scope of protection of the present invention, and is not limited to the above two methods.
[0048] A kit for identifying Aspergillus flavus, comprising the above primer pair.
[0049] Application of the above kit in species-level identification of Aspergillus flavus.
[0050] The above-mentioned kit is used in the diagnosis of aflatoxin pathogens, food safety and the detection of pollutants in traditional Chinese medicines.
[0051] It can be seen from the above technical solutions that, compared with the prior art, the present invention has the following beneficial effects:
[0052] (1) The present method constructs a library of over four million candidate target sequences for Aspergillus flavus. The selected specific target sequences are highly conserved within species and have interspecies specificity, enabling accurate species-level identification of a wide range of species and closely related species of the genus Aspergillus. This method fully considers the impact of intraspecific variation and optimizes bioinformatics analysis strategies and processes, enabling analysis results to be applied across a wide range of species, not just within the analysis scope, thus offering broader application prospects.
[0053] (2) Sequencing, as the gold standard for sequence identification, not only proves the conservation of each specific target sequence within the species of A. flavus, but also the differences with the homologous sequences of the closely related species A. oryzae also prove the specificity of the sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0055] Figure 1 This is a phylogenetic tree of the genomes of 238 different Aspergillus flavus strains based on the maximum likelihood method in Example 1 of the present invention;
[0056] Figure 2 is the number of 25mers (target sequences) containing PAM sequences at different frequencies in Example 1 of the present invention, wherein A is the number of target sequences in different frequency intervals; B is the proportion of target sequences in different frequency intervals in the total target sequences;
[0057] Figure 3 is the number of specific candidate target sequences at different frequencies in Example 1 of the present invention;
[0058] Figure 4 The BLAST result of the highest-frequency specific candidate target sequence in Example 1 of the present invention;
[0059] Figure 5 This is the sequencing verification result in Example 2 of the present invention;
[0060] Figure 6 This is the identification result in Example 3 of the present invention. DETAILED DESCRIPTION
[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0062] The reagents required for the present invention are conventional experimental reagents, purchased from commercial channels; the experimental methods not mentioned are conventional experimental methods and will not be described in detail here.
[0063] Example 1
[0064] Screening of Aspergillus flavus-specific target sequences
[0065] 1 Genome acquisition
[0066] The publicly available genomes of 238 different strains of A. flavus were downloaded from the NCBI database, along with their corresponding strains and submitting institutions. In addition, the genomes of Aspergillus fumigatus, Aspergillus oryzae, and Aspergillus parasiticus were downloaded for A. flavus genome correction.
[0067] 2. Screening of the whole genome of Aspergillus flavus
[0068] Based on the maximum likelihood method, JolyTree (v.1.1b.191021ac) was used to construct the genomes of the above 238 A. flavus strains. Aspergillus fumigatus, Aspergillus parasiticus, and Aspergillus oryzae were selected as external groups to filter out potential errors or unreliable A. flavus genomes. Figure 1 shown.
[0069] Depend on Figure 1 As can be seen, most A. flavus genomes clustered together, while a few showed strong similarity to genomes from other species. To ensure that the results were applicable to the identification of the majority of A. flavus individuals and to avoid interference from individual erroneous genomes, we filtered out genomes that clustered with other species before analyzing the A. flavus genomes. Ultimately, 217 individual A. flavus genomes were included in the analysis, representing 91.2% of all A. flavus genomes. Therefore, our results are representative of the overall A. flavus species-level situation.
[0070] 3 Whole genome analysis of Aspergillus flavus
[0071] The final screened A. flavus genomes were broken down into 25bp fragments using Jellyfish (v1.1.12), and the frequency of each 25bp fragment in different genomes was calculated. Subsequently, all 25-mers containing PAM sequences (starting with TTTV, where V represents G, C, or A) were extracted and mapped to their genomes using Bowtie (v1.1.0) with default parameters to determine their location in the genome. The details are as follows:
[0072] By cutting the genomes of 217 different strains of Aspergillus flavus, a total of 25mers (deduplicated) containing 4,141,014 PAM sequences were obtained, namely candidate target sequences (Target). Among them, there are 955,272 sequences with a repetition frequency of more than 50% in different genomes, accounting for 23.09% of the total sequences. Through the statistics of the number of candidate target sequences with different frequencies, it can be seen that sequences with a frequency below 0.1 account for the largest proportion, which shows that there is rich intraspecific variation in Aspergillus flavus, and that our study covers a sufficient number of different individuals within the Aspergillus flavus species. At the same time, there are also more sequences with a frequency of more than 0.9, which is consistent with the actual situation, that is, different individuals of the same species have some of the same characteristics and will be classified as the same species. This also shows that the current bioinformatics analysis strategy can successfully screen relatively conserved candidate target sequences within the Aspergillus flavus species for subsequent specific target sequence screening, such as Figure 2 shown.
[0073] 4 Screening of specific candidate target sequences within Aspergillus
[0074] Although Aspergillus flavus exhibits extensive intraspecific variation, providing species-specific target sequences is more suitable for accurate and rapid identification of Aspergillus flavus in current applications such as pathogen diagnosis, food safety, and detection of contaminants in traditional Chinese medicines. Therefore, we aimed to identify sequences that are highly conserved within species and highly specific across species for identification of Aspergillus flavus species.
[0075] For each species, all 25-mer sequences were aligned with the whole genomes of 76 other Aspergillus species, and sequences with at least three nucleotide differences from other sequences were selected to construct a library of species-specific candidate target sequences. The details are as follows:
[0076] In the specificity analysis, we first retrieved the genomic data of other species of Aspergillus based on the results of previous research, and successfully screened 544,604 specific sequences that differed by more than 3 bases from other species of the same genus through genome comparison. These sequences are called specific candidate target sequences within the genus. It should be noted that in our analysis strategy, when the difference is at the PAM position, even if there is only one, it will be considered specific. This is because, in principle, PAM is a prerequisite for Cas12a to function, so these sequences are also considered specific candidate target sequences within the genus.
[0077] 5 Construction of aflatoxin-specific candidate target sequence library
[0078] For Aspergillus flavus, compared with distinguishing it from other species of the same genus, Aspergillus oryzae, which belongs to the same subgenus of Aspergillus, poses a greater challenge to the identification of Aspergillus flavus at the species level. The genomes of Aspergillus oryzae and Aspergillus flavus are highly similar, which also explains why morphological identification and existing molecular markers cannot accurately distinguish the two.
[0079] The key to accurate species identification lies in whether it can be accurately distinguished from closely related species. Therefore, the obtained genus-specific candidate target sequences were further compared with the genomes of 10 standard strains of Aspergillus oryzae, and sequences with at least three nucleotide differences from the homologous sequences in Aspergillus oryzae were selected to form a library of Aspergillus flavus-specific candidate target sequences. Finally, 1,388 specific candidate target sequences were screened, of which 28 had an occurrence frequency of more than 90%, 132 were between 80% and 90%, and 1,228 were between 50% and 80%. Figure 3 shown.
[0080] The most frequent specific target sequence was (TTTCTCTAGGTTAAACGGTGACTCT, SEQ ID NO. 1), which was present in 211 individuals (97.2%). The NCBI BLAST results showed that the sequence was located on chromosome 4 and was completely consistent with the genomes of multiple Aspergillus flavus individuals, while the homologous sequences in other species were different from it, such as Figure 4 Further investigation revealed that this sequence, located in the CDS region, has the potential to encode a protein of unknown function. This result confirms that the strategy based on whole-genome analysis can identify highly conserved sequences within species and highly specific sequences between species for species-level identification of A. flavus.
[0081] 1,388 specific candidate target sequences were compared with the whole genomes of a wide range of taxa, and it was found that these sequences still maintained high specificity. These results show that when the method of the present invention is used to analyze the genomes of target species from multiple different individuals, specific target sequences that are highly conserved within the species can be preferentially screened out, thereby eliminating the need for subsequent confirmatory analysis of these sequences in more different samples. This optimized bioinformatics strategy improves the accuracy of species-level identification of specific target sequences, and is not limited to species identification within the scope of the analysis. By using NCBI's BLAST tool for analysis, it was found that these sequences are highly conserved within the species and show high specificity between species.
[0082] 6 Determination of aflatoxin-specific target sequences
[0083] The top 10 sequences with the highest frequency and good specificity in NCBI BLAST results were selected from the specific candidate target sequence library as specific target sequences, and specific primers were designed, as shown in the following table.
[0084] Table 1 Specific target sequences and primer information
[0085]
[0086] Example 2
[0087] Verification of Aspergillus flavus-specific target sequences by Sanger sequencing
[0088] Extract DNA from the sample to be tested, perform PCR amplification, and sequence the sample using the primer sets listed in Table 1. If the sequencing result is identical to the A. flavus-specific target sequence, the sample to be tested is A. flavus; if the sequencing result is different, the sample to be tested is not A. flavus.
[0089] The information of the samples to be tested is shown in the following table, where numbers 1 to 10 are Aspergillus flavus and numbers 11 to 13 are Aspergillus oryzae.
[0090] Table 2 Information of samples to be tested
[0091]
[0092] The sequencing results are shown in Table 3. Figure 5 shown.
[0093] Table 3 Sequencing statistical results
[0094]
[0095]
[0096] The results showed that the specific primers successfully amplified the specific target sequences. Sequencing analysis of the purified and recovered amplified products revealed that these sequences were highly conserved within the species, with each sequence sharing 100% identity across 10 different A. flavus samples. Furthermore, no exact matches were found in A. oryzae. Sequencing results showed that all sequences were present in all 10 A. flavus samples, confirming the intraspecific conservation of the specific candidate target sequences. Further BLAST analysis of the NCBI nucleic acid database revealed that each sequence was unique to different individuals of A. flavus and absent from other species. This demonstrates that by analyzing the genomes of a specific species and its more closely related species, we can identify target sequences that are both highly conserved within the species and highly specific across species. These target sequences directly impact the ability to identify A. flavus at the species level, ultimately impacting the accuracy and potential of genomic analysis methods. Therefore, we used A. oryzae, a closely related species within the subgenus Aspergillus, to validate the specificity of the target sequences and conducted a detailed analysis of their sequence differences through sequencing. It should be noted that in our bioinformatics analysis, when there is a difference between any PAM site of the candidate target sequence and its homologous sequence, the sequence will be directly identified as a specific sequence. Therefore, the sequencing results show that there are some homologous sequences in the specific target sequences screened in the Aspergillus oryzae sample, and their differences are limited to the PAM site. In short, sequencing, as the gold standard for sequence identification, not only proves the conservation of each specific target sequence within the Aspergillus species, but also proves the specificity of the sequence based on the differences with the homologous sequences of the closely related species Aspergillus oryzae. In addition, sequencing also provides mismatch information of the homologous sequence at this position, thereby achieving accurate identification of Aspergillus species based on the specific target sequence.
[0097] Example 3
[0098] PCR amplification of Aspergillus flavus and Aspergillus oryzae samples was performed using ITS universal primers (ITS1F: TCCGTAGGTGAACCTGCGG, SEQ ID NO. 27; ITS 4R: TCCTCCGCTTATTGATATGC, SEQ ID NO. 28). A 25 μL reaction system was used, including 12.5 μL of 2× Taq Master Mix, upstream and downstream primers (2.5 μmol·L -1 ) 1 μL each, genomic DNA 2 μL (50 ng·μL -1 ), and finally add 8.5 μL of sterile double-distilled water.
[0099] The PCR amplification program was 94°C, 5 min; 94°C, 60 s, 55°C, 30 s, 72°C, 45 s, 35 cycles; 72°C, 7 min.
[0100] 1.5% agarose gel was prepared to observe the PCR amplification results. The electrophoresis conditions were 100V for 30min and compared with MarkerDL 2000 to determine the presence and sequence length range of the amplified target band. The successfully amplified PCR product was purified and bidirectionally sequenced using ABI3730XL sequencer.
[0101] The results are as follows Figure 6 shown.
[0102] Depend on Figure 6 It can be seen that traditional molecular identification methods cannot distinguish and identify Aspergillus oryzae and Aspergillus flavus.
[0103] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0104] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A primer pair for amplifying a specific target nucleotide for species-level identification of Aspergillus flavus, characterized in that: The nucleotide sequences of the primer pairs are shown in SEQ ID NO.11 to SEQ ID NO.26; wherein, SEQ ID NO.11 to SEQ ID NO.12 are used to amplify the specific target sequence with the nucleotide sequence of SEQ ID NO.1; SEQ ID NO.13 to SEQ ID NO.14 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.2; SEQ ID NO.15 to SEQ ID NO.16 are used to amplify the specific target sequence with the nucleotide sequence of SEQ ID NO.3; SEQ ID NO.17 to SEQ ID NO.18 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.4; SEQ ID NO.17 to SEQ ID NO.18 are used to amplify the specific target sequence with the nucleotide sequence of SEQ ID NO.5; SEQ ID NO.19 to SEQ ID NO.20 are used to amplify the specific target sequence with the nucleotide sequence of SEQ ID NO.6; SEQ ID NO.21 to SEQ ID NO.22 are used to amplify the specific target sequence with the nucleotide sequence of SEQ ID NO.7; SEQ ID NO.17 to SEQ ID NO.18 are used to amplify the specific target sequence with the nucleotide sequence of SEQ ID NO.8; SEQ ID NO. 23 to SEQ ID NO. 24 are used to amplify the specific target sequence with the nucleotide sequence of SEQ ID NO. 9; SEQ ID NO.25 to SEQ ID NO.26 are used to amplify the specific target sequence whose nucleotide sequence is SEQ ID NO.
10.
2. A kit for identifying Aspergillus flavus, characterized in that: Comprising the primer pair according to claim 1.
3. Use of the kit according to claim 2 in species-level identification of Aspergillus flavus.
4. Use of the kit according to claim 2 in the detection of pollutants in grains and Chinese medicinal materials.
Citation Information
Patent Citations
Fungus species identification method based on time-treasure method, target nucleotide, primer pair, kit and application thereof
CN116287416A